VLDB 2026 Research / reviewers in the wild / expert
Anamaria Crisan
dblp:62/7969
· DBLP profile ↗
17ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0003-3445-3414ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-authorArtificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DSCode Comparator: An Interactive Interface for Comparing Models and Evaluating Code for Data Science TasksabstractCode-generating models are increasingly used to support data science tasks. Yet reviewing their outputs, both to understand how the code works and to assess its quality, remains largely manual and time-consuming. Instead of eliminating effort, these models shift the burden from writing code to verifying it. Complicating matters further, different models often produce divergent solutions of varying efficacy, creating additional challenges for code interrogation. To address this unmet need, we introduce DSCode Comparator, an interactive interface designed to support code understanding, evaluation, refinement, and comparison in data science workflows. DSCode Comparator allows code to be viewed from different levels of granularity, from individual lines of code to comparisons across prompts and tasks. The individual code views automatically annotate lines of code via an agentic pipeline we developed to facilitate quick functional overviews. The individual views also automatic diagnosis code quality according to efficiency, readability, and resource computation. The comparison and historical views leverage the annotations to create compact visual summaries of code, allowing for direct comparisons of its functionality, length, and efficiency across multiple models, data science tasks, and prompts. To evaluate DSCode Comparator, we conducted a user study with 22 participants, all with varying levels of proficiency in writing Data science code. Our findings show that, especially for non-expert users, DSCode Comparator sped up the pace of code comprehension and increased participant confidence. The majority report finding DSCode Comparator easier to use and more efficient than manual efforts in reviewing and refining code. Overall, our systems and their findings present an intelligent, human-centered approach to address the verification gap when using code generation models for Data Science. Victor Zhong, Anamaria Crisan |
IUI | 3 |
| 2026 | A Scoping Review of Mixed Initiative Visual Analytics in the Automation RenaissanceabstractAbstract Artificial agents are increasingly integrated into data analysis workflows, carrying out tasks that were primarily done by humans. Our research explores how the introduction of automation recalibrates the dynamic between humans and automating technology. To explore this question, we conducted a scoping review encompassing twenty years of mixed‐initiative visual analytic systems. To describe and contrast the relationship between humans and automation, we developed an integrated taxonomy to delineate the objectives of these mixed‐initiative visual analytics tools, how much automation they support, and the assumed roles of humans. Here, we describe our qualitative approach of integrating existing theoretical frameworks with new codes we developed. Our analysis shows that the visualization research literature lacks consensus on the definition of mixed‐initiative systems and explores a limited potential of the collaborative interaction landscape between people and automation. Our research provides a scaffold to advance the discussion of human‐AI collaboration during visual data analysis. Our integrated taxonomy is available in the form of a web application on https://smonadjemi.github.io/miva . Shayan Monadjemi, Yugan Guo, Kai Xu 0003, Alex Endert, Anamaria Crisan |
Comput. Graph. Forum | 5 |
| 2026 | Probing the Visualization Literacy of Vision Language Models: The Good, the Bad, and the UglyabstractVision Language Models (VLMs) demonstrate promising chart comprehension capabilities. Yet, prior explorations of their visualization literacy have been limited to assessing their response correctness and fail to explore their internal reasoning. To address this gap, we adapted attention-guided class activation maps (AG-CAM) for VLMs, to visualize the influence and importance of input features (image and text) on model responses. Using this approach, we conducted an examination of four open-source (ChartGemma, Janus 1B and 7B, and LLaVA) and two closed-source (GPT-4o, Gemini) models comparing their performance and, for the open-source models, their AG-CAM results. Overall, we found that ChartGemma, a 3B parameter VLM fine-tuned for chart question-answering (QA), outperformed other open-source models and exhibited performance on par with significantly larger closed-source VLMs. We also found that VLMs exhibit spatial reasoning by accurately localizing key chart features, and semantic reasoning by associating visual elements with corresponding data values and query tokens. Our approach is the first to demonstrate the use of AG-CAM on early fusion VLM architectures, which are widely used, and for chart QA. We also show preliminary evidence that these results can align with human reasoning. Our promising open-source VLMs results pave the way for transparent and reproducible research in AI visualization literacy. Code and Supplemental Materials: https://osf.io/fp3rg. Lianghan Dong, Anamaria Crisan |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Data Has Entered the Chat: How Data Workers Conduct Exploratory Visual Analytic Conversations with GenAI AgentsabstractWe investigate the potential of leveraging the code-generating capabilities of Large Language Models (LLMs) to support exploratory visual analysis (EVA) via conversational user interfaces (CUIs). We developed a technology probe that was deployed through two studies with a total of 50 data workers to explore the structure and flow of visual analytic conversations during EVA. We analyzed conversations from both studies using thematic analysis and derived a state transition diagram summarizing the conversational flow between four states of participant utterances ( Analytic Tasks , Editing Operations , Elaborations and Enrichments , and Directive Commands ) and two states of Generative AI (GenAI) agent responses (visualization, text). We describe the capabilities and limitations of GenAI agents according to each state and transitions between states as three co-occurring loops: analysis elaboration, refinement, and explanation. We discuss our findings as future research trajectories to improve the experiences of data workers using GenAI. The code and data are available at https://osf.io/6wxpa . Matt-Heun Hong, Anamaria Crisan |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2025 | From Dashboard Zoo to Census: A Case Study With Tableau PublicabstractDashboards remain ubiquitous tools for analyzing data and disseminating the findings. Understanding the range of dashboard designs, from simple to complex, can support development of authoring tools that enable end-users to meet their analysis and communication goals. Yet, there has been little work that provides a quantifiable, systematic, and descriptive overview of dashboard design patterns. Instead, existing approaches only consider a handful of designs, which limits the breadth of patterns that can be surfaced. More quantifiable approaches, inspired by machine learning (ML), are presently limited to single visualizations or capture narrow features of dashboard designs. To address this gap, we present an approach for modeling the content and composition of dashboards using a graph representation. The graph decomposes dashboard designs into nodes featuring content "blocks'; and uses edges to model "relationships", such as layout proximity and interaction, between nodes. To demonstrate the utility of this approach, and its extension over prior work, we apply this representation to derive a census of 25,620 dashboards from Tableau Public, providing a descriptive overview of the core building blocks of dashboards in the wild and summarizing prevalent dashboard design patterns. We discuss concrete applications of both a graph representation for dashboard designs and the resulting census to guide the development of dashboard authoring tools, making dashboards accessible, and for leveraging AI/ML techniques. Our findings underscore the importance of meeting users where they are by broadly cataloging dashboard designs, both common and exotic. Arjun Srinivasan, Joanna Purich, Michael Correll, Leilani Battle, Vidya Setlur, Anamaria Crisan |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | Groot: A System for Editing and Configuring Automated Data InsightsabstractVisualization tools now commonly present automated insights highlighting salient data patterns, including correlations, distributions, outliers, and differences, among others. While these insights are valuable for data exploration and chart interpretation, users currently only have a binary choice of accepting or rejecting them, lacking the flexibility to refine the system logic or customize the insight generation process. To address this limitation, we present Groot, a prototype system that allows users to proactively specify and refine automated data insights. The system allows users to directly manipulate chart elements to receive insight recommendations based on their selections. Additionally, Groot provides users with a manual editing interface to customize, reconfigure, or add new insights to individual charts and propagate them to future explorations. We describe a usage scenario to illustrate how these features collectively support insight editing and configuration and discuss opportunities for future work, including incorporating Large Language Models (LLMs), improving semantic data and visualization search, and supporting insight management. Sneha Gathani, Anamaria Crisan, Vidya Setlur, Arjun Srinivasan |
IEEE VIS | 2 |
| 2024 | Eliciting Model Steering Interactions From Users via Data and Visual Design ProbesabstractVisual and interactive machine learning systems (IML) are becoming ubiquitous as they empower individuals with varied machine learning expertise to analyze data. However, it remains complex to align interactions with visual marks to a user's intent for steering machine learning models. We explore using data and visual design probes to elicit users' desired interactions to steer ML models via visual encodings within IML interfaces. We conducted an elicitation study with 20 data analysts with varying expertise in ML. We summarize our findings as pairs of target-interaction, which we compare to prior systems to assess the utility of the probes. We additionally surfaced insights about factors influencing how and why participants chose to interact with visual encodings, including refraining from interacting. Finally, we reflect on the value of gathering such formative empirical evidence via data and visual design probes ahead of developing IML prototypes. Anamaria Crisan, Maddie Shang, Eric Brochu |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | Tracing and Visualizing Human-ML/AI Collaborative Processes through Artifacts of Data WorkabstractAutomated Machine Learning (AutoML) technology can lower barriers in data work yet still requires human intervention to be functional. However, the complex and collaborative process resulting from humans and machines trading off work makes it difficult to trace what was done, by whom (or what), and when. In this research, we construct a taxonomy of data work artifacts that captures AutoML and human processes. We present a rigorous methodology for its creation and discuss its transferability to the visual design process. We operationalize the taxonomy through the development of AutoML Trace a visual interactive sketch showing both the context and temporality of human-ML/AI collaboration in data work. Finally, we demonstrate the utility of our approach via a usage scenario with an enterprise software development team. Collectively, our research process and findings explore challenges and fruitful avenues for developing data visualization tools that interrogate the sociotechnical relationships in automated data work. Jennifer Rogers, Anamaria Crisan |
CHI | 2 |
| 2022 | Visualization in Data Science VDS @ KDD 2022abstractData science is the practice of deriving insight from data, enabled by modeling, computational methods, interactive visual analysis, and domain-driven problem solving. Data science draws from methodology developed in such fields as applied mathematics, statistics, machine learning, data mining, data management, visualization, and HCI. It drives discoveries in business, economy, biology, medicine, environmental science, the physical sciences, the humanities and social sciences, and beyond. Machine learning and data mining and visualization are integral parts of data science, and essential to enable sophisticated analysis of data. Nevertheless, both research areas are currently still rather separated and investigated by different communities rather independently. The goal of this workshop is to bring researchers from both communities together in order to discuss common interests, to talk about practical issues in application-related projects, and to identify open research problems. This summary gives a brief overview of the ACM KDD Workshop on Visualization in Data Science (VDS at ACM KDD and IEEE VIS), which will take place virtually on Aug 14-18, 2022 (Held in conjunction with KDD'22). The workshop website is available at http://www.visualdatascience.org/2022/ Claudia Plant, Nina C. Hubig, Junming Shao, Alvitta Ottley, Liang Gou, Torsten Möller, Adam Perer, Alexander Lex, Anamaria Crisan |
KDD | 9 |
| 2022 | GEViTRec: Data Reconnaissance Through Recommendation Using a Domain-Specific Visualization Prevalence Design SpaceabstractGenomic Epidemiology (genEpi) is a branch of public health that uses many different data types including tabular, network, genomic, and geographic, to identify and contain outbreaks of deadly diseases. Due to the volume and variety of data, it is challenging for genEpi domain experts to conduct data reconnaissance; that is, have an overview of the data they have and make assessments toward its quality, completeness, and suitability. We present an algorithm for data reconnaissance through automatic visualization recommendation, GEViTRec. Our approach handles a broad variety of dataset types and automatically generates visually coherent combinations of charts, in contrast to existing systems that primarily focus on singleton visual encodings of tabular datasets. We automatically detect linkages across multiple input datasets by analyzing non-numeric attribute fields, creating a data source graph within which we analyze and rank paths. For each high-ranking path, we specify chart combinations with positional and color alignments between shared fields, using a gradual binding approach to transform initial partial specifications of singleton charts to complete specifications that are aligned and oriented consistently. A novel aspect of our approach is its combination of domain-agnostic elements with domain-specific information that is captured through a domain-specific visualization prevalence design space. Our implementation is applied to both synthetic data and real Ebola outbreak data. We compare GEViTRec's output to what previous visualization recommendation systems would generate, and to manually crafted visualizations used by practitioners. We conducted formative evaluations with ten genEpi experts to assess the relevance and interpretability of our results. Code, Data, and Study Materials Availability: https://github.com/amcrisan/GEVitRec. Anamaria Crisan, Shannah Fisher, Jennifer L. Gardy, Tamara Munzner |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2021 | User Ex Machina : Simulation as a Design Probe in Human-in-the-Loop Text AnalyticsabstractTopic models are widely used analysis techniques for clustering documents and surfacing thematic elements of text corpora. These models remain challenging to optimize and often require a “human-in-the-loop” approach where domain experts use their knowledge to steer and adjust. However, the fragility, incompleteness, and opacity of these models means even minor changes could induce large and potentially undesirable changes in resulting model. In this paper we conduct a simulation-based analysis of human-centered interactions with topic models, with the objective of measuring the sensitivity of topic models to common classes of user actions. We find that user interactions have impacts that differ in magnitude but often negatively affect the quality of the resulting modelling in a way that can be difficult for the user to evaluate. We suggest the incorporation of sensitivity and “multiverse” analyses to topic model interfaces to surface and overcome these deficiencies. Code and Data Availability: https://osf.io/zgqaw Anamaria Crisan, Michael Correll |
CHI | 1 |
| 2021 | Fits and Starts: Enterprise Use of AutoML and the Role of Humans in the LoopabstractAutoML systems can speed up routine data science work and make machine learning available to those without expertise in statistics and computer science. These systems have gained traction in enterprise settings where pools of skilled data workers are limited. In this study, we conduct interviews with 29 individuals from organizations of different sizes to characterize how they currently use, or intend to use, AutoML systems in their data science work. Our investigation also captures how data visualization is used in conjunction with AutoML systems. Our findings identify three usage scenarios for AutoML that resulted in a framework summarizing the level of automation desired by data workers with different levels of expertise. We surfaced the tension between speed and human oversight and found that data visualization can do a poor job balancing the two. Our findings have implications for the design and implementation of human-in-the-loop visual analytics approaches. Anamaria Crisan, Brittany Fiore-Gartland |
CHI | 1 |
| 2021 | Passing the Data Baton : A Retrospective Analysis on Data Science Work and WorkersabstractData science is a rapidly growing discipline and organizations increasingly depend on data science work. Yet the ambiguity around data science, what it is, and who data scientists are can make it difficult for visualization researchers to identify impactful research trajectories. We have conducted a retrospective analysis of data science work and workers as described within the data visualization, human computer interaction, and data science literature. From this analysis we synthesis a comprehensive model that describes data science work and breakdown to data scientists into nine distinct roles. We summarise and reflect on the role that visualization has throughout data science work and the varied needs of data scientists themselves for tooling support. Our findings are intended to arm visualization researchers with a more concrete framing of data science with the hope that it will help them surface innovative opportunities for impacting data science work.Data availability:https://osf.io/z2xpd/?view_only=87fa24be486a473884adb9ffbe8db4ec Anamaria Crisan, Brittany Fiore-Gartland, Melanie Tory |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2019 | A systematic method for surveying data visualizations and a resulting genomic epidemiology visualization typology: GEViTabstractMOTIVATION: Data visualization is an important tool for exploring and communicating findings from genomic and healthcare datasets. Yet, without a systematic way of organizing and describing the design space of data visualizations, researchers may not be aware of the breadth of possible visualization design choices or how to distinguish between good and bad options. RESULTS: We have developed a method that systematically surveys data visualizations using the analysis of both text and images. Our method supports the construction of a visualization design space that is explorable along two axes: why the visualization was created and how it was constructed. We applied our method to a corpus of scientific research articles from infectious disease genomic epidemiology and derived a Genomic Epidemiology Visualization Typology (GEViT) that describes how visualizations were created from a series of chart types, combinations and enhancements. We have also implemented an online gallery that allows others to explore our resulting design space of visualizations. Our results have important implications for visualization design and for researchers intending to develop or use data visualization tools. Finally, the method that we introduce is extensible to constructing visualizations design spaces across other research areas. AVAILABILITY AND IMPLEMENTATION: Our browsable gallery is available at http://gevit.net and all project code can be found at https://github.com/amcrisan/gevitAnalysisRelease. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Anamaria Crisan, Jennifer L. Gardy, Tamara Munzner |
Bioinform. | 1 |
| 2019 | Adjutant: an R-based tool to support topic discovery for systematic and literature reviewsabstractSUMMARY: Adjutant is an open-source, interactive and R-based application to support mining PubMed for literature reviews. Given a PubMed-compatible search query, Adjutant downloads the relevant articles and allows the user to perform an unsupervised clustering analysis to identify data-driven topic clusters. Following clustering, users can also sample documents using different strategies to obtain a more manageable dataset for further analysis. Adjutant makes explicit trade-offs between speed and accuracy, which are modifiable by the user, such that a complete analysis of several thousand documents can take a few minutes. All analytic datasets generated by Adjutant are saved, allowing users to easily conduct other downstream analyses that Adjutant does not explicitly support. AVAILABILITY AND IMPLEMENTATION: Adjutant is implemented in R, using Shiny, and is available at https://github.com/amcrisan/Adjutant. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Anamaria Crisan, Tamara Munzner, Jennifer L. Gardy |
Bioinform. | 1 |
| 2012 | JointSNVMix: a probabilistic model for accurate detection of somatic mutations in normal/tumour paired next-generation sequencing dataabstractMOTIVATION: Identification of somatic single nucleotide variants (SNVs) in tumour genomes is a necessary step in defining the mutational landscapes of cancers. Experimental designs for genome-wide ascertainment of somatic mutations now routinely include next-generation sequencing (NGS) of tumour DNA and matched constitutional DNA from the same individual. This allows investigators to control for germline polymorphisms and distinguish somatic mutations that are unique to the tumour, thus reducing the burden of labour-intensive and expensive downstream experiments needed to verify initial predictions. In order to make full use of such paired datasets, computational tools for simultaneous analysis of tumour-normal paired sequence data are required, but are currently under-developed and under-represented in the bioinformatics literature. RESULTS: In this contribution, we introduce two novel probabilistic graphical models called JointSNVMix1 and JointSNVMix2 for jointly analysing paired tumour-normal digital allelic count data from NGS experiments. In contrast to independent analysis of the tumour and normal data, our method allows statistical strength to be borrowed across the samples and therefore amplifies the statistical power to identify and distinguish both germline and somatic events in a unified probabilistic framework. AVAILABILITY: The JointSNVMix models and four other models discussed in the article are part of the JointSNVMix software package available for download at http://compbio.bccrc.ca CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Andrew Roth, Jiarui Ding, Ryan D. Morin, Anamaria Crisan, Gavin Ha, Ryan Giuliany, Ali Bashashati, Martin Hirst, Gulisa Turashvili, Arusha Oloumi, Marco A. Marra, Samuel Aparicio, Sohrab P. Shah |
Bioinform. | 4 |
| 2010 | SNVMix: predicting single nucleotide variants from next-generation sequencing of tumorsabstractMOTIVATION: Next-generation sequencing (NGS) has enabled whole genome and transcriptome single nucleotide variant (SNV) discovery in cancer. NGS produces millions of short sequence reads that, once aligned to a reference genome sequence, can be interpreted for the presence of SNVs. Although tools exist for SNV discovery from NGS data, none are specifically suited to work with data from tumors, where altered ploidy and tumor cellularity impact the statistical expectations of SNV discovery. RESULTS: We developed three implementations of a probabilistic Binomial mixture model, called SNVMix, designed to infer SNVs from NGS data from tumors to address this problem. The first models allelic counts as observations and infers SNVs and model parameters using an expectation maximization (EM) algorithm and is therefore capable of adjusting to deviation of allelic frequencies inherent in genomically unstable tumor genomes. The second models nucleotide and mapping qualities of the reads by probabilistically weighting the contribution of a read/nucleotide to the inference of a SNV based on the confidence we have in the base call and the read alignment. The third combines filtering out low-quality data in addition to probabilistic weighting of the qualities. We quantitatively evaluated these approaches on 16 ovarian cancer RNASeq datasets with matched genotyping arrays and a human breast cancer genome sequenced to >40x (haploid) coverage with ground truth data and show systematically that the SNVMix models outperform competing approaches. AVAILABILITY: Software and data are available at http://compbio.bccrc.ca CONTACT: [email protected] SUPPLEMANTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rodrigo Goya, Mark G. F. Sun, Ryan D. Morin, Gillian Leung, Gavin Ha, Kimberley C. Wiegand, Janine Senz, Anamaria Crisan, Marco A. Marra, Martin Hirst, David G. Huntsman, Kevin Murphy 0002, Samuel Aparicio, Sohrab P. Shah |
Bioinform. | 8 |