EDBT 2026 Demo / reviewers in the wild / expert
Cristina Sarasua
dblp:91/7572
· DBLP profile ↗
19ranked-venue papers
3as first author
12since 2021 · last 2025
0000-0002-2076-9584ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 6 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HypER: Literature-grounded Hypothesis Generation and Distillation with ProvenanceabstractLarge Language models have demonstrated promising performance in research ideation across scientific domains.Hypothesis development, the process of generating a highly specific declarative statement connecting a research idea with empirical validation, has received relatively less attention.Existing approaches trivially deploy retrieval augmentation and focus only on the quality of the final output ignoring the underlying reasoning process behind ideation.We present HypER (Hypothesis Generation with Explanation and Reasoning), a small language model (SLM) trained for literature-guided reasoning and evidence-based hypothesis generation.HypER is trained in a multi-task setting to discriminate between valid and invalid scientific reasoning chains in presence of controlled distractions.We find that HypER outperforms the base model, distinguishing valid from invalid reasoning chains (+22% average absolute F1), generates better evidence-grounded hypotheses (0.327 vs. 0.305 base model) with high feasibility and impact as judged by human experts (>3.5 on 5-point Likert scale).Resource at . Example of a valid reasoning chainTitle: Evidence suggesting that a chronic disease self-management program can improve health status while reducing hospitalization Abstract: This study evaluated the effectiveness (changes in health behaviors, health status, and health service utilization) of a self-management program for chronic disease ... Rosni Vasu, Chandrayee Basu, Bhavana Dalvi, Cristina Sarasua, Peter Clark, Abraham Bernstein |
EMNLP | 4 |
| 2025 | AI support for data scientists: An empirical study on workflow and alternative code recommendationsabstractDespite the popularity of AI assistants for coding activities, there is limited empirical work on whether these coding assistants can help users complete data science tasks. Moreover, in data science programming, exploring alternative paths has been widely advocated, as such paths may lead to diverse understandings and conclusions (Gelman and Loken 2013; Kale et al. 2019). Whether existing AI-based coding assistants can support data scientists in exploring the relevant alternative paths remains unexplored. To fill this gap, we conducted a mixed-methods study to understand how data scientists solved different data science tasks with the help of an AI-based coding assistant that provides explicit alternatives as recommendations throughout the data science workflow. Specifically, we quantitatively investigated whether the users accept the code recommendations, including alternative recommendations, by the AI assistant and whether the recommendations are helpful when completing descriptive and predictive data science tasks. Through the empirical study, we also investigated if including information about the data science step (e.g., data exploration) they seek recommendations for in a prompt leads to helpful recommendations. In our study, we found that including the data science step in a prompt had a statistically significant improvement in the acceptance of recommendations, whereas the presence of alternatives did not lead to any significant differences. Our study also shows a statistically significant difference in the acceptance and usefulness of recommendations between descriptive and predictive tasks. Participants generally had positive sentiments regarding AI assistance and our proposed interface. We share further insights on the interactions that emerged during the study and the challenges that our users encountered while solving their data science tasks. Supplementary Information: The online version contains supplementary material available at 10.1007/s10664-025-10622-4. Dhivyabharathi Ramasamy, Cristina Sarasua, Abraham Bernstein |
Empir. Softw. Eng. | 2 |
| 2024 | A Unified Benchmark for Argument Mining
Florian Ruosch, John Lawrence, Cristina Sarasua, Abraham Bernstein |
COMMA | 3 |
| 2024 | Toward the Argument Web of Science
Florian Ruosch, Cristina Sarasua, Chris Reed 0001, Abraham Bernstein |
COMMA | 2 |
| 2024 | Documents with Integrated Visually Annotated Arguments
Florian Ruosch, Joel Watter, Cristina Sarasua, Abraham Bernstein |
COMMA | 3 |
| 2024 | Estimating the Semantic Density of Visual MediaabstractImage descriptions provide precious information for a myriad of visual media management tasks ranging from image classification to image search. The value of such curated collections comes from their diverse content and their accompanying extensive annotations. Such annotations are typically supplied by communities, where users (often volunteers) curate labels and/or descriptions of images. Supporting users in their quest to increase (overall) description completeness where possible is, therefore, of utmost importance. Luca Rossetto, Cristina Sarasua, Abraham Bernstein |
ACM Multimedia | 2 |
| 2024 | Fast and Adaptive Questionnaires for Voting Advice Applications
Fynn Bachmann, Cristina Sarasua, Abraham Bernstein |
ECML/PKDD (10) | 2 |
| 2024 | SciHyp: A Fine-Grained Dataset Describing Hypotheses and Their Components from Scientific Articles
Rosni Vasu, Cristina Sarasua, Abraham Bernstein |
ISWC (3) | 2 |
| 2023 | DREAM: Deployment of Recombination and Ensembles in Argument MiningabstractCurrent approaches to Argument Mining (AM) tend to take a holistic or black-box view of the overall pipeline.This paper, in contrast, aims to provide a solution to achieve increased performance based on current components instead of independent all-new solutions.To that end, it presents the Deployment of Recombination and Ensemble methods for Argument Miners (DREAM) framework that allows for the (automated) combination of AM components.Using ensemble methods, DREAM combines sets of AM systems to improve accuracy for the four tasks in the AM pipeline.Furthermore, it leverages recombination by using different argument miners elements throughout the pipeline.Experiments with five systems previously included in a benchmark show that the systems combined with DREAM can outperform the previous best single systems in terms of accuracy measured by an AM benchmark. Florian Ruosch, Cristina Sarasua, Abraham Bernstein |
EMNLP | 2 |
| 2023 | Workflow analysis of data science code in public GitHub repositoriesabstractAbstract Despite the ubiquity of data science, we are far from rigorously understanding how coding in data science is performed. Even though the scientific literature has hinted at the iterative and explorative nature of data science coding, we need further empirical evidence to understand this practice and its workflows in detail. Such understanding is critical to recognise the needs of data scientists and, for instance, inform tooling support. To obtain a deeper understanding of the iterative and explorative nature of data science coding, we analysed 470 Jupyter notebooks publicly available in GitHub repositories. We focused on the extent to which data scientists transition between different types of data science activities, or steps (such as data preprocessing and modelling), as well as the frequency and co-occurrence of such transitions. For our analysis, we developed a dataset with the help of five data science experts, who manually annotated the data science steps for each code cell within the aforementioned 470 notebooks. Using the first-order Markov chain model, we extracted the transitions and analysed the transition probabilities between the different steps. In addition to providing deeper insights into the implementation practices of data science coding, our results provide evidence that the steps in a data science workflow are indeed iterative and reveal specific patterns. We also evaluated the use of the annotated dataset to train machine-learning classifiers to predict the data science step(s) of a given code cell. We investigate the representativeness of the classification by comparing the workflow analysis applied to (a) the predicted data set and (b) the data set labelled by experts, finding an F1-score of about 71% for the 10-class data science step prediction problem. Dhivyabharathi Ramasamy, Cristina Sarasua, Alberto Bacchelli, Abraham Bernstein |
Empir. Softw. Eng. | 2 |
| 2023 | Visualising data science workflows to support third-party notebook comprehension: an empirical studyabstractAbstract Data science is an exploratory and iterative process that often leads to complex and unstructured code. This code is usually poorly documented and, consequently, hard to understand by a third party. In this paper, we first collect empirical evidence for the non-linearity of data science code from real-world Jupyter notebooks, confirming the need for new approaches that aid in data science code interaction and comprehension. Second, we propose a visualisation method that elucidates implicit workflow information in data science code and assists data scientists in navigating the so-called garden of forking paths in non-linear code. The visualisation also provides information such as the rationale and the identification of the data science pipeline step based on cell annotations. We conducted a user experiment with data scientists to evaluate the proposed method, assessing the influence of (i) different workflow visualisations and (ii) cell annotations on code comprehension. Our results show that visualising the exploration helps the users obtain an overview of the notebook, significantly improving code comprehension. Furthermore, our qualitative analysis provides more insights into the difficulties faced during data science code comprehension. Dhivyabharathi Ramasamy, Cristina Sarasua, Alberto Bacchelli, Abraham Bernstein |
Empir. Softw. Eng. | 2 |
| 2021 | The Impact of Task Abandonment in CrowdsourcingabstractCrowdsourcing has become a standard methodology to collect manually annotated data such as relevance judgments at scale. On crowdsourcing platforms like Amazon MTurk or FigureEight, crowd workers select tasks to work on based on different dimensions such as task reward and requester reputation. Requesters then receive the judgments of workers who self-selected into the tasks and completed them successfully. Several crowd workers, however, preview tasks, begin working on them, reaching varying stages of task completion without finally submitting their work. Such behavior results in unrewarded effort which remains invisible to requesters. In this paper, we conduct an investigation of the phenomenon of task abandonment, the act of workers previewing or beginning a task and deciding not to complete it. We follow a three-fold methodology which includes 1) investigating the prevalence and causes of task abandonment by means of a survey over different crowdsourcing platforms, 2) data-driven analysis of logs collected during a large-scale relevance judgment experiment, and 3) controlled experiments measuring the effect of different dimensions on abandonment. Our results show that task abandonment is a widely spread phenomenon. Apart from accounting for a considerable amount of wasted human effort, this bears important implications on the hourly wages of workers as they are not rewarded for tasks that they do not complete. We also show how task abandonment may have strong implications on the use of collected data (for example, on the evaluation of Information Retrieval systems). Lei Han 0003, Kevin Roitero, Ujwal Gadiraju, Cristina Sarasua, Alessandro Checco, Eddy Maddalena, Gianluca Demartini |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2020 | Crowd Worker Strategies in Relevance Judgment TasksabstractCrowdsourcing is a popular technique to collect large amounts of human-generated labels, such as relevance judgments used to create information retrieval (IR) evaluation collections. Previous research has shown how collecting high quality labels from a crowdsourcing platform can be challenging. Existing quality assurance techniques focus on answer aggregation or on the use of gold questions where ground-truth data allows to check for the quality of the responses. Lei Han 0003, Eddy Maddalena, Alessandro Checco, Cristina Sarasua, Ujwal Gadiraju, Kevin Roitero, Gianluca Demartini |
WSDM | 4 |
| 2019 | Non-parametric Class Completeness Estimators for Collaborative Knowledge Graphs - The Case of Wikidata
Michael Luggen, Djellel Eddine Difallah, Cristina Sarasua, Gianluca Demartini, Philippe Cudré-Mauroux |
ISWC (1) | 3 |
| 2019 | All Those Wasted Hours: On Task Abandonment in CrowdsourcingabstractCrowdsourcing has become a standard methodology to collect manually annotated data such as relevance judgments at scale. On crowdsourcing platforms like Amazon MTurk or FigureEight, crowd workers select tasks to work on based on different dimensions such as task reward and requester reputation. Requesters then receive the judgments of workers who self-selected into the tasks and completed them successfully. Several crowd workers, however, preview tasks, begin working on them, reaching varying stages of task completion without finally submitting their work. Such behavior results in unrewarded effort which remains invisible to requesters. In this paper, we conduct the first investigation into the phenomenon of task abandonment, the act of workers previewing or beginning a task and deciding not to complete it. We follow a three-fold methodology which includes 1) investigating the prevalence and causes of task abandonment by means of a survey over different crowdsourcing platforms, 2) data-driven analyses of logs collected during a large-scale relevance judgment experiment, and 3) controlled experiments measuring the effect of different dimensions on abandonment. Our results show that task abandonment is a widely spread phenomenon. Apart from accounting for a considerable amount of wasted human effort, this bears important implications on the hourly wages of workers as they are not rewarded for tasks that they do not complete. We also show how task abandonment may have strong implications on the use of collected data (for example, on the evaluation of IR systems). Lei Han 0003, Kevin Roitero, Ujwal Gadiraju, Cristina Sarasua, Alessandro Checco, Eddy Maddalena, Gianluca Demartini |
WSDM | 4 |
| 2019 | The Evolution of Power and Standard Wikidata Editors: Comparing Editing Behavior over Time to Predict Lifespan and Volume of Edits
Cristina Sarasua, Alessandro Checco, Gianluca Demartini, Djellel Eddine Difallah, Michael Feldman 0001, Lydia Pintscher |
Comput. Support. Cooperative Work. | 1 |
| 2017 | Methods for Intrinsic Evaluation of Links in the Web of Data
Cristina Sarasua, Steffen Staab, Matthias Thimm |
ESWC (1) | 1 |
| 2012 | CrowdMap: Crowdsourcing Ontology Alignment with Microtasks
Cristina Sarasua, Elena Simperl, Natasha F. Noy |
ISWC (1) | 1 |
| 2009 | MPEG-7 Compliant Indexation Tool for Multimedia Tourist Content
María Teresa Linaza, Cristina Sarasua, Yolanda Cobos |
ENTER | 2 |