VLDB 2026 Research / reviewers in the wild / expert
Ana Trisovic
dblp:250/2739
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0003-1991-0533ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mapping the Impact of Foundation Models on the UN Sustainable Development GoalsabstractUnderstanding how Foundation Models contribute to sustainability-focused science is difficult due to the broad scope of AI applications and the diverse language used across disciplines. In this work, we analyze the use of Foundation Models in scientific publications aligned with the Sustainable Development Goals (SDGs). Starting from a large citation network of 269K papers, we isolate those that actively adopt Foundation Models and classify their relevance to SDGs using zero-shot language model classification. By linking citation intent to research themes and excluding methodological literature, we construct a filtered view of real-world AI applications in sciences relevant to the SDGs. Our findings reveal that Foundation Model usage is heavily concentrated on a few SDGs (most notably SDG 3 and SDG 9, while many are virtually unaddressed), highlighting uneven alignment and suggesting gaps where AI research could be directed to better support global sustainability efforts. Janakan Sivaloganathan, Ana Trisovic, Neil C. Thompson |
eScience | 2 |
| 2025 | Improving FAIR Compliance for High-Dimensional Data via Automated Metadata ExtractionabstractHigh-dimensional datasets are becoming an increasingly vital asset in the machine learning and AI domains due to their ability to capture large volumes of structured, self-descriptive information. Research data repositories play a key role in supporting the documentation, discoverability, and reuse of such data. This paper highlights the growing presence of high-dimensional formats—particularly NetCDF and HDF5—underscoring the need for better infrastructure to support them. We use a large language model to assess the quality of metadata associated with these files and find that user-provided metadata often falls short when compared to the richness of embedded metadata within the files themselves. In response, we implement an automated metadata extraction process during file ingestion, offering a practical pathway to FAIR-ify high-dimensional data. Our empirical analysis and technical solution are demonstrated through integration with the Dataverse data repository platform. Ana Trisovic, Jan Range, Philip Durbin, Amber Leahey, Danielle Braun |
eScience | 1 |
| 2024 | SpaCE: The Spatial Confounding EnvironmentabstractSpatial confounding poses a significant challenge in scientific studies involving spatial data, where unobserved spatial variables can influence both treatment and outcome, possibly leading to spurious associations. To address this problem, we introduce SpaCE: The Spatial Confounding Environment, the first toolkit to provide realistic benchmark datasets and tools for systematically evaluating causal inference methods designed to alleviate spatial confounding. Each dataset includes training data, true counterfactuals, a spatial graph with coordinates, and smoothness and confounding scores characterizing the effect of a missing spatial confounder. It also includes realistic semi-synthetic outcomes and counterfactuals, generated using state-of-the-art machine learning ensembles, following best practices for causal inference benchmarks. The datasets cover real treatment and covariates from diverse domains, including climate, health and social sciences. SpaCE facilitates an automated end-to-end pipeline, simplifying data loading, experimental setup, and evaluating machine learning and causal inference models. The SpaCE project provides several dozens of datasets of diverse sizes and spatial complexity. It is publicly available as a Python package, encouraging community feedback and contributions. Mauricio Tec, Ana Trisovic, Michelle Audirac, Sophie Woodward, Jie Kate Hu, Naeem Khoshnevis, Francesca Dominici |
ICLR | 2 |
| 2022 | Toward Reusable Science with Readable Code and ReproducibilityabstractAn essential part of research and scientific communication is researchers' ability to reproduce the results of others. While there have been increasing standards for authors to make data and code available, many of these files are hard to re-execute in practice, leading to a lack of research reproducibility. This poses a major problem for students and researchers in the same field who cannot leverage the previously published findings for study or further inquiry. To address this, we propose an open-source platform named RE3 that helps improve the reproducibility and readability of research projects involving R code. Our platform incorporates assessing code readability with a machine learning model trained on a code readability survey and an automatic containerization service that executes code files and warns users of reproducibility errors. This process helps ensure the reproducibility and readability of projects and therefore fast-track their verification and reuse. Layan Bahaidarah, Ethan Hung, Andreas Francisco De Melo Oliveira, Jyotsna Penumaka, Lukas Rosario, Ana Trisovic |
e-Science | 6 |
| 2022 | Cluster Analysis of Open Research Data and a Case for Replication MetadataabstractResearch data are often released upon journal publication to enable result verification and reproducibility. For that reason, research dissemination infrastructures typically support diverse datasets coming from numerous disciplines, from tabular data and program code to audio-visual files. Metadata, or data about data, is critical to making research outputs adequately documented and FAIR. Aiming to contribute to the discussions on the development of metadata for research outputs, I conducted an exploratory analysis to determine how research datasets cluster based on what researchers organically deposit together. I use the content of over 40,000 datasets from the Harvard Dataverse research data repository as my sample for the cluster analysis. I find that the majority of the clusters are formed by single-type datasets, while in the rest of the sample, no meaningful clusters can be identified. For the result interpretation, I use the metadata standard employed by DataCite, a leading organization for documenting a scholarly record, and map existing resource types to my results. About 65% of the sample can be described with a single-type metadata (such as Dataset, Software or Report), while the rest would require aggregate metadata types. Though DataCite supports an aggregate type such as a Collection, I argue that a significant number of datasets, in particular those containing both data and code files (about 20% of the sample) would be more accurately described as a Replication resource metadata type. Such resource type would be particularly useful in facilitating research reproducibility. Ana Trisovic |
e-Science | 1 |