VLDB 2026 Research / reviewers in the wild / expert
Veronica Guidetti
dblp:47/5792
· DBLP profile ↗
7ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0003-2233-0097ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Trustworthy machine learning · 33% Knowledge representation and reasoning · 25% Efficient and distributed learning · 25% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
federated learning |
0.9 | 1 | 2025 | Mining Trustworthy Symbolic Regression Models in Federated Settings · ICDM 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | Mining Trustworthy Symbolic Regression Models in Federated Settings · ICDM 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
symbolic regression |
0.9 | 1 | 2025 | Mining Trustworthy Symbolic Regression Models in Federated Settings · ICDM 2025 |
Machine learning › Optimization for machine learning
multi-objective optimization |
0.6 | 1 | 2022 | Multi-Objective Symbolic Regression for Data-Driven Scoring System Management · ICDM 2022 |
Machine learning › Trustworthy machine learning
robustness |
0.3 | 1 | 2025 | Mining Trustworthy Symbolic Regression Models in Federated Settings · ICDM 2025 |
Methods — techniques the papers use, named apart from their topics
symbolic regression · 2.0multi-objective optimization · 1.1federated learning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCoRE: Streamlined corpus-based relation extraction using multi-label contrastive learning and Bayesian kNNabstractThe growing demand for efficient knowledge graph (KG) enrichment leveraging external corpora has intensified interest in relation extraction (RE), particularly under low-supervision settings. To address the need for adaptable and noise-resilient RE solutions that integrate seamlessly with pre-trained large language models (PLMs), we introduce SCoRE, a modular and cost-effective sentence-level RE system. SCoRE enables easy PLM switching, requires no finetuning, and adapts smoothly to diverse corpora and KGs. By combining supervised contrastive learning with a Bayesian k-Nearest Neighbors (kNN) classifier for multi-label classification, it delivers robust performance despite the noisy annotations of distantly supervised corpora. To improve RE evaluation, we propose two novel metrics: Correlatin Structure Distance (CSD), measuring the alignment between learned relational patterns and KG structures, and Precision at R (P@R), assessing utility as a recommender system. We also release Wiki20d, a benchmark dataset replicating real-world RE conditions where only KG-derived annotations are available. Experiments on five benchmarks demonstrate that SCoRE matches or slightly surpasses state-of-the-art methods (average gains of +3.2 in micro-F1 and +5.9 in macro-F1 against fully reproducible baselines), while reducing the training burden by more than an order of magnitude ( ≈ 99 % lower energy consumption in kWh) Further analyses reveal that increasing model complexity, as seen in prior work, degrades performance, highlighting the advantages of SCoRE’s minimal design. Combining efficiency, modularity, and scalability, SCoRE stands as an optimal choice for real-world RE applications. Luca Mariotti, Veronica Guidetti, Federica Mandreoli |
Knowl. Based Syst. | 2 |
| 2025 | Mining Trustworthy Symbolic Regression Models in Federated Settings
Mattia Billa, Veronica Guidetti, Luca La Rocca, Federica Mandreoli |
ICDM | 2 |
| 2023 | Death After Liver Transplantation: Mining Interpretable Risk Factors for Survival PredictionabstractThis study introduces a novel approach to mine risk factors for short-term death after liver transplantation (LT). The method outputs intelligible survival models by combining Cox’s regression with a genetic programming technique known as multi-objective symbolic regression (MOSR). We consider 485 Electronic Health Records (EHRs) of patients who underwent LT, containing information on hospitalization and preoperative conditions, with a focus on infections and colonizations by multi-resistant Gram-negative bacteria. We evaluate MOSR outcomes against several performance metrics and demonstrate that they are well-calibrated, predictive, safe, and parsimonious. Finally, we select the most promising post-LT early survival risk score based on information criteria, performance, and out-of-distribution safety. Validating this technique at a multicenter level could improve service pipeline logistics through a trustworthy machine-learning method. Veronica Guidetti, Giovanni Dolci, Erica Franceschini, Erica Bacca, Giulia Jole Burastero, Davide Ferrari 0004, Valentina Serra, Fabrizio Di Benedetto, Cristina Mussini, Federica Mandreoli |
DSAA | 1 |
| 2023 | Generalizing Similarity in Noisy Setups: The DIBS PhenomenonabstractThis work uncovers an interplay among data density, noise, and the generalization ability in similarity learning. We consider Siamese Neural Networks (SNNs), which are the basic form of contrastive learning, and explore two types of noise that can impact SNNs, Pair Label Noise (PLN) and Single Label Noise (SLN). Our investigation reveals that SNNs exhibit double descent behaviour regardless of the training setup and that it is further exacerbated by noise. We demonstrate that the density of data pairs is crucial for generalization. When SNNs are trained on sparse datasets with the same amount of PLN or SLN, they exhibit comparable generalization properties. However, when using dense datasets, PLN cases generalize worse than SLN ones in the overparametrized region, leading to a phenomenon we call Density-Induced Break of Similarity (DIBS). In this regime, PLN similarity violation becomes macroscopical, corrupting the dataset to the point where complete interpolation cannot be achieved, regardless of the number of model parameters. Our analysis also delves into the correspondence between online optimization and offline generalization in similarity learning. The results show that this equivalence fails in the presence of label noise in all the scenarios considered. Nayara Fonseca, Veronica Guidetti |
ECAI | 2 |
| 2022 | Multi-objective Symbolic Regression to Generate Data-driven, Non-fixed Structure and Intelligible Mortality Predictors using EHR: Binary Classification Methodology and Comparison with State-of-the-art
Veronica Guidetti, Vasa Curcin, Yanzhong Wang |
AMIA | 2 |
| 2022 | Multi-Objective Symbolic Regression for Data-Driven Scoring System ManagementabstractScores are mathematical combinations of elementary indicators (EIs) widely used to measure complex phenomena. Upon the theoretical framework definition, score construction requires a method to aggregate EIs. Aggregation is usually chosen among known methodologies fixing its shape through a try and error approach. Only then are the predictive power, the distribution of the index, and its ability to stratify the population measured. In this paper, we propose a novel data-driven approach that generates analytic aggregation methods relying on multi-objective symbolic regression. We translate the properties that the index must exhibit into optimization goals so that optimal index candidates replicate target variables, data balancing, and stratification. We run experiments on real data sets to solve three main score management problems: data-driven score simplification, generation, and combination. The results obtained show the effectiveness and robustness of the proposed approach. Davide Ferrari 0004, Veronica Guidetti, Federica Mandreoli |
ICDM | 2 |
| 2008 | Semantic Bookmarking and Search in the Earth Observation Domain
Francesca Fallucchi, Maria Teresa Pazienza, Noemi Scarpato, Armando Stellato, Luigi Fusco, Veronica Guidetti |
KES (3) | 6 |