VLDB 2026 Research / reviewers in the wild / expert
Alessandra Pascale
dblp:136/5179
· DBLP profile ↗
14ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-4448-2024ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Language models and text generation · 29% Generative modeling · 15% Trustworthy machine learning · 13% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 63% Knowledge graphs · 37% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 50% Medical and health informatics · 25% Computational social science and digital humanities · 25% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 16 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph reasoning |
1.0 | 1 | 2026 | FactCorrector: A Graph-Inspired Approach to Long-Form Factuality Correction of Large Language Models · ACL (1) 2026 |
Natural language and speech › Language models and text generation › large language model
large language model applications |
0.9 | 1 | 2025 | Forging Time Series with Language: A Large Language Model Approach to Synthetic Data Generation · NeurIPS 2025 |
Machine learning › Generative modeling
synthetic data generation |
0.9 | 1 | 2025 | Forging Time Series with Language: A Large Language Model Approach to Synthetic Data Generation · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
time series generation |
0.9 | 1 | 2025 | Forging Time Series with Language: A Large Language Model Approach to Synthetic Data Generation · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection |
0.8 | 1 | 2024 | A New Computationally Efficient Algorithm to solve Feature Selection for Functional Data Classification in High-dimensional Spaces · ICML 2024 |
Machine learning › Learning theory › statistical pattern recognition
functional data classification |
0.8 | 1 | 2024 | A New Computationally Efficient Algorithm to solve Feature Selection for Functional Data Classification in High-dimensional Spaces · ICML 2024 |
Machine learning › Graph learning › graph neural network
graph convolutional network |
0.8 | 1 | 2024 | Functional Graph Convolutional Networks: A Unified Multi-task and Multi-modal Learning Framework to Facilitate Health and Social-Care Insights · IJCAI 2024 |
Machine learning › Trustworthy machine learning › hallucination
hallucination evaluation |
0.8 | 1 | 2024 | WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia · NeurIPS 2024 |
Natural language and speech › Language models and text generation › retrieval-augmented generation
knowledge conflict |
0.8 | 1 | 2024 | WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia · NeurIPS 2024 |
Natural language and speech › Question answering and dialogue systems
question answering evaluation |
0.8 | 1 | 2024 | WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia · NeurIPS 2024 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.8 | 1 | 2024 | WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia · NeurIPS 2024 |
Information retrieval › document retrieval › domain-specific retrieval › biomedical information retrieval
biomedical literature search |
0.7 | 1 | 2023 | Graph-based Tool for Exploring PubMed Knowledge Base · ICDE 2023 |
Bioinformatics and computational biology
biomedical text mining |
0.3 | 1 | 2025 | Query-driven Document-level Scientific Evidence Extraction from Biomedical Studies · ACL (1) 2025 |
Information retrieval › document retrieval
passage retrieval |
0.2 | 1 | 2024 | WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia · NeurIPS 2024 |
Information retrieval
retrieval models |
0.2 | 1 | 2024 | WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
query-driven retrieval · 1.7document-level extraction · 1.7multimodal learning · 1.5multi-task learning · 1.5logistic loss · 1.5functional principal components · 1.5automated evaluation · 1.5graph-based navigation · 1.3full-text search · 1.3concept extraction · 1.3graph-inspired correction · 1.0tabular embeddings · 0.9autoregressive language model fine-tuning · 0.9large language model · 0.8human evaluation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FactCorrector: A Graph-Inspired Approach to Long-Form Factuality Correction of Large Language ModelsabstractJavier Carnerero-Cano, Massimiliano Pronesti, Radu Marinescu, Tigran T. Tchrakian, James Barry, Jasmina Gajcin, Yufang Hou, Alessandra Pascale, Elizabeth M. Daly. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Javier Carnerero-Cano, Massimiliano Pronesti, Radu Marinescu 0002, Tigran T. Tchrakian, James Barry, Jasmina Gajcin, Yufang Hou 0001, Alessandra Pascale, Elizabeth Daly |
ACL (1) | 8 |
| 2025 | Query-driven Document-level Scientific Evidence Extraction from Biomedical StudiesabstractMassimiliano Pronesti, Joao H Bettencourt-Silva, Paul Flanagan, Alessandra Pascale, Oisín Redmond, Anya Belz, Yufang Hou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Massimiliano Pronesti, Joao H. Bettencourt-Silva, Paul Flanagan, Alessandra Pascale, Oisin Redmond, Anya Belz, Yufang Hou 0001 |
ACL (1) | 4 |
| 2025 | Forging Time Series with Language: A Large Language Model Approach to Synthetic Data GenerationabstractSDForger is a flexible and efficient framework for generating high-quality multivariate time series using LLMs.
Leveraging a compact data representation, SDForger provides synthetic time series generation from a few samples and low-computation fine-tuning of any autoregressive LLM. Specifically, the framework transforms univariate and multivariate signals into tabular embeddings, which are then encoded into text and used to fine-tune the LLM.
At inference, new textual embeddings are sampled and decoded into synthetic time series that retain the original data's statistical properties and temporal dynamics. Across a diverse range of datasets, SDForger outperforms existing generative models in many scenarios, both in similarity-based evaluations and downstream forecasting tasks. By enabling textual conditioning in the generation process, SDForger paves the way for multimodal modeling and the streamlined integration of time series with textual information. The model is open-sourced at https://github.com/IBM/fms-dgt/tree/main/fms_dgt/public/databuilders/time_series. Cécile Rousseau, Tobia Boschi, Giandomenico Cornacchia, Dhaval Salwala, Alessandra Pascale, Juan Bernabé-Moreno |
NeurIPS | 5 |
| 2024 | A New Computationally Efficient Algorithm to solve Feature Selection for Functional Data Classification in High-dimensional SpacesabstractThis paper introduces a novel methodology for Feature Selection for Functional Classification, FSFC, that addresses the challenge of jointly performing feature selection and classification of functional data in scenarios with categorical responses and multivariate longitudinal features. FSFC tackles a newly defined optimization problem that integrates logistic loss and functional features to identify the most crucial variables for classification. To address the minimization procedure, we employ functional principal components and develop a new adaptive version of the Dual Augmented Lagrangian algorithm. The computational efficiency of FSFC enables handling high-dimensional scenarios where the number of features may considerably exceed the number of statistical units. Simulation experiments demonstrate that FSFC outperforms other machine learning and deep learning methods in computational time and classification accuracy. Furthermore, the FSFC feature selection capability can be leveraged to significantly reduce the problem’s dimensionality and enhance the performances of other classification algorithms. The efficacy of FSFC is also demonstrated through a real data application, analyzing relationships between four chronic diseases and other health and demographic factors. FSFC source code is publicly available at https://github.com/IBM/funGCN. Tobia Boschi, Francesca Bonin, Rodrigo Ordonez-Hurtado, Alessandra Pascale, Jonathan Epperlein |
ICML | 4 |
| 2024 | Functional Graph Convolutional Networks: A Unified Multi-task and Multi-modal Learning Framework to Facilitate Health and Social-Care Insights
Tobia Boschi, Francesca Bonin, Rodrigo Ordonez-Hurtado, Cécile Rousseau, Alessandra Pascale, John Dinsmore |
IJCAI | 5 |
| 2024 | WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from WikipediaabstractRetrieval-augmented generation (RAG) has emerged as a promising solution to mitigate the limitations of large language models (LLMs), such as hallucinations and outdated information. However, it remains unclear how LLMs handle knowledge conflicts arising from different augmented retrieved passages, especially when these passages originate from the same source and have equal trustworthiness. In this work, we conduct a comprehensive evaluation of LLM-generated answers to questions that have varying answers based on contradictory passages from Wikipedia, a dataset widely regarded as a high-quality pre-training resource for most LLMs. Specifically, we introduce WikiContradict, a benchmark consisting of 253 high-quality, human-annotated instances designed to assess the performance of LLMs in providing a complete perspective on conflicts from the retrieved documents, rather than choosing one answer over another, when augmented with retrieved passages containing real-world knowledge conflicts. We benchmark a diverse range of both closed and open-source LLMs under different QA scenarios, including RAG with a single passage, and RAG with 2 contradictory passages. Through rigorous human evaluations on a subset of WikiContradict instances involving 5 LLMs and over 3,500 judgements, we shed light on the behaviour and limitations of these models. For instance, when provided with two passages containing contradictory facts, all models struggle to generate answers that accurately reflect the conflicting nature of the context, especially for implicit conflicts requiring reasoning. Since human evaluation is costly, wealso introduce an automated model that estimates LLM performance using a strong open-source language model, achieving an F-score of 0.8. Using this automated metric, we evaluate more than 1,500 answers from seven LLMs across all WikiContradict instances. Yufang Hou 0001, Alessandra Pascale, Javier Carnerero-Cano, Tigran T. Tchrakian, Radu Marinescu 0002, Elizabeth Daly, Inkit Padhi, Prasanna Sattigeri |
NeurIPS | 2 |
| 2023 | Graph-based Tool for Exploring PubMed Knowledge BaseabstractStudies have shown that data retrieval and visualization tools can help health professionals to improve their understanding and communication with patients, their relationship with stakeholders, and their decision-making process. However, not many efforts have been made in this direction. In this paper, we present a prototype system for the indexing, annotation, and visualization of the PubMed knowledge base to enable the search and retrieval of health-related evidence. The proposed tool builds and keeps updated an enriched graph based on PubMed articles associating them with concepts extracted from the Unified Medical Language System (UMLS) Metathesaurus. Moreover, it allows a full-text search and graph-based navigation and supports an overview of concepts and related publications. The proposed architecture enables scale-up thanks to its containerized nature and parallelization capabilities. The code is open-source under the Apache V2 license. Simone Bottoni, Alberto Trombetta, Flavio Bertini 0001, Danilo Montesi, Francesca Bonin, Alessandra Pascale, Martin Gleize, Pierpaolo Tommasi |
ICDE | 6 |
| 2022 | Literature Knowledge Graphs for Risk Modeling: a use case on Diabetes
Joao H. Bettencourt-Silva, Natalia Mulligan, Francesca Bonin, Pierpaolo Tommasi, Alessandra Pascale, Vanessa López |
AMIA | 5 |
| 2021 | Outcome Prediction from Behaviour Change Intervention Evaluations using a Combination of Node and Word Embedding
Debasis Ganguly, Martin Gleize, Yufang Hou 0001, Charles Jochim, Francesca Bonin, Alessandra Pascale, Pierpaolo Tommasi, Pol Mac Aonghusa, Marie Johnston, Mike Kelly, Susan Michie |
AMIA | 6 |
| 2020 | Knowledge Extraction and Prediction from Behavior Science Randomized Controlled Trials: A Case Study in Smoking Cessation
Francesca Bonin, Martin Gleize, Yufang Hou 0001, Debasis Ganguly, Ailbhe Finnerty, Charles Jochim, Alessandra Pascale, Pierpaolo Tommasi, Pol Mac Aonghusa, Susan Michie |
AMIA | 7 |
| 2020 | HBCP Corpus: A New Resource for the Analysis of Behavioural Change Intervention ReportsabstractDue to the fast pace at which research reports in behaviour change are published, researchers, consultants and policymakers would benefit from more automatic ways to process these reports. Automatic extraction of the reports’ intervention content, population, settings and their results etc. are essential in synthesising and summarising the literature. However, to the best of our knowledge, no unique resource exists at the moment to facilitate this synthesis. In this paper, we describe the construction of a corpus of published behaviour change intervention evaluation reports aimed at smoking cessation. We also describe and release the annotation of 57 entities, that can be used as an off-the-shelf data resource for tasks such as entity recognition, etc. Both the corpus and the annotation dataset are being made available to the community. Francesca Bonin, Martin Gleize, Ailbhe Finnerty, Candice Moore, Charles Jochim, Emma Norris, Yufang Hou 0001, Alison J. Wright, Debasis Ganguly, Emily Hayes, Silje Zink, Alessandra Pascale, Pol Mac Aonghusa, Susan Michie |
LREC | 12 |
| 2019 | Building a Risk Model for the Patient-centred Care of Multiple Chronic DiseasesabstractWith the increase of multimorbidity due to population ageing, managing multiple chronic health conditions is a rising challenge. Machine-learning can contribute to a better understanding of persons with multimorbidity (PwMs) and how to design an effective framework of care and support for them. We present a risk model of older PwMs that was derived from the TILDA dataset, a longitudinal study of the ageing Irish population. This model is based on a 26-nodes Bayesian network that represents patients possibly having one or more chronic conditions among diabetes, chronic obstructive pulmonary disease and arthritis, through a joint probability distribution of demographic, symptomatic and behavioral dimensions. We describe our method, give an exploratory analysis of the risk model, and assess its prediction accuracy in a cross-validation experiment. Finally we discuss its use in supporting management of care for PwMs, drawing on comments from health practitioners on the model. Stéphane Deparis, Pierpaolo Tommasi, Alessandra Pascale, Hicham Rifai, Julie Doyle, John Dinsmore |
BIBM | 3 |
| 2014 | Cooperative Bayesian Estimation of Vehicular Traffic in Large-Scale NetworksabstractIntelligent transportation systems have enormous potential for improving the quality of our lives. They rely on traffic monitoring and control infrastructures to enable an efficient management of mobility. A crucial task is the estimation or prediction of traffic flows by large-scale sensor networks, which is a topic that has been attracting increasing attention in recent years because of its relevance in traffic control over urban areas or freeways. In this paper, we propose an innovative stochastic method for vehicular traffic estimation based on a distributed reconstruction of the density field through the cooperation of smaller monitoring subnetworks. The method guarantees high accuracy (because of information sharing) and, at the same time, moderate computational cost (due to distributed processing). Moreover, subnetworks do not need to exchange sensitive information (e.g., raw data) but simply traffic beliefs. We evaluate the performance of the method on simulated single-lane road scenarios, highlighting the potential benefits of the cooperative approach. As an example of application, we consider a fragmented monitoring scenario characterized by several sensor failures and we show how the proposed approach can overcome the problem related to the sensor malfunctions leveraging on information shared with neighboring subnetworks. Alessandra Pascale, Monica Nicoli, Umberto Spagnolini |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2013 | Estimation of highway traffic from sparse sensors: Stochastic modeling and particle filteringabstractTraffic control is essential for the achievement of a sustainable and safe mobility. Monitoring systems deployed over the roads collect a great amount of traffic data that must be efficiently processed by statistical methods to draw traffic macroparameters that are needed for control operations. In this paper we propose a particle filtering approach to estimate the density over a road network starting from noisy and sparse measurements provided by road-embedded sensors. We propose a new Bayesian framework based on the link-node cell transmission model to take into account the stochastic behavior of traffic and the hysteresis phenomenon that are typically observed in real data. Numerical tests show that the estimation method is able to reliably reconstruct the traffic field even in case of very sparse sensor deployments. Alessandra Pascale, Gabriel Gomes, Monica Nicoli |
ICASSP | 1 |