VLDB 2026 Research / reviewers in the wild / expert
Leonardo F. R. Ribeiro
dblp:245/8769 · also Leonardo Filipe Rodrigues Ribeiro
· DBLP profile ↗
15ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0003-2639-942XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 7 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DeepFact: Co-Evolving Benchmarks and Agents for Deep Research FactualityabstractYukun Huang, Leonardo F. R. Ribeiro, Momchil Hardalov, Bhuwan Dhingra, Markus Dreyer, Venkatesh Saligrama. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Leonardo F. R. Ribeiro, Momchil Hardalov, Bhuwan Dhingra, Markus Dreyer, Venkatesh Saligrama |
ACL (1) | 2 |
| 2026 | Benchmarking Deflection and Hallucination in Large Vision-Language ModelsabstractNicholas Moratelli, Christopher Davis, Leonardo F. R. Ribeiro, Bill Byrne, Gonzalo Iglesias. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Nicholas Moratelli, Leonardo F. R. Ribeiro, William J. Byrne, Gonzalo Iglesias |
ACL (1) | 3 |
| 2026 | What Does LLM Refinement Actually Improve? A Systematic Study on Document-Level Literary TranslationabstractShaomu Tan, Dawei Zhu, Ke Tran, Michael Denkowski, Sony Trenous, Leonardo F. R. Ribeiro, Bill Byrne, Felix Hieber. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shaomu Tan, Ke Tran, Michael J. Denkowski, Sony Trenous, Leonardo F. R. Ribeiro, William J. Byrne, Felix Hieber |
ACL (1) | 6 |
| 2025 | RAGferee: Building Contextual Reward Models for Retrieval-Augmented GenerationabstractAndrei Catalin Coman, Ionut Teodor Sorodoc, Leonardo F. R. Ribeiro, Bill Byrne, James Henderson, Adrià de Gispert. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Andrei Catalin, Ionut Sorodoc, Leonardo F. R. Ribeiro, William J. Byrne, Adrià de Gispert |
EMNLP | 3 |
| 2024 | REFINESUMM: Self-Refining MLLM for Generating a Multimodal Summarization DatasetabstractMultimodal Large Language Models (MLLMs) excel at synthesizing key information from diverse sources.However, generating accurate and faithful multimodal summaries is challenging, primarily due to the lack of appropriate multimodal datasets for fine-tuning that meaningfully integrate textual and visual modalities.To address this gap, we present a new dataset specifically designed for image-text multimodal summarization, harnessing the capabilities of state-of-the-art MLLMs.We generate summaries from Wikipedia sections and corresponding images and evaluate them across text-based, visual and multimodal dimensions, employing reference-free metrics.To refine the dataset, we: (1) filter the MLLM-generated summaries by training a critic model on human annotations and using its predictions to remove low-quality summaries; (2) fine-tune the MLLM with the filtered high-quality summaries; (3) use the fine-tuned model in turn to regenerate the summaries.This self-refinement process notably improves summary quality, as measured by human judgments and automatic multimodal metrics, resulting in a valuable dataset for multimodal summarization research.1 * Work done as an intern at Amazon AGI. 1 The dataset is publicly available at https://github. com/amazon-science/refinesumm.The Italian wall lizard or ruin lizard (Podarcis siculus, from the Greek meaning agile and feet) is a species of lizard in the family Lacertidae.P. siculus is native to Bosnia and Herzegovina, Croatia, France, Italy, Serbia, Montenegro, Slovenia, and Switzerland, but has also been introduced to Spain, .... Vaidehi Patil, Leonardo F. R. Ribeiro, Mengwen Liu, Mohit Bansal, Markus Dreyer |
ACL (1) | 2 |
| 2024 | Speechworthy Instruction-tuned Language ModelsabstractHyundong Justin Cho, Nicolaas Paul Jedema, Leonardo F. R. Ribeiro, Karishma Sharma, Pedro Szekely, Alessandro Moschitti, Ruben Janssen, Jonathan May. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Hyundong Cho, Nicolaas Paul Jedema, Leonardo F. R. Ribeiro, Karishma Sharma, Pedro A. Szekely, Alessandro Moschitti, Ruben Janssen, Jonathan May |
EMNLP | 3 |
| 2023 | Generating Summaries with Controllable Readability LevelsabstractReadability refers to how easily a reader can understand a written text.Several factors affect the readability level, such as the complexity of the text, its subject matter, and the reader's background knowledge.Generating summaries based on different readability levels is critical for enabling knowledge consumption by diverse audiences.However, current text generation approaches lack refined control, resulting in texts that are not customized to readers' proficiency levels.In this work, we bridge this gap and study techniques to generate summaries at specified readability levels.Unlike previous methods that focus on a specific readability level (e.g., lay summarization), we generate summaries with fine-grained control over their readability.We develop three text generation techniques for controlling readability:(1) instruction-based readability control, (2) reinforcement learning to minimize the gap between requested and observed readability and (3) a decoding approach that uses lookahead to estimate the readability of upcoming decoding steps.We show that our generation methods significantly improve readability control on news summarization (CNN/DM dataset), as measured by various readability metrics and human judgement, establishing strong baselines for controllable readability in summarization. 1 Leonardo F. R. Ribeiro, Mohit Bansal, Markus Dreyer |
EMNLP | 1 |
| 2022 | Incorporating Relevance Feedback for Information-Seeking Retrieval using Few-Shot Document Re-RankingabstractPairing a lexical retriever with a neural reranking model has set state-of-the-art performance on large-scale information retrieval datasets.This pipeline covers scenarios like question answering or navigational queries, however, for information-seeking scenarios, users often provide information on whether a document is relevant to their query in form of clicks or explicit feedback.Therefore, in this work, we explore how relevance feedback can be directly integrated into neural re-ranking models by adopting few-shot and parameterefficient learning techniques.Specifically, we introduce a kNN approach that re-ranks documents based on their similarity with the query and the documents the user considers relevant.Further, we explore Cross-Encoder models that we pre-train using meta-learning and subsequently fine-tune for each query, training only on the feedback documents.To evaluate our different integration strategies, we transform four existing information retrieval datasets into the relevance feedback scenario.Extensive experiments demonstrate that integrating relevance feedback directly in neural re-ranking models improves their performance, and fusing lexical ranking with our best performing neural reranker outperforms all other methods by 5.2% nDCG@20. Tim Baumgärtner, Leonardo F. R. Ribeiro, Nils Reimers 0001, Iryna Gurevych |
EMNLP | 2 |
| 2022 | FactGraph: Evaluating Factuality in Summarization with Semantic Graph RepresentationsabstractLeonardo Ribeiro, Mengwen Liu, Iryna Gurevych, Markus Dreyer, Mohit Bansal. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Leonardo F. R. Ribeiro, Mengwen Liu, Iryna Gurevych, Markus Dreyer, Mohit Bansal |
NAACL-HLT | 1 |
| 2021 | Smelting Gold and Silver for Improved Multilingual AMR-to-Text GenerationabstractRecent work on multilingual AMR-to-text generation has exclusively focused on data augmentation strategies that utilize silver AMR.However, this assumes a high quality of generated AMRs, potentially limiting the transferability to the target task.In this paper, we investigate different techniques for automatically generating AMR annotations, where we aim to study which source of information yields better multilingual results.Our models trained on gold AMR with silver (machine translated) sentences outperform approaches which leverage generated silver AMR.We find that combining both complementary sources of information further improves multilingual AMR-to-text generation.Our models surpass the previous state of the art for German, Italian, Spanish, and Chinese by a large margin. 1 Leonardo F. R. Ribeiro, Jonas Pfeiffer, Iryna Gurevych |
EMNLP (1) | 1 |
| 2021 | Structural Adapters in Pretrained Language Models for AMR-to-Text GenerationabstractPretrained language models (PLM) have recently advanced graph-to-text generation, where the input graph is linearized into a sequence and fed into the PLM to obtain its representation.However, efficiently encoding the graph structure in PLMs is challenging because such models were pretrained on natural language, and modeling structured data may lead to catastrophic forgetting of distributional knowledge.In this paper, we propose STRUCTADAPT, an adapter method to encode graph structure into PLMs.Contrary to prior work, STRUCTADAPT effectively models interactions among the nodes based on the graph connectivity, only training graph structure-aware adapter parameters.In this way, we incorporate task-specific knowledge while maintaining the topological structure of the graph.We empirically show the benefits of explicitly encoding graph structure into PLMs using STRUCTADAPT, outperforming the state of the art on two AMR-to-text datasets, training only 5.1% of the PLM parameters. 1 Leonardo F. R. Ribeiro, Iryna Gurevych |
EMNLP (1) | 1 |
| 2020 | Modeling Global and Local Node Contexts for Text Generation from Knowledge GraphsabstractRecent graph-to-text models generate text from graph-based data using either global or local aggregation to learn node representations. Global node encoding allows explicit communication between two distant nodes, thereby neglecting graph topology as all nodes are directly connected. In contrast, local node encoding considers the relations between neighbor nodes capturing the graph structure, but it can fail to capture long-range relations. In this work, we gather both encoding strategies, proposing novel neural models that encode an input graph combining both global and local node contexts, in order to learn better contextualized node embeddings. In our experiments, we demonstrate that our approaches lead to significant improvements on two graph-to-text datasets achieving BLEU scores of 18.01 on the AGENDA dataset, and 63.69 on the WebNLG dataset for seen categories, outperforming state-of-the-art models by 3.7 and 3.1 points, respectively. 1 Leonardo F. R. Ribeiro, Claire Gardent, Iryna Gurevych |
Trans. Assoc. Comput. Linguistics | 1 |
| 2019 | Ranking Generated Summaries by Correctness: An Interesting but Challenging Application for Natural Language InferenceabstractWhile recent progress on abstractive summarization has led to remarkably fluent summaries, factual errors in generated summaries still severely limit their use in practice.In this paper, we evaluate summaries produced by state-of-the-art models via crowdsourcing and show that such errors occur frequently, in particular with more abstractive models.We study whether textual entailment predictions can be used to detect such errors and if they can be reduced by reranking alternative predicted summaries.That leads to an interesting downstream application for entailment models.In our experiments, we find that outof-the-box entailment models trained on NLI datasets do not yet offer the desired performance for the downstream task and we therefore release our annotations as additional test data for future extrinsic evaluations of NLI. Tobias Falke, Leonardo F. R. Ribeiro, Prasetya Ajie Utama, Ido Dagan, Iryna Gurevych |
ACL (1) | 2 |
| 2019 | Enhancing AMR-to-Text Generation with Dual Graph RepresentationsabstractLeonardo F. R. Ribeiro, Claire Gardent, Iryna Gurevych. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Leonardo F. R. Ribeiro, Claire Gardent, Iryna Gurevych |
EMNLP/IJCNLP (1) | 1 |
| 2017 | struc2vec: Learning Node Representations from Structural IdentityabstractStructural identity is a concept of symmetry in which network nodes are identified according to the network structure and their relationship to other nodes. Structural identity has been studied in theory and practice over the past decades, but only recently has it been addressed with representational learning techniques. This work presents struc2vec, a novel and flexible framework for learning latent representations for the structural identity of nodes. struc2vec uses a hierarchy to measure node similarity at different scales, and constructs a multilayer graph to encode structural similarities and generate structural context for nodes. Numerical experiments indicate that state-of-the-art techniques for learning node representations fail in capturing stronger notions of structural identity, while struc2vec exhibits much superior performance in this task, as it overcomes limitations of prior approaches. As a consequence, numerical experiments indicate that struc2vec improves performance on classification tasks that depend more on structural identity. Leonardo F. R. Ribeiro, Pedro H. P. Saverese, Daniel R. Figueiredo 0001 |
KDD | 1 |