EDBT 2026 Demo / reviewers in the wild / expert
Leonardo Ranaldi
dblp:278/7831
· DBLP profile ↗
14ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0001-8488-4146ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 10 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evolving AgentsabstractThe inability of current agents to autonomously generate abstractions limits their operation in dynamic environments to knowledge fixed during training.We propose EVA (Evolving Verifiable Agents), an architecture for autonomous learning that distills textual trajectories into pseudo-symbolic abstractions.EVA orchestrates world modeling, policy execution, and meta-control through three interacting components: a Perceptor, an Actor, and a Conductor.By mapping raw interaction data into semi-structured causal predicates, the system bridges the gap between free-form natural language and rigid logic, enabling the agent to construct an internal curriculum and perform verifiable self-correction during deployment.We evaluate EVA through a bi-level training scheme across three open-weight backbones on Dynamic ScienceWorld and Pseudo-Symbolic Logic Maze.Our results show that EVA significantly reduces logical-error rates and improves sample efficiency.Notably, the model recovers performance after mid-episode distribution shifts that cause baseline agents to collapse, demonstrating that dynamic abstraction enables rapid adaptation to unforeseen scenarios. Leonardo Ranaldi |
ACL (1) | 1 |
| 2026 | Agentic Oversight via Dialectic ReasoningabstractDebate has emerged as a promising oversight mechanism for Large Language Models (LLMs) amid rising systemic complexity, particularly where models outperform human evaluators.Yet, Debate provides little verifiable evidence for its final judgments, and its scalability remains largely unexplored.To make oversight grounded and scale as capabilities extend, we introduce an Agentic Oversight framework.By using Dialectic Argumentation as a reasoning function, we extend this paradigm to multilingual and multimodal spaces.We employ a weak-to-strong oversight approach based on two expert models that evaluate and defend contesting answers, while a third blind judge determines the winner using Dialectic Argumentation.Experts argue only for belief-consistent answers, founding the Debate on disagreements.We experimented with six tasks on our framework in both multilingual and multimodal scenarios, and dialectic argumentation consistently outperforms singleexpert baselines.Moreover, we show that dialectic judgements from a weaker model deliver argument-mediated supervision that, via finetuning, instils unsupervised reasoning signals in expert models.Q: What is dish called? Leonardo Ranaldi, Federico Ranaldi |
ACL (1) | 1 |
| 2026 | Thinking in Schemas: Robust Syllogistic Reasoning in LLMsabstractLLMs often mistake what sounds true for what is formally valid.This limitation is especially evident in syllogistic reasoning, where plausible arguments can lead models to endorse conclusions that are logically invalid, a phenomenon known as content effect (CE).We present Boethius, a schema-guided framework for syllogistic reasoning that disentangles semantic plausibility from logical validity.Boethius adopts an auditable, quasi-formal reasoning process with two complementary stages: a Schema Module, which deduces the underlying logical form by analysing the formal structure of the premises, and an Instantiation Module, which instantiates this form over the concrete argument and evaluates validity independently of content-level semantics.Our results show that Boethius consistently outperforms existing approaches, improving syllogistic reasoning accuracy while substantially reducing CE.These gains hold for both large models in a pure in-context learning setting and smaller models trained via schema-guided trajectories using supervised fine-tuning and optimisation-based refinement."Some winged animals are not chickadees.It is certain that all chickadees are birds.Therefore, some winged animals are not birds." SCHEMA MODULE INSTANTIATION MODULE PREMISES IDENTIFICATION:P1: Some winged animals are not chickadees.P2: It is certain that all chickadees are birds. Federico Ranaldi, Leonardo Ranaldi, Fabio Massimo Zanzotto, Shay B. Cohen |
ACL (1) | 2 |
| 2025 | R2-MultiOmnia: Leading Multilingual Multimodal Reasoning via Self-TrainingabstractReasoning is an intricate process that transcends both language and vision; because of its inherently modality-agnostic nature, developing effective multilingual and multimodal reasoning capabilities is a substantial challenge for Multimodal Large Language Models (MLLMs). They struggle to activate complex reasoning behaviours, delivering step-wise explanation, questioning and reflection, particularly in multilingual settings where high-quality supervision across languages is lacking. Recent works have introduced eclectic strategies to enhance MLLMs' reasoning; however, they remain related to a single language. To make MLLMs' reasoning capabilities aligned among languages and improve modality performances, we propose R2-MultiOmnia, a modular approach that instructs the models to abstract key elements of the reasoning process and then refine reasoning trajectories via self-correction. Specifically, we instruct the models producing multimodal synthetic demonstrations by bridging modalities and then self-improving their capabilities. To stabilise learning and the reasoning processes structure, we propose Curriculum Learning Reasoning Stabilisation with structured output rewards to gradually refine the models' capabilities to learn and deliver robust reasoning processes. Experiments show that R2-MultiOmnia improves multimodal reasoning, gets aligned performances among the languages approaching strong models. Leonardo Ranaldi, Federico Ranaldi, Giulia Pucci |
ACL (1) | 1 |
| 2025 | Improving Chain-of-Thought Reasoning via Quasi-Symbolic AbstractionsabstractChain-of-Thought (CoT) represents a common strategy for reasoning in Large Language Models (LLMs) by decomposing complex tasks into intermediate inference steps.However, explanations generated via CoT are susceptible to content biases that negatively affect their robustness and faithfulness.To mitigate existing limitations, recent work has proposed the use of logical formalisms coupled with external symbolic solvers.However, fully symbolically formalised approaches introduce the bottleneck of requiring a complete translation from natural language to formal languages, a process that affects efficiency and flexibility.To achieve a trade-off, this paper investigates methods to disentangle content from logical reasoning without a complete formalisation.In particular, we present QuaSAR (for Quasi-Symbolic Abstract Reasoning), a variation of CoT that guides LLMs to operate at a higher level of abstraction via quasi-symbolic explanations.Our framework leverages the capability of LLMs to formalise only relevant variables and predicates, enabling the coexistence of symbolic elements with natural language.We show the impact of QuaSAR for in-context learning and for constructing demonstrations to improve the reasoning capabilities of smaller models.Our experiments show that quasi-symbolic abstractions can improve CoTbased methods by up to 8% accuracy, enhancing robustness and consistency on challenging adversarial variations on both natural language (i.e.MMLU-Redux) and symbolic reasoning tasks (i.e., GSM-Symbolic). Leonardo Ranaldi, Marco Valentino, André Freitas |
ACL (1) | 1 |
| 2025 | Advancing Oversight Reasoning across Languages for Audit Sycophantic Behaviour via X-AgentabstractLarge language models (LLMs) have demonstrated capabilities that are highly satisfactory to a wide range of users by adapting to their culture and wisdom.Yet, this can translate into a propensity to produce responses that align with users' viewpoints, even when the latter are wrong.This behaviour is known as sycophancy, the tendency of LLMs to generate misleading responses as long as they align with the user's, inducing bias and reducing reliability.To make interactions consistent, reliable and safe, we introduce X-Agent, an Oversight Reasoning framework that audits human-model dialogues, reasons about them, captures sycophancy and corrects the final outputs.Concretely, X-Agent extends debate-based frameworks by (i) auditing user-model conversations, (ii) applying a defence layer that steers model behaviour and goes beyond user beliefs, and (iii) extracting reasoning traces from evaluations that serve as training signals for mitigating sycophancy, all in a completely unsupervised way.We evaluate X-Agent across diverse scenarios and languages, showing that it consistently detects sycophancy, reduces unwarranted agreement, and improves crossturn consistency, advancing a reasoning-asoverview paradigm for safer user-model and model-model interaction.Our approach introduces a novel paradigm in which reasoning is not merely a means to solve problems, but as a mechanism for overseeing the problem-solving processes of other models. Giulia Pucci, Leonardo Ranaldi |
EMNLP | 2 |
| 2025 | Improving Multilingual Retrieval-Augmented Language Models through Dialectic Reasoning ArgumentationsabstractRetrieval-augmented generation (RAG) is key to improving large language models (LLMs) in systematically accessing richer factual knowledge.Yet, using RAG mechanisms brings intrinsic challenges, as LLMs must deal with conflicting knowledge, especially in multilingual retrieval, where the heterogeneity of knowledge retrieved may deliver different outlooks.To make RAG more analytical, critical and grounded, we introduce Dialectic-RAG (D-RAG), a modular approach guided by Argumentative Explanations, i.e., structured reasoning process that systematically evaluates retrieved information by comparing, contrasting, and resolving conflicting perspectives.Given a query and a set of multilingual related documents, D-RAG selects and exemplifies relevant knowledge for delivering dialectic explanations that, by critically weighing opposing arguments and filtering extraneous content, clearly determine the final response.We show the impact of our framework both as an incontext learning strategy and for constructing demonstrations to instruct smaller models.Our experiments demonstrate that D-RAG significantly improves RAG approaches, requiring low-impact computational effort and providing robustness to knowledge perturbations.英格兰有多少位女王[1] [英国曾有 13 位君主... Retrieval Augmented LLM Baseline Dialectic RAG -[3] states that eight queens ruled the Kingdom of England, naming Matilda, Lady Jane Grey, Mary I, Elizabeth I, Mary II, Anne, Victoria, and Elizabeth II.-[4] focuses specifically on Queen Anne, who ruled from 1702 and became the first monarch of the UK in 1707 following the union of England and Scotland.-[1] outlines the formation of the UK and mention 13 monarchs.-[5] mentions Elizabeth II and her rule over the UK but does not address the total number of queens in England.-Doc [3]: Claims there were eight queens in England's history.(Relevant): -Doc [4]: Discusses Queen Anne's reign and her role as the monarch when England and Scotland unified in 1707.(Partially Relevant) -Doc [1]: Explains the timeline of the United Kingdom's formation but does not reference the specific number of queens.(Partially Relevant) -Doc [5]: Mentions Elizabeth II's reign over the United Kingdom but focuses on her role in the Commonwealth.(Irrelevant) [3] In less than 200 years, eight queens ruled the Kingdom of England. Leonardo Ranaldi, Federico Ranaldi, Fabio Massimo Zanzotto, Barry Haddow, Alexandra Birch |
EMNLP | 1 |
| 2025 | Multilingual Reasoning via Self-trainingabstractLeonardo Ranaldi, Giulia Pucci. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Leonardo Ranaldi, Giulia Pucci |
NAACL (Long Papers) | 1 |
| 2025 | Eliciting Critical Reasoning in Retrieval-Augmented Generation via Contrastive ExplanationsabstractLeonardo Ranaldi, Marco Valentino, Andre Freitas. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Leonardo Ranaldi, Marco Valentino, André Freitas |
NAACL (Long Papers) | 1 |
| 2024 | Does the Language Matter? Curriculum Learning over Neo-Latin Languages
Giulia Pucci, Leonardo Ranaldi |
LREC/COLING | 2 |
| 2024 | Aligning Large and Small Language Models via Chain-of-Thought ReasoningabstractChain-of-Thought (CoT) prompting empowers the reasoning abilities of Large Language Models (LLMs), eliciting them to solve complex reasoning tasks in a step-wise manner.However, these abilities appear only in models with billions of parameters, which represent an entry barrier for many users who are constrained to operate on a smaller model scale, i.e., Small Language Models (SLMs).Although many companies are releasing LLMs of the same family with fewer parameters, these models tend not to preserve all the reasoning capabilities of the original models, including CoT reasoning.In this paper, we propose a method for aligning and transferring reasoning abilities between larger to smaller Language Models.By using an Instruction-tuning-CoT method, an Instruction-tuning designed around CoT-Demonstrations, we enable the SLMs to generate multi-step controlled reasoned answers when elicited with the CoT mechanism.Hence, we instruct a smaller Language Model using outputs generated by more robust models belonging to the same family or not, evaluating the impact across different types of models.Results obtained on question-answering and mathematical reasoning benchmarks show that LMs instructed via the Instruction-tuning CoT method produced by LLMs outperform baselines within both in-domain and out-domain scenarios. Leonardo Ranaldi, André Freitas |
EACL (1) | 1 |
| 2024 | Self-Refine Instruction-Tuning for Aligning Reasoning in Language ModelsabstractThe alignment of reasoning abilities between smaller and larger Language Models are largely conducted via supervised fine-tuning using demonstrations generated from robust Large Language Models (LLMs).Although these approaches deliver more performant models, they do not show sufficiently strong generalization as the training only relies on the provided demonstrations.In this paper, we propose a self-refine Instruction-tuning method that allows for Smaller Language Models to self-improve their reasoning abilities.Our approach is based on a two-stage process, where reasoning abilities are first transferred between LLMs and Small Language Models (SLMs) via Instruction-tuning on synthetic demonstrations provided by LLMs, and then the instructed models self-improve through preference optimization strategies.In particular, the second phase operates refinement heuristics based on Direct Preference Optimization, where the SLMs are prompted to deliver a series of reasoning paths by automatically sampling the generated responses and providing rewards using ground truths from the LLMs.Results obtained on commonsense and math reasoning tasks show that this approach consistently outperforms Instructiontuning in both in-domain and out-domain scenarios, aligning the reasoning abilities of smaller and larger language models. Leonardo Ranaldi, André Freitas |
EMNLP | 1 |
| 2024 | Empowering Multi-step Reasoning across Languages via Program-Aided Language ModelsabstractIn-context learning methods are commonly employed as inference strategies, where Large Language Models (LLMs) are elicited to solve a task by leveraging provided demonstrations without requiring parameter updates.Among these approaches are the reasoning methods, exemplified by Chain-of-Thought (CoT) and Program-Aided Language Models (PAL), which encourage LLMs to generate reasoning steps, leading to improved accuracy.Despite their success, the ability to deliver multi-step reasoning remains limited to a single language, making it challenging to generalize to other languages and hindering global development.In this work, we propose Cross-lingual Program-Aided Language Models (Cross-PAL), a method for aligning reasoning programs across languages.Our method delivers programs as intermediate reasoning steps in different languages through a double-step cross-lingual prompting mechanism inspired by the Program-Aided approach.Moreover, we introduce Self-consistent Cross-PAL (SCross-PAL) to ensemble different reasoning paths across languages.Our experimental evaluations show that Cross-PAL outperforms existing methods, reducing the number of interactions and achieving state-of-the-art performance. Leonardo Ranaldi, Giulia Pucci, Barry Haddow, Alexandra Birch |
EMNLP | 1 |
| 2020 | KERMIT: Complementing Transformer Architectures with Encoders of Explicit Syntactic InterpretationsabstractFabio Massimo Zanzotto, Andrea Santilli, Leonardo Ranaldi, Dario Onorati, Pierfrancesco Tommasino, Francesca Fallucchi. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Fabio Massimo Zanzotto, Andrea Santilli, Leonardo Ranaldi, Dario Onorati, Pierfrancesco Tommasino, Francesca Fallucchi |
EMNLP (1) | 3 |