VLDB 2026 Research / reviewers in the wild / expert
Marco Valentino
dblp:212/3533
· DBLP profile ↗
25ranked-venue papers
6as first author
23since 2021 · last 2026
0000-0002-9959-8385ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 6 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating Content Effects on Reasoning in Language Models Through Fine-Grained Activation SteeringabstractLarge language models (LLMs) exhibit reasoning biases, often conflating content plausibility with formal logical validity. This can lead to wrong inferences in critical domains, where plausible arguments are incorrectly deemed logically valid or vice versa. This paper investigates how content biases on reasoning can be mitigated through activation steering, an inference-time technique that modulates internal activations. Specifically, after localising the layers responsible for formal and plausible inference, we investigate activation steering on a controlled syllogistic reasoning task, designed to disentangle formal validity from content plausibility. An extensive empirical analysis reveals that contrastive steering methods consistently support linear control over content biases. However, a static approach is insufficient to debias all the tested models. We then investigate how to control content effects by dynamically determining the steering parameters through fine-grained conditional methods. By introducing a novel kNN-based conditional approach (K-CAST), we demonstrate that conditional steering can effectively reduce biases on unresponsive models, achieving up to 15% absolute improvement in formal reasoning accuracy. Finally, we found that steering for content effects is robust to prompt variations, incurs minimal side effects on multilingual language modeling capabilities, and can partially generalize to different reasoning tasks. In practice, we demonstrate that activation-level interventions offer a scalable inference-time strategy for enhancing the robustness of LLMs, contributing towards more systematic and unbiased reasoning capabilities Marco Valentino, Geonhee Kim, Dhairya Dalal, Zhixue Zhao, André Freitas |
AAAI | 1 |
| 2026 | Learning to Disentangle Latent Reasoning Rules with Language VAEs: A Systematic Study
Yingji Zhang, Marco Valentino, Danilo S. Carvalho, André Freitas |
AAAI | 2 |
| 2025 | Controlling Equational Reasoning in Large Language Models with Prompt InterventionsabstractThis paper investigates how hallucination rates in Large Language Models (LLMs) may be controlled via a symbolic data generation framework, exploring a fundamental relationship between the rate of certain mathematical errors and types of input intervention. Specifically, we systematically generate data for a derivation generation task using a symbolic engine, applying targeted interventions to prompts to perturb features of mathematical derivations such as the surface forms of symbols, equational tree structures, and mathematical context. We then evaluate the effect of prompt interventions across a range of LLMs including fine-tuned T5 models, GPT, and LLaMa-based models. Our experiments suggest that T5-Large can outperform the few-shot performance of GPT-4 on various evaluation sets generated via the framework. However, an extensive evaluation based on human analysis, template-based error detection, and text generation metrics reveals model weaknesses beyond what the reference-based metrics singularly describe. We use these results to tie characteristic distributional footprints of interventions to the human evaluation of LLM derivation quality, potentially leading to significant control over fine-grained mathematical capabilities of language models with respect to specific types of errors. Jordan Meadows, Marco Valentino, André Freitas |
AAAI | 2 |
| 2025 | Unravelling the Logic: Investigating the Generalisation of Transformers in Numerical Satisfiability ProblemsabstractTharindu Madusanka, Marco Valentino, Iqra Zahid, Ian Pratt-Hartmann, Riza Batista-Navarro. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Tharindu Madusanka, Marco Valentino, Iqra Zahid, Ian Pratt-Hartmann, Riza Theresa Batista-Navarro |
ACL (1) | 2 |
| 2025 | Faithful and Robust LLM-Driven Theorem Proving for NLI ExplanationsabstractNatural language explanations play a fundamental role in Natural Language Inference (NLI) by revealing how premises logically entail hypotheses.Recent work has shown that the interaction of large language models (LLMs) with theorem provers (TPs) can help verify and improve the validity of NLI explanations.However, TPs require translating natural language into machine-verifiable formal representations, a process that introduces the risk of semantic information loss and unfaithful interpretation, an issue compounded by LLMs' challenges in capturing critical logical structures with sufficient precision.Moreover, LLMs are still limited in their capacity for rigorous and robust proof construction within formal verification frameworks.To mitigate issues related to faithfulness and robustness, this paper investigates strategies to (1) alleviate semantic loss during autoformalisation, (2) efficiently identify and correct syntactic errors in logical representations, (3) explicitly use logical expressions to guide LLMs in generating structured proof sketches, and (4) increase LLMs' capacity of interpreting TP's feedback for iterative refinement.Our empirical results on e-SNLI, QASC and WorldTree using different LLMs demonstrate that the proposed strategies yield significant improvements in autoformalisation (+18.46%,+34.2%, +39.77%) and explanation refinement (+29.5%,+51.5%, +41.25%) over the state-of-the-art model.Moreover, we show that specific interventions on the hybrid LLM-TP architecture can substantially improve efficiency, drastically reducing the number of iterations required for successful verification.1 Marco Valentino, Louise A. Dennis, André Freitas |
ACL (1) | 2 |
| 2025 | Improving Chain-of-Thought Reasoning via Quasi-Symbolic AbstractionsabstractChain-of-Thought (CoT) represents a common strategy for reasoning in Large Language Models (LLMs) by decomposing complex tasks into intermediate inference steps.However, explanations generated via CoT are susceptible to content biases that negatively affect their robustness and faithfulness.To mitigate existing limitations, recent work has proposed the use of logical formalisms coupled with external symbolic solvers.However, fully symbolically formalised approaches introduce the bottleneck of requiring a complete translation from natural language to formal languages, a process that affects efficiency and flexibility.To achieve a trade-off, this paper investigates methods to disentangle content from logical reasoning without a complete formalisation.In particular, we present QuaSAR (for Quasi-Symbolic Abstract Reasoning), a variation of CoT that guides LLMs to operate at a higher level of abstraction via quasi-symbolic explanations.Our framework leverages the capability of LLMs to formalise only relevant variables and predicates, enabling the coexistence of symbolic elements with natural language.We show the impact of QuaSAR for in-context learning and for constructing demonstrations to improve the reasoning capabilities of smaller models.Our experiments show that quasi-symbolic abstractions can improve CoTbased methods by up to 8% accuracy, enhancing robustness and consistency on challenging adversarial variations on both natural language (i.e.MMLU-Redux) and symbolic reasoning tasks (i.e., GSM-Symbolic). Leonardo Ranaldi, Marco Valentino, André Freitas |
ACL (1) | 2 |
| 2025 | Enhancing Logical Reasoning in Language Models via Symbolically-Guided Monte Carlo Process SupervisionabstractLarge language models (LLMs) have shown strong performance in many reasoning benchmarks. However, recent studies have pointed to memorization, rather than generalization, as one of the leading causes for such performance. LLMs, in fact, are susceptible to content variations, demonstrating a lack of robust planning or symbolic abstractions supporting their reasoning process. To improve reliability, many attempts have been made to combine LLMs with symbolic methods. Nevertheless, existing approaches fail to effectively leverage symbolic representations due to the challenges involved in developing reliable and scalable verification mechanisms. In this paper, we propose to overcome such limitations by synthesizing high-quality symbolic reasoning trajectories with stepwise pseudo-labels at scale via Monte Carlo estimation. A Process Reward Model (PRM) can be efficiently trained based on the synthesized data and then used to select more symbolic trajectories. The trajectories are then employed with Direct Preference Optimization (DPO) and Supervised Fine-Tuning (SFT) to improve logical reasoning and generalization. Our results on benchmarks (i.e., FOLIO and LogicAsker) show the effectiveness of the proposed method with gains on frontier and open-weight models. Moreover, additional experiments on claim verification data reveal that fine-tuning on the generated symbolic reasoning trajectories enhances out-of-domain generalizability, suggesting the potential impact of the proposed method in enhancing planning and logical reasoning. Xingwei Tan, Marco Valentino, Mahmud Elahi Akhter, Maria Liakata, Nikolaos Aletras |
EMNLP | 2 |
| 2025 | Autoformalization in the Wild: Assessing LLMs on Real-World Mathematical DefinitionsabstractThanks to their linguistic capabilities, LLMs offer an opportunity to bridge the gap between informal mathematics and formal languages through autoformalization.However, it is still unclear how well LLMs generalize to sophisticated and naturally occurring mathematical statements.To address this gap, we investigate the task of autoformalizing real-world mathematical definitions: a critical component of mathematical discourse.Specifically, we introduce two novel resources for autoformalization, collecting definitions from Wikipedia (Def_Wiki) and arXiv papers (Def_ArXiv).We then systematically evaluate a range of LLMs, analyzing their ability to formalize definitions into Isabelle/HOL.Furthermore, we investigate strategies to enhance LLMs' performance including refinement through external feedback from Proof Assistants, and formal definition grounding, where we augment LLMs' formalizations through relevant contextual elements from formal mathematical libraries.Our findings reveal that definitions present a greater challenge compared to existing benchmarks, such as miniF2F.In particular, we found that LLMs still struggle with self-correction, and aligning with relevant mathematical libraries.At the same time, structured refinement methods and definition grounding strategies yield notable improvements of up to 16% on selfcorrection capabilities and 43% on the reduction of undefined errors, highlighting promising directions for enhancing LLM-based autoformalization in real-world scenarios.1 Marco Valentino, André Freitas |
EMNLP | 2 |
| 2025 | Eliciting Critical Reasoning in Retrieval-Augmented Generation via Contrastive ExplanationsabstractLeonardo Ranaldi, Marco Valentino, Andre Freitas. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Leonardo Ranaldi, Marco Valentino, André Freitas |
NAACL (Long Papers) | 2 |
| 2025 | SylloBio-NLI: Evaluating Large Language Models on Biomedical Syllogistic ReasoningabstractMagdalena Wysocka, Danilo Carvalho, Oskar Wysocki, Marco Valentino, Andre Freitas. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Magdalena Wysocka, Danilo S. Carvalho, Oskar Wysocki, Marco Valentino, André Freitas |
NAACL (Long Papers) | 4 |
| 2024 | Inference to the Best Explanation in Large Language ModelsabstractWhile Large Language Models (LLMs) have found success in real-world applications, their underlying explanatory process is still poorly understood.This paper proposes IBE-Eval, a framework inspired by philosophical accounts on Inference to the Best Explanation (IBE) to advance the interpretation and evaluation of LLM explanations.IBE-Eval estimates the plausibility of natural language explanations through a combination of explicit logical and linguistic features including: consistency, parsimony, coherence, and uncertainty.Extensive experiments are conducted on Causal Question Answering (CQA), where IBE-Eval is tasked to select the most plausible causal explanation amongst competing ones generated by the LLM (e.g.GPT 3.5 or LLaMA 2).The experiments reveal that IBE-Eval can successfully identify the best explanation with up to 77% accuracy (≈ 27% above random), improving upon a GPT 3.5-as-a-judge baseline (≈ +17%) while being intrinsically more efficient and interpretable.Additional analysis suggests that, despite LLM-specific variances, generated explanations tend to conform to IBE criteria and that IBE-Eval is significantly correlated with human judgment, opening up opportunities for future development of automated explanation verification tools. Inference to the Best Explanation (IBE) Dhairya Dalal, Marco Valentino, André Freitas, Paul Buitelaar |
ACL (1) | 2 |
| 2024 | Estimating the Causal Effects of Natural Logic Features in Transformer-Based NLI ModelsabstractRigorous evaluation of the causal effects of semantic features on language model predictions can be hard to achieve for natural language reasoning problems. However, this is such a desirable form of analysis from both an interpretability and model evaluation perspective, that it is valuable to investigate specific patterns of reasoning with enough structure and regularity to identify and quantify systematic reasoning failures in widely-used models. In this vein, we pick a portion of the NLI task for which an explicit causal diagram can be systematically constructed: the case where across two sentences (the premise and hypothesis), two related words/terms occur in a shared context. In this work, we apply causal effect estimation strategies to measure the effect of context interventions (whose effect on the entailment label is mediated by the semantic monotonicity characteristic) and interventions on the inserted word-pair (whose effect on the entailment label is mediated by the relation between these words). Extending related work on causal analysis of NLP models in different settings, we perform an extensive interventional study on the NLI task to investigate robustness to irrelevant changes and sensitivity to impactful changes of Transformers. The results strongly bolster the fact that similar benchmark accuracy scores may be observed for models that exhibit very different behaviour. Moreover, our methodology reinforces previously suspected biases from a causal perspective, including biases in favour of upward-monotone contexts and ignoring the effects of negation markers. Julia Rozanova, Marco Valentino, André Freitas |
LREC/COLING | 2 |
| 2024 | A Differentiable Integer Linear Programming Solver for Explanation-Based Natural Language InferenceabstractInteger Linear Programming (ILP) has been proposed as a formalism for encoding precise structural and semantic constraints for Natural Language Inference (NLI). However, traditional ILP frameworks are non-differentiable, posing critical challenges for the integration of continuous language representations based on deep learning. In this paper, we introduce a novel approach, named Diff-Comb Explainer, a neuro-symbolic architecture for explanation-based NLI based on Differentiable BlackBox Combinatorial Solvers (DBCS). Differently from existing neuro-symbolic solvers, Diff-Comb Explainer does not necessitate a continuous relaxation of the semantic constraints, enabling a direct, more precise, and efficient incorporation of neural representations into the ILP formulation. Our experiments demonstrate that Diff-Comb Explainer achieves superior performance when compared to conventional ILP solvers, neuro-symbolic black-box solvers, and Transformer-based encoders. Moreover, a deeper analysis reveals that Diff-Comb Explainer can significantly improve the precision, consistency, and faithfulness of the constructed explanations, opening new opportunities for research on neuro-symbolic architectures for explainable and transparent NLI in complex domains. Mokanarangan Thayaparan, Marco Valentino, André Freitas |
LREC/COLING | 2 |
| 2024 | Enhancing Ethical Explanations of Large Language Models through Iterative Symbolic RefinementabstractAn increasing amount of research in Natural Language Inference (NLI) focuses on the application and evaluation of Large Language Models (LLMs) and their reasoning capabilities.Despite their success, however, LLMs are still prone to factual errors and inconsistencies in their explanations, offering limited control and interpretability for inference in complex domains.In this paper, we focus on ethical NLI, investigating how hybrid neurosymbolic techniques can enhance the logical validity and alignment of ethical explanations produced by LLMs.Specifically, we present an abductive-deductive framework named Logic-Explainer, which integrates LLMs with an external backward-chaining solver to refine step-wise natural language explanations and jointly verify their correctness, reduce incompleteness and minimise redundancy.An extensive empirical analysis demonstrates that Logic-Explainer can improve explanations generated via in-context learning methods and Chainof-Thought (CoT) on challenging ethical NLI tasks, while, at the same time, producing formal proofs describing and supporting models' reasoning.As ethical NLI requires commonsense reasoning to identify underlying moral violations, our results suggest the effectiveness of neuro-symbolic methods for multi-step NLI more broadly, opening new opportunities to enhance the logical consistency, reliability, and alignment of LLMs. Marco Valentino, Louise A. Dennis, André Freitas |
EACL (1) | 2 |
| 2024 | Multi-Relational Hyperbolic Word Embeddings from Natural Language DefinitionsabstractNatural language definitions possess a recursive, self-explanatory semantic structure that can support representation learning methods able to preserve explicit conceptual relations and constraints in the latent space.This paper presents a multi-relational model that explicitly leverages such a structure to derive word embeddings from definitions.By automatically extracting the relations linking defined and defining terms from dictionaries, we demonstrate how the problem of learning word embeddings can be formalised via a translational framework in Hyperbolic space and used as a proxy to capture the global semantic structure of definitions.An extensive empirical analysis demonstrates that the framework can help imposing the desired structural constraints while preserving the semantic mapping required for controllable and interpretable traversal.Moreover, the experiments reveal the superiority of the Hyperbolic word embeddings over the Euclidean counterparts and demonstrate that the multirelational approach can obtain competitive results when compared to state-of-the-art neural models, with the advantage of being intrinsically more efficient and interpretable 1 . Marco Valentino, Danilo S. Carvalho, André Freitas |
EACL (1) | 1 |
| 2024 | Verification and Refinement of Natural Language Explanations through LLM-Symbolic Theorem ProvingabstractNatural language explanations represent a proxy for evaluating explanation-based and multi-step Natural Language Inference (NLI) models.However, assessing the validity of explanations for NLI is challenging as it typically involves the crowd-sourcing of apposite datasets, a process that is time-consuming and prone to logical errors.To address existing limitations, this paper investigates the verification and refinement of natural language explanations through the integration of Large Language Models (LLMs) and Theorem Provers (TPs).Specifically, we present a neuro-symbolic framework, named Explanation-Refiner, that integrates TPs with LLMs to generate and formalise explanatory sentences and suggest potential inference strategies for NLI.In turn, the TP is employed to provide formal guarantees on the logical validity of the explanations and to generate feedback for subsequent improvements.We demonstrate how Explanation-Refiner can be jointly used to evaluate explanatory reasoning, autoformalisation, and error correction mechanisms of state-of-the-art LLMs as well as to automatically enhance the quality of explanations of variable complexity in different domains. 1 Marco Valentino, Louise A. Dennis, André Freitas |
EMNLP | 2 |
| 2024 | A Symbolic Framework for Evaluating Mathematical Reasoning and Generalisation with TransformersabstractJordan Meadows, Marco Valentino, Damien Teney, Andre Freitas. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jordan Meadows, Marco Valentino, Damien Teney, André Freitas |
NAACL-HLT | 2 |
| 2024 | Multi-Operational Mathematical Derivations in Latent SpaceabstractMarco Valentino, Jordan Meadows, Lan Zhang, Andre Freitas. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Marco Valentino, Jordan Meadows, André Freitas |
NAACL-HLT | 1 |
| 2023 | NLI4CT: Multi-Evidence Natural Language Inference for Clinical Trial ReportsabstractHow can we interpret and retrieve medical evidence to support clinical decisions?Clinical trial reports (CTR) amassed over the years contain indispensable information for the development of personalized medicine.However, it is practically infeasible to manually inspect over 400,000+ clinical trial reports in order to find the best evidence for experimental treatments.Natural Language Inference (NLI) offers a potential solution to this problem, by allowing the scalable computation of textual entailment.However, existing NLI models perform poorly on biomedical corpora, and previously published datasets fail to capture the full complexity of inference over CTRs.In this work, we present a novel resource to advance research on NLI for reasoning on CTRs.The resource includes two main tasks.Firstly, to determine the inference relation between a natural language statement, and a CTR.Secondly, to retrieve supporting facts to justify the predicted relation.We provide NLI4CT, a corpus of 2400 statements and CTRs, annotated for these tasks.Baselines on this corpus expose the limitations of existing NLI approaches, with 6 state-of-the-art NLI models achieving a maximum F1 score of 0.627.To the best of our knowledge, we are the first to design a task that covers the interpretation of full CTRs.To encourage further work on this challenging dataset, we make the corpus, competition leaderboard, and website, available on CodaLab, and code to replicate the baseline experiments on GitHub 1 . Maël Jullien, Marco Valentino, Hannah Frost, Paul O'Regan, Donal Landers, André Freitas |
EMNLP | 2 |
| 2022 | Hybrid Autoregressive Inference for Scalable Multi-Hop Explanation RegenerationabstractRegenerating natural language explanations in the scientific domain has been proposed as a benchmark to evaluate complex multi-hop and explainable inference. In this context, large language models can achieve state-of-the-art performance when employed as cross-encoder architectures and fine-tuned on human-annotated explanations. However, while much attention has been devoted to the quality of the explanations, the problem of performing inference efficiently is largely under studied. Cross-encoders, in fact, are intrinsically not scalable, possessing limited applicability to real-world scenarios that require inference on massive facts banks. To enable complex multi-hop reasoning at scale, this paper focuses on bi-encoder architectures, investigating the problem of scientific explanation regeneration at the intersection of dense and sparse models. Specifically, we present SCAR (for Scalable Autoregressive Inference), a hybrid framework that iteratively combines a Transformer-based bi-encoder with a sparse model of explanatory power, designed to leverage explicit inference patterns in the explanations. Our experiments demonstrate that the hybrid framework significantly outperforms previous sparse models, achieving performance comparable with that of state-of-the-art cross-encoders while being approx 50 times faster and scalable to corpora of millions of facts. Further analyses on semantic drift and multi-hop question answering reveal that the proposed hybridisation boosts the quality of the most challenging explanations, contributing to improved performance on downstream inference tasks. Marco Valentino, Mokanarangan Thayaparan, Deborah Ferreira, André Freitas |
AAAI | 1 |
| 2022 | Case-Based Abductive Natural Language InferenceabstractMost of the contemporary approaches for multi-hop Natural Language Inference (NLI) construct explanations considering each test case in isolation. However, this paradigm is known to suffer from semantic drift, a phenomenon that causes the construction of spurious explanations leading to wrong conclusions. In contrast, this paper proposes an abductive framework for multi-hop NLI exploring the retrieve-reuse-refine paradigm in Case-Based Reasoning (CBR). Specifically, we present Case-Based Abductive Natural Language Inference (CB-ANLI), a model that addresses unseen inference problems by analogical transfer of prior explanations from similar examples. We empirically evaluate the abductive framework on commonsense and scientific question answering tasks, demonstrating that CB-ANLI can be effectively integrated with sparse and dense pre-trained encoders to improve multi-hop inference, or adopted as an evidence retriever for Transformers. Moreover, an empirical analysis of semantic drift reveals that the CBR paradigm boosts the quality of the most challenging explanations, a feature that has a direct impact on robustness and accuracy in downstream inference tasks. Marco Valentino, Mokanarangan Thayaparan, André Freitas |
COLING | 1 |
| 2022 | Diff-Explainer: Differentiable Convex Optimization for Explainable Multi-hop InferenceabstractAbstract This paper presents Diff-Explainer, the first hybrid framework for explainable multi-hop inference that integrates explicit constraints with neural architectures through differentiable convex optimization. Specifically, Diff- Explainer allows for the fine-tuning of neural representations within a constrained optimization framework to answer and explain multi-hop questions in natural language. To demonstrate the efficacy of the hybrid framework, we combine existing ILP-based solvers for multi-hop Question Answering (QA) with Transformer-based representations. An extensive empirical evaluation on scientific and commonsense QA tasks demonstrates that the integration of explicit constraints in a end-to-end differentiable framework can significantly improve the performance of non- differentiable ILP solvers (8.91%–13.3%). Moreover, additional analysis reveals that Diff-Explainer is able to achieve strong performance when compared to standalone Transformers and previous multi-hop approaches while still providing structured explanations in support of its predictions. Mokanarangan Thayaparan, Marco Valentino, Deborah Ferreira, Julia Rozanova, André Freitas |
Trans. Assoc. Comput. Linguistics | 2 |
| 2021 | Unification-based Reconstruction of Multi-hop Explanations for Science QuestionsabstractThis paper presents a novel framework for reconstructing multi-hop explanations in science Question Answering (QA).While existing approaches for multi-hop reasoning build explanations considering each question in isolation, we propose a method to leverage explanatory patterns emerging in a corpus of scientific explanations.Specifically, the framework ranks a set of atomic facts by integrating lexical relevance with the notion of unification power, estimated analysing explanations for similar questions in the corpus.An extensive evaluation is performed on the Worldtree corpus, integrating k-NN clustering and Information Retrieval (IR) techniques.We present the following conclusions: (1) The proposed method achieves results competitive with Transformers, yet being orders of magnitude faster, a feature that makes it scalable to large explanatory corpora (2) The unificationbased mechanism has a key role in reducing semantic drift, contributing to the reconstruction of many hops explanations (6 or more facts) and the ranking of complex inference facts (+12.0Mean Average Precision) (3) Crucially, the constructed explanations can support downstream QA models, improving the accuracy of BERT by up to 10% overall. Marco Valentino, Mokanarangan Thayaparan, André Freitas |
EACL | 1 |
| 2020 | A Framework for Evaluation of Machine Reading Comprehension Gold StandardsabstractMachine Reading Comprehension (MRC) is the task of answering a question over a paragraph of text. While neural MRC systems gain popularity and achieve noticeable performance, issues are being raised with the methodology used to establish their performance, particularly concerning the data design of gold standards that are used to evaluate them. There is but a limited understanding of the challenges present in this data, which makes it hard to draw comparisons and formulate reliable hypotheses. As a first step towards alleviating the problem, this paper proposes a unifying framework to systematically investigate the present linguistic features, required reasoning and background knowledge and factual correctness on one hand, and the presence of lexical cues as a lower bound for the requirement of understanding on the other hand. We propose a qualitative annotation schema for the first and a set of approximative metrics for the latter. In a first application of the framework, we analyse modern MRC gold standards and present our findings: the absence of features that contribute towards lexical ambiguity, the varying factual correctness of the expected answers and the presence of lexical cues, all of which potentially lower the reading comprehension complexity and quality of the evaluation data. Viktor Schlegel, Marco Valentino, André Freitas, Goran Nenadic, Riza Theresa Batista-Navarro |
LREC | 2 |
| 2018 | Adaptive Workflows of Home-Care ServicesabstractWith the increased number of elderly people in developed countries, assistive robotics is gaining more attention allowing to support home care assistance. Here, assistive robotics is adopted to monitor the activities of daily living (ADL) of patients with mild neurological disorders to limit the human monitoring, usually representing a burden for family members. In order to improve the effectiveness and user acceptance level of the robotic system, a middleware layer, able to automatically generate monitoring plans for home care patients, is proposed. The plans are generated as workflow of services, each one representing a monitoring task that can be executed by different devices, including humans, in different ways. We show that a service-oriented approach allows generating adaptive monitoring plans for patients with different levels of neurological disorders, taking into account the dynamic nature of their personality profiles, as well as of the environment they live in. Claudia Di Napoli, Marco Valentino, Luca Sabatucci, Massimo Cossentino |
WETICE | 2 |