Pasquale Minervini

dblp:58/10142 · DBLP profile ↗
← Back
62ranked-venue papers
13as first author
45since 2021 · last 2026
0000-0002-8442-602XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 54 · 9 first-author · 42 since 2021Databases, data management, data science and information retrieval · 9 · 6 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learning and Reasoning on Knowledge and Heterogeneous Graphs in the era of Graph Foundation and Large Language Models
abstract
Knowledge Graphs (KGs) and heterogeneous graphs (HGs) offer a principled way to represent multi-entity, multi-relational systems, while also revealing a persistent tension between expressive modeling, scalable learning, and faithful reasoning.Two trends are rapidly reshaping the field: graph foundation models (GFMs), which seek transfer across graphs, tasks, and domains via large-scale pretraining, and the growing integration of large language models (LLMs) with graph-structured knowledge to improve grounding, interaction, and reasoning.Temporal settings add further challenges, as evolving facts and interactions demand time-consistent modeling and evaluation.This tutorial provides a structured survey of these directions: we introduce a unified background and notation for typed heterogeneous graphs, (temporal) KGs, and event-based temporal heterogeneous graphs; we then formalize the main task families (KG completion, query answering, node/graph prediction, and temporal variants), emphasizing evaluation protocols and leakage pitfalls.Finally, we review recent advances in GFMs and LLM-graph integration, and summarize the state of the art in learning over temporal heterogeneous graphs and temporal KGs.
Matteo Zignani, Pasquale Minervini, Roberto Interdonato, Manuel Dileo
ESANN2
2026 Low-Rank Compression of Language Models via Differentiable Rank Selection
abstract
Approaches for compressing large-language models using low-rank decomposition have made strides, particularly with the introduction of activation and loss-aware SVD, which improves the trade-off between decomposition rank and downstream task performance. Despite these advancements, a persistent challenge remains--selecting the optimal ranks for each layer to jointly optimise compression rate and downstream task accuracy. Current methods either rely on heuristics that can yield sub-optimal results due to their limited discrete search space or are gradient-based but are not as performant as heuristic approaches without post-compression fine-tuning. To address these issues, we propose Learning to Low-Rank Compress (LLRC), a gradient-based approach which directly learns the weights of masks that select singular values in a fine-tuning-free setting. Using a calibration dataset, we train only the mask weights to select fewer and fewer singular values while minimising the divergence of intermediate activations from the original model. Our approach outperforms competing ranking selection methods that similarly require no post-compression fine-tuning across various compression rates on common-sense reasoning and open-domain question-answering tasks. For instance, with a compression rate of 20% on Llama-2-13B, LLRC outperforms the competitive Sensitivity-based Truncation Rank Searching (STRS) on MMLU, BoolQ, and OpenbookQA by 12%, 3.5%, and 4.4%, respectively. Compared to other compression techniques, our approach consistently outperforms fine-tuning-free variants of SVD-LLM and LLM-Pruner across datasets and compression rates. Our fine-tuning-free approach also performs competitively with the fine-tuning variant of LLM-Pruner.
Sidhant Sundrani, Francesco Tudisco, Pasquale Minervini
LREC3
2026 Tensor factorization for temporal knowledge graph forecasting
Manuel Dileo, Pasquale Minervini, Matteo Zignani, Sabrina Gaito
Neurocomputing2
2025 Adaptive Computation Modules: Granular Conditional Computation for Efficient Inference
abstract
While transformer models have been highly successful, they are computationally inefficient. We observe that for each layer, the full width of the layer may be needed only for a small subset of tokens inside a batch and that the "effective" width needed to process a token can vary from layer to layer. Motivated by this observation, we introduce the Adaptive Computation Module (ACM), a generic module that dynamically adapts its computational load to match the estimated difficulty of the input on a per-token basis. An ACM consists of a sequence of learners that progressively refine the output of their preceding counterparts. An additional gating mechanism determines the optimal number of learners to execute for each token. We also propose a distillation technique to replace any pre-trained model with an "ACMized" variant. Our evaluation of transformer models in computer vision and speech recognition demonstrates that substituting layers with ACMs significantly reduces inference costs without degrading the downstream accuracy for a wide interval of user-defined budgets.
Bartosz Wójcik, Alessio Devoto, Karol Pustelnik, Pasquale Minervini, Simone Scardapane
AAAI4
2025 Mixtures of In-Context Learners
abstract
In-context learning (ICL) adapts LLMs by providing demonstrations without fine-tuning the model parameters; however, it is very sensitive to the choice of in-context demonstrations, and processing many demonstrations can be computationally demanding.We propose Mixtures of In-Context Learners (MOICL), a novel approach that uses subsets of demonstrations to train a set of experts via ICL and learns a weighting function to merge their output distributions via gradient-based optimisation.In our experiments, we show performance improvements on 5 out of 7 classification datasets compared to a set of strong baselines (e.g., up to +13% compared to ICL and LENS).Moreover, we improve the Pareto frontier of ICL by reducing the inference time needed to achieve the same performance with fewer demonstrations.Finally, MOICL is more robust to out-ofdomain (up to +11%), imbalanced (up to +49%) and perturbed demonstrations (up to +38%). 1
Giwon Hong, Emile van Krieken, Edoardo Maria Ponti, Nikolay Malkin, Pasquale Minervini
ACL (1)5
2025 SmaLLEXT: 1st Workshop on Small and Efficient Large Language Models for Knowledge Extraction
Felice Antonio Merra, Kristian Skracic, Daniele Malitesta, Jacek Golebiowski, Pasquale Minervini
CIKM5
2025 SynDARin: Synthesising Datasets for Automated Reasoning in Low-Resource Languages
abstract
Question Answering (QA) datasets have been instrumental in developing and evaluating Large Language Model (LLM) capabilities. However, such datasets are scarce for languages other than English due to the cost and difficulties of collection and manual annotation. This means that producing novel models and measuring the performance of multilingual LLMs in low-resource languages is challenging. To mitigate this, we propose SynDARin, a method for generating and validating QA datasets for low-resoucre languages. We utilize parallel content mining to obtain human-curated paragraphs between English and the target language. We use the English data as context to generate synthetic multiple-choice (MC) question-answer pairs, which are automatically translated and further validated for quality. Combining these with their designated non-English human-curated paragraphs form the final QA dataset. The method allows to maintain content quality, reduces the likelihood of factual errors, and circumvents the need for costly annotation. To test the method, we created a QA dataset with 1.2K samples for the Armenian language. The human evaluation shows that 98% of the generated English data maintains quality and diversity in the question types and topics, while the translation validation pipeline can filter out ~70% of data with poor quality. We use the dataset to benchmark state-of-the-art LLMs, showing their inability to achieve human accuracy with some model performances closer to random chance. This shows that the generated dataset is non-trivial and can be used to evaluate reasoning capabilities in low-resource language.
Gayane Ghazaryan, Erik Arakelyan, Isabelle Augenstein, Pasquale Minervini
COLING4
2025 FLARE: Faithful Logic-Aided Reasoning and Exploration
abstract
Modern Question Answering (QA) and Reasoning approaches with Large Language Models (LLMs) commonly use Chain-of-Thought (CoT) prompting but struggle with generating outputs faithful to their intermediate reasoning chains.While neuro-symbolic methods like Faithful CoT (F-CoT) offer higher faithfulness through external solvers, they require codespecialized models and struggle with ambiguous tasks.We introduce Faithful Logic-Aided Reasoning and Exploration (FLARE), which uses LLMs to plan solutions, formalize queries into logic programs, and simulate code execution through multi-hop search without external solvers.Our method achieves SOTA results on 7 out of 9 diverse reasoning benchmarks and 3 out of 3 logic inference benchmarks while enabling measurement of reasoning faithfulness.We demonstrate that model faithfulness correlates with performance and that successful reasoning traces show an 18.1% increase in unique emergent facts, 8.6% higher overlap between code-defined and execution-trace relations, and 3.6% reduction in unused relations.
Erik Arakelyan, Pasquale Minervini, Patrick S. H. Lewis, Patrick Verga, Isabelle Augenstein
EMNLP2
2025 GRADA: Graph-based Reranking against Adversarial Documents Attack
abstract
Retrieval Augmented Generation (RAG) frameworks can improve the factual accuracy of large language models (LLMs) by integrating external knowledge from retrieved documents, which is useful for overcoming the limitations of models' static intrinsic knowledge.However, these systems are susceptible to adversarial attacks that manipulate the retrieval process by introducing documents that are adversarial yet semantically similar to the query.Notably, while these adversarial documents resemble the query, they exhibit weak similarity to benign documents in the retrieval set.Thus, we propose a simple yet effective Graph-based Reranking against Adversarial Document Attacks (GRADA) framework aimed at preserving retrieval quality while significantly reducing the success of adversaries.Our study evaluates the effectiveness of our approach through experiments conducted on six LLMs: GPT-3.5-Turbo,GPT-4o, Llama3.1-8b-Instruct,Llama3.1-70b-Instruct,Qwen2.5-7b-Instruct, and Qwen2.5-14b-Instruct.We use three datasets to assess performance, with results from the Natural Questions dataset showing up to an 80% reduction in attack success rates while maintaining minimal loss in accuracy.
Jingjie Zheng, Aryo Pradipta Gema, Giwon Hong, Xuanli He, Pasquale Minervini, Youcheng Sun, Qiongkai Xu
EMNLP5
2025 Enhancing neural link predictors for temporal knowledge graphs with temporal regularisers
abstract
The problem of link prediction in temporal knowledge graphs (TKGs) consists of finding missing links in the knowledge base under temporal constraints.Recently, [4] and [8] proposed a solution to the problem inspired by the canonical decomposition of 4-order tensors, where they regularise the representations of time steps by learning similar transformation for adjacent timestamps.However, the impact of the choice of temporal regularisation terms is still poorly understood.In this work, we systematically analyse several choices of temporal regularisers using linear functions and recurrent architectures.In our experiments, we show that by carefully selecting the temporal regulariser and regularisation weight, a simple method like TNTComplEx [4] can produce comparable results with state-of-the-art methods and enhance its original performance.Specifically, we observe that linear regularisers for temporal smoothing based on specific nuclear norms can significantly improve the predictive accuracy of the base temporal link prediction methods.
Manuel Dileo, Pasquale Minervini, Matteo Zignani, Sabrina Gaito
ESANN2
2025 An Auditing Test to Detect Behavioral Shift in Language Models
abstract
As language models (LMs) approach human-level performance, a comprehensive understanding of their behavior becomes crucial. This includes evaluating capabilities, biases, task performance, and alignment with societal values. Extensive initial evaluations, including red teaming and diverse benchmarking, can establish a model’s behavioral profile. However, subsequent fine-tuning or deployment modifications may alter these behaviors in unintended ways. We present an efficient statistical test to tackle Behavioral Shift Auditing (BSA) in LMs, which we define as detecting distribution shifts in qualitative properties of the output distributions of LMs. Our test compares model generations from a baseline model to those of the model under scrutiny and provides theoretical guarantees for change detection while controlling false positives. The test features a configurable tolerance parameter that adjusts sensitivity to behavioral changes for different use cases. We evaluate our approach using two case studies: monitoring changes in (a) toxicity and (b) translation performance. We find that the test is able to detect meaningful changes in behavior distributions using just hundreds of examples.
Leo Richter, Xuanli He, Pasquale Minervini, Matt J. Kusner
ICLR3
2025 Is Complex Query Answering Really Complex?
abstract
Complex query answering (CQA) on knowledge graphs (KGs) is gaining momentum as a challenging reasoning task. In this paper, we show that the current benchmarks for CQA might not be as complex as we think, as the way they are built distorts our perception of progress in this field. For example, we find that in these benchmarks most queries (up to 98% for some query types) can be reduced to simpler problems, e.g., link prediction, where only one link needs to be predicted. The performance of state-of-the-art CQA models decreses significantly when such models are evaluated on queries that cannot be reduced to easier types. Thus, we propose a set of more challenging benchmarks composed of queries that require models to reason over multiple hops and better reflect the construction of real-world KGs. In a systematic empirical investigation, the new benchmarks show that current methods leave much to be desired from current CQA methods.
Cosimo Gregucci, Bo Xiong 0001, Daniel Hernández 0002, Lorenzo Loconte, Pasquale Minervini, Steffen Staab, Antonio Vergari
ICML5
2025 When Can Proxies Improve the Sample Complexity of Preference Learning?
abstract
We address the problem of reward hacking, where maximising a proxy reward does not necessarily increase the true reward. This is a key concern for Large Language Models (LLMs), as they are often fine-tuned on human preferences that may not accurately reflect a true objective. Existing work uses various tricks such as regularisation, tweaks to the reward model, and reward hacking detectors, to limit the influence that such proxy preferences have on a model. Luckily, in many contexts such as medicine, education, and law, a sparse amount of expert data is often available. In these cases, it is often unclear whether the addition of proxy data can improve policy learning. We outline a set of sufficient conditions on proxy feedback that, if satisfied, indicate that proxy data can provably improve the sample complexity of learning the ground truth policy. These conditions can inform the data collection process for specific tasks. The result implies a parameterisation for LLMs that achieves this improved sample complexity. We detail how one can adapt existing architectures to yield this improved sample complexity.
Daniel Augusto R. M. A. de Souza, Zhengyan Shi, Mengyue Yang, Pasquale Minervini, Matt J. Kusner, Alexander D'Amour
ICML5
2025 Are We Done with MMLU?
abstract
Aryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong, Alessio Devoto, Alberto Carlo Maria Mancino, Rohit Saxena, Xuanli He, Yu Zhao, Xiaotang Du, Mohammad Reza Ghasemi Madani, Claire Barale, Robert McHardy, Joshua Harris, Jean Kaddour, Emile Van Krieken, Pasquale Minervini. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Aryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong, Alessio Devoto, Alberto Carlo Maria Mancino, Rohit Saxena, Xuanli He, Yu Zhao 0043, Xiaotang Du, Mohammad Reza Ghasemi Madani, Claire Barale, Robert McHardy, Joshua Harris, Jean Kaddour, Emile van Krieken, Pasquale Minervini
NAACL (Long Papers)16
2025 Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering
abstract
Yu Zhao, Alessio Devoto, Giwon Hong, Xiaotang Du, Aryo Pradipta Gema, Hongru Wang, Xuanli He, Kam-Fai Wong, Pasquale Minervini. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yu Zhao 0043, Alessio Devoto, Giwon Hong, Xiaotang Du, Aryo Pradipta Gema, Hongru Wang 0003, Xuanli He, Kam-Fai Wong, Pasquale Minervini
NAACL (Long Papers)9
2025 Neurosymbolic Reasoning Shortcuts under the Independence Assumption
abstract
The ubiquitous independence assumption among symbolic concepts in neurosymbolic (NeSy) predictors is a convenient simplification: NeSy predictors use it to speed up probabilistic reasoning. Recent works like van Krieken et al. (2024) and Marconato et al. (2024) argued that the independence assumption can hinder learning of NeSy predictors and, more crucially, prevent them from correctly modelling uncertainty. There is, however, scepticism in the NeSy community around the scenarios in which the independence assumption actually limits NeSy systems (Faronius and Dos Martires, 2025). In this work, we settle this question by formally showing that assuming independence among symbolic concepts entails that a model can never represent uncertainty over certain concept combinations. Thus, the model fails to be aware of _reasoning shortcuts_, i.e., the pathological behaviour of NeSy predictors that predict correct downstream tasks but for the wrong reasons.
Emile van Krieken, Pasquale Minervini, Edoardo Maria Ponti, Antonio Vergari
NeSy2
2025 Neurosymbolic Diffusion Models
abstract
Neurosymbolic (NeSy) predictors combine neural perception with symbolic reasoning to solve tasks like visual reasoning. However, standard NeSy predictors assume conditional independence between the symbols they extract, thus limiting their ability to model interactions and uncertainty --- often leading to overconfident predictions and poor out-of-distribution generalisation. To overcome the limitations of the independence assumption, we introduce _neurosymbolic diffusion models_ (NeSyDMs), a new class of NeSy predictors that use discrete diffusion to model dependencies between symbols. Our approach reuses the independence assumption from NeSy predictors at each step of the diffusion process, enabling scalable learning while capturing symbol dependencies and uncertainty quantification. Across both synthetic and real-world benchmarks — including high-dimensional visual path planning and rule-based autonomous driving — NeSyDMs achieve state-of-the-art accuracy among NeSy predictors and demonstrate strong calibration.
Emile van Krieken, Pasquale Minervini, Edoardo Maria Ponti, Antonio Vergari
NeurIPS2
2025 MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
abstract
The rapid extension of context windows in large vision-language models has given rise to long-context vision-language models (LCVLMs), which are capable of handling hundreds of images with interleaved text tokens in a single forward pass. In this work, we introduce MMLongBench, the first benchmark covering a diverse set of long-context vision-language tasks, to evaluate LCVLMs effectively and thoroughly. MMLongBench is composed of 13,331 examples spanning five different categories of downstream tasks, such as Visual RAG and Many-Shot ICL. It also provides broad coverage of image types, including various natural and synthetic images. To assess the robustness of the models to different input lengths, all examples are delivered at five standardized input lengths (8K-128K tokens) via a cross-modal tokenization scheme that combines vision patches and text tokens. Through a thorough benchmarking of 46 closed-source and open-source LCVLMs, we provide a comprehensive analysis of the current models' vision-language long-context ability. Our results show that: i) performance on a single task is a weak proxy for overall long-context capability; ii) both closed-source and open-source models face challenges in long-context vision-language tasks, indicating substantial room for future improvement; iii) models with stronger reasoning ability tend to exhibit better long-context performance. By offering wide task coverage, various image types, and rigorous length control, MMLongBench provides the missing foundation for diagnosing and advancing the next generation of LCVLMs.
Zhaowei Wang 0003, Wenhao Yu 0002, Xiyu Ren, Yu Zhao 0043, Rohit Saxena, Ginny Y. Wong, Simon See, Pasquale Minervini, Yangqiu Song, Mark Steedman
NeurIPS10
2025 Adaptive layer and token selection for efficient fine-tuning of vision transformers
abstract
Foundation models for computer vision built on Vision Transformer (ViT) architectures have become increasingly widespread. However, their fine-tuning process is resource-intensive, slowing their adoption in edge or low-energy applications. We introduce ALaST ( Adaptive Layer Selection for ViT Fine-Tuning ), a novel approach that dynamically optimizes the fine-tuning process to significantly reduce computational cost, memory consumption, and training time. Our method is founded on the critical observation that during fine-tuning, the importance of individual layers and tokens varies substantially across training iterations and depends on the specific mini-batch being processed. ALaST leverages this insight by adaptively estimating layer importance at each fine-tuning step and allocating computational resources—or “compute budgets”—proportionally. Layers assigned lower budgets are either trained with a reduced token set or temporarily frozen. Through comprehensive empirical evaluation on standard benchmarks, we demonstrate that ALaST achieves substantial efficiency gains: up to 1.3 × reduction in training time, 1.5 × reduction in FLOPs, and 2 × decrease in memory requirements, all while maintaining model performance within 0.5 % of full fine-tuning. Notably, our approach provides an automatic schedule for distributing computational resources across layers and can be combined with existing parameter-efficient fine-tuning techniques, offering an orthogonal dimension of optimization for Vision Transformers.
Alessio Devoto, Federico Alvetreti, Jary Pomponi, Paolo Di Lorenzo, Pasquale Minervini, Simone Scardapane
Neurocomputing5
2024 Using Natural Language Explanations to Improve Robustness of In-context Learning
abstract
Recent studies demonstrated that large language models (LLMs) can excel in many tasks via in-context learning (ICL).However, recent works show that ICL-prompted models tend to produce inaccurate results when presented with adversarial inputs.In this work, we investigate whether augmenting ICL with natural language explanations (NLEs) improves the robustness of LLMs on adversarial datasets covering natural language inference and paraphrasing identification.We prompt LLMs with a small set of human-generated NLEs to produce further NLEs, yielding more accurate results than both a zero-shot-ICL setting and using only human-generated NLEs.Our results on five popular LLMs (GPT3.5-turbo,Llama2, Vicuna, Zephyr, and Mistral) show that our approach yields over 6% improvement over baseline approaches for eight adversarial datasets: HANS, ISCS, NaN, ST, PICD, PISP, ANLI, and PAWS.Furthermore, previous studies have demonstrated that prompt selection strategies significantly enhance ICL on in-distribution test sets.However, our findings reveal that these strategies do not match the efficacy of our approach for robustness evaluations, resulting in an accuracy drop of 8% compared to the proposed approach.1
Xuanli He, Yuxiang Wu, Oana-Maria Camburu, Pasquale Minervini, Pontus Stenetorp
ACL (1)4
2024 SparseFit: Few-shot Prompting with Sparse Fine-tuning for Jointly Generating Predictions and Natural Language Explanations
abstract
Models that generate natural language explanations (NLEs) for their predictions have recently gained increasing interest.However, this approach usually demands large datasets of human-written NLEs for the ground-truth answers at training time, which can be expensive and potentially infeasible for some applications.When only a few NLEs are available (a fewshot setup), fine-tuning pre-trained language models (PLMs) in conjunction with promptbased learning has recently shown promising results.However, PLMs typically have billions of parameters, making full fine-tuning expensive.We propose SPARSEFIT, a sparse few-shot finetuning strategy that leverages discrete prompts to jointly generate predictions and NLEs.We experiment with SPARSEFIT on three sizes of the T5 language model and four datasets and compare it against existing state-of-the-art Parameter-Efficient Fine-Tuning (PEFT) techniques.We find that fine-tuning only 6.8% of the model parameters leads to competitive results for both the task performance and the quality of the generated NLEs compared to full finetuning of the model and produces better results on average than other PEFT methods in terms of predictive accuracy and NLE quality.
Jesus Solano, Mardhiyah Sanni, Oana-Maria Camburu, Pasquale Minervini
ACL (1)4
2024 Analysing The Impact of Sequence Composition on Language Model Pre-Training
abstract
Yu Zhao, Yuanbin Qu, Konrad Staniszewski, Szymon Tworkowski, Wei Liu, Piotr Miłoś, Yuxiang Wu, Pasquale Minervini. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yu Zhao 0043, Yuanbin Qu, Konrad Staniszewski, Szymon Tworkowski, Piotr Milos, Yuxiang Wu, Pasquale Minervini
ACL (1)8
2024 A Simple and Effective L_2 Norm-Based Strategy for KV Cache Compression
abstract
The deployment of large language models (LLMs) is often hindered by the extensive memory requirements of the Key-Value (KV) cache, especially as context lengths increase.Existing approaches to reduce the KV cache size involve either fine-tuning the model to learn a compression strategy or leveraging attention scores to reduce the sequence length.We analyse the attention distributions in decoderonly Transformers-based models and observe that attention allocation patterns stay consistent across most layers.Surprisingly, we find a clear correlation between the L 2 norm and the attention scores over cached KV pairs, where a low L 2 norm of a key embedding usually leads to a high attention score during decoding.This finding indicates that the influence of a KV pair is potentially determined by the key embedding itself before being queried.Based on this observation, we compress the KV cache based on the L 2 norm of key embeddings.Our experimental results show that this simple strategy can reduce the KV cache size by 50% on language modelling and needle-in-a-haystack tasks and 90% on passkey retrieval tasks without losing accuracy.Moreover, without relying on the attention scores, this approach remains compatible with FlashAttention, enabling broader applicability.
Alessio Devoto, Yu Zhao 0043, Simone Scardapane, Pasquale Minervini
EMNLP4
2024 Atomic Inference for NLI with Generated Facts as Atoms
abstract
With recent advances, neural models can achieve human-level performance on various natural language tasks.However, there are no guarantees that any explanations from these models are faithful, i.e. that they reflect the inner workings of the model.Atomic inference overcomes this issue, providing interpretable and faithful model decisions.This approach involves making predictions for different components (or atoms) of an instance, before using interpretable and deterministic rules to derive the overall prediction based on the individual atom-level predictions.We investigate the effectiveness of using LLM-generated facts as atoms, decomposing Natural Language Inference premises into lists of facts.While directly using generated facts in atomic inference systems can result in worse performance, with 1) a multi-stage fact generation process, and 2) a training regime that incorporates the facts, our fact-based method outperforms other approaches. 1 Logical Rules for TrainingInstance
Joe Stacey, Pasquale Minervini, Haim Dubossarsky, Oana-Maria Camburu, Marek Rei
EMNLP2
2024 Unveiling and Consulting Core Experts in Retrieval-Augmented MoE-based LLMs
abstract
Xin Zhou, Ping Nie, Yiwen Guo, Haojie Wei, Zhanqiu Zhang, Pasquale Minervini, Ruotian Ma, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Xin Zhou 0012, Ping Nie, Yiwen Guo, Haojie Wei, Zhanqiu Zhang, Pasquale Minervini, Ruotian Ma, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
EMNLP6
2024 On the Independence Assumption in Neurosymbolic Learning
abstract
State-of-the-art neurosymbolic learning systems use probabilistic reasoning to guide neural networks towards predictions that conform to logical constraints. Many such systems assume that the probabilities of the considered symbols are conditionally independent given the input to simplify learning and reasoning. We study and criticise this assumption, highlighting how it can hinder optimisation and prevent uncertainty quantification. We prove that loss functions bias conditionally independent neural networks to become overconfident in their predictions. As a result, they are unable to represent uncertainty over multiple valid options. Furthermore, we prove that the minima of such loss functions are usually highly disconnected and non-convex, and thus difficult to optimise. Our theoretical analysis gives the foundation for replacing the conditional independence assumption and designing more expressive neurosymbolic probabilistic models.
Emile van Krieken, Pasquale Minervini, Edoardo Maria Ponti, Antonio Vergari
ICML2
2023 Adaptive Perturbation-Based Gradient Estimation for Discrete Latent Variable Models
abstract
The integration of discrete algorithmic components in deep learning architectures has numerous applications. Recently, Implicit Maximum Likelihood Estimation, a class of gradient estimators for discrete exponential family distributions, was proposed by combining implicit differentiation through perturbation with the path-wise gradient estimator. However, due to the finite difference approximation of the gradients, it is especially sensitive to the choice of the finite difference step size, which needs to be specified by the user. In this work, we present Adaptive IMLE (AIMLE), the first adaptive gradient estimator for complex discrete distributions: it adaptively identifies the target distribution for IMLE by trading off the density of gradient information with the degree of bias in the gradient estimates. We empirically evaluate our estimator on synthetic examples, as well as on Learning to Explain, Discrete Variational Auto-Encoders, and Neural Relational Inference tasks. In our experiments, we show that our adaptive gradient estimator can produce faithful estimates while requiring orders of magnitude fewer samples than other gradient estimators.
Pasquale Minervini, Luca Franceschi 0001, Mathias Niepert
AAAI1
2023 Combining Inductive and Deductive Reasoning for Query Answering over Incomplete Knowledge Graphs
abstract
Current methods for embedding-based query answering over incomplete Knowledge Graphs (KGs) only focus on inductive reasoning, i.e., predicting answers by learning patterns from the data, and lack the complementary ability to do deductive reasoning, which requires the application of domain knowledge to infer further information. To address this shortcoming, we investigate the problem of incorporating ontologies into embedding-based query answering models by defining the task of embedding-based ontology-mediated query answering. We propose various integration strategies into prominent representatives of embedding models that involve (1) different ontology-driven data augmentation techniques and (2) adaptation of the loss function to enforce the ontology axioms. We design novel benchmarks for the considered task based on the LUBM and the NELL KGs and evaluate our methods on them. The achieved improvements in the setting that requires both inductive and deductive reasoning are from 20% to 55% in HITS@3.
Medina Andresel, Trung Kien Tran, Csaba Domokos, Pasquale Minervini, Daria Stepanova 0001
CIKM4
2023 REFER: An End-to-end Rationale Extraction Framework for Explanation Regularization
abstract
Human-annotated textual explanations are becoming increasingly important in Explainable Natural Language Processing.Rationale extraction aims to provide faithful (i.e., reflective of the behavior of the model) and plausible (i.e., convincing to humans) explanations by highlighting the inputs that had the largest impact on the prediction without compromising the performance of the task model.In recent works, the focus of training rationale extractors was primarily on optimizing for plausibility using human highlights, while the task model was trained on jointly optimizing for task predictive accuracy and faithfulness.We propose REFER, a framework that employs a differentiable rationale extractor that allows to back-propagate through the rationale extraction process.We analyze the impact of using human highlights during training by jointly training the task model and the rationale extractor.In our experiments, REFER yields significantly better results in terms of faithfulness, plausibility, and downstream task accuracy on both in-distribution and out-of-distribution data.On both e-SNLI and CoS-E, our best setting produces better results in terms of composite normalized relative gain than the previous baselines by 11% and 3%, respectively.
Mohammad Reza Ghasemi Madani, Pasquale Minervini
CoNLL2
2023 Machine Learning Survival Models for Relapse Prediction in a Early Stage Lung Cancer Patient
abstract
Lung cancer is one of the leading health complications causing high mortality worldwide. The relapsing behavior of medically treated early-stage lung cancer makes this disease even more complicated. Thus predicting such relapse using a data-centric approach provides a complementary perspective for clinicians to understand the disease. In this preliminary work, we explored off-the-shelf survival models to predict the relapse of early-stage lung cancer patients. We analyzed the survival models on a cohort of 1348 early-stage non-small cell lung cancer (NSCLC) patients in different timestamps. Using the prediction explanation model SHAP (SHapley Additive exPlanations), we further explained the best-performing survival model's predictions. Our explainable predictive model is a potential tool for oncologists that address an unmet clinical need for post-treatment patient stratification based on the relapse hazard.
Mohan Timilsina, Samuele Buosi, Adrianna Janik, Pasquale Minervini, Luca Costabello, Maria Torrente, Mariano Provencio, Virginia Calvo, Carlos Camps, Ana L. Ortega, Bartomeu Massutí, M. Rosario Garcia Campelo, Edel del Barco, Joaquim Bosch-Barrera, Vít Novácek
IJCNN4
2023 Adapting Neural Link Predictors for Data-Efficient Complex Query Answering
abstract
Answering complex queries on incomplete knowledge graphs is a challenging task where a model needs to answer complex logical queries in the presence of missing knowledge. Prior work in the literature has proposed to address this problem by designing architectures trained end-to-end for the complex query answering task with a reasoning process that is hard to interpret while requiring data and resource-intensive training. Other lines of research have proposed re-using simple neural link predictors to answer complex queries, reducing the amount of training data by orders of magnitude while providing interpretable answers. The neural link predictor used in such approaches is not explicitly optimised for the complex query answering task, implying that its scores are not calibrated to interact together. We propose to address these problems via CQD$^{\mathcal{A}}$, a parameter-efficient score \emph{adaptation} model optimised to re-calibrate neural link prediction scores for the complex query answering task. While the neural link predictor is frozen, the adaptation component -- which only increases the number of model parameters by $0.03\%$ -- is trained on the downstream complex query answering task. Furthermore, the calibration component enables us to support reasoning over queries that include atomic negations, which was previously impossible with link predictors. In our experiments, CQD$^{\mathcal{A}}$ produces significantly more accurate results than current state-of-the-art methods, improving from $34.4$ to $35.1$ Mean Reciprocal Rank values averaged across all datasets and query types while using $\leq 30\%$ of the available training query types. We further show that CQD$^{\mathcal{A}}$ is data-efficient, achieving competitive results with only $1\%$ of the complex training queries and robust in out-of-domain evaluations. Source code and datasets are available at https://github.com/EdinburghNLP/adaptive-cqd.
Erik Arakelyan, Pasquale Minervini, Daniel Daza, Michael Cochez, Isabelle Augenstein
NeurIPS2
2023 No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language Models
abstract
The computation necessary for training Transformer-based language models has skyrocketed in recent years. This trend has motivated research on efficient training algorithms designed to improve training, validation, and downstream performance faster than standard training. In this work, we revisit three categories of such algorithms: dynamic architectures (layer stacking, layer dropping), batch selection (selective backprop., RHO-loss), and efficient optimizers (Lion, Sophia). When pre-training BERT and T5 with a fixed computation budget using such methods, we find that their training, validation, and downstream gains vanish compared to a baseline with a fully-decayed learning rate. We define an evaluation protocol that enables computation to be done on arbitrary machines by mapping all computation time to a reference machine which we call reference system time. We discuss the limitations of our proposed protocol and release our code to encourage rigorous research in efficient training procedures: https://github.com/JeanKaddour/NoTrainNoGain.
Jean Kaddour, Oscar Key, Piotr Nawrot, Pasquale Minervini, Matt J. Kusner
NeurIPS4
2023 Synergy between imputed genetic pathway and clinical information for predicting recurrence in early stage non-small cell lung cancer
abstract
OBJECTIVE: Lung cancer exhibits unpredictable recurrence in low-stage tumors and variable responses to different therapeutic interventions. Predicting relapse in early-stage lung cancer can facilitate precision medicine and improve patient survivability. While existing machine learning models rely on clinical data, incorporating genomic information could enhance their efficiency. This study aims to impute and integrate specific types of genomic data with clinical data to improve the accuracy of machine learning models for predicting relapse in early-stage, non-small cell lung cancer patients. METHODS: The study utilized a publicly available TCGA lung cancer cohort and imputed genetic pathway scores into the Spanish Lung Cancer Group (SLCG) data, specifically in 1348 early-stage patients. Initially, tumor recurrence was predicted without imputed pathway scores. Subsequently, the SLCG data were augmented with pathway scores imputed from TCGA. The integrative approach aimed to enhance relapse risk prediction performance. RESULTS: The integrative approach achieved improved relapse risk prediction with the following evaluation metrics: an area under the precision-recall curve (PR-AUC) score of 0.75, an area under the ROC (ROC-AUC) score of 0.80, an F1 score of 0.61, and a Precision of 0.80. The prediction explanation model SHAP (SHapley Additive exPlanations) was employed to explain the machine learning model's predictions. CONCLUSION: We conclude that our explainable predictive model is a promising tool for oncologists that addresses an unmet clinical need of post-treatment patient stratification based on the relapse risk while also improving the predictive power by incorporating proxy genomic data not available for specific patients.
Mohan Timilsina, Dirk Fey, Samuele Buosi, Adrianna Janik, Luca Costabello, Enric Carcereny, Delvys Rodriguez Abreu, Manuel Cobo, Rafael Castro, Reyes Bernabé, Pasquale Minervini, Maria Torrente, Mariano Provencio, Vít Novácek
J. Biomed. Informatics11
2022 Integration of Clinical Information and Imputed Aneuploidy Scores to Enhance Relapse Prediction in Early Stage Lung Cancer Patients
Mohan Timilsina, Samuele Bousi, Dirk Fey, Adrianna Janik, Maria Torrente, Mariano Provencio, Alberto Bermúdez, Enric Carcereny, Luca Costabello, Delvys Rodriguez Abreu, Manuel Cobo, Rafael Castro, Reyes Bernabé, Maria Guirado, Pasquale Minervini, Vít Novácek
AMIA15
2022 MedDistant19: Towards an Accurate Benchmark for Broad-Coverage Biomedical Relation Extraction
abstract
Relation extraction in the biomedical domain is challenging due to the lack of labeled data and high annotation costs, needing domain experts. Distant supervision is commonly used to tackle the scarcity of annotated data by automatically pairing knowledge graph relationships with raw texts. Such a pipeline is prone to noise and has added challenges to scale for covering a large number of biomedical concepts. We investigated existing broad-coverage distantly supervised biomedical relation extraction benchmarks and found a significant overlap between training and test relationships ranging from 26% to 86%. Furthermore, we noticed several inconsistencies in the data construction process of these benchmarks, and where there is no train-test leakage, the focus is on interactions between narrower entity types. This work presents a more accurate benchmark MedDistant19 for broad-coverage distantly supervised biomedical relation extraction that addresses these shortcomings and is obtained by aligning the MEDLINE abstracts with the widely used SNOMED Clinical Terms knowledge base. Lacking thorough evaluation with domain-specific language models, we also conduct experiments validating general domain relation extraction findings to biomedical relation extraction.
Saadullah Amin, Pasquale Minervini, Pontus Stenetorp, Günter Neumann
COLING2
2022 Logical Reasoning with Span-Level Predictions for Interpretable and Robust NLI Models
abstract
Current Natural Language Inference (NLI) models achieve impressive results, sometimes outperforming humans when evaluating on indistribution test sets.However, as these models are known to learn from annotation artefacts and dataset biases, it is unclear to what extent the models are learning the task of NLI instead of learning from shallow heuristics in their training data.We address this issue by introducing a logical reasoning framework for NLI, creating highly transparent model decisions that are based on logical rules.Unlike prior work, we show that improved interpretability can be achieved without decreasing the predictive accuracy.We almost fully retain performance on SNLI, while also identifying the exact hypothesis spans that are responsible for each model prediction.Using the e-SNLI human explanations, we verify that our model makes sensible decisions at a span level, despite not using any span labels during training.We can further improve model performance and span-level decisions by using the e-SNLI explanations during training.Finally, our model is more robust in a reduced data setting.When training with only 1,000 examples, out-of-distribution performance improves on the MNLI matched and mismatched validation sets by 13% and 16% relative to the baseline.Training with fewer observations yields further improvements, both in-distribution and out-ofdistribution.
Joe Stacey, Pasquale Minervini, Haim Dubossarsky, Marek Rei
EMNLP2
2022 An Efficient Memory-Augmented Transformer for Knowledge-Intensive NLP Tasks
abstract
Access to external knowledge is essential for many natural language processing tasks, such as question answering and dialogue.Existing methods often rely on a parametric model that stores knowledge in its parameters, or use a retrieval-augmented model that has access to an external knowledge source.Parametric and retrieval-augmented models have complementary strengths in terms of computational efficiency and predictive accuracy.To combine the strength of both approaches, we propose the Efficient Memory-Augmented Transformer (EMAT) -it encodes external knowledge into a key-value memory and exploits the fast maximum inner product search for memory querying.We also introduce pre-training tasks that allow EMAT to encode informative key-value representations, and to learn an implicit strategy to integrate multiple memory slots into the transformer.Experiments on various knowledge-intensive tasks such as question answering and dialogue datasets show that, simply augmenting parametric models (T5-base) using our method produces more accurate results (e.g., 25.8 → 44.3 EM on NQ) while retaining a high throughput (e.g., 1000 queries/s on NQ).Compared to retrievalaugmented models, EMAT runs substantially faster across the board and produces more accurate results on WoW and ELI5. 1
Yuxiang Wu, Yu Zhao 0043, Baotian Hu, Pasquale Minervini, Pontus Stenetorp, Sebastian Riedel 0001
EMNLP4
2022 Complex Query Answering with Neural Link Predictors (Extended Abstract)
abstract
Neural link predictors are useful for identifying missing edges in large scale Knowledge Graphs. However, it is still not clear how to use these models for answering more complex queries containing logical conjunctions (∧), disjunctions (∨), and existential quantifiers (∃). We propose a framework for efficiently answering complex queries on in- complete Knowledge Graphs. We translate each query into an end-to-end differentiable objective, where the truth value of each atom is computed by a pre-trained neural link predictor. We then analyse two solutions to the optimisation problem, including gradient-based and combinatorial search. In our experiments, the proposed approach produces more accurate results than state-of-the-art methods — black-box models trained on millions of generated queries — without the need for training on a large and diverse set of complex queries. Using orders of magnitude less training data, we obtain relative improvements ranging from 8% up to 40% in Hits@3 across multiple knowledge graphs. We find that it is possible to explain the outcome of our model in terms of the intermediate solutions identified for each of the complex query atoms. All our source code and datasets are available online (https://github.com/uclnlp/cqd).
Pasquale Minervini, Erik Arakelyan, Daniel Daza, Michael Cochez
IJCAI1
2022 ReFactor GNNs: Revisiting Factorisation-based Models from a Message-Passing Perspective
abstract
Factorisation-based Models (FMs), such as DistMult, have enjoyed enduring success for Knowledge Graph Completion (KGC) tasks, often outperforming Graph Neural Networks (GNNs). However, unlike GNNs, FMs struggle to incorporate node features and generalise to unseen nodes in inductive settings. Our work bridges the gap between FMs and GNNs by proposing ReFactor GNNs. This new architecture draws upon $\textit{both}$ modelling paradigms, which previously were largely thought of as disjoint. Concretely, using a message-passing formalism, we show how FMs can be cast as GNNs by reformulating the gradient descent procedure as message-passing operations, which forms the basis of our ReFactor GNNs. Across a multitude of well-established KGC benchmarks, our ReFactor GNNs achieve comparable transductive performance to FMs, and state-of-the-art inductive performance while using an order of magnitude fewer parameters.
Pushkar Mishra, Luca Franceschi 0001, Pasquale Minervini, Pontus Stenetorp, Sebastian Riedel 0001
NeurIPS4
2021 On Predicting Recurrence in Early Stage Non-small Cell Lung Cancer
Sameh K. Mohamed, Brian Walsh, Mohan Timilsina, Vít Novácek, Maria Torrente, Fabio Franco, Mariano Provencio, Adrianna Janik, Luca Costabello, Pontus Stenetorp, Pasquale Minervini
AMIA11
2021 First Workshop on Knowledge Injection in Neural Networks (KINN)
abstract
Deep learning (DL) has made rapid progress in the last decade, with neural network-based language and vision models achieving state-of-the-art performance in various tasks. Yet purely data-driven neural network models exhibit several issues impacting real-world deployment of such models adversely. These include reliance on large quantities of training data, poor robustness, lack of generalization, poor explainability, and glaring gaps in implicit and commonsense knowledge. The availability of rich structured (or semi-structured) knowledge sources has spurred the research community into exploring Knowledge Injection in Neural Networks (KINN) as a means of mitigating the above-mentioned challenges. This has led to the development of hybrid AI systems that combine the purely data-driven learning of the neural network models with an infusion of knowledge from external sources. Such KINN systems include the development of retrieval augmented neural models, neuro symbolic systems and a plethora of combinations of NNs and knowledge graphs and structured knowledge bases.
Vasudev Lal, Somak Aditya, Yezhou Yang, Pasquale Minervini, Sandya Mannarswamy
CIKM4
2021 Stereotype and Skew: Quantifying Gender Bias in Pre-trained and Fine-tuned Language Models
abstract
Daniel de Vassimon Manela, David Errington, Thomas Fisher, Boris van Breugel, Pasquale Minervini. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Daniel de Vassimon Manela, David Errington, Thomas Fisher, Boris van Breugel, Pasquale Minervini
EACL5
2021 Complex Query Answering with Neural Link Predictors
Erik Arakelyan, Daniel Daza, Pasquale Minervini, Michael Cochez
ICLR3
2021 Implicit MLE: Backpropagating Through Discrete Exponential Family Distributions
abstract
Combining discrete probability distributions and combinatorial optimization problems with neural network components has numerous applications but poses several challenges. We propose Implicit Maximum Likelihood Estimation (I-MLE), a framework for end-to-end learning of models combining discrete exponential family distributions and differentiable neural components. I-MLE is widely applicable as it only requires the ability to compute the most probable states and does not rely on smooth relaxations. The framework encompasses several approaches such as perturbation-based implicit differentiation and recent methods to differentiate through black-box combinatorial solvers. We introduce a novel class of noise distributions for approximating marginals via perturb-and-MAP. Moreover, we show that I-MLE simplifies to maximum likelihood estimation when used in some recently studied learning settings that involve combinatorial solvers. Experiments on several datasets suggest that I-MLE is competitive with and often outperforms existing approaches which rely on problem-specific relaxations.
Mathias Niepert, Pasquale Minervini, Luca Franceschi 0001
NeurIPS2
2021 PAQ: 65 Million Probably-Asked Questions and What You Can Do With Them
abstract
Abstract Open-domain Question Answering models that directly leverage question-answer (QA) pairs, such as closed-book QA (CBQA) models and QA-pair retrievers, show promise in terms of speed and memory compared with conventional models which retrieve and read from text corpora. QA-pair retrievers also offer interpretable answers, a high degree of control, and are trivial to update at test time with new knowledge. However, these models fall short of the accuracy of retrieve-and-read systems, as substantially less knowledge is covered by the available QA-pairs relative to text corpora like Wikipedia. To facilitate improved QA-pair models, we introduce Probably Asked Questions (PAQ), a very large resource of 65M automatically generated QA-pairs. We introduce a new QA-pair retriever, RePAQ, to complement PAQ. We find that PAQ preempts and caches test questions, enabling RePAQ to match the accuracy of recent retrieve-and-read models, whilst being significantly faster. Using PAQ, we train CBQA models which outperform comparable baselines by 5%, but trail RePAQ by over 15%, indicating the effectiveness of explicit retrieval. RePAQ can be configured for size (under 500MB) or speed (over 1K questions per second) while retaining high accuracy. Lastly, we demonstrate RePAQ’s strength at selective QA, abstaining from answering when it is likely to be incorrect. This enables RePAQ to “back-off” to a more expensive state-of-the-art model, leading to a combined system which is both more accurate and 2x faster than the state-of-the-art model alone.
Patrick S. H. Lewis, Yuxiang Wu, Linqing Liu, Pasquale Minervini, Heinrich Küttler, Aleksandra Piktus, Pontus Stenetorp, Sebastian Riedel 0001
Trans. Assoc. Comput. Linguistics4
2020 Differentiable Reasoning on Large Knowledge Bases and Natural Language
abstract
Reasoning with knowledge expressed in natural language and Knowledge Bases (KBs) is a major challenge for Artificial Intelligence, with applications in machine reading, dialogue, and question answering. General neural architectures that jointly learn representations and transformations of text are very data-inefficient, and it is hard to analyse their reasoning process. These issues are addressed by end-to-end differentiable reasoning systems such as Neural Theorem Provers (NTPs), although they can only be used with small-scale symbolic KBs. In this paper we first propose Greedy NTPs (GNTPs), an extension to NTPs addressing their complexity and scalability limitations, thus making them applicable to real-world datasets. This result is achieved by dynamically constructing the computation graph of NTPs and including only the most promising proof paths during inference, thus obtaining orders of magnitude more efficient models 1. Then, we propose a novel approach for jointly reasoning over KBs and textual mentions, by embedding logic facts and natural language sentences in a shared embedding space. We show that GNTPs perform on par with NTPs at a fraction of their cost while achieving competitive link prediction results on large datasets, providing explanations for predictions, and inducing interpretable models.
Pasquale Minervini, Matko Bosnjak, Tim Rocktäschel, Sebastian Riedel 0001, Edward Grefenstette
AAAI1
2020 Make Up Your Mind! Adversarial Generation of Inconsistent Natural Language Explanations
abstract
To increase trust in artificial intelligence systems, a promising research direction consists of designing neural models capable of generating natural language explanations for their predictions.In this work, we show that such models are nonetheless prone to generating mutually inconsistent explanations, such as "Because there is a dog in the image."and "Because there is no dog in the [same] image.",exposing flaws in either the decision-making process of the model or in the generation of the explanations.We introduce a simple yet effective adversarial framework for sanity checking models against the generation of inconsistent natural language explanations.Moreover, as part of the framework, we address the problem of adversarial attacks with full target sequences, a scenario that was not previously addressed in sequence-to-sequence attacks.Finally, we apply our framework on a state-of-the-art neural natural language inference model that provides natural language explanations for its predictions.Our framework shows that this model is capable of generating a significant number of inconsistent explanations.PREMISE: A guy in a red jacket is snowboarding in midair.
Oana-Maria Camburu, Brendan Shillingford, Pasquale Minervini, Thomas Lukasiewicz, Phil Blunsom
ACL3
2020 Avoiding the Hypothesis-Only Bias in Natural Language Inference via Ensemble Adversarial Training
abstract
Natural Language Inference (NLI) datasets contain annotation artefacts resulting in spurious correlations between the natural language utterances and their respective entailment classes.These artefacts are exploited by neural networks even when only considering the hypothesis and ignoring the premise, leading to unwanted biases.Belinkov et al. (2019b) proposed tackling this problem via adversarial training, but this can lead to learned sentence representations that still suffer from the same biases.We show that the bias can be reduced in the sentence representations by using an ensemble of adversaries, encouraging the model to jointly decrease the accuracy of these different adversaries while fitting the data.This approach produces more robust NLI models, outperforming previous de-biasing efforts when generalised to 12 other NLI datasets (Belinkov et al., 2019a;Mahabadi et al., 2020).In addition, we find that the optimal number of adversarial classifiers depends on the dimensionality of the sentence representations, with larger sentence representations being more difficult to de-bias while benefiting from using a greater number of adversaries.
Joe Stacey, Pasquale Minervini, Haim Dubossarsky, Sebastian Riedel 0001, Tim Rocktäschel
EMNLP (1)2
2020 Don't Read Too Much Into It: Adaptive Computation for Open-Domain Question Answering
abstract
Most approaches to Open-Domain Question Answering consist of a light-weight retriever that selects a set of candidate passages, and a computationally expensive reader that examines the passages to identify the correct answer.Previous works have shown that as the number of retrieved passages increases, so does the performance of the reader.However, they assume all retrieved passages are of equal importance and allocate the same amount of computation to them, leading to a substantial increase in computational cost.To reduce this cost, we propose the use of adaptive computation to control the computational budget allocated for the passages to be read.We first introduce a technique operating on individual passages in isolation which relies on anytime prediction and a per-layer estimation of an early exit probability.We then introduce SKY-LINEBUILDER, an approach for dynamically deciding on which passage to allocate computation at each step, based on a resource allocation policy trained via reinforcement learning.Our results on SQuAD-Open show that adaptive computation with global prioritisation improves over several strong static and adaptive methods, leading to a 4.3x reduction in computation while retaining 95% performance of the full model.
Yuxiang Wu, Sebastian Riedel 0001, Pasquale Minervini, Pontus Stenetorp
EMNLP (1)3
2020 Learning Reasoning Strategies in End-to-End Differentiable Proving
abstract
Attempts to render deep learning models interpretable, data-efficient, and robust have seen some success through hybridisation with rule-based systems, for example, in Neural Theorem Provers (NTPs). These neuro-symbolic models can induce interpretable rules and learn representations from data via back-propagation, while providing logical explanations for their predictions. However, they are restricted by their computational complexity, as they need to consider all possible proof paths for explaining a goal, thus rendering them unfit for large-scale applications. We present Conditional Theorem Provers (CTPs), an extension to NTPs that learns an optimal rule selection strategy via gradient-based optimisation. We show that CTPs are scalable and yield state-of-the-art results on the CLUTRR dataset, which tests systematic generalisation of neural models by learning to reason over smaller graphs and evaluating on larger ones. Finally, CTPs show better link prediction results on standard benchmarks in comparison with other neural-symbolic models, while being explainable.
Pasquale Minervini, Sebastian Riedel 0001, Pontus Stenetorp, Edward Grefenstette, Tim Rocktäschel
ICML1
2019 NLProlog: Reasoning with Weak Unification for Question Answering in Natural Language
abstract
Rule-based models are attractive for various tasks because they inherently lead to interpretable and explainable decisions and can easily incorporate prior knowledge.However, such systems are difficult to apply to problems involving natural language, due to its linguistic variability.In contrast, neural models can cope very well with ambiguity by learning distributed representations of words and their composition from data, but lead to models that are difficult to interpret.In this paper, we describe a model combining neural networks with logic programming in a novel manner for solving multi-hop reasoning tasks over natural language.Specifically, we propose to use a Prolog prover which we extend to utilize a similarity function over pretrained sentence encoders.We fine-tune the representations for the similarity function via backpropagation.This leads to a system that can apply rulebased reasoning to natural language, and induce domain-specific rules from training data.We evaluate the proposed system on two different question answering tasks, showing that it outperforms two baselines -BIDAF (Seo et al., 2016a) and FASTQA (Weissenborn et al., 2017b) on a subset of the WIKIHOP corpus and achieves competitive results on the MEDHOP data set (Welbl et al., 2017).
Leon Weber-Genzel, Pasquale Minervini, Jannes Münchmeyer, Ulf Leser, Tim Rocktäschel
ACL (1)2
2018 Convolutional 2D Knowledge Graph Embeddings
abstract
Link prediction for knowledge graphs is the task of predicting missing relationships between entities. Previous work on link prediction has focused on shallow, fast models which can scale to large knowledge graphs. However, these models learn less expressive features than deep, multi-layer models — which potentially limits performance. In this work we introduce ConvE, a multi-layer convolutional network model for link prediction, and report state-of-the-art results for several established datasets. We also show that the model is highly parameter efficient, yielding the same performance as DistMult and R-GCN with 8x and 17x fewer parameters. Analysis of our model suggests that it is particularly effective at modelling nodes with high indegree — which are common in highly-connected, complex knowledge graphs such as Freebase and YAGO3. In addition, it has been noted that the WN18 and FB15k datasets suffer from test set leakage, due to inverse relations from the training set being present in the test set — however, the extent of this issue has so far not been quantified. We find this problem to be severe: a simple rule-based model can achieve state-of-the-art results on both WN18 and FB15k. To ensure that models are evaluated on datasets where simply exploiting inverse relations cannot yield competitive results, we investigate and validate several commonly used datasets — deriving robust variants where necessary. We then perform experiments on these robust datasets for our own and several previously proposed models, and find that ConvE achieves state-of-the-art Mean Reciprocal Rank across all datasets.
Tim Dettmers, Pasquale Minervini, Pontus Stenetorp, Sebastian Riedel 0001
AAAI2
2018 Adversarially Regularising Neural NLI Models to Integrate Logical Background Knowledge
abstract
Adversarial examples are inputs to machine learning models designed to cause the model to make a mistake.They are useful for understanding the shortcomings of machine learning models, interpreting their results, and for regularisation.In NLP, however, most example generation strategies produce input text by using known, pre-specified semantic transformations, requiring significant manual effort and in-depth understanding of the problem and domain.In this paper, we investigate the problem of automatically generating adversarial examples that violate a set of given First-Order Logic constraints in Natural Language Inference (NLI).We reduce the problem of identifying such adversarial examples to a combinatorial optimisation problem, by maximising a quantity measuring the degree of violation of such constraints and by using a language model for generating linguisticallyplausible examples.Furthermore, we propose a method for adversarially regularising neural NLI models for incorporating background knowledge.Our results show that, while the proposed method does not always improve results on the SNLI and MultiNLI datasets, it significantly and consistently increases the predictive accuracy on adversarially-crafted datasets -up to a 79.6% relative improvement -while drastically reducing the number of background knowledge violations.Furthermore, we show that adversarial examples transfer among model architectures, and that the proposed adversarial training procedure improves the robustness of NLI models to adversarial examples.
Pasquale Minervini, Sebastian Riedel 0001
CoNLL1
2018 Adaptive Knowledge Propagation in Web Ontologies
abstract
We focus on the problem of predicting missing assertions in Web ontologies. We start from the assumption that individual resources that are similar in some aspects are more likely to be linked by specific relations: this phenomenon is also referred to as homophily and emerges in a variety of relational domains. In this article, we propose a method for (1) identifying which relations in the ontology are more likely to link similar individuals and (2) efficiently propagating knowledge across chains of similar individuals. By enforcing sparsity in the model parameters, the proposed method is able to select only the most relevant relations for a given prediction task. Our experimental evaluation demonstrates the effectiveness of the proposed method in comparison to state-of-the-art methods from the literature.
Pasquale Minervini, Volker Tresp, Claudia d'Amato, Nicola Fanizzi
ACM Trans. Web1
2017 Regularizing Knowledge Graph Embeddings via Equivalence and Inversion Axioms
Pasquale Minervini, Luca Costabello, Emir Muñoz, Vít Novácek, Pierre-Yves Vandenbussche
ECML/PKDD (1)1
2017 Adversarial Sets for Regularising Neural Link Predictors
Pasquale Minervini, Thomas Demeester, Tim Rocktäschel, Sebastian Riedel 0001
UAI1
2016 Efficient energy-based embedding models for link prediction in knowledge graphs
Pasquale Minervini, Claudia d'Amato, Nicola Fanizzi
J. Intell. Inf. Syst.1
2015 Scalable Learning of Entity and Predicate Embeddings for Knowledge Graph Completion
abstract
Knowledge Graphs (KGs) are a widely used formalism for representing knowledge in the Web of Data. We focus on the problem of link prediction, i.e. predicting missing links in large knowledge graphs, so to discover new facts about the world. Representation learning models that embed entities and relation types in continuous vector spaces recently were used to achieve new state-of-the-art link prediction results. A limiting factor in these models is that the process of learning the optimal embedding vectors can be really time-consuming, and might even require days of computations for large KGs. In this work, we propose a principled method for sensibly reducing the learning time, while converging to more accurate link prediction models. Furthermore, we employ the proposed method for training and evaluating a set of novel and scalable models. Our extensive evaluations show significant improvements over state-of-the-art link prediction methods on several datasets.
Pasquale Minervini, Nicola Fanizzi, Claudia d'Amato, Floriana Esposito
ICMLA1
2014 Adaptive Knowledge Propagation in Web Ontologies
Pasquale Minervini, Claudia d'Amato, Nicola Fanizzi, Floriana Esposito
EKAW1
2014 A Gaussian Process Model for Knowledge Propagation in Web Ontologies
abstract
We consider the problem of predicting missing class-memberships and property values of individual resources in Web ontologies. We first identify which relations tend to link similar individuals by means of a finite-set Gaussian Process regression model, and then efficiently propagate knowledge about individuals across their relations. Our experimental evaluation demonstrates the effectiveness of the proposed method.
Pasquale Minervini, Claudia d'Amato, Nicola Fanizzi, Floriana Esposito
ICDM1
2013 Transductive Inference for Class-Membership Propagation in Web Ontologies
Pasquale Minervini, Claudia d'Amato, Nicola Fanizzi, Floriana Esposito
ESWC1
2010 Can Real-Time Machine Translation Overcome Language Barriers in Distributed Requirements Engineering?
abstract
In global software projects work takes place over long distances, meaning that communication will often involve distant cultures with different languages and communication styles that, in turn, exacerbate communication problems. However, being aware of cultural distance is not sufficient to overcome many of the barriers that language differences bring in the way of global project success. In this paper, we investigate the adoption of machine translation (MT) services in synchronous text-based chat in order to overcome any language barrier existing among groups of stakeholders who are remotely negotiating software requirements. We report our findings from a simulated study that compares the efficiency and the effectiveness of two MT services, Google Translate and apertium-service, in translating the messages exchanged during four distributed requirements engineering workshops. The results show that (a) Google Translate produces significantly more adequate translations than Apertium from English to Italian; (b) both services can be used in text-based chat without disrupting real-time interaction.
Fabio Calefato, Filippo Lanubile, Pasquale Minervini
ICGSE3