Erik Arakelyan

dblp:175/1770 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Question answering and dialogue systems · 25% Graph learning · 20% Knowledge representation and reasoning · 17%
Databases, data mining, and information retrieval
2 papers
Knowledge graphs · 62% Data mining · 38%
Theoretical computer science
1 paper
Logic in computer science · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning
link prediction
1.222023
Adapting Neural Link Predictors for Data-Efficient Complex Query Answering · NeurIPS 2023
Complex Query Answering with Neural Link Predictors (Extended Abstract) · IJCAI 2022
Natural language and speech › Question answering and dialogue systems › knowledge base question answering
complex question answering
1.122022
Complex Query Answering with Neural Link Predictors (Extended Abstract) · IJCAI 2022
Complex Query Answering with Neural Link Predictors · ICLR 2021
Machine learning › Trustworthy machine learning › interpretability
faithful reasoning
0.912025
FLARE: Faithful Logic-Aided Reasoning and Exploration · EMNLP 2025
Logic in computer science
logic programming
0.912025
FLARE: Faithful Logic-Aided Reasoning and Exploration · EMNLP 2025
Machine learning › Transfer learning and domain adaptation › domain adaptation
multi-source domain adaptation
0.712023
Topic-Guided Sampling For Data-Efficient Multi-Domain Stance Detection · ACL (1) 2023
Natural language and speech › Information extraction and text analysis
stance detection
0.712023
Topic-Guided Sampling For Data-Efficient Multi-Domain Stance Detection · ACL (1) 2023
Knowledge graphs › knowledge graph querying
complex query answering
0.712023
Adapting Neural Link Predictors for Data-Efficient Complex Query Answering · NeurIPS 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph
0.612022
Complex Query Answering with Neural Link Predictors (Extended Abstract) · IJCAI 2022
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
deductive question answering
0.512021
Complex Query Answering with Neural Link Predictors · ICLR 2021
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph reasoning
0.512021
Complex Query Answering with Neural Link Predictors · ICLR 2021
Natural language and speech › Language models and text generation › prompting
chain-of-thought prompting
0.312025
FLARE: Faithful Logic-Aided Reasoning and Exploration · EMNLP 2025
Data mining › sampling
diversity sampling
0.212023
Topic-Guided Sampling For Data-Efficient Multi-Domain Stance Detection · ACL (1) 2023
Data mining
sampling
0.212023
Topic-Guided Sampling For Data-Efficient Multi-Domain Stance Detection · ACL (1) 2023

Methods — techniques the papers use, named apart from their topics

neural link predictor · 2.4multi-hop search · 1.7logic program formalization · 1.7large language model · 1.7topic-guided sampling · 1.3score adaptation · 1.3contrastive learning · 1.3gradient-based optimization · 0.6combinatorial search · 0.6query embedding · 0.5
YearPublicationVenuePosition
2025 SynDARin: Synthesising Datasets for Automated Reasoning in Low-Resource Languages
abstract
Question Answering (QA) datasets have been instrumental in developing and evaluating Large Language Model (LLM) capabilities. However, such datasets are scarce for languages other than English due to the cost and difficulties of collection and manual annotation. This means that producing novel models and measuring the performance of multilingual LLMs in low-resource languages is challenging. To mitigate this, we propose SynDARin, a method for generating and validating QA datasets for low-resoucre languages. We utilize parallel content mining to obtain human-curated paragraphs between English and the target language. We use the English data as context to generate synthetic multiple-choice (MC) question-answer pairs, which are automatically translated and further validated for quality. Combining these with their designated non-English human-curated paragraphs form the final QA dataset. The method allows to maintain content quality, reduces the likelihood of factual errors, and circumvents the need for costly annotation. To test the method, we created a QA dataset with 1.2K samples for the Armenian language. The human evaluation shows that 98% of the generated English data maintains quality and diversity in the question types and topics, while the translation validation pipeline can filter out ~70% of data with poor quality. We use the dataset to benchmark state-of-the-art LLMs, showing their inability to achieve human accuracy with some model performances closer to random chance. This shows that the generated dataset is non-trivial and can be used to evaluate reasoning capabilities in low-resource language.
Gayane Ghazaryan, Erik Arakelyan, Isabelle Augenstein, Pasquale Minervini
COLING2
2025 FLARE: Faithful Logic-Aided Reasoning and Exploration
abstract
Modern Question Answering (QA) and Reasoning approaches with Large Language Models (LLMs) commonly use Chain-of-Thought (CoT) prompting but struggle with generating outputs faithful to their intermediate reasoning chains.While neuro-symbolic methods like Faithful CoT (F-CoT) offer higher faithfulness through external solvers, they require codespecialized models and struggle with ambiguous tasks.We introduce Faithful Logic-Aided Reasoning and Exploration (FLARE), which uses LLMs to plan solutions, formalize queries into logic programs, and simulate code execution through multi-hop search without external solvers.Our method achieves SOTA results on 7 out of 9 diverse reasoning benchmarks and 3 out of 3 logic inference benchmarks while enabling measurement of reasoning faithfulness.We demonstrate that model faithfulness correlates with performance and that successful reasoning traces show an 18.1% increase in unique emergent facts, 8.6% higher overlap between code-defined and execution-trace relations, and 3.6% reduction in unused relations.
Erik Arakelyan, Pasquale Minervini, Patrick S. H. Lewis, Patrick Verga, Isabelle Augenstein
EMNLP1
2024 Semantic Sensitivities and Inconsistent Predictions: Measuring the Fragility of NLI Models
abstract
Recent studies of the emergent capabilities of transformer-based Natural Language Understanding (NLU) models have indicated that they have an understanding of lexical and compositional semantics.We provide evidence that suggests these claims should be taken with a grain of salt: we find that state-of-the-art Natural Language Inference (NLI) models are sensitive towards minor semantics preserving surface-form variations, which lead to sizable inconsistent model decisions during inference.Notably, this behaviour differs from valid and in-depth comprehension of compositional semantics, however does neither emerge when evaluating model accuracy on standard benchmarks nor when probing for syntactic, monotonic, and logically robust reasoning.We propose a novel framework to measure the extent of semantic sensitivity.To this end, we evaluate NLI models on adversarially generated examples containing minor semantics-preserving surface-form input noise.This is achieved using conditional text generation, with the explicit condition that the NLI model predicts the relationship between the original and adversarial inputs as a symmetric equivalence entailment.We systematically study the effects of the phenomenon across NLI models for in-and outof domain settings.Our experiments show that semantic sensitivity causes performance degradations of 12.92% and 23.71% average over in-and out-of-domain settings, respectively.We further perform ablation studies, analysing this phenomenon across models, datasets, and variations in inference and show that semantic sensitivity can lead to major inconsistency within model predictions.
Erik Arakelyan, Zhaoqi Liu, Isabelle Augenstein
EACL (1)1
2023 Topic-Guided Sampling For Data-Efficient Multi-Domain Stance Detection
abstract
Stance Detection is concerned with identifying the attitudes expressed by an author towards a target of interest. This task spans a variety of domains ranging from social media opinion identification to detecting the stance for a legal claim. However, the framing of the task varies within these domains, in terms of the data collection protocol, the label dictionary and the number of available annotations. Furthermore, these stance annotations are significantly imbalanced on a per-topic and inter-topic basis. These make multi-domain stance detection a challenging task, requiring standardization and domain adaptation. To overcome this challenge, we propose Topic Efficient StancE Detection (TESTED), consisting of a topic-guided diversity sampling technique and a contrastive objective that is used for fine-tuning a stance classifier. We evaluate the method on an existing benchmark of 16 datasets with in-domain, i.e. all topics seen and out-of-domain, i.e. unseen topics, experiments. The results show that our method outperforms the state-of-the-art with an average of 3.5 F1 points increase in-domain, and is more generalizable with an averaged increase of 10.2 F1 on out-of-domain evaluation while using ≤ 10% of the training data. We show that our sampling technique mitigates both inter- and per-topic class imbalances. Finally, our analysis demonstrates that the contrastive learning objective allows the model a more pronounced segmentation of samples with varying labels.
Erik Arakelyan, Arnav Arora, Isabelle Augenstein
ACL (1)1
2023 Adapting Neural Link Predictors for Data-Efficient Complex Query Answering
abstract
Answering complex queries on incomplete knowledge graphs is a challenging task where a model needs to answer complex logical queries in the presence of missing knowledge. Prior work in the literature has proposed to address this problem by designing architectures trained end-to-end for the complex query answering task with a reasoning process that is hard to interpret while requiring data and resource-intensive training. Other lines of research have proposed re-using simple neural link predictors to answer complex queries, reducing the amount of training data by orders of magnitude while providing interpretable answers. The neural link predictor used in such approaches is not explicitly optimised for the complex query answering task, implying that its scores are not calibrated to interact together. We propose to address these problems via CQD$^{\mathcal{A}}$, a parameter-efficient score \emph{adaptation} model optimised to re-calibrate neural link prediction scores for the complex query answering task. While the neural link predictor is frozen, the adaptation component -- which only increases the number of model parameters by $0.03\%$ -- is trained on the downstream complex query answering task. Furthermore, the calibration component enables us to support reasoning over queries that include atomic negations, which was previously impossible with link predictors. In our experiments, CQD$^{\mathcal{A}}$ produces significantly more accurate results than current state-of-the-art methods, improving from $34.4$ to $35.1$ Mean Reciprocal Rank values averaged across all datasets and query types while using $\leq 30\%$ of the available training query types. We further show that CQD$^{\mathcal{A}}$ is data-efficient, achieving competitive results with only $1\%$ of the complex training queries and robust in out-of-domain evaluations. Source code and datasets are available at https://github.com/EdinburghNLP/adaptive-cqd.
Erik Arakelyan, Pasquale Minervini, Daniel Daza, Michael Cochez, Isabelle Augenstein
NeurIPS1
2022 Complex Query Answering with Neural Link Predictors (Extended Abstract)
abstract
Neural link predictors are useful for identifying missing edges in large scale Knowledge Graphs. However, it is still not clear how to use these models for answering more complex queries containing logical conjunctions (∧), disjunctions (∨), and existential quantifiers (∃). We propose a framework for efficiently answering complex queries on in- complete Knowledge Graphs. We translate each query into an end-to-end differentiable objective, where the truth value of each atom is computed by a pre-trained neural link predictor. We then analyse two solutions to the optimisation problem, including gradient-based and combinatorial search. In our experiments, the proposed approach produces more accurate results than state-of-the-art methods — black-box models trained on millions of generated queries — without the need for training on a large and diverse set of complex queries. Using orders of magnitude less training data, we obtain relative improvements ranging from 8% up to 40% in Hits@3 across multiple knowledge graphs. We find that it is possible to explain the outcome of our model in terms of the intermediate solutions identified for each of the complex query atoms. All our source code and datasets are available online (https://github.com/uclnlp/cqd).
Pasquale Minervini, Erik Arakelyan, Daniel Daza, Michael Cochez
IJCAI2
2021 Complex Query Answering with Neural Link Predictors
Erik Arakelyan, Daniel Daza, Pasquale Minervini, Michael Cochez
ICLR1