Houman Mehrafarin

dblp:317/0323 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 44% Knowledge representation and reasoning · 44% Question answering and dialogue systems · 13%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Knowledge representation and reasoning
logic-based reasoning
0.812024
Reasoning or a Semblance of it? A Diagnostic Study of Transitive Reasoning in LLMs · EMNLP 2024
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
compositional question answering
0.212024
Reasoning or a Semblance of it? A Diagnostic Study of Transitive Reasoning in LLMs · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

diagnostic probing · 0.8
YearPublicationVenuePosition
2024 Reasoning or a Semblance of it? A Diagnostic Study of Transitive Reasoning in LLMs
abstract
Evaluating Large Language Models (LLMs) on reasoning benchmarks demonstrates their ability to solve compositional questions.However, little is known of whether these models engage in genuine logical reasoning or simply rely on implicit cues to generate answers.In this paper, we investigate the transitive reasoning capabilities of two distinct LLM architectures, LLaMA 2 and Flan-T5, by manipulating facts within two compositional datasets: QASC and Bamboogle.We controlled for potential cues that might influence the models' performance, including (a) word/phrase overlaps across sections of test input; (b) models' inherent knowledge during pre-training or fine-tuning; and (c) Named Entities.Our findings reveal that while both models leverage (a), Flan-T5 shows more resilience to experiments (b and c), having less variance than LLaMA 2. This suggests that models may develop an understanding of transitivity through fine-tuning on knowingly relevant datasets, a hypothesis we leave to future work 1 .
Houman Mehrafarin, Arash Eshghi, Ioannis Konstas
EMNLP1