VLDB 2026 Research / reviewers in the wild / expert
Houman Mehrafarin
dblp:317/0323
· DBLP profile ↗
1ranked-venue papers
1as first author
1since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Language models and text generation · 44% Knowledge representation and reasoning · 44% Question answering and dialogue systems · 13% |
Topics — the 2 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Knowledge representation and reasoning
logic-based reasoning |
0.8 | 1 | 2024 | Reasoning or a Semblance of it? A Diagnostic Study of Transitive Reasoning in LLMs · EMNLP 2024 |
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
compositional question answering |
0.2 | 1 | 2024 | Reasoning or a Semblance of it? A Diagnostic Study of Transitive Reasoning in LLMs · EMNLP 2024 |
Methods — techniques the papers use, named apart from their topics
diagnostic probing · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Reasoning or a Semblance of it? A Diagnostic Study of Transitive Reasoning in LLMsabstractEvaluating Large Language Models (LLMs) on reasoning benchmarks demonstrates their ability to solve compositional questions.However, little is known of whether these models engage in genuine logical reasoning or simply rely on implicit cues to generate answers.In this paper, we investigate the transitive reasoning capabilities of two distinct LLM architectures, LLaMA 2 and Flan-T5, by manipulating facts within two compositional datasets: QASC and Bamboogle.We controlled for potential cues that might influence the models' performance, including (a) word/phrase overlaps across sections of test input; (b) models' inherent knowledge during pre-training or fine-tuning; and (c) Named Entities.Our findings reveal that while both models leverage (a), Flan-T5 shows more resilience to experiments (b and c), having less variance than LLaMA 2. This suggests that models may develop an understanding of transitivity through fine-tuning on knowingly relevant datasets, a hypothesis we leave to future work 1 . Houman Mehrafarin, Arash Eshghi, Ioannis Konstas |
EMNLP | 1 |