VLDB 2026 Research / reviewers in the wild / expert
Farima Fatahi Bayat
dblp:320/0224
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0002-8738-411XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 80% Trustworthy machine learning · 20% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › interpretability
faithful reasoning |
1.0 | 1 | 2026 | From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models · ACL (1) 2026 |
Natural language and speech › Language models and text generation
mathematical reasoning |
1.0 | 1 | 2026 | From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models · ACL (1) 2026 |
Natural language and speech › Language models and text generation › agentic language model
tool-augmented language models |
1.0 | 1 | 2026 | From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models · ACL (1) 2026 |
Natural language and speech › Language models and text generation › agentic language model › tool-augmented language models
tool-augmented reasoning |
1.0 | 1 | 2026 | From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models · ACL (1) 2026 |
Natural language and speech › Language models and text generation › evaluation of language models
factuality evaluation |
0.9 | 1 | 2025 | FactBench: A Dynamic Benchmark for In-the-Wild Language Model Factuality Evaluation · ACL (1) 2025 |
Natural language and speech › Language models and text generation
hallucination detection |
0.9 | 1 | 2025 | FactBench: A Dynamic Benchmark for In-the-Wild Language Model Factuality Evaluation · ACL (1) 2025 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | FactBench: A Dynamic Benchmark for In-the-Wild Language Model Factuality Evaluation · ACL (1) 2025 |
Natural language and speech › Language models and text generation
decoding |
0.8 | 1 | 2024 | Enhancing Language Model Factuality via Activation-Based Confidence Calibration and Guided Decoding · EMNLP 2024 |
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability
factuality |
0.8 | 1 | 2024 | Enhancing Language Model Factuality via Activation-Based Confidence Calibration and Guided Decoding · EMNLP 2024 |
Machine learning › Trustworthy machine learning
uncertainty and calibration |
0.8 | 1 | 2024 | Enhancing Language Model Factuality via Activation-Based Confidence Calibration and Guided Decoding · EMNLP 2024 |
Methods — techniques the papers use, named apart from their topics
preference optimization · 1.0benchmark construction · 1.0web retrieval · 0.9evidence-based evaluation · 0.9confidence-guided decoding · 0.8activation-based calibration · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language ModelsabstractTool-augmented Language Models (TaLMs) can invoke external tools to solve problems beyond their parametric capacity.However, it remains unclear whether these tool-enabled gains reflect trustworthy reasoning.Focusing on the Code Interpreter tool, we show that even when tools are selected and executed correctly, TaLMs treat tool outputs as substitutes for reasoning, producing solutions that appear correct but lack coherent justification.We term this failure mode Tool-Induced Myopia (TIM), and study it using PYMATH, a benchmark of 1,679 competition-level mathematical problems for which Python code is helpful but not sufficient.We further develop a multi-dimensional evaluation suite to quantify reasoning degradation in TaLMs relative to their non-tool counterparts.Our findings reveal that while TaLMs achieve up to a 19.3 percentage point gain in final-answer accuracy, their reasoning behavior consistently deteriorates (e.g., non-tool language models win up to 41.5% more often in pairwise comparisons of reasoning processes).This degradation intensifies with tool use; the more frequently a model invokes tools, the less coherent its reasoning becomes.Moreover, tool use shifts errors from arithmetic mistakes toward global reasoning failures (logic, assumption, creativity).Finally, we propose a preference-optimization-based framework that realigns TaLMs to use tool outputs as assistive evidence, improving both final-answer accuracy and the reasoning depth under tool use. 1 Farima Fatahi Bayat, Pouya Pezeshkpour, Estevam Hruschka |
ACL (1) | 1 |
| 2025 | FactBench: A Dynamic Benchmark for In-the-Wild Language Model Factuality EvaluationabstractThe rapid adoption of language models (LMs) across diverse applications has raised concerns about their factuality, i.e., their consistency with real-world facts.We introduce VERIFY, an evidence-based evaluation pipeline that measures LMs' factuality in real-world user interactions.VERIFY considers the verifiability of LM-generated content and categorizes content units as Supported, Unsupported, or Undecidable based on Web-retrieved evidence.Importantly, factuality judgment by VERIFY more strongly correlates with human evaluations than existing methods.Using VER-IFY, we identify "hallucination prompts," i.e., those that frequently elicit factual errors in LM responses.These prompts form FACTBENCH, a dataset of 1K prompts spanning 150 topics and tiered into Easy, Moderate, and Hard prompts.We benchmark widely-used openweight and proprietary LMs from six families, yielding three key findings: (i) LMs' factual precision declines from Easy to Hard prompts, (ii) factuality does not necessarily improve with scale; Llama3.1-405B-Instructperforms comparably to or worse than its 70B variant, and (iii) Gemini1.5-Proshows a notably higher refusal rate, with over-refusal in 25% of cases. Farima Fatahi Bayat, Sheza Munir, Lu Wang 0008 |
ACL (1) | 1 |
| 2024 | Enhancing Language Model Factuality via Activation-Based Confidence Calibration and Guided DecodingabstractCalibrating language models (LMs) aligns their generation confidence with the actual likelihood of answer correctness, which can inform users about LMs' reliability and mitigate hallucinated content.However, prior calibration methods, such as self-consistency-based and logit-based approaches, are either limited in inference-time efficiency or fall short of providing informative signals.Moreover, simply filtering out low-confidence responses reduces the LM's helpfulness when the answers are correct.Therefore, effectively using calibration techniques to enhance an LM's factuality remains an unsolved challenge.In this paper, we first propose an activation-based calibration method, ACTCAB, which trains a linear layer on top of the LM's last-layer activations that can better capture the representations of knowledge.Built on top of ACTCAB, we further propose CODEC, a confidence-guided decoding strategy to elicit truthful answers with high confidence from LMs.By evaluating on five popular QA benchmarks, ACTCAB achieves superior calibration performance than all competitive baselines, e.g., by reducing the average expected calibration error (ECE) score by up to 39%.Further experiments on CODEC show consistent improvements in several LMs' factuality on challenging QA datasets, such as TruthfulQA, highlighting the value of confidence signals in enhancing the factuality. 1 Farima Fatahi Bayat, Lu Wang 0008 |
EMNLP | 2 |
| 2022 | CompactIE: Compact Facts in Open Information ExtractionabstractA major drawback of modern neural OpenIE systems and benchmarks is that they prioritize high coverage of information in extractions over compactness of their constituents.This severely limits the usefulness of OpenIE extractions in many downstream tasks.The utility of extractions can be improved if extractions are compact and share constituents.To this end, we study the problem of identifying compact extractions with neural-based methods.We propose COMPACTIE, an OpenIE system that uses a novel pipelined approach to produce compact extractions with overlapping constituents.It first detects constituents of the extractions and then links them to build extractions.We train our system on compact extractions obtained by processing existing benchmarks.Our experiments on CaRB and Wire57 datasets indicate that COMPACTIE finds 1.5x-2x more compact extractions than previous systems, with high precision, establishing a new state-of-the-art performance in OpenIE. Farima Fatahi Bayat, Nikita Bhutani, H. V. Jagadish |
NAACL-HLT | 1 |