Lanlan Ji

dblp:438/7588 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 50% Question answering and dialogue systems · 50%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational finance and economics · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
hallucination detection
0.912025
PHANTOM: A Benchmark for Hallucination Detection in Financial Long-Context QA · NeurIPS 2025
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
long-context question answering
0.912025
PHANTOM: A Benchmark for Hallucination Detection in Financial Long-Context QA · NeurIPS 2025
Computational finance and economics › financial data analysis
financial document analysis
0.312025
PHANTOM: A Benchmark for Hallucination Detection in Financial Long-Context QA · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

large language model fine-tuning · 1.7benchmark construction · 1.7
YearPublicationVenuePosition
2025 PHANTOM: A Benchmark for Hallucination Detection in Financial Long-Context QA
abstract
While Large Language Models (LLMs) show great promise, their tendencies to hallucinate pose significant risks in high-stakes domains like finance, especially when used for regulatory reporting and decision-making. Existing hallucination detection benchmarks fail to capture the complexities of financial benchmarks, which require high numerical precision, nuanced understanding of the language of finance, and ability to handle long-context documents. To address this, we introduce PHANTOM, a novel benchmark dataset for evaluating hallucination detection in long-context financial QA. Our approach first generates a seed dataset of high-quality "query-answer-document (chunk)" triplets, with either hallucinated or correct answers - that are validated by human annotators and subsequently expanded to capture various context lengths and information placements. We demonstrate how PHANTOM allows fair comparison of hallucination detection models and provides insights into LLM performance, offering a valuable resource for improving hallucination detection in financial applications. Further, our benchmarking results highlight the severe challenges out-of-the-box models face in detecting real-world hallucinations on long context data, and establish some promising directions towards alleviating these challenges, by fine-tuning open-source LLMs using PHANTOM.
Lanlan Ji, Dominic Seyler, Gunkirat Kaur, Manjunath Hegde, Koustuv Dasgupta, Bing Xiang
NeurIPS1