EDBT 2026 Demo / reviewers in the wild / expert
Faisal Ladhak
dblp:194/1214
· DBLP profile ↗
20ranked-venue papers
5as first author
15since 2021 · last 2025
0009-0009-5730-3866ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceabstractBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Griffin Thomas Adams, Jeremy Howard, Iacopo Poli. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Benjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Griffin Adams, Jeremy Howard, Iacopo Poli |
ACL (1) | 9 |
| 2025 | SWAN: An Efficient and Scalable Approach for Long-Context Language ModelingabstractKrishna C Puvvada, Faisal Ladhak, Santiago Akle Serano, Cheng-Ping Hsieh, Shantanu Acharya, Somshubra Majumdar, Fei Jia, Samuel Kriman, Simeng Sun, Dima Rekesh, Boris Ginsburg. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Krishna C. Puvvada, Faisal Ladhak, Santiago Akle Serano, Cheng-Ping Hsieh, Shantanu Acharya, Somshubra Majumdar, Fei Jia, Samuel Kriman, Simeng Sun, Dima Rekesh, Boris Ginsburg |
EMNLP | 2 |
| 2024 | STORYSUMM: Evaluating Faithfulness in Story SummarizationabstractHuman evaluation has been the gold standard for checking faithfulness in abstractive summarization.However, with a challenging source domain like narrative, multiple annotators can agree a summary is faithful, while missing details that are obvious errors only once pointed out.We therefore introduce a new dataset, STORYSUMM, comprising LLM summaries of short stories with localized faithfulness labels and error explanations.This benchmark is for evaluation methods, testing whether a given method can detect challenging inconsistencies.Using this dataset, we first show that any one human annotation protocol is likely to miss inconsistencies, and we advocate for pursuing a range of methods when establishing ground truth for a summarization dataset.We finally test recent automatic metrics and find that none of them achieve more than 70% balanced accuracy on this task, demonstrating that it is a challenging benchmark for future work in faithfulness evaluation. Melanie Subbiah, Faisal Ladhak, Akankshya Mishra, Griffin Adams, Lydia B. Chilton, Kathy McKeown |
EMNLP | 2 |
| 2024 | Proving Test Set Contamination in Black-Box Language ModelsabstractLarge language models are trained on vast amounts of internet data, prompting concerns that they have memorized public benchmarks. Detecting this type of contamination is challenging because the pretraining data used by proprietary models are often not publicly accessible.
We propose a procedure for detecting test set contamination of language models with exact false positive guarantees and without access to pretraining data or model weights. Our approach leverages the fact that when there is no data contamination, all orderings of an exchangeable benchmark should be equally likely. In contrast, the tendency for language models to memorize example order means that a contaminated language model will find certain canonical orderings to be much more likely than others. Our test flags potential contamination whenever the likelihood of a canonically ordered benchmark dataset is significantly higher than the likelihood after shuffling the examples.
We demonstrate that our procedure is sensitive enough to reliably detect contamination in challenging situations, including models as small as 1.4 billion parameters, on small test sets only 1000 examples, and datasets that appear only a few times in the pretraining corpus. Finally, we evaluate LLaMA-2 to apply our test in a realistic setting and find our results to be consistent with existing contamination evaluations. Yonatan Oren, Nicole Meister, Niladri S. Chatterji, Faisal Ladhak, Tatsunori B. Hashimoto |
ICLR | 4 |
| 2024 | Benchmarking Large Language Models for News SummarizationabstractAbstract Large language models (LLMs) have shown promise for automatic summarization but the reasons behind their successes are poorly understood. By conducting a human evaluation on ten LLMs across different pretraining methods, prompts, and model scales, we make two important observations. First, we find instruction tuning, not model size, is the key to the LLM’s zero-shot summarization capability. Second, existing studies have been limited by low-quality references, leading to underestimates of human performance and lower few-shot and finetuning performance. To better evaluate LLMs, we perform human evaluation over high-quality summaries we collect from freelance writers. Despite major stylistic differences such as the amount of paraphrasing, we find that LLM summaries are judged to be on par with human written summaries. Faisal Ladhak, Esin Durmus, Percy Liang, Kathy McKeown, Tatsunori B. Hashimoto |
Trans. Assoc. Comput. Linguistics | 2 |
| 2023 | Generating EDU Extracts for Plan-Guided Summary Re-RankingabstractTwo-step approaches, in which summary candidates are generated-then-reranked to return a single summary, can improve ROUGE scores over the standard single-step approach.Yet, standard decoding methods (i.e., beam search, nucleus sampling, and diverse beam search) produce candidates with redundant, and often low quality, content.In this paper, we design a novel method to generate candidates for re-ranking that addresses these issues.We ground each candidate abstract on its own unique content plan and generate distinct plan-guided abstracts using a model's top beam.More concretely, a standard language model (a BART LM) auto-regressively generates elemental discourse unit (EDU) content plans with an extractive copy mechanism.The top K beams from the content plan generator are then used to guide a separate LM, which produces a single abstractive candidate for each distinct plan.We apply an existing re-ranker (BRIO) to abstractive candidates generated from our method, as well as baseline decoding methods.We show large relevance improvements over previously published methods on widely used single document news article corpora, with ROUGE-2 F1 gains of 0.88, 2.01, and 0.38 on CNN / Dailymail, NYT, and Xsum, respectively.A human evaluation on CNN / DM validates these results.Similarly, on 1k samples from CNN / DM, we show that prompting GPT-3 to follow EDU plans outperforms sampling-based methods by 1.05 ROUGE-2 F1 points.Code to generate and realize plans is available at https: //github.com/griff4692/edu-sum. Griffin Adams, Alexander R. Fabbri, Faisal Ladhak, Noémie Elhadad, Kathy McKeown |
ACL (1) | 3 |
| 2023 | Contrastive Error Attribution for Finetuned Language ModelsabstractRecent work has identified noisy and misannotated data as a core cause of hallucinations and unfaithful outputs in Natural Language Generation (NLG) tasks.Consequently, identifying and removing these examples is a key open challenge in creating reliable NLG systems.In this work, we introduce a framework to identify and remove low-quality training instances that lead to undesirable outputs, such as faithfulness errors in text summarization.We show that existing approaches for error tracing, such as gradient-based influence measures, do not perform reliably for detecting faithfulness errors in NLG datasets.We overcome the drawbacks of existing error tracing methods through a new, contrast-based estimate that compares undesired generations to human-corrected outputs.Our proposed method can achieve a mean average precision of 0.93 at detecting known data errors across synthetic tasks with known ground truth, substantially outperforming existing approaches.Using this approach and re-training models on cleaned data leads to a 70% reduction in entity hallucinations on the NYT dataset and a 55% reduction in semantic errors on the E2E dataset. Faisal Ladhak, Esin Durmus, Tatsunori B. Hashimoto |
ACL (1) | 1 |
| 2023 | When Do Pre-Training Biases Propagate to Downstream Tasks? A Case Study in Text SummarizationabstractFaisal Ladhak, Esin Durmus, Mirac Suzgun, Tianyi Zhang, Dan Jurafsky, Kathleen McKeown, Tatsunori Hashimoto. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Faisal Ladhak, Esin Durmus, Mirac Suzgun, Daniel Jurafsky, Kathy McKeown, Tatsunori B. Hashimoto |
EACL | 1 |
| 2023 | Whose Opinions Do Language Models Reflect?abstractLanguage models (LMs) are increasingly being used in open-ended contexts, where the opinions they reflect in response to subjective queries can have a profound impact, both on user satisfaction, and shaping the views of society at large. We put forth a quantitative framework to investigate the opinions reflected by LMs – by leveraging high-quality public opinion polls. Using this framework, we create OpinionQA, a dataset for evaluating the alignment of LM opinions with those of 60 US demographic groups over topics ranging from abortion to automation. Across topics, we find substantial misalignment between the views reflected by current LMs and those of US demographic groups: on par with the Democrat-Republican divide on climate change. Notably, this misalignment persists even after explicitly steering the LMs towards particular groups. Our analysis not only confirms prior observations about the left-leaning tendencies of some human feedback-tuned LMs, but also surfaces groups whose opinions are poorly reflected by current LMs (e.g., 65+ and widowed individuals). Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, Tatsunori B. Hashimoto |
ICML | 3 |
| 2022 | Spurious Correlations in Reference-Free Evaluation of Text GenerationabstractModel-based, reference-free evaluation metrics have been proposed as a fast and cost-effective approach to evaluate Natural Language Generation (NLG) systems.Despite promising recent results, we find evidence that reference-free evaluation metrics of summarization and dialog generation may be relying on spurious correlations with measures such as word overlap, perplexity, and length.We further observe that for text summarization, these metrics have high error rates when ranking current state-ofthe-art abstractive summarization systems.We demonstrate that these errors can be mitigated by explicitly designing evaluation metrics to avoid spurious features in reference-free evaluation. Esin Durmus, Faisal Ladhak, Tatsunori B. Hashimoto |
ACL (1) | 2 |
| 2022 | Faithful or Extractive? On Mitigating the Faithfulness-Abstractiveness Trade-off in Abstractive SummarizationabstractDespite recent progress in abstractive summarization, systems still suffer from faithfulness errors.While prior work has proposed models that improve faithfulness, it is unclear whether the improvement comes from an increased level of extractiveness of the model outputs as one naive way to improve faithfulness is to make summarization models more extractive.In this work, we present a framework for evaluating the effective faithfulness of summarization systems, by generating a faithfulnessabstractiveness trade-off curve that serves as a control at different operating points on the abstractiveness spectrum.We then show that the baseline system as well as recently proposed methods for improving faithfulness, fail to consistently improve over the control at the same level of abstractiveness.Finally, we learn a selector to identify the most faithful and abstractive summary for a given document, and show that this system can attain higher faithfulness scores in human evaluations while being more abstractive than the baseline system on two datasets.Moreover, we show that our system is able to achieve a better faithfulnessabstractiveness trade-off than the control at the same level of abstractiveness. Faisal Ladhak, Esin Durmus, He He 0001, Claire Cardie, Kathy McKeown |
ACL (1) | 1 |
| 2022 | Constrained Regeneration for Cross-Lingual Query-Focused Extractive SummarizationabstractQuery-focused summaries of foreign-language, retrieved documents can help a user understand whether a document is actually relevant to the query term. A standard approach to this problem is to first translate the source documents and then perform extractive summarization to find relevant snippets. However, in a cross-lingual setting, the query term does not necessarily appear in the translations of relevant documents. In this work, we show that constrained machine translation and constrained post-editing can improve human relevance judgments by including a query term in a summary when its translation appears in the source document. We also present several strategies for selecting only certain documents for regeneration which yield further improvements Elsbeth Turcan, David Wan, Faisal Ladhak, Petra Galuscáková, Sukanta Sen, Svetlana Tchistiakova, Weijia Xu, Marine Carpuat, Kenneth Heafield, Douglas W. Oard, Kathy McKeown |
COLING | 3 |
| 2022 | ToKen: Task Decomposition and Knowledge Infusion for Few-Shot Hate Speech DetectionabstractBadr AlKhamissi, Faisal Ladhak, Srinivasan Iyer, Veselin Stoyanov, Zornitsa Kozareva, Xian Li, Pascale Fung, Lambert Mathias, Asli Celikyilmaz, Mona Diab. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Badr AlKhamissi, Faisal Ladhak, Srinivasan Iyer 0001, Veselin Stoyanov, Zornitsa Kozareva, Xian Li 0003, Pascale Fung, Lambert Mathias, Asli Celikyilmaz, Mona T. Diab |
EMNLP | 2 |
| 2022 | Improving Faithfulness by Augmenting Negative Summaries from Fake DocumentsabstractCurrent abstractive summarization systems tend to hallucinate content that is unfaithful to the source document, posing a risk of misinformation.To mitigate hallucination, we must teach the model to distinguish hallucinated summaries from faithful ones.However, the commonly used maximum likelihood training does not disentangle factual errors from other model errors.To address this issue, we propose a back-translation-style approach to augment negative samples that mimic factual errors made by the model.Specifically, we train an elaboration model that generates hallucinated documents given the reference summaries, and then generates negative summaries from the fake documents.We incorporate the negative samples into training through a controlled generator, which produces faithful/unfaithful summaries conditioned on the control codes.Additionally, we find that adding textual entailment data through multitasking further boosts the performance.Experiments on three datasets (XSum, GigaWord, and WikiHow) show that our method consistently improves faithfulness without sacrificing informativeness according to both human and automatic evaluation. 1 Faisal Ladhak, Esin Durmus, He He 0001 |
EMNLP | 2 |
| 2021 | Segmenting Subtitles for Correcting ASR Segmentation ErrorsabstractDavid Wan, Chris Kedzie, Faisal Ladhak, Elsbeth Turcan, Petra Galuscakova, Elena Zotkina, Zhengping Jiang, Peter Bell, Kathleen McKeown. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. David Wan, Chris Kedzie, Faisal Ladhak, Elsbeth Turcan, Petra Galuscáková, Elena Zotkina, Zhengping Jiang, Peter Bell 0001, Kathy McKeown |
EACL | 3 |
| 2020 | Exploring Content Selection in Summarization of Novel ChaptersabstractWe present a new summarization task, generating summaries of novel chapters using summary/chapter pairs from online study guides.This is a harder task than the news summarization task, given the chapter length as well as the extreme paraphrasing and generalization found in the summaries.We focus on extractive summarization, which requires the creation of a gold-standard set of extractive summaries.We present a new metric for aligning reference summary sentences with chapter sentences to create gold extracts and also experiment with different alignment methods.Our experiments demonstrate significant improvement over prior alignment approaches for our task as shown through automatic metrics and a crowd-sourced pyramid analysis. Faisal Ladhak, Bryan Li, Yaser Al-Onaizan, Kathy McKeown |
ACL | 1 |
| 2020 | To BERT or Not to BERT: Comparing Task-specific and Task-agnostic Semi-Supervised Approaches for Sequence TaggingabstractKasturi Bhattacharjee, Miguel Ballesteros, Rishita Anubhai, Smaranda Muresan, Jie Ma, Faisal Ladhak, Yaser Al-Onaizan. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Kasturi Bhattacharjee, Miguel Ballesteros, Rishita Anubhai, Smaranda Muresan, Jie Ma 0005, Faisal Ladhak, Yaser Al-Onaizan |
EMNLP (1) | 6 |
| 2019 | Determining Relative Argument Specificity and Stance for Complex Argumentative StructuresabstractSystems for automatic argument generation and debate require the ability to (1) determine the stance of any claims employed in the argument and (2) assess the specificity of each claim relative to the argument context.Existing work on understanding claim specificity and stance, however, has been limited to the study of argumentative structures that are relatively shallow, most often consisting of a single claim that directly supports or opposes the argument thesis.In this paper, we tackle these tasks in the context of complex arguments on a diverse set of topics.In particular, our dataset consists of manually curated argument trees for 741 controversial topics covering 95,312 unique claims; lines of argument are generally of depth 2 to 6.We find that as the distance between a pair of claims increases along the argument path, determining the relative specificity of a pair of claims becomes easier and determining their relative stance becomes harder. Esin Durmus, Faisal Ladhak, Claire Cardie |
ACL (1) | 2 |
| 2019 | The Role of Pragmatic and Discourse Context in Determining Argument ImpactabstractEsin Durmus, Faisal Ladhak, Claire Cardie. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Esin Durmus, Faisal Ladhak, Claire Cardie |
EMNLP/IJCNLP (1) | 2 |
| 2016 | LatticeRnn: Recurrent Neural Networks Over Lattices
Faisal Ladhak, Ankur Gandhe, Markus Dreyer, Lambert Mathias, Ariya Rastrow, Björn Hoffmeister |
INTERSPEECH | 1 |