Nishant Balepur

dblp:346/4871 · DBLP profile ↗
← Back
12ranked-venue papers
10as first author
12since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 10 first-author · 12 since 2021
YearPublicationVenuePosition
2026 Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
abstract
Nishant Balepur, Malachi Hamada, Varsha Kishore, Sergey Feldman, Amanpreet Singh, Pao Siangliulue, Joseph Chee Chang, Eunsol Choi, Jordan Lee Boyd-Graber, Aakanksha Naik. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Nishant Balepur, Malachi Hamada, Varsha Kishore, Sergey Feldman, Amanpreet Singh, Pao Siangliulue, Joseph Chee Chang, Eunsol Choi, Jordan L. Boyd-Graber, Aakanksha Naik
ACL (1)1
2026 BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
abstract
Nishant Balepur, Bhavya Rajasekaran, Hyunjin Jane Oh, Michael Xie, Atrey Desai, Vipul Gupta, Steven James Moore, Eunsol Choi, Rachel Rudinger, Jordan Lee Boyd-Graber. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Nishant Balepur, Bhavya Rajasekaran, Hyunjin Jane Oh, Michael Xie, Atrey Desai, Steven James Moore, Eunsol Choi, Rachel Rudinger, Jordan L. Boyd-Graber
ACL (1)1
2026 Measuring User's Mental Models of Speech Translation in Human-AI Collaboration
abstract
Millions of people use machine translation (MT) tools daily, yet little is known about their perception of what systems can and cannot do.This paper studies users' mental models of speech translation systems through a new framework based on cross-lingual question answering, where users either accept MT output or request professional re-translation to answer questions based on the information presented in a foreign language.By analyzing user behavior and accuracy trends across varying translation qualities, we examine to what extent they can predict where the system is likely to be wrong, and how this mental model evolves.Users develop stronger mental models with practice, especially when they have some knowledge of the source language, primarily by relying on surface-level error cues.Moreover, providing speech transcriptions can help users develop better mental models.Our results show the promise of cross-lingual question answering as a downstream task for studying MT mental models, and advancing our understanding of human-AI collaboration.
HyoJung Han 0001, Nishant Balepur, Jordan L. Boyd-Graber, Marine Carpuat
ACL (1)2
2025 Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
abstract
Nishant Balepur, Vishakh Padmakumar, Fumeng Yang, Shi Feng, Rachel Rudinger, Jordan Lee Boyd-Graber. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Nishant Balepur, Vishakh Padmakumar, Fumeng Yang, Shi Feng 0005, Rachel Rudinger, Jordan L. Boyd-Graber
ACL (1)1
2025 Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
abstract
Multiple choice question answering (MCQA) is popular for LLM evaluation due to its simplicity and human-like testing, but we argue for its reform.We first reveal flaws in MCQA's format, as it struggles to: 1) test generation/subjectivity; 2) match LLM use cases; and 3) fully test knowledge.We instead advocate for generative formats based on human testing-where LLMs construct and explain answers-better capturing user needs and knowledge while remaining easy to score.We then show even when MCQA is a useful format, its datasets suffer from: leakage; unanswerability; shortcuts; and saturation.In each issue, we give fixes from education, like rubrics to guide MCQ writing; scoring methods to bridle guessing; and Item Response Theory to build harder MCQs.Lastly, we discuss LLM errors in MCQA-robustness, biases, and unfaithful explanations-showing how our prior solutions better measure or address these issues.While we do not need to desert MCQA, we encourage more efforts in refining the task based on educational testing, advancing evaluations. Q1. What's wrong with MCQA's format?A) It doesn't apply to many tasks ( §3.1) B) It's misaligned with LLM use cases ( §3.2) C) It doesn't fully test knowledge ( §3.3) Q2.What's wrong with MCQA datasets?A) Test sets are contaminated ( §5.1) B) They have unanswerable questions ( §5.2) C) They contain shortcuts ( §5.3) D) They're too easy for LLMs ( §5.4) Q3.How do LLMs struggle with MCQA?A) They lack robustness ( §6.1) B) They exhibit biases ( §6.2) C) They give unfaithful explanations ( §6.3) Q4.How can insights from education improve MCQA?A) Improve knowledge testing via generative formats ( §4) B) Combat test set leakage with fresh questions ( §5.1) C) Write MCQs informed by educational rubrics ( §5.2) D) Use calibration scoring to curb guessing ( §5.3.1)E) Find harder MCQs with item response theory ( §5.4.1)
Nishant Balepur, Rachel Rudinger, Jordan L. Boyd-Graber
ACL (1)1
2025 A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
abstract
Nishant Balepur, Matthew Shu, Yoo Yeon Sung, Seraphina Goldfarb-Tarrant, Shi Feng, Fumeng Yang, Rachel Rudinger, Jordan Lee Boyd-Graber. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Nishant Balepur, Matthew Shu, Yoo Yeon Sung, Seraphina Goldfarb-Tarrant, Shi Feng 0005, Fumeng Yang, Rachel Rudinger, Jordan L. Boyd-Graber
EMNLP1
2025 MoDS: Moderating a Mixture of Document Speakers to Summarize Debatable Queries in Document Collections
abstract
Nishant Balepur, Alexa Siu, Nedim Lipka, Franck Dernoncourt, Tong Sun, Jordan Lee Boyd-Graber, Puneet Mathur. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Nishant Balepur, Alexa F. Siu, Nedim Lipka, Franck Dernoncourt, Tong Sun 0005, Jordan L. Boyd-Graber, Puneet Mathur
NAACL (Long Papers)1
2024 Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
abstract
Multiple-choice question answering (MCQA) is often used to evaluate large language models (LLMs).To see if MCQA assesses LLMs as intended, we probe if LLMs can perform MCQA with choices-only prompts, where models must select the correct answer only from the choices.In three MCQA datasets and four LLMs, this prompt bests a majority baseline in 11/12 cases, with up to 0.33 accuracy gain.To help explain this behavior, we conduct an in-depth, black-box analysis on memorization, choice dynamics, and question inference.Our key findings are threefold.First, we find no evidence that the choices-only accuracy stems from memorization alone.Second, priors over individual choices do not fully explain choicesonly accuracy, hinting that LLMs use the group dynamics of choices.Third, LLMs have some ability to infer a relevant question from choices, and surprisingly can sometimes even match the original question.Inferring the original question is an impressive reasoning strategy, but it cannot fully explain the high choices-only accuracy of LLMs in MCQA.Thus, while LLMs are not fully incapable of reasoning in MCQA, we still advocate for the use of stronger baselines in MCQA benchmarks, the design of robust MCQA datasets for fair evaluations, and further efforts to explain LLM decision-making. 1Question: Which of these contains only a solution?Answer: (B) Question: Which can be considered a solution?Choices: (A) a jar of pickles (B) a bottle of juice (C) a bag of peanuts (D) a can of mixed fruit Answer: (B) Choices: (A) a jar of pickles (B) a bottle of juice (C) a bag of peanuts (D) a can of mixed fruit Question: Which can be considered a solution?Step 2: Answer the Question from Step 1 Classify Choice (A) Correctness Classify Choice (B) Correctness ... Abductive Question Inference ( §6) Question: Which of these contains only a solution?Choices: (A) a jar of pickles (B) a bottle of juice (C) a bag of peanuts (D) a can of mixed fruit Answer: (B) Choices: (A) a jar of pickles (B) a bottle of juice (C) a bag of peanuts (D) a can of mixed fruit Answer: (B) Full MCQA Prompt Choices-only Prompt ( §3) LLMs Can Perform MCQA with no Question, but how? Classify Choice (D) Correctness Question: Which of these contains only a solution?Choices: (A) \n (B) \n (C) \n (D) \n Answer: (B) Step 1: Guess the Question No Choices Empty Choices Choice: a can of mixed fruit Answer: False Choice: a bottle of juice Answer: True Choice: a jar of pickles Answer: False
Nishant Balepur, Abhilasha Ravichander, Rachel Rudinger
ACL (1)1
2024 A SMART Mnemonic Sounds like "Glue Tonic": Mixing LLMs with Student Feedback to Make Mnemonic Learning Stick
abstract
Nishant Balepur, Matthew Shu, Alexander Hoyle, Alison Robey, Shi Feng, Seraphina Goldfarb-Tarrant, Jordan Lee Boyd-Graber. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Nishant Balepur, Matthew Shu, Alexander Miserlis Hoyle, Alison Robey, Shi Feng 0005, Seraphina Goldfarb-Tarrant, Jordan L. Boyd-Graber
EMNLP1
2024 KARL: Knowledge-Aware Retrieval and Representations aid Retention and Learning in Students
abstract
Flashcard schedulers rely on 1) student models to predict the flashcards a student knows; and 2) teaching policies to pick which cards to show next via these predictions.Prior student models, however, just use study data like the student's past responses, ignoring the text on cards.We propose content-aware scheduling, the first schedulers exploiting flashcard content.To give the first evidence that such schedulers enhance student learning, we build KAR 3 L, a simple but effective content-aware student model employing deep knowledge tracing (DKT), retrieval, and BERT to predict student recall.We train KAR 3 L by collecting a new dataset of 123,143 study logs on diverse trivia questions.KAR 3 L bests existing student models in AUC and calibration error.To ensure our improved predictions lead to better student learning, we create a novel delta-based teaching policy to deploy KAR 3 L online.Based on 32 study paths from 27 users, KAR 3 L improves learning efficiency over SOTA, showing KAR 3 L's strength and encouraging researchers to look beyond historical study data to fully capture student abilities.1 * Equal contribution.
Matthew Shu, Nishant Balepur, Shi Feng 0005, Jordan L. Boyd-Graber
EMNLP2
2023 Text Fact Transfer
abstract
Text style transfer is a prominent task that aims to control the style of text without inherently changing its factual content.To cover more text modification applications, such as adapting past news for current events and repurposing educational materials, we propose the task of text fact transfer, which seeks to transfer the factual content of a source text between topics without modifying its style.We find that existing language models struggle with text fact transfer, due to their inability to preserve the specificity and phrasing of the source text, and tendency to hallucinate errors.To address these issues, we design ModQGA, a framework that minimally modifies a source text with a novel combination of end-to-end question generation and specificity-aware question answering.Through experiments on four existing datasets adapted for text fact transfer, we show that ModQGA can accurately transfer factual content without sacrificing the style of the source text. 1
Nishant Balepur, Jie Huang 0009, Kevin Chen-Chuan Chang
EMNLP1
2023 Expository Text Generation: Imitate, Retrieve, Paraphrase
abstract
Expository documents are vital resources for conveying complex information to readers.Despite their usefulness, writing expository text by hand is a challenging process that requires careful content planning, obtaining facts from multiple sources, and the ability to clearly synthesize these facts.To ease these burdens, we propose the task of expository text generation, which seeks to automatically generate an accurate and stylistically consistent expository text for a topic by intelligently searching a knowledge source.We solve our task by developing IRP, a framework that overcomes the limitations of retrieval-augmented models and iteratively performs content planning, fact retrieval, and rephrasing.Through experiments on three diverse, newly-collected datasets, we show that IRP produces factual and organized expository texts that accurately inform readers. 1 1 Code is available at https://github.com/ nbalepur/expository-text-generation. Ground TruthUniversity of Denver is a private institution that was founded in 1864.It has a total undergraduate enrollment of 5,867 (fall 2021), its setting is city, and the campus size is 125 acres... University of Montana is a public institution that was founded in 1893.It has a total undergraduate enrollment of 7,223 (fall 2021), its setting is urban, and the campus size is 220 acres... RAGUniversity of Denver is a public institution founded in 1891.It has a total of 5,867 students (fall 2021), its location is urban, and the campus covers 120 acres... Our Model (IRP)University of Denver is a private institution founded in 1864.It has a total of 5,867 students (fall 2021), it is located in the city, and the campus covers 126 acres... LLaMA+RetrUniversity of Denver is a private institution founded in 1864.It has a total of 11,482 students (fall 2021), its location is urban, and the campus covers 125 acres...
Nishant Balepur, Jie Huang 0009, Kevin Chen-Chuan Chang
EMNLP1