Xi Ye 0003

dblp:56/8549-3 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0003-0331-7946ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 9 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Beyond Single-shot Writing: Deep Research Agents are Unreliable at Multi-turn Report Revision
abstract
Existing benchmarks for Deep Research Agents (DRAs) treat report generation as a single-shot writing task, which fundamentally diverges from how human researchers iteratively draft and revise reports via self-reflection or peer feedback.Whether DRAs can reliably revise reports with user feedback remains unexplored.We introduce MR DRE, an evaluation suite that establishes multi-turn report revision as a new evaluation axis for DRAs.MR DRE consists of (1) a unified long-form report evaluation protocol spanning comprehensiveness, factuality, and presentation, and (2) a human-verified feedback simulation pipeline for multi-turn revision.Our analysis of five diverse DRAs reveals a critical limitation: while agents can address most user feedback, they regress on 16-27% of previously covered content and citation quality.Over multiple revision turns, even the best-performing agent leaves significant headroom, as they continue to disrupt content outside the feedback's scope and fail to preserve earlier edits.We also show that these issues are not easily resolvable through inference-time fixes such as prompt engineering and a dedicated sub-agent for revision 1 . Comprehensiveness
Bingsen Chen, Ping Nie, Yuyu Zhang, Xi Ye 0003, Chen Zhao 0013
ACL (1)5
2025 Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models Reasoning
abstract
Instruction Fine-Tuning (IFT) significantly enhances the zero-shot capabilities of pretrained Large Language Models (LLMs). While coding data is known to boost LLM reasoning abilities during pretraining, its role in activating internal reasoning capacities during IFT remains understudied. This paper investigates a key question: How does coding data impact LLMs' reasoning capacities during IFT stage? To explore this, we thoroughly examine the impact of coding data across different coding data proportions, model families, sizes, and reasoning domains, from various perspectives. Specifically, we create three IFT datasets with increasing coding data proportions, fine-tune six LLM backbones across different families and scales on these datasets, evaluate the tuned models' performance across twelve tasks in three reasoning domains, and analyze the outcomes from three broad-to-granular perspectives: overall, domain-level, and task-specific. Our holistic analysis provides valuable insights into each perspective. First, coding data tuning enhances the overall reasoning capabilities of LLMs across different model families and scales. Moreover, while the impact of coding data varies by domain, it shows consistent trends within each domain across different model families and scales. Additionally, coding data generally provides comparable task-specific benefits across model families, with optimal proportions in IFT datasets being task-dependent.
Xinlu Zhang, Zhiyu Chen 0002, Xi Ye 0003, Xianjun Yang, Lichang Chen, William Yang Wang, Linda R. Petzold
AAAI3
2025 Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking
abstract
Recent work has identified retrieval heads (Wu et al., 2025b), a subset of attention heads responsible for retrieving salient information in long-context language models (LMs), as measured by their copy-paste behavior in Needlein-a-Haystack tasks.In this paper, we introduce QRHEAD (Query-Focused Retrieval Head), an improved set of attention heads that enhance retrieval from long context.We identify QRHEAD by aggregating attention scores with respect to the input query, using a handful of examples from real-world tasks (e.g., long-context QA).We further introduce QR-RETRIEVER, an efficient and effective retriever that uses the accumulated attention mass of QRHEAD as retrieval scores.We use QR-RETRIEVER for long-context reasoning by selecting the most relevant parts with the highest retrieval scores.On multi-hop reasoning tasks LongMemEval and CLIPPER, this yields over 10% performance gains over full context and outperforms strong dense retrievers.We also evaluate QRRETRIEVER as a re-ranker on the BEIR benchmark and find that it achieves strong zero-shot performance, outperforming other LLM-based re-rankers such as RankGPT.Further analysis shows that both the querycontext attention scoring and task selection are crucial for identifying QRHEAD with strong downstream utility.Overall, our work contributes a general-purpose retriever and offers interpretability insights into the long-context capabilities of LMs. 1
Wuwei Zhang, Fangcong Yin, Howard Yen, Danqi Chen 0001, Xi Ye 0003
EMNLP5
2025 To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
abstract
Chain-of-thought (CoT) via prompting is the de facto method for eliciting reasoning capabilities from large language models (LLMs). But for what kinds of tasks is this extra "thinking" really helpful? To analyze this, we conducted a quantitative meta-analysis covering over 100 papers using CoT and ran our own evaluations of 20 datasets across 14 models. Our results show that CoT gives strong performance benefits primarily on tasks involving math or logic, with much smaller gains on other types of tasks. On MMLU, directly generating the answer without CoT leads to almost identical accuracy as CoT unless the question or model's response contains an equals sign, indicating symbolic operations and reasoning. Following this finding, we analyze the behavior of CoT on these problems by separating planning and execution and comparing against tool-augmented LLMs. Much of CoT's gain comes from improving symbolic execution, but it underperforms relative to using a symbolic solver. Our results indicate that CoT can be applied selectively, maintaining performance while saving inference costs. Furthermore, they suggest a need to move beyond prompt-based CoT to new paradigms that better leverage intermediate computation across the whole range of LLM applications.
Zayne Sprague, Fangcong Yin, Juan Diego Rodriguez, Dongwei Jiang, Manya Wadhwa, Prasann Singhal, Xi Ye 0003, Kyle Mahowald, Greg Durrett
ICLR8
2024 MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
abstract
While large language models (LLMs) equipped with techniques like chain-of-thought prompting have demonstrated impressive capabilities, they still fall short in their ability to reason robustly in complex settings. However, evaluating LLM reasoning is challenging because system capabilities continue to grow while benchmark datasets for tasks like logical deduction have remained static. We introduce MuSR, a dataset for evaluating language models on multistep soft reasoning tasks specified in a natural language narrative. This dataset has two crucial features. First, it is created through a novel neurosymbolic synthetic-to-natural generation algorithm, enabling the construction of complex reasoning instances that challenge GPT-4 (e.g., murder mysteries roughly 1000 words in length) and which can be scaled further as more capable LLMs are released. Second, our data instances are free text narratives corresponding to real-world domains of reasoning; this makes it simultaneously much more challenging than other synthetically-crafted benchmarks while remaining realistic and tractable for human annotators to solve with high accuracy. We evaluate a range of LLMs and prompting techniques on this dataset and characterize the gaps that remain for techniques like chain-of-thought to perform robust reasoning.
Zayne Sprague, Xi Ye 0003, Kaj Bostrom, Swarat Chaudhuri, Greg Durrett
ICLR2
2024 Effective Large Language Model Adaptation for Improved Grounding and Citation Generation
abstract
Xi Ye, Ruoxi Sun, Sercan Arik, Tomas Pfister. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Xi Ye 0003, Ruoxi Sun 0002, Sercan Ö. Arik, Tomas Pfister
NAACL-HLT1
2024 LoFiT: Localized Fine-tuning on LLM Representations
abstract
Recent work in interpretability shows that large language models (LLMs) can be adapted for new tasks in a learning-free way: it is possible to intervene on LLM representations to elicit desired behaviors for alignment. For instance, adding certain bias vectors to the outputs of certain attention heads is reported to boost the truthfulness of models. In this work, we show that localized fine-tuning serves as an effective alternative to such representation intervention methods. We introduce a framework called Localized Fine-Tuning on LLM Representations (LoFiT), which identifies a subset of attention heads that are most important for learning a specific task, then trains offset vectors to add to the model's hidden representations at those selected heads. LoFiT localizes to a sparse set of heads (3%-10%) and learns the offset vectors from limited training data, comparable to the settings used for representation intervention. For truthfulness and reasoning tasks, we find that LoFiT's intervention vectors are more effective for LLM adaptation than vectors from representation intervention methods such as Inference-time Intervention. We also find that the localization step is important: selecting a task-specific set of attention heads can lead to higher performance than intervening on heads selected for a different task. Finally, across 7 tasks we study, LoFiT achieves comparable performance to other parameter-efficient fine-tuning methods such as LoRA, despite modifying 20x-200x fewer parameters than these methods.
Fangcong Yin, Xi Ye 0003, Greg Durrett
NeurIPS2
2023 EEL: Efficiently Encoding Lattices for Reranking
abstract
Standard decoding approaches for conditional text generation tasks typically search for an output hypothesis with high model probability, but this may not yield the best hypothesis according to human judgments of quality.Reranking to optimize for downstream metrics can better optimize for quality, but many metrics of interest are computed with pre-trained language models, which are slow to apply to large numbers of hypotheses.We explore an approach for reranking hypotheses by using Transformers to efficiently encode lattices of generated outputs, a method we call EEL.With a single Transformer pass over the entire lattice, we can approximately compute a contextualized representation of each token as if it were only part of a single hypothesis in isolation.We combine this approach with a new class of token-factored rerankers (TFRs) that allow for efficient extraction of high reranker-scoring hypotheses from the lattice.Empirically, our approach incurs minimal degradation error compared to the exponentially slower approach of encoding each hypothesis individually.When applying EEL with TFRs across three text generation tasks, our results show both substantial speedup compared to naive reranking and often better performance on downstream metrics than comparable approaches.1
Prasann Singhal, Jiacheng Xu 0001, Xi Ye 0003, Greg Durrett
ACL (1)3
2023 Assessing Out-of-Domain Language Model Performance from Few Examples
abstract
While pretrained language models have exhibited impressive generalization capabilities, they still behave unpredictably under certain domain shifts.In particular, a model may learn a reasoning process on in-domain training data that does not hold for out-of-domain test data.We address the task of predicting outof-domain (OOD) performance in a few-shot fashion: given a few target-domain examples and a set of models with similar training performance, can we understand how these models will perform on OOD test data?We start from the baseline of looking at model accuracy on the few-shot examples, then investigate how to incorporate analysis of the models' behavior using feature attributions to improve our understanding of generalization.Specifically, we explore a set of "factors" designed to reveal model agreement with certain pathological heuristics that may indicate worse generalization capabilities.On textual entailment, paraphrase recognition, and a synthetic classification task, we show that attribution-based factors can help rank relative model OOD performance.However, accuracy on a few-shot test set is a surprisingly strong baseline, particularly when the system designer does not have in-depth prior knowledge about the domain shift.
Prasann Singhal, Jarad Forristal, Xi Ye 0003, Greg Durrett
EACL3
2023 Explanation Selection Using Unlabeled Data for Chain-of-Thought Prompting
abstract
Recent work has shown how to prompt large language models with explanations to obtain strong performance on textual reasoning tasks, i.e., the chain-of-thought paradigm.However, subtly different explanations can yield widely varying downstream task accuracy.Explanations that have not been "tuned" for a task, such as off-the-shelf explanations written by nonexperts, may lead to mediocre performance.This paper tackles the problem of how to optimize explanation-infused prompts in a blackbox fashion.We first generate sets of candidate explanations for each example in the prompt using a leave-one-out scheme, then find an effective combination of these explanations with a two-stage framework.We first evaluate explanations for each in-context example in isolation according to two proxy metrics, log likelihood and accuracy on new examples.Then, we search over combinations of explanations to find one that yields high performance against a silver-labeled development set.Across four textual reasoning tasks spanning question answering, mathematical reasoning, and natural language inference, results show that our proxy metrics correlate with ground truth accuracy and our overall method can effectively improve prompts over crowdworker annotations and naive search strategies. 1
Xi Ye 0003, Greg Durrett
EMNLP1
2023 SatLM: Satisfiability-Aided Language Models Using Declarative Prompting
abstract
Prior work has combined chain-of-thought prompting in large language models (LLMs) with programmatic representations to perform effective and transparent reasoning. While such an approach works well for tasks that only require forward reasoning (e.g., straightforward arithmetic), it is less effective for constraint solving problems that require more sophisticated planning and search. In this paper, we propose a new satisfiability-aided language modeling (SatLM) approach for improving the reasoning capabilities of LLMs. We use an LLM to generate a declarative task specification rather than an imperative program and leverage an off-the-shelf automated theorem prover to derive the final answer. This approach has two key advantages. The declarative specification is closer to the problem description than the reasoning steps are, so the LLM can parse it out of the description more accurately. Furthermore, by offloading the actual reasoning task to an automated theorem prover, our approach can guarantee the correctness of the answer with respect to the parsed specification and avoid planning errors in the solving process. We evaluate SATLM on 8 different datasets and show that it consistently outperforms program-aided LMs in the imperative paradigm. In particular, SATLM outperforms program-aided LMs by 23% on a challenging subset of the GSM arithmetic reasoning dataset; SATLM also achieves a new SoTA on LSAT and BoardgameQA, surpassing previous models that are trained on the respective training sets.
Xi Ye 0003, Qiaochu Chen, Isil Dillig, Greg Durrett
NeurIPS1
2022 Can Explanations Be Useful for Calibrating Black Box Models?
abstract
NLP practitioners often want to take existing trained models and apply them to data from new domains.While fine-tuning or few-shot learning can be used to adapt a base model, there is no single recipe for making these techniques work; moreover, one may not have access to the original model weights if it is deployed as a black box.We study how to improve a black box model's performance on a new domain by leveraging explanations of the model's behavior.Our approach first extracts a set of features combining human intuition about the task with model attributions generated by black box interpretation techniques, then uses a simple calibrator, in the form of a classifier, to predict whether the base model was correct or not.We experiment with our method on two tasks, extractive question answering and natural language inference, covering adaptation from several pairs of domains with limited target-domain data.The experimental results across all the domain pairs show that explanations are useful for calibrating these models, boosting accuracy when predictions do not have to be returned on every example.We further show that the calibration model transfers to some extent between tasks. 1
Xi Ye 0003, Greg Durrett
ACL (1)1
2022 RNG-KBQA: Generation Augmented Iterative Ranking for Knowledge Base Question Answering
abstract
Existing KBQA approaches, despite achieving strong performance on i.i.d.test data, often struggle in generalizing to questions involving unseen KB schema items.Prior rankingbased approaches have shown some success in generalization, but suffer from the coverage issue.We present RnG-KBQA, a Rank-and-Generate approach for KBQA, which remedies the coverage issue with a generation model while preserving a strong generalization capability.Our approach first uses a contrastive ranker to rank a set of candidate logical forms obtained by searching over the knowledge graph.It then introduces a tailored generation model conditioned on the question and the top-ranked candidates to compose the final logical form.We achieve new state-ofthe-art results on GRAILQA and WEBQSP datasets.In particular, our method surpasses the prior state-of-the-art by a large margin on the GRAILQA leaderboard.In addition, RnG-KBQA outperforms all prior approaches on the popular WEBQSP benchmark, even including the ones that use the oracle entity linking.The experimental results demonstrate the effectiveness of the interplay between ranking and generation, which leads to the superior performance of our proposed approach across all settings with especially strong improvements in zero-shot generalization. 1 * Work done during internship at Salesforce Research. 1 Code available at https://github.com/salesforce/rng-kbqa.
Xi Ye 0003, Semih Yavuz, Kazuma Hashimoto, Yingbo Zhou 0002, Caiming Xiong
ACL (1)1
2022 The Unreliability of Explanations in Few-shot Prompting for Textual Reasoning
abstract
Does prompting a large language model (LLM) like GPT-3 with explanations improve in-context learning? We study this question on two NLP tasks that involve reasoning over text, namely question answering and natural language inference. We test the performance of four LLMs on three textual reasoning datasets using prompts that include explanations in multiple different styles. For these tasks, we find that including explanations in the prompts for OPT, GPT-3 (davinci), and InstructGPT (text-davinci-001) only yields small to moderate accuracy improvements over standard few-show learning. However, text-davinci-002 is able to benefit more substantially.We further show that explanations generated by the LLMs may not entail the models’ predictions nor be factually grounded in the input, even on simple tasks with extractive explanations. However, these flawed explanations can still be useful as a way to verify LLMs’ predictions post-hoc. Through analysis in our three settings, we show that explanations judged by humans to be good—logically consistent with the input and the prediction—more likely cooccur with accurate predictions. Following these observations, we train calibrators using automatically extracted scores that assess the reliability of explanations, allowing us to improve performance post-hoc across all of our datasets.
Xi Ye 0003, Greg Durrett
NeurIPS1
2022 Diagnosing Ensemble Few-Shot Classifiers
abstract
The base learners and labeled samples (shots) in an ensemble few-shot classifier greatly affect the model performance. When the performance is not satisfactory, it is usually difficult to understand the underlying causes and make improvements. To tackle this issue, we propose a visual analysis method, FSLDiagnotor. Given a set of base learners and a collection of samples with a few shots, we consider two problems: 1) finding a subset of base learners that well predict the sample collections; and 2) replacing the low-quality shots with more representative ones to adequately represent the sample collections. We formulate both problems as sparse subset selection and develop two selection algorithms to recommend appropriate learners and shots, respectively. A matrix visualization and a scatterplot are combined to explain the recommended learners and shots in context and facilitate users in adjusting them. Based on the adjustment, the algorithm updates the recommendation results for another round of improvement. Two case studies are conducted to demonstrate that FSLDiagnotor helps build a few-shot classifier efficiently and increases the accuracy by 12% and 21%, respectively.
Weikai Yang, Xi Ye 0003, Xingxing Zhang 0001, Lanxi Xiao, Jiazhi Xia, Zhongyuan Wang 0006, Jun Zhu 0001, Hanspeter Pfister, Shixia Liu
IEEE Trans. Vis. Comput. Graph.2
2021 Connecting Attributions and QA Model Behavior on Realistic Counterfactuals
abstract
When a model attribution technique highlights a particular part of the input, a user might understand this highlight as making a statement about counterfactuals (Miller, 2019): if that part of the input were to change, the model's prediction might change as well.This paper investigates how well different attribution techniques align with this assumption on realistic counterfactuals in the case of reading comprehension (RC).RC is a particularly challenging test case, as token-level attributions that have been extensively studied in other NLP tasks such as sentiment analysis are less suitable to represent the reasoning that RC models perform.We construct counterfactual sets for three different RC settings, and through heuristics that can connect attribution methods' outputs to high-level model behavior, we can evaluate how useful different attribution methods and even different formats are for understanding counterfactuals.We find that pairwise attributions are better suited to RC than tokenlevel attributions across these different RC settings, with our best performance coming from a modification that we propose to an existing pairwise attribution method. 1
Xi Ye 0003, Rohan Nair, Greg Durrett
EMNLP (1)1
2020 Benchmarking Multimodal Regex Synthesis with Complex Structures
abstract
Existing datasets for regular expression (regex) generation from natural language are limited in complexity; compared to regex tasks that users post on StackOverflow, the regexes in these datasets are simple, and the language used to describe them is not diverse.We introduce STRUCTUREDREGEX, a new regex synthesis dataset differing from prior ones in three aspects.First, to obtain structurally complex and realistic regexes, we generate the regexes using a probabilistic grammar with pre-defined macros observed from real-world StackOverflow posts.Second, to obtain linguistically diverse natural language descriptions, we show crowdworkers abstract depictions of the underlying regex and ask them to describe the pattern they see, rather than having them paraphrase synthetic language.Third, we augment each regex example with a collection of strings that are and are not matched by the ground truth regex, similar to how real users give examples.Our quantitative and qualitative analysis demonstrates the advantages of STRUCTUREDREGEX over prior datasets.Further experimental results using various multimodal synthesis techniques highlight the challenge presented by our dataset, including non-local constraints and multi-modal inputs. 1 CatTemp concat( Comp , Comp , Comp ) reprange( Expr ,1,2) Literal reprange( Expr ,1,2) • • • • • • • • • one or two digits then "." and one or two digits SepTemp concat( Seg , Delimiter , Seg , Delimiter , Seg ) CatTemp IntTemp IntTemp Const Const • • • • • • • • • • • • • • • three delimited values, first will be an integer, with other two being either numeric or a string IntTemp and( Cons , Cons ) startwith( Expr ) endwith( Expr ) • • • • • • starts with "C0" and end with 4 digits Cons start[end]with(Expr) | not(start[end]with(Expr)) # must (not) start/end with contain(Expr) | not(contain(Expr)) # must (not) contain rep( ,k) | repatleast( ,k) |reprange( ,k,k) # length constraints AdvStartwithCons | AdvEndwithCons # adversative macro (e.g., start with capitals except A) CondContainCons # conditional macro.(e.g.letter, if contained, must be after a digit) Comp Literal | or(Literal,Literal,...) # literals like digits, letters, strings, or set of literals.rep(Expr,k) | repatleast(Expr,k) | reprange(Expr,k,k)# e.g, 3 digits, 2 -5 letter, etc. optional(Comp) # components can be optional.
Xi Ye 0003, Qiaochu Chen, Isil Dillig, Greg Durrett
ACL1
2020 Multi-modal synthesis of regular expressions
abstract
In this paper, we propose a multi-modal synthesis technique for automatically constructing regular expressions (regexes) from a combination of examples and natural language. Using multiple modalities is useful in this context because natural language alone is often highly ambiguous, whereas examples in isolation are often not sufficient for conveying user intent. Our proposed technique first parses the English description into a so-called hierarchical sketch that guides our programming-by-example (PBE) engine. Since the hierarchical sketch captures crucial hints, the PBE engine can leverage this information to both prioritize the search as well as make useful deductions for pruning the search space.
Qiaochu Chen, Xinyu Wang 0006, Xi Ye 0003, Greg Durrett, Isil Dillig
PLDI3
2020 Sketch-Driven Regular Expression Generation from Natural Language and Examples
abstract
Recent systems for converting natural language descriptions into regular expressions (regexes) have achieved some success, but typically deal with short, formulaic text and can only produce simple regexes. Real-world regexes are complex, hard to describe with brief sentences, and sometimes require examples to fully convey the user’s intent. We present a framework for regex synthesis in this setting where both natural language (NL) and examples are available. First, a semantic parser (either grammar-based or neural) maps the natural language description into an intermediate sketch, which is an incomplete regex containing holes to denote missing components. Then a program synthesizer searches over the regex space defined by the sketch and finds a regex that is consistent with the given string examples. Our semantic parser can be trained purely from weak supervision based on correctness of the synthesized regex, or it can leverage heuristically derived sketches. We evaluate on two prior datasets (Kushman and Barzilay 2013 ; Locascio et al. 2016 ) and a real-world dataset from Stack Overflow. Our system achieves state-of-the-art performance on the prior datasets and solves 57% of the real-world dataset, which existing neural systems completely fail on. 1
Xi Ye 0003, Qiaochu Chen, Xinyu Wang 0006, Isil Dillig, Greg Durrett
Trans. Assoc. Comput. Linguistics1