Xingdi Yuan

dblp:40/10147 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-7660-0059ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
YearPublicationVenuePosition
2025 Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs
abstract
We observe a novel phenomenon, contextual entrainment, across a wide range of language models (LMs) and prompt settings, providing a new mechanistic perspective on how LMs become distracted by “irrelevant” contextual information in the input prompt. Specifically, LMs assign significantly higher logits (or probabilities) to any tokens that have previously appeared in the context prompt, even for random tokens. This suggests that contextual entrainment is a mechanistic phenomenon, occurring independently of the relevance or semantic relation of the tokens to the question or the rest of the sentence. We find statistically significant evidence that the magnitude of contextual entrainment is influenced by semantic factors. Counterfactual prompts have a greater effect compared to factual ones, suggesting that while contextual entrainment is a mechanistic phenomenon, it is modulated by semantic factors.We hypothesise that there is a circuit of attention heads — the entrainment heads — that corresponds to the contextual entrainment phenomenon. Using a novel entrainment head discovery method based on differentiable masking, we identify these heads across various settings. When we “turn off” these heads, i.e., set their outputs to zero, the effect of contextual entrainment is significantly attenuated, causing the model to generate output that capitulates to what it would produce if no distracting context were provided. Our discovery of contextual entrainment, along with our investigation into LM distraction via the entrainment heads, marks a key step towards the mechanistic analysis and mitigation of the distraction problem.
Jingcheng Niu, Xingdi Yuan, Hamidreza Saghir, Amir H. Abdi
ACL (1)2
2024 OPEx: A Component-Wise Analysis of LLM-Centric Agents in Embodied Instruction Following
abstract
Embodied Instruction Following (EIF) is a crucial task in embodied learning, requiring agents to interact with their environment through egocentric observations to fulfill natural language instructions.Recent advancements have seen a surge in employing large language models (LLMs) within a framework-centric approach to enhance performance in embodied learning tasks, including EIF.Despite these efforts, there exists a lack of a unified understanding regarding the impact of various components-ranging from visual perception to action execution-on task performance.To address this gap, we introduce OPEx, a comprehensive framework that delineates the core components essential for solving embodied learning tasks: Observer, Planner, and Executor.Through extensive evaluations, we provide a deep analysis of how each component influences EIF task performance.Furthermore, we innovate within this space by deploying a multi-agent LLM communication strategy on a TextWorld counterpart, further enhancing task performance.Our findings reveal that LLM-centric design markedly improves EIF outcomes, identify visual perception and low-level action execution as critical bottlenecks, and demonstrate that augmenting LLMs with a multi-agent framework further elevates performance.1
Xingdi Yuan, Marc-Alexandre Côté, Bang Liu 0003
ACL (1)3
2024 Language-guided Skill Learning with Temporal Variational Inference
abstract
We present an algorithm for skill discovery from expert demonstrations. The algorithm first utilizes Large Language Models (LLMs) to propose an initial segmentation of the trajectories. Following that, a hierarchical variational inference framework incorporates the LLM-generated segmentation information to discover reusable skills by merging trajectory segments. To further control the trade-off between compression and reusability, we introduce a novel auxiliary objective based on the Minimum Description Length principle that helps guide this skill discovery process. Our results demonstrate that agents equipped with our method are able to discover skills that help accelerate learning and outperform baseline skill learning approaches on new long-horizon tasks in BabyAI, a grid world navigation environment, as well as ALFRED, a household simulation environment.
Haotian Fu, Pratyusha Sharma, Elias Stengel-Eskin, George Dimitri Konidaris, Nicolas Le Roux, Marc-Alexandre Côté, Xingdi Yuan
ICML7
2024 Think Before You Act: Decision Transformers with Working Memory
abstract
Decision Transformer-based decision-making agents have shown the ability to generalize across multiple tasks. However, their performance relies on massive data and computation. We argue that this inefficiency stems from the forgetting phenomenon, in which a model memorizes its behaviors in parameters throughout training. As a result, training on a new task may deteriorate the model’s performance on previous tasks. In contrast to LLMs’ implicit memory mechanism, the human brain utilizes distributed memory storage, which helps manage and organize multiple skills efficiently, mitigating the forgetting phenomenon. Inspired by this, we propose a working memory module to store, blend, and retrieve information for different downstream tasks. Evaluation results show that the proposed method improves training efficiency and generalization in Atari games and Meta-World object manipulation tasks. Moreover, we demonstrate that memory fine-tuning further enhances the adaptability of the proposed architecture.
Jikun Kang, Romain Laroche, Xingdi Yuan, Adam Trischler, Xue (Steve) Liu, Jie Fu 0001
ICML3
2024 Policy Improvement using Language Feedback Models
abstract
We introduce Language Feedback Models (LFMs) that identify desirable behaviour --- actions that help achieve tasks specified in the instruction - for imitation learning in instruction following. To train LFMs, we obtain feedback from Large Language Models (LLMs) on visual trajectories verbalized to language descriptions. First, by using LFMs to identify desirable behaviour to imitate, we improve in task-completion rate over strong behavioural cloning baselines on three distinct language grounding environments (Touchdown, ScienceWorld, and ALFWorld). Second, LFMs outperform using LLMs as experts to directly predict actions, when controlling for the number of LLM output tokens. Third, LFMs generalize to unseen environments, improving task-completion rate by 3.5-12.0% through one round of adaptation. Finally, LFMs can be modified to provide human-interpretable feedback without performance loss, allowing human verification of desirable behaviour for imitation learning.
Victor Zhong, Dipendra Misra, Xingdi Yuan, Marc-Alexandre Côté
NeurIPS3
2023 ByteSized32: A Corpus and Challenge Task for Generating Task-Specific World Models Expressed as Text Games
abstract
In this work we investigate the capacity of language models to generate explicit, interpretable, and interactive world models of scientific and common-sense reasoning tasks.We operationalize this as a task of generating text games, expressed as hundreds of lines of PYTHON code.To facilitate this task, we introduce BYTESIZED32 1 , a corpus of 32 reasoning-focused text games totalling 20k lines of PYTHON code.We empirically demonstrate that GPT-4 can use these games as templates for single-shot in-context learning, successfully producing runnable games on unseen topics in 28% of cases.When allowed to selfreflect on program errors, game runnability substantially increases to 57%.While evaluating simulation fidelity is labor intensive, we introduce a suite of automated metrics to assess game fidelity, technical validity, adherence to task specifications, and winnability, showing a high-degree of agreement with expert human ratings.We pose this as a challenge task to spur further development at the juncture of world modeling and code generation.
Ruoyao Wang, Graham Todd, Xingdi Yuan, Ziang Xiao, Marc-Alexandre Côté, Peter A. Jansen
EMNLP3
2022 Asking for Knowledge (AFK): Training RL Agents to Query External Knowledge Using Language
abstract
To solve difficult tasks, humans ask questions to acquire knowledge from external sources. In contrast, classical reinforcement learning agents lack such an ability and often resort to exploratory behavior. This is exacerbated as few present-day environments support querying for knowledge. In order to study how agents can be taught to query external knowledge via language, we first introduce two new environments: the grid-world-based Q-BabyAI and the text-based Q-TextWorld. In addition to physical interactions, an agent can query an external knowledge source specialized for these environments to gather information. Second, we propose the ‘Asking for Knowledge’ (AFK) agent, which learns to generate language commands to query for meaningful knowledge that helps solve the tasks. AFK leverages a non-parametric memory, a pointer mechanism and an episodic exploration bonus to tackle (1) irrelevant information, (2) a large query language space, (3) delayed reward for making meaningful queries. Extensive experiments demonstrate that the AFK agent outperforms recent baselines on the challenging Q-BabyAI and Q-TextWorld environments.
Iou-Jen Liu, Xingdi Yuan, Marc-Alexandre Côté, Pierre-Yves Oudeyer, Alexander G. Schwing
ICML2
2021 Interactive Machine Comprehension with Dynamic Knowledge Graphs
abstract
Interactive machine reading comprehension (iMRC) is machine comprehension tasks where knowledge sources are partially observable.An agent must interact with an environment sequentially to gather necessary knowledge in order to answer a question.We hypothesize that graph representations are good inductive biases, which can serve as an agent's memory mechanism in iMRC tasks.We explore four different categories of graphs that can capture text information at various levels.We describe methods that dynamically build and update these graphs during information gathering, as well as neural models to encode graph representations in RL agents.Extensive experiments on iSQuAD suggest that graph representations can result in significant performance improvements for RL agents.1
Xingdi Yuan
EMNLP (1)1
2021 ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, Matthew J. Hausknecht
ICLR2
2021 An Empirical Study on Neural Keyphrase Generation
abstract
Rui Meng, Xingdi Yuan, Tong Wang, Sanqiang Zhao, Adam Trischler, Daqing He. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Xingdi Yuan, Tong Wang 0012, Sanqiang Zhao, Adam Trischler, Daqing He
NAACL-HLT2
2020 Interactive Fiction Games: A Colossal Adventure
abstract
A hallmark of human intelligence is the ability to understand and communicate with language. Interactive Fiction games are fully text-based simulation environments where a player issues text commands to effect change in the environment and progress through the story. We argue that IF games are an excellent testbed for studying language-based autonomous agents. In particular, IF games combine challenges of combinatorial action spaces, language understanding, and commonsense reasoning. To facilitate rapid development of language-based agents, we introduce Jericho, a learning environment for man-made IF games and conduct a comprehensive study of text-agents across a rich set of games, highlighting directions in which agents can improve.
Matthew J. Hausknecht, Prithviraj Ammanabrolu, Marc-Alexandre Côté, Xingdi Yuan
AAAI4
2020 Interactive Machine Comprehension with Information Seeking Agents
abstract
Existing machine reading comprehension (MRC) models do not scale effectively to realworld applications like web-level information retrieval and question answering (QA).We argue that this stems from the nature of MRC datasets: most of these are static environments wherein the supporting documents and all necessary information are fully observed.In this paper, we propose a simple method that reframes existing MRC datasets as interactive, partially observable environments.Specifically, we "occlude" the majority of a document's text and add context-sensitive commands that reveal "glimpses" of the hidden text to a model.We repurpose SQuAD and NewsQA as an initial case study, and then show how the interactive corpora can be used to train a model that seeks relevant information through sequential decision making.We believe that this setting can contribute in scaling models to web-level QA scenarios.1
Xingdi Yuan, Jie Fu 0001, Marc-Alexandre Côté, Yi Tay, Christopher Joseph Pal, Adam Trischler
ACL1
2020 One Size Does Not Fit All: Generating and Evaluating Variable Number of Keyphrases
abstract
Different texts shall by nature correspond to different number of keyphrases.This desideratum is largely missing from existing neural keyphrase generation models.In this study, we address this problem from both modeling and evaluation perspectives.We first propose a recurrent generative model that generates multiple keyphrases as delimiter-separated sequences.Generation diversity is further enhanced with two novel techniques by manipulating decoder hidden states.In contrast to previous approaches, our model is capable of generating diverse keyphrases and controlling number of outputs.We further propose two evaluation metrics tailored towards the variable-number generation.We also introduce a new dataset (ST A C KEX) that expands beyond the only existing genre (i.e., academic writing) in keyphrase generation tasks.With both previous and new evaluation metrics, our model outperforms strong baselines on all datasets.
Xingdi Yuan, Tong Wang 0012, Khushboo Thaker, Peter Brusilovsky, Daqing He, Adam Trischler
ACL1
2020 Learning Dynamic Belief Graphs to Generalize on Text-Based Games
abstract
Playing text-based games requires skills in processing natural language and sequential decision making. Achieving human-level performance on text-based games remains an open challenge, and prior research has largely relied on hand-crafted structured representations and heuristics. In this work, we investigate how an agent can plan and generalize in text-based games using graph-structured representations learned end-to-end from raw text. We propose a novel graph-aided transformer agent (GATA) that infers and updates latent belief graphs during planning to enable effective action selection by capturing the underlying game dynamics. GATA is trained using a combination of reinforcement and self-supervised learning. Our work demonstrates that the learned graph-based representations help agents converge to better policies than their text-only counterparts and facilitate effective generalization across game configurations. Experiments on 500+ unique games from the TextWorld suite show that our best agent outperforms text-based baselines by an average of 24.2%.
Ashutosh Adhikari, Xingdi Yuan, Marc-Alexandre Côté, Mikulas Zelinka, Marc-Antoine Rondeau, Romain Laroche, Pascal Poupart, Jian Tang 0005, Adam Trischler, William L. Hamilton
NeurIPS2
2020 Graph Policy Network for Transferable Active Learning on Graphs
abstract
Graph neural networks (GNNs) have been attracting increasing popularity due to their simplicity and effectiveness in a variety of fields. However, a large number of labeled data is generally required to train these networks, which could be very expensive to obtain in some domains. In this paper, we study active learning for GNNs, i.e., how to efficiently label the nodes on a graph to reduce the annotation cost of training GNNs. We formulate the problem as a sequential decision process on graphs and train a GNN-based policy network with reinforcement learning to learn the optimal query strategy. By jointly training on several source graphs with full labels, we learn a transferable active learning policy which can directly generalize to unlabeled target graphs. Experimental results on multiple datasets from different domains prove the effectiveness of the learned policy in promoting active learning performance in both settings of transferring between graphs in the same domain and across different domains.
Shengding Hu, Zheng Xiong, Meng Qu, Xingdi Yuan, Marc-Alexandre Côté, Zhiyuan Liu 0001, Jian Tang 0005
NeurIPS4
2019 Simple and Effective Curriculum Pointer-Generator Networks for Reading Comprehension over Long Narratives
abstract
Yi Tay, Shuohang Wang, Anh Tuan Luu, Jie Fu, Minh C. Phan, Xingdi Yuan, Jinfeng Rao, Siu Cheung Hui, Aston Zhang. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Yi Tay, Shuohang Wang, Anh Tuan Luu, Jie Fu 0001, Minh C. Phan, Xingdi Yuan, Jinfeng Rao, Siu Cheung Hui, Aston Zhang
ACL (1)6
2019 Interactive Language Learning by Question Answering
abstract
Xingdi Yuan, Marc-Alexandre Côté, Jie Fu, Zhouhan Lin, Chris Pal, Yoshua Bengio, Adam Trischler. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xingdi Yuan, Marc-Alexandre Côté, Jie Fu 0001, Zhouhan Lin, Christopher Joseph Pal, Yoshua Bengio, Adam Trischler
EMNLP/IJCNLP (1)1
2019 Building Dynamic Knowledge Graphs from Text using Machine Reading Comprehension
Rajarshi Das, Tsendsuren Munkhdalai, Xingdi Yuan, Adam Trischler, Andrew McCallum
ICLR (Poster)3
2018 Rapid Adaptation with Conditionally Shifted Neurons
abstract
We describe a mechanism by which artificial neural networks can learn rapid adaptation - the ability to adapt on the fly, with little data, to new tasks - that we call conditionally shifted neurons. We apply this mechanism in the framework of metalearning, where the aim is to replicate some of the flexibility of human learning in machines. Conditionally shifted neurons modify their activation values with task-specific shifts retrieved from a memory module, which is populated rapidly based on limited task experience. On metalearning benchmarks from the vision and language domains, models augmented with conditionally shifted neurons achieve state-of-the-art results.
Tsendsuren Munkhdalai, Xingdi Yuan, Soroush Mehri, Adam Trischler
ICML2
2016 A Parallel-Hierarchical Model for Machine Comprehension on Sparse Data
abstract
Understanding unstructured text is a major goal within natural language processing.Comprehension tests pose questions based on short text passages to evaluate such understanding.In this work, we investigate machine comprehension on the challenging MCTest benchmark.Partly because of its limited size, prior work on MCTest has focused mainly on engineering better features.We tackle the dataset with a neural approach, harnessing simple neural networks arranged in a parallel hierarchy.The parallel hierarchy enables our model to compare the passage, question, and answer from a variety of trainable perspectives, as opposed to using a manually designed, rigid feature set.Perspectives range from the word level to sentence fragments to sequences of sentences; the networks operate only on word-embedding representations of text.When trained with a methodology designed to help cope with limited training data, our Parallel-Hierarchical model sets a new state of the art for MCTest, outperforming previous feature-engineered approaches slightly and previous neural approaches by a significant margin (over 15 percentage points).* A. Trischler and Z. Ye contributed equally to this work.
Adam Trischler, Xingdi Yuan, Philip Bachman
ACL (1)3
2016 Natural Language Comprehension with the EpiReader
abstract
We present EpiReader, a novel model for machine comprehension of text.Machine comprehension of unstructured, real-world text is a major research goal for natural language processing.Current tests of machine comprehension pose questions whose answers can be inferred from some supporting text, and evaluate a model's response to the questions.EpiReader is an end-to-end neural model comprising two components: the first component proposes a small set of candidate answers after comparing a question to its supporting text, and the second component formulates hypotheses using the proposed candidates and the question, then reranks the hypotheses based on their estimated concordance with the supporting text.We present experiments demonstrating that EpiReader sets a new state-of-the-art on the CNN and Children's Book Test benchmarks, outperforming previous neural models by a significant margin.
Adam Trischler, Xingdi Yuan, Philip Bachman, Alessandro Sordoni, Kaheer Suleman
EMNLP3
2014 Semantic aware sport image resizing jointly using seam carving and warping
Lifang Wu, Xingdi Yuan, Xiuzhen Zhang 0001, Lianchao Cao
Multim. Tools Appl.3