Patrick Xia 0002

dblp:128/4897-2 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Generative Adapter: Contextualizing Language Models in Parameters with A Single Forward Pass
abstract
Large language models (LLMs) acquire substantial knowledge during pretraining but often need adaptation to new contexts, tasks, or domains, typically achieved through fine-tuning or prompting. However, fine-tuning incurs significant training costs, while prompting increases inference overhead. Inspired by fast weight memory, we introduce GenerativeAdapter, an effective and efficient adaptation method that encode test-time context into language model parameters with a single forward pass. GenerativeAdapter augments a frozen pretrained LM with a lightweight adapter generator, trained via self-supervised learning, to produce parameter-efficient adapters. Notably, our generator is general-purpose, i.e., one generator can adapt the corresponding base model for all langauge processing scenarios. We apply GenerativeAdapter to two pretrained LMs (Mistral-7B-Instruct and Llama2-7B-Chat) and evaluate the adapted models across knowledge acquisition from documents, learning from demonstrations, and personalization for users. In StreamingQA, our approach is effective in injecting knowledge into the LM's parameters, achieving a 63.5\% improvement in F1 score over the model with supervised fine-tuning (from $19.5$ to $31.5$) for contexts as long as 32K tokens. In the MetaICL in-context learning evaluation, our method achieves an average accuracy of $44.9$ across 26 tasks, outperforming the base model. On MSC, our method proves to be highly competitive in memorizing user information from conversations with a 4x reduction in computation and memory costs compared to prompting with full conversation history. Overall, GenerativeAdapter provides a viable solution for adapting large LMs to evolving information and providing tailored user experience, while reducing training and inference costs relative to traditional fine-tuning and prompting techniques.
Tong Chen 0005, Hao Fang 0002, Patrick Xia 0002, Xiaodong Liu 0003, Benjamin Van Durme, Luke Zettlemoyer, Jianfeng Gao 0001, Hao Cheng 0002
ICLR3
2025 Multi-Field Adaptive Retrieval
abstract
Document retrieval for tasks such as search and retrieval-augmented generation typically involves datasets that are _unstructured_: free-form text without explicit internal structure in each document. However, documents can have some structure, containing fields such as an article title, a message body, or an HTML header. To address this gap, we introduce Multi-Field Adaptive Retrieval (mFAR), a flexible framework that accommodates any number and any type of document indices on _semi-structured_ data. Our framework consists of two main steps: (1) the decomposition of an existing document into fields, each indexed independently through dense and lexical methods, and (2) learning a model which adaptively predicts the importance of a field by conditioning on the document query, allowing on-the-fly weighting of the most likely field(s). We find that our approach allows for the optimized use of dense versus lexical representations across field types, significantly improves in document ranking over a number of existing retrievers, and achieves state-of-the-art performance for multi-field structured data.
Millicent Li, Tongfei Chen, Benjamin Van Durme, Patrick Xia 0002
ICLR4
2024 Language-to-Code Translation with a Single Labeled Example
abstract
Kaj Bostrom, Harsh Jhamtani, Hao Fang, Sam Thomson, Richard Shin, Patrick Xia, Benjamin Van Durme, Jason Eisner, Jacob Andreas. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Kaj Bostrom, Harsh Jhamtani, Hao Fang 0002, Sam Thomson, Richard Shin, Patrick Xia 0002, Benjamin Van Durme, Jason Eisner, Jacob Andreas
EMNLP6
2024 Learning to Retrieve Iteratively for In-Context Learning
abstract
We introduce iterative retrieval, a novel framework that empowers retrievers to make iterative decisions through policy optimization.Finding an optimal portfolio of retrieved items is a combinatorial optimization problem, generally considered NP-hard.This approach provides a learned approximation to such a solution, meeting specific task requirements under a given family of large language models (LLMs).We propose a training procedure based on reinforcement learning, incorporating feedback from LLMs.We instantiate an iterative retriever for composing in-context learning (ICL) exemplars and apply it to various semantic parsing tasks that demand synthesized programs as outputs.By adding only 4M additional parameters for state encoding, we convert an offthe-shelf dense retriever into a stateful iterative retriever, outperforming previous methods in selecting ICL exemplars on semantic parsing datasets such as SMCALFLOW, TREEDST, and MTOP.Additionally, the trained iterative retriever generalizes across different inference LLMs beyond the one used during training.
Yunmo Chen, Tongfei Chen, Harsh Jhamtani, Patrick Xia 0002, Richard Shin, Jason Eisner, Benjamin Van Durme
EMNLP4
2024 Natural Language Decomposition and Interpretation of Complex Utterances
Harsh Jhamtani, Hao Fang 0002, Patrick Xia 0002, Eran Levy, Jacob Andreas, Benjamin Van Durme
IJCAI3
2023 Multilingual Coreference Resolution in Multiparty Dialogue
abstract
Abstract Existing multiparty dialogue datasets for entity coreference resolution are nascent, and many challenges are still unaddressed. We create a large-scale dataset, Multilingual Multiparty Coref (MMC), for this task based on TV transcripts. Due to the availability of gold-quality subtitles in multiple languages, we propose reusing the annotations to create silver coreference resolution data in other languages (Chinese and Farsi) via annotation projection. On the gold (English) data, off-the-shelf models perform relatively poorly on MMC, suggesting that MMC has broader coverage of multiparty coreference than prior datasets. On the silver data, we find success both using it for data augmentation and training from scratch, which effectively simulates the zero-shot cross-lingual setting.
Boyuan Zheng 0001, Patrick Xia 0002, Mahsa Yarmohammadi, Benjamin Van Durme
Trans. Assoc. Comput. Linguistics2
2022 Adapting Coreference Resolution Models through Active Learning
abstract
Neural coreference resolution models trained on one dataset may not transfer to new, lowresource domains.Active learning mitigates this problem by sampling a small subset of data for annotators to label.While active learning is well-defined for classification tasks, its application to coreference resolution is neither well-defined nor fully understood.This paper explores how to actively label coreference, examining sources of model uncertainty and document reading costs.We compare uncertainty sampling strategies and their advantages through thorough error analysis.In both synthetic and human experiments, labeling spans within the same document is more effective than annotating spans across documents.The findings contribute to a more realistic development of coreference resolution models.
Michelle Yuan, Patrick Xia 0002, Chandler May, Benjamin Van Durme, Jordan L. Boyd-Graber
ACL (1)2
2022 Automatic Document Selection for Efficient Encoder Pretraining
abstract
Building pretrained language models is considered expensive and data-intensive, but must we increase dataset size to achieve better performance?We propose an alternative to larger training sets by automatically identifying smaller yet domain-representative subsets.We extend Cynical Data Selection, a statistical sentence scoring method that conditions on a representative target domain corpus.As an example, we treat the OntoNotes corpus as a target domain and pretrain a RoBERTa-like encoder from a cynically selected subset of the Pile.On both perplexity and across several downstream tasks in the target domain, it consistently outperforms random selection with 20x less data, 3x fewer training iterations, and 2x less estimated cloud compute cost, validating the recipe of automatic document selection for LM pretraining.
Yukun Feng, Patrick Xia 0002, Benjamin Van Durme, João Sedoc
EMNLP2
2021 Moving on from OntoNotes: Coreference Resolution Model Transfer
abstract
Academic neural models for coreference resolution (coref) are typically trained on a single dataset, OntoNotes, and model improvements are benchmarked on that same dataset.However, real-world applications of coref depend on the annotation guidelines and the domain of the target dataset, which often differ from those of OntoNotes.We aim to quantify transferability of coref models based on the number of annotated documents available in the target dataset.We examine eleven target datasets and find that continued training is consistently effective and especially beneficial when there are few target documents.We establish new benchmarks across several datasets, including state-of-the-art results on PreCo.
Patrick Xia 0002, Benjamin Van Durme
EMNLP (1)1
2020 Multi-Sentence Argument Linking
abstract
We present a novel document-level model for finding argument spans that fill an event's roles, connecting related ideas in sentencelevel semantic role labeling and coreference resolution.Because existing datasets for cross-sentence linking are small, development of our neural model is supported through the creation of a new resource, Roles Across Multiple Sentences (RAMS), which contains 9,124 annotated events across 139 types.We demonstrate strong performance of our model on RAMS and other event-related datasets.1
Seth Ebner, Patrick Xia 0002, Ryan Culkin, Kyle Rawlins, Benjamin Van Durme
ACL2
2020 Incremental Neural Coreference Resolution in Constant Memory
abstract
We investigate modeling coreference resolution under a fixed memory constraint by extending an incremental clustering algorithm to utilize contextualized encoders and neural components.Given a new sentence, our endto-end algorithm proposes and scores each mention span against explicit entity representations created from the earlier document context (if any).These spans are then used to update the entity's representations before being forgotten; we only retain a fixed set of salient entities throughout the document.In this work, we successfully convert a highperforming model (Joshi et al., 2020), asymptotically reducing its memory usage to constant space with only a 0.3% relative loss in F1 on OntoNotes 5.0.
Patrick Xia 0002, João Sedoc, Benjamin Van Durme
EMNLP (1)1
2020 Which *BERT? A Survey Organizing Contextualized Encoders
abstract
Pretrained contextualized text encoders are now a staple of the NLP community.We present a survey on language representation learning with the aim of consolidating a series of shared lessons learned across a variety of recent efforts.While significant advancements continue at a rapid pace, we find that enough has now been discovered, in different directions, that we can begin to organize advances according to common themes.Through this organization, we highlight important considerations when interpreting recent contributions and choosing which model to use.
Patrick Xia 0002, Benjamin Van Durme
EMNLP (1)1
2020 UniMorph 3.0: Universal Morphology
abstract
The Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological paradigms for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation and a type-level resource of annotated data in diverse languages realizing that schema. We have implemented several improvements to the extraction pipeline which creates most of our data, so that it is both more complete and more correct. We have added 66 new languages, as well as new parts of speech for 12 languages. We have also amended the schema in several ways. Finally, we present three new community tools: two to validate data for resource creators, and one to make morphological data available from the command line. UniMorph is based at the Center for Language and Speech Processing (CLSP) at Johns Hopkins University in Baltimore, Maryland. This paper details advances made to the schema, tooling, and dissemination of project resources since the UniMorph 2.0 release described at LREC 2018.
Arya McCarthy, Christo Kirov, Matteo Grella, Amrit Nidhi, Patrick Xia 0002, Kyle Gorman, Ekaterina Vylomova, Sabrina J. Mielke, Garrett Nicolai, Miikka Silfverberg, Timofey Arkhangelskiy, Nataly Krizhanovsky, Andrew Krizhanovsky, Elena Klyachko, Alexey Sorokin, John Mansfield, Valts Ernstreits, Yuval Pinter, Cassandra L. Jacobs, Ryan Cotterell, Mans Hulden, David Yarowsky
LREC5
2019 Can You Tell Me How to Get Past Sesame Street? Sentence-Level Pretraining Beyond Language Modeling
abstract
Alex Wang, Jan Hula, Patrick Xia, Raghavendra Pappagari, R. Thomas McCoy, Roma Patel, Najoung Kim, Ian Tenney, Yinghui Huang, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Jan Hula, Patrick Xia 0002, Raghavendra Pappagari, Tom McCoy 0001, Roma Patel, Najoung Kim, Ian Tenney, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman
ACL (1)3
2019 What do you learn from context? Probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia 0002, Berlin Chen, Adam Poliak, Tom McCoy 0001, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das 0001, Ellie Pavlick
ICLR (Poster)2
2018 UniMorph 2.0: Universal Morphology
Christo Kirov, Ryan Cotterell, John Sylak-Glassman, Géraldine Walther, Ekaterina Vylomova, Patrick Xia 0002, Manaal Faruqui, Sabrina J. Mielke, Arya McCarthy, Sandra Kübler, David Yarowsky, Jason Eisner, Mans Hulden
LREC6