Yunmo Chen

dblp:252/7831 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
10since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021
YearPublicationVenuePosition
2024 MultiMUC: Multilingual Template Filling on MUC-4
abstract
William Gantt, Shabnam Behzad, Hannah An, Yunmo Chen, Aaron White, Benjamin Van Durme, Mahsa Yarmohammadi. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
William Gantt, Shabnam Behzad, Hannah Youngeun An, Yunmo Chen, Aaron Steven White, Benjamin Van Durme, Mahsa Yarmohammadi
EACL (1)4
2024 Learning to Retrieve Iteratively for In-Context Learning
abstract
We introduce iterative retrieval, a novel framework that empowers retrievers to make iterative decisions through policy optimization.Finding an optimal portfolio of retrieved items is a combinatorial optimization problem, generally considered NP-hard.This approach provides a learned approximation to such a solution, meeting specific task requirements under a given family of large language models (LLMs).We propose a training procedure based on reinforcement learning, incorporating feedback from LLMs.We instantiate an iterative retriever for composing in-context learning (ICL) exemplars and apply it to various semantic parsing tasks that demand synthesized programs as outputs.By adding only 4M additional parameters for state encoding, we convert an offthe-shelf dense retriever into a stateful iterative retriever, outperforming previous methods in selecting ICL exemplars on semantic parsing datasets such as SMCALFLOW, TREEDST, and MTOP.Additionally, the trained iterative retriever generalizes across different inference LLMs beyond the one used during training.
Yunmo Chen, Tongfei Chen, Harsh Jhamtani, Patrick Xia 0002, Richard Shin, Jason Eisner, Benjamin Van Durme
EMNLP1
2024 Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation
abstract
Moderate-sized large language models (LLMs) -- those with 7B or 13B parameters -- exhibit promising machine translation (MT) performance. However, they do not match the performance of state-of-the-art conventional encoder-decoder translation models or larger-scale LLMs such as GPT-4. In this study, we bridge this performance gap. We first assess the shortcomings of supervised fine-tuning for LLMs in the MT task, emphasizing the quality issues present in the reference data, despite being human-generated. Then, in contrast to supervised fine-tuning which mimics reference translations, we introduce Contrastive Preference Optimization (CPO), a novel approach that trains models to avoid generating adequate but not perfect translations. Applying CPO to ALMA models with only 22K parallel sentences and 0.1% parameters yields significant improvements. The resulting model, called ALMA-R, can match or exceed the performance of the WMT competition winners and GPT-4 on WMT'21, WMT'22 and WMT'23 test datasets.
Amr Sharaf, Yunmo Chen, Weiting Tan, Lingfeng Shen, Benjamin Van Durme, Kenton Murray, Young Jin Kim 0006
ICML3
2023 Iterative Document-level Information Extraction via Imitation Learning
abstract
We present a novel iterative extraction model, ITERX, for extracting complex relations, or templates, i.e., #-tuples representing a mapping from named slots to spans of text within a document.Documents may feature zero or more instances of a template of any given type, and the task of template extraction entails identifying the templates in a document and extracting each template's slot values.Our imitation learning approach casts the problem as a Markov decision process (MDP), and relieves the need to use predefined template orders to train an extractor.It leads to state-of-the-art results on two established benchmarks -4-ary relation extraction on SCIREX and template extraction on MUC-4 -as well as a strong baseline on the new BETTER Granular task. 1
Yunmo Chen, William Gantt, Weiwei Gu, Tongfei Chen, Aaron Steven White, Benjamin Van Durme
EACL1
2023 A Unified View of Evaluation Metrics for Structured Prediction
abstract
We present a conceptual framework that unifies a variety of evaluation metrics for different structured prediction tasks (e.g.event and relation extraction, syntactic and semantic parsing).Our framework requires representing the outputs of these tasks as objects of certain data types, and derives metrics through matching of common substructures, possibly followed by normalization.We demonstrate how commonly used metrics for a number of tasks can be succinctly expressed by this framework, and show that new metrics can be naturally derived in a bottom-up way based on an output structure.We release a library that enables this derivation to create new metrics. 1 Finally, we consider how specific characteristics of tasks motivate metric design decisions, and suggest possible modifications to existing metrics in line with those motivations.
Yunmo Chen, William Gantt, Tongfei Chen, Aaron Steven White, Benjamin Van Durme
EMNLP1
2023 When Do Decompositions Help for Machine Reading?
abstract
Answering complex questions often requires multi-step reasoning in order to obtain the final answer.Most research into decompositions of complex questions involves open-domain systems, which have shown success in using these decompositions for improved retrieval.In the machine reading setting, however, work to understand when decompositions are helpful is understudied.We conduct experiments on decompositions in machine reading to unify recent work in this space, using a range of models and datasets.We find that decompositions can be helpful in zero or limited-data settings, giving several points of improvement in exact match.However, we also show that when models are given access to around a few hundred or more examples, decompositions are not helpful (and can actually be detrimental).Thus, our analysis implies that models can learn decompositions implicitly even with limited data. 1
Kangda Wei, Dawn J. Lawrie, Benjamin Van Durme, Yunmo Chen, Orion Weller
EMNLP4
2023 Condensing Multilingual Knowledge with Lightweight Language-Specific Modules
abstract
Incorporating language-specific (LS) modules or Mixture-of-Experts (MoE) are proven methods to boost performance in multilingual model performance, but the scalability of these approaches to hundreds of languages or experts tends to be hard to manage.We present Language-specific Matrix Synthesis (LMS), a novel method that addresses the issue.LMS utilizes parameter-efficient and lightweight modules, reducing the number of parameters while outperforming existing methods, e.g., +1.73 BLEU over Switch Transformer on OPUS-100 multilingual translation.Additionally, we introduce Fuse Distillation (FD) to condense multilingual knowledge from multiple LS modules into a single shared module, improving model inference and storage efficiency.Our approach demonstrates superior scalability and performance compared to state-of-the-art methods. 1 * Equal contribution computational cost may only come from communication among devices (such as ALLToALL) or gate routing.
Weiting Tan, Shuyue Stella Li, Yunmo Chen, Benjamin Van Durme, Philipp Koehn, Kenton Murray
EMNLP4
2023 Differentiable Tree Operations Promote Compositional Generalization
abstract
In the context of structure-to-structure transformation tasks, learning sequences of discrete symbolic operations poses significant challenges due to their non-differentiability. To facilitate the learning of these symbolic sequences, we introduce a differentiable tree interpreter that compiles high-level symbolic tree operations into subsymbolic matrix operations on tensors. We present a novel Differentiable Tree Machine (DTM) architecture that integrates our interpreter with an external memory and an agent that learns to sequentially select tree operations to execute the target transformation in an end-to-end manner. With respect to out-of-distribution compositional generalization on synthetic semantic parsing and language generation tasks, DTM achieves 100% while existing baselines such as Transformer, Tree Transformer, LSTM, and Tree2Tree LSTM achieve less than 30%. DTM remains highly interpretable in addition to its perfect performance.
Paul Soulos, Edward J. Hu, Kate McCurdy, Yunmo Chen, Roland Fernandez, Paul Smolensky, Jianfeng Gao 0001
ICML4
2022 An Empirical Study on Finding Spans
abstract
We present an empirical study on methods for span finding, the selection of consecutive tokens in text for some downstream tasks.We focus on approaches that can be employed in training end-to-end information extraction systems, and find there is no definitive solution without considering task properties, and provide our observations to help with future design choices: 1) a tagging approach often yields higher precision while span enumeration and boundary prediction provide higher recall; 2) span type information can benefit a boundary prediction approach; 3) additional contextualization does not help span finding in most cases.
Weiwei Gu, Boyuan Zheng 0001, Yunmo Chen, Tongfei Chen, Benjamin Van Durme
EMNLP3
2021 Everything Is All It Takes: A Multipronged Strategy for Zero-Shot Cross-Lingual Information Extraction
abstract
Mahsa Yarmohammadi, Shijie Wu, Marc Marone, Haoran Xu, Seth Ebner, Guanghui Qin, Yunmo Chen, Jialiang Guo, Craig Harman, Kenton Murray, Aaron Steven White, Mark Dredze, Benjamin Van Durme. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Mahsa Yarmohammadi, Marc Marone, Seth Ebner, Guanghui Qin, Yunmo Chen, Jialiang Guo, Craig Harman, Kenton Murray, Aaron Steven White, Mark Dredze, Benjamin Van Durme
EMNLP (1)7
2020 Hierarchical Entity Typing via Multi-level Learning to Rank
abstract
We propose a novel method for hierarchical entity classification that embraces ontological structure at both training and during prediction.At training, our novel multi-level learning-to-rank loss compares positive types against negative siblings according to the type tree.During prediction, we define a coarseto-fine decoder that restricts viable candidates at each level of the ontology based on already predicted parent type(s).We achieve stateof-the-art across multiple datasets, particularly with respect to strict accuracy.1
Tongfei Chen, Yunmo Chen, Benjamin Van Durme
ACL2