Yuxian Meng

dblp:234/8585 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
13since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2023 Defending against Backdoor Attacks in Natural Language Generation
abstract
The frustratingly fragile nature of neural network models make current natural language generation (NLG) systems prone to backdoor attacks and generate malicious sequences that could be sexist or offensive. Unfortunately, little effort has been invested to how backdoor attacks can affect current NLG models and how to defend against these attacks. In this work, by giving a formal definition of backdoor attack and defense, we investigate this problem on two important NLG tasks, machine translation and dialog generation. Tailored to the inherent nature of NLG models (e.g., producing a sequence of coherent words given contexts), we design defending strategies against attacks. We find that testing the backward probability of generating sources given targets yields effective defense performance against all different types of attacks, and is able to handle the one-to-many issue in many NLG tasks such as dialog generation. We hope that this work can raise the awareness of backdoor risks concealed in deep NLG systems and inspire more future work (both attack and defense) towards this direction.
Xiaofei Sun 0001, Xiaoya Li 0001, Yuxian Meng, Xiang Ao 0001, Lingjuan Lyu, Jiwei Li 0001, Tianwei Zhang 0004
AAAI3
2022 Dependency Parsing as MRC-based Span-Span Prediction
abstract
Higher-order methods for dependency parsing can partially but not fully address the issue that edges in dependency trees should be constructed at the text span/subtree level rather than word level.In this paper, we propose a new method for dependency parsing to address this issue.The proposed method constructs dependency trees by directly modeling span-span (in other words, subtree-subtree) relations.It consists of two modules: the text span proposal module which proposes candidate text spans, each of which represents a subtree in the dependency tree denoted by (root, start, end); and the span linking module, which constructs links between proposed spans.We use the machine reading comprehension (MRC) framework as the backbone to formalize the span linking module, where one span is used as query to extract the text span/subtree it should be linked to.The proposed method has the following merits: (1) it addresses the fundamental problem that edges in a dependency tree should be constructed between subtrees;(2) the MRC framework allows the method to retrieve missing spans in the span proposal stage, which leads to higher recall for eligible spans.Extensive experiments on the PTB, CTB and Universal Dependencies (UD) benchmarks demonstrate the effectiveness of the proposed method. 1 2
Leilei Gan, Yuxian Meng, Kun Kuang 0001, Xiaofei Sun 0001, Chun Fan 0001, Fei Wu 0001, Jiwei Li 0001
ACL (1)2
2022 Summarize, Outline, and Elaborate: Long-Text Generation via Hierarchical Supervision from Extractive Summaries
abstract
The difficulty of generating coherent long texts lies in the fact that existing models overwhelmingly focus on the tasks of local word prediction, and cannot make high level plans on what to generate or capture the high-level discourse dependencies between chunks of texts. Inspired by how humans write, where a list of bullet points or a catalog is first outlined, and then each bullet point is expanded to form the whole article, we propose SOE, a pipelined system that involves of summarizing, outlining and elaborating for long text generation: the model first outlines the summaries for different segments of long texts, and then elaborates on each bullet point to generate the corresponding segment. To avoid the labor-intensive process of summary soliciting, we propose the reconstruction strategy, which extracts segment summaries in an unsupervised manner by selecting its most informative part to reconstruct the segment. The proposed generation system comes with the following merits: (1) the summary provides high-level guidance for text generation and avoids the local minimum of individual word predictions; (2) the high-level discourse dependencies are captured in the conditional dependencies between summaries and are preserved during the summary expansion process and (3) additionally, we are able to consider significantly more contexts by representing contexts as concise summaries. Extensive experiments demonstrate that SOE produces long texts with significantly better quality, along with faster convergence speed.
Xiaofei Sun 0001, Zijun Sun, Yuxian Meng, Jiwei Li 0001, Chun Fan 0001
COLING3
2022 Paraphrase Generation as Unsupervised Machine Translation
abstract
In this paper, we propose a new paradigm for paraphrase generation by treating the task as unsupervised machine translation (UMT) based on the assumption that there must be pairs of sentences expressing the same meaning in a large-scale unlabeled monolingual corpus. The proposed paradigm first splits a large unlabeled corpus into multiple clusters, and trains multiple UMT models using pairs of these clusters. Then based on the paraphrase pairs produced by these UMT models, a unified surrogate model can be trained to serve as the final model to generate paraphrases, which can be directly used for test in the unsupervised setup, or be finetuned on labeled datasets in the supervised setup. The proposed method offers merits over machine-translation-based paraphrase generation methods, as it avoids reliance on bilingual sentence pairs. It also allows human intervene with the model so that more diverse paraphrases can be generated using different filtering criteria. Extensive experiments on existing paraphrase dataset for both the supervised and unsupervised setups demonstrate the effectiveness the proposed paradigm.
Xiaofei Sun 0001, Yufei Tian, Yuxian Meng, Nanyun Peng 0001, Fei Wu 0001, Jiwei Li 0001, Chun Fan 0001
COLING3
2022 An MRC Framework for Semantic Role Labeling
abstract
Semantic Role Labeling (SRL) aims at recognizing the predicate-argument structure of a sentence and can be decomposed into two subtasks: predicate disambiguation and argument labeling. Prior work deals with these two tasks independently, which ignores the semantic connection between the two tasks. In this paper, we propose to use the machine reading comprehension (MRC) framework to bridge this gap. We formalize predicate disambiguation as multiple-choice machine reading comprehension, where the descriptions of candidate senses of a given predicate are used as options to select the correct sense. The chosen predicate sense is then used to determine the semantic roles for that predicate, and these semantic roles are used to construct the query for another MRC model for argument labeling. In this way, we are able to leverage both the predicate semantics and the semantic role semantics for argument labeling. We also propose to select a subset of all the possible semantic roles for computational efficiency. Experiments show that the proposed framework achieves state-of-the-art or comparable results to previous work.
Jiwei Li 0001, Yuxian Meng, Xiaofei Sun 0001, Han Qiu 0001, Guoyin Wang 0002, Jun He 0008
COLING3
2022 BadPre: Task-agnostic Backdoor Attacks to Pre-trained NLP Foundation Models
Kangjie Chen, Yuxian Meng, Xiaofei Sun 0001, Shangwei Guo, Tianwei Zhang 0004, Jiwei Li 0001, Chun Fan 0001
ICLR2
2022 GNN-LM: Language Modeling based on Global Contexts via GNN
Yuxian Meng, Shi Zong, Xiaoya Li 0001, Xiaofei Sun 0001, Tianwei Zhang 0004, Fei Wu 0001, Jiwei Li 0001
ICLR1
2022 Triggerless Backdoor Attack for NLP Tasks with Clean Labels
abstract
Leilei Gan, Jiwei Li, Tianwei Zhang, Xiaoya Li, Yuxian Meng, Fei Wu, Yi Yang, Shangwei Guo, Chun Fan. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Leilei Gan, Jiwei Li 0001, Tianwei Zhang 0004, Xiaoya Li 0001, Yuxian Meng, Fei Wu 0001, Yi Yang 0001, Shangwei Guo, Chun Fan 0001
NAACL-HLT5
2022 Sentence Similarity Based on Contexts
abstract
Abstract Existing methods to measure sentence similarity are faced with two challenges: (1) labeled datasets are usually limited in size, making them insufficient to train supervised neural models; and (2) there is a training-test gap for unsupervised language modeling (LM) based models to compute semantic scores between sentences, since sentence-level semantics are not explicitly modeled at training. This results in inferior performances in this task. In this work, we propose a new framework to address these two issues. The proposed framework is based on the core idea that the meaning of a sentence should be defined by its contexts, and that sentence similarity can be measured by comparing the probabilities of generating two sentences given the same context. The proposed framework is able to generate high-quality, large-scale dataset with semantic similarity scores between two sentences in an unsupervised manner, with which the train-test gap can be largely bridged. Extensive experiments show that the proposed framework achieves significant performance boosts over existing baselines under both the supervised and unsupervised settings across different datasets.
Xiaofei Sun 0001, Yuxian Meng, Xiang Ao 0001, Fei Wu 0001, Tianwei Zhang 0004, Jiwei Li 0001, Chun Fan 0001
Trans. Assoc. Comput. Linguistics2
2021 ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin Information
abstract
Zijun Sun, Xiaoya Li, Xiaofei Sun, Yuxian Meng, Xiang Ao, Qing He, Fei Wu, Jiwei Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Zijun Sun, Xiaoya Li 0001, Xiaofei Sun 0001, Yuxian Meng, Xiang Ao 0001, Qing He 0003, Fei Wu 0001, Jiwei Li 0001
ACL/IJCNLP (1)4
2021 Layer-wise Model Pruning based on Mutual Information
abstract
Inspired by mutual information (MI) based feature selection in SVMs and logistic regression, in this paper, we propose MI-based layer-wise pruning: for each layer of a multi-layer neural network, neurons with higher values of MI with respect to preserved neurons in the upper layer are preserved.Starting from the top softmax layer, layer-wise pruning proceeds in a top-down fashion until reaching the bottom word embedding layer.The proposed pruning strategy offers merits over weight-based pruning techniques: (1) it avoids irregular memory access since representations and matrices can be squeezed into their smaller but dense counterparts, leading to greater speedup; (2) in a manner of top-down pruning, the proposed method operates from a more global perspective based on training signals in the top layer, and prunes each layer by propagating the effect of global signals through layers, leading to better performances at the same sparsity level.Extensive experiments show that at the same sparsity level, the proposed strategy offers both greater speedup and higher performances than weight-based pruning methods (e.g., magnitude pruning, movement pruning).
Chun Fan 0001, Jiwei Li 0001, Tianwei Zhang 0004, Xiang Ao 0001, Fei Wu 0001, Yuxian Meng, Xiaofei Sun 0001
EMNLP (1)6
2021 kFolden: k-Fold Ensemble for Out-Of-Distribution Detection
abstract
Out-of-Distribution (OOD) detection is an important problem in natural language processing (NLP).In this work, we propose a simple yet effective framework kFolden, which mimics the behaviors of OOD detection during training without the use of any external data.For a task with k training labels, kFolden induces k sub-models, each of which is trained on a subset with k -1 categories with the left category masked unknown to the sub-model.Exposing an unknown label to the sub-model during training, the model is encouraged to learn to equally attribute the probability to the seen k -1 labels for the unknown label, enabling this framework to simultaneously resolve in-and out-distribution examples in a natural way via OOD simulations.Taking text classification as an archetype, we develop benchmarks for OOD detection using existing text classification datasets.By conducting comprehensive comparisons and analyses on the developed benchmarks, we demonstrate the superiority of kFolden against current methods in terms of improving OOD detection performances while maintaining improved in-domain classification accuracy.1
Xiaoya Li 0001, Jiwei Li 0001, Xiaofei Sun 0001, Chun Fan 0001, Tianwei Zhang 0004, Fei Wu 0001, Yuxian Meng
EMNLP (1)7
2021 ConRPG: Paraphrase Generation using Contexts as Regularizer
abstract
A long-standing issue with paraphrase generation is how to obtain reliable supervision signals.In this paper, we propose an unsupervised paradigm for paraphrase generation based on the assumption that the probabilities of generating two sentences with the same meaning given the same context should be the same.Inspired by this fundamental idea, we propose a pipelined system which consists of paraphrase candidate generation based on contextual language models, candidate filtering using scoring functions, and paraphrase model training based on the selected candidates.The proposed paradigm offers merits over existing paraphrase generation methods: (1) using the context regularizer on meanings, the model is able to generate massive amounts of high-quality paraphrase pairs; and (2) using human-interpretable scoring functions to select paraphrase pairs from candidates, the proposed framework provides a channel for developers to intervene with the data generation process, leading to a more controllable model.Experimental results across different tasks and datasets demonstrate that the effectiveness of the proposed model in both supervised and unsupervised setups.
Yuxian Meng, Xiang Ao 0001, Qing He 0003, Xiaofei Sun 0001, Qinghong Han, Fei Wu 0001, Chun Fan 0001, Jiwei Li 0001
EMNLP (1)1
2020 A Unified MRC Framework for Named Entity Recognition
abstract
The task of named entity recognition (NER) is normally divided into nested NER and flat NER depending on whether named entities are nested or not.Models are usually separately developed for the two tasks, since sequence labeling models are only able to assign a single label to a particular token, which is unsuitable for nested NER where a token may be assigned several labels.
Xiaoya Li 0001, Jingrong Feng, Yuxian Meng, Qinghong Han, Fei Wu 0001, Jiwei Li 0001
ACL3
2020 Dice Loss for Data-imbalanced NLP Tasks
abstract
Many NLP tasks such as tagging and machine reading comprehension (MRC) are faced with the severe data imbalance issue: negative examples significantly outnumber positive ones, and the huge number of easy-negative examples overwhelms training.The most commonly used cross entropy criteria is actually accuracy-oriented, which creates a discrepancy between training and test.At training time, each training instance contributes equally to the objective function, while at test time F1 score concerns more about positive examples.
Xiaoya Li 0001, Xiaofei Sun 0001, Yuxian Meng, Junjun Liang, Fei Wu 0001, Jiwei Li 0001
ACL3
2020 SAC: Accelerating and Structuring Self-Attention via Sparse Adaptive Connection
abstract
While the self-attention mechanism has been widely used in a wide variety of tasks, it has the unfortunate property of a quadratic cost with respect to the input length, which makes it difficult to deal with long inputs. In this paper, we present a method for accelerating and structuring self-attentions: Sparse Adaptive Connection (SAC). In SAC, we regard the input sequence as a graph and attention operations are performed between linked nodes. In contrast with previous self-attention models with pre-defined structures (edges), the model learns to construct attention edges to improve task-specific performances. In this way, the model is able to select the most salient nodes and reduce the quadratic complexity regardless of the sequence length. Based on SAC, we show that previous variants of self-attention models are its special cases. Through extensive experiments on neural machine translation, language modeling, graph representation learning and image classification, we demonstrate SAC is competitive with state-of-the-art models while significantly reducing memory cost.
Xiaoya Li 0001, Yuxian Meng, Mingxin Zhou, Qinghong Han, Fei Wu 0001, Jiwei Li 0001
NeurIPS2
2019 Is Word Segmentation Necessary for Deep Learning of Chinese Representations?
abstract
Segmenting a chunk of text into words is usually the first step of processing Chinese text, but its necessity has rarely been explored.In this paper, we ask the fundamental question of whether Chinese word segmentation (CWS) is necessary for deep learning-based Chinese Natural Language Processing.We benchmark neural word-based models which rely on word segmentation against neural char-based models which do not involve word segmentation in four end-to-end NLP benchmark tasks: language modeling, machine translation, sentence matching/paraphrase and text classification.Through direct comparisons between these two types of models, we find that charbased models consistently outperform wordbased models.Based on these observations, we conduct comprehensive experiments to study why wordbased models underperform char-based models in these deep learning-based NLP tasks.We show that it is because word-based models are more vulnerable to data sparsity and the presence of out-of-vocabulary (OOV) words, and thus more prone to overfitting.We hope this paper could encourage researchers in the community to rethink the necessity of word segmentation in deep learning-based Chinese Natural Language Processing. 1
Xiaoya Li 0001, Yuxian Meng, Xiaofei Sun 0001, Qinghong Han, Arianna Yuan, Jiwei Li 0001
ACL (1)2
2019 Glyce: Glyph-vectors for Chinese Character Representations
abstract
It is intuitive that NLP tasks for logographic languages like Chinese should benefit from the use of the glyph information in those languages. However, due to the lack of rich pictographic evidence in glyphs and the weak generalization ability of standard computer vision models on character data, an effective way to utilize the glyph information remains to be found. In this paper, we address this gap by presenting Glyce, the glyph-vectors for Chinese character representations. We make three major innovations: (1) We use historical Chinese scripts (e.g., bronzeware script, seal script, traditional Chinese, etc) to enrich the pictographic evidence in characters; (2) We design CNN structures (called tianzege-CNN) tailored to Chinese character image processing; and (3) We use image-classification as an auxiliary task in a multi-task learning setup to increase the model's ability to generalize. We show that glyph-based models are able to consistently outperform word/char ID-based models in a wide range of Chinese NLP tasks. When combing with BERT, we are able to set new state-of-the-art results for a variety of Chinese NLP tasks, including language modeling, tagging (NER, CWS, POS), sentence pair classification (BQ, LCQMC, XNLI, NLPCC-DBQA), single sentence classification tasks (ChnSentiCorp, the Fudan corpus, iFeng), dependency parsing, and semantic role labeling. For example, the proposed model achieves an F1 score of 81.6 on the OntoNotes dataset of NER, +1.5 over BERT; it achieves an almost perfect accuracy of 99.8\% on the the Fudan corpus for text classification.
Yuxian Meng, Wei Wu 0044, Fei Wang 0060, Xiaoya Li 0001, Ping Nie, Fan Yin, Muyu Li, Qinghong Han, Xiaofei Sun 0001, Jiwei Li 0001
NeurIPS1