EDBT 2026 Demo / reviewers in the wild / expert
Leyang Cui
dblp:247/6181
· DBLP profile ↗
33ranked-venue papers
5as first author
26since 2021 · last 2025
0000-0001-5072-884XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 5 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Lost in Literalism: How Supervised Training Shapes Translationese in LLMsabstractLarge language models (LLMs) have achieved remarkable success in machine translation, demonstrating impressive performance across diverse languages. However, translationese—characterized by overly literal and unnatural translations—remains a persistent challenge in LLM-based translation systems. Despite their pre-training on vast corpora of natural utterances, LLMs exhibit translationese errors and generate unexpected unnatural translations, stemming from biases introduced during supervised fine-tuning (SFT). In this work, we systematically evaluate the prevalence of translationese in LLM-generated translations and investigate its roots during supervised training. We introduce methods to mitigate these biases, including polishing golden references and filtering unnatural training instances. Empirical evaluations demonstrate that these approaches significantly reduce translationese while improving translation naturalness, validated by human evaluations and automatic metrics. Our findings highlight the need for training-aware adjustments to optimize LLM translation outputs, paving the way for more fluent and target-language-consistent translations. Yafu Li, Ronghao Zhang, Zhilin Wang, Leyang Cui, Yongjing Yin, Tong Xiao 0001, Yue Zhang 0004 |
ACL (1) | 5 |
| 2025 | Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP TasksabstractPrevious work adopts large language models (LLMs) as evaluators to evaluate natural language process (NLP) tasks. However, certain shortcomings, e.g., fairness, scope, and accuracy, persist for current LLM evaluators. To analyze whether LLMs can serve as reliable alternatives to humans, we examine the fine-grained alignment between LLM evaluators and human annotators, particularly in understanding the target evaluation tasks and conducting evaluations that meet diverse criteria. This paper explores both conventional tasks (e.g., story generation) and alignment tasks (e.g., math reasoning), each with different evaluation criteria. Our analysis shows that 1) LLM evaluators can generate unnecessary criteria or omit crucial criteria, resulting in a slight deviation from the experts. 2) LLM evaluators excel in general criteria, such as fluency, but face challenges with complex criteria, such as numerical reasoning. We also find that LLM-pre-drafting before human evaluation can help reduce the impact of human subjectivity and minimize annotation outliers in pure human evaluation, leading to more objective evaluation. All resources are available at https://github.com/qtli/CoEval. Qintong Li, Leyang Cui, Lingpeng Kong, Wei Bi |
COLING | 2 |
| 2025 | Thinking Out Loud: Do Reasoning Models Know When They're Right?abstractLarge reasoning models (LRMs) have recently demonstrated impressive capabilities in complex reasoning tasks by leveraging increased test-time computation and exhibiting behaviors reminiscent of human-like self-reflection.While LRMs show a clear capacity for valuable self-reflection, how this ability interacts with other model behaviors remains underexplored.We investigate this connection by analyzing verbalized confidence, how models articulate their certainty, as a lens into the nature of self-reflection in LRMs.We find that supervised fine-tuning on reasoning traces (i.e., distillation) and reinforcement learning can improve verbalized calibration in reasoningintensive settings in a progressive, laddered fashion.However, our results also indicate that reasoning models may possess a diminished awareness of their own knowledge boundaries, as evidenced by significantly lower "I don't know" response rates on factuality benchmarks.Moreover, we examine the relationship between verbalized confidence and reasoning chains, finding that models tend to express higher confidence when providing shorter or less elaborate reasoning.Our findings highlight how reasoning-oriented training can enhance performance in reasoning-centric tasks while potentially incurring a reasoning tax, a cost reflected in the model's reduced ability to accurately recognize the limits of its own knowledge in small-scale models.More broadly, our work showcases how this erosion of knowledge boundaries can compromise model faithfulness, as models grow more confident without a commensurate understanding of when they should abstain. Qingcheng Zeng, Weihao Xuan, Leyang Cui, Rob Voigt |
EMNLP | 3 |
| 2025 | ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM ReasoningabstractEvaluating large language models (LLMs) poses significant challenges, particularly due to issues of data contamination and the leakage of correct answers. To address these challenges, we introduce ThinkBench, a novel evaluation framework designed to robustly evaluate the reasoning capability of LLMs. ThinkBench proposes a dynamic data generation method for constructing out-of-distribution (OOD) datasets and offers an OOD dataset that contains 2,912 samples drawn from reasoning tasks. ThinkBench unifies the evaluation of reasoning models and non-reasoning models. We evaluate 16 LLMs and 4 PRMs under identical experimental conditions and show that most of the LLMs' performance are far from robust and they face a certain level of data leakage. By dynamically generating OOD datasets, ThinkBench effectively provides a reliable evaluation of LLMs and reduces data contamination impact. Our data and codes are available at https://github.com/huangshulin123/ThinkBench. Shulin Huang, Linyi Yang, Yan Song 0003, Shawn Chen, Leyang Cui, Ziyu Wan, Qingcheng Zeng, Ying Wen 0001, Kun Shao, Weinan Zhang 0001, Jun Wang 0012, Yue Zhang 0004 |
NeurIPS | 5 |
| 2024 | NaRuto: Automatically Acquiring Planning Models from Narrative TextsabstractDomain model acquisition has been identified as a bottleneck in the application of planning technology, especially within narrative planning. Learning action models from narrative texts in an automated way is essential to overcome this barrier, but challenging because of the inherent complexities of such texts. We present an evaluation of planning domain models derived from narrative texts using our fully automated, unsupervised system, NaRuto. Our system combines structured event extraction, predictions of commonsense event relations, and textual contradictions and similarities. Evaluation results show that NaRuto generates domain models of significantly better quality than existing fully automated methods, and even sometimes on par with those created by semi-automated methods, with human assistance. Ruiqi Li 0005, Leyang Cui, Songtuan Lin, Patrik Haslum |
AAAI | 2 |
| 2024 | Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized RehearsalabstractJianheng Huang, Leyang Cui, Ante Wang, Chengyi Yang, Xinting Liao, Linfeng Song, Junfeng Yao, Jinsong Su. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Jianheng Huang, Leyang Cui, Ante Wang, Xinting Liao, Linfeng Song, Junfeng Yao, Jinsong Su |
ACL (1) | 2 |
| 2024 | GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem SolversabstractLarge language models (LLMs) have achieved impressive performance across various mathematical reasoning benchmarks.However, there are increasing debates regarding whether these models truly understand and apply mathematical knowledge or merely rely on shortcuts for mathematical reasoning.One essential and frequently occurring evidence is that when the math questions are slightly changed, LLMs can behave incorrectly.This motivates us to evaluate the robustness of LLMs' math reasoning capability by testing a wide range of question variations.We introduce the adversarial grade school math (GSM-PLUS) dataset, an extension of GSM8K augmented with various mathematical perturbations.Our experiments on 25 LLMs and 4 prompting techniques show that while LLMs exhibit different levels of math reasoning abilities, their performances are far from robust.In particular, even for problems that have been solved in GSM8K, LLMs can make mistakes when new statements are added or the question targets are altered.We also explore whether more robust performance can be achieved by composing existing prompting methods, in which we try an iterative method that generates and verifies each intermediate thought based on its reasoning goal and calculation result. Qintong Li, Leyang Cui, Xueliang Zhao, Lingpeng Kong, Wei Bi |
ACL (1) | 2 |
| 2024 | MAGE: Machine-generated Text Detection in the WildabstractYafu Li, Qintong Li, Leyang Cui, Wei Bi, Zhilin Wang, Longyue Wang, Linyi Yang, Shuming Shi, Yue Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yafu Li, Qintong Li, Leyang Cui, Wei Bi, Zhilin Wang, Longyue Wang, Linyi Yang, Shuming Shi 0001, Yue Zhang 0004 |
ACL (1) | 3 |
| 2024 | Knowledge Verification to Nip Hallucination in the BudabstractWhile large language models (LLMs) have demonstrated exceptional performance across various tasks following human alignment, they may still generate responses that sound plausible but contradict factual knowledge, a phenomenon known as hallucination.In this paper, we demonstrate the feasibility of mitigating hallucinations by verifying and minimizing the inconsistency between external knowledge present in the alignment data and the intrinsic knowledge embedded within foundation LLMs.Specifically, we propose a novel approach called Knowledge Consistent Alignment (KCA), which employs a well-aligned LLM to automatically formulate assessments based on external knowledge to evaluate the knowledge boundaries of foundation LLMs.To address knowledge inconsistencies in the alignment data, KCA implements several specific strategies to deal with these data instances.We demonstrate the superior efficacy of KCA in reducing hallucinations across six benchmarks, utilizing foundation LLMs of varying backbones and scales.This confirms the effectiveness of mitigating hallucinations by reducing knowledge inconsistency.Our code, model weights, and data are openly accessible at https://github.com/fanqiwan/KCA.* Part of the work was done during his internship at Tencent AI Lab. Fanqi Wan, Xinting Huang, Leyang Cui, Xiaojun Quan, Wei Bi, Shuming Shi 0001 |
EMNLP | 3 |
| 2024 | Retrieval is Accurate GenerationabstractStandard language models generate text by selecting tokens from a fixed, finite, and standalone vocabulary. We introduce a novel method that selects context-aware phrases from a collection of supporting documents. One of the most significant challenges for this paradigm shift is determining the training oracles, because a string of text can be segmented in various ways and each segment can be retrieved from numerous possible documents. To address this, we propose to initialize the training oracles using linguistic heuristics and, more importantly, bootstrap the oracles through iterative self-reinforcement. Extensive experiments show that our model not only outperforms standard language models on a variety of knowledge-intensive tasks but also demonstrates improved generation quality in open-ended text generation. For instance, compared to the standard language model counterpart, our model raises the accuracy from 23.47% to 36.27% on OpenbookQA, and improves the MAUVE score from 42.61% to 81.58% in open-ended text generation. Remarkably, our model also achieves the best performance and the lowest latency among several retrieval-augmented baselines. In conclusion, we assert that retrieval is more accurate generation and hope that our work will encourage further research on this new paradigm shift. Bowen Cao, Deng Cai 0002, Leyang Cui, Xuxin Cheng, Wei Bi, Yuexian Zou, Shuming Shi 0001 |
ICLR | 3 |
| 2024 | Gated Slot Attention for Efficient Linear-Time Sequence ModelingabstractLinear attention Transformers and their gated variants, celebrated for enabling parallel training and efficient recurrent inference, still fall short in recall-intensive tasks compared to traditional Transformers and demand significant resources for training from scratch.
This paper introduces Gated Slot Attention (GSA), which enhances Attention with Bounded-memory-Control (ABC) by incorporating a gating mechanism inspired by Gated Linear Attention (GLA).
Essentially, GSA comprises a two-layer GLA linked via $\operatorname{softmax}$, utilizing context-aware memory reading and adaptive forgetting to improve memory capacity while maintaining compact recurrent state size.
This design greatly enhances both training and inference efficiency through GLA's hardware-efficient training algorithm and reduced state size.
Additionally, retaining the $\operatorname{softmax}$ operation is particularly beneficial in ``finetuning pretrained Transformers to RNNs'' (T2R) settings, reducing the need for extensive training from scratch.
Extensive experiments confirm GSA's superior performance in scenarios requiring in-context recall and in T2R settings. Yu Zhang 0092, Rui-Jie Zhu 0003, Yue Zhang 0004, Leyang Cui, Yiqiao Wang 0005, Bolun Wang, Freda Shi, Bailin Wang, Wei Bi, Peng Zhou 0017, Guohong Fu |
NeurIPS | 5 |
| 2023 | Enhancing Grammatical Error Correction Systems with ExplanationsabstractGrammatical error correction systems improve written communication by detecting and correcting language mistakes.To help language learners better understand why the GEC system makes a certain correction, the causes of errors (evidence words) and the corresponding error types are two key factors.To enhance GEC systems with explanations, we introduce EXPECT, a large dataset annotated with evidence words and grammatical error types.We propose several baselines and analysis to understand this task.Furthermore, human evaluation verifies our explainable GEC system's explanations can assist second-language learners in determining whether to accept a correction suggestion and in understanding the associated grammar rule. Yuejiao Fei, Leyang Cui, Sen Yang 0005, Wai Lam, Zhen-Zhong Lan, Shuming Shi 0001 |
ACL (1) | 2 |
| 2023 | Explicit Syntactic Guidance for Neural Text GenerationabstractMost existing text generation models follow the sequence-to-sequence paradigm.Generative Grammar suggests that humans generate natural language texts by learning language grammar.We propose a syntax-guided generation schema, which generates the sequence guided by a constituency parse tree in a topdown direction.The decoding process can be decomposed into two parts: (1) predicting the infilling texts for each constituent in the lexicalized syntax context given the source sentence;(2) mapping and expanding each constituent to construct the next-level syntax context.Accordingly, we propose a structural beam search method to find possible syntax structures hierarchically.Experiments on paraphrase generation and machine translation show that the proposed method outperforms autoregressive baselines, while also demonstrating effectiveness in terms of interpretability, controllability, and diversity. Yafu Li, Leyang Cui, Jianhao Yan, Yongjing Yin, Wei Bi, Shuming Shi 0001, Yue Zhang 0004 |
ACL (1) | 2 |
| 2023 | RobustGEC: Robust Grammatical Error Correction Against Subtle Context PerturbationabstractGrammatical Error Correction (GEC) systems play a vital role in assisting people with their daily writing tasks.However, users may sometimes come across a GEC system that initially performs well but fails to correct errors when the inputs are slightly modified.To ensure an ideal user experience, a reliable GEC system should have the ability to provide consistent and accurate suggestions when encountering irrelevant context perturbations, which we refer to as context robustness.In this paper, we introduce RobustGEC, a benchmark designed to evaluate the context robustness of GEC systems.RobustGEC comprises 5,000 GEC cases, each with one original error-correct sentence pair and five variants carefully devised by human annotators.Utilizing RobustGEC, we reveal that state-of-the-art GEC systems still lack sufficient robustness against context perturbations.In addition, we propose a simple yet effective method for remitting this issue. Yue Zhang 0004, Leyang Cui, Enbo Zhao, Wei Bi, Shuming Shi 0001 |
EMNLP | 2 |
| 2023 | Non-autoregressive Text Editing with Copy-aware Latent AlignmentsabstractRecent work has witnessed a paradigm shift from Seq2Seq to Seq2Edit in the field of text editing, with the aim of addressing the slow autoregressive inference problem posed by the former.Despite promising results, Seq2Edit approaches still face several challenges such as inflexibility in generation and difficulty in generalizing to other languages.In this work, we propose a novel non-autoregressive text editing method to circumvent the above issues, by modeling the edit process with latent CTC alignments.We make a crucial extension to CTC by introducing the copy operation into the edit space, thus enabling more efficient management of textual overlap in editing.We conduct extensive experiments on GEC and sentence fusion tasks, showing that our proposed method significantly outperforms existing Seq2Edit models and achieves similar or even better results than Seq2Seq with over 4ˆ speedup.Moreover, it demonstrates good generalizability on German and Russian.In-depth analyses reveal the strengths of our method in terms of the robustness under various scenarios and generating fluent and flexible outputs. Yu Zhang 0092, Yue Zhang 0004, Leyang Cui, Guohong Fu |
EMNLP | 3 |
| 2023 | EDeR: Towards Understanding Dependency Relations Between EventsabstractRelation extraction is a crucial task in natural language processing (NLP) and information retrieval (IR).Previous work on event relation extraction mainly focuses on hierarchical, temporal and causal relations.Such relationships consider two events to be independent in terms of syntax and semantics, but they fail to recognize the interdependence between events.To bridge this gap, we introduce a human-annotated Event Dependency Relation dataset (EDeR).The annotation is done on a sample of documents from the OntoNotes dataset, which has the additional benefit that it integrates with existing, orthogonal, annotations of this dataset.We investigate baseline approaches for EDeR's event dependency relation prediction.We show that recognizing such event dependency relations can further benefit critical NLP tasks, including semantic role labelling and co-reference resolution. Ruiqi Li 0005, Patrik Haslum, Leyang Cui |
EMNLP | 3 |
| 2023 | LogiQA 2.0 - An Improved Dataset for Logical Reasoning in Natural Language UnderstandingabstractNLP research on logical reasoning regains momentum with the recent releases of a handful of datasets, notably LogiQA and Reclor. Logical reasoning is exploited in many probing tasks over large Pre-trained Language Models (PLMs) and downstream tasks like question-answering and dialogue systems. In this paper, we release LogiQA 2.0. The dataset is an amendment and re-annotation of LogiQA in 2020, a large-scale logical reasoning reading comprehension dataset adapted from the Chinese Civil Service Examination. We increase the data size, refine the texts with manual translation by professionals, and improve the quality by removing items with distinctive cultural features like Chinese idioms. Furthermore, we conduct a fine-grained annotation on the dataset and turn it into a two-way natural language inference (NLI) task, resulting in 35k premise-hypothesis pairs with gold labels, making it the first large-scale NLI dataset for complex logical reasoning. Compared to Question Answering, Natural Language Inference excels in generalizability and helps downstream tasks better. We establish a baseline for logical reasoning in NLI and incite further research. Hanmeng Liu, Jian Liu 0030, Leyang Cui, Zhiyang Teng, Nan Duan 0001, Ming Zhou 0001, Yue Zhang 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Investigating Non-local Features for Neural Constituency ParsingabstractThanks to the strong representation power of neural encoders, neural chart-based parsers have achieved highly competitive performance by using local features.Recently, it has been shown that non-local features in CRF structures lead to improvements.In this paper, we investigate injecting non-local features into the training process of a local span-based parser, by predicting constituent n-gram non-local patterns and ensuring consistency between non-local patterns and local constituents.Results show that our simple method gives better results than the self-attentive parser on both PTB and CTB.Besides, our method achieves state-of-the-art BERT-based performance on PTB (95.92 F1) and strong performance on CTB (92.31 F1).Our parser also achieves better or competitive performance in multilingual and zero-shot cross-domain settings compared with the baseline. Leyang Cui, Sen Yang 0005, Yue Zhang 0004 |
ACL (1) | 1 |
| 2022 | FactMix: Using a Few Labeled In-domain Examples to Generalize to Cross-domain Named Entity RecognitionabstractFew-shot Named Entity Recognition (NER) is imperative for entity tagging in limited resource domains and thus received proper attention in recent years. Existing approaches for few-shot NER are evaluated mainly under in-domain settings. In contrast, little is known about how these inherently faithful models perform in cross-domain NER using a few labeled in-domain examples. This paper proposes a two-step rationale-centric data augmentation method to improve the model’s generalization ability. Results on several datasets show that our model-agnostic method significantly improves the performance of cross-domain NER tasks compared to previous state-of-the-art methods compared to the counterfactual data augmentation and prompt-tuning methods. Linyi Yang, Lifan Yuan, Leyang Cui, Wenyang Gao, Yue Zhang 0004 |
COLING | 3 |
| 2022 | Cross-domain Generalization for AMR ParsingabstractMeaning Representation (AMR) parsing aims to predict an AMR graph from textual input.Recently, there has been notable growth in AMR parsing performance.However, most existing work focuses on improving the performance in the specific domain, ignoring the potential domain dependence of AMR parsing systems.To address this, we extensively evaluate five representative AMR parsers on five domains and analyze challenges to cross-domain AMR parsing.We observe that challenges to cross-domain AMR parsing mainly arise from the distribution shift of words and AMR concepts.Based on our observation, we investigate two approaches to reduce the domain distribution divergence of text and AMR features, respectively.Experimental results on two out-of-domain test sets show the superiority of our method. Xuefeng Bai 0001, Sen Yang 0005, Leyang Cui, Linfeng Song, Yue Zhang 0004 |
EMNLP | 3 |
| 2022 | Multi-Granularity Optimization for Non-Autoregressive TranslationabstractDespite low latency, non-autoregressive machine translation (NAT) suffers severe performance deterioration due to the naive independence assumption.This assumption is further strengthened by cross-entropy loss, which encourages a strict match between the hypothesis and the reference token by token.To alleviate this issue, we propose multi-granularity optimization for NAT, which collects model behaviors on translation segments of various granularities and integrates feedback for backpropagation.Experiments on four WMT benchmarks show that the proposed method significantly outperforms the baseline models trained with cross-entropy loss, and achieves the best performance on WMT'16 En⇔Ro and highly competitive results on WMT'14 En⇔De for fully non-autoregressive translation. Yafu Li, Leyang Cui, Yongjing Yin, Yue Zhang 0004 |
EMNLP | 2 |
| 2022 | Label Attention Network for Structured PredictionabstractSequence labeling assigns a label to each token in a sequence, which is a fundamental problem in natural language processing (NLP). Many NLP tasks, including part-of-speech tagging and named entity recognition, can be solved in a form of sequence labeling problem. Other tasks such as constituency parsing and non-autoregressive machine translation can also be transformed into sequence labeling tasks. Neural models have been shown powerful for sequence labeling by employing a multi-layer sequence encoding network. Conditional random field (CRF) is proposed to enrich information over label sequences, yet it suffers large computational complexity and over-reliance on Marko assumption. To this end, we propose label attention network (LAN) to hierarchically refine representation of marginal label distributions bottom-up, enabling higher layers to learn more informed label sequence distribution based on information from lower layers. We demonstrate the effectiveness of LAN through extensive experiments on various NLP tasks including POS tagging, NER, CCG supertagging, constituency parsing and non-autoregressive machine translation. Empirical results show that LAN not only improves the overall tagging accuracy with similar number of parameters, but also significantly speeds up the training and testing compared to CRF. Leyang Cui, Yafu Li, Yue Zhang 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Natural Language Inference in Context - Investigating Contextual Reasoning over Long TextsabstractNatural language inference (NLI) is a fundamental NLP task, investigating the entailment relationship between two texts. Popular NLI datasets present the task at sentence-level. While adequate for testing semantic representations, they fall short for testing contextual reasoning over long texts, which is a natural part of the human inference process. We introduce ConTRoL, a new dataset for ConTextual Reasoning over Long texts. Consisting of 8,325 expert-designed "context-hypothesis" pairs with gold labels, ConTRoL is a passage-level NLI dataset with a focus on complex contextual reasoning types such as logical reasoning. It is derived from competitive selection and recruitment test (verbal reasoning test) for police recruitment, with expert level quality. Compared with previous NLI benchmarks, the materials in ConTRoL are much more challenging, involving a range of reasoning types. Empirical results show that state-of-the-art language models perform by far worse than educated humans. Our dataset can also serve as a testing-set for downstream tasks like checking the factual correctness of summaries. Hanmeng Liu, Leyang Cui, Jian Liu 0030, Yue Zhang 0004 |
AAAI | 2 |
| 2021 | Solving Aspect Category Sentiment Analysis as a Text Generation TaskabstractAspect category sentiment analysis has attracted increasing research attention.The dominant methods make use of pre-trained language models by learning effective aspect category-specific representations, and adding specific output layers to its pre-trained representation.We consider a more direct way of making use of pre-trained language models, by casting the ACSA tasks into natural language generation tasks, using natural language sentences to represent the output.Our method allows more direct use of pre-trained knowledge in seq2seq language models by directly following the task setting during pre-training.Experiments on several benchmarks show that our method gives the best reported results, having large advantages in few-shot and zero-shot settings. Jian Liu 0030, Zhiyang Teng, Leyang Cui, Hanmeng Liu, Yue Zhang 0004 |
EMNLP (1) | 3 |
| 2021 | Knowledge Enhanced Fine-Tuning for Better Handling Unseen Entities in Dialogue GenerationabstractAlthough pre-training models have achieved great success in dialogue generation, their performance drops dramatically when the input contains an entity that does not appear in pretraining and fine-tuning datasets (unseen entity).To address this issue, existing methods leverage an external knowledge base to generate appropriate responses.In real-world scenario, the entity may not be included by the knowledge base or suffer from the precision of knowledge retrieval.To deal with this problem, instead of introducing knowledge base as the input, we force the model to learn a better semantic representation by predicting the information in the knowledge base, only based on the input context.Specifically, with the help of a knowledge base, we introduce two auxiliary training objectives: 1) Interpret Masked Word, which conjectures the meaning of the masked entity given the context; 2) Hypernym Generation, which predicts the hypernym of the entity based on the context.Experiment results on two dialogue corpus verify the effectiveness of our methods under both knowledge available and unavailable settings. Leyang Cui, Yu Wu 0012, Shujie Liu 0001, Yue Zhang 0004 |
EMNLP (1) | 1 |
| 2021 | Improving Skip-Gram Embeddings Using BERTabstractContextualized embeddings such as BERT and GPT have been shown to give significant improvement in NLP tasks. On the other hand, static embeddings such as skip-gram and GloVe still have desirable characteristics such as low computational cost, easy deployment and freedom from severe contextualized variation in representation. There has been some recent attempt enhancing the skip-gram model by adding syntactic information of context using GCN. We investigate the use of BERT embeddings instead for stronger context representation, which contains not only syntactic and surface features, but also rich knowledge from large-scale pre-training. Results show that BERT-enhanced skip-gram embeddings outperform GCN-enhanced embeddings on a range of tasks. Such embeddings also outperform recent effort distilling BERT embeddings into context-independent vectors. Yile Wang 0001, Leyang Cui, Yue Zhang 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Evaluating Commonsense in Pre-Trained Language ModelsabstractContextualized representations trained over large raw text data have given remarkable improvements for NLP tasks including question answering and reading comprehension. There have been works showing that syntactic, semantic and word sense knowledge are contained in such representations, which explains why they benefit such tasks. However, relatively little work has been done investigating commonsense knowledge contained in contextualized representations, which is crucial for human question answering and reading comprehension. We study the commonsense ability of GPT, BERT, XLNet, and RoBERTa by testing them on seven challenging benchmarks, finding that language modeling and its variants are effective objectives for promoting models' commonsense ability while bi-directional context and larger training set are bonuses. We additionally find that current models do poorly on tasks require more necessary inference steps. Finally, we test the robustness of models by making dual test cases, which are correlated so that the correct prediction of one sample should lead to correct prediction of the other. Interestingly, the models show confusion on these test cases, which suggests that they learn commonsense at the surface rather than the deep level. We release a test set, named CATs publicly, for future research. Yue Zhang 0004, Leyang Cui, Dandan Huang |
AAAI | 3 |
| 2020 | MuTual: A Dataset for Multi-Turn Dialogue ReasoningabstractNon-task oriented dialogue systems have achieved great success in recent years due to largely accessible conversation data and the development of deep learning techniques.Given a context, current systems are able to yield a relevant and fluent response, but sometimes make logical mistakes because of weak reasoning capabilities.To facilitate the conversation reasoning research, we introduce Mu-Tual, a novel dataset for Multi-Turn dialogue Reasoning, consisting of 8,860 manually annotated dialogues based on Chinese student English listening comprehension exams.Compared to previous benchmarks for non-task oriented dialogue systems, MuTual is much more challenging since it requires a model that can handle various reasoning problems.Empirical results show that state-of-the-art methods only reach 71%, which is far behind the human performance of 94%, indicating that there is ample room for improving reasoning ability.MuTual is available at https://github. com/Nealcly/MuTual. * Contribution during internship at MSRA.M: Ma'am Leyang Cui, Yu Wu 0012, Shujie Liu 0001, Yue Zhang 0004, Ming Zhou 0001 |
ACL | 1 |
| 2020 | Does Chinese BERT Encode Word Structure?abstractContextualized representations give significantly improved results for a wide range of NLP tasks.Much work has been dedicated to analyzing the features captured by representative models such as BERT.Existing work finds that syntactic, semantic and word sense knowledge are encoded in BERT.However, little work has investigated word features for character-based languages such as Chinese.We investigate Chinese BERT using both attention weight distribution statistics and probing tasks, finding that (1) word information is captured by BERT; (2) word-level features are mostly in the middle representation layers; (3) downstream tasks make different use of word features in BERT, with POS tagging and chunking relying the most on word features, and natural language inference relying the least on such features. Yile Wang 0001, Leyang Cui, Yue Zhang 0004 |
COLING | 2 |
| 2020 | Making the Best Use of Review Summary for Sentiment AnalysisabstractSentiment analysis provides a useful overview of customer review contents.Many review websites allow a user to enter a summary in addition to a full review.Intuitively, summary information may give additional benefit for review sentiment analysis.In this paper, we conduct a study to exploit methods for better use of summary information.We start by finding out that the sentimental signal distribution of a review and that of its corresponding summary are in fact complementary to each other.We thus explore various architectures to better guide the interactions between the two and propose a hierarchically-refined review-centric attention model.Empirical results show that our review-centric model can make better use of user-written summaries for review sentiment analysis, and is also more effective compared to existing methods when the user summary is replaced with summary generated by an automatic summarization system. Sen Yang 0005, Leyang Cui, Yue Zhang 0004 |
COLING | 2 |
| 2020 | What Have We Achieved on Text Summarization?abstractDeep learning has led to significant improvement in text summarization with various methods investigated and improved ROUGE scores reported over the years.However, gaps still exist between summaries produced by automatic summarizers and human professionals.Aiming to gain more understanding of summarization systems with respect to their strengths and limits on a fine-grained syntactic and semantic level, we consult the Multidimensional Quality Metric 1 (MQM) and quantify 8 major sources of errors on 10 representative summarization models manually.Primarily, we find that 1) under similar settings, extractive summarizers are in general better than their abstractive counterparts thanks to strength in faithfulness and factual-consistency; 2) milestone techniques such as copy, coverage and hybrid extractive/abstractive methods do bring specific improvements but also demonstrate limitations; 3) pre-training techniques, and in particular sequence-to-sequence pre-training, are highly effective for improving text summarization, with BART giving the best results. Dandan Huang, Leyang Cui, Sen Yang 0005, Guangsheng Bao, Yue Zhang 0004 |
EMNLP (1) | 2 |
| 2020 | LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical ReasoningabstractMachine reading is a fundamental task for testing the capability of natural language understand- ing, which is closely related to human cognition in many aspects. With the rising of deep learning techniques, algorithmic models rival human performances on simple QA, and thus increasingly challenging machine reading datasets have been proposed. Though various challenges such as evidence integration and commonsense knowledge have been integrated, one of the fundamental capabilities in human reading, namely logical reasoning, is not fully investigated. We build a comprehensive dataset, named LogiQA, which is sourced from expert-written questions for testing human Logical reasoning. It consists of 8,678 QA instances, covering multiple types of deductive reasoning. Results show that state-of-the-art neural models perform by far worse than human ceiling. Our dataset can also serve as a benchmark for reinvestigating logical AI under the deep learning NLP setting. The dataset is freely available at https://github.com/lgw863/LogiQA-dataset. Jian Liu 0030, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang 0001, Yue Zhang 0004 |
IJCAI | 2 |
| 2019 | Hierarchically-Refined Label Attention Network for Sequence LabelingabstractLeyang Cui, Yue Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Leyang Cui, Yue Zhang 0004 |
EMNLP/IJCNLP (1) | 1 |