Yang Liu 0124

dblp:51/3710-124 · DBLP profile ↗
← Back
29ranked-venue papers
9as first author
16since 2021 · last 2026
0009-0004-5641-5114ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 9 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Enhancing Lexical Relation Mining with Structured Sememe Knowledge
abstract
Lexical Relation Mining (LRM) aims to identify and classify lexical relations between word pairs.In this paper, we focus on two subtypes of LRM: Lexical Relation Classification (LRC) and Lexical Entailment (LE).Existing top-performing methods for them rely heavily on Pre-trained Language Models (PLMs) yet fail to distinguish nuanced lexical relations.From a linguistic perspective, intralexical treestructured sememe information can reflect interlexical relations.Inspired by this, we are motivated to explore leveraging such structured knowledge to enhance LRC and LE.We first propose an automated Sememe Tree Construction (STC) pipeline to predict sememe trees; Then, we present the SememeLRM method to fully leverage structured sememe knowledge; Experimental results show that it achieves a notable 1.6% improvement on average across benchmarks, even outperforming Large Language Model (LLM)-based methods that contain 20 times more parameters on most benchmarks.Further results also suggest that sememe trees predicted by our pipeline can rival the gold-standard in HowNet, extending their applicability to lexico-semantic computing.Overall, this paper presents a potentially generalizable framework for leveraging complete sememe trees and makes significant progress, helping to unlock the value of such intralexical knowledge in downstream tasks 1 .
Hansi Wang, Qiliang Liang, Yang Liu 0124
ACL (1)4
2025 How Sememic Components Can Benefit Link Prediction for Lexico-Semantic Knowledge Graphs?
abstract
Link Prediction (LP) aims to predict missing triple information within a Knowledge Graph (KG).Existing LP methods have sought to improve the performance by integrating structural and textual information.However, for lexico-semantic KGs designed to document fine-grained sense distinctions, these types of information may not be sufficient to support effective LP.From a linguistic perspective, word senses within lexico-semantic relations usually show systematic differences in their sememic components.In light of this, we are motivated to enhance LP with sememe knowledge.We first construct a Sememe Prediction (SP) dataset, SememeDef, for learning such knowledge, and two Chinese datasets, HN7 and CWN5, for LP evaluation; Then, we propose a method, SememeLP, to leverage this knowledge for LP fully.It consistently and significantly improves the LP performance in both English and Chinese, achieving SOTA MRR of 75.1%, 80.5%, and 77.1% on WN18RR, HN7, and CWN5, respectively; Finally, an in-depth analysis is conducted, making clear how sememic components can benefit LP for lexico-semantic KGs, which provides promising progress for the completion of them 1 .
Hansi Wang, Qiliang Liang, Yang Liu 0124
EMNLP4
2024 Disambiguate Words like Composing Them: A Morphology-Informed Approach to Enhance Chinese Word Sense Disambiguation
abstract
In parataxis languages like Chinese, word meanings are highly correlated with morphological knowledge, which can help to disambiguate word senses.However, in-depth exploration of morphological knowledge in previous word sense disambiguation (WSD) methods is still lacking due to the absence of publicly available resources.In this paper, we are motivated to enhance Chinese WSD with full morphological knowledge, including both word-formations and morphemes.We first construct the largest releasable Chinese WSD resources, including the lexico-semantic inventories MorInv and WrdInv, a Chinese WSD dataset MiCLS, and an out-of-volcabulary (OOV) test set.Then, we propose a model, MorBERT, to fully leverage this morphology-informed knowledge for Chinese WSD and achieve a SOTA F1 of 92.18% on MiCLS dataset.Finally, we demonstrated the model's robustness in low-resource settings and generalizability to OOV senses.These resources and methods may bring new insights into and solutions for various downstream tasks in both computational and humanistic fields 1 .
Qiliang Liang, Yaqi Yin, Hansi Wang, Yang Liu 0124
ACL (1)5
2024 PEARL: Prompting Large Language Models to Plan and Execute Actions Over Long Documents
abstract
Simeng Sun, Yang Liu, Shuohang Wang, Dan Iter, Chenguang Zhu, Mohit Iyyer. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Simeng Sun, Yang Liu 0124, Shuohang Wang, Dan Iter, Chenguang Zhu 0001, Mohit Iyyer
EACL (1)2
2023 UniSumm and SummZoo: Unified Model and Diverse Benchmark for Few-Shot Summarization
abstract
Yulong Chen, Yang Liu, Ruochen Xu, Ziyi Yang, Chenguang Zhu, Michael Zeng, Yue Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yulong Chen 0001, Yang Liu 0124, Ruochen Xu, Ziyi Yang 0011, Chenguang Zhu 0001, Michael Zeng 0001, Yue Zhang 0004
ACL (1)2
2023 Z-Code++: A Pre-trained Language Model Optimized for Abstractive Summarization
abstract
Pengcheng He, Baolin Peng, Song Wang, Yang Liu, Ruochen Xu, Hany Hassan, Yu Shi, Chenguang Zhu, Wayne Xiong, Michael Zeng, Jianfeng Gao, Xuedong Huang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Baolin Peng, Song Wang 0012, Yang Liu 0124, Ruochen Xu, Hany Hassan, Yu Shi 0001, Chenguang Zhu 0001, Wayne Xiong, Michael Zeng 0001, Jianfeng Gao 0001, Xuedong Huang 0001
ACL (1)4
2023 Unifying Vision, Text, and Layout for Universal Document Processing
abstract
We propose Universal Document Processing (UDOP), a foundation Document AI model which unifies text, image, and layout modalities together with varied task formats, including document understanding and generation. UDOP leverages the spatial correlation between textual content and document image to model image, text, and layout modalities with one uniform representation. With a novel Vision-Text-Layout Transformer, UDOP unifies pretraining and multi-domain downstream tasks into a prompt-based sequence generation scheme. UDOP is pretrained on both large-scale unlabeled document corpora using innovative self-supervised objectives and diverse labeled data. UDOP also learns to generate document images from text and layout modalities via masked image reconstruction. To the best of our knowledge, this is the first time in the field of document AI that one model simultaneously achieves high-quality neural document editing and content customization. Our method sets the state-of-the-art on 8 Document AI tasks, e.g., document understanding and QA, across diverse data domains like finance reports, academic papers, and web-sites. UDOP ranks first on the leaderboard of the Document Understanding Benchmark.11Code and models: https://github.com/microsoft/i-Code/tree/main/i-Code-Doc
Zineng Tang, Ziyi Yang 0011, Yuwei Fang, Yang Liu 0124, Chenguang Zhu 0001, Michael Zeng 0001, Cha Zhang, Mohit Bansal
CVPR5
2023 MACSum: Controllable Summarization with Mixed Attributes
abstract
Abstract Controllable summarization allows users to generate customized summaries with specified attributes. However, due to the lack of designated annotations of controlled summaries, existing work has to craft pseudo datasets by adapting generic summarization benchmarks. Furthermore, most research focuses on controlling single attributes individually (e.g., a short summary or a highly abstractive summary) rather than controlling a mix of attributes together (e.g., a short and highly abstractive summary). In this paper, we propose MACSum, the first human-annotated summarization dataset for controlling mixed attributes. It contains source texts from two domains, news articles and dialogues, with human-annotated summaries controlled by five designed attributes (Length, Extractiveness, Specificity, Topic, and Speaker). We propose two simple and effective parameter-efficient approaches for the new task of mixed controllable summarization based on hard prompt tuning and soft prefix tuning. Results and analysis demonstrate that hard prompt models yield the best performance on most metrics and human evaluations. However, mixed-attribute control is still challenging for summarization tasks. Our dataset and code are available at https://github.com/psunlpgroup/MACSum.
Yusen Zhang 0001, Yang Liu 0124, Ziyi Yang 0011, Yuwei Fang, Yulong Chen 0001, Dragomir R. Radev, Chenguang Zhu 0001, Michael Zeng 0001, Rui Zhang 0037
Trans. Assoc. Comput. Linguistics2
2022 DialogLM: Pre-trained Model for Long Dialogue Understanding and Summarization
abstract
Dialogue is an essential part of human communication and cooperation. Existing research mainly focuses on short dialogue scenarios in a one-on-one fashion. However, multi-person interactions in the real world, such as meetings or interviews, are frequently over a few thousand words. There is still a lack of corresponding research and powerful tools to understand and process such long dialogues. Therefore, in this work, we present a pre-training framework for long dialogue understanding and summarization. Considering the nature of long conversations, we propose a window-based denoising approach for generative pre-training. For a dialogue, it corrupts a window of text with dialogue-inspired noise, and guides the model to reconstruct this window based on the content of the remaining conversation. Furthermore, to process longer input, we augment the model with sparse attention which is combined with conventional attention in a hybrid manner. We conduct extensive experiments on five datasets of long dialogues, covering tasks of dialogue summarization, abstractive question answering and topic segmentation. Experimentally, we show that our pre-trained model DialogLM significantly surpasses the state-of-the-art models across datasets and tasks. Source code and all the pre-trained models are available on our GitHub repository (https://github.com/microsoft/DialogLM).
Ming Zhong 0005, Yang Liu 0124, Yichong Xu, Chenguang Zhu 0001, Michael Zeng 0001
AAAI2
2022 Training Data is More Valuable than You Think: A Simple and Effective Method by Retrieving from Training Data
abstract
Shuohang Wang, Yichong Xu, Yuwei Fang, Yang Liu, Siqi Sun, Ruochen Xu, Chenguang Zhu, Michael Zeng. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Shuohang Wang, Yichong Xu, Yuwei Fang, Yang Liu 0124, Ruochen Xu, Chenguang Zhu 0001, Michael Zeng 0001
ACL (1)4
2022 ParaTag: A Dataset of Paraphrase Tagging for Fine-Grained Labels, NLG Evaluation, and Data Augmentation
abstract
Paraphrase identification has been formulated as a binary classification task to decide whether two sentences hold a paraphrase relationship.Existing paraphrase datasets only annotate a binary label for each sentence pair.However, after a systematical analysis of existing paraphrase datasets, we found that the degree of paraphrase cannot be well characterized by a single binary label.And the criteria of paraphrase are not even consistent within the same dataset.We hypothesize that such issues would limit the effectiveness of paraphrase models trained on these data.To this end, we propose a novel fine-grained paraphrase annotation schema that labels the minimum spans of tokens in a sentence that don't have the corresponding paraphrases in the other sentence.Under this setting, we frame paraphrasing as a sequence tagging task.We collect 30k sentence pairs in English with the new annotation schema, resulting in the ParaTag dataset.In addition to reporting baseline results on ParaTag using state-of-art language models, we show that ParaTag is especially useful for training an automatic scorer for language generation evaluation.Finally, we train a paraphrase generation model from ParaTag and achieve better data augmentation performance on the GLUE benchmark than other public paraphrasing datasets.1
Shuohang Wang, Ruochen Xu, Yang Liu 0124, Chenguang Zhu 0001, Michael Zeng 0001
EMNLP3
2021 DialogSum Challenge: Summarizing Real-Life Scenario Dialogues
abstract
We propose a shared task on summarizing reallife scenario dialogues, DialogSum Challenge, to encourage researchers to address challenges in dialogue summarization, which has been less studied by the summarization community.Real-life scenario dialogue summarization has a wide potential application prospect in chatbot and personal assistant.It contains unique challenges such as special discourse structure, coreference, pragmatics and social common sense, which require specific representation learning technologies to deal with.We carefully annotate a large-scale dialogue summarization dataset based on multiple public dialogue corpus, opening the door to all kinds of summarization models.
Yulong Chen 0001, Yang Liu 0124, Yue Zhang 0004
INLG2
2021 Noisy Self-Knowledge Distillation for Text Summarization
abstract
In this paper we apply self-knowledge distillation to text summarization which we argue can alleviate problems with maximumlikelihood training on single reference and noisy datasets.Instead of relying on one-hot annotation labels, our student summarization model is trained with guidance from a teacher which generates smoothed labels to help regularize training.Furthermore, to better model uncertainty during training, we introduce multiple noise signals for both teacher and student models.We demonstrate experimentally on three benchmarks that our framework boosts the performance of both pretrained and nonpretrained summarizers achieving state-of-theart results. 1
Yang Liu 0124, Sheng Shen 0001, Mirella Lapata
NAACL-HLT1
2021 Decompose, Fuse and Generate: A Formation-Informed Method for Chinese Definition Generation
abstract
Hua Zheng, Damai Dai, Lei Li, Tianyu Liu, Zhifang Sui, Baobao Chang, Yang Liu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Damai Dai, Lei Li 0039, Tianyu Liu 0001, Zhifang Sui, Baobao Chang, Yang Liu 0124
NAACL-HLT7
2021 QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization
abstract
Ming Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, Dragomir Radev. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Ming Zhong 0005, Da Yin, Tao Yu 0009, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Awadallah 0001, Asli Celikyilmaz, Yang Liu 0124, Xipeng Qiu, Dragomir R. Radev
NAACL-HLT9
2021 MediaSum: A Large-scale Media Interview Dataset for Dialogue Summarization
abstract
This paper introduces MEDIASUM 1 , a largescale media interview dataset consisting of 463.6K transcripts with abstractive summaries.To create this dataset, we collect interview transcripts from NPR and CNN and employ the overview and topic descriptions as summaries.Compared with existing public corpora for dialogue summarization, our dataset is an order of magnitude larger and contains complex multi-party conversations from multiple domains.We conduct statistical analysis to demonstrate the unique positional bias exhibited in the transcripts of televised and radioed interviews.We also show that MEDIASUM can be used in transfer learning to improve a model's performance on other dialogue summarization tasks.
Chenguang Zhu 0001, Yang Liu 0124, Michael Zeng 0001
NAACL-HLT2
2019 Dependency Grammar Induction with a Neural Variational Transition-Based Parser
abstract
Dependency grammar induction is the task of learning dependency syntax without annotated training data. Traditional graph-based models with global inference achieve state-ofthe-art results on this task but they require O(n3) run time. Transition-based models enable faster inference with O(n) time complexity, but their performance still lags behind. In this work, we propose a neural transition-based parser for dependency grammar induction, whose inference procedure utilizes rich neural features with O(n) time complexity. We train the parser with an integration of variational inference, posterior regularization and variance reduction techniques. The resulting framework outperforms previous unsupervised transition-based dependency parsers and achieves performance comparable to graph-based models, both on the English Penn Treebank and on the Universal Dependency Treebank. In an empirical comparison, we show that our approach substantially increases parsing speed over graphbased models.
Bowen Li 0002, Jianpeng Cheng 0001, Yang Liu 0124, Frank Keller
AAAI3
2019 Implanting Rational Knowledge into Distributed Representation at Morpheme Level
abstract
Previously, researchers paid no attention to the creation of unambiguous morpheme embeddings independent from the corpus, while such information plays an important role in expressing the exact meanings of words for parataxis languages like Chinese. In this paper, after constructing the Chinese lexical and semantic ontology based on word-formation, we propose a novel approach to implanting the structured rational knowledge into distributed representation at morpheme level, naturally avoiding heavy disambiguation in the corpus. We design a template to create the instances as pseudo-sentences merely from the pieces of knowledge of morphemes built in the lexicon. To exploit hierarchical information and tackle the data sparseness problem, the instance proliferation technique is applied based on similarity to expand the collection of pseudo-sentences. The distributed representation for morphemes can then be trained on these pseudo-sentences using word2vec. For evaluation, we validate the paradigmatic and syntagmatic relations of morpheme embeddings, and apply the obtained embeddings to word similarity measurement, achieving significant improvements over the classical models by more than 5 Spearman scores or 8 percentage points, which shows very promising prospects for adoption of the new source of knowledge.
Zi Lin, Yang Liu 0124
AAAI2
2019 Hierarchical Transformers for Multi-Document Summarization
abstract
In this paper, we develop a neural summarization model which can effectively process multiple input documents and distill Transformer architecture with the ability to encode documents in a hierarchical manner. We represent cross-document relationships via an attention mechanism which allows to share information as opposed to simply concatenating text spans and processing them as a flat sequence. Our model learns latent dependencies among textual units, but can also take advantage of explicit graph representations focusing on similarity or discourse relations. Empirical results on the WikiSum dataset demonstrate that the proposed architecture brings substantial improvements over several strong baselines.
Yang Liu 0124, Mirella Lapata
ACL (1)1
2019 Generating Summaries with Topic Templates and Structured Convolutional Decoders
abstract
Existing neural generation approaches create multi-sentence text as a single sequence.In this paper we propose a structured convolutional decoder that is guided by the content structure of target summaries.We compare our model with existing sequential decoders on three data sets representing different domains.Automatic and human evaluation demonstrate that our summaries have better content coverage.
Laura Perez-Beltrachini, Yang Liu 0124, Mirella Lapata
ACL (1)2
2019 Text Summarization with Pretrained Encoders
abstract
Yang Liu, Mirella Lapata. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yang Liu 0124, Mirella Lapata
EMNLP/IJCNLP (1)1
2018 Structured Alignment Networks for Matching Sentences
abstract
Many tasks in natural language processing involve comparing two sentences to compute some notion of relevance, entailment, or similarity.Typically, this comparison is done either at the word level or at the sentence level, with no attempt to leverage the inherent structure of the sentence.When sentence structure is used for comparison, it is obtained during a non-differentiable pre-processing step, leading to propagation of errors.We introduce a model of structured alignments between sentences, showing how to compare two sentences by matching their latent structures.Using a structured attention mechanism, our model matches candidate spans in the first sentence to candidate spans in the second sentence, simultaneously discovering the tree structure of each sentence.Our model is fully differentiable and trained only on the matching objective.We evaluate this model on two tasks, entailment detection and answer sentence selection, and find that modeling latent tree structures results in superior performance.Analysis of the learned sentence structures shows they can reflect some syntactic phenomena.
Yang Liu 0124, Matt Gardner 0001, Mirella Lapata
EMNLP1
2018 Learning Structured Text Representations
abstract
In this paper, we focus on learning structure-aware document representations from data without recourse to a discourse parser or additional annotations. Drawing inspiration from recent efforts to empower neural networks with a structural bias (Cheng et al., 2016; Kim et al., 2017), we propose a model that can encode a document while automatically inducing rich structural dependencies. Specifically, we embed a differentiable non-projective parsing algorithm into a neural model and use attention mechanisms to incorporate the structural biases. Experimental evaluations across different tasks and datasets show that the proposed model achieves state-of-the-art results on document modeling tasks while inducing intermediate structures which are both interpretable and meaningful.
Yang Liu 0124, Mirella Lapata
Trans. Assoc. Comput. Linguistics1
2017 Learning Contextually Informed Representations for Linear-Time Discourse Parsing
abstract
Recent advances in RST discourse parsing have focused on two modeling paradigms: (a) high order parsers which jointly predict the tree structure of the discourse and the relations it encodes; or (b) lineartime parsers which are efficient but mostly based on local features.In this work, we propose a linear-time parser with a novel way of representing discourse constituents based on neural networks which takes into account global contextual information and is able to capture long-distance dependencies.Experimental results show that our parser obtains state-of-the art performance on benchmark datasets, while being efficient (with time complexity linear in the number of sentences in the document) and requiring minimal feature engineering.
Yang Liu 0124, Mirella Lapata
EMNLP1
2017 Twitter summarization with social-temporal context
Ruifang He, Yang Liu 0124, Guangchuan Yu, Jiliang Tang, Qinghua Hu, Jianwu Dang 0001
World Wide Web2
2016 Implicit Discourse Relation Classification via Multi-Task Neural Networks
abstract
Without discourse connectives, classifying implicit discourse relations is a challenging task and a bottleneck for building a practical discourse parser. Previous research usually makes use of one kind of discourse framework such as PDTB or RST to improve the classification performance on discourse relations. Actually, under different discourse annotation frameworks, there exist multiple corpora which have internal connections. To exploit the combination of different discourse corpora, we design related discourse classification tasks specific to a corpus, and propose a novel Convolutional Neural Network embedded multi-task learning system to synthesize these tasks by learning both unique and shared representations for each task. The experimental results on the PDTB implicit discourse relation classification task demonstrate that our model achieves significant gains over baseline systems.
Yang Liu 0124, Sujian Li, Xiaodong Zhang 0022, Zhifang Sui
AAAI1
2016 Recognizing Implicit Discourse Relations via Repeated Reading: Neural Networks with Multi-Level Attention
abstract
Recognizing implicit discourse relations is a challenging but important task in the field of Natural Language Processing.For such a complex text processing task, different from previous studies, we argue that it is necessary to repeatedly read the arguments and dynamically exploit the efficient features useful for recognizing discourse relations.To mimic the repeated reading strategy, we propose the neural networks with multi-level attention (NNMA), combining the attention mechanism and external memories to gradually fix the attention on some specific words helpful to judging the discourse relations.Experiments on the PDTB dataset show that our proposed method achieves the state-ofart results.The visualization of the attention weights also illustrates the progress that our model observes the arguments on each level and progressively locates the important words.
Yang Liu 0124, Sujian Li
EMNLP1
2016 Relation Classification Via Modeling Augmented Dependency Paths
abstract
Previous research on relation classification has verified the effectiveness of using dependency shortest paths or dependency subtrees. How to efficiently unify these two kinds of dependency information in relation classification is still an open problem. In this paper, we propose a novel structure, termed augmented dependency path (ADP), which is composed of the shortest dependency path between two entities and the subtrees attached to the shortest path. To exploit the semantic representation behind the ADP structure, we develop the dependency-based neural networks (DepNN) model which combines the advantages of the recursive neural network (RNN) and the convolutional neural network (CNN). In DepNN, RNN is designed to model the dependency subtrees since it is good at capturing the hierarchical structures. Then, the semantic representation in subtrees is passed to the nodes on the shortest path and CNN is used to get the most important features on the ADP. Experiments on the SemEval-2010 dataset show that the ADP structure including both the shortest dependency path and the attached subtrees is helpful to classify the semantic relations between two entities and our proposed method can achieve the state-of-the-art performance.
Yang Liu 0124, Sujian Li, Furu Wei, Heng Ji 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 A Novel Neural Topic Model and Its Supervised Extension
abstract
Topic modeling techniques have the benefits of modeling words and documents uniformly under a probabilistic framework. However, they also suffer from the limitations of sensitivity to initialization and unigram topic distribution, which can be remedied by deep learning techniques. To explore the combination of topic modeling and deep learning techniques, we first explain the standard topic modelfrom the perspective of a neural network. Based on this, we propose a novel neural topic model (NTM) where the representation of words and documents are efficiently and naturally combined into a uniform framework. Extending from NTM, we can easily add a label layer and propose the supervised neural topic model (sNTM) to tackle supervised tasks. Experiments show that our models are competitive in both topic discovery and classification/regression tasks.
Ziqiang Cao, Sujian Li, Yang Liu 0124, Wenjie Li 0002, Heng Ji 0001
AAAI3