Nianwen Xue

dblp:01/2933 · DBLP profile ↗
← Back
58ranked-venue papers
16as first author
12since 2021 · last 2026
0000-0002-4364-3618ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 55 · 16 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 3Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 Modal Dependency Parsing as Structured Prediction over Source-Cue Scope
abstract
Modal dependency parsing-the task of identifying a semantic graph that represents who is responsible for an event-centered claim and with what degree of certainty-relies on recognizing source-introducing cues and correctly linking them to their associated content.However, prior work has largely focused on identifying sources only, treating cue expressions and their modal coverage as auxiliary signals.In this work, we propose a structured prediction framework that leverages large language models (LLMs) to explicitly identify source-cue pairs as well as their respective scope, which together define the modal contexts governing downstream source attribution for events.By concentrating learning at the source-cue level and constraining event-level decisions to a small, scope-defined candidate set, our topdown approach enables inference over a reduced candidate space in long, event-rich documents.Experiments show this approach surpasses prior state-of-the-art results by 3 and 4% for English and Chinese datasets, respectively.
Jayeol Chun, Nianwen Xue
ACL (1)2
2026 Reframing Responsibility: Framing-Aware Event Causality Identification
abstract
Causal explanations in political narratives are often framed and contested.Different sources may explain the same event by assigning responsibility to different actors, expressing different levels of certainty.Standard Event Causality Identification (ECI) focuses on detecting causal links and does not capture these distinctions.We introduce Framing-Aware Event Causality Identification (FrECI), a framing-aware extension of ECI that models causal explanations as structured claims including responsibility targets, evaluative framing, source type, and epistemic modality grounded in established framing theories.We construct a multilingual dataset aligned across English, Chinese, and Arabic narratives using shared event anchors.We evaluate FrECI using prompt-based large language model baselines and supervised neural models.Results show that prompt-based baselines struggle to recover complete framed causal claims, while joint supervised models perform substantially better.Finally, we demonstrate that FrECI enables quantitative analysis of divergent causal attribution across narratives.The dataset and code are publicly available.1
Jin Zhao 0009, Xinrui Hu, Nianwen Xue
ACL (1)4
2025 Seeing the Same Story Differently: Framing-Divergent Event Coreference for Computational Framing Analysis
abstract
News articles often describe the same realworld event in strikingly different ways, shaping perception through framing rather than factual disagreement.However, traditional computational framing approaches often rely on coarse-grained topic classification, limiting their ability to capture subtle, event-level differences in how the same occurrences are presented across sources.We introduce Framingdivergent Event Coreference (FRECO), a novel task that identifies pairs of event mentions referring to the same underlying occurrence but differing in framing across documents to provide a event-centric lens for computational framing analysis.To support this task, we construct the high-agreement and diverse FRECO corpus.We evaluate the FRECO task on the corpus through supervised and preference-based tuning of large language models, providing strong baseline performance.To scale beyond the annotated data, we develop a bootstrapped mining pipeline that iteratively expands the training set with high-confidence FRECO pairs.Our approach enables scalable, interpretable analysis of how media frame the same events differently, offering a new lens for contrastive framing analysis at the event level.The dataset and code will be made publicly available.1
Jin Zhao 0009, Xinrui Hu, Nianwen Xue
EMNLP3
2025 Beyond Benchmarks: Building a Richer Cross-Document Event Coreference Dataset with Decontextualization
abstract
Jin Zhao, Jingxuan Tu, Bingyang Ye, Xinrui Hu, Nianwen Xue, James Pustejovsky. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Jin Zhao 0009, Jingxuan Tu, Bingyang Ye, Xinrui Hu, Nianwen Xue, James Pustejovsky
NAACL (Long Papers)5
2024 Building a Broad Infrastructure for Uniform Meaning Representations
abstract
This paper reports the first release of the UMR (Uniform Meaning Representation) data set. UMR is a graph-based meaning representation formalism consisting of a sentence-level graph and a document-level graph. The sentence-level graph represents predicate-argument structures, named entities, word senses, aspectuality of events, as well as person and number information for entities. The document-level graph represents coreferential, temporal, and modal relations that go beyond sentence boundaries. UMR is designed to capture the commonalities and variations across languages and this is done through the use of a common set of abstract concepts, relations, and attributes as well as concrete concepts derived from words from invidual languages. This UMR release includes annotations for six languages (Arapaho, Chinese, English, Kukama, Navajo, Sanapana) that vary greatly in terms of their linguistic properties and resource availability. We also describe on-going efforts to enlarge this data set and extend it to other genres and modalities. We also briefly describe the available infrastructure (UMR annotation guidelines and tools) that others can use to create similar data sets.
Julia Bonn, Matthew J. Buchholz, Jayeol Chun, Andrew Cowell, William Croft 0001, Lukas Denk, Sijia Ge, Jan Hajic 0001, Kenneth Lai, James H. Martin, Skatje Myers, Alexis Palmer, Martha Palmer, Claire Benet Post, James Pustejovsky, Kristine Stenzel, Haibo Sun, Zdenka Uresová, Rosa Vallejos, Jens E. L. Van Gysel, Meagan Vigus, Nianwen Xue, Jin Zhao 0009
LREC/COLING22
2024 Anchor and Broadcast: An Efficient Concept Alignment Approach for Evaluation of Semantic Graphs
abstract
In this paper, we present AnCast, an intuitive and efficient tool for evaluating graph-based meaning representations (MR). AnCast implements evaluation metrics that are well understood in the NLP community, and they include concept F1, unlabeled relation F1, labeled relation F1, and weighted relation F1. The efficiency of the tool comes from a novel anchor broadcast alignment algorithm that is not subject to the trappings of local maxima. We show through experimental results that the AnCast score is highly correlated with the widely used Smatch score, but its computation takes only about 40% the time.
Haibo Sun, Nianwen Xue
LREC/COLING2
2024 Media Attitude Detection via Framing Analysis with Events and their Relations
abstract
Framing is used to present some selective aspects of an issue and make them more salient, which aims to promote certain values, interpretations, or solutions (Entman, 1993).This study investigates the nuances of media framing on public perception and understanding by examining how events are presented within news articles.Unlike previous research that primarily focused on word choice as a framing device, this work explores the comprehensive narrative construction through events and their relations.Our method integrates event extraction, Cross-Document Event Coreference (CDEC), and causal relationship among events to extract framing devices employed by the media to assess their role in framing the narrative.We evaluate our approach with a media attitude detection task and show that the use of event mentions, event cluster descriptors, and their causal relations effectively captures the subtle nuances of framing, thereby providing deeper insights into the attitudes conveyed by news articles.The experimental results show the framing device models surpass the baseline models and offer a more detailed and explainable analysis of media framing effects.We make the source code and dataset publicly available.1
Jin Zhao 0009, Jingxuan Tu, Nianwen Xue
EMNLP4
2023 Cross-Document Event Coreference Resolution: Instruct Humans or Instruct GPT?
abstract
This paper explores utilizing Large Language Models (LLMs) to perform Cross-Document Event Coreference Resolution (CDEC) annotations and evaluates how they fare against human annotators with different levels of training.Specifically, we formulate CDEC as a multiclass classification problem on pairs of events that are represented as decontextualized sentences, and compare the predictions of GPT-4 with the judgment of fully trained annotators and crowdworkers on the same dataset.Our study indicates that GPT-4 with zero-shot learning outperformed crowd-workers by a large margin and exhibits a level of performance comparable to trained annotators.Upon closer analysis, GPT-4 also exhibits tendencies of being overly confident, and forcing annotation decisions even when such decisions are not warranted due to insufficient information.Our results have implications on how to perform complicated annotations such as CDEC in the age of LLMs, and show that the best way to acquire such annotations might be to combine the strengths of LLMs and trained human annotators in the annotation process, and using untrained or undertrained crowdworkers is no longer a viable option to acquire high-quality data to advance the state of the art for such problems.We make our source and data publicly available.1
Nianwen Xue, Bonan Min
CoNLL2
2023 A Kind Introduction to Lexical and Grammatical Aspect, with a Survey of Computational Approaches
abstract
Aspectual meaning refers to how the internal temporal structure of situations is presented.This includes whether a situation is described as a state or as an event, whether the situation is finished or ongoing, and whether it is viewed as a whole or with a focus on a particular phase.This survey gives an overview of computational approaches to modeling lexical and grammatical aspect along with intuitive explanations of the necessary linguistic concepts and terminology.In particular, we describe the concepts of stativity, telicity, habituality, perfective and imperfective, as well as influential inventories of eventuality and situation types.Aspect is a crucial component of semantics, especially for precise reporting of the temporal structure of situations, and future NLP approaches need to be able to handle and evaluate it systematically.
Annemarie Friedrich, Nianwen Xue, Alexis Palmer
EACL2
2022 Modal Dependency Parsing via Language Model Priming
abstract
The task of modal dependency parsing aims to parse a text into its modal dependency structure, which is a representation for the factuality of events in the text.We design a modal dependency parser that is based on priming pre-trained language models, and evaluate the parser on two data sets.Compared to baselines, we show an improvement of 2.6% in F-score for English and 4.6% for Chinese.To the best of our knowledge, this is also the first work on Chinese modal dependency parsing.
Jiarui Yao, Nianwen Xue, Bonan Min
NAACL-HLT2
2021 A Joint Model for Dropped Pronoun Recovery and Conversational Discourse Parsing in Chinese Conversational Speech
abstract
Jingxuan Yang, Kerui Xu, Jun Xu, Si Li, Sheng Gao, Jun Guo, Nianwen Xue, Ji-Rong Wen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Kerui Xu, Jun Xu 0001, Si Li 0001, Sheng Gao 0001, Jun Guo 0002, Nianwen Xue, Ji-Rong Wen
ACL/IJCNLP (1)7
2021 Factuality Assessment as Modal Dependency Parsing
abstract
Jiarui Yao, Haoling Qiu, Jin Zhao, Bonan Min, Nianwen Xue. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jiarui Yao, Haoling Qiu, Bonan Min, Nianwen Xue
ACL/IJCNLP (1)5
2020 Annotating Temporal Dependency Graphs via Crowdsourcing
abstract
We present the construction of a corpus of 500 Wikinews articles annotated with temporal dependency graphs (TDGs) that can be used to train systems to understand temporal relations in text.We argue that temporal dependency graphs, built on previous research on narrative times and temporal anaphora, provide a representation scheme that achieves a good balance between completeness and practicality in temporal annotation.We also provide a crowdsourcing strategy to annotate TDGs, and demonstrate the feasibility of this approach with an evaluation of the quality of the annotation, and the utility of the resulting data set by training a machine learning model on this data set.This data set is publicly available 1 .
Jiarui Yao, Haoling Qiu, Bonan Min, Nianwen Xue
EMNLP (1)4
2020 Abstract Meaning Representation for MWE: A study of the mapping of aspectuality based on Mandarin light verb jiayi
Nianwen Xue, Chu-Ren Huang
PACLIC2
2019 SMART: A Stratified Machine Reading Test
Jiarui Yao, Minxuan Feng, Haixia Feng, Nianwen Xue
NLPCC (1)6
2019 A Survey of Discourse Representations for Chinese Discourse Annotation
abstract
A key element in computational discourse analysis is the design of a formal representation for the discourse structure of a text. With machine learning being the dominant method, it is important to identify a discourse representation that can be used to perform large-scale annotation. This survey provides a systematic analysis of existing discourse representation theories to evaluate whether they are suitable for annotation of Chinese text. Specifically, the two properties, expressiveness and practicality, are introduced to compare the representations of theories based on rhetorical relations and the representations of theories based on entity relations. The comparison systematically reveals linguistic and computational characteristics of the theories. After that, we conclude that none of the existing theories are quite suitable for scalable Chinese discourse annotation because they are not both expressive and practical. Therefore, a new discourse representation needs to be proposed, which should balance the expressiveness and practicality, and cover rhetorical relations and entity relations. Inspired by the conclusions, this survey discusses some preliminary proposals on how to represent the discourse structure that are worth pursuing.
Xiaomian Kang, Chengqing Zong, Nianwen Xue
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2018 Neural Ranking Models for Temporal Dependency Structure Parsing
abstract
We design and build the first neural temporal dependency parser.It utilizes a neural ranking model with minimal feature engineering, and parses time expressions and events in a text into a temporal dependency tree structure.We evaluate our parser on two domains: news reports and narrative stories.In a parsing-only evaluation setup where gold time expressions and events are provided, our parser reaches 0.81 and 0.70 f-score on unlabeled and labeled parsing respectively, a result that is very competitive against alternative approaches.In an end-to-end evaluation setup where time expressions and events are automatically recognized, our parser beats two strong baselines on both data domains.Our experimental results and discussions shed light on the nature of temporal dependency structures in different domains and provide insights that we believe will be valuable to future research in this area.
Nianwen Xue
EMNLP2
2018 Structured Interpretation of Temporal Relations
Nianwen Xue
LREC2
2017 Addressing the Data Sparsity Issue in Neural AMR Parsing
abstract
Neural attention models have achieved great success in different NLP tasks.However, they have not fulfilled their promise on the AMR parsing task due to the data sparsity issue.In this paper, we describe a sequence-to-sequence model for AMR parsing and present different ways to tackle the data sparsity problem.We show that our methods achieve significant improvement over a baseline neural attention model and our results are also competitive against state-of-the-art systems that do not use extra linguistic resources.
Xiaochang Peng, Daniel Gildea, Nianwen Xue
EACL (1)4
2017 A Systematic Study of Neural Discourse Models for Implicit Discourse Relation
abstract
Inferring implicit discourse relations in natural language text is the most difficult subtask in discourse parsing.Many neural network models have been proposed to tackle this problem.However, the comparison for this task is not unified, so we could hardly draw clear conclusions about the effectiveness of various architectures.Here, we propose neural network models that are based on feedforward and long-short term memory architecture and systematically study the effects of varying structures.To our surprise, the best-configured feedforward architecture outperforms LSTM-based model in most cases despite thorough tuning.Further, we compare our best feedforward system with competitive convolutional and recurrent networks and find that feedforward can actually be more effective.For the first time for this task, we compile and publish outputs from previous neural and nonneural systems to establish the standard for further comparison.
Attapol Rutherford, Vera Demberg, Nianwen Xue
EACL (1)3
2017 Getting the Most out of AMR Parsing
abstract
This paper proposes to tackle the AMR parsing bottleneck by improving two components of an AMR parser: concept identification and alignment.We first build a Bidirectional LSTM based concept identifier that is able to incorporate richer contextual information to learn sparse AMR concept labels.We then extend an HMM-based word-to-concept alignment model with graph distance distortion and a rescoring method during decoding to incorporate the structural information in the AMR graph.We show integrating the two components into an existing AMR parser results in consistently better performance over the state of the art on various datasets.
Nianwen Xue
EMNLP2
2017 Translation Divergences in Chinese-English Machine Translation: An Empirical Investigation
abstract
In this article, we conduct an empirical investigation of translation divergences between Chinese and English relying on a parallel treebank. To do this, we first devise a hierarchical alignment scheme where Chinese and English parse trees are aligned in a way that eliminates conflicts and redundancies between word alignments and syntactic parses to prevent the generation of spurious translation divergences. Using this Hierarchically Aligned Chinese–English Parallel Treebank (HACEPT), we are able to semi-automatically identify and categorize the translation divergences between the two languages and quantify each type of translation divergence. Our results show that the translation divergences are much broader than described in previous studies that are largely based on anecdotal evidence and linguistic knowledge. The distribution of the translation divergences also shows that some high-profile translation divergences that motivate previous research are actually very rare in our data, whereas other translation divergences that have previously received little attention actually exist in large quantities. We also show that HACEPT allows the extraction of syntax-based translation rules, most of which are expressive enough to capture the translation divergences, and point out that the syntactic annotation in existing treebanks is not optimal for extracting such translation rules. We also discuss the implications of our study for attempts to bridge translation divergences by devising shared semantic representations across languages. Our quantitative results lend further support to the observation that although it is possible to bridge some translation divergences with semantic representations, other translation divergences are open-ended, thus building a semantic representation that captures all possible translation divergences may be impractical.
Dun Deng, Nianwen Xue
Comput. Linguistics2
2016 Large Multi-lingual, Multi-level and Multi-genre Annotation Corpus
Xuansong Li, Martha Palmer, Nianwen Xue, Lance A. Ramshaw, Mohamed Maamouri, Ann Bies, Kathryn Conger, Stephen Grimes, Stephanie M. Strassel
LREC3
2015 Feature Optimization for Constituent Parsing via Neural Networks
abstract
Zhiguo Wang, Haitao Mi, Nianwen Xue. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Haitao Mi, Nianwen Xue
ACL (1)3
2015 Improving the Inference of Implicit Discourse Relations via Classifying Explicit Discourse Connectives
abstract
Discourse relation classification is an important component for automatic discourse parsing and natural language understanding. The performance bottleneck of a discourse parser comes from implicit discourse relations, whose discourse connectives are not overtly present. Explicit discourse connectives can potentially be exploited to collect more training data to collect more data and boost the performance. However, using them indiscriminately has been shown to hurt the performance because not all discourse connectives can be dropped arbitrarily. Based on this insight, we investigate the interaction between discourse connectives and the discourse relations and propose the criteria for selecting the discourse connectives that can be dropped independently of the context without changing the interpretation of the discourse. Extra training data collected only by the freely omissible connectives improve the performance of the system without additional features.
Attapol Rutherford, Nianwen Xue
HLT-NAACL2
2015 A Transition-based Algorithm for AMR Parsing
abstract
We present a two-stage framework to parse a sentence into its Abstract Meaning Representation (AMR). We first use a dependency parser to generate a dependency tree for the sentence. In the second stage, we design a novel transition-based algorithm that transforms the dependency tree to an AMR graph. There are several advantages with this approach. First, the dependency parser can be trained on a training set much larger than the training set for the tree-to-graph algorithm, resulting in a more accurate AMR parser overall. Our parser yields an improvement of 5% absolute in F-measure over the best previous result. Second, the actions that we design are linguistically intuitive and capture the regularities in the mapping between the dependency structure and the AMR of a sentence. Third, our parser runs in nearly linear time in practice in spite of a worst-case complexity ofO(n 2 ).
Nianwen Xue, Sameer Pradhan
HLT-NAACL2
2014 Joint POS Tagging and Transition-based Constituent Parsing in Chinese with Non-local Features
abstract
We propose three improvements to ad-dress the drawbacks of state-of-the-art transition-based constituent parsers. First, to resolve the error propagation problem of the traditional pipeline approach, we incorporate POS tagging into the syntac-tic parsing process. Second, to allevi-ate the negative influence of size differ-ences among competing action sequences, we align parser states during beam-search decoding. Third, to enhance the pow-er of parsing models, we enlarge the fea-ture set with non-local features and semi-supervised word cluster features. Exper-imental results show that these modifica-tions improve parsing performance signif-icantly. Evaluated on the Chinese Tree-Bank (CTB), our final performance reach-es 86.3 % (F1) when trained on CTB 5.1, and 87.1 % when trained on CTB 6.0, and these results outperform all state-of-the-art parsers. 1
Nianwen Xue
ACL (1)2
2014 Building a Hierarchically Aligned Chinese-English Parallel Treebank
Dun Deng, Nianwen Xue
COLING2
2014 Discovering Implicit Discourse Relations Through Brown Cluster Pair Representation and Coreference Patterns
abstract
Sentences form coherent relations in a discourse without discourse connectives more frequently than with connectives.Senses of these implicit discourse relations that hold between a sentence pair, however, are challenging to infer.Here, we employ Brown cluster pairs to represent discourse relation and incorporate coreference patterns to identify senses of implicit discourse relations in naturally occurring text.Our system improves the baseline performance by as much as 25%.Feature analyses suggest that Brown cluster pairs and coreference patterns can reveal many key linguistic characteristics of each type of discourse relation.
Attapol Rutherford, Nianwen Xue
EACL2
2014 Automatic Inference of the Tense of Chinese Events Using Implicit Linguistic Information
abstract
We address the problem of automatically inferring the tense of events in Chinese text. We use a new corpus annotated with Chinese semantic tense information and other implicit Chinese linguistic informa-tion using a “distant annotation ” method. We propose three improvements over a rel-atively strong baseline method – a statisti-cal learning method with extensive feature engineering. First, we add two sources of implicit linguistic information as fea-tures – eventuality type and modality of an event, which are also inferred automat-ically. Second, we perform joint learning on semantic tense, eventuality type, and modality of an event. Third, we train arti-ficial neural network models for this prob-lem and compare its performance with feature-based approaches. Experimental results show considerable improvements on Chinese tense inference. Our best per-formance reaches 68.6 % in accuracy, out-performing a strong baseline method. 1
Nianwen Xue
EMNLP2
2014 Not an Interlingua, But Close: Comparison of English AMRs to Chinese and Czech
Nianwen Xue, Ondrej Bojar, Jan Hajic 0001, Martha Palmer, Zdenka Uresová, Xiuhong Zhang
LREC1
2014 Buy one get one free: Distant annotation of Chinese tense, event type and modality
Nianwen Xue
LREC1
2013 Towards Robust Linguistic Analysis using OntoNotes
Sameer Pradhan, Alessandro Moschitti, Nianwen Xue, Hwee Tou Ng, Anders Björkelund, Olga Uryupina
CoNLL3
2013 Dependency-based empty category detection via phrase structure trees
Nianwen Xue, Yaqin Yang
HLT-NAACL1
2013 Temporal relation discovery between events and temporal expressions identified in clinical narrative
Peter Anick, Pengyu Hong, Nianwen Xue
J. Biomed. Informatics4
2012 Chinese Comma Disambiguation for Discourse Analysis
Yaqin Yang, Nianwen Xue
ACL (1)2
2012 PDTB-style Discourse Annotation of Chinese Text
Yuping Zhou, Nianwen Xue
ACL (1)2
2012 Towards an Annotation Schema for Cancer Trajectory State Detection
Jeremy L. Warner, Peter Anick, Kenneth Roach, Nianwen Xue, Robin M. Joyce, Charles Safran, Pengyu Hong
AMIA4
2012 Annotating dropped pronouns in Chinese newswire text
Elizabeth Baran, Yaqin Yang, Nianwen Xue
LREC3
2012 Parallel Aligned Treebanks at LDC: New Challenges Interfacing Existing Infrastructures
Xuansong Li, Stephanie M. Strassel, Stephen Grimes, Safa Ismael, Mohamed Maamouri, Ann Bies, Nianwen Xue
LREC7
2012 A corpus of full-text journal articles is a robust evaluation tool for revealing differences in performance of biomedical natural language processing tools
abstract
BACKGROUND: We introduce the linguistic annotation of a corpus of 97 full-text biomedical publications, known as the Colorado Richly Annotated Full Text (CRAFT) corpus. We further assess the performance of existing tools for performing sentence splitting, tokenization, syntactic parsing, and named entity recognition on this corpus. RESULTS: Many biomedical natural language processing systems demonstrated large differences between their previously published results and their performance on the CRAFT corpus when tested with the publicly available models or rule sets. Trainable systems differed widely with respect to their ability to build high-performing models based on this data. CONCLUSIONS: The finding that some systems were able to train high-performing models based on this corpus is additional evidence, beyond high inter-annotator agreement, that the quality of the CRAFT corpus is high. The overall poor performance of various systems indicates that considerable work needs to be done to enable natural language processing systems to work well when the input is full-text journal articles. The CRAFT corpus provides a valuable resource to the biomedical natural language processing community for evaluation and training of new models for biomedical full text publications.
Karin Verspoor, Kevin Cohen 0001, Arrick Lanfranchi, Colin Warner, Helen L. Johnson 0001, Christophe Roeder, Jinho D. Choi, Christopher S. Funk, Yuriy Malenkiy, Miriam Eckert, Nianwen Xue, William A. Baumgartner Jr., Michael Bada, Martha Palmer, Lawrence Hunter
BMC Bioinform.11
2011 Singular or Plural? Exploiting Parallel Corpora for Chinese Number Prediction
Elizabeth Baran, Nianwen Xue
MTSummit2
2011 Steven Bird, Evan Klein and Edward Loper. Natural Language Processing with Python. O'Reilly Media, Inc 2009. ISBN: 978-0-596-51649-9
Nianwen Xue
Nat. Lang. Eng.1
2011 Introduction to the Special Issue on Chinese Language Processing
abstract
introduction Share on Introduction to the Special Issue on Chinese Language Processing Authors: Keh-Jiann Chen Institute of Information Science, Academia Sinica Institute of Information Science, Academia SinicaView Profile , Qun Liu Institute of Computing Technology, Chinese Academy of Sciences Institute of Computing Technology, Chinese Academy of SciencesView Profile , Nianwen Xue Brandeis University Brandeis UniversityView Profile , Le Sun Institute of Software, Chinese Academy of Sciences Institute of Software, Chinese Academy of SciencesView Profile Authors Info & Claims ACM Transactions on Asian Language Information ProcessingVolume 10Issue 3September 2011 Article No.: 11pp 1–3https://doi.org/10.1145/2002980.2002981Published:01 September 2011Publication History 0citation284DownloadsMetricsTotal Citations0Total Downloads284Last 12 Months3Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Keh-Jiann Chen, Qun Liu 0001, Nianwen Xue, Le Sun 0001
ACM Trans. Asian Lang. Inf. Process.3
2009 Adding semantic roles to the Chinese Treebank
abstract
Abstract We report work on adding semantic role labels to the Chinese Treebank, a corpus already annotated with phrase structures. The work involves locating all verbs and their nominalizations in the corpus, and semi-automatically adding semantic role labels to their arguments, which are constituents in a parse tree. Although the same procedure is followed, different issues arise in the annotation of verbs and nominalized predicates. For verbs, identifying their arguments is generally straightforward given their syntactic structure in the Chinese Treebank as they tend to occupy well-defined syntactic positions. Our discussion focuses on the syntactic variations in the realization of the arguments as well as our approach to annotating dislocated and discontinuous arguments. In comparison, identifying the arguments for nominalized predicates is more challenging and we discuss criteria and procedures for distinguishing arguments from non-arguments. In particular we focus on the role of support verbs as well as the relevance of event/result distinctions in the annotation of the predicate-argument structure of nominalized predicates. We also present our approach to taking advantage of the syntactic structure in the Chinese Treebank to bootstrap the predicate-argument structure annotation of verbs. Finally, we discuss the creation of a lexical database of frame files and its role in guiding predicate-argument annotation. Procedures for ensuring annotation consistency and inter-annotator agreement evaluation results are also presented.
Nianwen Xue, Martha Palmer
Nat. Lang. Eng.1
2008 Automatic Inference of the Temporal Location of Situations in Chinese Text
Nianwen Xue
EMNLP1
2008 Annotating "tense" in a Tense-less Language
Nianwen Xue, Kai-Yun Chen
LREC1
2008 Labeling Chinese Predicates with Semantic Roles
abstract
In this article we report work on Chinese semantic role labeling, taking advantage of two recently completed corpora, the Chinese PropBank, a semantically annotated corpus of Chinese verbs, and the Chinese Nombank, a companion corpus that annotates the predicate-argument structure of nominalized predicates. Because the semantic role labels are assigned to the constituents in a parse tree, we first report experiments in which semantic role labels are automatically assigned to hand-crafted parses in the Chinese Treebank. This gives us a measure of the extent to which semantic role labels can be bootstrapped from the syntactic annotation provided in the treebank. We then report experiments using automatic parses with decreasing levels of human annotation in the input to the syntactic parser: parses that use gold-standard segmentation and POS-tagging, parses that use only gold-standard segmentation, and fully automatic parses. These experiments gauge how successful semantic role labeling for Chinese can be in more realistic situations. Our results show that when hand-crafted parses are used, semantic role labeling accuracy for Chinese is comparable to what has been reported for the state-of-the-art English semantic role labeling systems trained and tested on the English PropBank, even though the Chinese PropBank is significantly smaller in size. When an automatic parser is used, however, the accuracy of our system is significantly lower than the English state of the art. This indicates that an improvement in Chinese parsing is critical to high-performance semantic role labeling for Chinese.
Nianwen Xue
Comput. Linguistics1
2006 Aligning Features with Sense Distinction Dimensions
Nianwen Xue, Jinying Chen, Martha Palmer
ACL1
2006 Annotating the Predicate-Argument Structure of Chinese Nominalizations
Nianwen Xue
LREC1
2006 Semantic role labeling of nominalized predicates in Chinese
Nianwen Xue
HLT-NAACL1
2005 Automatic Semantic Role Labeling for Chinese Verbs
Nianwen Xue, Martha Palmer
IJCAI1
2005 The Penn Chinese TreeBank: Phrase structure annotation of a large corpus
abstract
With growing interest in Chinese Language Processing, numerous NLP tools (e.g., word segmenters, part-of-speech taggers, and parsers) for Chinese have been developed all over the world. However, since no large-scale bracketed corpora are available to the public, these tools are trained on corpora with different segmentation criteria, part-of-speech tagsets and bracketing guidelines, and therefore, comparisons are difficult. As a first step towards addressing this issue, we have been preparing a large bracketed corpus since late 1998. The first two installments of the corpus, 250 thousand words of data, fully segmented, POS-tagged and syntactically bracketed, have been released to the public via LDC ( www.ldc.upenn.edu ). In this paper, we discuss several Chinese linguistic issues and their implications for our treebanking efforts and how we address these issues when developing our annotation guidelines. We also describe our engineering strategies to improve speed while ensuring annotation quality.
Nianwen Xue, Fei Xia 0004, Fu-Dong Chiou, Martha Palmer
Nat. Lang. Eng.1
2004 Calibrating Features for Semantic Role Labeling
Nianwen Xue, Martha Palmer
EMNLP1
2004 Using a Smoothing Maximum Entropy Model for Chinese Nominal Entity Tagging
Jinying Chen, Nianwen Xue, Martha Palmer
IJCNLP2
2003 Automatic predicate argument structure analysis of the Penn Chinese Treebank
abstract
Recent work in machine translation and information extraction has demonstrated the utility of a level that represents the predicate-argument structure. It would be especially useful for machine translation to have two such Proposition Banks, one for each language under consideration. A Proposition Bank for English has been developed over the last few years, and we describe here our development of a tool for facilitating the development of a Chinese Proposition Bank. We also discuss some issues specific to the Chinese Treebank that complicate the matter of mapping syntactic representation to a predicate-argument level, and report on some preliminary evaluation of the accuracy of the semantic tagging tool.
Nianwen Xue, Seth Kulick
MTSummit1
2002 Building a Large-Scale Annotated Chinese Corpus
Nianwen Xue, Fu-Dong Chiou, Martha Palmer
COLING1
2000 Developing Guidelines and Ensuring Consistency for Chinese Text Annotation
Fei Xia 0004, Martha Palmer, Nianwen Xue, Mary Ellen Okurowski, John Kovarik, Fu-Dong Chiou, Shizhe Huang, Tony Kroch, Mitchell P. Marcus
LREC3