Rui Wang 0005

dblp:w/RuiWang5 · DBLP profile ↗
← Back
39ranked-venue papers
5as first author
11since 2021 · last 2023
0009-0000-7393-6959ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2023 AD-KD: Attribution-Driven Knowledge Distillation for Language Model Compression
abstract
Knowledge distillation has attracted a great deal of interest recently to compress pre-trained language models.However, existing knowledge distillation methods suffer from two limitations.First, the student model simply imitates the teacher's behavior while ignoring the underlying reasoning.Second, these methods usually focus on the transfer of sophisticated model-specific knowledge but overlook dataspecific knowledge.In this paper, we present a novel attribution-driven knowledge distillation approach, which explores the token-level rationale behind the teacher model based on Integrated Gradients (IG) and transfers attribution knowledge to the student model.To enhance the knowledge transfer of model reasoning and generalization, we further explore multi-view attribution distillation on all potential decisions of the teacher.Comprehensive experiments are conducted with BERT on the GLUE benchmark.The experimental results demonstrate the superior performance of our approach to several state-of-the-art methods.
Siyue Wu, Hongzhan Chen, Xiaojun Quan, Qifan Wang 0001, Rui Wang 0005
ACL (1)5
2023 Unsupervised Dialogue Topic Segmentation with Topic-aware Contrastive Learning
abstract
Dialogue Topic Segmentation (DTS) plays an essential role in a variety of dialogue modeling tasks. Previous DTS methods either focus on semantic similarity or dialogue coherence to assess topic similarity for unsupervised dialogue segmentation. However, the topic similarity cannot be fully identified via semantic similarity or dialogue coherence. In addition, the unlabeled dialogue data, which contains useful clues of utterance relationships, remains underexploited. In this paper, we propose a novel unsupervised DTS framework, which learns topic-aware utterance representations from unlabeled dialogue data through neighboring utterance matching and pseudo-segmentation. Extensive experiments on two benchmark datasets (i.e., DialSeg711 and Doc2Dial) demonstrate that our method significantly outperforms the strong baseline methods. For reproducibility, we provide our code and data at: https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/dial-start.
Rui Wang 0005, Ting-En Lin, Yuchuan Wu, Min Yang 0007, Fei Huang 0002, Yongbin Li 0001
SIGIR2
2023 Multi-modal multi-hop interaction network for dialogue response generation
Jie Zhou 0015, Rui Wang 0005, Yuanbin Wu, Ming Yan 0008, Liang He 0001, Xuanjing Huang 0001
Expert Syst. Appl.3
2022 WikiDiverse: A Multimodal Entity Linking Dataset with Diversified Contextual Topics and Entity Types
abstract
Xuwu Wang, Junfeng Tian, Min Gui, Zhixu Li, Rui Wang, Ming Yan, Lihan Chen, Yanghua Xiao. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Xuwu Wang, Min Gui, Zhixu Li, Rui Wang 0005, Ming Yan 0008, Yanghua Xiao
ACL (1)5
2022 Dial2vec: Self-Guided Contrastive Learning of Unsupervised Dialogue Embeddings
abstract
In this paper, we introduce the task of learning unsupervised dialogue embeddings.Trivial approaches such as combining pre-trained word or sentence embeddings and encoding through pre-trained language models (PLMs) have been shown to be feasible for this task.However, these approaches typically ignore the conversational interactions between interlocutors, resulting in poor performance.To address this issue, we proposed a selfguided contrastive learning approach named dial2vec.Dial2vec considers a dialogue as an information exchange process.It captures the conversational interaction patterns between interlocutors and leverages them to guide the learning of the embeddings corresponding to each interlocutor.The dialogue embedding is obtained by an aggregation of the embeddings from all interlocutors.To verify our approach, we establish a comprehensive benchmark consisting of six widely-used dialogue datasets.We consider three evaluation tasks: domain categorization, semantic relatedness, and dialogue retrieval.Dial2vec achieves on average 8.7, 9.0, and 13.8 points absolute improvements in terms of purity, Spearman's correlation, and mean average precision (MAP) over the strongest baseline on the three tasks respectively.Further analysis shows that dial2vec obtains informative and discriminative embeddings for both interlocutors under the guidance of the conversational interactions and achieves the best performance when aggregating them through the interlocutor-level pooling strategy.
Rui Wang 0005, Yongbin Li 0001, Fei Huang 0002
EMNLP2
2022 Sentiment-aware multimodal pre-training for multimodal sentiment analysis
Junjie Ye 0005, Jie Zhou 0015, Rui Wang 0005, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
Knowl. Based Syst.4
2021 Entity Relation Extraction as Dependency Parsing in Visually Rich Documents
abstract
Previous works on key information extraction from visually rich documents (VRDs) mainly focus on labeling the text within each bounding box (i.e., semantic entity), while the relations in-between are largely unexplored.In this paper, we adapt the popular dependency parsing model, the biaffine parser, to this entity relation extraction task.Being different from the original dependency parsing model which recognizes dependency relations between words, we identify relations between groups of words with layout information instead.We have compared different representations of the semantic entity, different VRD encoders, and different relation decoders.For the model training, we explore multi-task learning to combine entity labeling and relation extraction tasks; and for the evaluation, we conduct experiments on different datasets with filtering and augmentation.The results demonstrate that our proposed model achieves 65.96% F1 score on the FUNSD dataset.As for the realworld application, our model has been applied to the in-house customs data, achieving reliable performance in the production setting.
Yue Zhang 0004, Bo Zhang 0071, Rui Wang 0005, Junjie Cao 0003, Chen Li 0001, Zuyi Bao
EMNLP (1)3
2021 DialogueCSE: Dialogue-based Contrastive Learning of Sentence Embeddings
abstract
Learning sentence embeddings from dialogues has drawn increasing attention due to its low annotation cost and high domain adaptability.Conventional approaches employ the siamese-network for this task, which obtains the sentence embeddings through modeling the context-response semantic relevance by applying a feed-forward network on top of the sentence encoders.However, as the semantic textual similarity is commonly measured through the element-wise distance metrics (e.g.cosine and L2 distance), such architecture yields a large gap between training and evaluating.In this paper, we propose DialogueCSE, a dialogue-based contrastive learning approach to tackle this issue.DialogueCSE first introduces a novel matching-guided embedding (MGE) mechanism, which generates a contextaware embedding for each candidate response embedding (i.e. the context-free embedding) according to the guidance of the multi-turn context-response matching matrices.Then it pairs each context-aware embedding with its corresponding context-free embedding and finally minimizes the contrastive loss across all pairs.We evaluate our model on three multi-turn dialogue datasets: the Microsoft Dialogue Corpus, the Jing Dong Dialogue Corpus, and the E-commerce Dialogue Corpus.Evaluation results show that our approach significantly outperforms the baselines across all three datasets in terms of MAP and Spearman's correlation measures, demonstrating its effectiveness.Further quantitative experiments show that our approach achieves better performance when leveraging more dialogue context and remains robust when less training data is provided.
Rui Wang 0005, Jian Sun 0021, Fei Huang 0002, Luo Si
EMNLP (1)2
2021 Chinese Opinion Role Labeling with Corpus Translation: A Pivot Study
abstract
Opinion Role Labeling (ORL), aiming to identify the key roles of opinion, has received increasing interest.Unlike most of the previous works focusing on the English language, in this paper, we present the first work of Chinese ORL.We construct a Chinese dataset by manually translating and projecting annotations from a standard English MPQA dataset.Then, we investigate the effectiveness of cross-lingual transfer methods, including model transfer and corpus translation.We exploit multilingual BERT with Contextual Parameter Generator and Adapter methods to examine the potentials of unsupervised crosslingual learning and our experiments and analyses for both bilingual and multilingual transfers establish a foundation for the future research of this task 1 .
Ranran Zhen, Rui Wang 0005, Guohong Fu, Chengguo Lv, Meishan Zhang
EMNLP (1)2
2021 Prototypical Representation Learning for Relation Extraction
Ning Ding 0002, Xiaobin Wang, Rui Wang 0005, Pengjun Xie, Ying Shen 0001, Fei Huang 0002, Hai-Tao Zheng 0002, Rui Zhang 0003
ICLR5
2021 A Unified Span-Based Approach for Opinion Mining with Syntactic Constituents
abstract
Qingrong Xia, Bo Zhang, Rui Wang, Zhenghua Li, Yue Zhang, Fei Huang, Luo Si, Min Zhang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Qingrong Xia, Bo Zhang 0071, Rui Wang 0005, Zhenghua Li, Yue Zhang 0004, Fei Huang 0002, Luo Si, Min Zhang 0005
NAACL-HLT3
2020 Boundary Enhanced Neural Span Classification for Nested Named Entity Recognition
abstract
Named entity recognition (NER) is a well-studied task in natural language processing. However, the widely-used sequence labeling framework is usually difficult to detect entities with nested structures. The span-based method that can easily detect nested entities in different subsequences is naturally suitable for the nested NER problem. However, previous span-based methods have two main issues. First, classifying all subsequences is computationally expensive and very inefficient at inference. Second, the span-based methods mainly focus on learning span representations but lack of explicit boundary supervision. To tackle the above two issues, we propose a boundary enhanced neural span classification model. In addition to classifying the span, we propose incorporating an additional boundary detection task to predict those words that are boundaries of entities. The two tasks are jointly trained under a multitask learning framework, which enhances the span representation with additional boundary supervision. In addition, the boundary detection model has the ability to generate high-quality candidate spans, which greatly reduces the time complexity during inference. Experiments show that our approach outperforms all existing methods and achieves 85.3, 83.9, and 78.3 scores in terms of F1 on the ACE2004, ACE2005, and GENIA datasets, respectively.
Chuanqi Tan, Mosha Chen, Rui Wang 0005, Fei Huang 0002
AAAI4
2020 Relational Graph Attention Network for Aspect-based Sentiment Analysis
abstract
Aspect-based sentiment analysis aims to determine the sentiment polarity towards a specific aspect in online reviews.Most recent efforts adopt attention-based neural network models to implicitly connect aspects with opinion words.However, due to the complexity of language and the existence of multiple aspects in a single sentence, these models often confuse the connections.In this paper, we address this problem by means of effective encoding of syntax information.Firstly, we define a unified aspect-oriented dependency tree structure rooted at a target aspect by reshaping and pruning an ordinary dependency parse tree.Then, we propose a relational graph attention network (R-GAT) to encode the new tree structure for sentiment prediction.Extensive experiments are conducted on the SemEval 2014 and Twitter datasets, and the experimental results confirm that the connections between aspects and opinion words can be better established with our approach, and the performance of the graph attention network (GAT) is significantly improved as a consequence.
Weizhou Shen, Yunyi Yang, Xiaojun Quan, Rui Wang 0005
ACL5
2020 Multi-Domain Dialogue Acts and Response Co-Generation
abstract
Generating fluent and informative responses is of critical importance for task-oriented dialogue systems.Existing pipeline approaches generally predict multiple dialogue acts first and use them to assist response generation.There are at least two shortcomings with such approaches.First, the inherent structures of multi-domain dialogue acts are neglected.Second, the semantic associations between acts and responses are not taken into account for response generation.To address these issues, we propose a neural co-generation model that generates dialogue acts and responses concurrently.Unlike those pipeline approaches, our act generation module preserves the semantic structures of multi-domain dialogue acts and our response generation module dynamically attends to different acts as needed.We train the two modules jointly using an uncertainty loss to adjust their task weights adaptively.Extensive experiments are conducted on the largescale MultiWOZ dataset and the results show that our model achieves very favorable improvement over several state-of-the-art models in both automatic and human evaluations.
Rui Wang 0005, Xiaojun Quan, Jianxing Yu
ACL3
2020 Syntax-Aware Opinion Role Labeling with Dependency Graph Convolutional Networks
abstract
Opinion role labeling (ORL) is a fine-grained opinion analysis task and aims to answer "who expressed what kind of sentiment towards what?".Due to the scarcity of labeled data, ORL remains challenging for data-driven methods.In this work, we try to enhance neural ORL models with syntactic knowledge by comparing and integrating different representations.We also propose dependency graph convolutional networks (DEPGCN) to encode parser information at different processing levels.In order to compensate for parser inaccuracy and reduce error propagation, we introduce multi-task learning (MTL) to train the parser and the ORL model simultaneously.We verify our methods on the benchmark MPQA corpus.The experimental results show that syntactic information is highly valuable for ORL, and our final MTL model effectively boosts the F1 score by 9.29 over the syntaxagnostic baseline.In addition, we find that the contributions from syntactic knowledge do not fully overlap with contextualized word representations (BERT).Our best model achieves 4.34 higher F1 score than the current state-ofthe-art.
Bo Zhang 0071, Yue Zhang 0004, Rui Wang 0005, Zhenghua Li, Min Zhang 0005
ACL3
2020 Semantic Role Labeling with Heterogeneous Syntactic Knowledge
abstract
Recently, due to the interplay between syntax and semantics, incorporating syntactic knowledge into neural semantic role labeling (SRL) has achieved much attention.Most of the previous syntax-aware SRL works focus on explicitly modeling homogeneous syntactic knowledge over tree outputs.In this work, we propose to encode heterogeneous syntactic knowledge for SRL from both explicit and implicit representations.First, we introduce graph convolutional networks to explicitly encode multiple heterogeneous dependency parse trees.Second, we extract the implicit syntactic representations from syntactic parser trained with heterogeneous treebanks.Finally, we inject the two types of heterogeneous syntax-aware representations into the base SRL model as extra inputs.We conduct experiments on two widely-used benchmark datasets, i.e., Chinese Proposition Bank 1.0 and English CoNLL-2005 dataset.Experimental results show that incorporating heterogeneous syntactic knowledge brings significant improvements over strong baselines.We further conduct detailed analysis to gain insights on the usefulness of heterogeneous (vs.homogeneous) syntactic knowledge and the effectiveness of our proposed approaches for modeling such knowledge.
Qingrong Xia, Rui Wang 0005, Zhenghua Li, Yue Zhang 0004, Min Zhang 0005
COLING2
2020 SentiX: A Sentiment-Aware Pre-Trained Model for Cross-Domain Sentiment Analysis
abstract
Pre-trained language models have been widely applied to cross-domain NLP tasks like sentiment analysis, achieving state-of-the-art performance.However, due to the variety of users' emotional expressions across domains, fine-tuning the pre-trained models on the source domain tends to overfit, leading to inferior results on the target domain.In this paper, we pre-train a sentimentaware language model (SENTIX) via domain-invariant sentiment knowledge from large-scale review datasets, and utilize it for cross-domain sentiment analysis task without fine-tuning.We propose several pre-training tasks based on existing lexicons and annotations at both token and sentence levels, such as emoticons, sentiment words, and ratings, without human interference.A series of experiments are conducted and the results indicate the great advantages of our model.We obtain new state-of-the-art results in all the cross-domain sentiment analysis tasks, and our proposed SENTIX can be trained with only 1% samples (18 samples) and it achieves better performance than BERT with 90% samples.Code is available at
Jie Zhou 0015, Rui Wang 0005, Yuanbin Wu, Wenming Xiao, Liang He 0001
COLING3
2020 Large Scale Abstractive Multi-Review Summarization (LSARS) via Aspect Alignment
abstract
In an active e-commerce environment, customers process a large number of reviews when deciding on whether to buy a product or not. Abstractive Multi-Review Summarization aims to assist users to efficiently consume the reviews that are the most relevant to them. We propose the first large-scale abstractive multi-review summarization dataset that leverages more than 17.9 billion raw reviews and uses novel aspect-alignment techniques based on aspect annotations. Furthermore, we demonstrate that one can generate higher-quality review summaries by using a novel aspect-alignment-based model. Results from both automatic and human evaluation show that the proposed dataset plus the innovative aspect-alignment model can generate high-quality and trustful review summaries.
Haojie Pan, Rongqin Yang, Rui Wang 0005, Deng Cai 0001, Xiaozhong Liu 0001
SIGIR4
2019 Syntax-Aware Neural Semantic Role Labeling
abstract
Semantic role labeling (SRL), also known as shallow semantic parsing, is an important yet challenging task in NLP. Motivated by the close correlation between syntactic and semantic structures, traditional discrete-feature-based SRL approaches make heavy use of syntactic features. In contrast, deep-neural-network-based approaches usually encode the input sentence as a word sequence without considering the syntactic structures. In this work, we investigate several previous approaches for encoding syntactic trees, and make a thorough study on whether extra syntax-aware representations are beneficial for neural SRL models. Experiments on the benchmark CoNLL-2005 dataset show that syntax-aware SRL approaches can effectively improve performance over a strong baseline with external word representations from ELMo. With the extra syntax-aware representations, our approaches achieve new state-of-the-art 85.6 F1 (single model) and 86.6 F1 (ensemble) on the test data, outperforming the corresponding strong baselines with ELMo by 0.8 and 1.0, respectively. Detailed error analysis are conducted to gain more insights on the investigated approaches.
Qingrong Xia, Zhenghua Li, Min Zhang 0005, Meishan Zhang, Guohong Fu, Rui Wang 0005, Luo Si
AAAI6
2019 A Deep Cascade Model for Multi-Document Reading Comprehension
abstract
A fundamental trade-off between effectiveness and efficiency needs to be balanced when designing an online question answering system. Effectiveness comes from sophisticated functions such as extractive machine reading comprehension (MRC), while efficiency is obtained from improvements in preliminary retrieval components such as candidate document selection and paragraph ranking. Given the complexity of the real-world multi-document MRC scenario, it is difficult to jointly optimize both in an end-to-end system. To address this problem, we develop a novel deep cascade learning model, which progressively evolves from the documentlevel and paragraph-level ranking of candidate texts to more precise answer extraction with machine reading comprehension. Specifically, irrelevant documents and paragraphs are first filtered out with simple functions for efficiency consideration. Then we jointly train three modules on the remaining texts for better tracking the answer: the document extraction, the paragraph extraction and the answer extraction. Experiment results show that the proposed method outperforms the previous state-of-the-art methods on two large-scale multidocument benchmark datasets, i.e., TriviaQA and DuReader. In addition, our online system can stably serve typical scenarios with millions of daily requests in less than 50ms.
Ming Yan 0008, Jiangnan Xia, Chen Wu 0006, Bin Bi, Zhongzhou Zhao, Ji Zhang 0011, Luo Si, Rui Wang 0005, Wei Wang 0225, Haiqing Chen
AAAI8
2019 Semi-supervised Domain Adaptation for Dependency Parsing
abstract
During the past decades, due to the lack of sufficient labeled data, most studies on crossdomain parsing focus on unsupervised domain adaptation, assuming there is no targetdomain training data.However, unsupervised approaches make limited progress so far due to the intrinsic difficulty of both domain adaptation and parsing.This paper tackles the semi-supervised domain adaptation problem for Chinese dependency parsing, based on two newly-annotated large-scale domain-specific datasets.1 We propose a simple domain embedding approach to merge the sourceand target-domain training data, which is shown to be more effective than both direct corpus concatenation and multi-task learning.In order to utilize unlabeled target-domain data, we employ the recent contextualized word representations and show that a simple fine-tuning procedure can further boost cross-domain parsing accuracy by large margins.
Zhenghua Li, Xue Peng, Min Zhang 0005, Rui Wang 0005, Luo Si
ACL (1)4
2019 BiSET: Bi-directional Selective Encoding with Template for Abstractive Summarization
abstract
The success of neural summarization models stems from the meticulous encodings of source articles.To overcome the impediments of limited and sometimes noisy training data, one promising direction is to make better use of the available training data by applying filters during summarization.In this paper, we propose a novel Bi-directional Selective Encoding with Template (BiSET) model, which leverages template discovered from training data to softly select key information from each source article to guide its summarization process.Extensive experiments on a standard summarization dataset were conducted and the results show that the template-equipped BiSET model manages to improve the summarization performance significantly with a new state of the art.
Xiaojun Quan, Rui Wang 0005
ACL (1)3
2019 Attention Optimization for Abstractive Document Summarization
abstract
Min Gui, Junfeng Tian, Rui Wang, Zhenglu Yang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Min Gui, Rui Wang 0005, Zhenglu Yang
EMNLP/IJCNLP (1)3
2019 Syntax-Enhanced Self-Attention-Based Semantic Role Labeling
abstract
Yue Zhang, Rui Wang, Luo Si. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yue Zhang 0004, Rui Wang 0005, Luo Si
EMNLP/IJCNLP (1)2
2019 Self-attentive Biaffine Dependency Parsing
abstract
The current state-of-the-art dependency parsing approaches employ BiLSTMs to encode input sentences.Motivated by the success of the transformer-based machine translation, this work for the first time applies the self-attention mechanism to dependency parsing as the replacement of the BiLSTM-based encoders, leading to competitive performance on both English and Chinese benchmark data. Based on the detailed error analysis, we then combine the power of both BiLSTM and self-attention via model ensembles, demonstrating their complementary capability of capturing contextual information. Finally, we explore the recently proposed contextualized word representations as extra input features, and further improve the parsing performance.
Ying Li 0065, Zhenghua Li, Min Zhang 0005, Rui Wang 0005, Sheng Li 0017, Luo Si
IJCAI4
2019 Overview of the NLPCC 2019 Shared Task: Cross-Domain Dependency Parsing
Xue Peng, Zhenghua Li, Min Zhang 0005, Rui Wang 0005, Yue Zhang 0004, Luo Si
NLPCC (2)4
2015 Design and realization of a modular architecture for textual entailment
abstract
Abstract A key challenge at the core of many Natural Language Processing (NLP) tasks is the ability to determine which conclusions can be inferred from a given natural language text. This problem, called theRecognition of Textual Entailment (RTE), has initiated the development of a range of algorithms, methods, and technologies. Unfortunately, research on Textual Entailment (TE), like semantics research more generally, is fragmented into studies focussing on various aspects of semantics such as world knowledge, lexical and syntactic relations, or more specialized kinds of inference. This fragmentation has problematic practical consequences. Notably, interoperability among the existing RTE systems is poor, and reuse of resources and algorithms is mostly infeasible. This also makes systematic evaluations very difficult to carry out. Finally, textual entailment presents a wide array of approaches to potential end users with little guidance on which to pick. Our contribution to this situation is the novel EXCITEMENT architecture, which was developed to enable and encourage the consolidation of methods and resources in the textual entailment area. It decomposes RTE into components with strongly typed interfaces. We specify (a) a modular linguistic analysis pipeline and (b) a decomposition of the ‘core’ RTE methods into top-level algorithms and subcomponents. We identify four major subcomponent types, including knowledge bases and alignment methods. The architecture was developed with a focus on generality, supporting all major approaches to RTE and encouraging language independence. We illustrate the feasibility of the architecture by constructing mappings of major existing systems onto the architecture. The practical implementation of this architecture forms the EXCITEMENT open platform. It is a suite of textual entailment algorithms and components which contains the three systems named above, including linguistic-analysis pipelines for three languages (English, German, and Italian), and comprises a number of linguistic resources. By addressing the problems outlined above, the platform provides a comprehensive and flexible basis for research and experimentation in textual entailment and is available as open source software under the GNU General Public License.
Sebastian Padó, Tae-Gil Noh, Asher Stern 0001, Rui Wang 0005, Roberto Zanoli
Nat. Lang. Eng.4
2014 Aligning Predicate-Argument Structures for Paraphrase Fragment Extraction
Michaela Regneri, Rui Wang 0005, Manfred Pinkal
LREC2
2014 Senti-LSSVM: Sentiment-Oriented Multi-Relation Extraction with Latent Structural SVM
abstract
Extracting instances of sentiment-oriented relations from user-generated web documents is important for online marketing analysis. Unlike previous work, we formulate this extraction task as a structured prediction problem and design the corresponding inference as an integer linear program. Our latent structural SVM based model can learn from training corpora that do not contain explicit annotations of sentiment-bearing expressions, and it can simultaneously recognize instances of both binary (polarity) and ternary (comparative) relations with regard to entity mentions of interest. The empirical evaluation shows that our approach significantly outperforms state-of-the-art systems across domains (cameras and movies) and across genres (reviews and forum posts). The gold standard corpus that we built will also be a valuable resource for the community.
Lizhen Qu, Yi Zhang 0003, Rui Wang 0005, Lili Jiang 0002, Rainer Gemulla, Gerhard Weikum
Trans. Assoc. Comput. Linguistics3
2012 Using Discourse Information for Paraphrase Extraction
Michaela Regneri, Rui Wang 0005
EMNLP-CoNLL2
2012 Constructing a Question Corpus for Textual Semantic Relations
Rui Wang 0005
LREC1
2012 Joint Grammar and Treebank Development for Mandarin Chinese with HPSG
Yi Zhang 0003, Rui Wang 0005, Yu Chen 0012
LREC2
2010 Constructing a Textual Semantic Relation Corpus Using a Discourse Treebank
Rui Wang 0005, Caroline Sporleder
LREC1
2010 Hybrid Constituent and Dependency Parsing with Tsinghua Chinese Treebank
Rui Wang 0005, Yi Zhang 0003
LREC1
2009 Cross-Domain Dependency Parsing Using a Deep Linguistic Grammar
Yi Zhang 0003, Rui Wang 0005
ACL/IJCNLP2
2009 Inference Rules and their Application to Recognizing Textual Entailment
Georgiana Dinu, Rui Wang 0005
EACL2
2009 Recognizing Textual Relatedness with Predicate-Argument Structures
Rui Wang 0005, Yi Zhang 0003
EMNLP1
2008 Hybrid Learning of Dependency Structures from Heterogeneous Linguistic Resources
Yi Zhang 0003, Rui Wang 0005, Hans Uszkoreit
CoNLL2
2007 Recognizing Textual Entailment Using a Subsequence Kernel Method
Rui Wang 0005, Günter Neumann
AAAI1