Yuanbin Wu

dblp:17/7186 · DBLP profile ↗
← Back
60ranked-venue papers
5as first author
33since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 56 · 5 first-author · 31 since 2021Databases, data management, data science and information retrieval · 8 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Guiding LLMs to decode text via aligning semantics in EEG signals and language
Huanran Zheng, Yuanbin Wu, Tianwen Qian, Wenjing Yue, Xiaoling Wang 0004
Expert Syst. Appl.2
2025 Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
abstract
Tao Ji, Bin Guo, Yuanbin Wu, Qipeng Guo, Shenlixing Shenlixing, Chenzhan Chenzhan, Xipeng Qiu, Qi Zhang, Tao Gui. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yuanbin Wu, Qipeng Guo, Shenlixing Shenlixing, Chenzhan Chenzhan, Xipeng Qiu, Qi Zhang 0001, Tao Gui
ACL (1)3
2025 On Support Samples of Next Word Prediction
abstract
Language models excel in various tasks by making complex decisions, yet understanding the rationale behind these decisions remains a challenge.This paper investigates data-centric interpretability in language models, focusing on the next-word prediction task.Using representer theorem, we identify two types of support samples-those that either promote or deter specific predictions.Our findings reveal that being a support sample is an intrinsic property, predictable even before training begins.Additionally, while non-support samples are less influential in direct predictions, they play a critical role in preventing overfitting and shaping generalization and representation learning.Notably, the importance of non-support samples increases in deeper layers, suggesting their significant role in intermediate representation formation.These insights shed light on the interplay between data and model decisions, offering a new dimension to understanding language model behavior and interpretability. 1 * Equal contribution. 1 Our source code is publicly available at https://github.com/liyuqian44/ On-Support-Samples-of-Next-Word-Prediction.
Yupei Du, Yufang Liu, Feifei Feng, Mou Xiao Feng, Yuanbin Wu
ACL (1)6
2025 The Role of Visual Modality in Multimodal Mathematical Reasoning: Challenges and Insights
abstract
Yufang Liu, Yao Du, Tao Ji, Jianing Wang, Yang Liu, Yuanbin Wu, Aimin Zhou, Mengdi Zhang, Xunliang Cai. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yufang Liu, Yuanbin Wu, Aimin Zhou, Mengdi Zhang 0002
ACL (1)6
2025 Logic-Regularized Verifier Elicits Reasoning from LLMs
abstract
Xinyu Wang, Changzhi Sun, Lian Cheng, Yuanbin Wu, Dell Zhang, Xiaoling Wang, Xuelong Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xinyu Wang 0022, Changzhi Sun, Lian Cheng, Yuanbin Wu, Dell Zhang, Xiaoling Wang 0004, Xuelong Li 0001
ACL (1)4
2025 TASO: Task-Aligned Sparse Optimization for Parameter-Efficient Model Adaptation
abstract
Daiye Miao, Yufang Liu, Jie Wang, Changzhi Sun, Yunke Zhang, Demei Yan, Shaokang Dong, Qi Zhang, Yuanbin Wu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Daiye Miao, Yufang Liu, Jie Wang 0126, Changzhi Sun, Yunke Zhang, Demei Yan, Shaokang Dong, Qi Zhang 0001, Yuanbin Wu
EMNLP9
2025 Protein Design with Dynamic Protein Vocabulary
abstract
Protein design is a fundamental challenge in biotechnology, aiming to design novel sequences with specific functions within the vast space of possible proteins. Recent advances in deep generative models have enabled function-based protein design from textual descriptions, yet struggle with structural plausibility. Inspired by classical protein design methods that leverage natural protein structures, we explore whether incorporating fragments from natural proteins can enhance foldability in generative models. Our empirical results show that even random incorporation of fragments improves foldability. Building on this insight, we introduce ProDVa, a novel protein design approach that integrates a text encoder for functional descriptions, a protein language model for designing proteins, and a fragment encoder to dynamically retrieve protein fragments based on textual functional descriptions. Experimental results demonstrate that our approach effectively designs protein sequences that are both functionally aligned and structurally plausible. Compared to state-of-the-art models, ProDVa achieves comparable function alignment using less than 0.04% of the training data, while designing significantly more well-folded proteins, with the proportion of proteins having pLDDT above 70 increasing by 7.38% and those with PAE below 10 increasing by 9.62%.
Nuowei Liu, Jiahao Kuang, Changzhi Sun, Man Lan, Yuanbin Wu
NeurIPS7
2025 Overview of the NLPCC 2025 Shared Task 3: Comprehensive Argument Analysis for Chinese Argumentative Essay
Zheqin Yin, Yupei Ren, Man Lan, Yuanbin Wu, Aimin Zhou, Xiaopeng Bai
NLPCC (4)4
2024 Generation with Dynamic Vocabulary
abstract
We introduce a new dynamic vocabulary for language models.It can involve arbitrary text spans during generation.These text spans act as basic generation bricks, akin to tokens in the traditional static vocabularies.We show that, the ability to generate multi-tokens atomically improve both generation quality and efficiency (compared to the standard language model, the MAUVE metric is increased by 25%, the latency is decreased by 20%).The dynamic vocabulary can be deployed in a plug-and-play way, thus is attractive for various downstream applications.For example, we demonstrate that dynamic vocabulary can be applied to different domains in a training-free manner.It also helps to generate reliable citations in question answering tasks (substantially enhancing citation results without compromising answer accuracy).
Changzhi Sun, Yuanbin Wu
EMNLP4
2024 Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models
abstract
Large Vision-Language Models (LVLMs) have achieved impressive performance, yet research has pointed out a serious issue with object hallucinations within these models.However, there is no clear conclusion as to which part of the model these hallucinations originate from.In this paper, we present an in-depth investigation into the object hallucination problem specifically within the CLIP model, which serves as the backbone for many state-of-the-art visionlanguage systems.We unveil that even in isolation, the CLIP model is prone to object hallucinations, suggesting that the hallucination problem is not solely due to the interaction between vision and language modalities.To address this, we propose a counterfactual data augmentation method by creating negative samples with a variety of hallucination issues.We demonstrate that our method can effectively mitigate object hallucinations for the CLIP model, and we show that the enhanced model can be employed as a visual encoder, effectively alleviating the object hallucination issue in LVLMs. 1
Yufang Liu, Changzhi Sun, Yuanbin Wu, Aimin Zhou
EMNLP4
2024 Overview of the NLPCC 2024 Shared Task 5: Argument Mining for Chinese Argumentative Essay
Zheqin Yin, Yupei Ren, Man Lan, Yuanbin Wu, Aimin Zhou, Xiaopeng Bai
NLPCC (5)4
2024 Overview of the NLPCC 2024 Shared Task: Chinese Essay Discourse Logic Evaluation and Integration
Hongyi Wu, Xinshu Shen, Man Lan, Yuanbin Wu, Xiaopeng Bai, Shaoguang Mao, Tao Ge 0001, Yan Xia 0005
NLPCC (5)5
2023 CodeIE: Large Code Generation Models are Better Few-Shot Information Extractors
abstract
Peng Li, Tianxiang Sun, Qiong Tang, Hang Yan, Yuanbin Wu, Xuanjing Huang, Xipeng Qiu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Tianxiang Sun, Qiong Tang, Hang Yan 0001, Yuanbin Wu, Xuanjing Huang 0001, Xipeng Qiu
ACL (1)5
2023 Rehearsal-free Continual Language Learning via Efficient Parameter Isolation
abstract
Zhicheng Wang, Yufang Liu, Tao Ji, Xiaoling Wang, Yuanbin Wu, Congcong Jiang, Ye Chao, Zhencong Han, Ling Wang, Xu Shao, Wenqiu Zeng. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yufang Liu, Yuanbin Wu, Congcong Jiang, Ye Chao, Zhencong Han, Xu Shao, Wenqiu Zeng
ACL (1)5
2023 Connective Prediction for Implicit Discourse Relation Recognition via Knowledge Distillation
abstract
Implicit discourse relation recognition (IDRR) remains a challenging task in discourse analysis due to the absence of connectives.Most existing methods utilize one-hot labels as the sole optimization target, ignoring the internal association among connectives.Besides, these approaches spend lots of effort on template construction, negatively affecting the generalization capability.To address these problems, we propose a novel Connective Prediction via Knowledge Distillation (CP-KD) approach to instruct large-scale pre-trained language models (PLMs) mining the latent correlations between connectives and discourse relations, which is meaningful for IDRR.Experimental results on the PDTB 2.0/3.0 and CoNLL 2016 datasets show that our method significantly outperforms the state-of-the-art models on coarse-grained and fine-grained discourse relations.Moreover, our approach can be transferred to explicit discourse relation recognition (EDRR) and achieve acceptable performance.Our code is released in https://github.com/cubenlp/CP_KD-for-IDRR.
Hongyi Wu, Man Lan, Yuanbin Wu
ACL (1)4
2023 A Multi-Task Dataset for Assessing Discourse Coherence in Chinese Essays: Structure, Theme, and Logic Analysis
abstract
This paper introduces the Chinese Essay Discourse Coherence Corpus (CEDCC), a multi-task dataset for assessing discourse coherence.Existing research tends to focus on isolated dimensions of discourse coherence, a gap which the CEDCC addresses by integrating coherence grading, topical continuity, and discourse relations.This approach, alongside detailed annotations, captures the subtleties of real-world texts and stimulates progress in Chinese discourse coherence analysis.Our contributions include the development of the CEDCC, the establishment of baselines for further research, and the demonstration of the impact of coherence on discourse relation recognition and automated essay scoring.The dataset and related codes is available at https: //github.com/cubenlp/CEDCC_corpus.
Hongyi Wu, Xinshu Shen, Man Lan, Shaoguang Mao, Xiaopeng Bai, Yuanbin Wu
EMNLP6
2023 Overview of the NLPCC 2023 Shared Task: Chinese Essay Discourse Coherence Evaluation
Hongyi Wu, Xinshu Shen, Man Lan, Xiaopeng Bai, Yuanbin Wu, Aimin Zhou, Shaoguang Mao, Tao Ge 0001, Yan Xia 0005
NLPCC (3)5
2023 Multi-modal multi-hop interaction network for dialogue response generation
Jie Zhou 0015, Rui Wang 0005, Yuanbin Wu, Ming Yan 0008, Liang He 0001, Xuanjing Huang 0001
Expert Syst. Appl.4
2022 Understanding Gender Bias in Knowledge Base Embeddings
abstract
Knowledge base (KB) embeddings have been shown to contain gender biases (Fisher et al., 2020b).In this paper, we study two questions regarding these biases: how to quantify them, and how to trace their origins in KB? Specifically, first, we develop two novel bias measures respectively for a group of person entities and an individual person entity.Evidence of their validity is observed by comparison with real-world census data.Second, we use influence function to inspect the contribution of each triple in KB to the overall group bias.To exemplify the potential applications of our study, we also present two strategies (by adding and removing KB triples) to mitigate gender biases in KB embeddings.
Yupei Du, Yuanbin Wu, Man Lan, Yan Yang 0008, Meirong Ma
ACL (1)3
2022 Few Clean Instances Help Denoising Distant Supervision
abstract
Existing distantly supervised relation extractors usually rely on noisy data for both model training and evaluation, which may lead to garbage-in-garbage-out systems. To alleviate the problem, we study whether a small clean dataset could help improve the quality of distantly supervised models. We show that besides getting a more convincing evaluation of models, a small clean dataset also helps us to build more robust denoising models. Specifically, we propose a new criterion for clean instance selection based on influence functions. It collects sample-level evidence for recognizing good instances (which is more informative than loss-level evidence). We also propose a teacher-student mechanism for controlling purity of intermediate results when bootstrapping the clean set. The whole approach is model-agnostic and demonstrates strong performances on both denoising real (NYT) and synthetic noisy datasets.
Yufang Liu, Ziyin Huang, Changzhi Sun, Man Lan, Yuanbin Wu, Xiaofeng Mou
COLING6
2022 LightEA: A Scalable, Robust, and Interpretable Entity Alignment Framework via Three-view Label Propagation
abstract
Entity Alignment (EA) aims to find equivalent entity pairs between KGs, which is the core step of bridging and integrating multi-source KGs.In this paper, we argue that existing GNNbased EA methods inherit the inborn defects from their neural network lineage: weak scalability and poor interpretability.Inspired by recent studies, we reinvent the Label Propagation algorithm to effectively run on KGs and propose a non-neural EA framework -LightEA, consisting of three efficient components: (i) Random Orthogonal Label Generation, (ii) Three-view Label Propagation, and (iii) Sparse Sinkhorn Iteration.According to the extensive experiments on public datasets, LightEA has impressive scalability, robustness, and interpretability.With a mere tenth of time consumption, LightEA achieves comparable results to state-of-the-art methods across all datasets and even surpasses them on many.
Xin Mao 0002, Yuanbin Wu, Man Lan
EMNLP3
2021 Generating CCG Categories
Yufang Liu, Yuanbin Wu, Man Lan
AAAI3
2021 UniRE: A Unified Label Space for Entity Relation Extraction
abstract
Yijun Wang, Changzhi Sun, Yuanbin Wu, Hao Zhou, Lei Li, Junchi Yan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Changzhi Sun, Yuanbin Wu, Hao Zhou 0012, Lei Li 0005, Junchi Yan
ACL/IJCNLP (1)3
2021 Are Negative Samples Necessary in Entity Alignment?: An Approach with High Performance, Scalability and Robustness
abstract
Entity alignment (EA) aims to find the equivalent entities in different KGs, which is a crucial step in integrating multiple KGs. However, most existing EA methods have poor scalability and are unable to cope with large-scale datasets. We summarize three issues leading to such high time-space complexity in existing EA methods: (1) Inefficient graph encoders, (2) Dilemma of negative sampling, and (3) "Catastrophic forgetting" in semi-supervised learning. To address these challenges, we propose a novel EA method with three new components to enable high Performance, high Scalability, and high Robustness (PSR): (1) Simplified graph encoder with relational graph sampling, (2) Symmetric negative-free alignment loss, and (3) Incremental semi-supervised learning. Furthermore, we conduct detailed experiments on several public datasets to examine the effectiveness and efficiency of our proposed method. The experimental results show that PSR not only surpasses the previous SOTA in performance but also has impressive scalability and robustness.
Xin Mao 0002, Yuanbin Wu, Man Lan
CIKM3
2021 ENPAR: Enhancing Entity and Entity Pair Representations for Joint Entity Relation Extraction
abstract
Current state-of-the-art systems for joint entity relation extraction (Luan et al., 2019;Wadden et al., 2019) usually adopt the multi-task learning framework.However, annotations for these additional tasks such as coreference resolution and event extraction are always equally hard (or even harder) to obtain.In this work, we propose a pre-training method ENPAR to improve the joint extraction performance.EN-PAR requires only the additional entity annotations that are much easier to collect.Unlike most existing works that only consider incorporating entity information into the sentence encoder, we further utilize the entity pair information.Specifically, we devise four novel objectives, i.e., masked entity typing, masked entity prediction, adversarial context discrimination, and permutation prediction, to pretrain an entity encoder and an entity pair encoder.Comprehensive experiments show that the proposed pre-training method achieves significant improvement over BERT on ACE05, SciERC, and NYT, and outperforms current state-of-the-art on ACE05.
Changzhi Sun, Yuanbin Wu, Hao Zhou 0012, Lei Li 0005, Junchi Yan
EACL3
2021 Is "hot pizza" Positive or Negative? Mining Target-aware Sentiment Lexicons
abstract
Modelling a word's polarity in different contexts is a key task in sentiment analysis.Previous works mainly focus on domain dependencies, and assume words' sentiments are invariant within a specific domain.In this paper, we relax this assumption by binding a word's sentiment to its collocation words instead of domain labels.This finer view of sentiment contexts is particularly useful for identifying commonsense sentiments expressed in neutral words such as "big" and "long".Given a target (e.g., an aspect), we propose an effective "perturb-and-see" method to extract sentiment words modifying it from large-scale datasets.The reliability of the obtained targetaware sentiment lexicons is extensively evaluated both manually and automatically.We also show that a simple application of the lexicon is able to achieve highly competitive performances on the unsupervised opinion relation extraction task.
Jie Zhou 0015, Yuanbin Wu, Changzhi Sun, Liang He 0001
EACL2
2021 Word Reordering for Zero-shot Cross-lingual Structured Prediction
abstract
Adapting word order from one language to another is a key problem in cross-lingual structured prediction.Current sentence encoders (e.g., RNN, Transformer with position embeddings) are usually word order sensitive.Even with uniform word form representations (MUSE, mBERT), word order discrepancies may hurt the adaptation of models.This paper builds structured prediction models with bag-of-words inputs.It introduces a new reordering module to organize words following the source language order, which learns taskspecific reordering strategies from a generalpurpose order predictor model.Experiments on zero-shot cross-lingual dependency parsing, POS tagging, and morphological tagging show that our model can significantly improve target language performances, especially for languages that are distant from the source language.1
Yong Jiang 0005, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Yuanbin Wu
EMNLP (1)6
2021 A Unified Encoding of Structures in Transition Systems
abstract
Transition systems usually contain various dynamic structures (e.g., stacks, buffers).An ideal transition-based model should encode these structures completely and efficiently.Previous works relying on templates or neural network structures either only encode partial structure information or suffer from computation efficiency.In this paper, we propose a novel attention-based encoder unifying representation of all structures in a transition system.Specifically, we separate two views of items on structures, namely structure-invariant view and structure-dependent view.With the help of parallel-friendly attention network, we are able to encoding transition states with O(1) additional complexity (with respect to basic feature extractors).Experiments on the PTB and UD show that our proposed method significantly improves the test speed and achieves the best transition-based model, and is comparable to state-of-the-art methods. 1
Yong Jiang 0005, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Yuanbin Wu
EMNLP (1)6
2021 From Alignment to Assignment: Frustratingly Simple Unsupervised Entity Alignment
abstract
Cross-lingual entity alignment (EA) aims to find the equivalent entities between crosslingual KGs (Knowledge Graphs), which is a crucial step for integrating KGs.Recently, many GNN-based EA methods are proposed and show decent performance improvements on several public datasets.However, existing GNN-based EA methods inevitably inherit poor interpretability and low efficiency from neural networks.Motivated by the isomorphic assumption of GNN-based methods, we successfully transform the cross-lingual EA problem into an assignment problem.Based on this re-definition, we propose a frustratingly Simple but Effective Unsupervised entity alignment method (SEU) without neural networks.Extensive experiments have been conducted to show that our proposed unsupervised approach even beats advanced supervised methods across all public datasets while having high efficiency, interpretability, and stability.
Xin Mao 0002, Yuanbin Wu, Man Lan
EMNLP (1)3
2021 A Dual-Attention Neural Network for Pun Location and Using Pun-Gloss Pairs for Interpretation
Meirong Ma, Jianguo Zhu 0001, Yuanbin Wu, Man Lan
NLPCC (1)5
2021 Boosting the Speed of Entity Alignment 10 ×: Dual Attention Matching Network with Normalized Hard Sample Mining
abstract
Seeking the equivalent entities among multi-source Knowledge Graphs (KGs) is the pivotal step to KGs integration, also known as entity alignment (EA). However, most existing EA methods are inefficient and poor in scalability. A recent summary points out that some of them even require several days to deal with a dataset containing 200,000 nodes (DWY100K). We believe over-complex graph encoder and inefficient negative sampling strategy are the two main reasons. In this paper, we propose a novel KG encoder — Dual Attention Matching Network (Dual-AMN), which not only models both intra-graph and cross-graph information smartly, but also greatly reduces computational complexity. Furthermore, we propose the Normalized Hard Sample Mining Loss to smoothly select hard negative samples with reduced loss shift. The experimental results on widely used public datasets indicate that our method achieves both high accuracy and high efficiency. On DWY100K, the whole running process of our method could be finished in 1,100 seconds, at least 10 × faster than previous work. The performances of our method also outperform previous works across all datasets, where [email protected] and MRR have been improved from 6% to 13%.
Xin Mao 0002, Yuanbin Wu, Man Lan
WWW3
2021 Aligned variational autoencoder for matching danmaku and video storylines
Qingchun Bai, Yuanbin Wu, Jie Zhou 0015, Liang He 0001
Neurocomputing2
2021 Entity-level sentiment prediction in Danmaku video interaction
Qingchun Bai, Jie Zhou 0015, Yuanbin Wu, Xin Lin 0001, Liang He 0001
J. Supercomput.5
2020 A Span-based Linearization for Constituent Trees
abstract
We propose a novel linearization of a constituent tree, together with a new locally normalized model.For each split point in a sentence, our model computes the normalizer on all spans ending with that split point, and then predicts a tree span from them.Compared with global models, our model is fast and parallelizable.Different from previous local models, our linearization method is tied on the spans directly and considers more local features when performing span prediction, which is more interpretable and effective.Experiments on PTB (95.8 F1) and CTB (92.1 F1) show that our model significantly outperforms existing local models and efficiently achieves competitive results with global models.
Yuanbin Wu, Man Lan
ACL2
2020 Relational Reflection Entity Alignment
abstract
Entity alignment aims to identify equivalent entity pairs from different Knowledge Graphs (KGs), which is essential in integrating multi-source KGs. Recently, with the introduction of GNNs into entity alignment, the architectures of recent models have become more and more complicated. We even find two counter-intuitive phenomena within these methods: (1) The standard linear transformation in GNNs is not working well. (2) Many advanced KG embedding models designed for link prediction task perform poorly in entity alignment. In this paper, we abstract existing entity alignment methods into a unified framework, Shape-Builder & Alignment, which not only successfully explains the above phenomena but also derives two key criteria for an ideal transformation operation. Furthermore, we propose a novel GNNs-based method, Relational Reflection Entity Alignment (RREA). RREA leverages Relational Reflection Transformation to obtain relation specific embeddings for each entity in a more efficient way. The experimental results on real-world datasets show that our model significantly outperforms the state-of-the-art methods, exceeding by 5.8%-10.9% on [email protected]
Xin Mao 0002, Yuanbin Wu, Man Lan
CIKM4
2020 SentiX: A Sentiment-Aware Pre-Trained Model for Cross-Domain Sentiment Analysis
abstract
Pre-trained language models have been widely applied to cross-domain NLP tasks like sentiment analysis, achieving state-of-the-art performance.However, due to the variety of users' emotional expressions across domains, fine-tuning the pre-trained models on the source domain tends to overfit, leading to inferior results on the target domain.In this paper, we pre-train a sentimentaware language model (SENTIX) via domain-invariant sentiment knowledge from large-scale review datasets, and utilize it for cross-domain sentiment analysis task without fine-tuning.We propose several pre-training tasks based on existing lexicons and annotations at both token and sentence levels, such as emoticons, sentiment words, and ratings, without human interference.A series of experiments are conducted and the results indicate the great advantages of our model.We obtain new state-of-the-art results in all the cross-domain sentiment analysis tasks, and our proposed SENTIX can be trained with only 1% samples (18 samples) and it achieves better performance than BERT with 90% samples.Code is available at
Jie Zhou 0015, Rui Wang 0005, Yuanbin Wu, Wenming Xiao, Liang He 0001
COLING4
2020 Pre-training Entity Relation Encoder with Intra-span and Inter-span Information
abstract
In this paper, we integrate span-related information into pre-trained encoder for entity relation extraction task.Instead of using generalpurpose sentence encoder (e.g., existing universal pre-trained models), we introduce a span encoder and a span pair encoder to the pre-training network, which makes it easier to import intra-span and inter-span information into the pre-trained model.To learn the encoders, we devise three customized pretraining objectives from different perspectives, which target on tokens, spans, and span pairs.In particular, a span encoder is trained to recover a random shuffling of tokens in a span, and a span pair encoder is trained to predict positive pairs that are from the same sentences and negative pairs that are from different sentences using contrastive loss.Experimental results show that the proposed pre-training method outperforms distantly supervised pretraining, and achieves promising performance on two entity relation extraction benchmark datasets (ACE05, SciERC).
Changzhi Sun, Yuanbin Wu, Junchi Yan, Peng Gao 0015, Guo Tong Xie
EMNLP (1)3
2020 MRAEA: An Efficient and Robust Entity Alignment Approach for Cross-lingual Knowledge Graph
abstract
Entity alignment to find equivalent entities in cross-lingual Knowledge Graphs (KGs) plays a vital role in automatically integrating multiple KGs. Existing translation-based entity alignment methods jointly model the cross-lingual knowledge and monolingual knowledge into one unified optimization problem. On the other hand, the Graph Neural Network (GNN) based methods either ignore the node differentiations, or represent relation through entity or triple instances. They all fail to model the meta semantics embedded in relation nor complex relations such as n-to-n and multi-graphs. To tackle these challenges, we propose a novel Meta Relation Aware Entity Alignment (MRAEA) to directly model cross-lingual entity embeddings by attending over the node's incoming and outgoing neighbors and its connected relations' meta semantics. In addition, we also propose a simple and effective bi-directional iterative strategy to add new aligned seeds during training. Our experiments on all three benchmark entity alignment datasets show that our approach consistently outperforms the state-of-the-art methods, exceeding by 15%-58% on [email protected] Through an extensive ablation study, we validate that the proposed meta relation aware representations, relation aware self-attention and bi-directional iterative strategy of new seed selection all make contributions to significant performance improvement. The code is available at https://github.com/MaoXinn/MRAEA.
Xin Mao 0002, Man Lan, Yuanbin Wu
WSDM5
2019 Distantly Supervised Entity Relation Extraction with Adapted Manual Annotations
abstract
We investigate the task of distantly supervised joint entity relation extraction. It’s known that training with distant supervision will suffer from noisy samples. To tackle the problem, we propose to adapt a small manually labelled dataset to the large automatically generated dataset. By developing a novel adaptation algorithm, we are able to transfer the high quality but heterogeneous entity relation annotations in a robust and consistent way. Experiments on the benchmark NYT dataset show that our approach significantly outperforms state-ofthe-art methods.
Changzhi Sun, Yuanbin Wu
AAAI2
2019 Graph-based Dependency Parsing with Graph Neural Networks
abstract
We investigate the problem of efficiently incorporating high-order features into neural graph-based dependency parsing.Instead of explicitly extracting high-order features from intermediate parse trees, we develop a more powerful dependency tree node representation which captures high-order information concisely and efficiently.We use graph neural networks (GNNs) to learn the representations and discuss several new configurations of GNN's updating and aggregation functions.Experiments on PTB show that our parser achieves the best UAS and LAS on PTB (96.0%, 94.3%) among systems without using any external resources.
Yuanbin Wu, Man Lan
ACL (1)2
2019 Joint Type Inference on Entities and Relations via Graph Convolutional Networks
abstract
We develop a new paradigm for the task of joint entity relation extraction.It first identifies entity spans, then performs a joint inference on entity types and relation types.To tackle the joint type inference task, we propose a novel graph convolutional network (GCN) running on an entity-relation bipartite graph.By introducing a binary relation classification task, we are able to utilize the structure of entity-relation bipartite graph in a more efficient and interpretable way.Experiments on ACE05 show that our model outperforms existing joint models in entity performance and is competitive with the state-of-the-art in relation performance.
Changzhi Sun, Yeyun Gong, Yuanbin Wu, Ming Gong 0001, Daxin Jiang, Man Lan, Shiliang Sun, Nan Duan 0001
ACL (1)3
2019 Exploring Human Gender Stereotypes with Word Association Test
abstract
Yupei Du, Yuanbin Wu, Man Lan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yupei Du, Yuanbin Wu, Man Lan
EMNLP/IJCNLP (1)2
2018 Inference on Syntactic and Semantic Structures for Machine Comprehension
abstract
Hidden variable models are important tools for solving open domain machine comprehension tasks and have achieved remarkable accuracy in many question answering benchmark datasets. Existing models impose strong independence assumptions on hidden variables, which leaves the interaction among them unexplored. Here we introduce linguistic structures to help capturing global evidence in hidden variable modeling. In the proposed algorithms, question-answer pairs are scored based on structured inference results on parse trees and semantic frames, which aims to assign hidden variables in a global optimal way. Experiments on the MCTest dataset demonstrate that the proposed models are highly competitive with state-of-the-art machine comprehension systems.
Yuanbin Wu, Man Lan
AAAI2
2018 An Adversarial Joint Learning Model for Low-Resource Language Semantic Textual Similarity
Man Lan, Yuanbin Wu, Jingang Wang, Long Qiu, Sheng Li 0017, Jun Lang 0001, Luo Si
ECIR3
2018 Extracting Entities and Relations with Joint Minimum Risk Training
abstract
We investigate the task of joint entity relation extraction.Unlike prior efforts, we propose a new lightweight joint learning paradigm based on minimum risk training (MRT).Specifically, our algorithm optimizes a global loss function which is flexible and effective to explore interactions between the entity model and the relation model.We implement a strong and simple neural network where the MRT is executed.Experiment results on the benchmark ACE05 and NYT datasets show that our model is able to achieve state-of-the-art joint extraction performances.
Changzhi Sun, Yuanbin Wu, Man Lan, Shiliang Sun, Kuang-chih Lee, Kewen Wu 0003
EMNLP2
2018 Memory-Based Model with Multiple Attentions for Multi-turn Response Selection
Xingwu Lu, Man Lan, Yuanbin Wu
ICONIP (2)3
2018 A Neural Generation-based Conversation Model Using Fine-grained Emotion-guide Attention
abstract
Human emotion interaction is crucial to social communications. However, existing generation-based conversation systems mainly put emphasis on the content of responses in terms of naturalness, diversity and coherence without consideration of the emotion interaction between conversation. In order to reduce the gap between human-generated and computer-generated responses, in this work we present a human-like Emotional Conversation Generation Model, named ECGM, by imitating human conversation. Specifically, ECGM applies an emotion-guide attention which captures and integrates the emotion of the given post into neural response generation. Comparative experiments evaluated by computerised and manual methods show that our proposed model is capable of generating more human-like emotional responses and relevant content as well.
Man Lan, Yuanbin Wu
IJCNN3
2018 Memory-Based Matching Models for Multi-turn Response Selection in Retrieval-Based Chatbots
Xingwu Lu, Man Lan, Yuanbin Wu
NLPCC (1)3
2017 Large-scale Opinion Relation Extraction with Distantly Supervised Neural Network
abstract
We investigate the task of open domain opinion relation extraction. Different from works on manually labeled corpus, we propose an efficient distantly supervised framework based on pattern matching and neural network classifiers. The patterns are designed to automatically generate training data, and the deep learning model is design to capture various lexical and syntactic features. The result algorithm is fast and scalable on large-scale corpus. We test the system on the Amazon online review dataset. The result shows that our model is able to achieve promising performances without any human annotations.
Changzhi Sun, Yuanbin Wu, Man Lan, Shiliang Sun, Qi Zhang 0001
EACL (1)2
2017 Multi-task Attention-based Neural Networks for Implicit Discourse Relationship Representation and Identification
abstract
We present a novel multi-task attentionbased neural network model to address implicit discourse relationship representation and identification through two types of representation learning, an attentionbased neural network for learning discourse relationship representation with two arguments and a multi-task framework for learning knowledge from annotated and unannotated corpora.The extensive experiments have been performed on two benchmark corpora (i.e., PDTB and CoNLL-2016 datasets).Experimental results show that our proposed model outperforms the state-of-the-art systems on benchmark corpora.
Man Lan, Jianxiang Wang, Yuanbin Wu, Zhengyu Niu, Haifeng Wang 0001
EMNLP3
2017 An Effective Gated and Attention-Based Neural Network Model for Fine-Grained Financial Target-Dependent Sentiment Analysis
Mengxiao Jiang, Jianxiang Wang, Man Lan, Yuanbin Wu
KSEM4
2017 A Learning Error Analysis for Structured Prediction with Approximate Inference
abstract
In this work, we try to understand the differences between exact and approximate inference algorithms in structured prediction. We compare the estimation and approximation error of both underestimate and overestimate models. The result shows that, from the perspective of learning errors, performances of approximate inference could be as good as exact inference. The error analyses also suggest a new margin for existing learning algorithms. Empirical evaluations on text classification, sequential labelling and dependency parsing witness the success of approximate inference and the benefit of the proposed margin.
Yuanbin Wu, Man Lan, Shiliang Sun, Qi Zhang 0001, Xuanjing Huang 0001
NIPS1
2017 Effective Semantic Relationship Classification of Context-Free Chinese Words with Simple Surface and Embedding Features
Yunxiao Zhou, Man Lan, Yuanbin Wu
NLPCC3
2016 Building mutually beneficial relationships between question retrieval and answer ranking to improve performance of community question answering
abstract
In community-based question answering (CQA) domain, there are two main tasks, i.e., question retrieval and answer ranking. Previous studies addressed these two tasks in an independent manner or in a sequential fashion without information communication. In this work we propose a novel method to improve the performance of CQA by mutually promoting the two tasks with the help of each other. Specifically, we propose two methods to improve question retrieval task by utilizing the rank of answers or extracting novel features from Q-A pairs respectively. Meanwhile, to improve answer ranking, we also present novel features with the help of similar questions. Experimental results on benchmark dataset showed that this mutually beneficial strategy between question retrieval and answer ranking not only improved the individual performance of these two tasks but also improved the overall performance of CQA through reducing errors propagating from question retrieval to answer ranking.
Man Lan, GuoShun Wu, Chunyun Xiao, Yuanbin Wu, Ju Wu
IJCNN4
2015 An Online Learning Algorithm for Bilinear Models
abstract
We investigate the bilinear model, which is a matrix form linear model with the rank 1 constraint. A new online learning algorithm is proposed to train the model parameters. Our algorithm runs in the manner of online mirror descent, and gradients are computed by the power iteration. To analyze it, we give a new second order approximation of the squared spectral norm, which helps us to get a regret bound. Experiments on two sequential labelling tasks give positive results.
Yuanbin Wu, Shiliang Sun
ICML1
2013 Grammatical Error Correction Using Integer Linear Programming
Yuanbin Wu, Hwee Tou Ng
ACL (1)1
2011 Structural Opinion Mining for Graph-based Sentiment Representation
Yuanbin Wu, Qi Zhang 0001, Xuanjing Huang 0001, Lide Wu
EMNLP1
2011 Opinion Mining with Sentiment Graph
abstract
Opinion mining became an active research topic in recent years due to its wide range of applications. A number of companies offer opinion mining services. One problem that has not been well studied so far is the representation model. In this paper, we propose a novel sentence level sentiment representation model. By taking the observation that lots of sentences which have complicated opinion relations can not be represented well by slots filling or feature-based model, the novel representation model sentiment graph is described in this paper. A supervised structural learning method is presented and used to construct sentiment graphs from sentences. Experimental results in a manually labeled corpus are given to show the effectiveness of the proposed approach.
Qi Zhang 0001, Yuanbin Wu, Xuanjing Huang 0001
Web Intelligence2
2009 Phrase Dependency Parsing for Opinion Mining
Yuanbin Wu, Qi Zhang 0001, Xuanjing Huang 0001, Lide Wu
EMNLP1
2009 Mining product reviews based on shallow dependency parsing
abstract
This paper presents a novel method for mining product reviews, where it mines reviews by identifying product features, expressions of opinions and relations between them. By taking advantage of the fact that most of product features are phrases, a concept of shallow dependency parsing is introduced, which extends traditional dependency parsing to phrase level. This concept is then implemented for extracting relation between product features and expressions of opinions. Experimental evaluations show that the mining task can benefit from shallow dependency parsing.
Qi Zhang 0001, Yuanbin Wu, Tao Li 0001, Mitsunori Ogihara, Xuanjing Huang 0001
SIGIR2