Xin Mao 0002

dblp:72/4332-2 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
9since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Bread: A Hybrid Approach for Instruction Data Mining Through Balanced Retrieval and Dynamic Data Sampling
Xinlin Zhuang, Xin Mao 0002, Hongyi Wu, Shangqing Zhao, Yuxiang Song, Chenghao Jia, Man Lan
NLPCC (2)2
2023 An Effective and Efficient Time-aware Entity Alignment Framework via Two-aspect Three-view Label Propagation
abstract
Entity alignment (EA) aims to find the equivalent entity pairs between different knowledge graphs (KGs), which is crucial to promote knowledge fusion. With the wide use of temporal knowledge graphs (TKGs), time-aware EA (TEA) methods appear to enhance EA. Existing TEA models are based on Graph Neural Networks (GNN) and achieve state-of-the-art (SOTA) performance, but it is difficult to transfer them to large-scale TKGs due to the scalability issue of GNN. In this paper, we propose an effective and efficient non-neural EA framework between TKGs, namely LightTEA, which consists of four essential components: (1) Two-aspect Three-view Label Propagation, (2) Sparse Similarity with Temporal Constraints, (3) Sinkhorn Operator, and (4) Temporal Iterative Learning. All of these modules work together to improve the performance of EA while reducing the time consumption of the model. Extensive experiments on public datasets indicate that our proposed model significantly outperforms the SOTA methods for EA between TKGs, and the time consumed by LightTEA is only dozens of seconds at most, no more than 10% of the most efficient TEA method.
Xin Mao 0002, Youshao Xiao, Changxu Wu, Man Lan
IJCAI2
2022 An Effective and Efficient Entity Alignment Decoding Algorithm via Third-Order Tensor Isomorphism
abstract
Xin Mao, Meirong Ma, Hao Yuan, Jianchao Zhu, ZongYu Wang, Rui Xie, Wei Wu, Man Lan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Xin Mao 0002, Meirong Ma, Jianchao Zhu, Zongyu Wang, Man Lan
ACL (1)1
2022 A Simple Temporal Information Matching Mechanism for Entity Alignment between Temporal Knowledge Graphs
abstract
Entity alignment (EA) aims to find entities in different knowledge graphs (KGs) that refer to the same object in the real world. Recent studies incorporate temporal information to augment the representations of KGs. The existing methods for EA between temporal KGs (TKGs) utilize a time-aware attention mechanisms to incorporate relational and temporal information into entity embeddings. The approaches outperform the previous methods by using temporal information. However, we believe that it is not necessary to learn the embeddings of temporal information in KGs since most TKGs have uniform temporal representations. Therefore, we propose a simple GNN model combined with a temporal information matching mechanism, which achieves better performance with less time and fewer parameters. Furthermore, since alignment seeds are difficult to label in real-world applications, we also propose a method to generate unsupervised alignment seeds via the temporal information of TKG. Extensive experiments on public datasets indicate that our supervised method significantly outperforms the previous methods and the unsupervised one has competitive performance.
Xin Mao 0002, Meirong Ma, Jianchao Zhu, Man Lan
COLING2
2022 LightEA: A Scalable, Robust, and Interpretable Entity Alignment Framework via Three-view Label Propagation
abstract
Entity Alignment (EA) aims to find equivalent entity pairs between KGs, which is the core step of bridging and integrating multi-source KGs.In this paper, we argue that existing GNNbased EA methods inherit the inborn defects from their neural network lineage: weak scalability and poor interpretability.Inspired by recent studies, we reinvent the Label Propagation algorithm to effectively run on KGs and propose a non-neural EA framework -LightEA, consisting of three efficient components: (i) Random Orthogonal Label Generation, (ii) Three-view Label Propagation, and (iii) Sparse Sinkhorn Iteration.According to the extensive experiments on public datasets, LightEA has impressive scalability, robustness, and interpretability.With a mere tenth of time consumption, LightEA achieves comparable results to state-of-the-art methods across all datasets and even surpasses them on many.
Xin Mao 0002, Yuanbin Wu, Man Lan
EMNLP1
2021 Target-dependent Event Detection: A New Task to Event Extraction from News
abstract
Event extraction aims to detect events and extract event arguments. However, various events are not only too nuanced and complex to distinguish, but also involve multiple entities in the real-world scenario, especially in the financial field. This brings a great challenge to the current event extraction. To address these problems, previous event-centric methods detect events first and then extract arguments. Due to the diversity and complexity of events, event detection has a low performance, which is unfit for the huge amount of news in the real world. Given that the performance of named entity recognition (NER) is satisfactory, we shift our perspective from event-centric to target-centric view. In this paper, we propose a new task: target-dependent event detection (TDED), which aims to extract target entities and detect their corresponding events. We also propose a semantic and syntactic aware approach to support thousands of target entity extraction first and dozens of event types detection, that can be applied to massive corpora. Experimental results on a real-world Chinese financial dataset demonstrate that our model outperforms previous methods, especially in complex scenarios.
Xin Mao 0002, Meirong Ma, Jianchao Zhu, Man Lan
IEEE BigData2
2021 Are Negative Samples Necessary in Entity Alignment?: An Approach with High Performance, Scalability and Robustness
abstract
Entity alignment (EA) aims to find the equivalent entities in different KGs, which is a crucial step in integrating multiple KGs. However, most existing EA methods have poor scalability and are unable to cope with large-scale datasets. We summarize three issues leading to such high time-space complexity in existing EA methods: (1) Inefficient graph encoders, (2) Dilemma of negative sampling, and (3) "Catastrophic forgetting" in semi-supervised learning. To address these challenges, we propose a novel EA method with three new components to enable high Performance, high Scalability, and high Robustness (PSR): (1) Simplified graph encoder with relational graph sampling, (2) Symmetric negative-free alignment loss, and (3) Incremental semi-supervised learning. Furthermore, we conduct detailed experiments on several public datasets to examine the effectiveness and efficiency of our proposed method. The experimental results show that PSR not only surpasses the previous SOTA in performance but also has impressive scalability and robustness.
Xin Mao 0002, Yuanbin Wu, Man Lan
CIKM1
2021 From Alignment to Assignment: Frustratingly Simple Unsupervised Entity Alignment
abstract
Cross-lingual entity alignment (EA) aims to find the equivalent entities between crosslingual KGs (Knowledge Graphs), which is a crucial step for integrating KGs.Recently, many GNN-based EA methods are proposed and show decent performance improvements on several public datasets.However, existing GNN-based EA methods inevitably inherit poor interpretability and low efficiency from neural networks.Motivated by the isomorphic assumption of GNN-based methods, we successfully transform the cross-lingual EA problem into an assignment problem.Based on this re-definition, we propose a frustratingly Simple but Effective Unsupervised entity alignment method (SEU) without neural networks.Extensive experiments have been conducted to show that our proposed unsupervised approach even beats advanced supervised methods across all public datasets while having high efficiency, interpretability, and stability.
Xin Mao 0002, Yuanbin Wu, Man Lan
EMNLP (1)1
2021 Boosting the Speed of Entity Alignment 10 ×: Dual Attention Matching Network with Normalized Hard Sample Mining
abstract
Seeking the equivalent entities among multi-source Knowledge Graphs (KGs) is the pivotal step to KGs integration, also known as entity alignment (EA). However, most existing EA methods are inefficient and poor in scalability. A recent summary points out that some of them even require several days to deal with a dataset containing 200,000 nodes (DWY100K). We believe over-complex graph encoder and inefficient negative sampling strategy are the two main reasons. In this paper, we propose a novel KG encoder — Dual Attention Matching Network (Dual-AMN), which not only models both intra-graph and cross-graph information smartly, but also greatly reduces computational complexity. Furthermore, we propose the Normalized Hard Sample Mining Loss to smoothly select hard negative samples with reduced loss shift. The experimental results on widely used public datasets indicate that our method achieves both high accuracy and high efficiency. On DWY100K, the whole running process of our method could be finished in 1,100 seconds, at least 10 × faster than previous work. The performances of our method also outperform previous works across all datasets, where [email protected] and MRR have been improved from 6% to 13%.
Xin Mao 0002, Yuanbin Wu, Man Lan
WWW1
2020 Relational Reflection Entity Alignment
abstract
Entity alignment aims to identify equivalent entity pairs from different Knowledge Graphs (KGs), which is essential in integrating multi-source KGs. Recently, with the introduction of GNNs into entity alignment, the architectures of recent models have become more and more complicated. We even find two counter-intuitive phenomena within these methods: (1) The standard linear transformation in GNNs is not working well. (2) Many advanced KG embedding models designed for link prediction task perform poorly in entity alignment. In this paper, we abstract existing entity alignment methods into a unified framework, Shape-Builder & Alignment, which not only successfully explains the above phenomena but also derives two key criteria for an ideal transformation operation. Furthermore, we propose a novel GNNs-based method, Relational Reflection Entity Alignment (RREA). RREA leverages Relational Reflection Transformation to obtain relation specific embeddings for each entity in a more efficient way. The experimental results on real-world datasets show that our model significantly outperforms the state-of-the-art methods, exceeding by 5.8%-10.9% on [email protected]
Xin Mao 0002, Yuanbin Wu, Man Lan
CIKM1
2020 MRAEA: An Efficient and Robust Entity Alignment Approach for Cross-lingual Knowledge Graph
abstract
Entity alignment to find equivalent entities in cross-lingual Knowledge Graphs (KGs) plays a vital role in automatically integrating multiple KGs. Existing translation-based entity alignment methods jointly model the cross-lingual knowledge and monolingual knowledge into one unified optimization problem. On the other hand, the Graph Neural Network (GNN) based methods either ignore the node differentiations, or represent relation through entity or triple instances. They all fail to model the meta semantics embedded in relation nor complex relations such as n-to-n and multi-graphs. To tackle these challenges, we propose a novel Meta Relation Aware Entity Alignment (MRAEA) to directly model cross-lingual entity embeddings by attending over the node's incoming and outgoing neighbors and its connected relations' meta semantics. In addition, we also propose a simple and effective bi-directional iterative strategy to add new aligned seeds during training. Our experiments on all three benchmark entity alignment datasets show that our approach consistently outperforms the state-of-the-art methods, exceeding by 15%-58% on [email protected] Through an extensive ablation study, we validate that the proposed meta relation aware representations, relation aware self-attention and bi-directional iterative strategy of new seed selection all make contributions to significant performance improvement. The code is available at https://github.com/MaoXinn/MRAEA.
Xin Mao 0002, Man Lan, Yuanbin Wu
WSDM1
2019 Scaling up Open Tagging from Tens to Thousands: Comprehension Empowered Attribute Value Extraction from Product Title
abstract
Supplementing product information by extracting attribute values from title is a crucial task in e-Commerce domain.Previous studies treat each attribute only as an entity type and build one set of NER tags (e.g., BIO) for each of them, leading to a scalability issue which unfits to the large sized attribute system in real world e-Commerce.In this work, we propose a novel approach to support value extraction scaling up to thousands of attributes without losing performance: (1) We propose to regard attribute as a query and adopt only one global set of BIO tags for any attributes to reduce the burden of attribute tag or model explosion;(2) We explicitly model the semantic representations for attribute and title, and develop an attention mechanism to capture the interactive semantic relations in-between to enforce our framework to be attribute comprehensive.We conduct extensive experiments in real-life datasets.The results show that our model not only outperforms existing state-of-the-art N-ER tagging models, but also is robust and generates promising results for up to 8, 906 attributes.
Xin Mao 0002, Man Lan
ACL (1)3