Xueling Lin

dblp:83/1479 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 4 first-author · 4 since 2021Computer networks · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Dynamic Service Knowledge Base Construction at WeChat
Haoyang Li 0002, Alexander Zhou 0001, Fengmei Jin, Qing Li 0001, Ziyuan Zhao, Hao Xin, Qiang Yan 0001, Tiezheng Mao, Xueling Lin, Zijian Li 0002, Lei Chen 0002
ADMA (4)9
2021 Fine-Grained Entity Typing via Label Noise Reduction and Data Augmentation
Haoyang Li 0002, Xueling Lin, Lei Chen 0002
DASFAA (1)2
2021 CaSIE: Canonicalize and Informative Selection of the OpenIE system
abstract
Knowledge extraction has become a hot topic recently with the increasing number of applications needed for large-scale knowledge bases (KBs), such as semantic search and QA systems. The goal of knowledge extraction is to extract relations and their arguments from natural language text. Recent research proposes two kinds of solutions. The first one, called Closed IE, tries to construct KB through predefined features or rules with respect to a specific domain. It requires specifying the interested predicates in advance, which restricts its application to the domains where prior knowledge about the interested predicates must be given. The second one, called Open IE, tries to extract facts by using the parsing structure from the unstructured text. However, they cannot avoid extracting redundant facts. Such extractions can hardly be directly used to populate the existing KB. Moreover, many correct extractions are not relevant to the document, which limits the applications to understand the essential information that the document conveys. In this paper, we propose an end-to-end system which takes a target incomplete KB and documents as input. It first performs joint entity and relation linking to the existing KB based on both contexts of document and background KB information. Then it summarizes the extracted facts by considering the relevance to the document and the diversity between them. Extensive experiments over real datasets demonstrate the effectiveness and efficiency of the proposed methods.
Hao Xin, Xueling Lin, Lei Chen 0002
ICDE2
2021 TENET: Joint Entity and Relation Linking with Coherence Relaxation
abstract
The joint entity and relation linking task aims to connect the noun phrases (resp., relational phrases) extracted from natural language documents to the entities (resp., predicates) in general knowledge bases (KBs). This task benefits numerous downstream systems, such as question answering and KB population. Previous works on entity and relation linking rely on the global coherence assumption, i.e., entities and predicates within the same document are highly correlated with each other. However, this assumption is not always valid in many real-world scenarios. Due to KB incompleteness or data sparsity, sparse coherence among the entities and predicates within the same document is common. Moreover, there may exist isolated entities or predicates that are not related to any other linked concepts. In this paper, we propose TENET, a joint entity and relation linking technique, which relaxes the coherence assumption in an unsupervised manner. Specifically, we formulate the joint entity and relation linking task as a minimum-cost rooted tree cover problem on the knowledge coherence graph constructed based on the document. We then propose effective approximation algorithms with pruning strategies to solve this problem and derive the linking results. Extensive experiments on real-world datasets demonstrate the superior effectiveness and efficiency of our method against the state-of-the-art techniques.
Xueling Lin, Lei Chen 0002, Chaorui Zhang
SIGMOD Conference1
2020 TransN: Heterogeneous Network Representation Learning by Translating Node Embeddings
abstract
Learning network embeddings has attracted growing attention in recent years. However, most of the existing methods focus on homogeneous networks, which cannot capture the important type information in heterogeneous networks. To address this problem, in this paper, we propose TransN, a novel multi-view network embedding framework for heterogeneous networks. Compared with the existing methods, TransN is an unsupervised framework which does not require node labels or user-specified meta-paths as inputs. In addition, TransN is capable of handling more general types of heterogeneous networks than the previous works. Specifically, in our framework TransN, we propose a novel algorithm to capture the proximity information inside each single view. Moreover, to transfer the learned information across views, we propose an algorithm to translate the node embeddings between different views based on the dual-learning mechanism, which can both capture the complex relations between node embeddings in different views, and preserve the proximity information inside each view during the translation. We conduct extensive experiments on real-world heterogeneous networks, whose results demonstrate that the node embeddings generated by TransN outperform those of competitors in various network mining tasks.
Zijian Li 0002, Wenhao Zheng 0001, Xueling Lin, Ziyuan Zhao, Zhe Wang 0019, Yue Wang 0012, Xun Jian 0001, Lei Chen 0002, Qiang Yan 0001, Tiezheng Mao
ICDE3
2020 KBPearl: A Knowledge Base Population System Supported by Joint Entity and Relation Linking
abstract
Nowadays, most openly available knowledge bases (KBs) are incomplete, since they are not synchronized with the emerging facts happening in the real world. Therefore, knowledge base population (KBP) from external data sources, which extracts knowledge from unstructured text to populate KBs, becomes a vital task. Recent research proposes two types of solutions that partially address this problem, but the performance of these solutions is limited. The first solution, dynamic KB construction from unstructured text, requires specifications of which predicates are of interest to the KB, which needs preliminary setups and is not suitable for an in-time population scenario. The second solution, Open Information Extraction (Open IE) from unstructured text, has limitations in producing facts that can be directly linked to the target KB without redundancy and ambiguity. In this paper, we present an end-to-end system, KBPearl, for KBP, which takes an incomplete KB and a large corpus of text as input, to (1) organize the noisy extraction from Open IE into canonicalized facts; and (2) populate the KB by joint entity and relation linking, utilizing the context knowledge of the facts and the side information inferred from the source text. We demonstrate the effectiveness and efficiency of KBPearl against the state-of-the-art techniques, through extensive experiments on real-world datasets.
Xueling Lin, Haoyang Li 0002, Hao Xin, Zijian Li 0002, Lei Chen 0002
Proc. VLDB Endow.1
2020 Circa: collaborative code offloading among multiple mobile devices
Xueling Lin, Jingjie Jiang, Calvin Hong Yi Li, Bo Li 0001, Baochun Li
Wirel. Networks1
2019 Canonicalization of Open Knowledge Bases with Side Information from the Source Text
abstract
Nowadays Open Information Extraction (Open IE) approaches, which extracttriples from unstructured text, contribute to the construction of large Open Knowledge Bases (Open KBs). However, one crucial problem is that the noun phrases and relation phrases in the extracted triples are not well canonicalized, which leads to a large number of redundant and ambiguous facts. For example, bothandmay be extracted and stored in Open KBs. Recent research proposes to solve this problem by clustering over manually-defined feature spaces based on the similarity of the noun phrases and relation phrases. However, the performance of such techniques is limited, since only the information contained in the triples is utilized to measure their similarity. In this paper, we propose to perform canonicalization over Open IE triples by incorporating the side information from the original data sources, including the candidate entities of the noun phrases detected in the source text, the types of the candidate entities and the domain knowledge of the source text. We model the canonicalization problem of noun phrases and relation phrases jointly based on such side information, and demonstrate the effectiveness of our approach through extensive experiments on two real-world datasets.
Xueling Lin, Lei Chen 0002
ICDE1
2018 Domain-Aware Multi-Truth Discovery from Conflicting Sources
abstract
In the Big Data era, truth discovery has served as a promising technique to solve conflicts in the facts provided by numerous data sources. The most significant challenge for this task is to estimate source reliability and select the answers supported by high quality sources. However, existing works assume that one data source has the same reliability on any kinds of entity, ignoring the possibility that a source may vary in reliability on different domains. To capture the influence of various levels of expertise in different domains, we integrate domain expertise knowledge to achieve a more precise estimation of source reliability. We propose to infer the domain expertise of a data source based on its data richness in different domains. We also study the mutual influence between domains, which will affect the inference of domain expertise. Through leveraging the unique features of the multi-truth problem that sources may provide partially correct values of a data item, we assign more reasonable confidence scores to value sets. We propose an integrated Bayesian approach to incorporate the domain expertise of data sources and confidence scores of value sets, aiming to find multiple possible truths without any supervision. Experimental results on two real-world datasets demonstrate the feasibility, efficiency and effectiveness of our approach.
Xueling Lin, Lei Chen 0002
Proc. VLDB Endow.1
2015 Circa: Offloading collaboratively in the same vicinity with iBeacons
abstract
Code offloading to remote infrastructures has been a common practice for mobile users who seek extra power or computing resources to perform computation-intensive tasks. Existing works, however, have so far mainly focused on code offloading from a single mobile device to remote cloud servers, which restricts the potential of code offloading only to devices with available Internet access. In this paper, we propose Circa, a new framework that demonstrates the feasibility of code offloading among multiple mobile devices in close proximity to one another, leveraging the presence of iBeacons. Our objective is to eliminate the costs incurred by running virtual machine instances in the cloud, and the need to connect remotely to the cloud as well. With the assistance of iBeacons, devices in the same vicinity can discover and support one another through collaborative code offloading with short-range communication, obviating the need for centralized servers. We have implemented Circa on the iOS platform and validated its feasibility using iOS devices. According to our experimental results, with more than two collaborators, Circa is capable of reducing the total execution time of an offloaded task substantially, while preserving satisfactory performance of the mobile application.
Xueling Lin, Jingjie Jiang, Bo Li 0001, Baochun Li
ICC1