Xin Wang 0030

dblp:10/5630-30 · DBLP profile ↗
← Back
88ranked-venue papers in the field
8as first author
57since 2021 · last 2026
0000-0001-9651-0651ORCID · conflict

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 29 (1 first)Information Retrieval & Web Search · 29 (3 first)Knowledge Engineering, Semantic Web & Information Systems · 16Data Mining & Knowledge Discovery · 7Other / Interdisciplinary · 5 (4 first)Big Data, Cloud & Distributed Data Systems · 1Business Process & Enterprise Data · 1
YearPublicationVenuePosition
2026 ConvD: Attention Enhanced Dynamic Convolutional Embeddings for Knowledge Graph Completion (Extended Abstract)
Zhao Li 0009, Xin Wang 0030, Ye Yuan 0001
ICDE3
2026 HyCubE: Efficient Knowledge Hypergraph 3D Circular Convolutional Embedding (Extended Abstract)
Zhao Li 0009, Xin Wang 0030
ICDE2
2026 KG-BiLM: Knowledge Graph Embedding via Bidirectional Language Models
Xin Wang 0030, Zhao Li 0009, Dongxiao He, Yanbing Li, Wushour Slamu
WWW2
2026 Many Hands Make Light Work: Group-based Information Diffusion Prediction over Long-Context Cascades
abstract
Information diffusion prediction aims to forecast the temporal spread of opinions and behaviors by identifying potential adopters. Existing methods typically treat information diffusion as a sequence of individual adoptions and rely on computationally expensive pairwise (one-to-one) influence computations, often restricting predictions to just the next adopter. This individual-level paradigm both misrepresents real-world collective (many-to-many) influences and suffers a critical efficiency trade-off: to remain feasible, such models must truncate long diffusion histories, thereby overlooking early initiators and opinion leaders. To overcome these limitations, we formalize a more practical task: Group-based Information Diffusion Prediction, and propose an effective and scalable GRID framework. Specifically, GRID first learns group-oriented graph embeddings via a task-regularized information bottleneck objective, which amplifies key influence pathways and produces reliable user embeddings for group identification. Built on these embeddings, the core GroupAttn module captures inter-group influence while reducing complexity from quadratic to linear in cascade length. This enables the modeling of ultra-long cascades (exceeding 10,000 users) without truncation while preserving representational fidelity within a provable error bound. Finally, a group-wise objective guides the model to predict semantically meaningful future groups. Extensive experiments on four real-world datasets show that GRID outperforms ten state-of-the-art baselines by an average of 10.65% in accuracy, while achieving an order-of-magnitude gain in efficiency and extending the supported cascade length by up to 10 times.
Zihan Feng 0001, Yajun Yang, Xin Huang 0001, Xin Wang 0030, Hong Gao 0001, Qinghua Hu
WWW4
2026 ReaLM: Residual Quantization Bridges Knowledge Graph Embeddings and Large Language Models
abstract
Large Language Models (LLMs) have recently emerged as a powerful paradigm for Knowledge Graph Completion (KGC), offering strong reasoning and generalization capabilities beyond traditional embedding-based approaches. However, existing LLM-based methods often struggle to fully exploit structured semantic representations, as the continuous embedding space of pretrained KG models is fundamentally misaligned with the discrete token space of LLMs. This discrepancy hinders effective semantic transfer and limits their performance. To address this challenge, we propose ReaLM, a novel and effective framework that bridges the gap between KG embeddings and LLM tokenization through the mechanism of residual vector quantization. ReaLM discretizes pretrained KG embeddings into compact code sequences and integrates them as learnable tokens within the LLM vocabulary, enabling seamless fusion of symbolic and contextual knowledge. Furthermore, we incorporate ontology-guided class constraints to enforce semantic consistency, refining entity predictions based on class-level compatibility. Extensive experiments on two widely used benchmark datasets demonstrate that ReaLM achieves state-of-the-art performance, confirming its effectiveness in aligning structured knowledge with large-scale language models. The implementation is publicly available at https://github.com/xiumu-gg/ReaLM.
Xin Wang 0030, Jiaoyan Chen 0001, Lingbing Guo, Zhao Li 0009
WWW2
2026 DTransKT: A Dual Transferable Knowledge Tracing Framework for Cross-Disciplinary Self-Adaptation
abstract
Knowledge Tracing (KT) is pivotal in intelligent tutoring systems, as it models the dynamic evolution of student knowledge from their learning interactions. However, the cross-disciplinary generalization of existing KT models is subjected to a dual constraint: the heterogeneity in students' cognitive abilities and the divergent disciplinary-specific knowledge structures.To address this challenge, we propose DTransKT, a dual transferable knowledge tracing framework tailored for cross-disciplinary adaptability. DTransKT enhances existing knowledge tracing models by dynamically aligning student representations and integrating external knowledge semantics.Specifically, the framework incorporates a Cross-disciplinary Graph-matching (CG) module, which captures meta-skill representations based on students' learning trajectories. Through cross-disciplinary node matching, the CG module aligns student-specific features, thereby improving tracing accuracy. Additionally, the Cross-disciplinary Attention-assisting (CA) module leverages pre-trained language models to extract meta-semantic from textual content, enhancing transferability.Extensive experimental evaluations demonstrate that DTransKT consistently enhances the performance of seven prominent KT models under direct transfer settings, achieving average improvements of 14.2% in accuracy (ACC) and 4.5% in area under the curve (AUC) across diverse datasets. These findings affirm the efficacy of our approach in enabling cross-disciplinary transfer for knowledge tracing. Code and pre-trained models are available at: https://github.com/Dual-KT/DTransKT.
Kun Liang 0002, Jiake Ge, Xin Wang 0030, Yiying Zhang 0004
WWW3
2026 LLM-Driven Semantic ID for Information Diffusion Prediction
Haoshuang Liu, Zihan Feng 0001, Yajun Yang, Xin Wang 0030, Hong Gao 0001, Qinghua Hu
WWW4
2025 DO: An Efficient Deep Reinforcement Learning Approach for Optimal Route with Collective Spatial Keywords
abstract
Given a source, destination, and required keywords, the Optimal Route with Collective Spatial Keywords ( ORCSK ) query aims to find the shortest route covering all keywords. Existing Point of Interest (POI) candidate set-based and path expansion-based methods frequently produce inferior route quality or excessive time overhead, particularly under large-scale query keywords. To address this challenge, we introduce the DO framework, which pioneers the employ Deep Reinforcement Learning for the ORCSK. Specifically, DO first integrates the spatial index with the H2H index to generate and refine high-quality candidate sets. Subsequently, DO utilizes a Transformer-based model to determine the optimal route from the sets. To effectively combine spatial distance and POI attributes, we propose a novel dual-cross encoder architecture. Furthermore, leveraging this architecture, we introduce a multi-route generating strategy, exploiting parallel computing to enhance route quality. Our experiments on real-life road networks demonstrate superior route quality and response time compared to the state-of-the-art method, with an average improvement of 1-2 orders of magnitude in response time, and maintain high efficiency even under large-scale query keywords or dynamic POI attributes scenarios.
Jiajia Li 0003, Jiming Dong, Lei Li 0003, Yu Yang 0012, Xin Wang 0030, Mengxuan Zhang 0001
CIKM5
2025 QiboGraph: A Knowledge Graph for Traditional Chinese Medicine
Xin Wang 0030, Yongzhe Jia
DASFAA (6)3
2025 MetaEformer: Unveiling and Leveraging Meta-Patterns for Complex and Dynamic Systems Load Forecasting
Shaoyuan Huang, Tiancheng Zhang 0009, Zhongtian Zhang, Xiaofei Wang 0001, Lanjun Wang, Xin Wang 0030
KDD (2)6
2025 Ontology-Enhanced Knowledge Graph Completion Using Large Language Models
Xin Wang 0030, Jiaoyan Chen 0001, Zhao Li 0009
ISWC (1)2
2025 Segmentation Similarity Enhanced Semantic Related Entity Fusion for Multi-modal Knowledge Graph Completion
abstract
Multi-modal Knowledge Graph Completion (MKGC) aims at leveraging multi-modal information to infer missing objective facts in incomplete multi-modal knowledge graphs, thereby significantly enhancing their expressive capabilities. The segmentation of semantic data, including image segmentation and word-level descriptions, often contain implicit relationships between entities that are frequently overlooked by existing methodologies, thus limiting the effectiveness of reasoning tasks. Therefore, we propose a novel completion inference method based on fine-grained semantic segmentation, which enhances reasoning capability by utilizing implicit relationships between entities. Primarily, we introduce the concept of Semantic Related Entity (SRE) and a novel SRE selection algorithm, which captures the semantic neighboring relationships of entities based on segmentation semantic similarity to fully exploit the semantic association information. Subsequently, we propose a Multi-modal Related Entity Fusion Transformer (M-REFT) model to effectively utilize SREs from semantic modalities and neighbors from structural modality for completion inference. The M-REFT employs a hierarchical Transformer architecture to encode the fusion modality representation between each entity and its SREs, and then decode the triplet representation with the neighbor information to identify missing entities in incomplete triplets. We conducted extensive comparative experiments with several state-of-the-art models on three datasets, demonstrating the significant performance advantages of M-REFT. A series of ablation experiments and case studies further validate the rationality and necessity of the SRE concept and the SRE selection algorithm.
Bo Ning 0002, Xin Wang 0030, Chengfei Liu
SIGIR3
2025 TCM-Eval: A Multi-dimensional Benchmark Framework for Evaluating Large Language Models in Traditional Chinese Medicine
Meihan Zhang, Xin Wang 0030, Hanshuo Xing, Pengwei Zhuang
WISE (2)2
2025 HySAE: An Efficient Semantic-Enhanced Representation Learning Model for Knowledge Hypergraph Link Prediction
abstract
Representation learning technique is an effective link prediction paradigm to alleviate the incompleteness of knowledge hypergraphs. However, the n-ary complex semantic information inherent in knowledge hypergraphs causes existing methods to face the dual limitations of weak effectiveness and low efficiency. In this paper, we propose a novel knowledge hypergraph representation learning model, HySAE, which can achieve a satisfactory trade-off between effectiveness and efficiency. Concretely, HySAE builds an efficient semantic-enhanced 3D scalable end-to-end embedding architecture to sufficiently capture knowledge hypergraph n-ary complex semantic information with fewer parameters, which can significantly reduce the computational cost of the model. In particular, we also design an efficient position-aware entity role semantic embedding way and two enhanced semantic learning strategies to further improve the effectiveness and scalability of our proposed method. Extensive experimental results on all datasets demonstrate that HySAE consistently outperforms state-of-the-art baselines, with an average improvement of 9.15%, a maximum improvement of 39.44%, an average 10.39x faster, and 75.79% fewer parameters.
Zhao Li 0009, Xin Wang 0030, Jianxin Li 0001
WWW2
2025 ShapeShifter: Workload-Aware Adaptive Evolving Index Structures Based on Learned Models
abstract
In real-world tasks like data management and Web search, index operations often exhibit strong skewness, unlike standard benchmarks with uniform data distribution. While learned indexes improve query and update efficiency, they typically fail to address the skewed workload access, often prioritizing a single performance metric at the cost of overall index effectiveness. Additionally, the full reliance on learned models can increase vulnerability to attacks, compromising system stability. To address these challenges, we propose ShapeShifter, an adaptive evolutionary structure based on traditional indexes, capable of dynamically adjusting node structures according to the workload. ShapeShifter introduces a node evolution strategy with workload-skew-aware policies to adaptively adjust and optimize the partial index structure, leveraging a hybrid mechanism that combines traditional and learned structures for robust performance and optimal time-space tradeoff under skewed workloads and extreme data conditions. The evaluation results show that ShapeShifter achieves the optimal tradeoff while maintaining robustness.
Hui Wang 0074, Xin Wang 0030, Jiake Ge, Lei Liang 0002
WWW2
2025 Large Language Model Enhanced Knowledge Representation Learning: A Survey
abstract
Abstract Knowledge Representation Learning (KRL) is crucial for enabling applications of symbolic knowledge from Knowledge Graphs (KGs) to downstream tasks by projecting knowledge facts into vector spaces. Despite their effectiveness in modeling KG structural information, KRL methods are suffering from the sparseness of KGs. The rise of Large Language Models (LLMs) built on the Transformer architecture presents promising opportunities for enhancing KRL by incorporating textual information to address information sparsity in KGs. LLM-enhanced KRL methods, including three key approaches, encoder-based methods that leverage detailed contextual information, encoder-decoder-based methods that utilize a unified Seq2Seq model for comprehensive encoding and decoding, and decoder-based methods that utilize extensive knowledge from large corpora, have significantly advanced the effectiveness and generalization of KRL in addressing a wide range of downstream tasks. This work provides a broad overview of downstream tasks while simultaneously identifying emerging research directions in these evolving domains.
Xin Wang 0030, Haofen Wang, Leong Hou U, Zhao Li 0009
Data Sci. Eng.1
2025 High Performance or Low Memory? An Updatable Learned Index Framework for Time-Space Tradeoff
abstract
The first generation of learned indexes inherently achieved lower space overhead than traditional index structures, establishing this advantage as one of the pivotal research directions in index optimization. However, in their pursuit of peak performance, designers often significantly increase space overhead, which becomes infeasible in scenarios with limited storage space. Furthermore, the design of current learned indexes optimized for time-space tradeoff is flawed, as they collapse catastrophically under prevalent dense or duplicate insertion workloads. To address these challenges, we first quantitatively analyze the time-space correlation characteristics of learned indexes from a theoretical perspective and identify the core influencing factors. Based on this, time-space cost minimization function models are established and an updatable learned index framework, LIFT, is constructed. Furthermore, LIFT incorporates specifically designed structural adjustment mechanisms to effectively counter existing poisoning attacks, significantly enhancing index robustness without increasing time-space cost. Evaluation results demonstrate that LIFT consistently achieves the optimal time-space tradeoff across various workloads and datasets, outperforming all other state-of-the-art indexes.
Hui Wang 0074, Xin Wang 0030, Jiake Ge, Yunpeng Chai, Lei Liang 0002
Proc. ACM Manag. Data2
2025 ConvD: Attention Enhanced Dynamic Convolutional Embeddings for Knowledge Graph Completion
abstract
Knowledge graphs often suffer from incompleteness issues, which can be alleviated through information completion. However, current state-of-the-art deep knowledge convolutional embedding models rely on external convolution kernels and conventional convolution processes, which limits the feature interaction capability of the model. This paper introduces a novel dynamic convolutional embedding model, named ConvD, which directly reshapes relation embeddings into multiple internal convolution kernels. This approach effectively enhances the feature interactions between relation embeddings and entity embeddings. Simultaneously, we incorporate a priori knowledgeoptimized attention mechanism that assigns distinct contribution weights to multiple relational convolution kernels during dynamic convolution, further boosting the expressive power of the model. Extensive experiments on various datasets show that our proposed model consistently outperforms the state-of-the-art baseline methods, with average improvements ranging from 3.28% to 14.69% across all the evaluation metrics, while the number of parameters is reduced by 50.66% to 85.40% compared to other state-of-the-art models.
Zhao Li 0009, Xin Wang 0030, Jianxin Li 0001, Ye Yuan 0001
IEEE Trans. Knowl. Data Eng.3
2025 HyCubE: Efficient Knowledge Hypergraph 3D Circular Convolutional Embedding
abstract
Knowledge hypergraph embedding models are usually computationally expensive due to the inherent complex semantic information. However, existing works mainly focus on improving the effectiveness of knowledge hypergraph embedding, making the model architecture more complex and redundant. It is desirable and challenging for knowledge hypergraph embedding to reach a trade-off between model effectiveness and efficiency. In this paper, we propose an end-to-end efficient knowledge hypergraph embedding model, HyCubE, which designs a novel3D circular convolutional neural networkand thealternate mask stackstrategy to enhance the interaction and extraction of feature information comprehensively. Furthermore, our proposed model achieves a better trade-off between effectiveness and efficiency by adaptively adjusting the 3D circular convolutional layer structure to handle$n$-ary knowledge tuples of different arities with fewer parameters. In addition, we use a knowledge hypergraph 1-N multilinear scoring way to accelerate the model training efficiency further. Finally, extensive experimental results on all datasets demonstrate that our proposed model consistently outperforms state-of-the-art baselines, with an average improvement of 8.22% and a maximum improvement of 33.82% across all metrics. Meanwhile, HyCubE is 6.12x faster, GPU memory usage is 52.67% lower, and the number of parameters is reduced by 85.21% compared with the average metric of the latest state-of-the-art baselines.
Zhao Li 0009, Xin Wang 0030, Jianxin Li 0001
IEEE Trans. Knowl. Data Eng.2
2024 Multi-level Contrastive Learning on Weak Social Networks for Information Diffusion Prediction
Zihan Feng 0001, Yajun Yang, Hong Gao 0001, Xin Wang 0030, Qinghua Hu
DASFAA (6)5
2024 Graphologue: Bridging RDBMS and Graph Databases with Natural Language Interfaces
Yongzhe Jia, Jianguo Wei, Xin Wang 0030, Xintian Zuo, Yuxuan Yang 0006
DASFAA (7)3
2024 Incremental Graph Computation: Anchored Vertex Tracking in Dynamic Social Networks (Extended Abstract)
abstract
User engagement has recently received significant attention in understanding the decay and expansion of communities in many online social networking platforms. Many user engagement studies have been conducted to find a set of critical (anchored) users in the static social network. However, social networks are highly dynamic and their structures are continuously evolving. In this paper, we target a new research problem called Anchored Vertex Tracking (AVT), aiming to track the anchored users at each timestamp of evolving networks. To address the AVT problem, we develop a greedy algorithm inspired by the previous anchored k-core study in the static networks. Furthermore, we design an incremental algorithm to efficiently solve the AVT problem by utilizing the smoothness of the network structure's evolution. The extensive experiments demonstrate the performance of our proposed algorithms.
Taotao Cai, Shuiqiao Yang, Jianxin Li 0001, Quan Z. Sheng, Jian Yang 0001, Xin Wang 0030, Wei Zhang 0098, Longxiang Gao
ICDE6
2024 Two Birds One Stone: Dual-Role Path Based Subgraph Matching Using Partial Evaluation
Chengguo Li, Xin Wang 0030, Yongqi Yin, Hui Wang 0074
WISE (2)2
2024 RPQBench: A Benchmark for Regular Path Queries on Graph Data
Hui Wang 0074, Xin Wang 0030, Menglu Ma, Yiheng You
WISE (2)2
2024 VQFT: A Visual Query Approach Based on Full-Text Search for Knowledge Graphs
abstract
Existing knowledge graph query approaches, whether traditional textual query languages or visual query languages, have steep learning curves that are unfriendly for non-expert users. This demonstration presents a Visual Query approach based on Full-Text search for knowledge graphs, called VQFT, which simplifies the process of querying knowledge graphs for users. Inspired by full-text search techniques, VQFT aims to combine the user-friendliness of visual query with the intuitiveness of full-text search , enabling users to query knowledge graphs as straightforward as using a search engine. Faceted full-text indexes, visual query constructor , and an interactive user interface are designed to achieve this goal. User tests and surveys have demonstrated that VQFT is more user-friendly and easier to learn than existing methods, which simplifies the construction of knowledge graph queries for non-expert users.
Zhaozhuo Li, Xin Wang 0030, Meng Wang 0009, Yajun Yang, Bohan Li 0001
Proc. VLDB Endow.2
2024 Extending Graph Rules with Oracles
abstract
This paper proposes a class of graph rules for deducing associations between entities, referred to as Graph Rules with Oracles and denoted by GROs. As opposed to previous graph rules, GROs support oracle functions to import (a) external knowledge, and (b) internal computations such as aggregate operators and machine learning predicates, and so on. Moreover, the semantics of GROs are defined in terms of pivoted dual simulation, in contrast to the subgraph isomorphism. We show how GROs can be used to predict links and catch anomalies, among other things. We formalize the association deduction problem with GROs in terms of the chase, and prove their Church-Rosser property. We show that both the deduction and incremental deduction problems with GROs are in PTIME, as opposed to the intractability of their counterparts with prior graph rules. We also provide sequential and parallel algorithms for association deduction and incremental deduction. Using real-life and synthetic graphs, we experimentally verify the effectiveness, scalability, and efficiency of the algorithms.
Bowen Dong 0004, Wenzhi Fu, Xin Wang 0030, Wenjun Wang 0002
Proc. VLDB Endow.5
2024 HJE: Joint Convolutional Representation Learning for Knowledge Hypergraph Completion
abstract
Knowledge hypergraph representation learning, which projects entities and$n$-ary relations into a low-dimensional vector space, remains a challenging area to be explored despite the ubiquity of$n$-ary relational facts in the real world. Current methods are always extensions of those used for knowledge graphs with shallow or deep structures. However, shallow and linear models limit the extraction capacity of the latent knowledge, while deep and non-linear models lead to the overabundance of parameters. In this paper, we propose a novel knowledge hypergraph completion model called HJE, which utilizes the powerful capability of convolutional neural networks for efficient representation learning. Interaction-enhanced 3D convolution and relation-aware 2D convolution are jointly utilized by HJE to extract explicit and implicit global knowledge and semantic information effectively without compromising the translation property of the model. Moreover, HJE constructs a unified learnable embedding matrix to capture entity position information in knowledge tuples. The entity mask mechanism can naturally couple the multilinear scoring approach for$n$-ary facts to speed up the training convergence of the model. Extensive experimental results on real datasets of knowledge hypergraphs and knowledge graphs demonstrate the superior performance of HJE compared with state-of-the-art baselines.
Zhao Li 0009, Chenxu Wang 0014, Xin Wang 0030, Jianxin Li 0001
IEEE Trans. Knowl. Data Eng.3
2023 TKGAT: Temporal Knowledge Graph Representation Learning Using Attention Network
Zhao Li 0009, Xin Wang 0030
ADMA (2)3
2023 A Cross-Region-based Framework for Supporting Car-Sharing
Rui Zhu 0003, Xuexin Zhang, Xin Wang 0030, Jiajia Li 0003, Anzhen Zhang, Chuanyu Zong
ADMA (1)3
2023 Hierarchical Label Inference Incorporating Attribute Semantics in Attributed Networks
abstract
Node attribute label inference is an important problem in attributed networks. Most existing works assume that node labels are at a single level, but in practice, the attribute labels can always be organized in a hierarchical structure according to their semantics. In this paper, we propose a novel hierarchical label inference model for attributed networks. Specifically, we propose a triple attention mechanism to extract fine-grained label semantics from three levels: hierarchical, sibling and global. Next, we propose the semantic fully-connected layer to explicitly exploit label semantics for attribute inference. We also propose semantic label propagation to enhance the interaction between the label semantics and the attributed network, and this interaction enables nodes in the attributed network to realise the proximity assumption at the label semantic level. Finally, we combine the semantic fully-connected layer with semantic label propagation for top-down hierarchical attribute inference. Extensive experiments demonstrate the superiority of our model.
Yajun Yang, Qinghua Hu, Xin Wang 0030, Hong Gao 0001
ICDM4
2023 HyConvE: A Novel Embedding Model for Knowledge Hypergraph Link Prediction with Convolutional Neural Networks
abstract
Knowledge hypergraph embedding, which projects entities and n-ary relations into a low-dimensional continuous vector space to predict missing links, remains a challenging area to be explored despite the ubiquity of n-ary relational facts in the real world. Currently, knowledge hypergraph link prediction methods are essentially simple extensions of those used in knowledge graphs, where n-ary relational facts are decomposed into different subelements. Convolutional neural networks have been shown to have remarkable information extraction capabilities in previous work on knowledge graph link prediction. In this paper, we propose a novel embedding-based knowledge hypergraph link prediction model named HyConvE, which exploits the powerful learning ability of convolutional neural networks for effective link prediction. Specifically, we employ 3D convolution to capture the deep interactions of entities and relations to efficiently extract explicit and implicit knowledge in each n-ary relational fact without compromising its translation property. In addition, appropriate relation and position-aware filters are utilized sequentially to perform two-dimensional convolution operations to capture the intrinsic patterns and position information in each n-ary relation, respectively. Extensive experimental results on real datasets of knowledge hypergraphs and knowledge graphs demonstrate the superior performance of HyConvE compared with state-of-the-art baselines.
Chenxu Wang 0014, Xin Wang 0030, Zhao Li 0009, Jianxin Li 0001
WWW2
2023 PosKHG: A Position-Aware Knowledge Hypergraph Model for Link Prediction
abstract
Abstract Link prediction in knowledge hypergraphs is essential for various knowledge-based applications, including question answering and recommendation systems. However, many current approaches simply extend binary relation methods from knowledge graphs to n-ary relations, which does not allow for capturing entity positional and role information in n-ary tuples. To address this issue, we introduce PosKHG, a method that considers entities’ positions and roles within n-ary tuples. PosKHG uses an embedding space with basis vectors to represent entities’ positional and role information through a linear combination, which allows for similar representations of entities with related roles and positions. Additionally, PosKHG employs a relation matrix to capture the compatibility of both information with all associated entities and a scoring function to measure the plausibility of tuples made up of entities with specific roles and positions. PosKHG achieves full expressiveness and high prediction efficiency. In experimental results, PosKHG achieved an average improvement of 4.1% on MRR compared to other state-of-the-art knowledge hypergraph embedding methods. Our code is available at https://anonymous.4open.science/r/PosKHG-C5B3/ .
Xin Wang 0030, Chenxu Wang 0014, Zhao Li 0009
Data Sci. Eng.2
2023 Special Issue of DASFAA 2023
abstract
We are pleased to present a special issue of Data Science and Engineering (DSE), which contains a collection of six extended papers from the DASFAA 2023 conference.The International Conference on Database Systems for Advanced Applications (DASFAA) is a well-established international conference series that provides a forum for technical presentations and discussions among database researchers, developers, and users from academia, business, and industry, which showcases state-of-the-art research and development activities in the general areas of database systems, Web information systems, and their advanced applications.The conference's long history has established the event as the premier research conference in the database area.
Xin Wang 0030, Maria Luisa Sapino, Wook-Shin Han, Yingxiao Shao, Hongzhi Yin
Data Sci. Eng.1
2023 HR-Index: An Effective Index Method for Historical Reachability Queries over Evolving Graphs
abstract
Reachability query is a fundamental problem and has been well studied on static graphs. However, in the real world, the graphs are not static but always evolving over time. In this paper, we study the problem of historical reachability query on evolving graphs. We propose a novel index, named HR-Index, which integrates complete and correct historical reachability information of the evolving graph. A historical reachability query on an evolving graph can be converted into a static reachability query on its HR-Index and thus query efficiency can be improved significantly. We also propose two optimization techniques to reduce the size of HR-Index effectively. We confirm the effectiveness and efficiency of our method through conducting extensive experiments on real-life datasets. Experimental results show both vertex and edge size of HR-Index are far smaller than that of the evolving graphs and our method has at least an order of magnitude improvement in time and space efficiency compared to the state-of-the-art method.
Yajun Yang, Xiangju Zhu, Junhu Wang, Xin Wang 0030, Hong Gao 0001
Proc. ACM Manag. Data5
2023 KGNav: A Knowledge Graph Navigational Visual Query System
abstract
Visual query is a vital technique for comprehending and analyzing knowledge graphs, which provides an effective method to lower the barrier of querying knowledge graphs for non-professional users. Nevertheless, visual query techniques for knowledge graphs and ontologies that have emerged in recent years cannot bridge the gap between global information provided by the knowledge graph schema and underlying data of knowledge graph. Thus it cannot fully exploit the global information to navigate users for querying knowledge graphs. This demonstration showcases KGNav, a Knowledge Graph Navigational visual query system. KGNav (1) redefines the minimal unit of operation to abstract the conceptual hierarchy, i.e., Knowledge Graph Schema, in the domain from the original knowledge graph in an offline semi-automatic way through the equivalence relations between these units; it also (2) provides a series of operators and an interactive GUI to capture user query intentions, guiding users to explore the Knowledge Graph Schema to achieve in-depth analysis of knowledge graphs. We will demonstrate the capability of KGNav in reducing tedious queries, enabling users to swiftly grasp the structure of the knowledge graph, and performing queries through several fundamental scenarios.
Xin Wang 0030, Zhaozhuo Li
Proc. VLDB Endow.2
2023 Incremental Graph Computation: Anchored Vertex Tracking in Dynamic Social Networks
abstract
User engagement has recently received significant attention in understanding the decay and expansion of communities in many online social networking platforms. When a user chooses to leave a social networking platform, it may cause a cascading dropping out among her friends. In many scenarios, it would be a good idea to persuade critical users to stay active in the network and prevent such a cascade because critical users can have significant influence on user engagement of the whole network. Many user engagement studies have been conducted to find a set of critical(anchored)users in the static social network. However, social networks are highly dynamic and their structures are continuously evolving. In order to fully utilize the power of anchored users in evolving networks, existing studies have to mine multiple sets of anchored users at different times, which incurs an expensive computational cost. To better understand user engagement in evolving network, we target a new research problem calledAnchored Vertex Tracking(AVT) in this paper, aiming to track the anchored users at each timestamp of evolving networks. Nonetheless, it is nontrivial to handle the AVT problem which we have proved to be NP-hard. To address the challenge, we develop a greedy algorithm inspired by the previous anchored$k$-core study in the static networks. Furthermore, we design an incremental algorithm to efficiently solve the AVT problem by utilizing the smoothness of the network structure's evolution. The extensive experiments conducted on real and synthetic datasets demonstrate the performance of our proposed algorithms and the effectiveness in solving the AVT problem.
Taotao Cai, Shuiqiao Yang, Jianxin Li 0001, Quan Z. Sheng, Jian Yang 0001, Xin Wang 0030, Wei Zhang 0098, Longxiang Gao
IEEE Trans. Knowl. Data Eng.6
2023 Rethink the Linearizability Constraints of Raft for Distributed Systems
abstract
With the deployment of modern hardware such as Flash-based SSDs and the high-speed network in distributed systems, the distributed consensus and consistency module (e.g., Raft) is typically the most time-consuming part. The reason lies in that Raft introduces some very strict constraints to ensure the linearizability. Therefore, in this paper, we rethink these constraints in-depth and find that some of them are not necessary, and can be broken to accelerate the performance significantly without breaking the linear consistency for distributed systems. An improved distributed consensus algorithm calledBUC-Raft(Breaking Unnecessary Constraints of Raft) is proposed in this paper and implemented in an industry-level distributed system. The experimental results suggest that both the write and the read performance can be accelerated significantly by BUC-Raft.
Yangyang Wang 0008, Zikai Wang 0003, Yunpeng Chai, Xin Wang 0030
IEEE Trans. Knowl. Data Eng.4
2023 A Measurement-Driven Analysis and Prediction of Content Propagation in the Device-to-Device Social Networks
abstract
In the 5 G era, data traffic has been growing rapidly. A small number of popular data files may dominate the network traffic and lead to heavy network congestion. Device-to-Device (D2D) communication can be used for caching and offloading significant data traffic. D2D social networks are instantiated paradigms of D2D communication. Existing studies maximize the performances of caching and offloading in D2D social networks by predicting potential content propagation paths. However, predicting such paths still faces many challenges, such as limitation of user spatial-temporal features, fragility of D2D social networks, and uncertainty of participants. As a solution, we first measure users' multi-dimensional features and content propagation paths to explore the distributions of D2D activities. Then we propose a D2D-LSTM model to predict complete content propagation paths hierarchically and design a prototype-user model for new participants. Experimental results demonstrate the state-of-the-art performances of D2D-LSTM. D2D-LSTM achieves at most 95% and at least 84.6% average precision in predicting terminal prototype-user class. Tree generation tests show that the generated trees have at most 64% and at least 17% similarity with ground-truth trees.
Heng Zhang 0032, Shaoyuan Huang, Xin Wang 0030, Jianxin Li 0001, Xiaofei Wang 0001, Victor C. M. Leung
IEEE Trans. Knowl. Data Eng.3
2023 MLI: A Multi-level Inference Mechanism for User Attributes in Social Networks
abstract
In the social network, each user has attributes for self-description called user attributes, which are semantically hierarchical. Attribute inference has become an essential way for social platforms to realize user classifications and targeted recommendations. Most existing approaches mainly focus on the flat inference problem neglecting the semantic hierarchy of user attributes, which will cause serious inconsistency in multi-level tasks. In this article, we propose a multi-level model MLI, where information propagation part collects attribute information by mining the global graph structure, and the attribute correction part realizes the mutual correction between different levels of attributes. Further, we put forward the concept of generalized semantic tree, a way of representing the hierarchical structure of user attributes, whose nodes are allowed to have multiple parent nodes unlike the regular tree. Both regular and generalized semantic trees are commonly used in practice, and can be handled by our model. Besides, by making the inference start from sub-networks with sufficient attribute information, we design a “Ripple” algorithm to improve the efficiency and effectiveness of our model. For evaluation purposes, we conduct extensive verification experiments on DBLP datasets. The experimental results show the superior effect of MLI, compared with the state-of-the-art methods.
Yajun Yang, Xin Wang 0030, Hong Gao 0001, Qinghua Hu
ACM Trans. Inf. Syst.3
2022 LocRDF: An Ontology-Aware Key-Value Store for Massive RDF Data
Jinghan Li, Xueyang Liu, Rong Cheng, Yiran Hu, Xin Wang 0030
WISA6
2022 Explainable Link Prediction in Knowledge Hypergraphs
abstract
Link prediction in knowledge hypergraphs has been recognized as a critical issue in various downstream tasks for knowledge-enabled applications, from question answering to recommender systems. However, most existing approaches are primarily performed in a black-box fashion, which learn low-dimensional embeddings for inference, thus cannot provide human-understandable interpretation. In this paper, we present HyperMLN, an n-ary, mixed, and explainable framework that interprets the path-reasoning process with first-order logic, which provides a knowledge-enhanced interpretable prediction framework, in which domain knowledge in the logic rules improves the performance of embedding models, while semantic information in the embedding space can optimize the weight of the logic rules in turn. To provide benchmark rule sets for explainable link prediction methods, three types of meta-logic rules in each popular dataset are mined for interpreting results. While achieving explainability, our framework also realizes an average improvement of 3.2% on [email protected] compared to the state-of-the-art knowledge hypergraph embedding method. Our code is available at https://github.com/zirui-chen/HyperMLN.
Xin Wang 0030, Chenxu Wang 0014, Jianxin Li 0001
CIKM2
2022 HET-KG: Communication-Efficient Knowledge Graph Embedding Training via Hotness-Aware Cache
abstract
With the popularization and application of Artificial Intelligence technology, knowledge graph embedding methods are widely used for a variety of machine learning tasks. However, most of the current knowledge graph embedding models are trained with a large number of parameters and high computational time complexity. This becomes a main obstacle to apply these existing models to large-scale knowledge graphs. To address this challenge, we propose HET-KG, a distributed system for training knowledge graph embedding efficiently. HET-KG can reduce the communication overheads by introducing a cache embedding table structure to maintain hot-embeddings at each worker. To improve the effectiveness of the cache mechanism, we design a prefetching algorithm and a filtering algorithm for adaptively selecting hot-embeddings, and provide two kinds of hot-embedding table construction strategies. To address the issue of inconsistency between the local cached hot-embeddings and the global embeddings, we also develop a hot-embedding synchronization algorithm for dynamically updating the cache embedding table, which can guarantee the inconsistency bounded within a given threshold. Finally, extensive experiments are conducted on three knowledge graph datasets FB15k, WN18, and Freebase-86m. The experimental results show that HET-KG achieves 3.7x and 1.1x speedup over the state-of-the-art systems PyTorch-BigGraph and DGL-KE, respectively.
Sicong Dong, Xupeng Miao, Pengkai Liu, Xin Wang 0030, Bin Cui 0001, Jianxin Li 0001
ICDE4
2022 Adaptive Lower-Level Driven Compaction to Optimize LSM-Tree Key-Value Stores
abstract
Log-structured merge (LSM) tree key-value (KV) stores have been widely deployed in many NoSQL and SQL systems, serving online big data applications such as social networking, graph processing, machine learning, etc. The batch processing of sorted data merging (i.e., compaction) in LSM-tree key-value stores improves the write efficiency, and some lazy compaction methods have been proposed to accumulate more data within a batch. However, these batched writing methods lead to significant tail latency, which is unacceptable for online processing. Aiming to optimize both latency and throughput, we propose a novel Lower-level Driven Compaction (LDC) method which breaks the limitations of the traditional upper-level driven compaction manner and triggers practical compaction actions bottom-up, with the benefits of both decreasing the compaction granularity for smaller latency and reducing write amplification for higher throughput. Furthermore, we extend LDC to Adaptive LDC (ALDC) by adding an adaptive policy to adjust the key compaction threshold to fit the changes of workloads’ features. The experimental results indicate that ALDC reduces the tail latency significantly and meanwhile achieves a much higher and stable throughput compared with existing approaches.
Yunpeng Chai, Yanfeng Chai, Xin Wang 0030, Haocheng Wei, Yangyang Wang 0008
IEEE Trans. Knowl. Data Eng.3
2021 Constructing Chinese Historical Literature Knowledge Graph Based on BERT
Qingyan Guo, Guanzhong Liu, Zijing Ji, Yuxin Shen, Xin Wang 0030
WISA7
2021 Incremental Validation of RDF Graphs
Xin Wang 0030, Baozhu Liu
WISA2
2021 An X-Architecture SMT Algorithm Based on Competitive Swarm Optimizer
Ruping Zhou, Genggeng Liu, Wenzhong Guo, Xin Wang 0030
WISA4
2021 CANCN-BERT: A Joint Pre-Trained Language Model for Classical and Modern Chinese
abstract
Pre-Trained Models (PTMs) can learn general knowledge representations and perform well in Natural Language Processing (NLP) tasks. For the Chinese language, several PTMs are developed, however, most existing methods concentrate on modern Chinese and are not ideal for processing classical Chinese due to the differences in grammars and semantics between these two forms. In this paper, in order to process two forms of Chinese uniformly, we propose a novel Classical and Modern Chinese pre-trained language model (CANCN-BERT), with the advantage of effectively processing both classical and modern Chinese, which is an extension of BERT. Form-aware pre-training tasks are elaborately designed to train our model, so as to better adapt it to classical and modern Chinese corpus. Moreover, we define a joint model, proposing dedicated optimization methods through different paths with the control of the switch mechanism. Our model merges characteristics of both classical and modern Chinese, which can adequately and efficiently enhance the representation ability for both forms. Extensive experiments show that our model outperforms baseline models on processing classical and modern Chinese and achieves significant and consistent improvements. Also, the results of ablation experiments demonstrate the effectiveness of each module.
Zijing Ji, Xin Wang 0030, Yuxin Shen, Guozheng Rao
CIKM2
2021 DataType-Aware Knowledge Graph Representation Learning in Hyperbolic Space
abstract
Knowledge Graph (KG) representation learning aims to encode both entities and relations into a continuous low-dimensional vector space. Most existing methods only concentrate on learning representations from structural triples in Euclidean space, which cannot well exploit the rich semantic information with hierarchical structure in KGs. In this paper, we propose a novel DataType-aware hyperbolic knowledge representation learning model called DT-GCN, which has the advantage of fully embedding attribute values of data types information. We refine data types into five primitive modalities, including integer, double, Boolean, temporal, and textual. For each modality, an encoder is specifically designed to learn its embedding. In addition, we define a unified space based on Euclidean, spherical, and hyperbolic space, which is a continuous curvature space that combines advantages of three different spaces. Extensive experiments on both synthetic and real-world datasets show that our model is consistently better than the state-of-the-art models. The average performance is improved by 2.19% and 3.46% than the optimal baseline model on node classification and link prediction tasks, respectively. The results of ablation experiments demonstrate the advantages of embedding data types information and leveraging the unified space.
Yuxin Shen, Zhao Li 0009, Xin Wang 0030, Jianxin Li 0001, Xiaowang Zhang
CIKM3
2021 OntoCSM: Ontology-Aware Characteristic Set Merging for RDF Type Discovery
Pengkai Liu, Shunting Cai, Baozhu Liu, Xin Wang 0030
DASFAA (1)4
2021 A Multilevel Inference Mechanism for User Attributes over Social Networks
Yajun Yang, Xin Wang 0030, Hong Gao 0001, Qinghua Hu, Dan Yin
DASFAA (2)3
2021 UniKG: A Unified Interoperable Knowledge Graph Database System
abstract
Knowledge graph currently has two main data models: RDF graph and property graph. The query language on RDF graph is SPARQL, while the query language on property graph is mainly Cypher. Different data models and query languages hinder the wider application of knowledge graphs. In this demonstration, we propose a unified interoperable knowledge graph database system, UniKG. (1) Based on the relational model, a unified storage scheme is utilized to efficiently store RDF graphs and property graphs, and support the query requirements of knowledge graphs. (2) Using the characteristicset-based method, the storage problem of untyped entities is addressed in UniKG. (3) UniKG realizes the interoperability of SPARQL and Cypher, and enables them to interchangeably operate on the same knowledge graph. (4) With a unified Web interface, users are allowed to query with two different languages over the same knowledge graph and visualize query results and explanations.
Baozhu Liu, Xin Wang 0030, Pengkai Liu, Sizhuo Li, Yunpeng Chai
ICDE2
2021 Rethink the Linearizability Constraints of Raft for Distributed Key-Value Stores
abstract
Distributed key-value stores have been widely used as NoSQL systems or the storage layer of distributed relational databases for various big data applications (e.g., social networking, graph processing, machine learning, etc.) due to their excellent scalability and adaptability. Although modern hardware such as Flash-based SSDs and the high-speed network is commonly deployed in key-value stores to promote performance, the distributed consensus and consistency module (e.g., Raft) is typically the most time-consuming part in distributed systems. The reason lies in that Raft introduces some very strict constraints to ensure the linearizability. Therefore, in this paper, we rethink these constraints in-depth and find that some of them are not necessary, and can be broken to accelerate the performance significantly without breaking the linear consistency for distributed key-value storage systems. An improved distributed consensus algorithm called KV-Raft is proposed in this paper and implemented in an industry-level distributed key-value system, i.e., TiKV. The experimental results suggest that both the write and the read performance can be accelerated significantly by KV-Raft. For example, in the typical read/write-balanced case, KV-Raft promotes the system throughput by 53.6%, and reduce the average write and read latency by 37.8% and 29.4%, respectively.
Yangyang Wang 0008, Zikai Wang 0003, Yunpeng Chai, Xin Wang 0030
ICDE4
2021 XTuning: Expert Database Tuning System Based on Reinforcement Learning
Yanfeng Chai, Jiake Ge, Yunpeng Chai, Xin Wang 0030, Boxuan Zhao
WISE (1)4
2021 OntoSP: Ontology-Based Semantic-Aware Partitioning on RDF Graphs
Sizhuo Li, Weixue Chen, Baozhu Liu, Pengkai Liu, Xin Wang 0030, Yuan-Fang Li
WISE (1)5
2021 Comparison the Performance of Classification Methods for Diagnosis of Heart Disease and Chronic Conditions
Jiarui Si, Haohan Zou, Chuanyi Huang, Huan Feng, Shuaijun Hu, Xin Wang 0030
WISE (2)9
2021 Optimal Subgraph Matching Queries over Distributed Knowledge Graphs Based on Partial Evaluation
Jiao Xing, Baozhu Liu, Jianxin Li 0001, Farhana Murtaza Choudhury, Xin Wang 0030
WISE (1)5
2021 Special Issue of APWeb‑WAIM 2020
abstract
We are pleased to present a special issue of Data Science and Engineering (DSE), which contains a collection of six extended papers from the APWeb-WAIM 2020 conference.We also include a regular submission paper in this issueAPWeb-WAIM conferences focus on research, development, and applications in relation to Web information management, including a wide range of topics, such as text analysis, graph data processing, social networks, recommender systems, information retrieval, data streams, knowledge graph, data mining and application, query processing, machine learning, database and Web applications, big data, and blockchain.
Xin Wang 0030, Bohan Li 0001, Shiyu Yang 0002
Data Sci. Eng.1
2020 SLPSO-Based X-Architecture Steiner Minimum Tree Construction
Xiaohua Chen 0003, Ruping Zhou, Genggeng Liu, Xin Wang 0030
WISA4
2020 A Text Representation Model Based on Convolutional Neural Network and Variational Auto Encoder
Canyang Guo, Genggeng Liu, Xin Wang 0030
WISA4
2020 BERT-Based Named Entity Recognition in Chinese Twenty-Four Histories
Xin Wang 0030
WISA2
2020 A Block-Level RNN Model for Resume Block Classification
abstract
Resume block classification is the most significant step in resume information extraction. However, the existing algorithms applied to resume block classification are all the general text classification algorithms, which failed to consider the contextual order of each block within a resume. In order to improve the performance of resume block classification, we propose in this paper a block-level bidirectional recurrent neural network model that makes full use of the contextual order relationship among different resume blocks. The experimental results show that the average F1-score value of our model on three 1,400 real resume datasets is 6% to 9% higher than the existing methods.
Qiqiang Xu, Ji Zhang 0001, Youwen Zhu, Bohan Li 0001, Donghai Guan, Xin Wang 0030
IEEE BigData6
2020 PDKE: An Efficient Distributed Embedding Framework for Large Knowledge Graphs
Sicong Dong, Xin Wang 0030, Lele Chai, Jianxin Li 0001, Yajun Yang
DASFAA (2)2
2020 Predicting Workplace Injuries Using Machine Learning Algorithms
abstract
Predicting workplace injury using automated techniques opens newer possibilities in evidence-based research. This paper presents our preliminary research in a PhD project in predicting workplace incidents using machine learning algorithms. The analysis on the model performance using several mainstream machine learning algorithms including random forest, k-nearest neighbor and decision tree indicated that the general performance of the decision tree model was found to be statistically higher than that of the other two algorithms.
Divya Sukumar, Ji Zhang 0001, Xiaohui Tao 0001, Xin Wang 0030, Wenbin Zhang 0002
DSAA4
2020 A Knowledge Enhanced Ensemble Learning Model for Mental Disorder Detection on Social Media
Guozheng Rao, Chengxia Peng, Li Zhang 0059, Xin Wang 0030, Zhiyong Feng 0002
KSEM (2)4
2020 Seeds Selection for Influence Maximization Based on Device-to-Device Social Knowledge by Reinforcement Learning
Xu Tong, Xiaofei Wang 0001, Jianxin Li 0001, Xin Wang 0030
KSEM (2)5
2019 The Air Quality Prediction Based on a Convolutional LSTM Network
Canyang Guo, Wenzhong Guo, Chi-Hua Chen 0002, Xin Wang 0030, Genggeng Liu
WISA4
2019 A Unified Relational Storage Scheme for RDF and Property Graphs
Pengkai Liu, Xiefan Guo, Sizhuo Li, Xin Wang 0030
WISA5
2019 COEA: An Efficient Method for Entity Alignment in Online Encyclopedias
Yimin Lv, Xin Wang 0030, Runpu Yue, Fuchuan Tang, Xue Xiang
ADMA2
2019 KG3D: An Interactive 3D Visualization Tool for Knowledge Graphs
Lin Wang 0079, Xin Wang 0030, Dianquan Li, Jianpeng Duan, Yongzhe Jia
ADMA3
2019 LDC: A Lower-Level Driven Compaction Method to Optimize SSD-Oriented Key-Value Stores
abstract
Log-structured merge (LSM) tree key-value (KV) stores have been widely deployed in many NoSQL and SQL systems, serving online big data applications such as social networking, bioinfomatics, graph processing, machine learning, etc. The batch processing of sorted data merging (i.e., compaction) in LSM-tree KV stores greatly improves the efficiency of writing, leading to good write performance and high space efficiency. Recently, some lazy compaction methods were proposed to further promote the system throughput through delaying the compaction to accumulate more data within a compaction batch. However, the batched writing manner also leads to significant tail latency, which is unacceptable for online processing, and the newly proposed lazy approaches worsen the tail latency problem. Furthermore, the unbalanced read/write performance of the widely deployed SSDs make the performance optimization harder. Aiming to optimize both the tail latency and the system throughput, in this paper, we propose a novel Lower-level Driven Compaction (LDC) method for LSM-tree KV stores. LDC breaks the limitations of the traditional upper-level driven compaction manner and triggers practical compaction actions by lower-level data. It has the benefits of both decreasing the compaction granularity effectively for smaller tail latency and reducing the write amplification of LSM-tree compaction for higher throughput. We have implemented LDC in LevelDB; the experimental results indicate that LDC can reduce the 99.9th percentile latency for 2.62 times compared with the traditional upper-level driven compaction mechanism, and achieve 56.7% ~ 72.3% higher system throughput at the same time.
Yunpeng Chai, Yanfeng Chai, Xin Wang 0030, Haocheng Wei, Ning Bao, Yushi Liang
ICDE3
2019 Structural Role Enhanced Attributed Network Embedding
Zhao Li 0009, Xin Wang 0030, Jianxin Li 0001, Qingpeng Zhang
WISE2
2019 OntoDS: An Ontology-Aware Distributed Storage Scheme for RDF Graphs
Baozhu Liu, Xin Wang 0030, Yajun Yang, Yunpeng Chai
WISE2
2019 Learning Relational Fractals for Deep Knowledge Graph Embedding in Online Social Networks
Ji Zhang 0001, Leonard Tan, Xiaohui Tao 0001, Dianwei Wang, Jia-Ching Ying, Xin Wang 0030
WISE6
2019 Efficient Subgraph Matching on Large RDF Graphs Using MapReduce
abstract
With the popularity of knowledge graphs growing rapidly, large amounts of RDF graphs have been released, which raises the need for addressing the challenge of distributed subgraph matching queries. In this paper, we propose an efficient distributed method to answer subgraph matching queries on big RDF graphs using MapReduce. In our method, query graphs are decomposed into a set of stars that utilize the semantic and structural information embedded RDF graphs as heuristics. Two optimization techniques are proposed to further improve the efficiency of our algorithms. One algorithm, called RDF property filtering , filters out invalid input data to reduce intermediate results; the other is to improve the query performance by postponing the Cartesian product operations. The extensive experiments on both synthetic and real-world datasets show that our method outperforms the close competitors S2X and SHARD by an order of magnitude on average.
Xin Wang 0030, Lele Chai, Yajun Yang, Jianxin Li 0001, Junhu Wang, Yunpeng Chai
Data Sci. Eng.1
2018 An Evolutionary Analysis of DBpedia Datasets
Weixi Li, Lele Chai, Chaozhou Yang, Xin Wang 0030
WISA4
2018 Distributed Efficient Provenance-Aware Regular Path Queries on Large RDF Graphs
Yueqi Xin, Xin Wang 0030, Di Jin 0001, Simiao Wang
DASFAA (1)2
2018 PROSE: A Plugin-Based Framework for Paraconsistent Reasoning on Semantic Web
abstract
The study of paraconsistent reasoning with ontologies is especially important for the Semantic Web since knowledge is not always perfect within it. However, classical OWL reasoners cannot support reasoning with inconsistent ontologies. In this article, the authors present a plugin-based framework called prose to provide rich paraconsistent reasoning services for OWL ontologies, whose architecture contains the three following parts: a classical OWL reasoner, a multi-valued transformer, and an OWL API connecting with them. Within the proposed framework prose, they implement different multi-valued paraconsistent reasoning in the OWL. Moreover, they select three popular classical OWL reasoners and two typical kinds of reasoning services for users. As the authors excepted, prose does exactly enable current classical OWL reasoners to tolerate inconsistency in a simple and convenient way. Finally, they evaluate the three reasoners in a united framework (prose) and, as a result, those results can amend the analysis of the three reasoners on inconsistent ontologies.
Xiaowang Zhang, Zhiyong Feng 0002, Wenrui Wu, Xin Wang 0030, Guozheng Rao
Int. J. Semantic Web Inf. Syst.4
2017 Visualization of Linked Biomedical Data Using Cluster Chart
abstract
With the continuous increasing of biomedical data, how to effectively use these large-scale data sets has become an urgent problem. It is also an essential issue to make benefit to users by consuming these biomedical data on the Semantic Web in a reasonable way. We present a visualization approach based on a tree-like layered interactive user interface, realize the queries of the relationships between targets, compounds, and diseases, and show the width of the path between the two biological entities according to their correlations. Furthermore, we design an iterative query method, which can find not only direct results of the input entity, but also extended results with some similarities of the input entity. Thus, the potential relationships among the extended results can be further investigated by biomedical scientists. Therefore, we have developed a user-friendly visualization system that can leverage the rich sets of the linked biomedical data.
Yiran Shan, Xin Wang 0030
WISA2
2016 RORS: Enhanced Rule-Based OWL Reasoning on Spark
Zhiyong Feng 0002, Xiaowang Zhang, Xin Wang 0030, Guozheng Rao
APWeb (2)4
2016 Efficient Distributed Regular Path Queries on RDF Graphs Using Partial Evaluation
abstract
We propose an efficient distributed method for answering regular path queries (RPQs) on large-scale RDF graphs using partial evaluation. In local computation, we devise a dynamic programming approach to evaluate local and partial answers of an RPQ on each computing site in parallel. In the assembly phase, an automata-based algorithm is proposed to assemble the partial answers of the RPQ into the final results. The experiments on benchmark RDF graphs show that our method outperforms the state-of-the-art message passing methods by up to an order of magnitude.
Xin Wang 0030, Junhu Wang, Xiaowang Zhang
CIKM1
2016 Context-Free Path Queries on RDF Graphs
Xiaowang Zhang, Zhiyong Feng 0002, Xin Wang 0030, Guozheng Rao, Wenrui Wu
ISWC (1)3
2016 On the statistical analysis of practical SPARQL queries
abstract
In this paper, we analyze some basic features of SPARQL queries from practical world in a statistical way. In particular, we focus on three statistic features including the occurrence frequency of triple patterns, fragments, and well-designed patterns and four semantic features including monotonicity, non-monotonicity, weak monotonicity and satisfiability. All the features contribute to characterize SPARQL queries in different dimensions. We hope that this statistical analysis would provide some useful observations for researchers and engineers who are interested in what real-word SPARQL queries look like, so that they could develop some practical heuristics for processing SPARQL queries, as well as build SPARQL query processing engines and benchmarks. In addition, our research facilitates to reduce scope of the problems by avoiding some cases that may not occur in practice.
Xingwang Han, Zhiyong Feng 0002, Xiaowang Zhang, Xin Wang 0030, Guozheng Rao
WebDB4
2015 GraSS: An Efficient Method for RDF Subgraph Matching
Xuedong Lyu, Xin Wang 0030, Yuan-Fang Li, Zhiyong Feng 0002, Junhu Wang
WISE (1)2
2014 Ontology-Based Spelling Suggestion for RDF Keyword Search
Junhu Wang, Xin Wang 0030
ER3
2014 TraPath: Fast Regular Path Query Evaluation on Large-Scale RDF Graphs
Xin Wang 0030, Guozheng Rao, Longxiang Jiang, Xuedong Lyu, Yajun Yang, Zhiyong Feng 0002
WAIM1
2013 Ontology-Based Semantic Search for Large-Scale RDF Data
Xin Wang 0030, Zhiyong Feng 0002, Longxiang Jiang
WAIM2
2012 Jingwei+: A Distributed Large-Scale RDF Data Server
Xin Wang 0030, Longxiang Jiang, Zhiyong Feng 0002, Pufeng Du
APWeb1
2008 BPEL4RBAC: An Authorisation Specification for WS-BPEL
Xin Wang 0030, Yanchun Zhang, Jian Yang 0001
WISE1