EDBT 2026 Demo / reviewers in the wild / expert
Jiuyang Tang
dblp:69/5457
· DBLP profile ↗
35ranked-venue papers in the field
1as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 15 (1 first)Database Systems & Data Management · 14Other / Interdisciplinary · 3Big Data, Cloud & Distributed Data Systems · 2Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LACHT: An LLM-Aligned Cross-Hyperedge Tracer for Personalized Learning Path Planning in Programming
Jiuyang Tang, Yaqing Sheng, Jinzhi Liao, Xiang Zhao 0002 |
DASFAA (5) | 2 |
| 2026 | Unified Entity Matching under Scarce Supervision via Meta-Rule Induction and RetrievalabstractEntity matching is a fundamental task in a wide range of retrieval and knowledge applications, aiming to identify whether two objects correspond to the same real-world entity across heterogeneous sources. Typical variants include entity resolution (ER), entity linking (EL), and entity alignment (EA). While recent unified matchers have made progress through multi-task training with comprehensive annotations, real-world pipelines often operate under scarce supervision, where labeled data is incomplete and fails to cover the full spectrum of matching scenarios. In this regime, supervised unified models degrade substantially, and deployable compact LLMs remain unreliable: lightweight fine-tuning and in-context learning yield inconsistent behavior and can even exhibit negative effects under scenario shifts. To fill in this gap, we propose øurs, a meta-rule induction and retrieval framework for unified entity matching under scarce supervision. Instead of relying on parametric adaptation, øurs converts limited supervision into explicit natural-language rules, abstracts them into reusable meta-rules via hierarchical clustering, and retrieves the most relevant meta-rules to guide the LLM's inference for each input instance. This design improves robustness by grounding decisions on explicit and reusable evidence, instead of relying solely on implicit adaptation or prompt demonstrations. Extensive experiments show that øurs achieves state-of-the-art performance on unified entity matching under scarce supervision. Weixin Zeng, Jiuyang Tang, Xiang Zhao 0002 |
SIGIR | 3 |
| 2026 | Temporal Heterogeneous Network Representation Learning With Dynamic Influence ModelingabstractTemporal heterogeneous network representation learning is a pivotal approach for encapsulating the diversity of nodes and edges along with their temporal evolution into concise, low‐dimensional node representations. This technique has demonstrated remarkable efficacy in various network analysis and inference tasks. However, existing approaches study network evolution mainly by analyzing snapshots of temporal networks, while neglecting the intrinsic formation mechanisms of temporal heterogeneous networks. Few dynamic models delve into the intrinsic factors propelling network evolution. To fill this research gap, we introduce a novel learning framework for temporal heterogeneous network representation learning with dynamic influence modeling, denoted as THNRD. THNRD pioneers the application of the Hawkes process to temporal heterogeneous networks, utilizing the linking process of dynamic events to emulate the network’s formation mechanism, capturing the intrinsic dynamic progression of temporal heterogeneous networks. Subsequently, THNRD introduces a multilayer spatiotemporal aggregation model under a unified spatiotemporal framework, which is designed to harmoniously integrate the semantic and dynamic attributes of the networks. We also take node influence into consideration to further describe the temporal emergent phenomena. We verify the effectiveness of our proposed method via extensive experimental evaluations on real‐world datasets. The results consistently demonstrate that THNRD outperforms current state‐of‐the‐art methods. Haodan Ran, Yang Fang 0001, Xiang Zhao 0002, Jiuyang Tang, Weiming Zhang 0003 |
Int. J. Intell. Syst. | 4 |
| 2026 | INKER: Adaptive dynamic retrieval augmented generation with internal-external knowledge integration
Jiuyang Tang, Weixin Zeng, Xiang Zhao 0002 |
Inf. Process. Manag. | 2 |
| 2025 | Dynamic Graph Learning via Historical Information Perception and Multi-Granular Temporal Curriculum LearningabstractDynamic graph representation learning has emerged as a pivotal paradigm for modeling time-varying relational patterns in complex systems ranging from social networks to urban mobility. While existing methods achieve notable progress in temporal modeling, one critical challenge remains insufficiently addressed: identifying dual temporal evolution, i,e., instantaneous states and evolutionary trajectories. To address the challenge, we propose HMGNN, a novel dynamic graph learning framework that harmoniously integrates temporal dynamics modeling with stable structural representation learning, allowing adaptive pattern discovery while preserving feature consistency in evolving environments. Firstly, we propose a dynamic model that integrates a historical information perception module and a temporal aggregation module. The module converts the historical information into the model and adaptively measures the impact of the instantaneous and historical information effectively through the aggregation function. Secondly, we devise a dual-component model learning framework comprising contrastive learning and multi-granular temporal curriculum learning to holistically capture evolutionary dynamics. The contrastive learning component employs continuous-view contrastive alignment to preserve stable node feature across temporal evolution. Complementarily, our multi-granular temporal curriculum learning introduces masking mechanism to explicitly learn different time interval evolution patterns. Extensive experiments demonstrate the significant superiority of HMGNN against state-of-the-art dynamic graph learning methods in terms of all evaluation metrics. Yuehang Cao, Xiang Zhao 0002, Yang Fang 0001, Yan Pan 0003, Jiuyang Tang |
CIKM | 5 |
| 2025 | Yes is Harder than No: A Behavioral Study of Framing Effects in Large Language Models Across Downstream TasksabstractFraming effect is a well-known cognitive bias in which individuals' responses to the same underlying question vary depending on how the question is phrased. Recent studies suggest that large language models (LLMs) also exhibit framing effects, but existing work has primarily replicated psychological experiments using hand-crafted prompts, leaving their impact on practical downstream tasks underexplored. To fill in the gap, in this paper, we conduct a systematic empirical investigation into framing effects in LLMs across multiple real-world downstream tasks. We construct semantically equivalent prompts with positive and negative framings and evaluate a wide range of LLMs under these conditions. We uncover several behavioral regularities of framing effects in LLMs, among which the most notable one is a consistent response asymmetry: LLMs find answering ''yes'' harder than ''no''. That is, LLMs tend to issue affirmative responses (i.e., ''yes'') only when they are highly confident, while they incline to answer negatively (i.e., ''no'') under uncertainty. We interpret this asymmetry through the lens of Error Management Theory (EMT), which posits that rational agents adopt risk-averse strategies to minimize the more costly error. We empirically show that this behavior is partially attributable to a statistical imbalance in the frequency of positive versus negative framing cues in pretraining corpora. Furthermore, we demonstrate that the framing-induced bias in LLMs can inform prompt engineering and active in-context learning, i.e., using framing-sensitive samples as demonstrations can improve model performance. Finally, we offer a preliminary strategy to mitigate the framing effect, i.e., injecting debiasing instructions, which shows promise. In all, our work uncovers a fundamental behavioral bias in LLMs and offers practical guidance for their reliable deployment across downstream tasks. Weixin Zeng, Jiuyang Tang, Ji Wang 0002, Xiang Zhao 0002 |
CIKM | 3 |
| 2025 | PRIM: Encoding Propagation Probability and Role-Aware Representation for Influence Maximization
Niran Deng, Jiuyang Tang, Yang Fang 0001, Tianyang Shao, Jinzhi Liao, Xiang Zhao 0002 |
DASFAA (4) | 2 |
| 2025 | Dual-Prompting Based Event Anomaly Detection in Dynamic Graphs
Haodan Ran, Yang Fang 0001, Jiuyang Tang, Weiming Zhang 0003, Jinzhi Liao, Xiang Zhao 0002 |
DASFAA (3) | 3 |
| 2025 | Hyperedge Graph Contrastive Learning [Extended abstract]abstractAlthough various graph contrastive learning (GCL) techniques have been employed to generate augmented views and maximize their mutual information, current solutions only consider the pairwise relationships based on edges, neglecting the high-order information that can help generate more informative augmented views and make better contrast. To fill in this gap, we propose to leverage hyperedge to facilitate GCL, as it connects two or more nodes and can model high-order relationships among multiple nodes. More specifically, hyperedges are constructed based on the original graph. Then, we conduct node-level Page Rank based on hyperedges and hyperedge-level PageRank based on nodes to generate augmented views. As to the contrasting stage, different from existing GCL methods that simply treat the corresponding nodes of the anchor in different views as positives and overlook certain nodes strongly associated with the anchor, we build the positives and negatives based on hyperedges, where whether a node is a positive is determined by the number of hyperedges it coexists with the anchor. We compare our hyperedge GCL with state-of-the-art methods on downstream tasks, and the empirical results validate the superiority of our proposal. Further experiments on graph augmentation and graph contrastive loss also demonstrate the effectiveness of the proposed modules. Weixin Zeng, Jiuyang Tang, Xiang Zhao 0002 |
ICDE | 3 |
| 2025 | Towards Unsupervised Entity Alignment for Highly Heterogeneous Knowledge GraphsabstractHighly Heterogeneous Entity Alignment (HHEA) represents a more realistic application scenario of Entity Alignment (EA). This challenging task aims to align equivalent entities between highly heterogeneous knowledge graphs (HHKGs) with significant differences in structure, scale, and overlap. In practice, obtaining labeled data for HHEA is often difficult, necessitating research into unsupervised HHEA. This involves addressing several challenges, including the difficulty in capturing structural and semantic associations between HHKGs, the absence of explicit HHEA paradigms, and the high time and computational costs. Unfortunately, there is no solution for unsupervised HHEA. To bridge this gap, this paper formally investigates the unsupervised HHEA problem and proposes an effective unsupervised HHEA solution, AdaCoAgentEA, which addresses the challenges of unsupervised HHEA from the perspective of multi-agent collaboration. Specifically, we design an adaptive collaboration framework with three functional areas powered by multi-agent LLMs and small models, effectively eliminating dependence on labeled data while capturing structural and semantic correlations between HHKGs. Furthermore, we design a suite of optimization tools for AdaCoAgentEA, including meta-alignment mechanisms and communication protocols, which facilitate effective associations between HHKGs and provide explicit HHEA paradigms while reducing time and computational costs. Extensive experiments demonstrate that our proposed framework achieves state-of-the-art performance in both unsupervised HHEA and classic EA tasks across five datasets, rivaling fully supervised models while maintaining high efficiency and scalability. Runhao Zhao, Weixin Zeng, Jiuyang Tang, Yawen Li 0001, Guanhua Ye, Junping Du 0001, Xiang Zhao 0002 |
ICDE | 3 |
| 2025 | Dual Sequence Modeling for Knowledge TracingabstractAbstract Knowledge tracing (KT) refers to the problem of predicting a learner’s future performance based on their past performance in education. Recently, attention-based sequence modeling methods achieve impressive predictive performance. However, existing solutions merely consider one single sequence modeling method, which might fail to capture the comprehensive state of knowledge across long sequences. In this paper, we propose D ual S equence M odeling for K nowledge T racing (DSMKT). DSMKT aims to enhance the modeling of a learner’s long-term profile by collaborating two sequence modeling methods, i.e., the masked self-attention mechanism and the gated recurrent unit. To further exploit the synergy between two sequence models, we adopt the idea of online knowledge distillation and adaptively combine two branches to form a stronger teacher model, which in turn provides predictions as extra supervision for better modeling ability. Extensive experiments on four real-world benchmark datasets show that DSMKT performs excellently in predicting future learner responses. Qian Ning, Kunjia Liu, Jiuyang Tang, Shiqi Zhang 0011, Weixin Zeng, Xiang Zhao 0002 |
Data Sci. Eng. | 4 |
| 2025 | Confusing negative commonsense knowledge generation with hierarchy modeling and LLM-enhanced filtering
Yaqing Sheng, Weixin Zeng, Jiuyang Tang, Lihua Liu 0002, Xiang Zhao 0002 |
Inf. Process. Manag. | 3 |
| 2025 | Towards human-like questioning: Knowledge base question generation with bias-corrected reinforcement learning from human feedback
Runhao Zhao, Jiuyang Tang, Weixin Zeng, Yunxiao Guo, Xiang Zhao 0002 |
Inf. Process. Manag. | 2 |
| 2024 | Zero-shot Knowledge Graph Question Generation via Multi-agent LLMs and Small Models SynthesisabstractKnowledge Graph Question Generation (KGQG) is the task of generating natural language questions based on the given knowledge graph (KG). Although extensively explored in recent years, prevailing models predominantly depend on labelled data for training deep learning models or employ large parametric frameworks, e.g., Large Language Models (LLMs), which can incur significant deployment costs and pose practical implementation challenges. To address these issues, in this work, we put forward a zero-shot, multi-agent KGQG framework. This framework integrates the capabilities of LLMs with small models to facilitate cost-effective, high-quality question generation. In specific, we develop a professional editorial team architecture accompanied by two workflow optimization tools to reduce unproductive collaboration among LLMs-based agents and enhance the robustness of the system. Extensive experiments demonstrate that our proposed framework derives the new state-of-the-art performance on the zero-shot KGQG tasks, with relative gains of 20.24% and 13.57% on two KGQG datasets, respectively, which rival fully supervised state-of-the-art models. Runhao Zhao, Jiuyang Tang, Weixin Zeng, Xiang Zhao 0002 |
CIKM | 2 |
| 2024 | Matching Knowledge Graphs in Entity Embedding Spaces: An Experimental Study [Extended Abstract]abstractEntity alignment (EA) identifies equivalent entities that locate in different knowledge graphs (KGs), and has attracted growing research interests over the last few years with the advancement of KG embedding techniques. Although a pile of embedding-based EA frameworks have been developed, they mainly focus on improving the performance of entity representation learning, while largely overlook the subsequent stage that matches$KGs$in entity embedding spaces. Nevertheless, accurately matching entities based on learned entity representations is crucial to the overall alignment performance, as it coordinates individual alignment decisions and determines the global matching result. Hence, it is essential to understand how well existing solutions for matching KGs in entity embedding spaces perform on present benchmarks, as well as their strengths and weaknesses. To this end, in this article we provide a comprehensive survey and evaluation of matching algorithms for KGs in entity embedding spaces in terms of effectiveness and efficiency on both classic settings and new scenarios that better mirror real-life challenges. Based on in-depth analysis, we provide useful insights into the design trade-offs and good paradigms of existing works, and suggest promising directions for future development. Weixin Zeng, Xiang Zhao 0002, Jiuyang Tang, Xueqi Cheng 0001 |
ICDE | 4 |
| 2024 | Hyperedge Graph Contrastive LearningabstractAlthough various graph contrastive learning (GCL) techniques have been employed to generate augmented views and maximize their mutual information, current solutions only consider the pairwise relationships based on edges, neglecting the high-order information that can help generate more informative augmented views and make better contrast. To fill in this gap, we propose to leverage hyperedge to facilitate GCL, as it connects two or more nodes and can model high-order relationships among multiple nodes. More specifically, hyperedges are constructed based on the original graph. Then, we conduct node-level PageRank based on hyperedges and hyperedge-level PageRank based on nodes to generate augmented views. As to the contrasting stage, different from existing GCL methods that simply treat the corresponding nodes of the anchor in different views as positives and overlook certain nodes strongly associated with the anchor, we build the positives and negatives based on hyperedges, where whether a node is a positive is determined by the number of hyperedges it coexists with the anchor. We compare our hyperedge GCL with state-of-the-art methods on downstream tasks, and the empirical results validate the superiority of our proposal. Further experiments on graph augmentation and graph contrastive loss also demonstrate the effectiveness of the proposed modules. Weixin Zeng, Jiuyang Tang, Xiang Zhao 0002 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Knowledge Graph Completion with Fused Factual and Commonsense Information
Changsen Liu, Jiuyang Tang, Weixin Zeng, Jibing Wu, Hongbin Huang |
WISA | 2 |
| 2023 | Interpretable Fake News Detection with Graph EvidenceabstractAutomatic detection of fake news has received widespread attentions over recent years. A pile of efforts has been put forward to address the problem with high accuracy, while most of them lack convincing explanations, making it difficult to curb the continued spread of false news in real-life cases. Although some models leverage external resources to provide preliminary interpretability, such external signals are not always available. To fill in this gap, in this work, we put forward an interpretable fake news detection model IKA by making use of the historical evidence in the form of graphs. Specifically, we establish both positive and negative evidence graphs by collecting the signals from the historical news, i.e., training data. Then, given a piece of news to be detected, in addition to the common features used for detecting false news, we compare the news and evidence graphs to generate both the matching vector and the related graph evidence for explaining the prediction. We conduct extensive experiments on both Chinese and English datasets. The experiment results show that the detection accuracy of IKA exceeds the state-of-the-art approaches and IKA can provide useful explanations for the prediction results. Besides, IKA is general and can be applied on other models to improve their interpretability. Weixin Zeng, Jiuyang Tang, Xiang Zhao 0002 |
CIKM | 3 |
| 2023 | Matching Knowledge Graphs in Entity Embedding Spaces: An Experimental StudyabstractEntity alignment (EA) identifies equivalent entities that locate in different knowledge graphs (KGs), and has attracted growing research interests over the last few years with the advancement of KG embedding techniques. Although a pile of embedding-based EA frameworks have been developed, they mainly focus on improving the performance ofentity representation learning, while largely overlook the subsequent stage thatmatches KGs in entity embedding spaces. Nevertheless, accurately matching entities based on learned entity representations is crucial to the overall alignment performance, as it coordinates individual alignment decisions and determines the global matching result. Hence, it is essential to understand how well existing solutions for matching KGs in entity embedding spaces perform on present benchmarks, as well as their strengths and weaknesses. To this end, in this article we provide a comprehensive survey and evaluation of matching algorithms for KGs in entity embedding spaces in terms of effectiveness and efficiency on both classic settings and new scenarios that better mirror real-life challenges. Based on in-depth analysis, we provide useful insights into the design trade-offs and good paradigms of existing works, and suggest promising directions for future development. Weixin Zeng, Xiang Zhao 0002, Jiuyang Tang, Xueqi Cheng 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | PTAU: Prompt Tuning for Attributing Unanswerable QuestionsabstractCurrent question answering systems are insufficient when confronting real-life scenarios, as they can hardly be aware of whether a question is answerable given its context. Hence, there is a recent pursuit of unanswerability of a question and its attribution. Attribution of unanswerability requires the system to choose an appropriate cause for an unanswerable question. As the task is sophisticated for even human beings, it is expensive to acquire labeled data, which makes it a low-data regime problem. Moreover, the causes themselves are semantically abstract and complex, and the process of attribution is heavily question- and context-dependent. Thus, a capable model has to carefully appreciate the causes, and then, judiciously contrast the question with its context, in order to cast it into the right cause. In response to the challenges, we present PTAU, which refers to and implements a high-level human reading strategy such that one reads with anticipation. In specific, PTAU leverages the recent prompt-tuning paradigm, and is further enhanced with two innovatively conceived modules: 1) a cause-oriented template module that constructs continuous templates towards certain attributing class in high dimensional vector space; and 2) a semantics-aware label module that exploits label semantics through contrastive learning to render the classes distinguishable. Extensive experiments demonstrate that the proposed design better enlightens not only the attribution model, but also current question answering models, leading to superior performance. Jinzhi Liao, Xiang Zhao 0002, Jianming Zheng, Xinyi Li 0001, Jiuyang Tang |
SIGIR | 6 |
| 2022 | Toward Entity Alignment in the Open World: An Unsupervised Approach with Confidence ModelingabstractAbstract Entity alignment (EA) aims to discover the equivalent entities in different knowledge graphs (KGs). It is a pivotal step for integrating KGs to increase knowledge coverage and quality. Recent years have witnessed a rapid increase of EA frameworks. However, state-of-the-art solutions tend to rely on labeled data for model training. Additionally, they work under the closed-domain setting and cannot deal with entities that are unmatchable. To address these deficiencies, we offer an unsupervised framework that performs entity alignment in the open world. Specifically, we first mine useful features from the side information of KGs. Then, we devise an unmatchable entity prediction module to filter out unmatchable entities and produce preliminary alignment results. These preliminary results are regarded as the pseudo-labeled data and forwarded to the progressive learning framework to generate structural representations, which are integrated with the side information to provide a more comprehensive view for alignment. Finally, the progressive learning framework gradually improves the quality of structural embeddings and enhances the alignment performance. Furthermore, noticing that the pseudo-labeled data are of various qualities, we introduce the concept of confidence to measure the probability of an entity pair of being true and develop a confidence-based unsupervised EA framework . Our solutions do not require labeled data and can effectively filter out unmatchable entities. Comprehensive experimental evaluations validate the superiority of our proposals . Xiang Zhao 0002, Weixin Zeng, Jiuyang Tang, Xinyi Li 0001, Minnan Luo |
Data Sci. Eng. | 3 |
| 2022 | An Experimental Study of State-of-the-Art Entity Alignment ApproachesabstractEntity alignment (EA) finds equivalent entities that are located in different knowledge graphs (KGs), which is an essential step to enhance the quality of KGs, and hence of significance to downstream applications (e.g., question answering and recommendation). Recent years have witnessed a rapid increase of EA approaches, yet the relative performance of them remains unclear, partly due to the incomplete empirical evaluations, as well as the fact that comparisons were carried out under different settings (i.e., datasets, information used as input, etc.). In this paper, we fill in the gap by conducting a comprehensive evaluation and detailed analysis of state-of-the-art EA approaches. We first propose a general EA framework that encompasses all the current methods, and then group existing methods into three major categories. Next, we judiciously evaluate these solutions on a wide range of use cases, based on their effectiveness, efficiency and robustness. Finally, we construct a new EA dataset to mirror the real-life challenges of alignment, which were largely overlooked by existing literature. This study strives to provide a clear picture of the strengths and weaknesses of current EA approaches, so as to inspire quality follow-up research. Xiang Zhao 0002, Weixin Zeng, Jiuyang Tang, Wei Wang 0011, Fabian M. Suchanek |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | On entity alignment at scale
Weixin Zeng, Xiang Zhao 0002, Xinyi Li 0001, Jiuyang Tang, Wei Wang 0011 |
VLDB J. | 4 |
| 2021 | Reinforced Active Entity AlignmentabstractEntity alignment (EA) is the task of detecting equivalent entities from different knowledge graphs (KGs). Although this problem has been intensively studied during the last few years, the majority of the state-of-the-arts heavily rely on the labeled data, which are difficult to obtain in practice. Therefore, it calls for the study of EA with scarce supervision. To resolve this issue, we put forward a reinforced active entity alignment framework to select the entities to be manually labeled with the aim of enhancing alignment performance with minimal labeling efforts. Under this framework, we further devise an unsupervised contrastive loss to contrast different views of entity representations and augment the limited supervision signals by exploiting the vast unlabeled data. We empirically evaluate our proposal on eight popular KG pairs, and the results demonstrate that our proposed model and its components consistently boost the alignment performance under scarce supervision. Weixin Zeng, Xiang Zhao 0002, Jiuyang Tang, Changjun Fan |
CIKM | 3 |
| 2021 | Towards Entity Alignment in the Open World: An Unsupervised Approach
Weixin Zeng, Xiang Zhao 0002, Jiuyang Tang, Xinyi Li 0001, Minnan Luo |
DASFAA (1) | 3 |
| 2021 | Learning Discriminative Neural Representations for Event DetectionabstractRetrieving event instances from texts is pivotal to various natural language processing applications (e.g., automatic question answering and dialogue systems), and the first task to perform is event detection. There are two related sub-tasks therein-trigger identification and type classification, and the former is considered to play a dominant role. Nevertheless, it is notoriously challenging to predict event triggers right. To handle the task, existing work has made tremendous progress by incorporating manual features, data augmentation and neural networks, etc. Due to the scarcity of data and insufficient representation of trigger words, however, they still fail to precisely determine the spans of triggers (coined as trigger span detection problem). To address the challenge, we propose to learn discriminative neural representations (DNR) from texts. Specifically, our DNR model tackles the trigger span detection problem by exploiting two novel techniques: 1) a contrastive learning strategy, which enlarges the discrepancy between representations of words inside and outside triggers; and 2) a Mixspan strategy, which better trains the model to differentiate words nearby triggers' span boundaries. Extensive experiments on benchmarks-ACE2005 and TAC2015-demonstrate the superiority of our DNR model, leading to state-of-the-art performance. Jinzhi Liao, Xiang Zhao 0002, Xinyi Li 0001, Lingling Zhang 0005, Jiuyang Tang |
SIGIR | 5 |
| 2021 | Reinforcement Learning-based Collective Entity Alignment with Adaptive FeaturesabstractEntity alignment (EA) is the task of identifying the entities that refer to the same real-world object but are located in different knowledge graphs (KGs). For entities to be aligned, existing EA solutions treat them separately and generate alignment results as ranked lists of entities on the other side. Nevertheless, this decision-making paradigm fails to take into account the interdependence among entities. Although some recent efforts mitigate this issue by imposing the 1-to-1 constraint on the alignment process, they still cannot adequately model the underlying interdependence and the results tend to be sub-optimal. To fill in this gap, in this work, we delve into the dynamics of the decision-making process, and offer a reinforcement learning (RL)–based model to align entities collectively. Under the RL framework, we devise the coherence and exclusiveness constraints to characterize the interdependence and restrict collective alignment. Additionally, to generate more precise inputs to the RL framework, we employ representative features to capture different aspects of the similarity between entities in heterogeneous KGs, which are integrated by an adaptive feature fusion strategy. Our proposal is evaluated on both cross-lingual and mono-lingual EA benchmarks and compared against state-of-the-art solutions. The empirical results verify its effectiveness and superiority. Weixin Zeng, Xiang Zhao 0002, Jiuyang Tang, Xuemin Lin 0001, Paul Groth |
ACM Trans. Inf. Syst. | 3 |
| 2020 | Collective Entity Alignment via Adaptive FeaturesabstractEntity alignment (EA) identifies entities that refer to the same real-world object but locate in different knowledge graphs (KGs), and has been harnessed for KG construction and integration. When generating EA results, current solutions treat entities independently and fail to take into account the interdependence between entities. To fill this gap, we propose a collective EA framework. We first employ three representative features, i.e., structural, semantic and string signals, which are adapted to capture different aspects of the similarity between entities in heterogeneous KGs. In order to make collective EA decisions, we formulate EA as the classical stable matching problem, which is further effectively solved by deferred acceptance algorithm. Our proposal is evaluated on both cross-lingual and mono-lingual EA benchmarks against state-of-the-art solutions, and the empirical results verify its effectiveness and superiority. Weixin Zeng, Xiang Zhao 0002, Jiuyang Tang, Xuemin Lin 0001 |
ICDE | 3 |
| 2020 | Degree-Aware Alignment for Entities in TailabstractEntity alignment (EA) is to discover equivalent entities in knowledge graphs (KGs), which bridges heterogeneous sources of information and facilitates the integration of knowledge. Existing EA solutions mainly rely on structural information to align entities, typically through KG embedding. Nonetheless, in real-life KGs, only a few entities are densely connected to others, and the rest majority possess rather sparse neighborhood structure. We refer to the latter as long-tail entities, and observe that such phenomenon arguably limits the use of structural information for EA. Weixin Zeng, Xiang Zhao 0002, Wei Wang 0011, Jiuyang Tang |
SIGIR | 4 |
| 2014 | An Appliance-Driven Approach to Detection of Corrupted Load Curve DataabstractLoad curve data in power systems refers to users' electrical energy consumption data periodically collected with meters. It has become one of the most important assets for modern power systems. Many operational decisions are made based on the information discovered in the data. Load curve data, however, usually suffers from corruptions caused by various factors, such as data transmission errors or malfunctioning meters. To solve the problem, tremendous research efforts have been made on load curve data cleansing. Most existing approaches apply outlier detection methods from the supply side (i.e., electricity service providers), which may only have aggregated load data. In this paper, we propose to seek aid from the demand side (i.e., electricity service users). With the help of readily available knowledge on consumers' appliances, we present an appliance-driven approach to load curve data cleansing. This approach utilizes data generation rules and a Sequential Local Optimization Algorithm (SLOA) to solve the Corrupted Data Identification Problem (CDIP). We evaluate the performance of SLOA with real-world trace data and synthetic data. The results indicate that, comparing to existing load data cleansing methods, such as B-spline smoothing, our approach has an overall better performance and can effectively identify consecutive corrupted data. Experimental results also show that our method is robust in various tests. Guoming Tang, Kui Wu 0001, Jian Pei 0001, Jiuyang Tang, Jingsheng Lei |
CIKM | 4 |
| 2014 | Improving Performance of Graph Similarity Joins Using Selected Substructures
Xiang Zhao 0002, Chuan Xiao 0001, Wenjie Zhang 0001, Xuemin Lin 0001, Jiuyang Tang |
DASFAA (1) | 5 |
| 2013 | Core-based community evolution in mobile social networksabstractCommunity evolution in social networks attracts a lot of attention in recent years. Existing methods always depict the relationship of two nodes using the temporary connection. However, these temporary connections cannot be fully recognized as the real relationships when the history connections among nodes are considered. Cumulative stable contacts are proposed to depict the correlation among nodes. The whole process is divided into timestamps. At each timestamp, the community cores will be detected due to the variation of nodes and links firstly. Then, all nodes will be divided into a few of communities due to the community cores. Meanwhile, communities can be tracked through the incremental computing, which can help to recognize the evolving of community structure. Empirical studies on real-world social networks demonstrate that our proposed method can effectively detect stable community in mobile social networks. Hao Xu 0038, Weidong Xiao 0003, Daquan Tang, Jiuyang Tang, Zhenwen Wang |
IEEE BigData | 4 |
| 2013 | Provenance comparison for large-scale knowledge discoveryabstractProvenance is a record that describes entities and processes involved in producing, delivering and influencing a resource. Provenance management and reuse can enable interesting applications for knowledge discovery and analytics. One crucial component of a provenance management system is the comparison between provenances. In the era of big data, provenance management systems are in need of a scalable algorithmic solution for efficient comparison. Existing solutions to the problem have large memory footprint and require overlong system response time. In this paper, we present a new solution to threshold-based provenance comparison. We model provenance directly as graph, and propose to measure provenance similarity using provenance edit distance. Following the depth-first search paradigm, we design an algorithm PEDSim based on an encoding technique specific to provenance graphs and quantifiable heuristics. Extensive experiments on real data demonstrate the superiority of our method to other alternatives. Xiang Zhao 0002, Bin Ge 0006, Jiuyang Tang, Weidong Xiao 0003, Haichuan Shang |
IEEE BigData | 3 |
| 2008 | Advanced Star CoordinatesabstractWith the development of data collection technology, effective visualization tools are needed urgently to understand the abundant multidimensional and multivariate data and information in the science, engineering and commerce fields. Star Coordinates is a traditional multivariate data visualization technique, but there are some limitations of it. In the paper we propose the advanced star coordinates (ASC), which addresses these drawbacks. ASC uses the diameter instead of the radius as the dimension axis, projects the multidimensional information object to low dimension visual space, which is meaningful to users, and designs the dimension configuration strategy to optimize the order and angle of the dimension axes. The experiment results show that the dimension configuration strategy reduces the user operation burden greatly and helps them explore the connotative characteristics of the multidimensional information aggregation quickly and exactly. The visualization result is easily understandable and expresses the dimension distribution information effectively. Jiuyang Tang, Daquan Tang |
WAIM | 2 |
| 2005 | An Algebra for Capability Object Interoperability of Heterogeneous Data Integration Systems
Jiuyang Tang, Weiming Zhang 0003, Weidong Xiao 0003 |
APWeb | 1 |