Hui Xu 0011

dblp:90/3055-11 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2024
0000-0002-4617-9814ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Distributional Learning for Network Alignment with Global Constraints
abstract
Network alignment, pairing corresponding nodes across the source and target networks, plays an important role in many data mining tasks. Extensive studies focus on learning node embeddings across different networks in a unified space. However, these methods have not taken the large structural discrepancy between aligned nodes into account and, thus, are largely confined by the deterministic representations of nodes. In this work, we propose a novel network alignment framework highlighted by distributional learning and globally optimal alignment. By modeling the uncertainty of each node by Gaussian distribution, our framework builds similarity matrices on the Wasserstein distance between distributions and applies Sinkhorn operation, which learns the globally optimal mapping in an end-to-end fashion. We show that each integrated part of the framework contributes to the overall performance. Under a variety of experimental settings, our alignment framework shows superior accuracy and efficiency to the state-of-the-art.
Hui Xu 0011, Liyao Xiang, Xiaoying Gan, Luoyi Fu, Xinbing Wang, Chenghu Zhou
ACM Trans. Knowl. Discov. Data1
2024 Open-World Graph Active Learning for Node Classification
abstract
The great power of Graph Neural Networks (GNNs) relies on a large number of labeled training data, but obtaining the labels can be costly in many cases. Graph Active Learning (GAL) is proposed to reduce such annotation costs, but the existing methods mainly focus on improving labeling efficiency with fixed classes, and are limited to handle the emergence of novel classes. We term the problem as Open-World Graph Active Learning (OWGAL) and propose a framework of the same name. The key is to recognize novel-class as well as informative nodes in a unified framework. Instead of a fully connected neural network classifier, OWGAL employs prototype learning and label propagation to assign high uncertainty scores to the targeted nodes in the representation and topology space, respectively. Weighted sampling further suppresses the impact of unimportant classes by weighing both the node and class importance. Experimental results on four large-scale datasets demonstrate that our framework achieves a substantial improvement of 5.97% to 16.57% on Macro-F1 over state-of-the-art methods.
Hui Xu 0011, Liyao Xiang, Junjie Ou, Yuting Weng, Xinbing Wang, Chenghu Zhou
ACM Trans. Knowl. Discov. Data1
2023 Temporal Knowledge Graph Reasoning with Historical Contrastive Learning
abstract
Temporal knowledge graph, serving as an effective way to store and model dynamic relations, shows promising prospects in event forecasting. However, most temporal knowledge graph reasoning methods are highly dependent on the recurrence or periodicity of events, which brings challenges to inferring future events related to entities that lack historical interaction. In fact, the current moment is often the combined effect of a small part of historical information and those unobserved underlying factors. To this end, we propose a new event forecasting model called Contrastive Event Network (CENET), based on a novel training framework of historical contrastive learning. CENET learns both the historical and non-historical dependency to distinguish the most potential entities that can best match the given query. Simultaneously, it trains representations of queries to investigate whether the current moment depends more on historical or non-historical events by launching contrastive learning. The representations further help train a binary classifier whose output is a boolean mask to indicate related entities in the search space. During the inference process, CENET employs a mask-based strategy to generate the final results. We evaluate our proposed model on five benchmark graphs. The results demonstrate that CENET significantly outperforms all existing methods in most metrics, achieving at least 8.3% relative improvement of Hits@1 over previous state-of-the-art baselines on event-based datasets.
Yi Xu 0004, Junjie Ou, Hui Xu 0011, Luoyi Fu
AAAI3
2023 Grace: Graph Self-Distillation and Completion to Mitigate Degree-Related Biases
abstract
Due to the universality of graph data, node classification shows its great importance in a wide range of real-world applications. Despite the successes of Graph Neural Networks (GNNs), GNN based methods rely heavily on rich connections and perform poorly on low-degree nodes. Since many real-world graphs follow a long-tailed distribution in node degrees, they suffer from a substantial performance bottleneck as a significant fraction of nodes is of low degree. In this paper, we point out that under-represented self-representations and low neighborhood homophily ratio of low-degree nodes are two main culprits. Based on that, we propose a novel method Grace which improves the node representation by self-distillation, and increases neighborhood homophily ratio of low-degree nodes by graph completion. To avoid error propagation of graph completion, label propagation is further leveraged. Experimental evidence has shown that our method well supports real-world graphs, and is superior in balancing degree-related bias and overall performance on node classification tasks.
Hui Xu 0011, Liyao Xiang, Femke Huang, Yuting Weng, Ruijie Xu 0005, Xinbing Wang, Chenghu Zhou
KDD1
2023 Tracing Truth and Rumor Diffusions Over Mobile Social Networks: Who are the Initiators?
abstract
With the increasing popularity of mobile devices, each user is able to conveniently acquire messages from others, and share diverse forms of information, like texts, images, or videos through online mobile apps. The full freedom of speech makes a great amount of truth (i.e., true information) and rumor (i.e., false information) propagate rapidly in a hybrid way through mobile platforms. As a huge variety of information floods pouring over us each day, identifying the authenticity of massive events becomes a necessary task to maintain the stability of Mobile Social Networks (MSNs). An important way to realize it is to trace their diffusions and make judgements according to the reliability of sources. With this regard, this paper proposes a diffusion model that characterizes the simultaneous diffusion of both truth and rumor in realistic MSNs, and makes the first attempt to figure out their respective sources. The problem of interest can be stated as: Given an outcome of cascade of both truth and rumor in MSNs, i.e., a set of nodes that might be the ignorant, the spreader of truth or rumor, or simply the silent receiver, how can we infer both truth sources and rumor sources? Different from previous sources detection works considering single type of nodes, the interplay between truth diffusions and rumor diffusions makes the conventional methods not work. To answer this question, we aim to maximize thesimilarity index, i.e., the number of nodes possessing the same states between the resulting network triggered by our estimated sources with the proposed diffusion model and the given observation network. Compared with existing techniques to trace diffusions of truth or rumor, it is much harder to find two kinds of sets at the same time, including truth sources and rumor sources, due to two primary reasons: (i) our biset optimization makes the submodularity techniques fail; (ii) our objective function is proven to be non-bisubmodular. To overcome above limitations, we first convert the objectivesimilarity indexinto a bisubmodular function by virtue of set covering. Based on this, we propose an approximation algorithm called Truth and Rumor Sources Detection (TRSD) algorithm via multiple reverse samplings with a provable$\frac{1}{4(1+\epsilon)^2}$approximation ratio. Further, a novel “time reversal” sources optimization strategy is proposed to converge the number of output sources from TRSD to a steady state. The effectiveness of our models and algorithms are empirical validated in two various datasets, from which we observe an up to 15% ofsimilarity indexgain as well as a narrowed down gap 0.6% to the ground truth.
Shan Qu, Hui Xu 0011, Luoyi Fu, Huan Long, Xinbing Wang, Guihai Chen, Chenghu Zhou
IEEE Trans. Mob. Comput.2
2021 Speedup Robust Graph Structure Learning with Low-Rank Information
abstract
Recent studies have shown that graph neural networks (GNNs) are vulnerable to unnoticeable adversarial perturbations, which largely confines their deployment in many safety-critical domains. Robust graph structure learning has been proposed to improve the GNN performance in the face of adversarial attacks. In particular, the low-rank methods are utilized to purify the perturbed graphs. However, these methods are mostly computationally expensive with O(n3) time complexity and O(n2) space complexity. We propose LRGNN, a fast and robust graph structure learning framework, which exploits the low-rank property as prior knowledge to speed up optimization. To eliminate adversarial perturbation, LRGNN decouples the adjacency matrix into a low-rank component and a sparse one, and learns by minimizing the rank of the first part while suppressing the second part. Its sparse variant is formed to reduce the memory footprint further. Experimental results on various attack settings have shown LRGNN acquires comparable robustness with the state-of-the-art much more efficiently, boasting a significant advantage on large-scale graphs.
Hui Xu 0011, Liyao Xiang, Jiahao Yu 0001, Xinbing Wang
CIKM1
2021 GAKG: A Multimodal Geoscience Academic Knowledge Graph
abstract
The research of geoscience plays a strong role in helping people gain a better understanding of the Earth. To effectively represent the knowledge (KG) from enormous geoscience research papers, knowledge graphs can be a powerful means. In the face of enormous geoscience research papers, knowledge graphs can be a powerful means to manage the relationships of data and integrate knowledge extracted from them. However, the existing geoscience KGs mainly focus on the external connection between concepts, whereas the potential abundant information contained in the internal multimodal data of the paper is largely overlooked for more fine-grained knowledge mining. To this end, we propose GAKG, a large-scale multimodal academic KG based on 1.12 million papers published in various geoscience-related journals. In addition to the bibliometrics elements, we also extracted the internal illustrations, tables, and text information of the articles, and dig out the knowledge entities of the papers and the era and spatial attributes of the articles, coupling multimodal academic data and features. Specifically, GAKG realizes knowledge entity extraction under our proposed Human-In-the-Loop framework, the novelty of which is to combine the techniques of machine reading and information retrieval with manual annotation of geoscientists in the loop. Considering the fact that literature of geoscience often contains more abundant illustrations and time scale information compared with that of other disciplines, we extract all the geographical information and era from the geoscience papers' text and illustrations, mapping papers to the atlas and chronology. Based on GAKG, we build several knowledge discovery benchmarks for finding geoscience communities and predicting potential links. GAKG and its services have been made publicly available and user-friendly.
Cheng Deng 0001, Yuting Jia, Hui Xu 0011, Luoyi Fu, Weinan Zhang 0001, Haisong Zhang, Xinbing Wang, Chenghu Zhou
CIKM3