EDBT 2026 Demo / reviewers in the wild / expert
Xiaorui Jiang
dblp:89/8617
· DBLP profile ↗
12ranked-venue papers in the field
5as first author
7since 2021 · last 2025
0000-0003-4255-5445ORCID · conflict
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 7 (5 first)Database Systems & Data Management · 2Other / Interdisciplinary · 2Data Mining & Knowledge Discovery · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Boosting Bot Detection via Heterophily-Aware Representation Learning and Prototype-Guided Cluster DiscoveryabstractDetecting social media bots is essential for maintaining the security and trustworthiness of social networks. While contemporary graph-based detection methods demonstrate promising results, their practical application is limited by label reliance and poor generalization capability across diverse communities. Generative Graph Self-Supervised Learning (GSL) presents a promising paradigm to overcome these limitations, yet existing approaches predominantly follow the homophily assumption and fail to capture the global patterns in the graph, which potentially diminishes their effectiveness when facing the challenges of interaction camouflage and distributed deployment in bot detection scenarios. To this end, we propose BotHP, a generative GSL framework tailored to boost graph-based bot detectors through heterophily-aware representation learning and prototype-guided cluster discovery. Specifically, BotHP leverages a dual-encoder architecture, consisting of a graph-aware encoder to capture node commonality and a graph-agnostic encoder to preserve node uniqueness. This enables the simultaneous modeling of both homophily and heterophily, effectively countering the interaction camouflage issue. Additionally, BotHP incorporates a prototype-guided cluster discovery pretext task to model the latent global consistency of bot clusters and identify spatially dispersed yet semantically aligned bot collectives. Extensive experiments on two real-world bot detection benchmarks demonstrate that BotHP consistently boosts graph-based bot detectors, improving detection performance, alleviating label reliance, and enhancing generalization capability. Buyun He, Xiaorui Jiang, Qi Wu 0021, Hao Liu 0007, Yingguang Yang, Yong Liao 0003 |
KDD (2) | 2 |
| 2025 | Fusion-Augmented Deep Multi-view Clustering via Contrastive View ExpansionabstractMulti-view clustering has garnered increasing research attention due to its capacity to learn common semantics across views for enhanced performance. Numerous methods have been proposed to effectively learn multi-view common semantics, among which deep learning-based approaches have gradually become mainstream owing to their superior representation capabilities. However, most existing methods only fuse multi-view information at the final clustering stage, causing underutilization of multi-view data during training. To address this limitation, we propose a novel framework that treats fused multi-view features as an additional view during model training (FADE). Our autoencoder-based method first integrates deep features from multiple views to generate new representations, which are incorporated as the (V+1)-th view. Through dedicated MLPs, we extract high-level features and cluster assignments from all views. To learn common semantics, contrastive learning is first conducted between the (V+1)-th view and other views, followed by comprehensive cross-view contrastive learning to enforce multi-view consistency. Finally, we fine-tune the model by jointly optimizing high-level semantic features and cluster assignments. Comparative experiments demonstrate that our method effectively extracts common information while mitigating the impact of view-private information, outperforming multiple approaches. Zhongyi Ma, Xiaorui Jiang, Yong Liao 0003 |
MMAsia | 2 |
| 2025 | Federated Deep Incomplete Multi-View Clustering with Heterogeneity-Matching and Attention-Based ImputationabstractRecently, federated multi-view clustering has gained attention as an effective approach for exploring the clustering structure of multi-view/multi-modal data distributed across multiple clients. However, most existing methods primarily focus on sample-related client scenarios with limited research on sample-unrelated settings. These scenarios face severe data heterogeneity, while missing data further reduces available knowledge, making the problem more challenging. To address these challenges, we propose FedHMAI, a novel horizontal federated incomplete multi-view clustering method. Specifically, on the server side, we perform heterogeneity matching by leveraging Maximum Mean Discrepancy (MMD) to identify and distribute to each client the features with the least heterogeneity. Then on the client side, local cluster structure learning is conducted to alleviate data heterogeneity and reinforce cluster discriminability. Furthermore, we design an attention-based dual-level imputation mechanism that jointly leverages data- and feature-level imputation, along with matching features, to ensure imputation accuracy and consistency. Extensive experimental results validate that FedHMAI effectively overcomes the challenges of incomplete multi-view data in horizontal federated settings, consistently achieving superior clustering performance. Xiaorui Jiang, Yong Liao 0003 |
MMAsia | 2 |
| 2025 | Estimating the quality of published medical research with ChatGPT
Mike Thelwall, Xiaorui Jiang, Peter A. Bath |
Inf. Process. Manag. | 2 |
| 2025 | Is OpenAlex suitable for research quality evaluation and which citation indicator is best?abstractAbstract This article compares (1) citation analysis with OpenAlex and Scopus, testing their citation counts, document type/coverage, and subject classifications and (2) three citation‐based indicators: raw counts, (field and year) Normalized Citation Scores (NCS), and Normalized Log‐transformed Citation Scores (NLCS). Methods (1&2): The indicators calculated from 28.6 million articles were compared through 8704 correlations on two gold standards for 97,816 UK Research Excellence Framework (REF) 2021 articles. The primary gold standard is ChatGPT scores, and the secondary is the average REF2021 expert review score for the department submitting the article. Results: (1) OpenAlex provides better citation counts than Scopus, and its inclusive document classification/scope does not seem to cause substantial field normalization problems. The broadest OpenAlex classification scheme provides the best indicators. (2) Counterintuitively, raw citation counts are at least as good as nearly all field normalized indicators and better for single years, and NCS is better than NLCS. (1&2) There are substantial field differences. Thus, (1) OpenAlex is suitable for citation analysis in most fields and (2) the major citation‐based indicators seem to work counterintuitively compared to quality judgments. Field normalization seems ineffective because more cited fields tend to produce higher quality work, affecting interdisciplinary research or within‐field topic differences. Mike Thelwall, Xiaorui Jiang |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2023 | Extracting the evolutionary backbone of scientific domains: The semantic main path network analysis approach based on citation context analysisabstractAbstract Main path analysis is a popular method for extracting the scientific backbone from the citation network of a research domain. Existing approaches ignored the semantic relationships between the citing and cited publications, resulting in several adverse issues, in terms of coherence of main paths and coverage of significant studies. This paper advocated the semantic main path network analysis approach to alleviate these issues based on citation function analysis. A wide variety of SciBERT‐based deep learning models were designed for identifying citation functions. Semantic citation networks were built by either including important citations, for example, extension, motivation, usage and similarity, or excluding incidental citations like background and future work. Semantic main path network was built by merging the top‐ K main paths extracted from various time slices of semantic citation network. In addition, a three‐way framework was proposed for the quantitative evaluation of main path analysis results. Both qualitative and quantitative analysis on three research areas of computational linguistics demonstrated that, compared to semantics‐agnostic counterparts, different types of semantic main path networks provide complementary views of scientific knowledge flows. Combining them together, we obtained a more precise and comprehensive picture of domain evolution and uncover more coherent development pathways between scientific ideas. Xiaorui Jiang |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2021 | An Empirical Study of Span Modeling in Science NER
Xiaorui Jiang |
TPDL | 1 |
| 2020 | Main path analysis on cyclic citation networksabstractMain path analysis is a famous network‐based method for understanding the evolution of a scientific domain. Most existing methods have two steps, weighting citation arcs based on search path counting and exploring main paths in a greedy fashion, with the assumption that citation networks are acyclic. The only available proposal that avoids manual cycle removal is to preprint transform a cyclic network to an acyclic counterpart. Through a detailed discussion about the issues concerning this approach, especially deriving the “de‐preprinted” main paths for the original network, this article proposes an alternative solution with two‐fold contributions. Based on the argument that a publication cannot influence itself through a citation cycle, the SimSPC algorithm is proposed to weight citation arcs by counting simple search paths. A set of algorithms are further proposed for main path exploration and extraction directly from cyclic networks based on a novel data structure main path tree . The experiments on two cyclic citation networks demonstrate the usefulness of the alternative solution. In the meanwhile, experiments show that publications in strongly connected components may sit on the turning points of main path networks, which signifies the necessity of a systematic way of dealing with citation cycles. Xiaorui Jiang, Xinghao Zhu, Jingqiang Chen |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2016 | Exploiting heterogeneous scientific literature networks to combat ranking bias: Evidence from the computational linguistics areaabstractIt is important to help researchers find valuable papers from a large literature collection. To this end, many graph‐based ranking algorithms have been proposed. However, most of these algorithms suffer from the problem of ranking bias. Ranking bias hurts the usefulness of a ranking algorithm because it returns a ranking list with an undesirable time distribution. This paper is a focused study on how to alleviate ranking bias by leveraging the heterogeneous network structure of the literature collection. We propose a new graph‐based ranking algorithm, MutualRank, that integrates mutual reinforcement relationships among networks of papers, researchers, and venues to achieve a more synthetic, accurate, and less‐biased ranking than previous methods. MutualRank provides a unified model that involves both intra‐ and inter‐network information for ranking papers, researchers, and venues simultaneously. We use the ACL Anthology Network as the benchmark data set and construct the gold standard from computer linguistics course websites of well‐known universities and two well‐known textbooks. The experimental results show that MutualRank greatly outperforms the state‐of‐the‐art competitors, including PageRank, HITS, CoRank, Future Rank, and P‐Rank, in ranking papers in both improving ranking effectiveness and alleviating ranking bias. Rankings of researchers and venues by MutualRank are also quite reasonable. Xiaorui Jiang, Xiaoping Sun, Hai Zhuge |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2014 | Scaling Hop-Based Reachability Indexing for Fast Graph Pattern Query ProcessingabstractGraphs are becoming increasingly dominant in modeling real-life networked data including social and biological networks, the WWW and the Semantic Web, etc. Graph pattern queries are useful for gathering information with expressive semantics from these graph-structured data. Current methods for graph pattern query processing have performance deficiency caused by inefficiencies of the underlying reachability index and costly merge-join operations on huge amounts of tuple-formatted intermediate results. To overcome the above problems, this paper contributes in the following aspects to boost graph pattern query evaluation. First, we propose an improved hop-based reachability indexing scheme 3-Hop which gains faster reachability query evaluation, less indexing costs and better scalabilities than state-of-the-art hop-based methods. Second, we propose a two-stage node filtering algorithm based on 3-Hop to answer tree pattern queries more efficiently. Tree pattern queries serve as the underlying facility for graph pattern query evaluation. Furthermore, we use a graph representation of the intermediate results during node filtering and final results enumeration. Experiments on real-life and synthetic datasets demonstrate the effectiveness of the proposed methods. Ronghua Liang, Hai Zhuge, Xiaorui Jiang, Qiang Zeng 0002, Xiaofei He 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2012 | Towards an effective and unbiased ranking of scientific literature through mutual reinforcementabstractIt is important to help researchers find valuable scientific papers from a large literature collection containing information of authors, papers and venues. Graph-based algorithms have been proposed to rank papers based on networks formed by citation and co-author relationships. This paper proposes a new graph-based ranking framework MutualRank that integrates mutual reinforcement relationships among networks of papers, researchers and venues to achieve a more synthetic, accurate and fair ranking result than previous graph-based methods. MutualRank leverages the network structure information among papers, authors, and their venues available from a literature collection dataset and sets up a unified mutual reinforcement model that involves both intra- and inter-network information for ranking papers, authors and venues simultaneously. To evaluate, we collect a set of recommended papers from websites of graduate-level computational linguistics courses of 15 top universities as the benchmark and apply different methods to estimate paper importance. The results show that MutualRank greatly outperforms the competitors including Pag-eRank, HITS and CoRank in ranking papers as well as researchers. The experimental results also demonstrate that venues ranked by MutualRank are reasonable. Xiaorui Jiang, Xiaoping Sun, Hai Zhuge |
CIKM | 1 |
| 2012 | Adding Logical Operators to Tree Pattern Queries on Graph-Structured DataabstractAs data are increasingly modeled as graphs for expressing complex relationships, the tree pattern query on graph-structured data becomes an important type of queries in real-world applications. Most practical query languages, such as XQuery and SPARQL, support logical expressions using logical-AND/OR/NOT operators to define structural constraints of tree patterns. In this paper, (1) we propose generalized tree pattern queries (GTPQs) over graph-structured data, which fully support propositional logic of structural constraints. (2) We make a thorough study of fundamental problems including satisfiability, containment and minimization, and analyze the computational complexity and the decision procedures of these problems. (3) We propose a compact graph representation of intermediate results and a pruning approach to reduce the size of intermediate results and the number of join operations -- two factors that often impair the efficiency of traditional algorithms for evaluating tree pattern queries. (4) We present an efficient algorithm for evaluating GTPQs using 3-hop as the underlying reachability index. (5) Experiments on both real-life and synthetic data sets demonstrate the effectiveness and efficiency of our algorithm, from several times to orders of magnitude faster than state-of-the-art algorithms in terms of evaluation time, even for traditional tree pattern queries with only conjunctive operations. Qiang Zeng 0002, Xiaorui Jiang, Hai Zhuge |
Proc. VLDB Endow. | 2 |