EDBT 2026 Demo / reviewers in the wild / expert
Jiang Xie 0002
dblp:x/JXie-2
· DBLP profile ↗
14ranked-venue papers
8as first author
9since 2021 · last 2025
0000-0003-0286-3662ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Deliberation on Priors: Trustworthy Reasoning of Large Language Models on Knowledge GraphsabstractKnowledge graph-based retrieval-augmented generation seeks to mitigate hallucinations in Large Language Models (LLMs) caused by insufficient or outdated knowledge. However, existing methods often fail to fully exploit the prior knowledge embedded in knowledge graphs (KGs), particularly their structural information and explicit or implicit constraints. The former can enhance the faithfulness of LLMs' reasoning, while the latter can improve the reliability of response generations. Motivated by these, we propose a trustworthy reasoning framework, termed Deliberation over Priors (\texttt{DP}), which sufficiently utilizes the priors contained in KGs. Specifically, \texttt{DP} adopts a progressive knowledge distillation strategy that integrates structural priors into LLMs through a combination of supervised fine-tuning and Kahneman-Tversky Optimization, thereby improving the faithfulness of relation path generation. Furthermore, our framework employs a reasoning-introspection strategy, which guides LLMs to perform refined reasoning verification based on extracted constraint priors, ensuring the reliability of response generation. Extensive experiments on three benchmark datasets demonstrate that \texttt{DP} achieves new state-of-the-art performance, especially a H@1 improvement of 13% on the ComplexWebQuestions dataset, and generates highly trustworthy responses. We also conduct various analyses to verify its flexibility and practicality. Code is available at [https://github.com/mira-ai-lab/Deliberation-on-Priors](https://github.com/mira-ai-lab/Deliberation-on-Priors). Jie Ma 0001, Ning Qu, Zhitao Gao 0003, Jun Liu 0002, Hongbin Pei, Jiang Xie 0002, Lingyun Song, Pinghui Wang |
NeurIPS | 7 |
| 2025 | AW-GBGAE: An Adaptive Weighted Graph Autoencoder Based on Granular-Balls for General Data ClusteringabstractIn the current scenario, a vast amount of unlabeled high-dimensional data exhibits intrinsic relationships, making it suitable for information extraction through graph-based clustering methods. However, these datasets often lack edge structure information and contain numerous irrelevant features. To address these challenges, we propose a comprehensive solution that involves: (1) applying a feature weighting approach to manage features, (2) constructing edges based on weighted granular-balls, and (3) integrating graph convolutional networks (GCNs) with edge generation to develop an autoencoder network. Our method significantly enhances the extraction of relevant information from high-dimensional, unlabeled data, improving the overall performance and reliability of the clustering process. Extensive experimental results demonstrate that our model, AW-GBGAE, excels in clustering tasks and exhibits strong competitiveness compared to baseline models. The code is publicly available at https://github.com/xjnine/AWGBGAE. Jiang Xie 0002, Yuxin Cheng, Shuyin Xia, Chunfeng Hua, Guoyin Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | GBCT: Efficient and Adaptive Clustering via Granular-Ball Computing for Complex DataabstractTraditional clustering algorithms often focus on the most fine-grained information and achieve clustering by calculating the distance between each pair of data points or implementing other calculations based on points. This way is not inconsistent with the cognitive mechanism of "global precedence" in the human brain, resulting in those methods' bad performance in efficiency, generalization ability, and robustness. To address this problem, we propose a new clustering algorithm called granular-ball clustering via granular-ball computing. First, clustering algorithm based on granular-ball (GBCT) generates a smaller number of granular-balls to represent the original data and forms clusters according to the relationship between granular-balls, instead of the traditional point relationship. At the same time, its coarse-grained characteristics are not susceptible to noise, and the algorithm is efficient and robust; besides, as granular-balls can fit various complex data, GBCT performs much better in nonspherical datasets than other traditional clustering methods. The completely new coarse granularity representation method of GBCT and cluster formation mode can also be used to improve other traditional methods. All codes can be available at https://github.com/wylbdthxbw/GBC. Shuyin Xia, Bolun Shi, Jiang Xie 0002, Guoyin Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | W-GBC: An Adaptive Weighted Clustering Method Based on Granular-Ball StructureabstractExisting weighted clustering algorithms often heavily rely on specific parameters. Specifically, in addition to the number of clusters (k), several other parameters need to be manually tuned, which greatly limits their practical applicability. The fundamental issue lies in the fact that most weighted clustering methods derive feature weights through global iterations. To address this challenge, this paper introduces a novel weighted granular-ball structure, continually optimizing weights during the ball splitting process and restricting the calculation of local data point weights to the corresponding weighted granular-ball. We employ local iterations within this structure as an approximation to global weight calculations. This method eliminates the need for parameter tuning during the weight calculation process and incidentally addresses the “curse of dimensionality” in traditional granular-ball computing model. When applied to complex real-world datasets, this method accurately represents high-dimensional data, thereby improving clustering precision and extending the adaptability of the granular-ball computing model in high-dimensional spaces. Comprehensive experimental analysis demonstrates that our W-GBC algorithm performs well in terms of clustering results and competes strongly with baseline algorithms. The code has been released and is now available at https://github.com/xjnine/W-GBC. Jiang Xie 0002, Chunfeng Hua, Shuyin Xia, Yuxin Cheng, Guoyin Wang 0001, Xinbo Gao 0001 |
ICDE | 1 |
| 2024 | An Efficient Fuzzy Stream Clustering Method Based on Granular-Ball StructureabstractCurrent data stream clustering algorithms face low efficiency in both the online and offline phases, and struggle to address the problem of cluster boundary overlap caused by concept drift. Specifically, in the online phase, the majority of existing data stream clustering algorithms require each newly arriving sample to be scanned and inserted into the appropriate micro-clusters. In offline clustering, algorithms typically require all sample points as input. Moreover, most data stream clustering algorithms struggle to effectively deal with the problem of cluster boundary overlap caused by the concept drift. To tackle these challenges, we use a granular-ball structure for the coarse-grained representation of data stream. This structure eliminates the need for computations on all data points in both the online and offline phases. Additionally, we introduce fuzziness into the granular-ball structure to resolve the issue of cluster boundary overlap caused by the concept drift. Experimental results on both synthetic and real-world datasets demonstrate that our approach achieves efficient and accurate clustering performance when compared to existing data stream clustering algorithms. Our source code is publicly available at https://github.com/xjnine/GBFuzzyStream. Jiang Xie 0002, Minggao Dai, Shuyin Xia, Jinajinz Zhang, Guoyin Wang 0001, Xinbo Gao 0001 |
ICDE | 1 |
| 2024 | GB-DBSCAN: A fast granular-ball based DBSCAN clustering algorithm
Dongdong Cheng, Shuyin Xia, Guoyin Wang 0001, Sulan Zhang, Jiang Xie 0002 |
Inf. Sci. | 8 |
| 2024 | MGNR: A Multi-Granularity Neighbor Relationship and Its Application in KNN Classification and Clustering MethodsabstractIn the real world, data distributions often exhibit multiple granularities. However, the majority of existing neighbor-based machine-learning methods rely on manually setting a single-granularity for neighbor relationships. These methods typically handle each data point using a single-granularity approach, which severely affects their accuracy and efficiency. This paper adopts a dual-pronged approach: it constructs a multi-granularity representation of the data using the granular-ball computing model, thereby boosting the algorithm's time efficiency. It leverages the multi-granularity representation of the data to create tailored, multi-granularity neighborhood relationships for different task scenarios, resulting in improved algorithmic accuracy. The experimental results convincingly demonstrate that the proposed multi-granularity neighbor relationship effectively enhances KNN classification and clustering methods. Jiang Xie 0002, Xuexin Xiang, Shuyin Xia, Lian Jiang, Guoyin Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | K-Means Clustering With Natural Density Peaks for Discovering Arbitrary-Shaped ClustersabstractDue to simplicity, K-means has become a widely used clustering method. However, its clustering result is seriously affected by the initial centers and the allocation strategy makes it hard to identify manifold clusters. Many improved K-means are proposed to accelerate it and improve the quality of initialize cluster centers, but few researchers pay attention to the shortcoming of K-means in discovering arbitrary-shaped clusters. Using graph distance (GD) to measure the dissimilarity between objects is a good way to solve this problem, but computing the GD is time-consuming. Inspired by the idea that granular ball uses a ball to represent the local data, we select representatives from a local neighborhood, called natural density peaks (NDPs). On the basis of NDPs, we propose a novel K-means algorithm for identifying arbitrary-shaped clusters, called NDP-Kmeans. It defines neighbor-based distance between NDPs and takes advantage of the neighbor-based distance to compute the GD between NDPs. Afterward, an improved K-means with high-quality initial centers and GD is used to cluster NDPs. Finally, each remaining object is assigned according to its representative. The experimental results show that our algorithms can not only recognize spherical clusters but also manifold clusters. Therefore, NDP-Kmeans has more advantages in detecting arbitrary-shaped clusters than other excellent algorithms. Dongdong Cheng, Sulan Zhang, Shuyin Xia, Guoyin Wang 0001, Jiang Xie 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | An Efficient Spectral Clustering Algorithm Based on Granular-BallabstractIn order to solve the problem that the traditional spectral clustering algorithm is time-consuming and resource consuming when applied to large-scale data, resulting in poor clustering effect or even unable to cluster, this paper proposes a spectral clustering algorithm based on granular-ball(GBSC). The algorithm changes the construction method of the similarity matrix. Based on granular-ball, the size of the similarity matrix is greatly reduced, and the construction of the similarity matrix is more reasonable. Experimental results show that the proposed algorithm achieves better speedup ratio, less memory consumption and stronger anti noise performance while achieving similar clustering results to the traditional spectral clustering algorithm. Suppose the number of granular-balls is$m$,$n$is the number of points in the dataset, and$m< < n$, the time complexity of GBSC is$O(m^{3})$. It is proved that GBSC has good adaptability to large-scale datasets. All codes have been released athttps://github.com/xjnine/GBSC. Jiang Xie 0002, Weiyu Kong, Shuyin Xia, Guoyin Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | A new internal index based on density core for clustering validation
Jiang Xie 0002, Zhongyang Xiong, Qizhu Dai, Yu-Fang Zhang |
Inf. Sci. | 1 |
| 2020 | A local-gravitation-based method for the detection of outliers and boundary points
Jiang Xie 0002, Zhongyang Xiong, Qizhu Dai, Yu-Fang Zhang |
Knowl. Based Syst. | 1 |
| 2020 | A density-core-based clustering algorithm with local resultant force
Yu-Fang Zhang, Jiang Xie 0002, Qizhu Dai, Zhongyang Xiong, Jing-Pei Dan |
Soft Comput. | 3 |
| 2019 | A novel clustering algorithm based on the natural reverse nearest neighbor structure
Qizhu Dai, Zhongyang Xiong, Jiang Xie 0002, Yu-Fang Zhang, Jia-Xing Shang |
Inf. Syst. | 3 |
| 2018 | Density core-based clustering algorithm with dynamic scanning radius
Jiang Xie 0002, Zhongyang Xiong, Yu-Fang Zhang, Yong Feng 0002, Jie Ma 0001 |
Knowl. Based Syst. | 1 |