EDBT 2026 Demo / reviewers in the wild / expert
Yongda Cai
dblp:190/5563
· DBLP profile ↗
10ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0002-3321-879XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Approximate approach for frequent itemsets mining on massive distributed data beyond computing capacity
Alladoumbaye Ngueilbaye, Sibagatullin Ratmir, Yongda Cai, Mohammad Sultan Mahmud, Xudong Sun 0004, Andrey Nechesov, Sergey Goncharov 0002, Joshua Zhexue Huang |
Expert Syst. Appl. | 3 |
| 2025 | Spectral ensemble clustering with doubly stochastic co-association matrix
Yongda Cai, Mohammad Sultan Mahmud, Jingsheng Xu, Xudong Sun 0004, Joshua Zhexue Huang |
Inf. Sci. | 1 |
| 2024 | Non-MapReduce computing for intelligent big data analysis
Xudong Sun 0004, Lingxiang Zhao, Yongda Cai, Dingming Wu 0001, Joshua Zhexue Huang |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | CDFRS: A scalable sampling approach for efficient big data analysisabstractThe sampling-based approximation method has demonstrated its potential in various domains such as machine learning, query processing, and data analysis. Most preceding sampling algorithms generate samples at the record level, making it impractical to apply them to very large datasets using a single machine. Even distributed solutions encounter efficiency issues when dealing with terabyte-scale datasets. In this paper, we introduce a scalable sampling approach named CDFRS, which can generate samples with a distribution-preserving guarantee from extensive datasets. CDFRS exhibits significantly improved speed compared to existing sampling algorithms when dealing with terabyte-scale datasets. We provide theoretical guarantees and empirical justifications, demonstrating that samples generated by the CDFRS approach maintain the distribution characteristics of the original dataset. Additionally, we propose a sample size determination algorithm, denoted as A2. Experiment results indicate that the running time of CDFRS shows at least an order of magnitude improvement over other distributed sampling methods. Notably, sampling a 10TB dataset using CDFRS only takes hundreds of seconds, while the compared method requires more than ten thousand seconds. In the context of big data analysis, including tasks such as classification and clustering, models trained with samples generated by CDFRS closely match those trained with the entire training set. Furthermore, the proposed A2 algorithm efficiently determines an appropriate sample size compared with traditional methods. Yongda Cai, Dingming Wu 0001, Xudong Sun 0004, Siyue Wu, Jingsheng Xu, Joshua Zhexue Huang |
Inf. Process. Manag. | 1 |
| 2024 | A scalable and flexible basket analysis system for big transaction data in SparkabstractBasket analysis is a prevailing technique to help retailers uncover patterns and associations of sold products in customer shopping transactions. However, as the size of transaction databases grows, the traditional basket analysis techniques and systems become less effective because of two issues in the applications of the big data age: data scalability and flexibility to adapt different application tasks. This paper proposes a scalable distributed frequent itemset mining (ScaDistFIM) algorithm for basket analysis on big transaction data to solve these two problems. ScaDistFIM is performed in two stages. The first stage uses the FP-Growth algorithm to compute the local frequent itemsets from each random subset of the distributed transaction dataset, and all random subsets are computed in parallel. The second stage uses an approximation method to aggregate all local frequent itemsets to the final approximate set of frequent itemsets where the support values of the frequent itemsets are estimated. We further elaborate on implementing the ScaDistFIM algorithm and a flexible basket analysis system using Spark SQL queries to demonstrate the system’s flexibility in real applications. The experiment results on synthetic and real-world transaction datasets demonstrate that compared to the Spark FP-Growth algorithm, the ScaDistFIM algorithm can achieve time savings of at least 90% while ensuring nearly 100% accuracy. Hence, the ScaDistFIM algorithm exhibits superior scalability. On dataset GenD with 1 billion records, the ScaDistFIM algorithm requires only 360 s to achieve 100% precision and recall. In contrast, due to memory limitations, Spark FP-Growth cannot complete the computation task. Xudong Sun 0004, Alladoumbaye Ngueilbaye, Kaijing Luo, Yongda Cai, Dingming Wu 0001, Joshua Zhexue Huang |
Inf. Process. Manag. | 4 |
| 2024 | Joint learning of fuzzy embedded clustering and non-negative spectral clustering
Wujian Ye, Jiada Wang, Yongda Cai, Chin-Chen Chang 0001 |
Multim. Tools Appl. | 3 |
| 2023 | Fast spectral clustering with self-adapted bipartite graph learning
Mingjun Zhu, Yongda Cai, Zheng Wang 0037, Feiping Nie 0001 |
Inf. Sci. | 3 |
| 2022 | Safety monitoring of machinery equipment and fault diagnosis method based on support vector machine and improved evidence theory
Xingtong Zhu, Jianbin Xiong, Yeh-Cheng Chen, Yongda Cai |
Int. J. Inf. Comput. Secur. | 4 |
| 2022 | A new method to build the adaptive k-nearest neighbors similarity graph matrix for spectral clustering
Yongda Cai, Joshua Zhexue Huang, Jianfei Yin |
Neurocomputing | 1 |
| 2020 | Fast adaptive neighbors clustering via embedded clustering
Yongda Cai, Feiping Nie 0001, Wujian Ye |
Neurocomputing | 2 |