VLDB 2026 Research / reviewers in the wild / expert
Rustam Mussabayev
dblp:225/3699
· DBLP profile ↗
9ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0001-7283-5144ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | One-pass .ѵ -means for big data clustering
Bassem Jarboui, Nenad Mladenovic, Fatmah Almathkour, Rustam Mussabayev |
Pattern Recognit. Lett. | 4 |
| 2025 | Automatic Creation of Multilingual Knowledge Graph with Large Language Models
Gulmira Tolegen, Alymzhan Toleu, Rustam Mussabayev, Alexander Krassovitskiy, Nurbakyt Zhuldyzbayuly |
ACIIDS (1) | 3 |
| 2024 | Superior Parallel Big Data Clustering Through Competitive Stochastic Sample Size Optimization in Big-MeansabstractAbstract This paper introduces a novel K-means clustering algorithm, an advancement on the conventional Big-means methodology. The proposed method efficiently integrates parallel processing, stochastic sampling, and competitive optimization to create a scalable variant designed for big data applications. It addresses scalability and computation time challenges typically faced with traditional techniques. The algorithm adjusts sample sizes dynamically for each worker during execution, optimizing performance. Data from these sample sizes are continually analyzed, facilitating the identification of the most efficient configuration. By incorporating a competitive element among workers using different sample sizes, efficiency within the Big-means algorithm is further stimulated. In essence, the algorithm balances computational time and clustering quality by employing a stochastic, competitive sampling strategy in a parallel computing setting. Rustam Mussabayev, Ravil Mussabayev |
ACIIDS (2) | 1 |
| 2024 | Enhancing Low-Resource NER via Knowledge Transfer from LLM
Gulmira Tolegen, Alymzhan Toleu, Rustam Mussabayev |
ICCCI (1) | 3 |
| 2024 | Distributed random swap: An efficient algorithm for minimum sum-of-squares clusteringabstractThe clustering model known as Minimum Sum-of-Squares Clustering (MSSC) is widely used, with the popular k-means algorithm serving as its local minimizer. It is well-known that solutions of k-means can result in substantial deviations from the true global optimum of MSSC. While numerous heuristics and metaheuristics have been proposed to overcome this limitation, none have gained dominant acceptance in academic literature. This is likely related to challenges such as intricate implementations and a multitude of tunable parameters. In this paper, we dispute the belief that simplifying an algorithm for MSSC inherently means sacrificing quality. We present the Distributed Random Swap (DRS-means) algorithm, which is designed to enhance clustering performance for the MSSC problem. This algorithm can be interpreted as an iterative method that refines the solution generated by the k-means algorithm during each iteration. The enhancement is achieved by selecting points based on specific probability distributions. These distributions are carefully designed to improve and speed up the exploration phase. The proposed algorithm is straightforward to implement. DRS-means offers a user-friendly solution with state-of-the-art results, making it suitable for a wide range of research fields. Olzhas Kozbagarov, Rustam Mussabayev |
Inf. Sci. | 2 |
| 2023 | How to Use K-means for Big Data Clustering?
Rustam Mussabayev, Nenad Mladenovic, Bassem Jarboui, Ravil Mussabayev |
Pattern Recognit. | 1 |
| 2022 | Language-Independent Approach for Morphological DisambiguationabstractThis paper presents a language-independent approach for morphological disambiguation which has been regarded as an extension of POS tagging, jointly predicting complex morphological tags. In the proposed approach, all words, roots, POS and morpheme tags are embedded into vectors, and contexts representations from surface word and morphological contexts are calculated. Then the inner products between analyses and the context’s representations are computed to perform the disambiguation. The underlying hypothesis is that the correct morphological analysis should be closer to the context in a vector space. Experimental results show that the proposed approach outperforms the existing models on seven different language datasets. Concretely, compared with the baselines of MarMot and a sophisticated neural model (Seq2Seq), the proposed approach achieves around 6% improvement in average accuracy for all languages while running about 6 and 33 times faster than MarMot and Seq2Seq, respectively. Alymzhan Toleu, Gulmira Tolegen, Rustam Mussabayev |
COLING | 3 |
| 2019 | Neural Named Entity Recognition for Kazakh
Gulmira Tolegen, Alymzhan Toleu, Orken J. Mamyrbayev, Rustam Mussabayev |
CICLing (2) | 4 |
| 2018 | Energy-Based Centroid Identification and Cluster Propagation with Noise Detection
Alexander Krassovitskiy, Rustam Mussabayev |
ICCCI (1) | 2 |