Rustam Mussabayev

dblp:225/3699 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0001-7283-5144ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 One-pass .ѵ -means for big data clustering
Bassem Jarboui, Nenad Mladenovic, Fatmah Almathkour, Rustam Mussabayev
Pattern Recognit. Lett.4
2025 Automatic Creation of Multilingual Knowledge Graph with Large Language Models
Gulmira Tolegen, Alymzhan Toleu, Rustam Mussabayev, Alexander Krassovitskiy, Nurbakyt Zhuldyzbayuly
ACIIDS (1)3
2024 Superior Parallel Big Data Clustering Through Competitive Stochastic Sample Size Optimization in Big-Means
abstract
Abstract This paper introduces a novel K-means clustering algorithm, an advancement on the conventional Big-means methodology. The proposed method efficiently integrates parallel processing, stochastic sampling, and competitive optimization to create a scalable variant designed for big data applications. It addresses scalability and computation time challenges typically faced with traditional techniques. The algorithm adjusts sample sizes dynamically for each worker during execution, optimizing performance. Data from these sample sizes are continually analyzed, facilitating the identification of the most efficient configuration. By incorporating a competitive element among workers using different sample sizes, efficiency within the Big-means algorithm is further stimulated. In essence, the algorithm balances computational time and clustering quality by employing a stochastic, competitive sampling strategy in a parallel computing setting.
Rustam Mussabayev, Ravil Mussabayev
ACIIDS (2)1
2024 Enhancing Low-Resource NER via Knowledge Transfer from LLM
Gulmira Tolegen, Alymzhan Toleu, Rustam Mussabayev
ICCCI (1)3
2024 Distributed random swap: An efficient algorithm for minimum sum-of-squares clustering
abstract
The clustering model known as Minimum Sum-of-Squares Clustering (MSSC) is widely used, with the popular k-means algorithm serving as its local minimizer. It is well-known that solutions of k-means can result in substantial deviations from the true global optimum of MSSC. While numerous heuristics and metaheuristics have been proposed to overcome this limitation, none have gained dominant acceptance in academic literature. This is likely related to challenges such as intricate implementations and a multitude of tunable parameters. In this paper, we dispute the belief that simplifying an algorithm for MSSC inherently means sacrificing quality. We present the Distributed Random Swap (DRS-means) algorithm, which is designed to enhance clustering performance for the MSSC problem. This algorithm can be interpreted as an iterative method that refines the solution generated by the k-means algorithm during each iteration. The enhancement is achieved by selecting points based on specific probability distributions. These distributions are carefully designed to improve and speed up the exploration phase. The proposed algorithm is straightforward to implement. DRS-means offers a user-friendly solution with state-of-the-art results, making it suitable for a wide range of research fields.
Olzhas Kozbagarov, Rustam Mussabayev
Inf. Sci.2
2023 How to Use K-means for Big Data Clustering?
Rustam Mussabayev, Nenad Mladenovic, Bassem Jarboui, Ravil Mussabayev
Pattern Recognit.1
2022 Language-Independent Approach for Morphological Disambiguation
abstract
This paper presents a language-independent approach for morphological disambiguation which has been regarded as an extension of POS tagging, jointly predicting complex morphological tags. In the proposed approach, all words, roots, POS and morpheme tags are embedded into vectors, and contexts representations from surface word and morphological contexts are calculated. Then the inner products between analyses and the context’s representations are computed to perform the disambiguation. The underlying hypothesis is that the correct morphological analysis should be closer to the context in a vector space. Experimental results show that the proposed approach outperforms the existing models on seven different language datasets. Concretely, compared with the baselines of MarMot and a sophisticated neural model (Seq2Seq), the proposed approach achieves around 6% improvement in average accuracy for all languages while running about 6 and 33 times faster than MarMot and Seq2Seq, respectively.
Alymzhan Toleu, Gulmira Tolegen, Rustam Mussabayev
COLING3
2019 Neural Named Entity Recognition for Kazakh
Gulmira Tolegen, Alymzhan Toleu, Orken J. Mamyrbayev, Rustam Mussabayev
CICLing (2)4
2018 Energy-Based Centroid Identification and Cluster Propagation with Noise Detection
Alexander Krassovitskiy, Rustam Mussabayev
ICCCI (1)2