Askat Kuzdeuov

dblp:240/1124 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0001-6169-8252ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Real-Time Multispectral Human Pose Estimation
abstract
Human pose estimation (HPE) is essential in human motion analysis. Nowadays, numerous RGB datasets are available to train deep learning-based HPE models. However, poor lighting and privacy issues pose challenges in the visible domain. Thermal cameras can address these issues as they are illumination invariant. However, few annotated thermal HPE datasets exist for training deep learning models. Also, HPE models trained for either the thermal or visual domain, with little exploration of cross-domain knowledge transfer. In this work, we trained the YOLO11-pose models for multispectral human pose estimation by fusing COCO and OpenThermalPose2 datasets. The results show that our models achieve high accuracy in both domains, even outperforming models specialized for each domain. The largest model, YOLO11x-pose, achieved an $AP_{pose}^{50:95}$ of 95.23% on the test set of OpenThermalPose2, establishing a new benchmark for this dataset. Also, the model achieved an $AP_{pose}^{50:95}$ of 69.89% on the COCO validation set, slightly improving the results of the original YOLO11x-pose model. We optimized the models and deployed them on an NVIDIA Jetson AGX Orin 64GB. The models in TensorRT format with half-precision floating-point achieved the best balance of speed and accuracy, making them suitable for real-time applications. We have made the pre-trained models publicly available at https://github.com/IS2AI/multispectral-motion-analysis to support research in this area.
Askat Kuzdeuov, Huseyin Atakan Varol
IECON1
2025 Multilingual Speech Command Recognition with Language Identification
abstract
Multilingual Speech Command Recognition (SCR) facilitates voice interaction in environments where multiple languages are used interchangeably, a common characteristic of multilingual regions. In such settings, SCR and language identification (LID) are handled by separate models. This separation increases inference time, energy consumption, and memory usage. To address this gap, we propose a unified multitask model that performs SCR and LID simultaneously using a shared encoder and two task-specific output heads. We tested our approach using 15 languages: English, Kazakh, Tatar, Russian, Arabic, Turkish, French, German, Catalan, Spanish, Polish, Dutch, Persian, Kinyarwanda, and Italian. We trained and compared a monolingual SCR model for each language, a multilingual SCR model without LID, a multitask multilingual SCR model with LID, and an LID-only model. The multitask model achieves an average accuracy of 90.73% for SCR and 90.99% for LID, outperforming both the multilingual SCR model without LID and the LID-only model. We have made the source code and pretrained models available at https://github.com/IS2AI/Keyword-MLP-LangID to promote research in this area.
Artur Muratov, Askat Kuzdeuov, Huseyin Atakan Varol
IECON2
2024 OpenThermalPose: An Open-Source Annotated Thermal Human Pose Dataset and Initial YOLOv8-Pose Baselines
abstract
Human pose estimation has a variety of applications in action recognition, human-robot interaction, motion capture, augmented reality, sports analytics, and healthcare. There is a substantial stream of datasets and deep learning-based models to attain robust human pose estimation within the visible domain. Nonetheless, there are certain obstacles in this domain, including insufficient illumination and privacy concerns. These issues can be addressed using thermal cameras. However, only a limited number of annotated thermal human pose datasets are available to train data-hungry deep learning models. In this regard, we introduce a novel open-source thermal human pose dataset named OpenThermalPose. The dataset contains 6,090 thermal images and 14,315 annotated human instances. The annotations include bounding boxes and 17 anatomical keypoints, following the annotation format of the MS COCO dataset. The dataset covers various fitness exercises, multiple-person activities, and outdoor walking in different locations and weather conditions. As a baseline, we trained and evaluated YOLOv8-pose models on our dataset. We have made the dataset, source code, and pretrained models publicly available at https://github.com/IS2AI/OpenThermalPose to bolster research in this area.
Askat Kuzdeuov, Darya Taratynova, Alim Tleuliyev, Huseyin Atakan Varol
FG1
2024 An Open-Source Tatar Speech Commands Dataset for IoT and Robotics Applications
abstract
Speech command recognition (SCR) is the task of detecting predefined voice commands in an audio stream. SCR is used widely in voice assistants, augmented reality, and robotics. Nowadays, Google Speech Commands (GSC) is a benchmark dataset for training and testing SCR models for English. In this work, we developed an open-source speech commands dataset for the low-resourced Tatar language. The dataset employs 35 commands in the GSC dataset and contains 3,547 utterances. To prove the efficacy of our dataset, we trained and evaluated the Keyword-MLP model on our dataset. The model achieved an accuracy of 98.28% on the test set. We optimized the model by converting it to the ONNX Runtime format and deployed it on an NVIDIA Jetson Orin NX 16GB. The optimized model showed an average inference time of 0.003 s on the device, making it suitable for real-time IoT and robotics applications. We have made the dataset, source code, and pretrained models publicly available at https://github.com/IS2AI/TatarSCR to promote the development of voice-controlled systems for the Tatar language.
Askat Kuzdeuov, Rinat Gilmullin, Bulat Khakimov, Huseyin Atakan Varol
IECON1
2022 TFW: Annotated Thermal Faces in the Wild Dataset
abstract
Face detection and subsequent localization of facial landmarks are the primary steps in many face applications. Numerous algorithms and benchmark datasets have been introduced to develop robust models for the visible domain. However, varying conditions of illumination still pose challenging problems. In this regard, thermal cameras are employed to address this problem, because they operate on longer wavelengths. However, thermal face and facial landmark detection in the wild is an open research problem because most of the existing thermal datasets were collected in controlled environments. In addition, many of them were not annotated with face bounding boxes and facial landmarks. In this work, we present a thermal face dataset with manually labeled bounding boxes and facial landmarks to address these problems. The dataset contains 9,982 images of 147 subjects collected under controlled and uncontrolled conditions. As a baseline, we trained the YOLOv5 object detection model and its adaptation for face detection, YOLO5Face, on our dataset. In addition to our test set, we evaluated the models on the external RWTH-Aachen thermal face dataset to show the efficacy of our dataset. We have made the dataset, source code, and pre-trained models publicly available at https://github.com/IS2AI/TFW to bolster research in thermal face analysis.
Askat Kuzdeuov, Dana Aubakirova, Darina Koishigarina, Huseyin Atakan Varol
IEEE Trans. Inf. Forensics Secur.1
2021 A Vaccination Simulator for COVID-19: Effective and Sterilizing Immunization Cases
abstract
In this work, we present a particle-based SEIR epidemic simulator as a tool to assess the impact of different vaccination strategies on viral propagation and to model sterilizing and effective immunization outcomes. The simulator includes modules to support contact tracing of the interactions amongst individuals and epidemiological testing of the general population. The particles are distinguished by age to represent more accurately the infection and mortality rates. The tool can be calibrated by region of interest and for different vaccination strategies to enable locality-sensitive virus mitigation policy measures and resource allocation. Moreover, the vaccination policy can be simulated based on the prioritization of certain age groups or randomly vaccinating individuals across all age groups. The results based on the experience of the province of Lecco, Italy, indicate that the simulator can evaluate vaccination strategies in a way that incorporates local circumstances of viral propagation and demographic susceptibilities. Further, the simulator accounts for modeling the distinction between sterilizing immunization, where immunized people are no longer contagious, and effective immunization, where the individuals can transmit the virus even after getting immunized. The parametric simulation results showed that the sterilizing-age-based vaccination scenario results in the least number of deaths. Furthermore, it revealed that older people should be vaccinated first to decrease the overall mortality rate. Also, the results showed that as the vaccination rate increases, the mortality rate between the scenarios shrinks.
Aknur Karabay, Askat Kuzdeuov, Shyryn Ospanova, Michael Lewis 0005, Huseyin Atakan Varol
IEEE J. Biomed. Health Informatics2
2020 A Network-Based Stochastic Epidemic Simulator: Controlling COVID-19 With Region-Specific Policies
abstract
In this work, we present an open-source stochastic epidemic simulator calibrated with extant epidemic experience of COVID-19. The simulator models a country as a network representing each node as an administrative region. The transportation connections between the nodes are modeled as the edges of this network. Each node runs a Susceptible-Exposed-Infected-Recovered (SEIR) model and population transfer between the nodes is considered using the transportation networks which allows modeling of the geographic spread of the disease. The simulator incorporates information ranging from population demographics and mobility data to health care resource capacity, by region, with interactive controls of system variables to allow dynamic and interactive modeling of events. The single-node simulator was validated using the thoroughly reported data from Lombardy, Italy. Then, the epidemic situation in Kazakhstan as of 31 May 2020 was accurately recreated. Afterward, we simulated a number of scenarios for Kazakhstan with different sets of policies. We also demonstrate the effects of region-based policies such as transportation limitations between administrative units and the application of different policies for different regions based on the epidemic intensity and geographic location. The results show that the simulator can be used to estimate outcomes of policy options to inform deliberations on governmental interdiction policies.
Askat Kuzdeuov, Daulet Baimukashev, Aknur Karabay, Bauyrzhan Ibragimov, Almas Mirzakhmetov, Mukhamet Nurpeiissov, Michael Lewis 0005, Huseyin Atakan Varol
IEEE J. Biomed. Health Informatics1