EDBT 2026 Demo / reviewers in the wild / expert
Huseyin Atakan Varol
dblp:36/11511
· DBLP profile ↗
33ranked-venue papers
1as first author
26since 2021 · last 2026
0000-0002-4042-425XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 12 since 2021Systems, architecture and hardware · 11 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mixed Reality and Desktop Hand Hygiene Training With Deep Learning-Based Step Recognition and Real-Time Decision SupportabstractHand hygiene (HH) is essential for preventing healthcare-associated infections, yet conventional monitoring approaches primarily capture event occurrence and provide limited insight into procedural quality, timing, and individualized feedback. To address these limitations, we present a real-time HH training and assessment framework that combines deep-learning-based WHO step recognition with a protocol-aware decision-support engine deployed on both a Desktop LED display and a Mixed Reality (MR) headset. The system supports two complementary modes: Concurrent-Feedback Coaching (CFC), which provides real-time sequence guidance and corrective prompts, and Uncued Retention Assessment (URA), which evaluates unguided execution and summarizes detected steps and errors. To support real-time deployment, we retrained and evaluated YOLOv12+MV, TimeSformer, and TSM on four heterogeneous HH datasets (PSCUH, Jurmala, METC, and Kaggle). While several YOLOv12 variants achieved strong recognition performance, the compact YOLOv12-n+MV model provided the most favorable accuracy-efficiency trade-off for deployment, achieving F1-scores of 0.99, 0.87, 0.72, and 0.58 across Kaggle, Jurmala, METC, and PSCUH, respectively, with low computational cost. This lightweight recognizer was integrated with temporal majority voting and a protocol-aware controller to support stable closed-loop interaction on Desktop and HoloLens 2. We evaluated the framework in a controlled mixed-methods study with $N=20$N=20 participants using a $2\times 2$2×2 design (Desktop vs. MR; CFC vs. URA). Desktop yielded significantly faster, more temporally stable, and less error-prone HH performance than MR, whereas CFC reduced total completion time and URA reduced weighted mistake scores, indicating a speed-accuracy trade-off. Subjective results showed higher perceived usability for Desktop than for MR and for CFC than for URA. NASA-TLX further showed a higher workload for MR than Desktop across five subscales under counterbalancing, while URA increased perceived effort relative to CFC. Overall, these findings suggest that Desktop is more suitable when efficiency, stability, and lower workload are priorities, whereas CFC and URA can be selectively used to emphasize guided acquisition or independent recall within Desktop and MR HH training workflows. Syed Muhammad Umair Arif, Tomiris Rakhimzhanova, Abylaikhan Myrzakhanov, Huseyin Atakan Varol |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | Deep Object Recognition-Based Analysis of Diverse Culinary LandscapesabstractThe rapid development of AI offers promising opportunities to promote healthier habits and address critical public health challenges. Advances in computer vision (CV) and the widespread use of smartphones and social media have sparked interest in food computation from visual data as a scalable and efficient solution. However, developing robust CV models for food object detection requires large-scale, high-quality datasets. To meet this demand, we created the Global Gastronomic Culinary Dataset (GGCD), a comprehensive resource for food object detection, building on our prior work with Central Asian Food Scenes Dataset and also annotation and preprocessing of a publicly available dataset of globally consumed food items. To evaluate the quality and utility of the dataset, we conducted comprehensive experiments using state-of-the-art object detection models. Specifically, the YOLOv11x model achieved a mAP50 of 0.711 on our testing set. To validate the generalizability of the trained model, we tested the performance on the East China University of Science and Technology Food Dataset (ECUSTFD) and our model trained on GGCD achieved a mAP of 0.683, demonstrating its effectiveness in food object detection across diverse settings. Aibota Sanatbyek, Aknur Karabay, Huseyin Atakan Varol, Mei-Yen Chan |
ICIP | 3 |
| 2025 | Multi-Modal Vision and Language Models for Real-Time Emergency ResponseabstractRecent advancements in ambient assisted living (AAL) technologies leverage machine learning (ML) and deep learning (DL) for improved emergency response and preventive care. This research introduces a multi-modal system with an advanced vision-language model (VLM) to enhance detection capabilities in AAL settings. Using DL, the system interprets scenes to generate captions, answer visual questions, and facilitate commonsense reasoning. An interactive chatbot with a large language model (LLM) and text-to-speech and speech-to-text capabilities enables real-time assessments of abnormal behavior. The system uses prompt engineering to refine anomaly detection without extensive retraining. It autonomously dispatches ambulances and generates alerts. Qualitative analysis confirms high usability among study participants, while quantitative assessments show a detection accuracy of 93.44 %, a recall rate of 95 %, and a specificity rate of 88.88 %. User interactions further enhance accuracy to 100%. This multi-modal system improves emergency recognition and response, providing caregivers with actionable insights in real time. Adil Zhiyenbayev, Rakhat Abdrakhmanov, Huseyin Atakan Varol, Adnan Yazici |
ICTAI | 3 |
| 2025 | Integrating Vision-Language Models and Multimodal Retrieval for Real-Time Emergency Response in HealthcareabstractInjuries and sudden health crises at home demand rapid medical response. We present a lightweight framework that integrates the PrismerZ vision-language model with a key-frame selection algorithm and a multimodal retrieval system to recognize emergencies from video data. The approach combines image captioning and visual question answering with efficient storage and search of embeddings to support healthcare professionals. Evaluated on the Kinetics benchmark (86.5 % image captioning accuracy, 92.5 % visual question answering accuracy) and on a self-collected dataset of emergency scenarios (85.8 % image captioning accuracy, 87.5 % visual question answering accuracy), the system operated within seconds on an embedded edge device. By uniting anomaly detection and multimedia retrieval, the framework extends human activity recognition toward actionable emergency detection in home environments. Adil Zhiyenbayev, Rakhat Abdrakhmanov, Huseyin Atakan Varol, Adnan Yazici |
ICTAI | 3 |
| 2025 | Real-Time Multispectral Human Pose EstimationabstractHuman pose estimation (HPE) is essential in human motion analysis. Nowadays, numerous RGB datasets are available to train deep learning-based HPE models. However, poor lighting and privacy issues pose challenges in the visible domain. Thermal cameras can address these issues as they are illumination invariant. However, few annotated thermal HPE datasets exist for training deep learning models. Also, HPE models trained for either the thermal or visual domain, with little exploration of cross-domain knowledge transfer. In this work, we trained the YOLO11-pose models for multispectral human pose estimation by fusing COCO and OpenThermalPose2 datasets. The results show that our models achieve high accuracy in both domains, even outperforming models specialized for each domain. The largest model, YOLO11x-pose, achieved an $AP_{pose}^{50:95}$ of 95.23% on the test set of OpenThermalPose2, establishing a new benchmark for this dataset. Also, the model achieved an $AP_{pose}^{50:95}$ of 69.89% on the COCO validation set, slightly improving the results of the original YOLO11x-pose model. We optimized the models and deployed them on an NVIDIA Jetson AGX Orin 64GB. The models in TensorRT format with half-precision floating-point achieved the best balance of speed and accuracy, making them suitable for real-time applications. We have made the pre-trained models publicly available at https://github.com/IS2AI/multispectral-motion-analysis to support research in this area. Askat Kuzdeuov, Huseyin Atakan Varol |
IECON | 2 |
| 2025 | Multilingual Speech Command Recognition with Language IdentificationabstractMultilingual Speech Command Recognition (SCR) facilitates voice interaction in environments where multiple languages are used interchangeably, a common characteristic of multilingual regions. In such settings, SCR and language identification (LID) are handled by separate models. This separation increases inference time, energy consumption, and memory usage. To address this gap, we propose a unified multitask model that performs SCR and LID simultaneously using a shared encoder and two task-specific output heads. We tested our approach using 15 languages: English, Kazakh, Tatar, Russian, Arabic, Turkish, French, German, Catalan, Spanish, Polish, Dutch, Persian, Kinyarwanda, and Italian. We trained and compared a monolingual SCR model for each language, a multilingual SCR model without LID, a multitask multilingual SCR model with LID, and an LID-only model. The multitask model achieves an average accuracy of 90.73% for SCR and 90.99% for LID, outperforming both the multilingual SCR model without LID and the LID-only model. We have made the source code and pretrained models available at https://github.com/IS2AI/Keyword-MLP-LangID to promote research in this area. Artur Muratov, Askat Kuzdeuov, Huseyin Atakan Varol |
IECON | 3 |
| 2025 | Improbability Roller-2: A Hybrid Mobile Robot with Variable Diameter Transformable WheelsabstractLocomotion on unstructured terrain poses a significant challenge for wheeled mobile robots lacking reconfigurable mechanisms. Achieving both stability and agile motion in such environments requires a hybrid approach that leverages their adaptable nature to varying surface conditions while ensuring efficient mobility. In this research, we present Improbability Roller 2, a refined iteration of our hybrid mobile robot with variable-diameter wheels, simplifying the design while improving its maneuverability and adaptability on unstructured terrain. The new compliant outer wheel structure and its folding mechanism allow for a higher wheel size change ratio. With the combination of multimode steering, which integrates both differential drive control and steering based on wheel size disparity, the robot can now optimize locomotion on diverse terrain while maintaining traction. The robot was tested across various obstacles and multiple surface conditions to validate the effectiveness of the new wheel design and the dual steering strategy. Experiments, including slope and step climbing, confined space traversal, and locomotion on loose gravel and snow, demonstrated the robot’s improved terrain adaptability, consistent traction, and control across varying surfaces. Gourav Moger, Huseyin Atakan Varol |
IROS | 2 |
| 2024 | KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech SynthesisabstractThis study focuses on the creation of the KazEmoTTS dataset, designed for emotional Kazakh text-to-speech (TTS) applications. KazEmoTTS is a collection of 54,760 audio-text pairs, with a total duration of 74.85 hours, featuring 34.23 hours delivered by a female narrator and 40.62 hours by two male narrators. The list of the emotions considered include “neutral”, “angry”, “happy”, “sad”, “scared”, and “surprised”. We also developed a TTS model trained on the KazEmoTTS dataset. Objective and subjective evaluations were employed to assess the quality of synthesized speech, yielding an MCD score within the range of 6.02 to 7.67, alongside a MOS that spanned from 3.51 to 3.57. To facilitate reproducibility and inspire further research, we have made our code, pre-trained model, and dataset accessible in our GitHub repository. Adal Abilbekov, Saida Mussakhojayeva, Rustem Yeshpanov, Huseyin Atakan Varol |
LREC/COLING | 4 |
| 2024 | KazParC: Kazakh Parallel Corpus for Machine TranslationabstractWe introduce KazParC, a parallel corpus designed for machine translation across Kazakh, English, Russian, and Turkish. The first and largest publicly available corpus of its kind, KazParC contains a collection of 371,902 parallel sentences covering different domains and developed with the assistance of human translators. Our research efforts also extend to the development of a neural machine translation model nicknamed Tilmash. Remarkably, the performance of Tilmash is on par with, and in certain instances, surpasses that of industry giants, such as Google Translate and Yandex Translate, as measured by standard evaluation metrics such as BLEU and chrF. Both KazParC and Tilmash are openly available for download under the Creative Commons Attribution 4.0 International License (CC BY 4.0) through our GitHub repository. Rustem Yeshpanov, Alina Polonskaya, Huseyin Atakan Varol |
LREC/COLING | 3 |
| 2024 | KazSAnDRA: Kazakh Sentiment Analysis Dataset of Reviews and AttitudesabstractThis paper presents KazSAnDRA, a dataset developed for Kazakh sentiment analysis that is the first and largest publicly available dataset of its kind. KazSAnDRA comprises an extensive collection of 180,064 reviews obtained from various sources and includes numerical ratings ranging from 1 to 5, providing a quantitative representation of customer attitudes. The study also pursued the automation of Kazakh sentiment classification through the development and evaluation of four machine learning models trained for both polarity classification and score classification. Experimental analysis included evaluation of the results considering both balanced and imbalanced scenarios. The most successful model attained an F1-score of 0.81 for polarity classification and 0.39 for score classification on the test sets. The dataset and fine-tuned models are open access and available for download under the Creative Commons Attribution 4.0 International License (CC BY 4.0) through our GitHub repository. Rustem Yeshpanov, Huseyin Atakan Varol |
LREC/COLING | 2 |
| 2024 | OpenThermalPose: An Open-Source Annotated Thermal Human Pose Dataset and Initial YOLOv8-Pose BaselinesabstractHuman pose estimation has a variety of applications in action recognition, human-robot interaction, motion capture, augmented reality, sports analytics, and healthcare. There is a substantial stream of datasets and deep learning-based models to attain robust human pose estimation within the visible domain. Nonetheless, there are certain obstacles in this domain, including insufficient illumination and privacy concerns. These issues can be addressed using thermal cameras. However, only a limited number of annotated thermal human pose datasets are available to train data-hungry deep learning models. In this regard, we introduce a novel open-source thermal human pose dataset named OpenThermalPose. The dataset contains 6,090 thermal images and 14,315 annotated human instances. The annotations include bounding boxes and 17 anatomical keypoints, following the annotation format of the MS COCO dataset. The dataset covers various fitness exercises, multiple-person activities, and outdoor walking in different locations and weather conditions. As a baseline, we trained and evaluated YOLOv8-pose models on our dataset. We have made the dataset, source code, and pretrained models publicly available at https://github.com/IS2AI/OpenThermalPose to bolster research in this area. Askat Kuzdeuov, Darya Taratynova, Alim Tleuliyev, Huseyin Atakan Varol |
FG | 4 |
| 2024 | Field of View Invariant Object Recognition Using Non-Linear Transformation AugmentationabstractVarious camera types, from linear perspective to wide field of view (FoV), panoramic, and hemispherical, are utilized in computer vision due to their distinct features. Images obtained from a perspective camera, which mimics the human eye are dominant in benchmark datasets such as ImageNet and MS COCO. Others obtained from wide FoV cameras are present in surveillance and person detection datasets, and autonomous driving. The proliferation of visual data acquisition through cameras with diverse FoVs has posed a significant challenge to the development of robust object recognition models. This heterogeneity necessitates the creation of an object recognition engine that can seamlessly adapt to the complexities introduced by different viewpoints, distortions, and angle variations. In this work, we leverage data augmentation to enhance the robustness of the YOLOv7 model when confronted with various camera types, such as perspective variations, distortions, and diverse viewing angles in images. We created the RMFV365 dataset, comprising 5.1 million images that encompass a broad spectrum of perspectives and distortions. This dataset serves as a valuable resource for training the model in generalized object recognition tasks. For testing purposes, we modified COHI, a benchmark fisheye testing dataset to COHI-365. The YOLOv7 model trained on RMFV365 (YOLOv7-T2) outperforms the YOLOv7 model trained on Objects365 (YOLOv7-Base), achieving a 4.8% improvement as measured by the mAP50 metric when tested on COHI-365. Also, YOLOv7-T2 achieves a close result as YOLOv7-Base when tested on the original Objects365 data. Muslim Alaran, Zarema Balgabekova, Huseyin Atakan Varol |
IECON | 3 |
| 2024 | A Tensegrity Hybrid Mobile Robot with Mechanical Impact ResistanceabstractIn this work, we present a hybrid mobile robot with legged and wheeled locomotion modes and tensegrity components. Our work uniquely integrates classical robot structures with tensegrities to create a lightweight robot capable of resisting fall impacts and moving efficiently. After describing the mechanical and embedded system implementation of the robot in detail, we present a set of experiments to demonstrate the efficacy of our robot. Specifically, we dropped the robot from different heights and conducted cost of transport measurements for wheeled and legged modes at different speeds. The findings demonstrate that the use of tensegrity structures could be beneficial in terms of impact resistance and in different challenging applications, while double-type locomotion system could serve for more efficient locomotion and power consumption. Daniil Filimonov, Altay Zhakataev, Huseyin Atakan Varol |
IECON | 3 |
| 2024 | An Open-Source Tatar Speech Commands Dataset for IoT and Robotics ApplicationsabstractSpeech command recognition (SCR) is the task of detecting predefined voice commands in an audio stream. SCR is used widely in voice assistants, augmented reality, and robotics. Nowadays, Google Speech Commands (GSC) is a benchmark dataset for training and testing SCR models for English. In this work, we developed an open-source speech commands dataset for the low-resourced Tatar language. The dataset employs 35 commands in the GSC dataset and contains 3,547 utterances. To prove the efficacy of our dataset, we trained and evaluated the Keyword-MLP model on our dataset. The model achieved an accuracy of 98.28% on the test set. We optimized the model by converting it to the ONNX Runtime format and deployed it on an NVIDIA Jetson Orin NX 16GB. The optimized model showed an average inference time of 0.003 s on the device, making it suitable for real-time IoT and robotics applications. We have made the dataset, source code, and pretrained models publicly available at https://github.com/IS2AI/TatarSCR to promote the development of voice-controlled systems for the Tatar language. Askat Kuzdeuov, Rinat Gilmullin, Bulat Khakimov, Huseyin Atakan Varol |
IECON | 4 |
| 2024 | Enhancing Human Pose Estimation Accuracy Using Synthetic DataabstractIn industrial applications, Human Pose Estimation (HPE) is crucial for enhancing both automation and human-computer interaction. This study investigates the impact of synthetic data on HPE model efficacy, particularly examining the performance of the YOLOv8 algorithm. Using Nvidia Omniverse Isaac Sim, we created a synthetic dataset called ISAAC, designed for various complex scenarios. This tool was chosen for its ability to simulate highly realistic and intricate industrial contexts with advanced physics and AI capabilities. The inclusion of this synthetic dataset significantly enhances model accuracy, evidenced by up to a 19% increase in mean Average Precision (mAP) at an Intersection over Union (IoU) of 0.5, and a 12% improvement across the 0.5-0.95 IoU range compared to traditional datasets. These results highlight the substantial advantages of synthetic data in training more accurate and robust HPE models, advocating for the integration of innovative data solutions in the field of computer vision. Rakhat Meiramov, Zarema Balgabekova, Huseyin Atakan Varol, Adnan Yazici |
IECON | 3 |
| 2023 | A Data-Centric Approach for Object Recognition in Hemispherical Camera ImagesabstractObject recognition is a machine learning problem that involves the correct classification and localization of objects in an image. Object recognition has found wide applications in Industry 4.0, surveillance, and autonomous driving. Prominent datasets such as Pascal VOC, ImageNet, and MS COCO have further spurred advances in object recognition. However, the images in these datasets were captured using ordinary perspective lenses. Unlike ordinary cameras, fisheye or hemispherical cameras have a wide FoV reaching 180°. Due to this property, they can acquire wide panoramic images, thus, capturing more objects in a scene as compared to ordinary perspective lenses. In this work, we provide the first benchmark testing dataset of real fisheye images containing commonly-occurring objects in the human environment obtained with hemispherical cameras. The dataset comprises 1,000 images containing 39 classes with a total number of 14,218 object instances. We trained the YOLOv7 model on the original MS COCO and transformed FisheyeCOCO datasets and tested the obtained models with our dataset. The results indicate that training the model on the combination of original and transformed datasets improves the performance by 2.5% compared to using these datasets separately. Zarema Balgabekova, Muslim Alaran, Huseyin Atakan Varol |
IECON | 3 |
| 2023 | Noise-Robust Automatic Speech Recognition for Industrial and Urban EnvironmentsabstractAutomatic Speech Recognition (ASR) models can achieve human parity, but their performance degrades significantly when used in noisy industrial and urban environments. In this paper, we present monolingual and multilingual ASR models, which can perform effectively even in extreme noise conditions. Specifically, we first generated a large synthetic noise-augmented dataset for Kazakh and English speech, and then used this data to train mono- and multilingual ASR models using state-of-the-art deep learning architectures. To evaluate our models, we have compared them to models trained on original data. The results showed that our models outperform the original ones in terms of word error rate (WER). For the monolingual case, the average WER of our model is 25.1%, while the original model yields 39.9% WER. For the multilingual case, we created a noise-robust customization of the Whisper model, which has a significantly improved performance on Kazakh (for both original and noisy data) and resulted on average 10% performance improvement on noisy English audios. Finally, we tested our models in an industrial setting, which has proven that our models are robust to unseen types of noise and can be used for real-life applications. Daniil Orel, Huseyin Atakan Varol |
IECON | 2 |
| 2023 | Augmented-Reality-Based Human Memory Enhancement Using Artificial IntelligenceabstractThis work presents a human memory augmentation system that uses augmented reality (AR), computer vision (CV), and artificial intelligence to replace the internal mental representation of objects in the environment with an external augmented representation. The system consists of two components: 1) an AR headset; and 2) a computing station. The AR headset runs an application that senses the indoor environment, sends data to the computing station for processing, receives the processed data, and updates the external representation of objects using a virtual 3-D object projected into the real environment in front of the user's eyes. The computing station performs CV-based indoor environment self-localization, object detection, and object-to-location binding using first-person view data received from the AR headset. We designed a behavioral study to evaluate the usability of the system. In a pilot study with 26 participants (12 females and 14 males), we investigated human performance in an experimental task that involved remembering the positions of objects in a physical space and displaying the positions of the learned objects on the 2-D map of the space. We conducted the studies under two conditions—that is, with and without using the AR system. We investigated the usability of the system, subjective workload, and performance variables under both conditions. The results showed that the AR-based augmentation of the mental representation of objects indoors reduced cognitive load and increased performance accuracy. Zhanat Makhataeva, Tolegen Akhmetov, Huseyin Atakan Varol |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2023 | An Augmented Reality-Based Warning System for Enhanced Safety in Industrial SettingsabstractOperator inattention can lead to accidents in industry, resulting in injury, death, and financial loss. To prevent this, we present a warning system that leverages artificial intelligence and augmented reality. In order to compare our warning system with the business-as-usual (safety mat) and no warning conditions, we conducted experiments with nine participants using quantitative measurements. Additionally, augmented reality warning types (auditory, visual, and audiovisual) were compared in an experiment with 30 participants. Several hypotheses were tested using subjective questionnaires: participants accept the warning system as beneficial overall (H1), participants prefer auditory warnings to visual warnings (H2), and the audiovisual system would be the most preferred warning system (H3). Participants' acceptance of the system was high, and they preferred audiovisual warnings to the other types of warnings. Participants ranked the visual warning as preferable to the auditory warning. The results can be utilized to design warning systems to reduce industrial accidents. Tolegen Akhmetov, Huseyin Atakan Varol |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | KSC2: An Industrial-Scale Open-Source Kazakh Speech CorpusabstractWe present the first industrial-scale open-source Kazakh speech corpus for automatic speech recognition research and development. Our corpus subsumes two previously presented corpora: 1) Kazakh speech corpus (KSC) and 2) Kazakh text-to-speech 2 (KazakhTTS2). We also provide additional data from other sources, including television news, television and radio programs, parliament speeches, and podcasts. Our corpus, which we have named KSC2, contains over a thousand hours of high-quality transcribed data, which is triple the size of KSC. KSC2 was manually transcribed with the help of native Kazakh speakers and validated via preliminary speech recognition experiments on various evaluation sets. Moreover, it contains utterances with Kazakh-Russian code-switching, a conversational practice common among Kazakh speakers. We believe that our corpus will facilitate speech processing research for Kazakh, which is widely considered an under-resourced language. To ensure the reproducibility of experiments, we share the KSC2 corpus, training recipes, and pretrained models. Saida Mussakhojayeva, Yerbolat Khassanov, Huseyin Atakan Varol |
INTERSPEECH | 3 |
| 2022 | KazakhTTS2: Extending the Open-Source Kazakh TTS Corpus With More Data, Speakers, and TopicsabstractWe present an expanded version of our previously released Kazakh text-to-speech (KazakhTTS) synthesis corpus. In the new KazakhTTS2 corpus, the overall size has increased from 93 hours to 271 hours, the number of speakers has risen from two to five (three females and two males), and the topic coverage has been diversified with the help of new sources, including a book and Wikipedia articles. This corpus is necessary for building high-quality TTS systems for Kazakh, a Central Asian agglutinative language from the Turkic family, which presents several linguistic challenges. We describe the corpus construction process and provide the details of the training and evaluation procedures for the TTS system. Our experimental results indicate that the constructed corpus is sufficient to build robust TTS models for real-world applications, with a subjective mean opinion score ranging from 3.6 to 4.2 for all the five speakers. We believe that our corpus will facilitate speech and language research for Kazakh and other Turkic languages, which are widely considered to be low-resource due to the limited availability of free linguistic data. The constructed corpus, code, and pretrained models are publicly available in our GitHub repository. Saida Mussakhojayeva, Yerbolat Khassanov, Huseyin Atakan Varol |
LREC | 3 |
| 2022 | KazNERD: Kazakh Named Entity Recognition DatasetabstractWe present the development of a dataset for Kazakh named entity recognition. The dataset was built as there is a clear need for publicly available annotated corpora in Kazakh, as well as annotation guidelines containing straightforward—but rigorous—rules and examples. The dataset annotation, based on the IOB2 scheme, was carried out on television news text by two native Kazakh speakers under the supervision of the first author. The resulting dataset contains 112,702 sentences and 136,333 annotations for 25 entity classes. State-of-the-art machine learning models to automatise Kazakh named entity recognition were also built, with the best-performing model achieving an exact match F1-score of 97.22% on the test set. The annotated dataset, guidelines, and codes used to train the models are freely available for download under the CC BY 4.0 licence from https://github.com/IS2AI/KazNERD. Rustem Yeshpanov, Yerbolat Khassanov, Huseyin Atakan Varol |
LREC | 3 |
| 2022 | TFW: Annotated Thermal Faces in the Wild DatasetabstractFace detection and subsequent localization of facial landmarks are the primary steps in many face applications. Numerous algorithms and benchmark datasets have been introduced to develop robust models for the visible domain. However, varying conditions of illumination still pose challenging problems. In this regard, thermal cameras are employed to address this problem, because they operate on longer wavelengths. However, thermal face and facial landmark detection in the wild is an open research problem because most of the existing thermal datasets were collected in controlled environments. In addition, many of them were not annotated with face bounding boxes and facial landmarks. In this work, we present a thermal face dataset with manually labeled bounding boxes and facial landmarks to address these problems. The dataset contains 9,982 images of 147 subjects collected under controlled and uncontrolled conditions. As a baseline, we trained the YOLOv5 object detection model and its adaptation for face detection, YOLO5Face, on our dataset. In addition to our test set, we evaluated the models on the external RWTH-Aachen thermal face dataset to show the efficacy of our dataset. We have made the dataset, source code, and pre-trained models publicly available at https://github.com/IS2AI/TFW to bolster research in thermal face analysis. Askat Kuzdeuov, Dana Aubakirova, Darina Koishigarina, Huseyin Atakan Varol |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2021 | A Crowdsourced Open-Source Kazakh Speech Corpus and Initial Speech Recognition BaselineabstractYerbolat Khassanov, Saida Mussakhojayeva, Almas Mirzakhmetov, Alen Adiyev, Mukhamet Nurpeiissov, Huseyin Atakan Varol. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Yerbolat Khassanov, Saida Mussakhojayeva, Almas Mirzakhmetov, Alen Adiyev, Mukhamet Nurpeiissov, Huseyin Atakan Varol |
EACL | 6 |
| 2021 | KazakhTTS: An Open-Source Kazakh Text-to-Speech Synthesis DatasetabstractThis paper introduces a high-quality open-source speech synthesis dataset for Kazakh, a low-resource language spoken by over 13 million people worldwide. The dataset consists of about 93 hours of transcribed audio recordings spoken by two professional speakers (female and male). It is the first publicly available large-scale dataset developed to promote Kazakh text-to-speech (TTS) applications in both academia and industry. In this paper, we share our experience by describing the dataset development procedures and faced challenges, and discuss important future directions. To demonstrate the reliability of our dataset, we built baseline end-to-end TTS models and evaluated them using the subjective mean opinion score (MOS) measure. Evaluation results show that the best TTS models trained on our dataset achieve MOS above 4 for both speakers, which makes them applicable for practical use. The dataset, training recipe, and pretrained TTS models are freely available. Saida Mussakhojayeva, Aigerim Janaliyeva, Almas Mirzakhmetov, Yerbolat Khassanov, Huseyin Atakan Varol |
Interspeech | 5 |
| 2021 | A Vaccination Simulator for COVID-19: Effective and Sterilizing Immunization CasesabstractIn this work, we present a particle-based SEIR epidemic simulator as a tool to assess the impact of different vaccination strategies on viral propagation and to model sterilizing and effective immunization outcomes. The simulator includes modules to support contact tracing of the interactions amongst individuals and epidemiological testing of the general population. The particles are distinguished by age to represent more accurately the infection and mortality rates. The tool can be calibrated by region of interest and for different vaccination strategies to enable locality-sensitive virus mitigation policy measures and resource allocation. Moreover, the vaccination policy can be simulated based on the prioritization of certain age groups or randomly vaccinating individuals across all age groups. The results based on the experience of the province of Lecco, Italy, indicate that the simulator can evaluate vaccination strategies in a way that incorporates local circumstances of viral propagation and demographic susceptibilities. Further, the simulator accounts for modeling the distinction between sterilizing immunization, where immunized people are no longer contagious, and effective immunization, where the individuals can transmit the virus even after getting immunized. The parametric simulation results showed that the sterilizing-age-based vaccination scenario results in the least number of deaths. Furthermore, it revealed that older people should be vaccinated first to decrease the overall mortality rate. Also, the results showed that as the vaccination rate increases, the mortality rate between the scenarios shrinks. Aknur Karabay, Askat Kuzdeuov, Shyryn Ospanova, Michael Lewis 0005, Huseyin Atakan Varol |
IEEE J. Biomed. Health Informatics | 5 |
| 2020 | A Network-Based Stochastic Epidemic Simulator: Controlling COVID-19 With Region-Specific PoliciesabstractIn this work, we present an open-source stochastic epidemic simulator calibrated with extant epidemic experience of COVID-19. The simulator models a country as a network representing each node as an administrative region. The transportation connections between the nodes are modeled as the edges of this network. Each node runs a Susceptible-Exposed-Infected-Recovered (SEIR) model and population transfer between the nodes is considered using the transportation networks which allows modeling of the geographic spread of the disease. The simulator incorporates information ranging from population demographics and mobility data to health care resource capacity, by region, with interactive controls of system variables to allow dynamic and interactive modeling of events. The single-node simulator was validated using the thoroughly reported data from Lombardy, Italy. Then, the epidemic situation in Kazakhstan as of 31 May 2020 was accurately recreated. Afterward, we simulated a number of scenarios for Kazakhstan with different sets of policies. We also demonstrate the effects of region-based policies such as transportation limitations between administrative units and the application of different policies for different regions based on the epidemic intensity and geographic location. The results show that the simulator can be used to estimate outcomes of policy options to inform deliberations on governmental interdiction policies. Askat Kuzdeuov, Daulet Baimukashev, Aknur Karabay, Bauyrzhan Ibragimov, Almas Mirzakhmetov, Mukhamet Nurpeiissov, Michael Lewis 0005, Huseyin Atakan Varol |
IEEE J. Biomed. Health Informatics | 8 |
| 2019 | Color-Coded Fiber-Optic Tactile Sensor for an Elastomeric Robot SkinabstractThe sense of touch is essential for reliable mapping between the environment and a robot which interacts physically with objects. Presumably, an artificial tactile skin would facilitate safe interaction of the robots with the environment. In this work, we present our color-coded tactile sensor, incorporating plastic optical fibers (POF), transparent silicone rubber and an off-the-shelf color camera. Processing electronics are placed away from the sensing surface to make the sensor robust to harsh environments. Contact localization is possible thanks to the lower number of light sources compared to the number of camera POFs. Classical machine learning techniques and a hierarchical classification scheme were used for contact localization. Specifically, we generated the mapping from stimulation to sensation of a robotic perception system using our sensor. We achieved a force sensing range up to 18 N with the force resolution of around 3.6 N and the spatial resolution of 8 mm. The color-coded tactile sensor is suitable for tactile exploration and might enable further innovations in robust tactile sensing. Zhanat Kappassov, Daulet Baimukashev, Zharaskhan Kuanyshuly, Yerzhan Massalin, Arshat Urazbayev, Huseyin Atakan Varol |
ICRA | 6 |
| 2018 | A Series Elastic Tactile Sensing Array for Tactile Exploration of Deformable and Rigid ObjectsabstractTactile sensing arrays are used to detect contacts of robotic systems with the environment. They are particularly useful for scenarios in which vision-based sensors cannot be used. Thanks to the presence of multiple sensing elements, tactile arrays also provide spatial information about the contact location. In this work, we present our series elastic tactile array to enable tactile exploration for position-controlled robot manipulators. Sixteen compliant sensing elements are arranged as a 4×4 array. This allows the position-controlled robot to explore objects via palpation. Tactile sensing was accomplished by measuring the change of the magnetic field caused by neodymium magnets embedded into the series elastic elements. We demonstrate the efficacy of our sensor with two sets of experiments involving physical interaction scenarios. Firstly, we show that the sensor can be used to differentiate between rigid and deformable objects. Secondly, we show that point clouds of objects can be generated quickly with our sensor module attached to a position-controlled robot manipulator as an end-effector. Zhanat Kappassov, Daulet Baimukashev, Olzhas Adiyatov, Shyngys Salakchinov, Yerzhan Massalin, Huseyin Atakan Varol |
IROS | 6 |
| 2014 | Inertial motion capture based reference trajectory generation for a mobile manipulatorabstractThis paper presents a human-robot interaction system integrating a mobile manipulator with a full-body inertial motion capture suit. The framework aims to provide an intuitive, effective and easy-to-deploy teleoperation interface. User control intent is acquired by the motion capture system in real-time, processed and relayed to the mobile manipulator wirelessly. Specifically, body center of mass and the right hand kinematic data of the user are used to generate position and orientation references for the robot base and manipulator, respectively. The left arm is employed to provide high-level user commands such as "Manipulator On/Off", "Base On/Off" and "Manipulator Pause/Resume". The efficacy of the presented system was demonstrated in real-time teleoperation experiments using KUKA youBot mobile manipulator accomplishing pick-and-place tasks. Yerbolat Khassanov, Nursultan Imanberdiyev, Huseyin Atakan Varol |
HRI | 3 |
| 2014 | Real-time gesture recognition for the high-level teleoperation interface of a mobile manipulatorabstractThis paper describes an inertial motion capture based arm gesture recognition system for the high-level control of a mobile manipulator. Left arm kinematic data of the user is acquired by an inertial motion capture system (Xsens MVN) in real-time and processed to extract supervisory user interface commands such as "Manipulator On/Off", "Base On/Off" and "Operation Pause/Resume" for a mobile manipulator system (KUKA youBot). Principal Component Analysis and Linear Discriminant Analysis are employed for dimension reduction and classification of the user kinematic data, respectively. The classification accuracy for the six class gesture recognition problem is 95.6 percent. In order to increase the reliability of the gesture recognition framework in real-time operation, a consensus voting scheme involving the last ten classification results is implemented. During the five-minute long teleoperation experiment, a total of 25 high-level commands were recognized correctly by the consensus voting enhanced gesture recognizer. The experimental subject stated that the user interface was easy to learn and did not require extensive mental effort to operate. Yerbolat Khassanov, Nursultan Imanberdiyev, Huseyin Atakan Varol |
HRI | 3 |
| 2014 | Medical decision support using machine learning for early detection of late-onset neonatal sepsisabstractOBJECTIVE: The objective was to develop non-invasive predictive models for late-onset neonatal sepsis from off-the-shelf medical data and electronic medical records (EMR). DESIGN: The data used in this study are from 299 infants admitted to the neonatal intensive care unit in the Monroe Carell Jr. Children's Hospital at Vanderbilt and evaluated for late-onset sepsis. Gold standard diagnostic labels (sepsis negative, culture positive sepsis, culture negative/clinical sepsis) were assigned based on all the laboratory, clinical and microbiology data available in EMR. Only data that were available up to 12 h after phlebotomy for blood culture testing were used to build predictive models using machine learning (ML) algorithms. MEASUREMENT: We compared sensitivity, specificity, positive predictive value and negative predictive value of sepsis treatment of physicians with the predictions of models generated by ML algorithms. RESULTS: The treatment sensitivity of all the nine ML algorithms and specificity of eight out of the nine ML algorithms tested exceeded that of the physician when culture-negative sepsis was included. When culture-negative sepsis was excluded both sensitivity and specificity exceeded that of the physician for all the ML algorithms. The top three predictive variables were the hematocrit or packed cell volume, chorioamnionitis and respiratory rate. CONCLUSIONS: Predictive models developed from off-the-shelf and EMR data using ML algorithms exceeded the treatment sensitivity and treatment specificity of clinicians. A prospective study is warranted to assess the clinical utility of the ML algorithms in improving the accuracy of antibiotic use in the management of neonatal sepsis. Subramani Mani, Asli Ozdas, Constantin F. Aliferis, Huseyin Atakan Varol, Qingxia Chen, Randy J. Carnevale, Yukun Chen 0001, Joann Romano-Keeler, Hui Nian, Jörn-Hendrik Weitkamp |
J. Am. Medical Informatics Assoc. | 4 |
| 2009 | Early Prediction of Reading Disability using Machine Learning
Huseyin Atakan Varol, Subramani Mani, Donald L. Compton, Lynn S. Fuchs, Douglas Fuchs |
AMIA | 1 |