VLDB 2026 Research / reviewers in the wild / expert
Munkhjargal Gochoo
dblp:149/9570
· DBLP profile ↗
23ranked-venue papers
8as first author
12since 2021 · last 2027
0000-0002-6613-7435ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 6 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 7 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | TCubes: Non-privacy invasive dataset for activities of daily living as thermal cubesabstractThis work introduces the first open multi–infrared (IR) thermal array sensor dataset for recognizing 19 activities of daily living (ADLs) performed by 74 subjects. Each activity sample forms a 32×32×32 thermal tensor cube-hence the name TCubes. The dataset surpasses existing resources in scale, featuring the highest number of activities, twice as many sensors, and three times as many participants as comparable datasets. Thermal patterns were captured using nine low-resolution (8×8) IR thermal sensors, providing a non-invasive and privacy-preserving means of activity recognition. A comprehensive benchmarking study evaluates both convolutional and transformer-based architectures—including C3D, R(2+1)D-18, MViTv2, and Swin-T—to assess their ability to learn spatiotemporal representations from coarse thermal imagery. Results are highly promising: R(2+1)D-18 achieves the most consistent performance with an F1-score of 0.900, while transformer models such as MViTv2 and Swin-T effectively capture subtle gestures and generalize well across 18 and 19 activity classes. The lightweight 3DCNN-Mixed model further demonstrates strong efficiency for resource-constrained applications, highlighting the trade-off between accuracy and computational cost. Analyses leveraging entropy, mutual confusion, and frequency-domain representations reveal how factors such as activity location, posture, temporal dynamics, and motion periodicity influence recognition accuracy. Overall, this dataset and benchmarking suite establish a robust foundation for future research in low-resolution, low-compute, non-invasive, and privacy-preserving human activity recognition, with broad implications for eldercare, healthcare monitoring, and smart environments. Luubaatar Badarch, Byambaa Dorj, Hansaem Park, Gantumur Tsogtgerel, Fady Shibata-Alnajjar, Munkhjargal Gochoo |
Expert Syst. Appl. | 6 |
| 2025 | Comparing Emotion Detection Methods in Online Classrooms: YOLO Models, Multimodal LLM, and Human BaselineabstractThe COVID-19 pandemic has transformed learning environments, challenging educators to understand students' behaviors during the online mode of learning, in particular emotions associated with students' attention during virtual classrooms. As learning transitions between physical and virtual spaces, the ability to interpret student attention and engagement has become complex. In response to this challenge, our research investigates the use of GPT-4o, a multimodal large language model, for identifying student emotions by analyzing images in diverse learning settings. The study involved analyzing online classroom images featuring 149 faces, utilizing three distinct approaches: a computer vision model (YOLO), the multimodal LLM (GPT-4o), and a human-annotated baseline. The analysis systematically categorized facial expressions into eight emotional categories: Happy, Sad, Angry, Neutral, Contempt, Disgust, Fear, and Surprise. The findings indicate that multimodal LLMs can effectively detect student emotions, achieving an average accuracy of 93.8%, which aligns with the human baseline accuracy of 97.0%. In contrast, YOLO models maintained an average accuracy of 81.9%, performing well for basic emotions but struggling with subtle expressions. This research contributes to enhancing educational practices by providing valuable insights regarding the application of multimodal LLMs to assist educators in comprehending student emotions within both physical and digital classroom settings. Medha Mohan Ambali Parambil, Salah Bouktif, Munkhjargal Gochoo, Fady Shibata-Alnajjar |
EDUCON | 3 |
| 2024 | CVQA: Culturally-diverse Multilingual Visual Question Answering BenchmarkabstractVisual Question Answering~(VQA) is an important task in multimodal AI, which requires models to understand and reason on knowledge present in visual and textual data. However, most of the current VQA datasets and models are primarily focused on English and a few major world languages, with images that are Western-centric. While recent efforts have tried to increase the number of languages covered on VQA datasets, they still lack diversity in low-resource languages. More importantly, some datasets extend the text to other languages, either via translation or some other approaches, but usually keep the same images, resulting in narrow cultural representation. To address these limitations, we create CVQA, a new Culturally-diverse Multilingual Visual Question Answering benchmark dataset, designed to cover a rich set of languages and regions, where we engage native speakers and cultural experts in the data collection process. CVQA includes culturally-driven images and questions from across 28 countries in four continents, covering 26 languages with 11 scripts, providing a total of 9k questions. We benchmark several Multimodal Large Language Models (MLLMs) on CVQA, and we show that the dataset is challenging for the current state-of-the-art models. This benchmark will serve as a probing evaluation suite for assessing the cultural bias of multimodal models and hopefully encourage more research efforts towards increasing cultural awareness and linguistic diversity in this field. Chenyang Lyu, Haryo Akbarianto Wibowo, Santiago Góngora, Aishik Mandal, Sukannya Purkayastha, Jesús-Germán Ortiz-Barajas, Emilio Villa-Cueva, Jinheon Baek, Soyeong Jeong, Injy Hamed, Zheng Wei Lim, Paula Mónica Silva, Jocelyn Dunstan, Mélanie Jouitteau, David Le Meur, Joan Nwatu, Ganzorig Batnasan, Munkh-Erdene Otgonbold, Munkhjargal Gochoo, Guido Ivetta, Luciana Benotti, Laura Alonso Alemany, Hernán Maina, Jiahui Geng, Tiago Timponi Torrent, Frederico Belcavello, Marcelo Viridiano, Jan Christian Blaise Cruz, Dan John Velasco, Oana Ignat, Zara Burzo, Chenxi Whitehouse, Artem Abzaliev, Teresa Clifford, Grainne Caulfield, Teresa Lynn, Christian Salamea Palacios, Vladimir Araujo, Yova Kementchedjhieva, Mihail Mihaylov, Israel Abebe Azime, Henok Biadglign Ademtew, Bontu Fufa Balcha, Naome A. Etori, David Ifeoluwa Adelani, Rada Mihalcea, Atnafu Lambebo Tonja, Maria Camila Buitrago Cabrera, Gisela Vallejo, Holy Lovenia, Ruochen Zhang 0001, Marcos Estecha-Garitagoitia, Mario Rodríguez-Cantelar, Toqeer Ehsan, Rendi Chevi, Muhammad Farid Adilazuarda, Ryandito Diandaru, Samuel Cahyawijaya, Fajri Koto, Tatsuki Kuribayashi, Haiyue Song, Aditya Khandavally, Thanmay Jayakumar, Raj Dabre, Mohamed Fazli Mohamed Imam, Kumaranage Ravindu Yasas Nagasinghe, Alina Dragonetti, Luis Fernando D'Haro, Olivier Niyomugisha, Jay Gala, Pranjal A. Chitale, Fauzan Farooqui, Thamar Solorio, Alham Fikri Aji |
NeurIPS | 21 |
| 2023 | Fisheye Multiple Object Tracking by Learning Distortions Without DewarpingabstractWe develop a new Multiple Object Tracking (MOT) scheme for fisheye cameras that can directly perform vehicle detection, re-identification, and tracking under fisheye distortions without explicit dewarping. Fisheye cameras provide omnidirectional coverage that is wider than traditional cameras, reducing fewer need of cameras to monitor road intersections. However, the problem of distorted views introduces new challenges for fisheye MOT. In this paper, we propose a Fish-Eye Multiple Object Tracking (FEMOT) approach with two novelties. We develop the Distorted Fisheye Image Augmentation (DFIA) method to improve object detection and re-identification on fisheye cameras, where fisheye model training can be performed on existing datasets of traditional cameras via fisheye data synthesis and augmentation. We also develop the Hybrid Data Association (HDA) method to perform tracking directly on fisheye views, without the need of de-warping. The developed FEMOT framework provides practical design and advancement that enables large-scale use of fisheye cameras in smart city and surveillance applications. Ping-Yang Chen, Jun-Wei Hsieh, Ming-Ching Chang, Munkhjargal Gochoo, Fang-Pang Lin, Yong-Sheng Chen |
ICIP | 4 |
| 2023 | Fine-Tuning Vision Transformer for Arabic Sign Language Video Recognition on Augmented Small-Scale DatasetabstractWith the rise of AI, the recognition of Sign Language (SL) through sign-to-text has gained significance in the field of computer vision and deep machine learning. However, there are only a few medium to large open datasets available for this task, as it requires a vast dataset of thousands of signs for words/phrases in different environments, which is a time-consuming and tedious process. Furthermore, there has been very little effort towards Arabic Sign Language Recognition (ArSLR). This research paper presents the results of fine-tuning the Vision Transformer (ViT) model on a small-scale in-house dataset of ArSL. The main goal is to attain satisfactory results by utilizing minimal computing power and a small dataset involving less than 10 individuals, with only one recording made for each sign in every environment. The dataset comprises 49 classes/signs, all of which were made with two hands and belong to the Level I category in terms of popularity. To enhance the dataset, three types of augmentations - translation, shear, and rotation were employed. The ViT model, pre-trained on the Kinetics dataset, was trained on the variation of augmented datasets with 2 to 40 times samples for each original video, where the training set includes original and augmented videos of 8 volunteers and the test set includes only original videos of one particular volunteer. Experimental results reveal that the combination of rotation and shear outperformed the others, achieving an accuracy of 93% on the 20 times augmented samples per class per signer dataset. We believe this study sheds light on small-scale dataset-based SLR tasks and video/action recognition in general. Munkhjargal Gochoo, Ganzorig Batnasan, Ahmed Abdelhadi Ahmed, Munkh-Erdene Otgonbold, Fady Shibata-Alnajjar, Timothy K. Shih, Tan-Hsu Tan, Khin Wee Lai |
SMC | 1 |
| 2023 | Maximum entropy scaled super pixels segmentation for multi-object detection and scene recognition via deep belief network
Adnan Ahmed Rafique, Munkhjargal Gochoo, Ahmad Jalal |
Multim. Tools Appl. | 2 |
| 2023 | A simulated measurement for COVID-19 pandemic using the effective reproductive number on an empirical portion of population: epidemiological models
Belal Alsinglawi, Omar Mubin, Fady Shibata-Alnajjar, Khalid Kheirallah, Mahmoud Elkhodr, Mohammed Al-Zobbi, Mauricio Novoa, Mudassar Arsalan, Tahmina Nasrin Poly, Munkhjargal Gochoo, Gulfaraz Khan, Kapal Dev |
Neural Comput. Appl. | 10 |
| 2022 | ArSL21L: Arabic Sign Language Letter Dataset Benchmarking and an Educational Avatar for Metaverse ApplicationsabstractIt is complicated for the PwHL (people with hearing loss) to make a relationship with social majority, which naturally demands an interactive auto computer systems that have ability to understand sign language. With a trending Metaverse applications using augmented reality (AR) and virtual reality (VR), it is easier and interesting to teach sign language remotely using an avatar that mimics the gesture of a person using AI (Artificial Intelligence)-based system. There are various proposed methods and datasets for English SL (sign language); however, it is limited for Arabic sign language. Therefore, we present our collected and annotated Arabic Sign Language Letters Dataset (ArSL21L) consisting of 14202 images of 32 letter signs with various backgrounds collected from 50 people. We benchmarked our ArSL21L dataset on state-of-the-art object detection models, i.e., 4 versions of YOLOv5. Among the models, YOLOv5l achieved the best result with COCOmAP of 0.83. Moreover, we provide comparison results of classification task between ArSL2018 dataset, the only Arabic sign language letter dataset for classification task, and our dataset by running classification task on in-house short video. The results revealed that the model trained on our dataset has a superior performance over the model trained on ArSL2018. Moreover, we have created our prototype avatar which can mimic the ArSL (Arabic Sign Language) gestures for Metaverse applications. Finally, we believe, ArSL21L and the ArSL avatar will offer an opportunity to enhance the research and educational applications for not only the PwHL, but also in general real and virtual world applications. Our ArSL21L benchmark dataset is publicly available for research use on the Mendeley. Ganzorig Batnasan, Munkhjargal Gochoo, Munkh-Erdene Otgonbold, Fady Shibata-Alnajjar, Timothy K. Shih |
EDUCON | 2 |
| 2022 | Mixed Stage Partial Network and Background Data Augmentation for Surveillance Object DetectionabstractState-of-the-art (SoTA) object detection models and their accuracy have been improved by a large margin via CNNs (Convolutional Neural Networks); however, these models still perform poorly for small road objects. Moreover, the SoTA models are mainly trained on public benchmark datasets such as MS COCO, which include more complicated backgrounds and thus make them robust for object detection. However, for surveillance or road videos, their monotone backgrounds make these SoTA detectors background-over-fitted. In applications such as autonomous driving or traffic flow estimation, the background-over-fitting problem will increase various challenges and lead to accuracy degradation in object detection. One novelty of this paper is to propose an MBA (Mixed Background Augmentation) method to improve detection accuracy without adding new labeling efforts and any pre-training processes. During the inference stage, only one input image is needed for vehicle detection without involving background subtraction. Another novelty of this paper is the design of an efficient MSP (Mixed Stage Partial) network to detect objects more accurately and efficiently from surveillance videos. Extensive experiments on KITTI and UA-DETRAC benchmarks show that the proposed method achieves the SoTA results for highly accurate and efficient vehicle detection. The detection accuracy is improved from 78.53% to 83.59% with 25.7$fps$on the UA-DETRAC data set. The implementation code is available athttps://github.com/pingyang1117/MSPNet. Ping-Yang Chen, Jun-Wei Hsieh, Munkhjargal Gochoo, Yong-Sheng Chen |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Light-Weight Mixed Stage Partial Network for Surveillance Object Detection with Background Data AugmentationabstractState-of-the-art (SoTA) models have improved object detection accuracy with a large margin via convolutional neural networks, however still with an inferior performance for small objects. Moreover, these models are trained mainly based on the COCO dataset, and its backgrounds are more complicated than road environments, and thus degrade the accuracy of small road object detection. Compared with the COCO dataset, the background of a surveillance video is relatively stable and can be used to enhance the accuracy of road object detection. This paper designs a computationally efficient mixed stage partial (MSP) network to detect road objects. Another novelty of this paper is to propose a mixed background data augmentation method to enhance the detection accuracy without adding new labelling efforts. During inference, only the input image is used to detect road objects without further using any subtraction information. Extensive experiments on KITTI and UA-DETRAC benchmarks show the proposed method achieves the SoTA results for highly-accurate and efficient road object detection. Ping-Yang Chen, Jun-Wei Hsieh, Munkhjargal Gochoo, Yong-Sheng Chen |
ICIP | 3 |
| 2021 | Ultra-Low Resolution Infrared Sensor-Based Wireless Sensor Network for Privacy-Preserved Recognition of Daily Activities of LivingabstractIn the last few decades, the number of elderly people living alone surge worldwide due to an increase in human life expectancy. Researchers suggest different types of methods to monitor elderly residents living alone to prevent them from being unable to get help on time after some incidents such as accidental falling or heart attack. High-resolution data generating methods in monitoring elderly people with RGB cameras or wearable devices are inconvenient for elderly residents. While the former raises a privacy concern, the latter is not practical to wear on and off or charge its battery frequently. One of the solutions that solve both aforementioned problems is ultra-low resolution infrared (IR) sensor arrays. We offer a dataset for Activities of Daily Living (ADL) collected with 8×8 IR sensor arrays from 74 volunteers. For ADL recognition, four types of deep learning models, Convolutional Neural Network (CNN), two types of Recurrent Neural Networks (RNN), and Transformer models are employed. Among them, CNN and Transformer models showed promising results. We believe the dataset is a good contribution to versatile data sources for researchers to accelerate their work on the development of privacy-preserved ADL recognition systems. Luubaatar Badarch, Munkhjargal Gochoo, Ganzorig Batnasan, Fady Shibata-Alnajjar, Tan-Hsu Tan |
NCA | 2 |
| 2021 | Privacy-Preserved Social Distancing System Using Low-Resolution Thermal Sensors and Deep LearningabstractThe COVID-19 pandemic brought drastic changes to daily routines and abiding by the guideline for social distancing is necessary to prevent the spread of the virus while easing back to normality. While camera-based social distancing solutions are widely available, they do not preserve the privacy of the user. To the best of our knowledge, there is no practical privacy-preserved social distancing technology in the market. This research employs IoT and Deep Learning to implement an indoor privacy-preserved social distancing wireless sensor network-based system that uses 8×8 low-resolution infrared sensors (AMG8833) to promote social distancing within companies and organizations. This research uses four top view wireless sensor nodes to collect a total of 6,606 low-resolution infrared images, creating a dataset that covers 200 cases of varying numbers of people in diverse locations. The YOLOv4-tiny model achieved a [email protected] of 95.4% and an inference time of 5.16ms when trained on the collected and labeled dataset to detect people. Thus, we conclude that our proposed novel social distancing system is as effective as other solutions that use high-resolution images while maintaining privacy. Aisha Fahad Alraeesi, Hanan Fekri Kharbash, Jawaher Saif Alghfeli, Shamma Sultan Alsaedi, Munkhjargal Gochoo |
SMC | 5 |
| 2020 | Drone-Based Vehicle Flow Estimation and its Application to Traffic Conflict Hotspot Detection at IntersectionsabstractDrones can provide a wider field of view, high mobility and flexibility for monitoring and analyzing traffic flows and safety conditions. In case of a perpendicular viewing angle to the ground, there will be a very less occlusion that can occur and make vehicle tracking be easier. Thus, a drone-based solution will be better for traffic conflict hotspot detection at an interaction. However, due to its observation far from the ground, limited battery time, and bandwidth, this solution should be edge-based and have a good recognition rate in small object detection. However, current edge-based SoTA (state-of-the-art) methods are weak in a small object detection. We propose CoBiF net (Concatenated Bi-Fusion feature pyramid network), a one-stage object detection model for a real-time small object detection, which consists of SPP (spatial pyramid pooling), FE (Feature Extractor), CF (Concatenated Feature) block, and BFM (Bottom-up Fusion Module). CoBiF net is memory-and-bandwidth saving for the most edge devices. Extensive experiments on UA VDT benchmark show the proposed method achieved the SoTA results for the small object detection task in terms of accuracy and efficiency. Ping-Yang Chen, Jun-Wei Hsieh, Munkhjargal Gochoo, Ming-Ching Chang, Chien-Yao Wang, Yong-Sheng Chen, Hong-Yuan Mark Liao |
ICIP | 3 |
| 2020 | Lownet: Privacy Preserved Ultra-Low Resolution Posture Image ClassificationabstractIndoor posture recognition is vital for monitoring/detecting exercises, activities of daily living, accidental falls, unusual behavior, etc. However, high-resolution image based systems have a high accuracy, they are considered as intrusive and most of the current state-of-the-art image classifiers (VGG, ImageNet, ResNext) are not applicable for ultra-low resolution (<; 32 pixels in extent) image classification due to their downsizing feature extraction architecture. Thus, we propose a shallow LowNet model for classifying privacy preserved 16x16 posture images with its feature preserving architecture, variable ReLU slopes, and a custom loss function. LowNet outperformed, with an Accuracy of 98.94% and F1-score of 79.86%, the existing models (LeNet, ResNet1, ResNet-2) which can run on our Ultra lowresolution Thermal Posture Image (UTPI38) dataset (offered here) with 38 classes (4374 samples) collected from 23 volunteers. More experimental results are discussed on the custom loss, and variable ReLU slopes which gave 8.2% performance increase. Thus, we conclude that LowNet is useful in a multiclass ultra-low-resolution thermal posture image classification task. Munkhjargal Gochoo, Tan-Hsu Tan, Fady Shibata-Alnajjar, Jun-Wei Hsieh, Ping-Yang Chen |
ICIP | 1 |
| 2020 | Deep Real-time Hand Detectoin Using CFPN on Embedded SystemsabstractReal-time HI (Human Interface) systems need accurate and efficient hand detection models to meet the limited resources in budget, dimension, memory, computing, and electric power. In recent years, object detection became a less challenging task with the latest deep CNN-based state-of-the-art models, i.e., RCNN, SSD, and YOLO; however, these models cannot provide the desired efficiency and accuracy for HI systems on embedded devices due to their complex time-consuming architecture. In addition, the detection of small hands ( pixels) is still a challenging task for all the above existing methods. Thus, we propose a shallow model named Concatenated Feature Pyramid Network (CFPN) to provide above mentioned performance for small hand detection. The superiority of CFPN is confirmed on a HandFlow dataset with mAP:0.5 of 95.6 and FPS of 33 on Nvidia TX2. The COCO dataset is also used to compare with other state-of-the-art method and shows the highest efficiency and accuracy with the proposed CFPN model. Thus we conclude that the proposed model is useful for real-life small hand detection on embedded devices. Pirdiansyah Hendri, Jun-Wei Hsieh, Ping-Yang Chen, Munkhjargal Gochoo, Yong-Sheng Chen |
ICPR | 4 |
| 2019 | Smaller Object Detection for Real-Time Embedded Traffic Flow Estimation Using Fish-Eye CamerasabstractReal-time embedded traffic flow estimation (RETFE) systems need accurate and efficient vehicle detection models to meet limited resources in budget, dimension, memory, and computing power. In recent years, object detection became a less challenging task with latest deep CNN-based state-of-the-art models, i.e., RCNN, SSD, and YOLO; however, these models cannot provide desired performance for RETFE systems due to their complex time-consuming architecture. In addition, small object (<; 30×30 pixels) detection is still a challenging task for existing methods. Thus, we propose a shallow model named Concatenated Feature Pyramid Network (CFPN) that inspired from YOLOv3 to provide above mentioned performance for the smaller object detection. Main contribution is a proposed concatenated block (CB) which has reduced number of convolutional layers and concatenations instead of time-consuming algebraic operations. The superiority of CFPN is confirmed on the COCO and an in-house CarFlow datasets on Nvidia TX2. Thus we conclude that CFPN is useful for real-time embedded smaller object detection task. Ping-Yang Chen, Jun-Wei Hsieh, Munkhjargal Gochoo, Chien-Yao Wang, Hong-Yuan Mark Liao |
ICIP | 3 |
| 2019 | Chronic Kidney Disease Stage Classification Using Renal Artery Doppler-Derived ParametersabstractIn renal medicine, Estimated Glomerular Filtration Rate (eGFR) based method is a standard for the diagnosis of chronic kidney disease. However, this method is invasive, uncomfortable, costly, and could be dangerous because it requires to draw blood from the artery vessels. Researchers have developed several non-invasive Doppler-derived measures based chronic kidney disease (CKD) stage diagnosing or prognosing approaches; however, there is no adequate automatic renal artery Doppler-derived CKD stage classification method in the literature. Thus, we propose a non-invasive, safer, faster, and low cost, SVM-based CKD stage classification method from a sonogram of the renal artery blood flow. The proposed method extracts kurtosis and curvature parameters of the probability distribution that generated from renal artery blood flow waveform. Kurtosis and curvatures are employed to measure the tailedness and curvedness of the probability distribution. We collected a total of 528 sonograms from 110 (49 males) CKD patients during 2010-2013. The experimental results revealed a statistically significant correlation between the parameters and CKD progress stages. Post-voting results revealed the best f1score of 0.956 for Positive (stages 1-5) CKD stages. Munkhjargal Gochoo, Jun-Wei Hsieh, Chien-Hung Lee, Yun-Chih Chen, Yu-Chi Shih |
SMC | 1 |
| 2019 | Novel IoT-Based Privacy-Preserving Yoga Posture Recognition System Using Low-Resolution Infrared Sensors and Deep LearningabstractIn recent years, the number of yoga practitioners has been drastically increased and there are more men and older people practice yoga than ever before. Internet of Things (IoT)-based yoga training system is needed for those who want to practice yoga at home. Some studies have proposed RGB/Kinect camera-based or wearable device-based yoga posture recognition methods with a high accuracy; however, the former has a privacy issue and the latter is impractical in the long-term application. Thus, this paper proposes an IoT-based privacy-preserving yoga posture recognition system employing a deep convolutional neural network (DCNN) and a low-resolution infrared sensor-based wireless sensor network (WSN). The WSN has three nodes (x, y, and z-axes) where each integrates 8 × 8 pixels' thermal sensor module and a Wi-Fi module for connecting the deep learning server. We invited 18 volunteers to perform 26 yoga postures for two sessions each lasted for 20 s. First, recorded sessions are saved as .csv files, then preprocessed and converted to grayscale posture images. Totally, 93200 posture images are employed for the validation of the proposed DCNN models. The tenfold cross-validation results revealed that F1-scores of the models trained with xyz (all 3-axes) and y (only y-axis) posture images were 0.9989 and 0.9854, respectively. An average latency for a single posture image classification on the server was 107 ms. Thus, we conclude that the proposed IoT-based yoga posture recognition system has a great potential in the privacy-preserving yoga training system. Munkhjargal Gochoo, Tan-Hsu Tan, Shih-Chia Huang, Tsedevdorj Batjargal, Jun-Wei Hsieh, Fady Shibata-Alnajjar, Yung-fu Chen |
IEEE Internet Things J. | 1 |
| 2019 | Unobtrusive Activity Recognition of Elderly People Living Alone Using Anonymous Binary Sensors and DCNNabstractElderly population (over the age of 60) is predicted to be 1.2 billion by 2025. Most of the elderly people would like to stay alone in their own house due to the high eldercare cost and privacy invasion. Unobtrusive activity recognition is the most preferred solution for monitoring daily activities of the elderly people living alone rather than the camera and wearable devices based systems. Thus, we propose an unobtrusive activity recognition classifier using deep convolutional neural network (DCNN) and anonymous binary sensors that are passive infrared motion sensors and door sensors. We employed Aruba annotated open data set that was acquired from a smart home where a voluntary single elderly woman was living inside for eight months. First, ten basic daily activities, namely, Eating, Bed_to_Toilet, Relax, Meal_Preparation, Sleeping, Work, Housekeeping, Wash_Dishes, Enter_Home, and Leave_Home are segmented with different sliding window sizes, and then converted into binary activity images. Next, the activity images are employed as the ground truth for the proposed DCNN model. The 10-fold cross-validation evaluation results indicated that our proposed DCNN model outperforms the existing models with F1-score of 0.79 and 0.951 for all ten activities and eight activities (excluding Leave_Home and Wash_Dishes), respectively. Munkhjargal Gochoo, Tan-Hsu Tan, Shing-Hong Liu, Fu-Rong Jean, Fady Shibata-Alnajjar, Shih-Chia Huang |
IEEE J. Biomed. Health Informatics | 1 |
| 2018 | Device-Free Non-Privacy Invasive Indoor Human Posture Recognition Using Low-Resolution Infrared Sensor-Based Wireless Sensor Networks and DCNNabstractHuman posture recognition is the foundation of the human activity monitoring. The activity monitoring system is in high demand for the elderly living alone to monitor their health status and accidental fall since the world elderly population will be doubled by 2050. Researchers have developed many camera or wearable device-based human recognition systems; however, they are considered to be privacy-invasive and/or not practical for the long-term monitoring. We propose a device-free unobtrusive indoor human posture recognition system leveraging a low-resolution infrared sensor-based wireless sensor network and deep convolutional neural network (DCNN). We integrated AMG8833 sensor module with 8×8 thermal sensors for sensing the human body temperature and WiFi module for a wireless sensor network. Three wireless sensor nodes are used to capture 3-axis human thermal image. Totally, 15063 samples are collected from four volunteers while they had performed eight human postures as the ground truth for the 10-fold cross-validation of DCNN models. Experimental results indicate that the highest average F1-score for the eight postures was 0.9981. Thus, the proposed system has the high potential for monitoring elderly daily activities, exercise, and fall in emergency cases. Moreover, we believe that our proposed system will be a milestone in the device-free unobtrusive sensing technology. Munkhjargal Gochoo, Tan-Hsu Tan, Tsedevdorj Batjargal, Oleg Seredin, Shih-Chia Huang |
SMC | 1 |
| 2017 | Device-free non-invasive front-door event classification algorithm for forget event detection using binary sensors in the smart houseabstractMany elderly persons prefer to stay alone in a single-resident house for seeking an independent life and reducing the cost of health care. However, the independent life cannot be maintained if the resident develops dementia. Thus, an early detection of dementia is essential for the elderly to extend their independent lifetime. One of the early symptoms of dementia is forgetting something when the person leaves the house. In this study, we introduce a novel front-door events (exit, enter, visitor, other, and brief-return-and-exit (BRE)) and their classification scheme that validated by using open datasets (n = 10) collected from ten single-resident testbeds by anonymous binary sensors. BRE events occur when four consecutive events (exit-enter-exit-enter) happen in some certain time intervals (t1, t2, and t3), and some of them may be the forget events. Each testbed had one older adult (aged 73 years and over) during the experimental period (μ = 583.1 ± 297.3 days). The algorithm automatically classifies the resident's front-door events and ignores visitor's entrance and exit events. The experimental results reveal the significance of the tiparameters for the number of BRE events. Since BRE events may include forget events, the proposed algorithm could be a useful tool for the forget event detection. Munkhjargal Gochoo, Tan-Hsu Tan, Fu-Rong Jean, Shih-Chia Huang, Sy-Yen Kuo |
SMC | 1 |
| 2016 | Design and application of novel morphological filter used in vehicle detectionabstractIn this paper we represent our proposed novel morphological filter developed under the scope of Taiwan-Mongolian co-project. We applied the implemented filter in vehicle detection from CCTV video signal. Our goalwas to develop a filter that can reduce the noise in background subtracted binary image, which created by camera shake, and unnecessary moving objects such as wave of the tree etc. We compared our filter performance with morphological open, close, erosion, dilation, and median filters. PSNR (Peak Signal to Noise Ratio) is employed for evaluating the performance of the filters, our filter's PSNR was relatively higher (21.39) than the other method. Furthermore, we used our filter for vehicle detection, and detection rate was 100% as the other methods. Thus, we conclude the new filter is sufficient for denoising binary image, and suitable for vehicle detection. Munkhjargal Gochoo, Damdinsuren Bayanduuren, Uyangaa Khuchit, Galbadrakh Battur, Tan-Hsu Tan, Sy-Yen Kuo, Shih-Chia Huang |
ICIS | 1 |
| 2016 | Improved global motion estimation via motion vector clustering for video stabilization
Andrey Kopylov, Shih-Chia Huang, Oleg Seredin, Roman Karpov, Sy-Yen Kuo, K. Robert Lai, Tan-Hsu Tan, Munkhjargal Gochoo, Damdinsuren Bayanduuren, Cihun-Siyong Alex Gong, Patrick C. K. Hung |
Eng. Appl. Artif. Intell. | 9 |