VLDB 2026 Research / reviewers in the wild / expert
Kang-Hyun Jo
dblp:43/1302 · also Kanghyun Jo
· DBLP profile ↗
217ranked-venue papers
4as first author
61since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 62 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 58 · 14 since 2021Human-computer interaction and ubiquitous computing · 50 · 2 first-author · 23 since 2021Systems, architecture and hardware · 46 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 18 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dilated multi-Layer perceptron mixer for faster neural networks
Van-Dung Hoang, Xuan-Thuy Vo, Kang-Hyun Jo |
Neural Networks | 3 |
| 2026 | Optimal Proxy Mining Contrastive Network for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) performance enhancement hinges on extracting the most informative features from unlabeled person datasets. In recent approaches, proxy-based contrastive learning with awareness of camera labels has been adopted for model training, thereby achieving highly promising results. However, inappropriate selections of contrastive pairs can significantly degrade the performance of these models. To address this issue, we propose the Optimal Proxy Mining Contrastive Network (OPMCN), a novel framework designed to strategically optimize the selection of proxies for positive and negative pair formation, thus enhancing the efficacy of contrastive training. The OPMCN framework proposes two specific contrastive losses: Hardest Camera Proxy Mining (HCPM) and False Negative Proxies Mining (FNPM), each essential for enhancing model performance in unsupervised settings. The HCPM loss targets proxies from the most challenging cameras to maximize semantic differences between pairs while ensuring minimal background shifts. In contrast, the FNPM loss counters noise in pseudo labels by prioritizing similarity rankings over clustering results to effectively identify and correct false negatives among proxies. Moreover, we have developed the Pyramid Kernel Global Context (PKGC) block, which employs an attention mechanism that focuses on identity-invariant semantic cues in instances. This module utilizes optimally sized convolutional kernels to enhance identity recognition consistency across camera-based variations, thereby improving the precision of feature extraction. Experimental results on several popular datasets prove that our work surpasses existing unsupervised person Re-ID approaches to a remarkable extent. Ge Cao, Qing Tang 0004, Xuan-Thuy Vo, Adri Priadana, Kang-Hyun Jo |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Large Vision-Language Models with PEFT for Generating Descriptive Annotations in Person Re-IdentificationabstractGenerating descriptive annotations for person re-identification (Re-ID) images is essential for bridging vision and language domain, improving both interpretability and cross-modal retrieval performance. However, large vision-language models (LVLMs), which trained on broad web-scale corpora, often struggle to generate accurate, context-relevant descriptions for ReID samples due to inherent challenges such as occlusions, low resolution, varying illumination, and diverse viewpoints. In this paper, we propose to apply Parameter-Efficient Fine-Tuning (PEFT) via Low-Rank Adaptation (LoRA) to tune Qwen2-VL for ReID-specific captioning tasks. Leveraging existing Re-ID datasets with paired image-text annotations, our fine-tuned model generates domain-aligned and discriminative captions. Experiments show significant improvements in caption relevance and identity descriptiveness, highlighting the potential of PEFT-tuned LVLMs for real-world ReID applications. Ge Cao, Qing Tang 0004, Adri Priadana, Tran Tien Dat, Ashraf Uddin Russo, Kang-Hyun Jo |
HSI | 6 |
| 2025 | Waste Detection on Low-Light EnvironmentabstractOn waste-sorting conveyor lines, illumination often fluctuates sharply or drops to very low levels, causing conventional object detectors to fail. To achieve lighting-robust performance without extra lamps, vision rooms, or site-specific retraining, we propose an Illumination-Invariant Convolution (IIC) block that can be inserted into any backbone network. Working in log-intensity space under a Lambertian model, IIC applies learnable zero-mean cross-channel filters that mute lighting artefacts and boost material cues, and then merges the resulting maps with base features to produce lighting-robust representations. We integrate IIC into the “nano” versions of YOLOv5/8/10/11 and lightweight RT-DETR, training on roughly 200 k conveyor-belt images (29 classes) from the AI-Hub waste dataset. IIC raises YOLO mAP50 by 2.5-4.5 pp and mAP50-95 by up to 4.6 pp; even the data-hungry, transformer-based RT-DETR gains up to 1.5 pp. In low-light video tests, the IIC-augmented models successfully detected objects the original networks missed, raising recall. This demonstrates that inserting the IIC module at a network's input provides a straightforward path to illumination-robust, field-ready waste-sorting systems. Jehwan Choi, Minseung Kim, Kang-Hyun Jo |
HSI | 3 |
| 2025 | Enhanced Breast Cancer Detection: A Transfer Learning Framework with Auxiliary Classification for Mammographic Image AnalysisabstractBreast cancer remains one of the most prevalent malignancies affecting women worldwide, with over 2.3 million new cases diagnosed annually. While conventional diagnostic methods like mammography have reduced mortality rates, artificial intelligence offers opportunities to further enhance detection accuracy. In this paper, we present a comprehensive framework for breast cancer identification using deep learning techniques applied to mammographic images. Our approach leverages transfer learning from ImageNet pretrained models and incorporates an auxiliary classification mechanism to improve diagnostic performance. We evaluated multiple state-of-the-art convolutional neural network architectures, including MobileNetV3, ResNet50, ConvNeXtV2-nano, ResNeXt50, MobileViT-small, and EfficientNet-B3, trained on four publicly available datasets (BMCD, CDD-CESM, CMMD, and MiniDDSM). Performance was assessed on separate external test datasets (VinDr and RSNA) to simulate real-world clinical deployment scenarios. Our experimental results demonstrate that EfficientNet-B3 with auxiliary classification achieved superior performance across most metrics, with an accuracy of 86.14%, F1-score of 85.80%, and AUC of 90.39% on validation data. We further introduce the Probabilistic F1-score as a clinically relevant evaluation metric that accounts for prediction confidence rather than binary decisions alone. The proposed framework, available as an open-source implementation, offers a promising approach for enhancing breast cancer detection while providing insights into the challenges of deploying AI systems in diverse clinical environments. We released our codebase at: https://github.com/thanhhnvnqb/mmbreast Van-Thanh Hoang, Vu Huu Dao, Tu Minh Phuong, Kang-Hyun Jo |
HSI | 4 |
| 2025 | Artificial Behavior Intelligence: Technology, Challenges, and Future DirectionsabstractUnderstanding and predicting human behavior has emerged as a core capability in various AI application domains such as autonomous driving, smart healthcare, surveillance systems, and social robotics. This paper defines the technical frame-work of Artificial Behavior Intelligence (ABI), which comprehensively analyzes and interprets human posture, facial expressions, emotions, behavioral sequences, and contextual cues. It details the essential components of ABI, including pose estimation, face and emotion recognition, sequential behavior analysis, and context-aware modeling. Furthermore, we highlight the transformative potential of recent advances in large-scale pretrained models, such as large language models (LLMs), vision foundation models, and multimodal integration models, in significantly improving the accuracy and interpretability of behavior recognition. Our research team has a strong interest in the ABI domain and is actively conducting research, particularly focusing on the development of intelligent lightweight models capable of efficiently inferring complex human behaviors. This paper identifies several technical challenges that must be addressed to deploy ABI in real-world applications including learning behavioral intelligence from limited data, quantifying uncertainty in complex behavior prediction, and optimizing model structures for low-power, real-time inference. To tackle these challenges, our team is exploring various optimization strategies including lightweight transformers, graph-based recognition architectures, energy-aware loss functions, and multimodal knowledge distillation, while validating their applicability in real-time environments. Kang-Hyun Jo, Jehwan Choi, Kwanho Kim, Duy-Linh Nguyen, Xuan-Thuy Vo, Adri Priadana, Tien-Dat Tran |
HSI | 1 |
| 2025 | Waste Object Detection Using Bright to Dark Feature AlignmentabstractFactory environments vary significantly in lighting and camera conditions, necessitating models that are robust to illumination changes. Although prior studies have addressed this issue by constructing separate low-light datasets, such approaches face scalability challenges due to the high cost of data collection. To overcome this issue, the proposed approach leverages DARK-ISP(Low-light Image Synthesis Pipeline) from the prior work DAI-Net to construct a synthetic low-light image dataset. This approach eliminates the need to collect real dark data. Retraining a model from scratch to handle low-light conditions is not cost-effective. A more practical approach is to leverage high-performance models pretrained in bright industrial environments. While fine-tuning such models is a possible solution, it often suffers from performance degradation due to domain shift. This study adopts a teacher-student framework to perform back-bone feature-level domain alignment between a teacher model trained on well-lit images and a student model trained on low-light images. The alignment is achieved using MMD(Maximum Mean Discrepancy) loss, which effectively mitigates the domain shift problem and reduces the representational gap between the two models. The detection model is based on RT-DETRv2, a lightweight ViT-based architecture that achieves both real-time performance and high object detection accuracy. Building upon the aligned features, a knowledge distillation method based on KD-DETR was applied. This method, specifically tailored for DETR architectures, further enhanced detection performance under low-light conditions. Experiments conducted on ROBOne recyclable waste dataset show that the proposed method achieves a 6.4% higher mAP compared to DAI-Net and a 2.2% improvement over simple fine-tuning on the target domain. Minseung Kim, Jehwan Choi, Jongchae Lee, Kyubin Hwang, Kang-Hyun Jo |
HSI | 6 |
| 2025 | Mathematical Analysis of High-Frequency Spatial Attention for Facial Expression RecognitionabstractFacial expression recognition (FER) refers to the task of matching geometric patterns in facial images to corresponding emotional states. Recent advances have shifted FER from traditional handcrafted approaches to deep learning-based models, where attention mechanisms have become key to improving performance. Our prior work introduced the High Frequency Spatial Attention (HFSA) module, inspired by psychological findings that humans utilize high-frequency components when recognizing facial expressions. HFSA was designed to guide CNN-based feature extractors in effectively leveraging high-frequency information. However, the architecture was heuristically designed without theoretical justification. This study aims to mathematically analyze the inner workings of the HFSA module, which were not addressed in previous research. First, we show that the singular value decomposition (SVD) of a feature map in HFSA corresponds to performing principal component analysis (PCA) on its row (or column) feature vectors. Second, we interpret the cross-correlation between the singular values of the feature map and its high-frequency component as an indirect comparison of their geometric distributions. To support these interpretations, we introduce the Principal Axis Hypothesis, which posits that the principal axes of the feature map and its high-frequency component are structurally similar. To validate this hypothesis, we measured the similarity of principal axes across all CNN blocks. Results showed that the similarity remained consistently in the range of le-4 to le-6 throughout training epochs. Moreover, deeper layers exhibited stronger alignment, suggesting greater similarity in high-level features. These findings also align with prior studies indicating that CNNs behave similarly to high-pass filters. Kang-Hyun Jo |
HSI | 2 |
| 2025 | Optimal Driving Path Analysis Based on Vehicle CharacteristicsabstractThis study provides fundamental data for autonomous vehicles by analyzing vehicles entering a roundabout and deriving the optimal driving path. Using state-of-the-art computer vision techniques, this paper analyzes traffic volumes and individual vehicle dwell times in the Gongeoptap, one of the busiest intersections in Ulsan, Korea. The traffic monitoring system is implemented to capture long-duration traffic footage, enabling the extraction of lane-level vehicle trajectories and time-series movement. The analysis focuses on vehicles entering from zone A which is the highest-volume entry zone. Additionally, dwell time is measured across different lanes and time slots to identify efficiency and congestion patterns. The results reveal that entry lane selection significantly affects internal travel time within the roundabout. Hence, selecting optimal lanes for each exit zone can reduce delays and prevent conflicts. Based on these findings, optimal route recommendations are proposed for each exit zone. These results offer practical implications for both human-driven and autonomous vehicles. Future work will extend this analysis by modeling roundabout congestion types and validating the derived paths through digital twin based simulations. Dokyung Kim, Kwanho Kim, Kang-Hyun Jo |
HSI | 3 |
| 2025 | Velocity-Variance-Based Dynamic Object Filtering for 4D Imaging Radar OdometryabstractAccurate ego-motion estimation is crucial for LiDAR odometry in autonomous navigation. However, in dynamic environments, conventional LiDAR-only SLAM systems often struggle to distinguish between static and dynamic objects, resulting in map degradation and localization drift. To address this, we propose an approach that integrates 4D imaging radar into the LiDAR odometry pipeline by leveraging ego-velocity estimation derived from radar Doppler measurements. Specifically, we introduce a velocity-variance-weighted least squares (WLS) framework that estimates ego-velocity using the variance of per-point velocity as an inverse confidence measure. This method inherently reduces the influence of outliers and dynamic targets without requiring explicit outlier removal. The estimated velocity is used to remove dynamic points from the radar point cloud, making it suitable for integration into LiDARbased odometry frameworks. We evaluate our method on realworld radar datasets collected in dynamic outdoor environments. Experimental results demonstrate that our WLS-based filtering significantly improves LiDAR odometry stability, reducing ghost artifacts and increasing convergence speed. compared to LSQonly methods, our approach achieves lower initial alignment error and stable pose estimation. Kang-Hyun Jo |
HSI | 2 |
| 2025 | 3D Reconstruction Using Colorized LiDAR Point CloudsabstractThis paper proposes a 3D Gaussian Splatting(3DGS) pipeline that directly utilizes LiDAR point data, replacing the sparse point cloud created with COLMAP. Only camera poses are extracted through COLMAP, LiDAR cloud point data is mapped to the 3D space, and downsampled. Each LiDAR point is added by sampling and averaging RGB values by projecting it as a multi-view image. Proposed method initializes the Gaussian Splatting model using the color LiDAR point cloud as the initial Gaussian position and color. As a result of the KITTI dataset experiment, the proposed method showed better performance and higher detail compared to the traditional COLMAP-based initialization, and also performance improvement was achieved in PSNR/SSIM/LPIPS measurements. Jongchae Lee, Minseung Kim, Kang-Hyun Jo |
HSI | 3 |
| 2025 | Probformer: Transformers with Probabilistic Auto-correlation for Acid Dosage PredictionabstractFroth flotation is a crucial separation technique in the metal beneficiation process. Among the various control variables, acid dosage plays a significant role in adjusting the slurry pH and directly impacts flotation efficiency. However, traditional acid dosage prediction methods struggle to achieve accurate modeling due to the nonlinear, time-varying nature of the flotation process and the complex coupling among variables, thereby hindering process optimization and stable control. This study proposes a deep learning-based time series prediction model, Probformer, to address the acid dosage forecasting task. Probformer incorporates a time series decomposition mechanism and a probabilistic autocorrelation attention mechanism to effectively capture long-term trends and periodic variations during prediction. Experiments conducted on real-world industrial data demonstrate that Probformer outperforms conventional models such as Autoformer, Transformer, Informer, and Reformer in terms of prediction accuracy, achieving lower prediction errors and higher fitting accuracy. The results highlight strong performance of Probformer in acid dosage forecasting for froth flotation, indicating its promising potential for practical applications. Yanyu Yin, Kang-Hyun Jo |
HSI | 3 |
| 2025 | Efficient Human Behavior Detector for Vision-based Emergency Evacuation SystemsabstractThe emergency evacuation systems are often installed in crowded places such as airports, train stations, and shopping malls to evacuate and protect people when incidents occur quickly. With the development of surveillance cameras, vision-based emergency evacuation systems have demonstrated their ability to observe and promptly warn flexibly. This paper proposes a human behavior detector by fine-tuning the YOLOv11n detection network with the Global Attention Mechanism (GAM) to enhance the individual human action recognition. Extensive experiments are trained and evaluated on the Human Behavior Detection Dataset (HBDset) using a NVIDIA Tesla V100 32GB GPU. The proposed network achieves 62.2% of mAP and an inference speed of 1.3 milliseconds (ms), and outperforms other networks of the same scale. Duy-Linh Nguyen, Xuan-Thuy Vo, Adri Priadana, Jehwan Choi, Kang-Hyun Jo |
HSI | 5 |
| 2025 | A Hybrid Vision Transformer and Convolutional Neural Network Architecture for Banana Leaf Disease ClassificationabstractThe classification of banana leaf disease plays a crucial role in early disease detection and in preventing the condition from worsening. To handle this task, this study proposes a hybrid Vision Transformer (ViT) architecture that leverages the strengths of both the convolutional and self-attention layers. By leveraging convolutional layers in the earlier stages and self-attention layers in the later stages, the proposed architecture aims to balance effective feature learning and computational cost while achieving better efficiency. Experimental results show that this model achieves an outstanding accuracy of up to 97.65% while maintaining a moderate tradeoff in computational complexity. Thi-Kim-Anh Pham, Duy-Linh Nguyen, Xuan-Thuy Vo, Kang-Hyun Jo |
HSI | 4 |
| 2025 | A Lightweight CNN-Based Framework for Infant Emotional State Recognition in Real-Time ScenariosabstractRecognizing infant emotion is challenging, mainly due to the lack of facial expression data specifically focused on infants. Most publicly available datasets are developed for the general public or adults, which may not accurately capture the distinct facial characteristics of infants. This paper presents the initial stages of developing a deep learning-based framework for infant emotion recognition, utilizing a state-of-the-art and lightweight CNN backbone. To construct a suitable dataset, this paper filters infant faces from an existing dataset using age-related characteristics to create a more focused training set. The curation dataset is then used to refine the current CNN architecture selection. During testing, facial regions are detected in real-time from input videos to localize areas of interest before classification. This approach aims to improve the reliability of infant emotion recognition while maintaining efficiency suitable for real-time applications. This paper describes the dataset preparation process, model evaluation, and the design of a real-time workflow. Future work will explore improvements in data quality, model performance, and comprehensive application scenarios. Rahmatullah Arrizal Pranatadesta, Adri Priadana, Kang-Hyun Jo |
HSI | 3 |
| 2025 | Efficiency-Accuracy Trade-Off of Facial Attribute Classifier Supporting Human-Robot InteractionabstractThe advancement of robotics has been driven by the integration of artificial intelligence, machine learning, and sophisticated sensing technologies, enabling more seamless Human-Robot Interaction (HRI). Facial Attribute Classifier (FAC) plays a crucial role in HRI by helping robots understand human emotions, intentions, and social cues, fostering personalized and intuitive interactions. However, while existing methods achieve high accuracy, their computational complexity limits real-time applications on low-cost or CPU-based devices, highlighting the need for lightweight models that balance accuracy and efficiency. This work proposes an Efficient Network (ENet) designed to achieve an optimal trade-off between efficiency and accuracy of FAC. ENet introduces an Enhanced Sequential Efficient Attention Module (ESEAM) to improve the quality of feature maps while maintaining high efficiency. Accordingly, ENet demonstrates a compromise between efficiency and accuracy on the CelebA and LFWA datasets. The proposed ENet is computationally efficient, generating a few parameters, making it well-suited for CPU-based applications. When combined with a face detector, the optimized FAC achieves a processing speed of 25.88 frames per second (FPS) on an Intel Core i7-9750H CPU, demonstrating its suitability for real-time use. Adri Priadana, Duy-Linh Nguyen, Xuan-Thuy Vo, Ge Cao, Jehwan Choi, Kang-Hyun Jo |
HSI | 6 |
| 2025 | Reinforcement Learning Autonomous Driving via Reward Function Design in SimulationabstractIn this paper, we present a deep-reinforcement learning(RL) approach to autonomous driving in a simulated environment, supported by a carefully crafted reward function. Using Unity's ML-Agents toolkit, we train a kart to follow 15 sequential waypoints along a track with six challenging turns. The agent must also learn to avoid five randomly placed static obstacles(large balls and bowling pins) scattered throughout the course. The agent monitors a rich set of state inputs: its current speed, position, and orientation; the direction and distance to the next waypoint; information about any obstacle located at that waypoint; and readings from six ray-cast sensors used for wall detection. A detailed reward function designed to encourage efficient route completion and safe driving behavior. he agent gets a small reward every time step for keeping a steady forward speed. Earn bigger rewards when it gets close to a waypoint and passes it, with an extra bonus for each waypoint hit in a row. Reaching the final waypoint brings a large reward. Crashing into a wall or obstacle gives a penalty. With this reward plan, the agent learned to drive the track, avoid most obstacles, and avoid walls. In 100 test runs it completed the whole course 64% of the time. These results show that a well-designed reward function can help a deep-RLcar balance speed, progress, and safety in a simulation. Jae-hyeon Sung, Kwanho Kim, Kang-Hyun Jo |
HSI | 3 |
| 2025 | Efficient Multi-Scale Spatial Interactions for Visual Recognition TasksabstractConvolution operation has local connectivity and translation equivalence while self-attention operation captures long-range spatial dependencies. Adopting the merits of convolution and self-attention operations in hierarchical networks can result in better visual representation and generalization performance. However, integrating self-attention layers into earlier stages is inefficient because self-attention operation has quadratic complexity with token lengths. In this work, we tackle this issue and propose an Efficient Multi-scale Spatial interaction Network (EMSNet) that takes advantage of hybrid networks. The EMSNet has key insights: (1) Each stage efficiently models both short-range and long-range spatial interactions via the design of the multi-scale tokens; (2) The novel convolution-based multi-head self-attention (C-MHSA) operation is introduced to learn spatial interactions inside local regions; (3) The efficient combination of the depthwise convolution, coordinate depthwise convolution, C-MHSA, and global multi-head self-attention (G-MHSA) are performed via channel splitting strategy, extracting wide ranges of frequencies and multi-order interactions. Extensive experiments on ImageNet-1K image classification, MS-COCO object detection, and segmentation tasks verify the effectiveness and generalization ability of the EMSNet. For instance, the EMS Net-XTiny gets 77.1% Top-1 accuracy on ImageNet-1K which is much greater than PVTvl-Tiny by 2% with only 22% parameters and 37% GFLOPs. Xuan-Thuy Vo, Duy-Linh Nguyen, Adri Priadana, Jehwan Choi, Kang-Hyun Jo |
HSI | 5 |
| 2025 | Geometric Feature Construction and Cluster Analysis of the Concha Region with 2D ImagesabstractThe auricular concha is a crucial component of the auricle and holds significant importance in areas such as auricular acupuncture point localization in traditional Chinese medicine, auricle recognition, and the design of ear-worn devices. In this study, a segmentation model was first employed to perform eight-region segmentation of auricle images, from which the auricular concha subregion images were extracted. Subsequently, considering the characteristics of ear-worn devices, key points within the auricular concha region were selected and detected using a keypoint detection model. A method for constructing geo-metric features of the auricular concha region was then designed. Finally, three clustering methods-K-means, Deep Clustering, and Autoencoder combined with K-means-were applied to the geometric features of the auricular concha region to generate multiple representative auricular concha templates, followed by comparative analysis. Experimental results demonstrate that the Autoencoder + K-means clustering method achieves superior performance in generating personalized design templates for ear-worn devices. This study provides robust technical support for applications in personalized intelligent wearable device design, auricular acupuncture point localization, and auricle recognition. Kang-Hyun Jo |
HSI | 3 |
| 2025 | Ear3D-PAF: PCA guided Adaptive Fusion Network for 3D Ear Point Cloud ReconstructionabstractThis study presents the Ear3D-PAF network, an advanced method for 3D ear point cloud reconstruction, addressing the challenges of data scarcity and complex structural geometry. The approach employs a PCA-guided encoder-decoder architecture to ensure global geometric coherence and high-fidelity reconstruction of local details. By synergistically integrating PCA guidance with a Point Cloud Encoder-Decoder framework, the encoder effectively captures both global and local feature representations. A Curvature-based Adaptive Feature Fusion mechanism enables the decoder to proficiently learn intricate ear geometries. In this strategy, high-curvature regions prioritize features derived from deep learning, while low-curvature regions emphasize PCA-guided features. Geometric consistency is optimized through a composite loss function incorporating Chamfer distance, mean squared error, and normal vector cosine distance. Evaluated on a dataset of 500 ear samples, the proposed method outperforms a PCA with 20 principal components, reducing Chamfer distance by 25.4% and mean squared error by 62%. Compared to PointNet++, it achieves reductions of 57% and 93.9%, respectively. The method exhibits superior reconstruction accuracy in high curvature areas, such as the helix and concavities, providing a robust and precise 3D reconstruction framework for applications in medical diagnostics, biometric authentication, and virtual reality. Hebin Zhou, Kang-Hyun Jo |
HSI | 4 |
| 2025 | A compact version of EfficientNet for skin disease diagnosis application
Van-Thanh Hoang, Nguyen Duy Quang, Tu Minh Phuong, Kang-Hyun Jo, Van-Dung Hoang |
Neurocomputing | 4 |
| 2025 | A High-Accuracy and Faster Face Recognizer Supporting Biometric Continuous Authentication for Smart Factory WorkersabstractSmart factories require secure and sustainable worker authentication for safe operations. Biometric continuous authentication based on facial recognition is one of the most convenient mechanisms. This method applies a face recognition task to verify the captured face as an authorized user. However, existing methods that employ large networks for high-accuracy face recognition incur high computational costs and slow down the process, rendering them unsuitable for continuous operation. This work proposes an efficient and rapid face recognizer with high accuracy. It offers a faster face residual network, containing efficient FasterFace blocks and efficient channel spatial attention for improved feature extraction. As a result, the proposed network achieves 97.08% based on average accuracy, outperforming the other networks on five benchmark datasets. It performs faster at 19.91 frames per second in real time on CPU-based hardware when integrated with a face detector, showcasing its capacity to support real-time biometric continuous authentication for smart factory workers. Adri Priadana, Duy-Linh Nguyen, Xuan-Thuy Vo, Muhamad Dwisnanto Putro, Ge Cao, Kang-Hyun Jo |
IEEE Trans. Ind. Informatics | 6 |
| 2025 | Local Self-Attention With Mixing Abstract Tokens for Urban Autonomous DrivingabstractAlthough local self-attentions exhibit translation equivariance and locality similar to convolution, the model has limited receptive fields and weak modeling ability. The main reason is that self-attention is computed within nonoverlapped windows. To overcome this issue, common methods need further operations to communicate the information across windows, such as window shifting, and sliding. These operations are memory unfriendly, not well supported, and optimized by modern deep-learning frameworks. Alternatively, this article exchanges information across nonoverlapped windows via efficiently mixing abstract tokens (MAT). The MAT block includes the following steps. First, the image tokens are partitioned into windows and each window is merged with an abstract token. Second, in each window, interactions of image tokens and the abstract token to image tokens are performed. Third, because the abstract token learns abstract information from each corresponding window, mixing all abstract tokens via transformer encoder helps to exchange information between local windows and result in global context modeling. Fourth, the global information of the mixed tokens is propagated back to the image tokens through transformer decoder. The MAT block is efficient and easy to implement, only containing matrix multiplications. In addition, this article also proposes a bilinear patch embedding that samples relevant regions of the input tokens based on learned offsets. Extensive experiments are conducted and evaluated with various tasks such as image classification, object detection, and segmentation. As a result, our method achieves promising performances across tasks. For example, MAT-2 accomplishes79.0%top-1 accuracy on ImageNet-1 K with0.7GFLOPs and outperforms the baseline Swin-0.7 G by4.6%while reducing15.2 mson CPU and0.53 mson GPU devices. The MAT-4 surpasses Swin-T by1.8%mIoU with only70%GFLOPs. Xuan-Thuy Vo, Duy-Linh Nguyen, Adri Priadana, Ge Cao, Jehwan Choi, Kang-Hyun Jo |
IEEE Trans. Ind. Informatics | 6 |
| 2024 | Efficient Vision Transformers with Partial Attention
Xuan-Thuy Vo, Duy-Linh Nguyen, Adri Priadana, Kang-Hyun Jo |
ECCV (83) | 4 |
| 2024 | Inverted Residual Bottlenecks with Large Kernel Attention for Remote Scene ClassificationabstractRemote sensing image classification plays a piv-otal role in environmental monitoring and urban planning, yet it faces the challenge of accurately interpreting complex and high-resolution images with fast inference speed for real time applications. To address this, we introduce the Mobile Large Kernel Attention Network (MLKANet), which integrates MobileNetV2's inverted residual structures with the large kernel attention mechanism from the Visual Attention Network (VAN). Our proposed MLKANet achieves a compelling balance of computational efficiency and sophisticated feature extraction, while maintaining the speed from the MobileNetV2 baseline. This study evaluates MLKANet's performance against state-of-the-art models using the Aerial Image Dataset (AID), demonstrating superior accuracy and efficiency. The architecture's effectiveness is further evidenced through an ablation study highlighting the scalability of our approach and class-wise performance analysis that showcases MLKANet's proficiency across various scene types. RussoMohammadAshraf Uddin, Adri Priadana, Ge Cao, Kang-Hyun Jo |
HSI | 4 |
| 2024 | EMPCNet: Facial Attribute Recognition Using Efficient Multi - Perspective Convolution for Human-Robot InteractionabstractHuman-robot interaction has evolved into a significant field in robotics. In this domain, facial attributes are essential as they enable robots to understand human emotions, intentions, and preferences. In robot applications, which typically involve low-cost devices, efficient recognition technology is crucial for promising real-time operation by robots. This work proposes EMPCNet to perform facial attribute recognition, consisting of an Efficient Multi-Perspective Convolution (EMPC) block used to efficiently extract and capture various information from multiple perspectives using different kernel sizes and shapes of convolutional operations. The proposed network, which only utilizes a few parameters and low computational operations, achieves competitive performance on the CelebA and LFWA datasets. Additionally, when integrated with face detection, the proposed EMPCNet operates efficiently in real-time on a CPU with Intel Core i7-9750H, achieving a frame rate of 21.27 frames per second (FPS) with an image input size of$224\times 224$consisting of a face area. Adri Priadana, Duy-Linh Nguyen, Xuan-Thuy Vo, RussoMohammadAshraf Uddin, Kang-Hyun Jo |
HSI | 5 |
| 2024 | Enhancing Unsupervised Domain Adaptive Person Re-identification Clustering Through Parsing-Based Attribute Labeling
Ge Cao, Kang-Hyun Jo |
ICIC (4) | 2 |
| 2024 | Vehicle Movement Status Network on Drone-Perspective View with Adaptive Adversarial LearningabstractThe rapidly developing autonomous driving field now needs a more secure transportation system through information between multiple mobility. Deep learning that can judge traffic conditions by convergence of various sensor data and in particular, research on the convolutional neural network using computer vision are being actively conducted. In addition, recognizing many objects at once in a large area through drone images and understanding the movement of the object is used as safe traffic assistance information. In this study, an image classification study is conducted to determine the status of the vehicle on the road through drone flight image data. The goal is to build a new image classification model robust to the proposed image classification network by applying the weighted adversarial learning method. Weight adversarial learning is a method of securing robust performance in image classification of various statuses while disturbing the model by forcibly reflecting the slope value in reverse when updating the network through the reverse gradient layer. In the experiment, model performance is evaluated through the collected drone flight data set. Youlkyeong Lee, Jehwan Choi, Kang-Hyun Jo |
IECON | 3 |
| 2024 | Simple Human Fall Surveillance System Based on Person DetectionabstractHuman fall is a common problem that often occurs with the elderly, disabled people, and people with bone diseases and neurological diseases. Sometimes, it also comes from human carelessness. Detecting and warning of human falls can minimize the unfortunate risks. Therefore, human fall detection has been widely applied in medical care and surveillance systems. This paper proposes a simple human fall surveillance system based on a person detection network. This system utilizes the pre-trained YOLOv8 network architecture with a related person body dataset. The proposed system reduces the computational complexity and simplifies the use of available datasets for building a surveillance system. As a result, the proposed system achieves the best speed at 206 Frames per second (FPS) when testing on a GeForce GTX 1080Ti 11GB GPU. Duy-Linh Nguyen, Xuan-Thuy Vo, Adri Priadana, Duc-Vuong Nguyen, Thi-Le-Hang Nguyen, Kang-Hyun Jo |
IECON | 6 |
| 2024 | Wider Neighborhood-Aware Attention in Improving YOLOv8n for One-Stage Human Fall DetectionabstractHuman fall detection has become a crucial technology in bolstering intelligent surveillance systems. A one-stage human fall detection model based on the YOLO network emerges as an ideal solution for implementation in limited resource environments, supporting real-time operation with faster speed. This work introduces a Wider Neighborhood-Aware Attention (WN2A) module to enhance YOLOv8n performance for one-stage human fall detection on a CPU device. WN2A enables the YOLOv8n network to focus on crucial information within the feature map based on the channel while considering a wider neighborhood area from a spatial point of view. As a result, the proposed WN2A applied on the YOLOv8n network outperforms the other methods based on the mean Average Precision (mAP) of two benchmark datasets. Moreover, the improved YOLOv8n network enables operating at 27.38 frames per second on an Intel Core i7-9750H CPU while providing higher mAP. Adri Priadana, Duy-Linh Nguyen, Xuan-Thuy Vo, Jehwan Choi, Kang-Hyun Jo |
IECON | 5 |
| 2024 | Lightweight CNN-Based Driver Eye Status Surveillance for Smart VehiclesabstractTraffic accidents are the leading death rate among accident categories. One of the major causes of road traffic accidents is driver drowsiness. Many studies have paid attention to this issue and developed driver assistance tools to reduce the risk. These methods mainly analyze driver behavior, vehicle behavior, and driver physiology. This article proposes a driver eye status surveillance system based on lightweight convolutional neural networks (CNNs). The overall system consists of the following three stages: Face detection, eye detection, and eye classification. In the first stage, the system utilizes a small real-time face detector, named nano YOLO5Face. The second stage focuses on exploiting the compact CNN network architecture combined with the inception network, and triplet attention mechanism. Finally, the system uses a simple classification network architecture to classify open or closed eye status. Additionally, this work also provides the datasets for the eye detection task comprised of 10 659 images and 21 318 labels. As a result, the real-time testing reached 33.12 frames per second (FPS) and 25.11 FPS on an Intel Core I7-4770 CPU @ 3.40 GHz [personal computer (PC)] and a 128-core Nvidia Maxwell GPU (Jetson Nano device), respectively. Duy-Linh Nguyen, Muhamad Dwisnanto Putro, Kang-Hyun Jo |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | VSNet: Vehicle State Classification for Drone Image with Mosaic Augmentation and Soft-Label Assignment
Youlkyeong Lee, Jehwan Choi, Kang-Hyun Jo |
ACIIDS (1) | 3 |
| 2023 | YOLOv5 with Combination of Coordinate Attention and CBAM for Object Detection on DroneabstractObject detection is an important study in computer vision to discriminate the position and class of an object in an image. Object detection in drone images is a technology that automatically detects and classifies objects using deep learning algorithms in flight images taken by drones. Object detection using drone images can rescue human life in disaster situations, grasp the situation at the disaster site, and identify the growth status of crops or pests in agriculture. In addition, it can be used in various fields such as infrastructure management, roads and railways, and city planning. A quick calculation is required. Although rapid computation is possible due to recent hardware development, there are many difficulties in using GPUs in industrial settings. In order to utilize drones in industrial sites, an object detection algorithm capable of real-time operation in a low-cost device is required. In this paper, we propose YOLOv5 with the combination of Coordinate Attention and CBAM for Object Detection on Drone for an algorithm capable of real-time operation in a low-cost device. The proposed architecture makes the model lighter by reducing the number of parameters and improves the object detection rate of the model through Coordinate Attention and CBAM. The model is trained using the VisDrone dataset, and the object detection rate, mAP, increased by about 10% to 22.2mAP, and the number of parameters decreased by about 70% to 2,147,589. Jinsu An, Muhamad Dwisnanto Putro, Adri Priadana, Youlkyeong Lee, Junmyeong Kim, Kang-Hyun Jo |
IECON | 6 |
| 2023 | CSA: Channel-Wise Similarity Attention for Vehicle State ClassificationabstractDeveloped for specific missions, CNNs have gradually improved the performance of object classification networks by using various architectures. The weight of the convolutional layer is a crucial factor in feature extraction. However, as the number of layers increases, performance degradation can occur due to problems such as the vanishing gradient. To overcome this problem, networks have evolved to continuously incorporate information from previous feature maps using various attention mechanisms. In this study, a Channel-wise Similarity Attention (CSA) method is proposed to measure the similarity of feature maps between channels and enhance positive information by highlighting it. Additionally, a deformable convolutional kernel is embedded to apply a flexible receptive field around the object area in the image, replacing the fixed receptive field of the conventional CNN layer. The network is trained end-to-end to classify the condition of vehicles on the road using collected drone flight images. The proposed model achieves an accuracy of 86.13% and 302 frames per second with a number of parameters of 1,273,504. Youlkyeong Lee, Jehwan Choi, Jinsu An, Kang-Hyun Jo |
IECON | 4 |
| 2023 | Vehicle Detector Based on Improved YOLOv5 Architecture for Traffic Management and Control SystemsabstractVehicle detection is an important module in traffic management and control systems. These systems require compactness, mobility, and high accuracy when deployed in a real-time context. Based on the YOLOv5 network architecture, this paper proposes several improvements to increase the performance and speed of the network when applied to vehicle detection. The research aims to redesign the backbone and neck modules with lightweight convolutional network architectures such as EfficientNet, PP-LCNet, and MobileNet. In addition, the Squeeze-and-Excitation (SE) attention architecture is also used inside the above-mentioned architectures to help the network focus on salient information during feature extraction. The network is trained and evaluated on a modified and normalized dataset of the UA-DETRAC dataset. As a result, the proposed network achieves 58.1% of [email protected] and 40.1% of [email protected]:0.95 with just over ten million network parameters. This result outperforms other methods and is comparable to the lightweight architectures of the YOLOv5 family. Duy-Linh Nguyen, Xuan-Thuy Vo, Adri Priadana, Kang-Hyun Jo |
IECON | 4 |
| 2023 | Facial Attribute Recognition Using Lightweight Multi-Label CNN-Transformer Architecture for Intelligent AdvertisingabstractIn modern cities, intelligent advertising platforms have been widely engaged in public areas. A facial attribute recognition technique is essential to assist these platforms in delivering suitable adverts for each audience. These platforms also require a recognition technology that can operate at least suitably on a CPU device to reduce implementation costs. This work proposed a lightweight multi-label CNN-Transformer architecture with an efficient inception block (EIB) and squeeze channel transformer encoder (SCTE) to perform facial attribute recognition efficiently. EIB is used to extract face features in multi-scale and levels supported by SCTE in improving its feature map's quality. The proposed architecture produces fewer parameters with low operations and gains competitive accuracy on the CelebA and LWFA datasets consisting of images with multi-label. Moreover, the proposed architecture integrated with face detection can perform sufficiently on a CPU configuration in real-time with 21 frames per second (FPS) using 224 × 224 input size of face area image. Adri Priadana, Muhamad Dwisnanto Putro, Jinsu An, Duy-Linh Nguyen, Xuan-Thuy Vo, Kang-Hyun Jo |
IECON | 6 |
| 2023 | Recent advances in automatic feature detection and classification of fruits including with a special emphasis on Watermelon (Citrillus lanatus): A review
Danilo Cáceres Hernández, Ricardo Gutierrez, Kelvin Kung, Juan Rodriguez, Oscar Lao, Kenji Contreras, Kang-Hyun Jo, Javier E. Sánchez-Galán |
Neurocomputing | 7 |
| 2022 | Fine-Grained Soccer Actions Classification Using Deep Neural NetworkabstractAction recognition from video data modal is contemplation about recognizing features in spatial position and characterizing the temporal changes across these extracted spatial features. Soccer video action recognition is a categorized field of video action recognition. Soccer sports have enormous viewers across the globe, and there is a significant commercial impact on the broadcasters. High inter-class similarities of various soccer actions symbolize the fine-grained nature of action recognition. Soccer matches are broadcasted live and captured using shifting cameras rather than stationary cameras. The fine-grained nature of soccer actions and shifting camera possessions make the recognition more intricate than the coarse-grained. There exist only limited researches that address this issue. To this extent, we propose a deep learning approach that successfully categorizes ten different soccer actions from our custom-developed SoccerAct10 dataset. Feature extraction is accomplished utilizing transfer learning from state-of-the-art convolutional neural networks (CNN) models such as DenseNet201, InceptionResNetV2, MobileNetV2, ResNet152V2, and Xception. All these models were trained on a massive ImageNet dataset. Long short-term memory (LSTM), with the input features from CNN, models the temporal changes of soccer actions. LSTM has already proven its success in tackling the vanishing gradient. At the final layer, the softmax activation function yields the distributions of probabilities of each soccer action. Empirical evaluation demystifies the effectiveness of our proposed approach, distinguishing ten distinct soccer actions with 90% accuracy. Anik Sen, Syed Mohammad Minhaz Hossain, RussoMohammadAshraf Uddin, Kaushik Deb, Kang-Hyun Jo |
HSI | 5 |
| 2022 | Person Search via Background and Foreground Contrastive LearningabstractThe specific person search is the foundation of a wide range of applications in intelligent security and surveillance systems. Although detection and re-id have been widely studied, they are difficult to apply to practical applications directly. Therefore, this paper focuses on person search, which aims to solve person detection and person re-identification (re-id) jointly. The common practice is to append the standard detection loss and re-id branches parallelly on Faster RCNN. The traditional re-id utilized Online Instance Matching (OIM) to pull a sample closer to its identity class. However, the relationship among RoIs of an image has not been fully explored in previous methods. To address this issue, we propose Background and Foreground Contrastive Loss (BFCL) to further boost re-id performance. We consider that RoIs from one image have a high probability of containing similar patterns, which might disturb the re-id performance. Therefore, we proposed BFCL to strengthen the learning of distinguishing similar background and foreground by leveraging inter-RoIs pairwise similarity. In summary, our method jointly optimizes the regression loss, classification loss, re-id loss, and the proposed BFCL for achieving optimal performances in person search model. Experiments are performed on two large-scale person search datasets, CUHK-SYSU and PRW. Results show that the proposed BFCL consistently boosts the performance of the baseline framework SeqNet in two datasets. The improved results demonstrate the effectiveness of the proposed BFCL and the necessity of exploring the relationship among RoIs. Qing Tang 0004, Kang-Hyun Jo |
HSI | 2 |
| 2022 | Low Computational Vehicle Re-Identification for Unlabeled Drone Flight ImagesabstractRecently advanced vehicle re-identification frameworks are mainly based on convolutional neural networks (CNN) and labeled information. Previous frameworks face two difficulties. First CNN includes complicated architectures, which require expensive GPU devices to perform computation. The second difficulty is that annotating vehicle identities for every frame is expensive and time-consuming. To tackle these two difficulties, this study proposes a simple but effective method to perform re-ID without CNN and labeled identities. The proposed method has two streams of vehicle re-identification. The object detector takes charge of detecting vehicles on the road. With the position of vehicles in the image, the condition module extracts the vehicle movement information and sets the condition to match the same vehicle between current and subsequent frames. To train the object detector and test the proposed algorithm, a set of drone flight images collect and annotate for studying the traffic road. It contains 9,776 train images and 2,200 test images for object detection. In the experiments, three different traffic video clips were applied for testing the proposed method. Youlkyeong Lee, Qing Tang 0004, Jehwan Choi, Kang-Hyun Jo |
IECON | 4 |
| 2022 | Unsupervised Object Re-identification via Instances Correlation LossabstractThis paper studies the fully unsupervised object re-identification (re-ID) problem which can learn re-ID without any human-annotated labeled data. Recent works show that self-supervised momentum contrastive learning is an effective method for unsupervised object re-ID, but they neglect to optimize one important component - the similarity relationships among instances. Previous works focus on enforcing instance-to-centroid learning, which does not fully utilize the inter-instances information. Thus, we propose an Instances Correlation Loss (ICL) to enforce instance-to-instance learning in each training iteration. Experimental results show that the proposed ICL effectively boost the performance, which demonstrates that learning strategy is also a central importance to unsupervised re-ID task. Extensive experiments are performed on three mainstream person re-ID datasets and one vehicle re-ID dataset. Qing Tang 0004, Kang-Hyun Jo |
INDIN | 2 |
| 2022 | Multi-level Feature Reweighting and Fusion for Instance SegmentationabstractAccurate instance segmentation requires high-resolution features for performing a dense pixel-wise prediction task. However, using high-resolution feature maps results in highly expensive model complexity and ineffective receptive fields. To overcome the problems of high-resolution features, conventional methods explore multi-level feature fusion that exchanges the information between low-level features at earlier layers and high-level features at top layers. Both low and high information is extracted by the hierarchical backbone network where high-level features contain more semantic cues and low-level features encompass more specific patterns. Thus, adopting these features to the training segmentation model is necessary, and designing a more efficient multi-level feature fusion is crucial. Existing methods balance such information by using top-down and bottom-up pathway connections with more inefficient convolution layers to produce richer multi-scale features. In this work, we contribute two folds: (1) a simple but effective multilevel feature reweighting layer is proposed to strengthen deep high-level features based on channel reweighting generated from multiple features of the backbone, and (2) an efficient fusion block is proposed to process low-resolution features in a depth-to-spatial manner and combine enhanced multi-level features together. These designs enable the segmentation models to predict instance kernels for mask generation on high-level feature maps. To verify the effectiveness of the proposed method, we conduct experiments on the challenging benchmark dataset MS-COCO. Surprisingly, our simple network outperforms the baseline in both accuracy and inference speed. More specifically, we achieve 35.4% APmaskat 19.5 FPS on a GPU device, becoming a state-of-the-art instance segmentation method. Xuan-Thuy Vo, Tien-Dat Tran, Duy-Linh Nguyen, Kang-Hyun Jo |
INDIN | 4 |
| 2022 | A review on anchor assignment and sampling heuristics in deep learning-based object detection
Xuan-Thuy Vo, Kang-Hyun Jo |
Neurocomputing | 2 |
| 2022 | Deep learning-based perception systems for autonomous driving: A comprehensive survey
Li-Hua Wen, Kang-Hyun Jo |
Neurocomputing | 2 |
| 2022 | A Fast CPU Real-Time Facial Expression Detector Using Sequential Attention Network for Human-Robot InteractionabstractFacial expression detection is a method to predict human facial emotions. This work is a trending research topic that can be implemented for human-robot interaction. More recently, deep convolutional neural network provides a robust extractor features but tends to be slow in real-time implementations and often requires a large memory and graphics processing units for fast execution. In this article, an efficient CPU-based facial expression detector is proposed using a sequential attention network to improve the baseline performance. The proposed attention network consists of three modules, global representation to capture the global features, channel representation, and dimension representation, which are focused on the channel and using spatial attention to discriminate local features. The efficient partial transfer module is also presented as a light backbone to extract facial features from an image. The entire module is trained and tested on several benchmarks to classify seven facial expressions. As a result, the proposed model reaches an accuracy of 98.18%, 98.75%, 95.63%, and 74.17% on CK+, JAFFE, KDEF, and FER-2013, respectively. It achieves competitive performance when compared to state-of-the-art methods. Lastly, it is integrated with a face detector and runs in real-time without a constraint at 69 frames per second on a CPU. Muhamad Dwisnanto Putro, Duy-Linh Nguyen, Kang-Hyun Jo |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | Accurate Bounding Box Prediction for Single-Shot Object DetectionabstractAccurate single-shot object detection is an extremely challenging task in real environments because of complex scenes, occlusion, ambiguities, blur, and shadow, i.e., these factors are called uncertainty problem. It leads to unreliable labeling of bounding box annotation and makes detectors arduous to learn bounding box localization. Previous methods viewed the ground truth box coordinates as a rigid distribution omitting localization uncertainty in real datasets. This article proposes a novel bounding box encoding algorithm integrated into the single-shot detector (BBENet) to consider the flexible distribution of bounding box localization. First, discretized ground truth labels are generated by decomposing each object’s boundary into multiple boundaries. The new representation of ground truth boxes is more arbitrary and flexible to cover any case of complex scenes. During training, the detector directly learns discretized box locations instead of continuous domain. Second, the bounding box encoding algorithm reorganizes bounding box predictions to be more accurate. Furthermore, another problem in existing methods is inconsistency in estimating detection quality. The single-shot detection consists of classification and localization tasks, but the popular detectors consider the classification score as the final detection quality. Thus, it lacks localization quality and hinders the overall performance because both tasks have a positive correlation. To overcome this problem, BBENet introduces detection quality by combining the localization and classification quality to rank detection during nonmaximum suppression. The localization quality is computed based on how uncertain the predicted boxes are, which is a new perspective in detection literature. The proposed BBENet is evaluated on three benchmark datasets, i.e., MS-COCO, Pascal VOC, and CrowdHuman. Without bells and whistles, BBENet outperforms the existing methods by a large margin with comparable speed, achieving the state-of-the-art single-shot detector. Xuan-Thuy Vo, Kang-Hyun Jo |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | Eye State Recognizer Using Light-Weight Architecture for Drowsiness Warning
Duy-Linh Nguyen, Muhamad Dwisnanto Putro, Kang-Hyun Jo |
ACIIDS | 3 |
| 2021 | Real-Time Multi-view Face Mask Detector on Edge Device for Supporting Service Robots in the COVID-19 Pandemic
Muhamad Dwisnanto Putro, Duy-Linh Nguyen, Kang-Hyun Jo |
ACIIDS | 3 |
| 2021 | Practical Analysis on Architecture of EfficientNetabstractConvolutional neural networks (CNNs) are now used in a variety of computer vision applications. However, it is quite hard to adopt them in real-time system due to the problem of increasing model size. Recently, some efficient networks which still have acceptable performance are proposed. Among them, EfficientNet is one of the state-of-the-art architectures. It can be considered a family of network models. EfficientNet could take its place among the state-of-the-art on the ImageNet challenge while still have much fewer parameters and computation cost. But given some of its subtleties, it is more efficient than most of its predecessors. It uses the inverted bottleneck residual blocks of MobileNetV2, in addition to squeeze-and-excitation modules (SE modules). This paper investigates the effect of SE modules on the performance of EfficientNet-B0, the fundamental network model in its family, by repositioning/removing the SE modules. Van-Thanh Hoang, Kang-Hyun Jo |
HSI | 2 |
| 2021 | Efficient Face Detector Using Spatial Attention Module in Real-Time Application on an Edge Device
Muhamad Dwisnanto Putro, Duy-Linh Nguyen, Kang-Hyun Jo |
ICIC (1) | 3 |
| 2021 | Recognition of Multiple Panamanian Watermelon Varieties Based on Feature Extraction Analysis
Javier E. Sánchez-Galán, Anel Henry, Fatima Rangel, Emmy Sáez, Kang-Hyun Jo, Danilo Cáceres Hernández |
ICIC (2) | 5 |
| 2021 | Regression-Aware Classification Feature for Pedestrian Detection and Tracking in Video Surveillance Systems
Xuan-Thuy Vo, Tien-Dat Tran, Duy-Linh Nguyen, Kang-Hyun Jo |
ICIC (1) | 4 |
| 2021 | A Lightweight Attention Fusion Module for Multi-sensor 3-D Object Detection
Li-Hua Wen, Ting-Yue Xu, Kang-Hyun Jo |
ICIC (1) | 3 |
| 2021 | Unsupervised Person Re-Identification Via Nearest Neighbor Collaborative Training StrategyabstractBecause of the lack of human-labeled data, the challenge of unsupervised person re-identification (re-ID) is to learn to generate correct pseudo labels for training. Unlike the human-labeled annotation, the generated pseudo labels contain the noise labels that harm the model’s performance. In this paper, we propose the Nearest Neighbors Collaborative Training (NNCT) strategy to mitigate the effects of noisy labels by utilizing information of the nearest neighbor of an image. The proposed NNCT trains the image and its nearest neighbor collaboratively, thereby enhancing the generalization capability of the network and shortening the distance with neighbors. To make training using the up-to-date nearest neighbor possible, we introduce a Pseudo Label Memory Bank (PLMB) to store the up-to-date labels of all images. The experimental results confirm the superiority of the proposed method, which surpasses state-of-the-arts on two mainstream person re-ID datasets, Market-1501, and DukeMTMC-reID in both fully unsupervised learning manner and Unsupervised Domain Adaptation (UDA) manner. Qing Tang 0004, Kang-Hyun Jo |
ICIP | 2 |
| 2021 | Attention based Object Classification for Drone ImageryabstractThis paper shows how to make the drone imagery for surveillance or tracking the object in the ground. To detect or classify objects on the ground, convolutional neural networks was adopted and compared with some existed methods and the proposed attention blocks in it. The objects on the ground from the drone images are relatively very small and diversity of the appearance from its perspective projections. This is mainly due to the arbitrary viewpoints from the bird eye views. Furthermore, the distance from its viewpoint in the sky is quite much changeable so that the image of the object is too diverse in appearance and its size. However, the drone is so useful to see widely while navigating in the sky. It is much more attentive to use for real application. Here, some proposed target objects are mainly located in the ground, like static and dynamic objects such as street lamps or trees, vehicles, trucks and pedestrians. These works were done for the national projects to establish the general AI services in Korea recently. For the experiments such as buildup the ground truth of target objects after taken in regulated distance and viewing angles and performed to detect exactly objects in an arbitrary image. For experiments of detection and classification of five categories of objects, attention based CNN architecture was adopted and compared comprehensively with the existed networks like MobileNet, VGG16, SqueezeNet, and ResNet. The experimental results outperformed for the archived drone image dataset with 87.12% in precision. The architecture shows almost 3 times faster with respect to VGG16 or 2 times faster than MobileNet in the speed but a half slimer and twice thicker respectively in the number of parameters. Thus, the Attention Block is useful while a drone navigates through a certain route according to the ground location regardless of the appearance and size of the target region in image. Jehwan Choi, Kang-Hyun Jo |
IECON | 2 |
| 2021 | Light-weight Convolutional Neural Network for Distracted Driver ClassificationabstractDriving is an activity that requires the coordination of many senses with complex manipulations. However, the driver can be affected by a several factors such as using a mobile phone, adjusting audio equipment, smoking, drinking, eating, talking to a passenger or drowsy. Therefore, the development of assistant applications to warn distracted driver is very necessary. Because of the limited space and mobility, the equipment also requires compact, energy-saving and efficient. This paper proposes a lightweight Convolutional Neural Network for a distracted driver warning system. The method is built based on a combination of standard convolution and Depthwise Separable Convolution operation to optimize the network parameters but still ensure the important information and speed. The network was trained and evaluated on two datasets, AUC (the American University in Cairo) and StateFarm dataset from Kaggle’s competition. As a result, the evaluation accuracy reached 95.36% and 99.95%, respectively. Duy-Linh Nguyen, Muhamad Dwisnanto Putro, Xuan-Thuy Vo, Kang-Hyun Jo |
IECON | 4 |
| 2021 | Dynamic Multi-Loss Weighting for Multiple People Tracking in Video Surveillance SystemsabstractMultiple people tracking is a fundamental yet challenging task in the computer vision field, which served as a primary process for high-level tasks such as human behaviors, action recognition, pose estimation. Person tracking is decomposed into detection and re-identification (re-ID) sub-tasks. Conventionally, the detection learns classification and regression objectives simultaneously; and the re-ID sub-task is treated as a classification task. Therefore, person tracking is multiple task learning corresponding to multiple loss functions (multiple objectives) with one bounding box regression and two classifications. The difference between various tasks is as follows: the ranges of each objective are inconsistent, the contribution of each task to the overall gradient is altered, and the learning pace of each task is different (level of difficulty). It leads to an objective imbalance in multi-task learning. Previous methods proposed weighting factors as new hyper-parameters to balance the ranges of each task. The dimension of search space for manually tuning these hyper-parameters is high, which depends on the number of tasks. Accordingly, selecting reasonable weighting factors is difficult and complicated. This paper introduces dynamic multi-loss weighting (DMW) with simple but effective in which the weighting factors are dynamically changed during training without introducing any hyper-parameters. The dynamic weights are optimized to balance regression and classification objectives, which depend on the difficulty level of each task and the correlation between each task. Additionally, the general convolution operations are spatially invariant to some degree, which hinders the network’s performance. Hence, this work employs the position-sensitive operation improving feature extraction. The proposed method is conducted on the MOT17 challenging benchmark, which outperforms the online multiple people trackers without using additional data. Xuan-Thuy Vo, Tien-Dat Tran, Duy-Linh Nguyen, Kang-Hyun Jo |
INDIN | 4 |
| 2021 | 3-D Facial Landmarks Detection for Intelligent Video SystemsabstractFacial landmark detection is a fundamental research topic in computer vision that is widely adopted in many applications. Recently, thanks to the development of convolutional neural networks, this topic has been largely improved. This article proposes facial-landmark detector, which is based on a state-of-the-art architecture for landmark localization called stacked hourglass network, to obtain accurate facial landmark-points. More specifically, this article uses residual networks as the backbone instead of a 7 × 7 convolution layer. Additionally, it modifies the hourglass modules by using the residual-dense blocks in the mainstream for capturing more efficient features and the 1 × 1 convolution layers in the branch streams for reducing the model size and computational time, instead of the original residual blocks. The proposed architecture also enhances the features from modified hourglass modules with finer-resolution features via a lateral connection to generate more accurate results. The proposed network can outperform other state-of-the-art methods on the AFLW2000-3D dataset and the LS3D-W dataset, the largest three-dimensional (3-D face) alignment dataset to date. Van-Thanh Hoang, De-Shuang Huang, Kang-Hyun Jo |
IEEE Trans. Ind. Informatics | 3 |
| 2021 | High Performance and Efficient Real-Time Face Detector on Central Processing Unit Based on Convolutional Neural NetworkabstractFace detection is crucial in the development of face recognition, expression, tracking, and classification. Conventional methods have accuracy constraints on several challenging conditions, including nonfrontal faces, occlusions, and complex backgrounds. However, the convolutional neural network (CNN) methods produce high performances despite a large amount of computation. Therefore, CNN requires expensive hardware and is not suitable for low-cost central processing units (CPUs). This article develops a light architecture for a CNN-based real-time face detector. The proposed architecture consists of two main modules, the backbone to extract distinctive facial features and multilevel detection to perform prediction at multiple scales. Furthermore, it utilizes several approaches to enhance the training result, including balancing loss and tweaks on the training configuration. The proposed detector has one stage and is trained using the input of images from WIDER FACE with challenges, which contains more challenging images than other datasets. As a result, the detector achieves state-of-the-art performance on several benchmark datasets compared with the other CPU-based models. Then, its efficiency is superior to that of competitors, as it runs at 53 frames per second on a CPU for video graphics array resolution images. Muhamad Dwisnanto Putro, Laksono Kurnianggoro, Kang-Hyun Jo |
IEEE Trans. Ind. Informatics | 3 |
| 2021 | Deep Atrous Spatial Features-Based Supervised Foreground Detection Algorithm for Industrial Surveillance SystemsabstractCamera-based surveillance systems largely perform an intrusion detection task for sensitive areas. The task may seem trivial but is quite challenging due to environmental changes and object behaviors such as those due to night-time, sunlight, IR camera, camouflage, and static foreground objects, etc. Convolutional neural network based algorithms have shown promise in dealing with these challenges. However, they are exclusively focused on accuracy. This article proposes an efficient supervised foreground detection (SFDNet) algorithm based on atrous deep spatial features. The features are extracted using atrous convolution kernels to enlarge the field-of-view of a kernel mask, thereby encoding rich context features without increasing the number of parameters. The network further benefits from a residual dense block strategy that mixes the mid and high-level features to retain the foreground information lost in low-resolution high-level features. The extracted features are expanded using a novel pyramid upsampling network. The feature maps are upsampled using bilinear interpolation and pass through a 3x3 convolutional kernel. The expanded feature maps are concatenated with the corresponding mid and low-level feature maps from an atrous feature extractor to further refine the expanded feature maps. The SFDNet showed better performance than high-ranked foreground detection algorithms on the three standard databases. The testing demo can be found at https://drive.google.com/file/d/1z_zEj9Yp7GZeM2gSIwYKvSzQlxMAiarw/view?usp=sharing. Ajmal Shahbaz, Kang-Hyun Jo |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | Three-Attention Mechanisms for One-Stage 3-D Object Detection Based on LiDAR and CameraabstractThis article studies one-stage 3-D object detection based on light detection and ranging (LiDAR) point clouds and red-green-blue (RGB) images that aims to boost 3-D object detection accuracy based on three attention mechanisms. Currently, most of the previous works converted LiDAR point clouds into bird's-eye-view (BEV) images, achieving a significant performance. However, they still have a problem due to partial height information (z-axis value) loss during the conversion. To eliminate this problem, the height information of the LiDAR point clouds is projected onto an RGB image and embedded into the original RGB image to generate a new image, named RGBD. This is the first attention mechanism to improve 3-D detection accuracy. Moreover, two other attention mechanisms extract more discriminative global and local features, respectively. Specifically, the global attention network is appended to a feature encoder, and the local attention network is used for the view-specific region of interest fusion. Massive experiments evaluated on the KITTI benchmark suite show that the proposed approach outperforms state-of-the-art LiDAR-Camera-based methods on the car class (easy, moderate, hard): 2-D (90.35%, 88.47%, 86.98%), 3-D (85.12%, 76.23%, 74.46%), and BEV (89.64%, 86.23%, 85.60%). Li-Hua Wen, Kang-Hyun Jo |
IEEE Trans. Ind. Informatics | 2 |
| 2020 | Slice Operator for Efficient Convolutional Neural Network Architecture
Van-Thanh Hoang, Kang-Hyun Jo |
ACIIDS (2) | 2 |
| 2020 | Lightweight Convolutional Neural Network for Real-Time Face Detector on CPU Supporting Interaction of Service RobotabstractFace detection plays an essential role in the success of the interaction between service robots and consumers. This method is the initial stage for face-related applications. Practical applications require face detection to work in real-time and can be implemented on low-cost devices such as CPU. Traditional methods have problems when the face is not frontal, blocked, and partially covered, but real-time speed is not an obstacle. On the other hand, deep learning has succeeded in accurately distinguishing facial features and backgrounds. Face sizes that tend to be medium and large when robot interaction with consumers so it can employ Convolutional Neural Networks (CNN) with light weights. In this paper, a real-time face detector is built that can work on the CPU. This detector will be implemented explicitly in service robots to support interactions with consumers. It can overcome the occlusion and not-frontal face. Detector architecture consists of the backbone as rapidly features extractor, transition module as a transformer of prediction map, and the dual-detection layer is head of a network prediction based on scale assignment. As a result, the detector can work at speeds of 301 frames per second on CPU without ignoring the accuracy. Muhamad Dwisnanto Putro, Duy-Linh Nguyen, Kang-Hyun Jo |
HSI | 3 |
| 2020 | Enhanced Feature Pyramid Networks by Feature Aggregation Module and Refinement ModuleabstractFeature pyramids executing refinements on the raw feature maps produced by the backbone (e.g., ResNet, VGG) are universally employed in object detection tasks (e.g., Faster R-CNN, Mask R-CNN, YOLO, SSD, RetinaNet) to mitigate scale variation problem. Although these object detections with feature pyramids accomplish a boost in accuracy without compromising speed, they have some limitations since that they only naturally design the feature pyramid with consecutive scales, the pyramidal architecture of the backbone, which are initially constructed for the classification task. This problem leads to the feature imbalance between high-level features and low-level features in object detection. In this work, the proposed method introduces Feature Aggregation Module (FAM) and Refinement Module (RM) to obtain more powerful feature pyramids for predicting objects of different scales. First, the multi-level feature maps (i.e., multiple layers) extracted by the backbone network are aggregated as the basic feature. Second, the basic feature is enhanced by a refinement module exploiting long-range dependency. Three, to construct a feature pyramid for object detection, the proposed FAM is used by converting the basic feature (after utilizing a refinement module) into multi-level features. Finally, refined multi-level features and raw features generated by the backbone could be enhanced through shortcut connections to capture more representative. To perform the efficiency, the proposed method integrates the FAM and the RAM into the architecture of Faster R-CNN called EFPN Faster R-CNN. Especially on the MS-COCO dataset, EFPN Faster R-CNN achieves 2.2 points higher Average Precision (AP) than FPN Faster R-CNN without bells and whistles. Xuan-Thuy Vo, Kang-Hyun Jo |
HSI | 2 |
| 2020 | Bidirectional Non-local Networks for Object Detection
Xuan-Thuy Vo, Li-Hua Wen, Tien-Dat Tran, Kang-Hyun Jo |
ICCCI | 4 |
| 2020 | Aggregated Deep Saliency Prediction by Self-attention Network
Ge Cao, Qing Tang 0004, Kang-Hyun Jo |
ICIC (3) | 3 |
| 2020 | Accurate and Efficient Traffic Sign Detection with a Guided Region Enlarging Algorithm
Qing Tang 0004, Ge Cao, Kang-Hyun Jo |
ICIC (1) | 3 |
| 2020 | LiDAR-Camera-Based Deep Dense Fusion for Robust 3D Object Detection
Li-Hua Wen, Kang-Hyun Jo |
ICIC (3) | 2 |
| 2020 | Eyes Status Detector Based on Light-weight Convolutional Neural Networks supporting for Drowsiness Detection SystemabstractThe drowsiness is the leading cause of many accidents on the road. These causes can be reduced by using the drowsiness alarm or drowsiness detection system. These systems monitor drivers while driving and alarm when they don't focus or have some abnormal signs in the driver's body. Currently, most methodologies use the analysis of human behaviors, vehicle behaviors, and human physiological conditions. This paper regards eyes status analysis based on deep learning method using proposed Convolutional Neural Networks (CNN) with two stages are face detection and eyes classification. The face detector employs a single detector module and shallow layer, then the eyes classifier using simple CNN without ignoring the accuracy. As a result, the average speed was tested in real-time by 50.03 fps (frames per second) on Intel Core I7-4770 CPU @ 3.40 GHz. Duy-Linh Nguyen, Muhamad Dwisnanto Putro, Kang-Hyun Jo |
IECON | 3 |
| 2020 | A Dual Attention Module for Real-time Facial Expression RecognitionabstractIn this paper, a real-time face expression based on a Convolutional Neural Network with a Dual Attention Module is presented for classifying various facial emotions. The system contains two main components. Firstly, local convolutional features of faces are extracted by the VGG13 baseline. Secondly, dual attention masks are automatically employed to enhance the backbone end based on the global probability of features. It represents the position and channel of the feature map. The local features from baseline are combined with the attention to infer the emotional label module. A single network is trained in an end-to-end scheme with five million parameters. Experiments on benchmark datasets show the attention module gives increased accuracy. Besides, this module provides lightweight and runs 60.20 frames per second when working in real-time on CPU devices. Muhamad Dwisnanto Putro, Duy-Linh Nguyen, Kang-Hyun Jo |
IECON | 3 |
| 2019 | Modified Stacked Hourglass Networks for Facial Landmarks Detection
Van-Thanh Hoang, Kang-Hyun Jo |
ACIIDS (2) | 2 |
| 2019 | Ensemble of Predictions from Augmented Input as Adversarial Defense for Face Verification System
Laksono Kurnianggoro, Kang-Hyun Jo |
ACIIDS (2) | 2 |
| 2019 | Real-Time Multiple Faces Tracking with Moving Camera for Support Service Robot
Muhamad Dwisnanto Putro, Kang-Hyun Jo |
ACIIDS (2) | 2 |
| 2019 | Practical Analysis on Running Speed for Efficient CNN Architecture DesignabstractConvolutional neural networks (CNNs) have shown significant performance in solving various artificial intelligence tasks in recent years. However, the increasing model size has raised challenges in adopting them in limited-resource applications. Recently, many research works try to build efficient networks which are as small as possible and have small computation time while still have acceptable performance. The state-of-the-art architectures are ShuffleNetV2 and MobileNet. They use Depthwise Separable Convolution (DWConvolution) in place of standard Convolution to reduce the model size. Their designs are very efficient which follows many practical guild-lines. However, the ShuffleNetV2 has higher memory access cost (MAC) than the MobileNet. This paper evaluates these two networks on GPU and CPU with and without tuning step to show the advantages of each network. The experiments show that the MobileNet is faster on high-computational-optimization devices like GPU, whereas the ShuffleNetV2 is faster on low-computational-optimization devices like CPU when they have similar FLOPs. Van-Thanh Hoang, Kang-Hyun Jo |
HSI | 2 |
| 2019 | Optimized Latent Features for Deep Image CompressionabstractThe work presented in this paper is focused on deep image compression where a neural network architecture is used to extract the small-sized latent features which encodes the full information of the input image. A common drawback in such system is low quality of the reconstructed image. This paper aims to alleviate this drawback by augmenting the latent feature using an optimization strategy. The values in a given latent features are updated iteratively to produce an output image with lower reconstruction loss. The proposed method has been evaluated MNIST dataset where the results shows that it could provide significant gain compared to the reconstruction results which are generated from the non optimized latent features. Laksono Kurnianggoro, Kang-Hyun Jo |
HSI | 2 |
| 2019 | Design and Implementation of a Smart System for Watermelon RecognitionabstractThe main goal of this paper was to develop a system for watermelon classification using smart systems, using the harvested fruit's external texture pigmentation or pixels for digital treatment. In order to achieve this goal, an algorithm was used in order to process the image of watermelons. Based on digital images, common patterns present in watermelons were compared with the images of watermelons recognition system for the treatment of images. Define the parameters that determine the characteristics that involve the product for the identification of Citrullus Lanatus (watermelon). Anel Henry, Kelvin Kung, Kang-Hyun Jo, Danilo Cáceres Hernández |
HSI | 3 |
| 2019 | Illegally Parked Vehicle Detection Based on Haar-Cascade Classifier
Aapan Mutsuddy, Kaushik Deb, Tahmina Khanam, Kang-Hyun Jo |
ICIC (2) | 4 |
| 2019 | Occluded Object Classification with Assistant Unit
Qing Tang 0004, Youlkyeong Lee, Kang-Hyun Jo |
ICIC (3) | 3 |
| 2019 | Fully Convolutional Neural Networks for 3D Vehicle Detection Based on Point Clouds
Li-Hua Wen, Kang-Hyun Jo |
ICIC (2) | 2 |
| 2019 | Deep Residual Networks with Pyramid Depthwise Separable ConvolutionabstractConvolutional neural networks (CNNs) have shown remarkable performance in various computer vision tasks in recent years. However, the increasing model size has raised challenges in adopting them in real-time applications as well as mobile or embedded vision applications. Many works try to build networks as small as possible while still have acceptable performance by using Depthwise Separable Convolution (DWConvolution) in place of standard Convolution. This paper proposes a network which uses an extended version of DWCon-volution, called Pyramid DW Convolution. Instead of using just a 3 × 3 kernel size for DWConvolution, the proposed network uses a pyramid kernel size to capture more spatial information. The proposed architecture is evaluated on the highly competitive object recognition benchmark datasets ImageNet and the two CIFAR (CIFAR-10, CIFAR-100). The experiments demonstrate that the proposed network achieves better performance compared with other state-of-the-art networks. Van-Thanh Hoang, Ajmal Shahbaz, Kang-Hyun Jo |
IECON | 3 |
| 2019 | Convolutional Neural Network based Foreground Segmentation for Video Surveillance SystemsabstractConvolutional Neural Networks (CNN) have shown astonishing results in the field of computer vision. This paper proposes a foreground segmentation algorithm based on CNN to tackle the practical challenges in the video surveillance system such as illumination changes, dynamic backgrounds, camouflage, and static foreground object, etc. The network is trained using the input of image sequences with respective ground-truth. The algorithm employs a CNN called VGG-16 to extract features from the input. The extracted feature maps are upsampled using a bilinear interpolation. The upsampled feature mask is passed through a sigmoid function and threshold to get the foreground mask. Binary cross entropy is used as the error function to compare the constructed foreground mask with the ground truth. The proposed algorithm was tested on two standard datasets and showed superior performance as compared to the top-ranked foreground segmentation methods. Ajmal Shahbaz, Van-Thanh Hoang, Kang-Hyun Jo |
IECON | 3 |
| 2019 | Attention-Guided Model for Robust Face Detection System
Laksono Kurnianggoro, Kang-Hyun Jo |
PSIVT | 2 |
| 2019 | 3-D Human Pose Estimation Using Cascade of Multiple Neural NetworksabstractEstimating three-dimensional (3-D) human poses from a given two-dimensional (2-D) shape is still an inherently ill-posed problem in computer vision. This paper proposes a method called cascade of multiple neural networks (CMNN) to solve this problem in following two steps: 1) create the initial estimated 3-D shape using the Zhou et al. method with a small number of basis shapes and 2) make this initial shape more alike to the original shape by using the CMNN. In comparing to existing works, the proposed method shows a significant outperformance in both accuracy and processing time. This paper also introduces a new system called Human3D that can estimate the 3-D pose of all people in a single RGB image. This system comprises two part: convolution pose machine (CPM) for estimating 2-D poses of all people in an RGB image and CMNN for reconstructing 3-D poses of them from outputs of the CPM. Van-Thanh Hoang, Kang-Hyun Jo |
IEEE Trans. Ind. Informatics | 2 |
| 2018 | Speed-Up 3D Human Pose Estimation Task Using Sub-spacing Approach
Van-Thanh Hoang, Kang-Hyun Jo |
ACIIDS (2) | 2 |
| 2018 | Stationary Object Detection for Vision-Based Smart Monitoring System
Wahyono, Reza Pulungan, Kang-Hyun Jo |
ACIIDS (2) | 3 |
| 2018 | Towards an Integrated Method of Detection and Description for Face Authentication SystemabstractThe work in this paper aims to construct a face authentication system based on the deep learning. It is consisted of face detection module, face description system, and retrieval method. Neural network is utilized for both of the detection and description modules as an initial attempt to unify both of the system. In this case, the single shot detection network is utilized as the face detector while the descriptor extractor network is trained by triplet embedding loss function. The proposed system was tested on a novel dataset with several identities to evaluate its robustness. Experiment shows that the result is promising and can be used in a new environment with novel faces without re-training the network. Laksono Kurnianggoro, Kang-Hyun Jo |
HSI | 2 |
| 2018 | Evaluation of IEEE 802.11n and IEEE 802.11p Based on Vehicle to Vehicle CommunicationsabstractWireless communication opens up new avenues in the field of intelligent vehicle systems, including monitoring, guidance and warning applications. In order to determine more precisely the feasibility of using the 802.11n standard in vehicular communication systems, this paper describes a process for evaluating the propagation performance of the 802.11n and 802.11p standards. To do that, the antennas are described and simulated using Win Pro Solutions which include Proman, Wallman softwares. The measurements were performed using an urban environment to prove the effectiveness of the simulation analysis. Edgar Ian Murillo, Hector Poveda, Kang-Hyun Jo, Danilo Cáceres Hernández |
HSI | 3 |
| 2018 | Entity Detection for Information Retrieval in Video Streams
Kang-Hyun Jo |
ICIC (3) | 2 |
| 2018 | Multi-Person Pose Estimation with Human Detection: A Parallel ApproachabstractHuman pose estimation is a fundamental research topic in computer vision. This topic has been largely improved recently thanks to the development of convolution neural network. This paper proposes a new CNN architecture which combines a key-points estimator and an object detector. This network can detect poses of all people and the around objects in the image in parallel. In general, to address the multi-person pose estimation, the network generates human key-points and bounding boxes simultaneously. Thus, it ensembles these key-points into full poses of multiple people based on the bounding boxes. Van-Thanh Hoang, Kang-Hyun Jo |
IECON | 2 |
| 2018 | Geometrical Feature Based Stairways Detection and Recognition Using Depth SensorabstractStairways detection and distance measurement have been a continuous challenge of research area in human-system interaction to reach topnotch solution with greater portability in assisting visually impaired people and guiding autonomous navigation system at smart environments in the real world. For that, a framework is proposed in this work to detect the stair region from depth stair image based on a unique geometrical feature of a stair. The unique geometrical feature is every stair step's height gradually decreases from bottom to top of the stair. For that initially, the depth image is preprocessed and extracted the Canny edge image. After that, a proposed edge linking procedure is utilized through the Brute-Force Search technique to improve the broken edges. Furthermore, a non-candidate edge elimination procedure is used to extract the longest potential concurrent horizontal edge segment by considering the orientation of the horizontal edges. Finally, the extracted potential concurrent horizontal edge segment is verified as stair edge segment by justifying the aforementioned unique feature of stair and detects the stair region of interest (ROI). Furthermore, one-dimensional depth feature is extracted from the ROI and sent to the support vector machine (SVM) for recognizing the up, down, and negative stair. The distance of the recognized stair region from the camera is estimated based on the depth feature. Stairs images captured under different lighting conditions have been used to test the proposed framework to evaluate the resultant accuracy of the system. Md. Khaliluzzaman, Kaushik Deb, Kang-Hyun Jo |
IECON | 3 |
| 2018 | A survey of 2D shape representation: Methods, evaluations, and future research directions
Laksono Kurnianggoro, Wahyono, Kang-Hyun Jo |
Neurocomputing | 3 |
| 2018 | Adaptive consensus control of output-constrained second-order nonlinear systems via neurodynamic optimization
Kang-Hyun Jo |
Neurocomputing | 3 |
| 2018 | Fast Smoke Detection for Video Surveillance Using CUDAabstractSmoke detection is a key component of disaster and accident detection. Despite the wide variety of smoke detection methods and sensors that have been proposed, none has been able to maintain a high frame rate while improving detection performance. In this paper, a smoke detection method for surveillance cameras is presented that relies on shape features of smoke regions as well as color information. The method takes advantage of the use of a stationary camera by using a background subtraction method to detect changes in the scene. The color of the smoke is used to assess the probability that pixels in the scene belong to a smoke region. Due to the variable density of the smoke, not all pixels of the actual smoke area appear in the foreground mask. These separate pixels are united by morphological operations and connected-component labeling methods. The existence of a smoke region is confirmed by analyzing the roughness of its boundary. The final step of the algorithm is to check the density of edge pixels within a region. Comparison of objects in the current and previous frames is conducted to distinguish fluid smoke regions from rigid moving objects. Some parts of the algorithm were boosted by means of parallel processing using compute unified device architecture graphics processing unit, thereby enabling fast processing of both low-resolution and high-definition videos. The algorithm was tested on multiple video sequences and demonstrated appropriate processing time for a realistic range of frame sizes. Alexander Filonenko, Danilo Cáceres Hernández, Kang-Hyun Jo |
IEEE Trans. Ind. Informatics | 3 |
| 2017 | Improving Traffic Sign Recognition Using Low Dimensional Features
Laksono Kurnianggoro, Wahyono, Kang-Hyun Jo |
ACIIDS (2) | 3 |
| 2017 | Comparative study of modern convolutional neural networks for smoke detection on image dataabstractThis work evaluates modern convolutional neural networks (CNN) for the task of smoke detection on image data. The networks that were tested are AlexNet, Inception-V3, Inception-V4, ResNet, VGG, and Xception. They all have shown high performance on huge ImageNet dataset, but the possibility of using such CNNs needed to be checked for a very specific task of smoke detection with a high diversity of possible scenarios and a small available dataset. Experimental results have shown that inception-based networks reach high performance when samples in the training dataset cover enough scenarios while accuracy dramatically drops when older networks are utilized. Alexander Filonenko, Laksono Kurnianggoro, Kang-Hyun Jo |
HSI | 3 |
| 2017 | Lane marking detection using image features and line fitting modelabstractThe lane marking detection task is an essential process in the field of semi-autonomous and autonomous navigation. This paper proposes a method that combines the color and edge information to robustly detect the lane marking within the image either located far on near to the vehicle. Firstly, the region of interest is extracted from the image. Secondly, the set of lane marking features are extracted. To do that, the change in color between road and marking surface is used along a probability density function to extract the set of candidates. Finally, a clustering method along a line fitting model is implemented. Preliminary results were performed and tested on a group of consecutive frames to prove the effectiveness of the proposed method. Danilo Cáceres Hernández, Alexander Filonenko, Ajmal Shahbaz, Kang-Hyun Jo |
HSI | 4 |
| 2017 | An improved method for 3D shape estimation using active shape modelabstractThis paper tackles the problem of reconstructing 3D human poses from 2D landmarks, which is still an ill-posed problem. A widely-used approach is active shape model (ASM) which considers an unknown 3D shape as a linear combination of predefined basis shapes. The existing methods often resolve an optimization problem to reckon the weights and viewpoints of basis shapes, but they could fall into a locally-optimal and/or not use in the real-time system. In this paper, we propose an improved method by doing categorize database into subspaces to reduce execution time and make reconstruction accuracy better in four steps: (i) Separating 3D shapes in training database into subspaces based on their features. (ii) Learning predefined basis shapes of each subspace. (iii) Reconstructing 3D human poses from basis shapes of all subspaces. (iv) Picking out the best shape among them as the final result. Van-Thanh Hoang, Kang-Hyun Jo |
HSI | 2 |
| 2017 | Welcome messageabstractWelcome to HSI2017, the 10th International Conference on Human System Interactions in 2017 was held at the University of Ulsan in Ulsan, Republic of Korea. The University of Ulsan have organized the conference and the conference is technically co-sponsored by IEEE Industrial Electronics Society. HSI conference series has been one of the most important academic meetings in the field of interactions between human and systems. Until now the HSI conference series have been held in Krakow (Poland) 2008, Catania (Italy) 2009, Rzeszow (Poland) 2010, Yokohama (Japan) 2011, Perth (Australia) 2012, Gdansk (Poland) 2013, Lisbon (Portugal) 2014, Warsaw (Poland) 2015, and Portsmouth (United Kingdom) 2016. Kang-Hyun Jo, Luís Gomes 0001, Milos Manic, Jacek Ruminski, Young Soo Suh |
HSI | 1 |
| 2017 | Classification of human fall in top Viewed kinect depth images using binary support vector machineabstractVision based human fall action classification from non fall has been given significant importance over the past decade since the rise of falling events related to elderly people living alone has increased. This paper proposes a method to classify falls from non fall action in top Viewed kinect camera depth images. The usage of depth camera images provides an effective solution regarding privacy concerns and the top Viewed camera output has an added advantage of reducing occlusion effect in the cluttered home environment. Our method considers a fixed background setting overall the experiments and foreground is obtained by frame differencing. Then the human silhouette is extracted by largest connected component selection. Ellipse Fit over the human silhouette is used to obtain feature vectors. A binary support vector machine(SVM)classifier is used to distinguish fall from non falling frames. The proposed method is tested over[6] UR fall detection dataset providing a platform for comparison to other researchers. Sowmya Kasturi, Kang-Hyun Jo |
HSI | 2 |
| 2017 | Baggage detection and classification using human body parameter & boosting techniqueabstractAutomatic Video Surveillance System (AVSS) has become important to computer vision researchers as crime in public places has increased in the twenty first century. As a new branch of AVSS, baggage detection and classification has a broad area of security applications. Some of them are, detecting carriage of illegal materials into baggage, detecting unclaimed baggage in public space that can be placed by terrorists for violence, detecting baggage in baggage restricted super shop etc. However, in this paper, a detection & classification framework of baggage is proposed using dynamic human body parameter with boosting strategy. Initially, background subtraction is performed instead of sliding window approach to speed up the system and HSI model is used to deal with uneven illumination condition. Then, to overcome the shadow effect a model is introduced. Extraction of rotational signal descriptor (RSD_HOG) from Region of Interest (ROI) added efficiency in HOG. Finally, dynamic approach in human body parameter setting enabled the system to detect & classify single or multiple carried baggages although some portions of human are absent. In baggage detection, boosting of similarity measure based cascade multilayer SVMs into HOG based SVM generated a strong classifier. This scheme has used to deal with various texture patterns of baggages. Experimental results discovered the system satisfactorily accurate and faster comparative to other alternatives. Tahmina Khanam, Kaushik Deb, Kang-Hyun Jo |
HSI | 3 |
| 2017 | Object classification for LIDAR data using encoded featuresabstractObject classification is an important task in vision-based systems. In this work, an intelligent system to perform detection and classification of road objects is presented. The proposed method utilize machine learning algorithm to classify group of points into various categories that represent several road objects. This classification system was trained using 50 features of 2D laser point which were encoded into smaller dimension in order to obtain efficiency. The method was evaluated on public dataset and the experiment results shows that the proposed method achieve quality improvements compared to the baseline. Comparison of several machine learning methods for object classification is also presented to emphasize the superiority of this proposed method. Laksono Kurnianggoro, Kang-Hyun Jo |
HSI | 2 |
| 2017 | Strategy for automatic person indexing and retrieval system in news interview video sequencesabstractThe needs for automatic semantic information indexing and retrieval systems have raised owing to the rapid growth of video content. Above all, in the information sources of the video, textual information has an important role in high level semantics. This paper presents the strategy for automatic person indexing and retrieval system using overlay text in the television news interview video sequences. For recognition the overlay text and extraction the information, the proposed methods are based on many of the accepted production rules of the television news and the temporality of the video sequences. The experimental results on Korean television news video show that the proposed method efficiently realizes the automatic indexing and retrieval system. Kang-Hyun Jo |
HSI | 2 |
| 2017 | Multi-layer superpixel-based MeshStereo for accurate stereo matchingabstractThis paper describes an approach to generate a dense disparity map between the stereo pairs using the multilayer superpixel with the different size. We assume the pixels within superpixel belong to the same 3D surface. The pixel-wise matching costs are computed with census and gradient features. Then, the input stereo images are segmented into M layers superpixel. SLIC algorithm is used to obtain superpixel. The disparity map for each superpixel image is generated by optimizing a superpixel grid MRF. There are the overlapping regions regardless of the size of the superpixel. The multiple disparity maps are merged by using the median filter. Dong-Wook Seo, Kang-Hyun Jo |
HSI | 2 |
| 2017 | Sterile zone monitoring with human verificationabstractThis paper proposes efficient real time method for sterile zone monitoring with human verification. The propose method consists of two main parts: Motion detection module and human verification module. The role of motion detection module is to segment out foreground object from background. Probabilistic Foreground Detector based on Gaussian Mixture Model(GMM) is used. Region of interest (ROI) obtained from motion detection module is fed into SVM classifier. SVM classifier is trained using HOG descriptor. The proposed method is tested on the standard datasets gives promising results. Ajmal Shahbaz, Wahyono, Kang-Hyun Jo |
HSI | 3 |
| 2017 | Smoke Detection on Video Sequences Using Convolutional and Recurrent Neural Networks
Alexander Filonenko, Laksono Kurnianggoro, Kang-Hyun Jo |
ICCCI (2) | 3 |
| 2017 | Shape Classification Using Combined Features
Laksono Kurnianggoro, Wahyono, Alexander Filonenko, Kang-Hyun Jo |
ICCCI (2) | 4 |
| 2017 | Human Carrying Baggage Classification Using Transfer Learning on CNN with Direction Attribute
Wahyono, Kang-Hyun Jo |
ICIC (1) | 2 |
| 2017 | Identification of pedestrian attributes using deep networkabstractIn the person re-identification across multiple camera research field, attributes of the pedestrian are important cues to differentiate the appearance of each identity. In this work, ten types of attributes are considered as defined in the DukeMTMC-attribute dataset. A custom deep network architecture is proposed to perform the identification process. Furthermore, experiments were carried out to assess the system compared to other pre-trained networks which are commonly used in other literature. The results show that the proposed network achieve better performance compared to the others.σ Laksono Kurnianggoro, Kang-Hyun Jo |
IECON | 2 |
| 2017 | Optimal background modeling for cluttered scenesabstractThis paper proposes optimal background modeling scheme for the cluttered scenes. The background initialization is the first step in the process of segmenting out moving information. Concrete background model ensures the proper segmentation of moving information from the scene. Each pixel is modeled as mixture of Gaussian. Using decision criteria, background/foreground pixels are differentiated. During the background maintenance step, three different learning rates were used to update model separately. They constitute short, medium, and long term background model. Background models obtained with each learning rate are augmented using temporal median filtering separately. Finally, three background models are compared with ground truth using loss function based on the Mean Square Error (MSE). The background model with minimum MSE value is selected as optimal background model. The proposed method is tested on the standard datasets available online. Ajmal Shahbaz, Kang-Hyun Jo |
IECON | 2 |
| 2017 | An improved method for 3D shape estimation using cascade of neural networksabstractThis paper tackles the problem of estimating 3D human poses from given 2D landmarks, which is still an ill-posed problem. The existing works have successfully applied Active Shape Model approach to estimate 3D human poses, but the error is still high. In this paper, we propose an improved method by using the cascade of neural networks to make the estimated shape more alike to the ground truth shape in two steps: (i) Creating the initial estimated 3D shape using existing methods. (ii) Making estimated shape more accuracy by using the cascade of neural networks. Compared to existing works, our method shows a significant improvement. Van-Thanh Hoang, Van-Dung Hoang, Kang-Hyun Jo |
INDIN | 3 |
| 2017 | Automatic person information extraction using overlay text in television news interview videosabstractThe growing amount of video data has raised the needs for automatic semantic information indexing and retrieval systems. Among the information sources of the video, the textual information becomes an important source of high-level semantics. This paper presents the framework for automatic person information extraction using overlay text in the television news interview video sequences. Therefore, one of the contribution of this paper is to help to build automatically indexing for broadcast news video archives. This system is based on many of the accepted production rules of the television news and the temporality of the video sequences. The experimental results on Korean television news video show that the proposed method efficiently realizes the automatic person information extraction in the news interview video sequences. Kang-Hyun Jo |
INDIN | 2 |
| 2017 | Detection of pedestrian crossing road: A study on pedestrian pose recognition
Joko Hariyono, Kang-Hyun Jo |
Neurocomputing | 2 |
| 2017 | Body part boosting model for carried baggage detection and classification
Wahyono, Joko Hariyono, Kang-Hyun Jo |
Neurocomputing | 3 |
| 2017 | Cumulative Dual Foreground Differences for Illegally Parked Vehicles DetectionabstractIllegally parked vehicles on the urban road may create a traffic flow problem as well as a potential traffic accident, such as crashing between parked and other vehicles. Thus, the intelligent traffic monitoring system should be able to prevent this situation by integrating an illegally parked vehicle detection module. However, implementing such a module becomes more challenging due to road environments, such as weather conditions, occlusion, and illumination changing. Hence, this work addresses a method to implement an illegally parked vehicle detection based on the cumulative dual foreground differences from the short- and long-term background models, temporal analysis, vehicle detector, and tracking. The extensive experiments were conducted using both iLIDS and our proposed datasets to evaluate the effectiveness of the proposed method by comparing with other methods. The results showed that the method is effective in detecting illegally parked vehicles and can be considered as part of the intelligent traffic monitoring system. Wahyono, Kang-Hyun Jo |
IEEE Trans. Ind. Informatics | 2 |
| 2016 | Accelerative Object Classification Using Cascade Structure for Vision Based Security Monitoring Systems
Van-Dung Hoang, Kang-Hyun Jo |
ACIIDS (1) | 2 |
| 2016 | Multiscale Car Detection Using Oriented Gradient Feature and Boosting Machine
Wahyono, Van-Dung Hoang, Kang-Hyun Jo |
ACIIDS (1) | 3 |
| 2016 | Robust lane marking detection based on multi-feature fusionabstractIn the field of intelligent vehicle systems (IVS), color and edge of lane markings are important features for vision-based applications. This paper proposes a method to detect lane marking based on a fusion approach which combine color and edge lane marking information. Firstly, by knowing the vehicle speed the road surface region of interest is extracted using the typical stopping distance. Secondly, a lane marking clustering method is introduced. This is done by combining the edge and color information of the lane marking. Finally, a fitting model is implemented. A line fitting model is used to extract the lane marking parameters. However for those regions in which lane can not described as a line, the algorithm computed the curve parameters using Lagrange interpolating polynomial. Danilo Cáceres Hernández, Dong-Wook Seo, Kang-Hyun Jo |
HSI | 3 |
| 2016 | Stairways detection and distance estimation approach based on three connected point and triangular similarityabstractDetecting stair region and estimating the distance from a camera to stair in a stair image is the fundamental step in the implementation of autonomous stair climbing navigation, as well as alarm systems for vision impaired people. In this paper, a framework is proposed for detecting the stair region from stair image utilizing some natural properties of stair. One unique property of them is, every stair step's beginning and ending horizontal edge point intersects with two vertical edge points creating three connected point. These vertical edges are stair step's height and its width edge. Another unique property is steps of a stair appear gradually increasing order from top to bottom of the stair in a parallel arrangement. For that initially, directional Gabor filter and Canny edge detector are employed on the stair image to eliminate the influence of illumination and for detecting stair edges. Non-candidate stair edges are removed by performing filtering operation. Then longest horizontal edges are extracted by using a proposed edge linking method on the edge image. After that, a search method is applied for finding stair step height and width edge point at the beginning and ending point of the longest horizontal edges. This operation is performed for detecting three connected points to validate the stair edges. In the next step, increasing longest horizontal edge segments are extracted by comparing x coordinate values of two consecutive edges end points to justify them as stair edges. Finally, from this set of horizontal edges, vanishing point is calculated to verify stair edges from other similar patterns and confirm the detection of stair candidate region. In addition, the triangular similarity is used for distance estimation from camera to stair. The proposed framework is tested using various stair images under a variety of conditions and results are presented to demonstrate the efficiency and effectiveness. Md. Khaliluzzaman, Kaushik Deb, Kang-Hyun Jo |
HSI | 3 |
| 2016 | Building facade detection using geometric planar constraintsabstractThis paper describes the approach for detecting the building facade using geometric constrain. A planar building facade is only considered. Most of the building facade has the form of a rectangle, which is produced by two or more vertical and horizontal lines, which are orthogonal to each other on the building. The parallel lines in the 3D space are projected onto lines in the image plane that are grouped along the direction of vanishing points. Line segments are extracted from an image for estimating vanishing points. In order to cluster the line segments corresponding to respective a hypothesized vanishing point, J-linkage algorithm is used for simultaneous estimation of multiple models. After clustering the line segments, vanishing point is estimated by minimizing error between vanishing point and line segments of an image. Building facade is detected by using the relationship between vertical and horizontal vanishing points. Building has a cubic structure with two perpendicular planes, which meet at a line. The intersection points between two horizontal vanishing points are located on the intersection line of two facades. Finally, building facades are extracted by the geometric relationships between vertical vanishing point corresponding to respective horizontal vanishing point. Dong-Wook Seo, Hyun-Deok Kang, Danilo Cáceres Hernández, Kang-Hyun Jo |
HSI | 4 |
| 2016 | Designing interface and integration framework for multi-channels intelligent surveillance systemabstractThis paper addresses design and implementation of a multi-channels intelligent surveillance system (ISS). The system is developed for handling multiple source data from four different camera networks. It processes the incoming frames simultaneously using a multi-threading strategy and automatically detect any suspicious event on the monitoring area. Different with the existing commercial systems, the ISS covers several different surveillance tasks in ensuring a public safety such as unattended object detection, fire and smoke detection, human detection and tracking, sterile zone monitoring, and illegally parked vehicle detection. Each individual task is implemented separately using advanced image processing and pattern recognition algorithms. The system combines these all surveillance tasks considering not only a high accuracy, but also a fast processing time. Wahyono, Alexander Filonenko, Kang-Hyun Jo |
HSI | 3 |
| 2016 | Online Background-Subtraction with Motion Compensation for Freely Moving Camera
Laksono Kurnianggoro, Wahyono, Danilo Cáceres Hernández, Kang-Hyun Jo |
ICIC (2) | 5 |
| 2016 | Beginning Frame and Edge Based Name Text Localization in News Interview Videos
Jungil Ahn, Youlkyeong Lee, Kang-Hyun Jo |
ICIC (3) | 4 |
| 2016 | A Similarity-Based Approach for Shape Classification Using Region Decomposition
Wahyono, Laksono Kurnianggoro, Kang-Hyun Jo |
ICIC (2) | 4 |
| 2016 | Online Programming Design of Distributed System Based on Multi-level Storage
Laksono Kurnianggoro, Wahyono, Kang-Hyun Jo |
ICIC (3) | 4 |
| 2016 | Estimation of collision risk for improving driver's safetyabstractThis paper introduces a method for analyzing the critical situation based on collision risk probability. Pedestrians in the scene are captured from a monocular camera mounted on the vehicle. Position information of object is extracted by projecting the centroid of bounding box to the ground plane. Five elements of collision criteria are used for our risk analysis. Pedestrian walking direction, its velocity and how aware pedestrian to the traffic are obtained from the pedestrian side. Car speed and relative distance of pedestrian from the car are extracted from car side. Then, with certain values of collision criteria, those elements are constructed. The critical situation is defined as joint probability of elements. Pedestrian are localized according to the critical situation as green for secure label, yellow for carefully and red for high priority to be alerted. A quantitative analysis is performed by measuring effectiveness of this approach. A real-world measurement and human perception survey are performed for evaluation. The performance evaluation shows our proposed method achieved average accuracy 87.5% and it significantly outperforms human perception survey with more than 30% improvement. Joko Hariyono, Ajmal Shahbaz, Laksono Kurnianggoro, Kang-Hyun Jo |
IECON | 4 |
| 2016 | Coarse-to-fine approach for fast correlation-based visual trackingabstractThis paper proposes a coarse-to-fine approach for fast image tracking. The tracking method is built based on correlation tracker which employs online learning and fast detection by utilizing Fourier transform principles. Firstly, a small patch is extracted from a region near the tracked pixel. This patch is divided into a number of cells and then features are extracted from each cell, providing a set of training data. Together with the target values which are set as maximum at the center of the patch and getting smaller as the cell position getting farther from the center, a training is performed to determine the filter values. In the successive frame, the filter response is calculated to determine the position of the tracked pixel which is co-located with the maximum response of the filter. Since the features are extracted from cells, the new position of the tracked pixel is not precisely known. By employing a second detection at finer resolution within the corresponding cell, the ambiguity of the tracked pixel is eliminated. The proposed method was evaluated on a public dataset and the result shows that this strategy achieves a faster computation time compared to the baseline method. Laksono Kurnianggoro, Dong-Wook Seo, Joko Hariyono, Ajmal Shahbaz, Kang-Hyun Jo |
IECON | 5 |
| 2016 | Parameter analysis of probabilistic foreground detector for intelligent surveillance systemabstractForeground detection can be considered as backbone of multistage computer vision systems. Foreground detection using Gaussian Mixture Models (GMM) is famous choice because of its good accuracy and low computational cost. There are several parameters (e.g., learning rate, mean, and variance) involved in the model and assigning appropriate values may lead to better foreground segmentation. This paper analyzes the effect of different parameters of GMM on the extraction of foreground. Furthermore, optimal parameter setting suitable in every background setting e.g. indoor, outdoor, and complex backgrounds, etc. is proposed. Standard datasets with indoor and outdoor sequences were tested. Ajmal Shahbaz, Laksono Kurnianggoro, Joko Hariyono, Kang-Hyun Jo |
IECON | 4 |
| 2016 | Smoke detection for surveillance cameras based on color, motion, and shapeabstractThis paper presents a smoke detection approach for surveillance cameras that uses color, shape, and motion characteristics. The fact a camera is immovable simplifies detection task by applying background subtraction. Color analysis emphasizes moving objects that have higher probability to be actual smoke. Due to limited performance of background subtraction, a real smoke region is represented as many separate pixels. These pixels are combined using density-based spatial clustering of applications with noise method and morphological operations. Shape of smoke candidate is evaluated using boundary roughness and area variability. Irregular density of smoke can be checked by edge density. The dynamic nature of smoke is confirmed by motion analysis. Tests on various datasets have shown consistency of the method. Alexander Filonenko, Danilo Cáceres Hernández, Wahyono, Kang-Hyun Jo |
INDIN | 4 |
| 2016 | Detecting illegally parked vehicle based on cumulative dual foreground differenceabstractAs one of the traffic monitoring tasks, detecting an illegally parked vehicle aims to prevent car crashing between parked and other vehicles. However, developing such a task becomes more complex due to weather conditions, occlusion, illumination changing, and other factors. This work addresses a framework to detect an illegally parked vehicle using a cumulative dual foreground difference. In our framework, two background models with different learning rates are generated based on a Gaussian mixture model, defined as short- and long-term models. Each model extracts foreground pixels and the stability of these pixels are then analyzed based on cumulative values and temporal positions over a certain period of time. Subsequently, the connected component labeling is performed on the static pixels to form stable regions. To determine whether the candidate region is vehicle, a rule-based filtering approach is performed. Finally, the detection-based tracking is applied to reduce false positives. The effectiveness of the proposed framework is evaluated using i-LIDS and ISLab dataset. The experiment results show that the proposed framework is efficient and robust to detect an illegally parked vehicle. Thus, it can be considered as one of the task solutions for a traffic monitoring system. Wahyono, Alexander Filonenko, Kang-Hyun Jo |
INDIN | 3 |
| 2016 | Joint components based pedestrian detection in crowded scenes using extended feature descriptors
Van-Dung Hoang, Kang-Hyun Jo |
Neurocomputing | 2 |
| 2016 | A Simplified Solution to Motion Estimation Using an Omnidirectional Camera and a 2-D LRF SensorabstractRecently, omnidirectional camera-based motion estimation has been improved as a result of planar-motion assumption, which helps reduce the computational time. In practice in outdoor terrains, vehicle motion does not satisfy this assumption. Motivated by this problem, this paper proposes a method that uses the minimal set of parameters of geometric constraint for estimating the pseudo-three-dimensional motion of the vehicle based on an omnidirectional camera and a laser range finder (LRF). To reduce the number of the parameters for accelerating the computational speed, it is supposed that the vehicle moves under the nonholonomic four-wheel motion model, which requires constraints between rotation and translation components. The method consists of two stages. First, the LRF is used to estimate the orientation and the translation magnitude of the vehicle movement. Second, a pair of corresponding points in sequential images is used along with the results of the first stage to estimate the motion of the vehicle. Furthermore, this paper also presents a closed-form solution for problem solving. The experimental results, using synthetic and real data in different terrain conditions, demonstrate that the proposed method provides the higher accuracy with the lower computational time as compared with the state-of-the-art methods. Van-Dung Hoang, Kang-Hyun Jo |
IEEE Trans. Ind. Informatics | 2 |
| 2016 | Unattended Object Identification for Intelligent Surveillance Systems Using Sequence of Dual Background DifferenceabstractImage-based surveillance systems are widely employed toward safety and security applications in many fields. Cameras, that are connected over an IP network for monitoring public areas, can produce large quantities of video footage. It is tedious for humans to simultaneously observe every type of event on several cameras. Thus, it is necessary to build a user-friendly intelligent system, enabling the analysis of video to detect suspicious events. One of the most important tasks of this system would be to identify unattended objects to prevent an unexpected accident such as the bombing of a public space. This paper presents a novel technique for such a task. The method is based on a sequence of dual background differences, which is obtained by computing the intensity difference between the current and reference background models within a time period. A clustering and an object detector are then integrated to identify the unattended objects. The effectiveness of the method was verified using public and our own databases. The results confirmed that the method is efficient to detect unattended objects and is suitable for implementation in video surveillance systems. Wahyono, Alexander Filonenko, Kang-Hyun Jo |
IEEE Trans. Ind. Informatics | 3 |
| 2015 | Human Detection from Omnidirectional Camera Using Feature Tracking and Motion Segmentation
Joko Hariyono, Van-Dung Hoang, Kang-Hyun Jo |
ACIIDS (2) | 3 |
| 2015 | Combined Motion Estimation and Tracking Control for Autonomous Navigation
Van-Dung Hoang, Kang-Hyun Jo |
ACIIDS (2) | 2 |
| 2015 | Iterative road detection based on vehicle speedabstractWhen moving towards fully autonomous navigation, safety plays the most important role for both pedestrian and driver. This paper proposes a method to estimate the lane road region of interest based on the stopping typical distance of a vehicle required by the current speed of the vehicle. This was achieved by taking advantage of the difference in color of the road surface given by the lane marking as well as the pavement road. This method was executed in three main steps. Firstly, a distance estimation method using the vehicle speed was presented. Secondly, a lane marking edge feature extraction method was proposed. Finally, in order to determine the road surface a curve fitting model was implemented. Preliminary results were performed and tested on a group of consecutive frames to prove the effectiveness of the proposed method. Danilo Cáceres Hernández, Alexander Filonenko, Laksono Kurnianggoro, Dong-Wook Seo, Kang-Hyun Jo |
HSI | 5 |
| 2015 | Detecting abandoned objects in crowded scenes of surveillance videos using adaptive dual background modelabstractDetecting an abandoned object in crowded scenes of surveillance videos becomes more complex task due to occlusions, lighting changes, and other factors. In this paper, a new framework to detect abandoned object using dual background model subtraction is presented. In our system, the adaptive background model is generated based on statistical information of pixel intensity that robust against lighting condition. Foreground analysis using geometrical properties is then applied in order to filter out false region. Human and vehicle detection are then integrated to verify the region as static object, human or vehicle. The robustness and efficiency of the proposed method are tested on several public databases such as i-LIDS and PETS2006 datasets. These are also tested using our own dataset, ISLab dataset. The test and evaluation result show that our method is efficient and robust to detect abandoned object in crowded scenes. Wahyono, Alexander Filonenko, Kang-Hyun Jo |
HSI | 3 |
| 2015 | Comparison of Edge Operators for Detection of Vanishing Points
Dong-Wook Seo, Danilo Cáceres Hernández, Kang-Hyun Jo |
ICCCI (1) | 3 |
| 2015 | Spatial-Based Joint Component Analysis Using Hybrid Boosting Machine for Detecting Human Carrying Baggage
Wahyono, Kang-Hyun Jo |
ICCCI (1) | 2 |
| 2015 | Distance Sensor Fusion for Obstacle Detection at Night Based on Kinect Sensors
Alexander Filonenko, Danilo Cáceres Hernández, Andrey Vavilin, Taeho Kim 0004, Kang-Hyun Jo |
ICIC (2) | 5 |
| 2015 | Localization of Pedestrian with Respect to Car Speed
Joko Hariyono, Danilo Cáceres Hernández, Kang-Hyun Jo |
ICIC (2) | 3 |
| 2015 | Carried Baggage Detection and Classification Using Part-Based Model
Wahyono, Kang-Hyun Jo |
ICIC (3) | 2 |
| 2015 | Detection of pedestrian crossing roadabstractDetection of pedestrian crossing road is described in this paper. Single camera is used to detect pedestrians, thus classify them as a pedestrian crossing road or not. The moving pedestrian is detected using improved sparse optical flow method. The proposed technique consists of three main components. First, overlapping blocks are applied in consecutive images. KLT tracker is used to find corresponding corner feature in consecutive images. Second, classify each block into motion region (foreground) and background, where each block is processed by a cascade composed of three classifiers. Third, probabilistic generation of the foreground mask is performed. The classification decisions for all blocks are integrated into final pixel-level foreground segmentation. In order to classify the pedestrian crossing road, a walking human model is proposed. It is calculated by the region volume of the detected bounding box. A walking human is defined as the ratio of the width divided by the height of the detected bounding box. To be sure that the moving object is a walking human, ratio of the centroid location from the ground plane divided by the height of bounding box should satisfy a constraint. The proposed algorithms are evaluated using publicly (Caltech and ETH) datasets and our real world driving data. The performance result shows the correct pedestrian detection rate is 99.50% at 0.09 false positive per image. The pedestrian crossing road classification shown correct detection rate is 98.10%. Joko Hariyono, Kang-Hyun Jo |
ICIP | 2 |
| 2015 | Smoke Detection for Autonomous Vehicles using Laser Range Finder and Camera
Alexander Filonenko, Danilo Cáceres Hernández, Van-Dung Hoang, Kang-Hyun Jo |
IEA/AIE | 4 |
| 2015 | Tracking Failure Detection using Time Reverse Distance Error for Human Tracking
Joko Hariyono, Van-Dung Hoang, Kang-Hyun Jo |
IEA/AIE | 3 |
| 2015 | Real-time flood detection for video surveillanceabstractThis paper introduces the real-time flash flood detection method for stationary surveillance cameras. It can be applied for rural and urban areas and capable of working during day time. The background subtraction was used to detect all changes appear in a scene. After this step, many pixel belonging to the same moving objects may be divided. They are united by morphological closing. Too small separate objects are then removed form the scene. Color probability was calculated for all the pixels belonging to a foreground mask and connected components with low probability value were filtered out. Finally, results were improved by edge density and boundary roughness. The most time consuming step was implemented in parallel using CUDA. Real-time performance was achieved in this way. The algorithm was tested on publicly accepted video. Alexander Filonenko, Wahyono, Danilo Cáceres Hernández, Dong-Wook Seo, Kang-Hyun Jo |
IECON | 5 |
| 2015 | Building detection based on facet for urban reconstructionabstractThis paper describes an approach to detect the building based on facet for urban reconstruction. The urban environment is composed of a variety objects. In particular, the building has a rich geometric structure. It consists of two or more facets. Each facet is presented with a quadrangle in the image. A quadrangle is produced by two or more vertical and horizontal lines. There are many lines which are orthogonal to each other on the building. The parallel lines in the 3D space are projected onto lines in the image plane that are grouped along the direction of vanishing points, provide strong cues for detecting the building facet in the urban environment. We detect the facet of building based on vanishing points. MSAC is used to estimate the dominant vanishing points in the image. These are separated into one vertical direction and two or more horizontal directions. An initial facet is made by a vertical and one of horizontal clusters. The building has a cubic structure with two orthogonal planes which meet at one line. The intersection points between two horizontal vanishing points are located on the intersection line of two facets. Finally, building facets are extracted by the geometric relationships between vertical and horizontal vanishing points. Dong-Wook Seo, Danilo Cáceres Hernández, Alexander Filonenko, Kang-Hyun Jo |
IECON | 4 |
| 2015 | Illegally parked vehicle detection using adaptive dual background modelabstractDetecting an illegally parked vehicle in urban scenes of traffic monitoring system becomes more complex task due to occlusions, lighting changes, and other factors. In this paper, a new framework to detect illegally parked vehicle using dual background model subtraction is presented. In our system, the adaptive background model is generated based on statistical information of pixel intensity that robust against lighting condition. Foreground analysis using geometrical properties is then applied in order to filter out false region. Vehicle detection is then integrated to verify the region as vehicle or not. Vehicle detection method is performed based on Scalable Histogram of Oriented Gradient feature and is trained using Support Vector Machine. The robustness and efficiency of the proposed method are tested on i-LIDS datasets. These are also tested using our own dataset, ISLab dataset. The test and evaluation result show that our method is efficient and robust to detect illegally parked vehicle in traffic scenes. Thus, it is very useful for traffic monitoring application system. Wahyono, Alexander Filonenko, Kang-Hyun Jo |
IECON | 3 |
| 2015 | Real-time smoke detection for surveillanceabstractThis paper introduces the smoke detection method for surveillance cameras. The background subtraction was used to determine moving objects. Color probability was utilized to find possible smoke pixels in a scene. Separate pixels, acquired by background subtraction, were united by morphological operations and connected components labeling methods. The existence of the smoke region is then confirmed by boundary roughness and edge density. In the last step, the current frame is compared to the previous one in order to check the behavior of objects. The most computationally expensive steps are processed in parallel using CUDA to achieve real-time performance. Computational time was decreased by more than 6 times comparing to the CPU processing only. Alexander Filonenko, Danilo Cáceres Hernández, Kang-Hyun Jo |
INDIN | 3 |
| 2015 | Crosswalk detection based on laser scanning from moving vehicleabstractThe safety plays the most important role for both pedestrian and driver in autonomous, semi-autonomous or non-autonomous vehicle. To improve pedestrian safety, this paper presents a new type of laser feature extraction methods for crosswalk marking through Laser Measurement System (LMS). The crosswalk detection is achieved in three stages as follows: Lane Surface Identification (LSI), Lane Marking Recognition (LMR), and finally Crosswalk Marking Detection (CMD). Preliminary results were performed and tested on a group of consecutive frames during the daylight condition to prove its effectiveness. Danilo Cáceres Hernández, Alexander Filonenko, Dong-Wook Seo, Kang-Hyun Jo |
INDIN | 4 |
| 2015 | Automatic LED text recognition method on electronic road sign using local spatial pattern and random forest classifierabstractIn the field of intelligent transportation systems (ITS), an electronic road sign (ERS) is an important device for giving a real-time traffic-related information. The ERSs generally display dynamic text information that each character consists of matrix of a light-emitting diodes lamp, named LED text. This paper addresses an LED text detection and recognition method, as an application of ITS for assisting the driver. Our method is divided into several main stages. First, the ERS is localized from the input image using color model on the RGB-color space. Second, LED text contained on the ERS are detected based on supporting points. supporting points representing as a center of LED segment on a binary map of the input image. Third, each character of LED text is recognized using local spatial pattern feature and random forest classifier. Last, the recognized characters are merged into text line. Experimental results verify that the proposed method is robust to detect and recognize the LED text. Wahyono, Alexander Filonenko, Kang-Hyun Jo |
Intelligent Vehicles Symposium | 3 |
| 2015 | Comparison of Vehicle Control Systems Based on Multiple FeaturesabstractIn striving to achieve autonomous navigation, the guidance system plays an essential role. In this article, the authors present two recently developed vision-based methods to estimate the heading angle of the vehicle. The authors propose a real-time guidance application based on edge and color information using an omnidirectional image. First, line segments were extracted. Second, the Random Sample Consensus (RANSAC) curve fitting method was implemented. Third, the set of intersection points for each pair of curves was extracted. Finally, by implementing the Density Based Spatial Clustering of Applications with Noise (DBSCAN) algorithms, the heading angle was computed. The main contribution of this work is in the evaluation of the methods applied, which included the “edge-based extraction in road scenes” method, which uses the spatial density information around the vehicle (curbs, barrier, gutters, side strip areas), and the “edge-based lane marking segmentation.” Both of these methods were evaluated in terms of the performance assessment of the length of the line segments in relation to the heading angle measurement. In that sense, the experiments were conducted for the purpose of testing the processing time and the number of extracted line segments. The preliminary results were gathered and tested on a group of consecutive frames to demonstrate the effectiveness of the proposed methods. Danilo Cáceres Hernández, Dong-Wook Seo, Hyun-Uk Chae, Kang-Hyun Jo |
Cybern. Syst. | 4 |
| 2015 | LED Dot matrix text recognition method in natural scene
Wahyono, Kang-Hyun Jo |
Neurocomputing | 2 |
| 2014 | Human Detection from Mobile Omnidirectional Camera Using Ego-Motion Compensated
Joko Hariyono, Van-Dung Hoang, Kang-Hyun Jo |
ACIIDS (1) | 3 |
| 2014 | Methods for Vanishing Point Estimation by Intersection of Curves from Omnidirectional Image
Danilo Cáceres Hernández, Van-Dung Hoang, Kang-Hyun Jo |
ACIIDS (1) | 3 |
| 2014 | Simple and Efficient Method for Calibration of a Camera and 2D Laser Rangefinder
Van-Dung Hoang, Danilo Cáceres Hernández, Kang-Hyun Jo |
ACIIDS (1) | 3 |
| 2014 | Laser based obstacle avoidance strategy for autonomous robot navigation using DBSCAN for versatile distanceabstractTowards fully autonomous navigation, guidance plays an important task for successful autonomous navigation. In this paper, the authors propose an obstacle avoidance strategy based on distance clustering analysis for safe autonomous robot navigation. Autonomous navigation systems must be able to recognize objects in order to perform a collision free motion in both unknown indoor/outdoor environments. Firstly, it was proposed to detect objects using the Density-based spatial clustering of applications with noise (DBSCAN) method through a dynamic density-reachable implementation. Secondly, in order to determine an optimal path for collision avoidance a distance clustering analysis was implemented. Subsequently, a set of possible waypoints were extracted in order to estimate the best path candidate. Preliminary results were gathered and tested on a group of consecutive frames. These specific methods of measurement were chosen to prove their effectiveness. Danilo Cáceres Hernández, Van-Dung Hoang, Kang-Hyun Jo |
HSI | 3 |
| 2014 | Global path planning for unmanned ground vehicle based on road map imagesabstractIn automatic navigation of mobile systems, first, they require providing a path network for robot/vehicle motion. Therefore, path planning is an important task of autonomous vehicle systems. To deal with the problem, this paper presents a method for constructing the shortest path, which support for vehicle auto-navigation in outdoor environments. The method using online road map images to estimate not only the shape of road network but also the directed road network, which could not be estimated by the use of only aerial/satellite images. The proposed method to solve this problem includes three stages. First, a raw network of path for motion is detected using the road map images. Second, the path network is converted to the Global coordinates, which provides a convenience for online auto-navigation task. Third, the shortest path for motion is estimated based on the A* algorithm. The experimental results demonstrate robustness and effectiveness of the method for path networks estimation under the large scene of outdoor environments. Van-Dung Hoang, Danilo Cáceres Hernández, Joko Hariyono, Kang-Hyun Jo |
HSI | 4 |
| 2014 | Motion Segmentation Using Optical Flow for Pedestrian Detection from Moving Vehicle
Joko Hariyono, Van-Dung Hoang, Kang-Hyun Jo |
ICCCI | 3 |
| 2014 | Optimal Partial Rotation Error for Vehicle Motion Estimation Based on Omnidirectional Camera
Van-Dung Hoang, Kang-Hyun Jo |
ICCCI | 2 |
| 2014 | Augmented Reality Surveillance System for Road Traffic Monitoring
Alexander Filonenko, Andrey Vavilin, Taeho Kim 0004, Kang-Hyun Jo |
ICIC (2) | 4 |
| 2014 | Partially Obscured Human Detection Based on Component Detectors Using Multiple Feature Descriptors
Van-Dung Hoang, Danilo Cáceres Hernández, Kang-Hyun Jo |
ICIC (1) | 3 |
| 2014 | Ego-Motion Compensated for Moving Object Detection in a Mobile Robot
Joko Hariyono, Laksono Kurnianggoro, Wahyono, Danilo Cáceres Hernández, Kang-Hyun Jo |
IEA/AIE (2) | 5 |
| 2014 | Fuzzy Logic Guidance Control Systems for Autonomous Navigation Based on Omnidirectional Sensing
Danilo Cáceres Hernández, Van-Dung Hoang, Alexander Filonenko, Kang-Hyun Jo |
IEA/AIE (1) | 4 |
| 2014 | Similarity-Based Classification of 2-D Shape Using Centroid-Based Tree-Structured Descriptor
Wahyono, Laksono Kurnianggoro, Joko Hariyono, Kang-Hyun Jo |
IEA/AIE (2) | 4 |
| 2014 | Smoke detection on roads for autonomous vehiclesabstractThis paper describes the smoke detection algorithm for autonomous vehicles equipped with camera and lidar. The main feature is the ability to detect smoke with ego motion of the camera. Color characteristics of smoke are used to detect regions of interest by similarity of pixels between the current frame and the training data. The following metrics are used: red, green, blue, cyan, saturation channels and spatial entropy. Each region of interest is then enhanced by removing small objects and by filling holes. Sky region is removed by checking edge density of the region. Other rigid objects are expelled by the boundary roughness feature. By knowing the fact that smoke tends to change its shape in frame sequence, the angle-radius shape descriptor is introduced. Cross-correlation of this descriptor between regions in consequent frames will show objects with not appropriate behavior. Data from the camera and lidar are fused to make the final decision. Alexander Filonenko, Van-Dung Hoang, Kang-Hyun Jo |
IECON | 3 |
| 2014 | Local path planning strategy: A practical implementation for versatile distanceabstractIn this paper we present a local path planning for autonomous robot navigation using spatial clustering based object detection. Autonomous navigation systems must be able to recognize objects in order to perform a collision free motion. First, obstacle detection strategy was proposed by using the Density-based spatial clustering of applications with noise (DBSCAN) method through a dynamic density-reachable implementation. Second, in order to determine an optimal path for collision avoidance a distance clustering analysis it was implemented. Then, the system follows the optimal path until a collision is detected or new path should be detected due to environmental uncertainty. Finally, to control the mobile robot, a fuzzy logic controller was implemented. These specific methods of measurement were chosen to prove their effectiveness. Danilo Cáceres Hernández, Van-Dung Hoang, Alexander Filonenko, Kang-Hyun Jo |
IECON | 4 |
| 2014 | Camera and laser range finder fusion for real-time car detectionabstractThis paper describes a car detection method by combining data obtained from a laser and a camera. Data from the camera and the laser range finder (LRF) are combined after a calibration method has been performed. The calibration method defines the relative pose between camera and LRF. Car candidates are then extracted from the LRF data. The car candidate regions on the image are generated based on the filtered LRF data based on its size. To filter out the bad candidates, a verification method is performed on the car candidate regions. This method eliminates the needs of checking over several positions and scales, enables a speed enhancement over the general object detection strategy. Laksono Kurnianggoro, Wahyono, Danilo Cáceres Hernández, Kang-Hyun Jo |
IECON | 4 |
| 2014 | Traffic sign recognition system for autonomous vehicle using cascade SVM classifierabstractIn past two decades, developing a system that can navigate vehicle autonomously becomes more interesting problem. The vehicle is equipped by sensors, such as radar, laser, GPS, and camera for sensing the surrounding. Among them, utilization of camera with computer vision technique is the most adopted method for constructing such a system. It is because camera provides a lot of information and is low-cost device rather than other sensors. Traffic road sign, as one of the important information from camera, carries a lot of useful information that are required for navigating. Thus, in this work, traffic sign detection and recognition is addressed. First, the input image is converted into normalize red and blue color space, as traffic sign usually appear with red and blue color. Second, maximally extremal stable region is then performed for extracting candidate region. Using heuristic rule of geometry properties, the false region will be excluded. Third, histogram of oriented gradient method is applied in order to extract feature from candidate region. Lastly, cascade support vector machine classifier is then processed to classify region belong to certain class of traffic sign. The extensive experiment would be carried out over German traffic sign recognition database and video. The experimental results demonstrate the effectiveness of our systems. Wahyono, Laksono Kurnianggoro, Joko Hariyono, Kang-Hyun Jo |
IECON | 4 |
| 2014 | Hybrid cascade boosting machine using variant scale blocks based HOG features for pedestrian detection
Van-Dung Hoang, Le-My Ha, Kang-Hyun Jo |
Neurocomputing | 3 |
| 2014 | 3D scene reconstruction enhancement method based on automatic context analysis and convex optimization
Le-My Ha, Andrey Vavilin, Kang-Hyun Jo |
Neurocomputing | 3 |
| 2013 | Iterative vanishing point estimation based on DBSCAN for omnidirectional imageabstractRegarding the autonomous of robot navigation, vanishing point (VP) plays an important role in visual robot applications such as iterative estimation of rotation angle for automatic control as well as scene understanding. Autonomous navigation systems must be able to recognize feature descriptors. Consequently, this navigating ability can help the system to identify roads, corridors, and stairs; ensuring autonomous navigation along the environments mentioned before the vanishing point detection is proposed. In this paper, the authors propose solutions for finding the vanishing point in real time based density-based spatial clustering of applications with noise (DBSCAN). First, we proposed to extract the longest segments of lines from the edge frame. Second, the set of intersection points for each pair of line segments are extracted by computing Lagrange coefficients. Finally, by using DBSCAN the VP is estimated. Preliminary results are performed and tested on a group of consecutive frames undertaken at Nam-gu Ulsan, South Korea to prove its effectiveness. Danilo Cáceres Hernández, Van-Dung Hoang, Kang-Hyun Jo |
HSI | 3 |
| 2013 | Planar motion estimation using omnidirectional camera and laser rangefinderabstractThe vision based motion estimation has been investigated in the last few years. Although some progresses have been made in the assumption of planar, still there are no methods satisfying the high accuracy and real-time with absolute translation. This paper proposes a method to estimate vehicle motion by the fusion of omnidirectional camera and laser rangefinder to overcome the drawbacks mentioned above. The vehicle motion contains of rotation and translation components. The rotation is estimated based on the simple but efficient edge matching method by using camera. The absolute translation problem is solved based on ICP method by using laser rangefinder. The experiments were carried out using an electric vehicle with a camera mounted on the roof and a laser device mounted on the bumper. In order to evaluate the motion estimation, the vehicle positions were compared with GPS information and superimposed onto aerial images collected by Google map API. The experimental results showed that the error is 1.1 times smaller and the computational cost is 10.4 times faster than the 1-Point RANSAC method. Also, the error is 4.1 times smaller than the method based on appearance color features. Van-Dung Hoang, Le-My Ha, Kang-Hyun Jo |
HSI | 3 |
| 2013 | Vanishing Point Based Image Segmentation and Clustering for Omnidirectional Image
Danilo Cáceres Hernández, Van-Dung Hoang, Kang-Hyun Jo |
ICIC (2) | 3 |
| 2013 | Combining Edge and One-Point RANSAC Algorithm to Estimate Visual Odometry
Van-Dung Hoang, Danilo Cáceres Hernández, Kang-Hyun Jo |
ICIC (1) | 3 |
| 2013 | Detecting and Recognizing LED Dot Matrix Text in Natural Scene Images
Wahyono, Kang-Hyun Jo |
ICIC (3) | 2 |
| 2013 | Visual surveillance with sensor network for accident detectionabstractThis paper describes an autonomous monitoring system to detect environmental accidents such as fire and gas leaks. This system is designed as a set of sensor nodes mounted on statical and dynamical objects and connected via a wireless network. Each sensor node measures such important environmental parameters as temperature, humidity, poisonous gases concentration, etc. Video surveillance is used to increase the probability the accident is detected. The fire detection vision technique based on the color information and flame behavior properties is used. Video stream is distributed via the Internet and can be viewed on a personal computer (PC) and on a mobile device. The graphical user interface (GUI) based on the sensor network and vision data helps an operator to make correct inference about the threat level. Alexander Filonenko, Kang-Hyun Jo |
IECON | 2 |
| 2013 | Localization estimation based on Extended Kalman filter using multiple sensorsabstractThis paper describes a method for localization estimation based on Extended Kalman filter using an omnidirectional camera and a laser rangefinder. Laser rangefinder information is used for predicting absolute motion of the vehicle. The geometric constraint of sequence pairwise omnidirectional images is used to correct the error and construct the mapping. The advantage of omnidirectional camera is a large of field-of-view, which is helpful for long distance tracking feature landmarks. For motion estimation based on vision, the absolute translation of vehicle is approximated posterior information at previous step. The structure from motion based on bearing and range sensors can yield the corrected local position at short distance of movements but it will be accumulative errors overtime. To utilize the advantages of two sensors, Extended Kalman Filter framework is applied for integrating multiple sensors for localization estimation. The experiments were carried out using an electric vehicle with the omnidirectional camera mounted on the roof and the laser device mounted on the bumper. The simulation results will demonstrate the effectiveness of this method from large field-of-view scene images of outdoor environment. Van-Dung Hoang, Le-My Ha, Danilo Cáceres Hernández, Kang-Hyun Jo |
IECON | 4 |
| 2013 | Omnidirectional stereo vision based vehicle detection and distance measurement for driver assistance systemabstractThis paper proposes the driver assistance system based on the omnidirectional stereo vision. This system informs a driver to prepare unexpected situation during driving. This information is a relative distance between a preceding car and driving vehicle and obtained by using the omnidirectional stereo vision. The omnidirectional camera sees 360 degrees unlike a conventional perspective camera. Therefore it is useful to search the road conditions nearby a driving vehicle at every sequence. For detecting a preceding car, we use histogram of oriented gradient (HOG) which has is robustness to for illumination change. Because of stereo vision, the proposed algorithm measures the relative distance between preceding car and moving vehicle whenever preceding car is detected. In experiments, the ratio of detecting preceding car is 99.78% and 94.59% for left and right cameras respectively. The measurement process of relative distance is started when preceding car is detected both left and right cameras and the system detects a center point of tail right. In the measurement process, if a preceding car is not detected from left or right image scene, the system estimates the center point using optical flow consider with previous frame. The maximum error of estimated distance is less than 25cm shown in the Fig 9. Therefore the error of estimated distance is quite ignorable value because the normal relative distance between two moving car is over 2m. Hansung Park, Kang-Hyun Jo, Kangik Eom, Sungmin Yang, Taeho Kim 0004 |
IECON | 3 |
| 2013 | 3D motion estimation based on pitch and azimuth from respective camera and laser rangefinder sensingabstractThis paper proposes a new method to estimate the 3D motion of a vehicle based on car-like structured motion model using an omnidirectional camera and a laser rangefinder. In recent years, motion estimation using vision sensor has improved by assuming planar motion in most conventional research to reduce requirement parameters and computational cost. However, for real applications in environment of outdoor terrain, the motion does not satisfy this condition. In contrast, our proposed method uses one corresponding image point and motion orientation to estimate the vehicle motion in 3D. In order to reduce requirement parameters for speedup computational systems, the vehicle moves under car-like structured motion model assumption. The system consists of a camera and a laser rangefinder mounted on the vehicle. The laser rangefinder is used to estimate motion orientation and absolute translation of the vehicle. An omnidirectional image-based one-point correspondence is used for combining with motion orientation and absolute translation to estimate rotation components of yaw, pitch angles and three translation components of Tx, Ty, and Tz. Real experiments in sloping terrain demonstrate the accuracy of vehicle localization estimation using the proposed method. The error at the end of travel position of our method, one-point RANSAC are 1.1%, 5.1%, respectively. Van-Dung Hoang, Danilo Cáceres Hernández, Le-My Ha, Kang-Hyun Jo |
IROS | 4 |
| 2013 | Automatic context analysis for image classification and retrieval based on optimal feature subset selection
Andrey Vavilin, Kang-Hyun Jo |
Neurocomputing | 2 |
| 2012 | Robust Human Detection Using Multiple Scale of Cell Based Histogram of Oriented Gradients and AdaBoost Learning
Van-Dung Hoang, Le-My Ha, Kang-Hyun Jo |
ICCCI (1) | 3 |
| 2012 | Enhancing 3D Scene Models Based on Automatic Context Analysis and Optimization Algorithm
Le-My Ha, Andrey Vavilin, Kang-Hyun Jo |
ICIC (3) | 3 |
| 2012 | Enhancing Point Clouds Accuracy of Small Baseline Images Based on Convex Optimization
Le-My Ha, Andrey Vavilin, Sungmin Yang, Kang-Hyun Jo |
IEA/AIE | 4 |
| 2012 | Camera Motion Estimation and Moving Object Detection Based on Local Feature Tracking
Andrey Vavilin, Le-My Ha, Kang-Hyun Jo |
IEA/AIE | 3 |
| 2012 | Fast human detection based on parallelogram haar-like featuresabstractInspired by a recent image descriptors for object detection, this paper proposed the feature description method based on set of modified Haar-like features which have parallelogram shapes. Using the proposed feature descriptors to develop a rapid detection system for human detection based on cascade structure used for boosting classifier. Specially, human detection in omnidirectional image as well as unwrap omnidirectional to panoramic image were described in this paper. The experimental results showed that the proposed method could produce high accuracy detection rate with lower false positive rate and higher recall rate than Haar-like features, and faster than HOG feature. It is efficiency with different resolutions and poses under a variety conditional such as flare illumination, clutter backgrounds, and so on. Van-Dung Hoang, Andrey Vavilin, Kang-Hyun Jo |
IECON | 3 |
| 2012 | Removing outliers of large scale scene models based on automatic context analysis and convex optimizationabstractThis paper proposes a method for removing outliers of large scale scene model. First, the context of the scene images are analyzed. Some objects which may have negative effect should be removed. For instance, the sky often appear as background and moving object appear in most of scene images. They are also one of reasons that cause the outliers. Second, the constraints of image pair-wise are computed based on invariant features. The correspondence problem is solved by iterative method which remove the outliers. To avoid the disadvantage of incremental structure from motion, the global rotation of cameras are estimated by a robust method. These global rotations are fed to the point clouds generation procedure in third step. In contrast with using only canonical bundle adjustment which gain unstable structure in small baseline geometry and local minima, the proposed method utilized known-rotation framework combined bundle adjustment to generate accurate point clouds and camera positions with single global minimum. The patch based multi-view stereopsis is applied to dense point cloud upgrading. The simulation results will demonstrate the accuracy of this method from large scale scene images in outdoor environment. Le-My Ha, Andrey Vavilin, Kang-Hyun Jo |
IECON | 3 |
| 2012 | Moving object detection and camera motion analysis from moving cameraabstractThis paper describes an approach for the scene motion analysis in complex scenes from moving camera. Proposed method consists of two parts. First image of the sequence is used to compose a triangular grid which vertices are used as a feature points in further processing. Grid is optimized in order to increase number of elements in the regions with higher level of details. Description vector based on color and edge distribution associated with each vertex. These feature vectors are used to track vertices in a frame sequence. Grid is then updated in order by removing mismatched vertices. New vertices are added in grid segments with high level of details. A difference between vertices coordinates in consequent frames form vector field which is used for motion analysis. On the second step motion field is analyzed in order to find regions with similar motion vectors. These regions are then classified as a moving object of as a part of the background. Vertices corresponding to one moving object are tracked as a group preserving geometrical relationship between connected points. Andrey Vavilin, Kang-Hyun Jo |
IECON | 2 |
| 2011 | Building Detection and 3D Reconstruction from Two-View of Monocular Camera
Le-My Ha, Kang-Hyun Jo |
ICCCI (1) | 2 |
| 2011 | Building Face Reconstruction from Sparse View of Monocular Camera
Le-My Ha, Kang-Hyun Jo |
ICIC (2) | 2 |
| 2011 | Automatic Context Analysis for Image Classification and Retrieval
Andrey Vavilin, Kang-Hyun Jo, Moon-Ho Jeong, Jong-Eun Ha, Dong Joong Kang |
ICIC (1) | 2 |
| 2011 | Automatic Vehicle Identification by Plate Recognition for Intelligent Transportation System Applications
Kaushik Deb, Le-My Ha, Byung-Seok Woo, Kang-Hyun Jo |
IEA/AIE (2) | 4 |
| 2011 | Stairway Detection Based on Single Camera by Motion Stereo
Danilo Cáceres Hernández, Taeho Kim 0004, Kang-Hyun Jo |
IEA/AIE (1) | 3 |
| 2010 | Entrance Detection of Buildings Using Multiple Cues
Suk-Ju Kang, Hoang-Hon Trinh, Dae-Nyeon Kim, Kang-Hyun Jo |
ACIIDS (1) | 4 |
| 2010 | License Plate Tilt Correction Based on the Straight Line Fitting Method and Projection
Kaushik Deb, Andrey Vavilin, Jung-Won Kim, Kang-Hyun Jo |
ICCCI (3) | 4 |
| 2010 | Entrance Detection of Building Component Based on Multiple Cues
Dae-Nyeon Kim, Hoang-Hon Trinh, Kang-Hyun Jo |
ICIC (3) | 3 |
| 2010 | Fast HDR Image Generation Technique Based on Exposure Blending
Andrey Vavilin, Kaushik Deb, Kang-Hyun Jo |
IEA/AIE (3) | 3 |
| 2010 | Supervised training database for building recognition by using cross ratio invariance and SVD-based method
Hoang-Hon Trinh, Dae-Nyeon Kim, Kang-Hyun Jo |
Appl. Intell. | 3 |
| 2009 | Appearance Feature Based Human Correspondence under Non-overlapping Views
Hyun-Uk Chae, Kang-Hyun Jo |
ICIC (1) | 2 |
| 2009 | Vehicle License Plate Detection Algorithm Based on Color Space and Geometrical Properties
Kaushik Deb, Vasily V. Gubarev, Kang-Hyun Jo |
ICIC (1) | 3 |
| 2009 | Auto-surveillance for Object to Bring In/Out Using Multiple Camera
Taeho Kim 0004, Dong-Wook Seo, Hyun-Uk Chae, Kang-Hyun Jo |
ICIC (1) | 4 |
| 2009 | Object Analysis for Outdoor Environment Perception Using Multiple Features
Dae-Nyeon Kim, Hoang-Hon Trinh, Kang-Hyun Jo |
ICIC (1) | 3 |
| 2009 | A Novel User Created Message Application Service Design for Bidirectional TPEG
Kang-Hyun Jo |
ICIC (2) | 2 |
| 2009 | Window Extraction Using Geometrical Characteristics of Building Surface
Hoang-Hon Trinh, Dae-Nyeon Kim, Suk-Ju Kang, Kang-Hyun Jo |
ICIC (1) | 4 |
| 2009 | Building-Based Structural Data for Core Functions of Outdoor Scene Analysis
Hoang-Hon Trinh, Dae-Nyeon Kim, Suk-Ju Kang, Kang-Hyun Jo |
ICIC (1) | 4 |
| 2009 | An Efficient Method of Vehicle License Plate Detection Based on HSI Color Model and Histogram
Kaushik Deb, Heechul Lim, Suk-Ju Kang, Kang-Hyun Jo |
IEA/AIE | 4 |
| 2009 | A Vehicle License Plate Detection Method for Intelligent Transportation System ApplicationsabstractDetecting license plates is crucial and inevitable in the vehicle license plate recognition system. In this article, a Hue-Saturation-Intensity (HSI) color model is adopted to select automatically statistical threshold value for detecting candidate regions. The focus of this article is on the implementation of a new method to detect candidate regions when vehicle bodies and license plates (LP) have similar color. The proposed method is able to deal with candidate regions under independent orientation and scale of the plate. For the decomposing candidate regions, predetermined LP alphanumeric characters are used by position histogram to verify and detect vehicle LP regions. Various LP images were used with a variety of conditions to test the proposed method and results proved its effectiveness. Kaushik Deb, Kang-Hyun Jo |
Cybern. Syst. | 2 |
| 2008 | Generation of Multiple Background Model by Estimated Camera Motion Using Edge Segments
Taeho Kim 0004, Kang-Hyun Jo |
ICIC (1) | 2 |
| 2008 | Region Segmentation of Outdoor Scene Using Multiple Features and Context Information
Dae-Nyeon Kim, Hoang-Hon Trinh, Kang-Hyun Jo |
ICIC (3) | 3 |
| 2008 | Cross Ratio-Based Refinement of Local Features for Building Recognition
Hoang-Hon Trinh, Dae-Nyeon Kim, Kang-Hyun Jo |
ICIC (1) | 3 |
| 2008 | Building Surface Refinement Using Cluster of Repeated Local Features by Cross Ratio
Hoang-Hon Trinh, Dae-Nyeon Kim, Kang-Hyun Jo |
IEA/AIE | 3 |
| 2007 | Object Recognition of Outdoor Environment by Segmented Regions for Robot Navigation
Dae-Nyeon Kim, Hoang-Hon Trinh, Kang-Hyun Jo |
ICIC (1) | 3 |
| 2007 | Robust Human Face Detection for Moving Pictures Based on Cascade-Typed Hybrid Classifier
Phuong-Trinh Pham-Ngoc, Taeho Kim 0004, Kang-Hyun Jo |
ICIC (2) | 3 |
| 2007 | Automated Junction Structure Recognition from Road Guidance Signs
Andrey Vavilin, Kang-Hyun Jo |
ICIC (1) | 2 |
| 2007 | Remote control of a moving robot using the virtual linkabstractA new remote control method of a moving robot is proposed, where a moving robot is moved according to the pointing position and orientation of the remote controller. The remote controller consists of the camera and gyroscopes. The landmark in a moving robot is recognized by the camera in the remote controller and a robot is moved in the camera's pointing position and orientation. The 'virtual link' term is used since the robot is moved as if there is a link between the robot and the remote controller. Gyroscopes are also used in the remote controller so that fast estimation of the camera's pointing position and orientation is possible. The proposed method is verified through experiments. Young Soo Suh, Sang Kyeong Park, Dae-Nyeon Kim, Kang-Hyun Jo |
ICRA | 4 |
| 1998 | Context-Based Recognition of Manipulative Hand Gestures for Human Computer Interaction
Kang-Hyun Jo, Yoshinori Kuno, Yoshiaki Shirai |
ACCV (2) | 1 |
| 1998 | Manipulative Hand Gesture Recognition Using Task Knowledge for Human Computer Interaction
Kang-Hyun Jo, Yoshinori Kuno, Yasuhiro Shirai |
FG | 1 |
| 1995 | Human-robot interface using uncalibrated stereo visionabstractThis paper presents a human-robot interface system that enables a user to move a robot by moving his hand. The system adopts uncalibrated stereo vision based on the affine invariants from multiple views. The system can interpret hand gestures both in the user-centered frame and in the world-fixed frame. Suppose the user wants to indicate the direction of a mobile robot motion by the direction of his hand. If he moves his hand forward, the robot moves straight ahead regardless of his body position, that is, regardless of the hand direction in the world by the former interpretation. By the latter interpretation, however, the robot turns towards the direction indicated by the hand. The former is useful when we operate a teleoperation robot while watching the images sent from the robot. Operation experiments show the usefulness of the system. Yoshinori Kuno, Kang-Hyun Jo, Yoshiaki Shirai |
IROS (1) | 3 |