Kang Ryoung Park

dblp:98/5714 · DBLP profile ↗
← Back
72ranked-venue papers
9as first author
29since 2021 · last 2026
0000-0002-1214-9510ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 53 · 4 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-authorHuman-computer interaction and ubiquitous computing · 6 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 2
YearPublicationVenuePosition
2026 Covariance attention and correlation-based knowledge distillation for semantic segmentation of lens flare-degraded road scenes
abstract
With the continued advancement of autonomous driving technology, semantic segmentation methods for precisely recognizing forward-facing objects are now essential. However, lens flare, which frequently arises in road environments because of strong light sources such as the sun, street lamps, and vehicle headlights, diminishes image quality and significantly reduces semantic segmentation performance. To address this issue, prior research on frontal-viewing camera images integrated lens flare detection and restoration within the network, but this approach introduces substantial inference time and computational overhead. To solve these problems, this study presents a novel covariance attention and correlation information fusion-based knowledge distillation (CACKD) framework for semantic segmentation of lens flare-degraded road-scene images captured by frontal-viewing camera. The teacher model in CACKD aggregates multi-layer gradient-weighted class activation mapping (Grad-CAM) information to precisely localize both global and fine-grained regions affected by lens flare, and then concentrates processing on these targeted flare regions to restore the image. During semantic segmentation, it employs channel-wise covariance attention based on a covariance projection, and the student model reproduces this attention at corresponding locations; by re-aligning inter-channel correlations, meaningful co-activations are reinforced and unnecessary dependencies are attenuated, substantially improving prediction consistency. In parallel, inter-channel correlation information and covariance-based attention features from multiple teacher stages are transferred to a lightweight, single-stage student model through a hybrid-based knowledge distillation (KD) strategy, to achieve teacher-level accuracy while dramatically reducing the computational cost and inference time of the student model. Experiments conducted on the synthesized flare (Syn-flare) Cambridge-driving labeled video (CamVid) and Syn-flare Karlsruhe Institute of Technology and Toyota Technological Institute at Chicago (KITTI) datasets, respectively, demonstrate that the CACKD student model achieves mean intersection over union (mIoU) scores of 74.48 % and 66.70 %, respectively, while requiring only 32.52 million parameters and outperforming state-of-the-art methods despite its smaller model size. Moreover, the proposed model operates effectively on embedded systems with limited computational resources, underscoring its suitability for real-world autonomous driving environments.
Hyun Woo Song, Seon Jong Kang, Yong Ho Lee, Min Su Jeong, Seong In Jeong, Ho Won Lee, Kang Ryoung Park
Eng. Appl. Artif. Intell.7
2026 Unmanned aerial vehicle-based weed segmentation from multispectral imagery in an edge computing environment
abstract
Unmanned aerial vehicles (UAVs) equipped with computer vision techniques offer a practical and scalable solution for large-scale weed mapping in agricultural fields. Accurate pixel-level discrimination of soil, crops, and weeds from UAV imagery remains challenging due to severe class imbalance and the high visual similarity between crop and weed structures. Many existing segmentation approaches are either computationally demanding or depend on ensemble strategies, which restrict their applicability on edge or real-time platforms. To address these limitations, this study introduces a progressive receptive-field context-aware network (PRC-Net) designed for efficient deployment under constrained computational resources. The proposed progressive receptive-field fusion (PRF) module incrementally enlarges the receptive field across multiple scales, enabling improved identification of small crop and weed regions within predominantly soil backgrounds. A channel synergy-focused attention (CSA) mechanism is incorporated to selectively enhance informative spectral bands and their inter-channel relationships in multispectral data, thereby improving class separability. Furthermore, a lightweight context-aware feature extraction (LCFE) module establishes a compact encoder–decoder bottleneck that facilitates effective contextual refinement while reducing parameter redundancy and overfitting risk. Collectively, these components retain discriminative features of minority classes and enhance spectral segmentation quality with minimal computational overhead. PRC-Net is validated on the publicly available WeedMap (RedEdge-M, Sequoia) and Sesame Aerial datasets, achieving mean intersection over union (MIOU) values of 0.8397, 0.6520, and 0.6961, respectively. Experimental performance highlights that PRC-Net achieves superior performance and efficiency compared to existing methods, confirming its suitability for UAV-based weed detection in precision farming.
Jung Soo Kim, Seong In Jeong, Rehan Akram, Hafiz Ali Hamza Gondal, Muhammad Hamza Tariq, Kang Ryoung Park
Expert Syst. Appl.7
2026 Deep learning-based image outpainting of finger-vein image
abstract
Despite fast authentication and user convenience, the lack of a fixed frame in contactless finger-vein acquisition causes missing regions and discrepancies between enrolled and query images, thereby degrading recognition performance. Existing image outpainting-based methods restore missing regions but often contain a large number of parameters, making them slow and unsuitable for real-time applications. To overcome these issues, this paper proposes a lightweight image outpainting network called knowledge distilled adaptive frequency attention network (KD-AFA-Net). KD-AFA-Net is based on a lightweight model that uses thinner separable U-Net with knowledge distillation (KD) from a high-performance teacher. In addition, to compensate for the limitations of convolutional neural networks (CNNs) in capturing global information, a novel adaptive frequency attention (AFA) module is designed. The AFA module decomposes intermediate features via a two-dimensional fast Fourier transform (FFT), learns the importance of high-frequency and low-frequency components, and emphasizes the important ones. Furthermore, this paper also proposes the AFA KD loss which enables the student model to effectively learn the frequency-domain refined outputs of the teacher’s AFA module. Moreover, we analyze recognition performance and use large language models (LLMs), ChatGPT-4o and ChatGPT-5 to prioritize experiments and to examine utilization strategies for future image-based tasks. Experiments on the Hong Kong Polytechnic University finger-image database version 1, the Shandong University machine learning and applications-homologous multi-modal traits (SDUMLA-HMT) finger-vein database, and the MMCBNU_6000 database show that KD-AFA-Net achieves equal error rates (EERs) of 2.56%, 3.49%, and 1.78% respectively, outperforming state-of-the-art image outpainting and KD methods while supporting real-time efficiency.
Jun Seo Kim, Jin Seong Hong, Jung Soo Kim, Seong In Jeong, Seok Jun Lim, Won Ho Jang, Kang Ryoung Park
Expert Syst. Appl.7
2026 Deep learning-based classification and segmentation of brain tumor progression with clinical pipeline by generative artificial intelligence
abstract
Distinguishing between early-stage brain tumors (EBTs) and progressive-stage brain tumors (PBTs) from magnetic resonance imaging (MRI) scans is pivotal, as interpretation complexity challenges neuro-oncologists and impacts timely treatment decisions. Existing deep learning (DL)-based approaches typically handle brain tumor classification and segmentation separately, emphasizing multi-level feature extraction, but neglecting the benefits of integrated feature fusion. This study introduces a novel DL-based clinical pipeline, enhanced with generative artificial intelligence (GenAI), that explicitly fuses features from classification and segmentation models for stage-specific tumor analysis. Our fusion-based framework first classifies brain tumors into EBTs and PBTs using a dedicated classification network (C-Net), incorporating an enhanced-pooled attention block and a dilated fusion block to selectively extract and fuse multi-level features, balancing computational efficiency with feature relevance. Subsequently, stage-specific segmentation network 1 (S1-Net) and segmentation network 2 (S2-Net) leverage hierarchical feature fusion through progressive upsampling, capturing distinct tumor characteristics at multiple abstraction levels. Finally, a clinician-validated GenAI module provides post-segmentation semantic interpretation by analyzing segmentation masks and patient metadata to describe tumor morphology and explicitly report limitations. Aligned with information fusion principles, this integration combines the precision of task-specific models (S1-Net, S2-Net) with expert-guided GenAI reasoning, enhancing interpretability while preserving clinical safety. The effectiveness of the proposed pipeline is validated with open dataset of brain tumor progression: 1) C-Net achieves an accuracy of 85.06%, precision of 86.73%, recall of 85%, and harmonic mean of recall and precision (F1-score) of 85.84%; 2) S1-Net attains a Dice score (DS) of 79.96% and intersection over union (IoU) of 71.22%; and 3) S2-Net achieves 81.14% DS and 72.44% IoU, significantly outperforming state-of-the-art methods.
Haseeb Sultan, Zeeshan Ullah, Kang Ryoung Park, Jihie Kim
Expert Syst. Appl.3
2025 Weak saliency ensemble network for person Re-identification using infrared light images
abstract
In recent years, person re-identification (re-id) has primarily been studied using visible light (VL) images. However, the challenges of employing VL images in nighttime environments have prompted research into using infrared light (IR) images. Yet, the utilization of both VL and IR images in person re-id has resulted in increased computational cost and processing time in multi-modality systems, leading to studies focusing solely on IR images. Nevertheless, IR images, lacking color and texture information, generally yield lower recognition performance in existing person re-id studies. In addition, previous studies have shown that person re-id performance suffers in the presence of complex background noise. To tackle these challenges, this study proposes a new weak saliency ensemble network (WSE-Net) for person re-id using IR images. WSE-Net incorporates a channel reduction of feature (CRF) method to reduce computational cost in the ensemble network, a technique for converting input images into group of patch images and feeding them into the ensemble model to enhance the reduced feature information, and a grouped convolution ensemble network (GCE-Net) that enables the fusion of features extracted from original and attention-guided ensemble models. The performance of person re-id using WSE-Net was evaluated on the Dongguk body-based person recognition database version 1 (DBPerson-Recog-DB1) and the Sun Yat-sen university multiple modality re-identification version 1 (SYSU-MM01). Experimental results demonstrated that on DBPerson-Recog-DB1, WSE-Net achieved 93.65% in rank 1, 95.28% in mean average precision (mAP), and 93.52% in the harmonic mean of precision and recall. Additionally, on SYSU-MM01, WSE-Net achieved 86.85% in rank 1, 44.58% in mAP, and 40.06% in the harmonic mean of precision and recall. Furthermore, the accuracy of WSE-Net on both datasets surpassed that of state-of-the-art methods.
Min Su Jeong, Seong In Jeong, Dong Chan Lee, Seung Yong Jung, Kang Ryoung Park
Eng. Appl. Artif. Intell.5
2025 Parallel-wise global and local attention vision transformer-based generative adversarial network using fourier transform loss for generating fake iris image
Jung Soo Kim, Jin Seong Hong, Seung Gu Kim, Kang Ryoung Park
Eng. Appl. Artif. Intell.4
2025 Convolutional self-attention with adaptive channel-attention network for obstructive sleep apnea detection using limited training data
abstract
Obstructive sleep apnea (OSA) is a chronic sleep disorder caused by blockage of the upper airway for at least 10 s due to the collapsing of the tongue and soft palate. OSA can cause serious health problems including hypertension and coronary heart. Polysomnography is a technique to simultaneously record physiological signals such as electroencephalograms, electrooculograms, electrocardiograms (ECGs) etc., to diagnose various diseases including OSA. However, the process is time-consuming and tedious. Therefore, detecting OSA from ECGs (electrical signals recording heart variability using electrodes) is an alternative that can be extended to wearable devices. However, two challenges hinder their real-world applications: 1) Performance is directly proportional to the data size, and 2) algorithms are not robust for cross-dataset evaluation. We propose a novel deep-learning model called convolutional self-attention with adaptive channel-attention network (CSAC-Net) to address these issues. Specifically, the first issue is addressed by using the proposed Convolutional self-attention module in a multi-scale projection approach and fusing the features at the end. This enables the exploitation of long-range dependencies with diverse feature vectors. The second issue is addressed by leveraging invariant mapping through the proposed adaptive channel-attention (ACA) and inter-feature attention (IFA) modules. ACA module fuses multi-level features to embed adaptive characteristics while the IFA module exploits features from different stage to preserve the originality of features. To the best of our knowledge, this is the first study to address the underlying issues. Extensive experiments validate the effectiveness of CSAC-Net using two open databases: physiologic signal network apnea electrocardiogram (PhysioNet Apnea-ECG) and national sleep research resource best apnea interventions in research (NSRR-BestAIR). Their respective accuracies are respectively 93.4 % and 76.1 %, outperforming the state-of-the-art methods. Furthermore, the robustness of the CSAC-Net is validated through cross-database evaluation using various open databases.
Nadeem Ullah, Haseeb Sultan, Jin Seong Hong, Seung Gu Kim, Rehan Akram, Kang Ryoung Park
Eng. Appl. Artif. Intell.6
2025 Artificial-Intelligence-Based Low-Light Marine Image Enhancement for Semantic Segmentation in Edge-Intelligence-Empowered Internet of Things Environment
abstract
For accurate detection of marine life to utilize marine resources while ensuring protection of ecosystem, marine animal segmentation (MAS) has been widely researched. Furthermore, development of autonomous underwater vehicle (AUV) has expanded the scope of marine ecosystem research into deep sea where AUV utilizes artificial light sources to address the problem of low-light conditions. However, these light sources can disturb the ecosystem. In addition, extremely low-light images are acquired in areas distant from AUV due to the limitations of the light sources, such as limited field of view, resulting in poor quality of underwater images. Therefore, we propose multiscale features and residual dual attention-based low-light image enhancement network (MRLE-Net) for semantic segmentation of marine images. To preserve fine-grained information under low-light environment and reduce noise, MRLE-Net introduces dual feature extraction, multiscale feature extraction, and residual dual attention blocks. Furthermore, to improve the semantic segmentation accuracy, it employs a discrete wavelet transform-based loss function. In experiments using two open databases of MAS3K and DeepFish, the mean intersection of union values of semantic segmentation by our method are 78.72% and 83.62%, respectively, showing superior accuracy to the state-of-the-art methods. In addition, our MRLE-Net demonstrates its ability to operate on embedded system with low-computational resources as edge computing. From them, we confirm that it can be adopted to AUV in edge intelligence empowered Internet of Things environment by removing communication overheads caused by transmitting lots of images from AUV’s camera to and receiving the segmentation result from high-computing cloud by 5G technology.
Su Jin Im, Chaeyeong Yun, Kang Ryoung Park
IEEE Internet Things J.4
2025 DCR-KD: Dynamic Class Relation Knowledge Distillation for Semantic Segmentation With the Frontal-Viewing Camera of Limited Field of View in an Internet of Things Environment
abstract
In autonomous driving with Internet of Things (IoT) devices, real-time road perception and semantic segmentation are essential for intelligent transportation systems. However, deploying models on IoT devices is challenging due to their limited computational power, memory, and energy availability. Additionally, limited field of view (FoV) occurs frequently in real-world scenarios, such as constrained camera angles or occlusion by large objects, leading to significant degradations in segmentation performance. To address these challenges, we propose Dynamic Class Relation Knowledge Distillation (DCR-KD), a framework that generates lightweight models by transferring knowledge from high-performance teacher models. Central to our method is the Limited FoV Edge Module (LFEM), which extracts edge-aware features of the teacher to refine the learning of the student. LFEM is designed to capture edge-based features in regions with limited FoV, effectively representing critical object boundaries. A channel attention mechanism further enhances semantic features, allowing the student to focus on key information in constrained visual contexts. A dynamic class relation map captures global semantic relationships among classes, enriching the scene understanding of the student. The final student model is independent of the teacher during inference, enabling efficient deployment in resource-constrained environments. Extensive evaluations demonstrate the effectiveness of DCR-KD, including segmentation performance and feature visualizations. Our method bridges the performance gap between resource-intensive teacher models and efficient student models, providing a practical solution for IoT-based real-time road perception, particularly under limited FoV conditions. The code and pre-trained models are publicly available at.
Seong In Jeong, Min Su Jeong, Eun Som Jeon, Kang Ryoung Park
IEEE Internet Things J.4
2024 Multiscale triplet spatial information fusion-based deep learning method to detect retinal pigment signs with fundus images
abstract
Inherited retinal diseases (IRDs) are genetic disorders that cause progressive deterioration of the photoreceptors associated with vision loss or blindness. Retinitis pigmentosa (RP) is a rare hereditary ophthalmic disease that initially causes night blindness owing to continuous retinal pigment deterioration. A computer-aided diagnosis (CAD)-based RP diagnosis solution by pigment sign detection can help ophthalmologists to analyze and treat the disease timely. At present, most of the research addresses retinal disease CAD using expensive optical coherence tomography (OCT); however, fundus imaging-based solutions are quick, convenient, and inexpensive for massive screening. This study proposes two convolutional neural networks (CNNs)-based segmentation that combines multiscale features by spatial information fusion: a single spatial fusion network (SSF-Net) and a triplet spatial fusion network (TSF-Net). SSF-Net fuses four multiscale spatial information streams. TSF-Net exploits triplet spatial information fusion by early, intermediate, and late fusion to ensure the fine segmentation of retinal pigment signs without preprocessing. TSF-Net creates a valuable difference in performance over SSF-Net. To evaluate SSF-Net and TSF-Net, the open dataset, named Retinal Images for Pigment Signs is utilized with 4-fold cross-validation. The experiment results confirm that SSF-Net and TSF-Net demonstrate superior performance compared to the state-of-the-art methods for the screening and analysis of RP disease.
Adnan Haider, Chanhun Park, Jin Seong Hong, Kang Ryoung Park
Eng. Appl. Artif. Intell.5
2024 Deep learning-based restoration of multi-degraded finger-vein image by non-uniform illumination and noise
abstract
The recognition performance deteriorates if degradation factors including blur, noise, and non-uniform illumination exist in the image when acquiring a finger-vein image. Especially, multiple degradation factors can occur when acquiring the finger-vein image, and they require the image restoration. However, previous flow-based model produced lower image quality than the other restoration models, and diffusion-based model had the disadvantage of slow inference speed. Therefore, this study suggests a deep learning-based generative adversarial network for multi-degraded finger-vein image restoration by non-uniform illumination and noise (MFNN-GAN). It considers multiple degradation factors such as non-uniform illumination and noise. Unlike the existing finger-vein image restoration model, MFNN-GAN is capable of adaptive restoration to multiple degradations. Therefore, even if the illumination by near-infrared (NIR) illuminator of finger-vein recognition device is weak or non-uniform, or the consequent captured image is noisy, good recognition performance can be achieved only by our method without replacing the illuminator or camera sensor. The experimental results obtained using finger-vein open datasets, session 1 images from database version 1 of the Hong Kong Polytechnic University finger-image (HKPU-DB) and finger-vein database of SDUMLA-HMT (SDUMLA-HMT-DB)-based degraded databases. The experimental results show that we obtained the lower equal error rate (EER) of finger-vein recognition using MFNN-GAN compared to other state-of-the-art algorithms.
Jin Seong Hong, Seung Gu Kim, Jung Soo Kim, Kang Ryoung Park
Eng. Appl. Artif. Intell.4
2024 CNCAN: Contrast and normal channel attention network for super-resolution image reconstruction of crops and weeds
abstract
Numerous studies have been performed to apply camera vision technologies in robot-based agriculture and smart farms. In particular, to obtain high accuracy, it is essential to procure high-resolution (HR) images, which requires a high-performance camera. However, due to high costs it is difficult to widely apply the camera in agricultural robots. To overcome this limitation, we propose contrast and normal channel attention network (CNCAN) for super-resolution reconstruction (SR), which is the first research for the accurate semantic segmentation of crops and weeds even with low-resolution (LR) images captured by low-cost and LR camera. Attention block and activation function that considers high frequency and contrast information of images are used in CNCAN, and the residual connection method is applied to improve the learning stability. As a result of experimenting with three open datasets, namely, Bonirob, rice seedling and weed, and crop/weed field image (CWFID) datasets, the mean intersection of union (MIOU) results of semantic segmentation for crops and weeds with SR images through CNCAN were 0.7685, 0.6346, and 0.6931 in the Bonirob, rice seedling and weed, and CWFID datasets, respectively, confirming higher accuracy than other state-of-the-art methods for SR.
Chaeyeong Yun, Su Jin Im, Kang Ryoung Park
Eng. Appl. Artif. Intell.4
2024 A novel convolution transformer-based network for histopathology-image classification using adaptive convolution and dynamic attention
abstract
Renal cell carcinoma (RCC), which is the primary subtype of kidney cancer, is among the leading causes of cancer. Recent breakthroughs in computer vision, particularly deep learning, have revolutionized the analysis of histopathology images, thus providing potential solutions for tasks such as the grading of renal cell carcinoma. Nevertheless, the multitude of available neural network architectures and the absence of systematic evaluations render it challenging to identify optimal models and training configurations for distinct histopathology classification tasks. Hence, we propose a novel hybrid model that effectively combines the advantages of vision transformers and convolutional neural networks. The proposed method, which is named the renal cancer grading network, comprises two essential components: an adaptive convolution (AC) block and a dynamic attention (DA) block. The AC block emphasizes efficient feature extraction and spatial representation learning via intelligently designed convolutional operations. The DA block, which is constructed on the features of the AC block, is a crucial module for histopathology-image classification. It introduces a dynamic attention mechanism and employs a transformer encoder to refine learned representations. Experiments were conducted on four publicly available histopathology datasets: RCC dataset of Kasturba medical college (KMC), colorectal cancer histology (CRCH), break cancer histology (BreakHis) and colon cancer histopathology dataset (CCH). The proposed method demonstrated an accuracy of 90.62%, precision of 91.23%, recall of 90.63%, and a weighted harmonic mean of precision and recall (F1-score) of 90.92 on the KMC dataset. Similarly, the proposed method demonstrates consistent accuracy (weighted average F1-score of 99%) on the CRCH dataset, recognition rate of 88.30% on the BreakHis dataset, and an accuracy of 99.7% on CCH dataset. These results confirm that our method outperforms the state-of-the-art methods, thus demonstrating its effectiveness and robustness across various datasets.
Tahir Mahmood 0003, Abdul Wahid 0007, Jin Seong Hong, Seung Gu Kim, Kang Ryoung Park
Eng. Appl. Artif. Intell.5
2024 CN4SRSS: Combined network for super-resolution reconstruction and semantic segmentation in frontal-viewing camera images of vehicle
abstract
Recently, the importance of semantic segmentation research for scene understanding in frontal viewing camera images of autonomous vehicles has increased. The existing state-of-the-art (SOTA) methods for semantic segmentation exhibit high accuracy for high-resolution images and low-resolution (LR) images without degradation factors of blur and noise. Owing to the nature of vehicles, the need is increasing for the pre-judgment of emergencies through the accurate semantic segmentation of LR images with the degradation factors acquired by low-cost camera at far distance. However, no research exists on super-resolution reconstruction (SR)-based semantic segmentation of LR images with degradation factors. Therefore, this study proposes a novel combined network for a super-resolution reconstruction and semantic segmentation (CN4SRSS) framework based on attention and re-focus network (ARNet), which exhibits low computational cost and high semantic segmentation accuracy. The experimental results using LR image datasets based on CamVid and Minicity datasets, which are open databases, show that the semantic segmentation accuracy (pixel accuracy) based on the proposed CN4SRSS and DeepLab v3 + is 93.14% and 89.48%, respectively. Particularly, the proposed method shows higher accuracy when compared to the SOTA methods. Furthermore, the proposed method has been confirmed that requires lower computational cost in terms of the number of parameters, memory usage, number of multi-adds calculation, and floating-point operations per second (FLOPs) than the SOTA methods.
Kyung Bong Ryu, Seon Jong Kang, Seong In Jeong, Min Su Jeong, Kang Ryoung Park
Eng. Appl. Artif. Intell.5
2024 Dilated multilevel fused network for virus classification using transmission electron microscopy images
abstract
Previous studies have demonstrated significant performance in the field of virus classification; however, they focused on the classification of a small number of virus classes, with a maximum of 16 classes. To address this limitation, this study aims to create a deep learning-based network that outperforms the state-of-the-art (SOTA) models for the classification of 22 different virus classes with the fewest possible trainable parameters. We introduce an automatic identification system for virus classes based on our classification-driven retrieval framework. The proposed dilated multilevel fused network (DMLF-Net) utilizes the multilevel feature fusion concept within a network to exploit more abstract features for microscopic data analysis. A multi-stage training strategy was applied to achieve optimal model convergence without overfitting the training data. We evaluated the performance of the DMLF-Net on three open databases including two virus datasets and one bacteria species dataset. The results demonstrated an accuracy of 89.89%, a weighted harmonic mean of precision and recall (F1-score) of 83.39%, and an area under the curve (AUC) of 92.50% for the 1st virus dataset. For the 2nd virus dataset, the accuracy was 80.70%, the F1-score was 81.20%, and the AUC was 86.20%. For the 3rd bacteria species dataset, the accuracy was 95.93% and the F1-score was 96.24%. DMLF-Net outperforms SOTA methods in terms of classification accuracy while utilizing nearly 5.3 times fewer trainable parameters (25.5 million) compared to the second-best model, visual geometry group (VGG)16 (134.3 million).
Haseeb Sultan, Jin Seong Hong, Seung Gu Kim, Rehan Akram, Hafiz Ali Hamza Gondal, Muhammad Hamza Tariq, Kang Ryoung Park
Eng. Appl. Artif. Intell.8
2024 Multi-path residual attention network for cancer diagnosis robust to a small number of training data of microscopic hyperspectral pathological images
abstract
Duct cancer is a malignant disease with higher mortality rates in males than in females, emphasizing the need for early diagnosis to improve treatment outcomes. Although various imaging modalities such as magnetic resonance imaging (MRI) and computed tomography scan (CT-scan) have been used for pathological analysis, hyperspectral imaging stands out as a promising approach, especially when combined with deep learning techniques. Hyperspectral imaging provides detailed information on tissue composition and biochemical properties, enabling better distinction between cancerous and healthy tissues. Although previous research based on hyperspectral imaging shows high accuracy, no previous research has used a small amount of training data, despite this being the usual case in medical image applications. Therefore, we propose a multi-path residual attention network (MRA-Net) with chunked residual channel attention (CRCA), which is a novel deep learning model specifically designed to address the challenges posed by limited training data, with a particular focus on using hyperspectral images. By leveraging the unique spectral information provided by hyperspectral imaging, MRA-Net extracts distinctive features, enhancing its ability to differentiate between cancerous and healthy tissues. We conducted the training and validation of our model using a publicly accessible dataset, resulting in an accuracy of 84.31% and a weighted harmonic mean of precision and recall (F1 score) of 84.29%, demonstrating its state-of-the-art performance compared to existing methods.
Abdul Wahid 0007, Tahir Mahmood 0003, Jin Seong Hong, Seung Gu Kim, Nadeem Ullah, Rehan Akram, Kang Ryoung Park
Eng. Appl. Artif. Intell.7
2023 Multi-scale feature retention and aggregation for colorectal cancer diagnosis using gastrointestinal images
abstract
Colonoscopy is considered the gold standard for colorectal cancer diagnosis and prognosis. However, existing methods are less accurate and prone to overlooking lesions during gastrointestinal endoscopic examinations. Computer-assisted diagnosis combined with robot-assisted minimally invasive surgery (RMIS) can significantly help medical practitioners detect and treat lesions. Therefore, two novel architectures are developed for polyp and surgical instrument segmentation to aid colorectal cancer diagnosis, assessment, and treatment. Colorectal cancer segmentation network (CCS-Net) is the base network used in this study. It uses the maximum convolutional layers near the input image for effective feature extraction from low-level information. In addition, CCS-Net uses an efficient feature upsampling unit to efficiently increase the input spatial features’ map size. Hence, CCS-Net is capable of providing a fair performance with satisfactory computational efficiency The multi-scale feature retention and aggregation network (MFRA-Net) is the final network in this study. MFRA-Net is developed to improve the segmentation accuracy of the CCS-Net further as it uses multi-scale feature retention to retain low-level spatial features and transfers them to deep stages of the network. MFRA-Net also combines multi-scale high-strided low-level information with high-level information to boost network segmentation performance. Finally, all the transferred multi-scale features from the early stages of the network are aggregated with high-level features in the deep levels of the network. This multi-scale feature retention and aggregation mechanism enables the network to maintain a better segmentation performance compared with other methods even with challenging blur, specular reflection, low contrast, and high variation cases. We evaluated both architectures on four challenging datasets: Kvasir-SEG, CVC-ClinicDB, Kvasir-Instrument, and the UW-Sinus-Surgery-Live dataset. The proposed method achieves dice similarity coefficients of 95.98%, 94.19%, 92.81%, and 88.57% for the CVC-ClinicDB, Kvasir-SEG, Kvasir-Instrument, and UW-Sinus-Surgery-Live datasets. The proposed method achieves superior segmentation performance compared with state-of-the-art methods and requires only 4.9 million trainable parameters for complete training. Therefore, the proposed networks can effectively assist health professionals in surgical procedures and colorectal cancer diagnosis through surgical instruments and polyp segmentation, respectively.
Adnan Haider, Se Hyun Nam, Jin Seong Hong, Haseeb Sultan, Kang Ryoung Park
Eng. Appl. Artif. Intell.6
2023 LCW-Net: Low-light-image-based crop and weed segmentation network using attention module in two decoders
abstract
Crop segmentation using cameras is commonly used in large agricultural areas, but the time and duration of crop harvesting varies in large farms. Considering this situation, there is a need for low-light image-based segmentation of crop and weed images for late-time harvesting, but no prior research has considered this. As a first study on this topic, we propose a low-light image-based crop and weed segmentation network (LCW-Net) that uses an attention module in two decoders to perform only one step without restoration of low-light images. We also design a loss function to accurately segment regions of objects, crops, and weeds in low-light images to avoid training overfitting and balance the learning task for object, crop, and weed segmentation. There are no existing low-light public databases, and it is difficult to obtain ground truth segmentation information for self-collected database in low-light environments. Therefore, we experimented with converting two public databases, the crop and weed field image dataset (CWFID) and BoniRob dataset, into low-light datasets. The experimental results showed that the mean intersection of unions (mIoUs) of segmentation for crops and weeds were 0.8718 and 0.8693 for the BoniRob dataset, respectively, and 0.8337 and 0.8221 for the CWFID dataset, respectively, indicating that LCW-Net outperforms the state-of-the-art methods.
Yu Hwan Kim, Chaeyeong Yun, Su Jin Im, Kang Ryoung Park
Eng. Appl. Artif. Intell.5
2023 CFFR-Net: A channel-wise features fusion and recalibration network for surgical instruments segmentation
abstract
Surgical instrument segmentation plays a crucial role in robot-assisted surgery by furnishing essential information about instrument location and orientation. This information not only enhances surgical planning but also augments the precision and safety of procedures. Despite promising strides in recent research on surgical instrument segmentation, accuracy still faces obstacles due to local feature processing limitations, surgical environment complexity, and instrument morphological variability. To address these challenges, we introduced the channel-wise features fusion and recalibration network (CFFR-Net). This network utilizes a dual-stream mechanism, combining a context-guided block and dense block for feature extraction. The context-guided block captures a variety of contextual information by using different dilation rates. Additionally, CFFR-Net employs a fusion mechanism that harmonizes context-guided and dense streams. This integration, along with the inclusion of Squeeze-and-Excitation attention, enhances both the precision and robustness of semantic instrument segmentation. We performed experiments using two publicly available datasets for surgical instrument segmentation: the Kvasir-instrument and Endovis2017 datasets. The results of these experiments were highly encouraging, as our proposed model exhibited remarkable performance on both datasets compared to the state-of-the-art methods. On the Kvasir-instrument set, our model achieved a Dice score of 95.84% and mean intersection over union (mIOU) value of 92.40%. Similarly, on the Endovis2017 set, it obtained a Dice score of 95.47% and mIOU value of 93.02%.
Tahir Mahmood 0003, Jin Seong Hong, Nadeem Ullah, Abdul Wahid 0007, Kang Ryoung Park
Eng. Appl. Artif. Intell.6
2023 DCDA-Net: Dual-convolutional dual-attention network for obstructive sleep apnea diagnosis from single-lead electrocardiograms
abstract
Obstructive sleep apnea (OSA) is a breathing-related chronic disease in which the soft palate and tongue collapse and block the upper airway for at least 10 s during sleep. It can lead to many heart diseases such as hypertension, myocardial infarction, and coronary heart syndrome if not detected early. Artificial intelligence has facilitated the diagnosis of many diseases in healthcare. Polysomnography is a widely used but unpleasant, time-consuming, technically demanding, and financially expensive procedure to detect OSA. Some previous methods have detected OSA using time-domain information from an electrocardiogram (ECG), whereas others have used frequency-domain information. The limitations of these two approaches can be handled using the data’s time–frequency representation. Nevertheless, there is room for enhancing the detection accuracy of OSA using the time–frequency representation approach. Therefore, we propose a novel technique that takes the ECG signal and detects R-peaks from the QRS complexes. Afterward, we interpolate those R-peaks by linear interpolation and get an interpolated-R signal. Then we magnify the interpolated-R signal corresponding to the apnea and normal frequency ranges. After magnification in the time domain, we transformed the magnified version into a scalogram. We also transformed the original one-minute ECG signal into a spectrogram after denoising. Overall, we used ECG signals to generate scalograms and spectrograms for 2 dimensional convolutional neural network (2D CNN) to classify obstructive sleep apnea. For apnea classification, we proposed a dual convolutional dual attention network (DCDA-Net) that includes a dual convolutionally modified inception module, a spatial attention module, and a channel attention module. Finally, we apply a support vector machine to the probability scores obtained from DCDA-Net based on the scalogram and spectrogram. Extensive experimental results using the open PhysioNet apnea ECG dataset confirm the effectiveness of our method in terms of accuracy and F1 score of 98% and 97.5%, respectively, which outperforms state-of-the-art methods.
Nadeem Ullah, Tahir Mahmood 0003, Seung Gu Kim, Se Hyun Nam, Haseeb Sultan, Kang Ryoung Park
Eng. Appl. Artif. Intell.6
2023 CAM-CAN: Class activation map-based categorical adversarial network
abstract
Numerous studies have investigated image classification. In particular, recent methods based on deep learning have exhibited high accuracies. However, various existing state-of-the-art methods based on deep learning show different accuracies depending on the database and environment. Accordingly, different deep learning models need to be used in image classification studies according to the database, environment, and research field. This study investigated a technique to increase the accuracy of the existing deep learning-based models. The proposed method was applied to various existing state-of-the-art methods. In the proposed method, a convolution neural network (CNN) is trained using the classification activation map (CAM) to focus on specific areas in the input image. The CAM image is used as the ground-truth image. Furthermore, the concept of the CAM-based categorical adversarial network (CAM-CAN), in which the CNN is trained based on a generative adversarial network, is proposed in this paper. An action recognition experiment was performed using the self-collected Dongguk thermal image database (DTh-DB) and open database, and the results revealed that the accuracies of the existing state-of-the-art methods significantly increased after applying the proposed method. For instance, the accuracies obtained using the DTh-DB, TPR, PPV, ACC, and F1 with the conventional DenseNet201 model were 80.14%, 75.28%, 96.0%, and 75.91%, respectively. After applying the proposed method, the accuracies increased to 86.53%, 89.90%, 97.64%, and 85.84%, respectively.
Ganbayar Batchuluun, Jiho Choi, Kang Ryoung Park
Expert Syst. Appl.3
2023 Volumetric Model Genesis in Medical Domain for the Analysis of Multimodality 2-D/3-D Data Based on the Aggregation of Multilevel Features
abstract
The automatic and accurate classification of medical imaging data has potential applications in computer-aided disease diagnosis, prognosis, and treatment. However, it remains a challenge to optimize recent deep learning algorithms in medical domain for the accurate classification of large-scale 3D volumetric data. To address these challenges, we propose an efficient deep volumetric classification network based on the aggregation of multilevel deep features for accurate classification of large-scale medical 2D/3D imaging data. To perform a detailed quantitative analysis of our method, 26 different datasets were fused to construct a single large-scale multimodal database that comprises a total of seventy different classes, including 151,095 data samples. Additionally, 15 different baseline methods were configured under the same experimental protocol for volumetric model genesis and extensive performance comparison with our method. The experimental results of our method exhibited promising performance as area under the curve of 93.66% and outperformed various state-of-the-art methods.
Muhammad Owais, Se Woon Cho, Kang Ryoung Park
IEEE Trans. Ind. Informatics3
2022 Detecting retinal vasculature as a key biomarker for deep Learning-based intelligent screening and analysis of diabetic and hypertensive retinopathy
Adnan Haider, Young Won Lee, Kang Ryoung Park
Expert Syst. Appl.4
2022 Artificial Intelligence-based computer-aided diagnosis of glaucoma using retinal fundus images
abstract
Glaucoma is one of the most common chronic diseases that may lead to irreversible vision loss. The number of patients with permanent vision loss due to glaucoma is expected to increase at an alarming rate in the near future. A considerable amount of research is being conducted on computer-aided diagnosis for glaucoma. Segmentation of the optic cup (OC) and optic disc (OD) is usually performed to distinguish glaucomatous and non-glaucomatous cases in retinal fundus images. However, the OC boundaries are quite non-distinctive; consequently, the accurate segmentation of the OC is substantially challenging, and the OD segmentation performance also needs to be improved. To overcome this problem, we propose two networks, separable linked segmentation network (SLS-Net) and separable linked segmentation residual network (SLSR-Net), for accurate pixel-wise segmentation of the OC and OD. In SLS-Net and SLSR-Net, a large final feature map can be maintained in our networks, which enhances the OC and OD segmentation performance by minimizing the spatial information loss. SLSR-Net employs external residual connections for feature empowerment. Both proposed networks comprise a separable convolutional link to enhance computational efficiency and reduce the cost of network. Even with a few trainable parameters, the proposed architecture is capable of providing high segmentation accuracy. The segmentation performances of the proposed networks were evaluated on four publicly available retinal fundus image datasets: Drishti-GS, REFUGE, Rim-One-r3, and Drions-DB which confirmed that our networks outperformed the state-of-the-art segmentation architectures.
Adnan Haider, Min Beom Lee, Muhammad Owais, Tahir Mahmood 0003, Haseeb Sultan, Kang Ryoung Park
Expert Syst. Appl.7
2022 GRA-GAN: Generative adversarial network for image style transfer of Gender, Race, and age
abstract
Despite a large amount of available data, the datasets that have been recently used in studies on age estimation still entail the age class imbalance problem owing to different age distributions of race or gender. This results in overfitting in which training data aligns toward one side and ultimately reduces the generality of age estimation. Same problems can occur in the cases of race and gender recognition. This problem can be solved if age images that were insufficient in a previously trained distribution or race and gender information that was not considered in the previously trained distribution can be newly created as images that are identical to the previously trained distribution. Therefore, we propose a race, age, and gender image transformation technique by a generative adversarial network for image style transfer of gender, race, and age (GRA-GAN) based on channel-wise and multiplication-based information fusion of encoder and decoder features. Experiments using four open databases (MORPH, AAF, AFAD, and UTK) indicated that our method outperformed the state-of-the-art methods.
Yu Hwan Kim, Se Hyun Nam, Seung Baek Hong, Kang Ryoung Park
Expert Syst. Appl.4
2022 DSRD-Net: Dual-stream residual dense network for semantic segmentation of instruments in robot-assisted surgery
abstract
In conventional robot-assisted minimally invasive procedures (RMIS), surgeons have narrow visual and complex working spaces, along with specular reflection, blood, camera-lens fogging, and complex backgrounds, which increase the risk of human error and tissue damage. The use of deep learning-based techniques can decrease these risks by providing segmented instruments, real-time tracking, pose estimation, and surgeons’ skill assessment. Recently, several deep learning-based methods have been proposed for surgical instrument segmentation. These methods have shown significant performance for the RMIS. However, we found that most of these methods still have scope for improvement in terms of accuracy, robustness, and computational cost. In addition, gastrointestinal pathologies have not been explored in previous studies. Therefore, we propose a dual-stream residual dense network (DSRD-Net), an accurate and robust deep learning-based surgical instrument segmentation method that mainly utilizes the strength of residual, dense, and atrous spatial pyramid pooling architectures. Our proposed method was tested on publicly available gastrointestinal endoscopy (the Kvasir-Instrument Dataset) and abdominal porcine procedures datasets (The 2017 Robotic Instrument Segmentation Challenge Dataset). The experimental results show that the proposed method outperforms the state-of-the-art methods.
Tahir Mahmood 0003, Se Woon Cho, Kang Ryoung Park
Expert Syst. Appl.3
2022 DMDF-Net: Dual multiscale dilated fusion network for accurate segmentation of lesions related to COVID-19 in lung radiographic scans
abstract
The recent disaster of COVID-19 has brought the whole world to the verge of devastation because of its highly transmissible nature. In this pandemic, radiographic imaging modalities, particularly, computed tomography (CT), have shown remarkable performance for the effective diagnosis of this virus. However, the diagnostic assessment of CT data is a human-dependent process that requires sufficient time by expert radiologists. Recent developments in artificial intelligence have substituted several personal diagnostic procedures with computer-aided diagnosis (CAD) methods that can make an effective diagnosis, even in real time. In response to COVID-19, various CAD methods have been developed in the literature, which can detect and localize infectious regions in chest CT images. However, most existing methods do not provide cross-data analysis, which is an essential measure for assessing the generality of a CAD method. A few studies have performed cross-data analysis in their methods. Nevertheless, these methods show limited results in real-world scenarios without addressing generality issues. Therefore, in this study, we attempt to address generality issues and propose a deep learning-based CAD solution for the diagnosis of COVID-19 lesions from chest CT images. We propose a dual multiscale dilated fusion network (DMDF-Net) for the robust segmentation of small lesions in a given CT image. The proposed network mainly utilizes the strength of multiscale deep features fusion inside the encoder and decoder modules in a mutually beneficial manner to achieve superior segmentation performance. Additional pre- and post-processing steps are introduced in the proposed method to address the generality issues and further improve the diagnostic performance. Mainly, the concept of post-region of interest (ROI) fusion is introduced in the post-processing step, which reduces the number of false-positives and provides a way to accurately quantify the infected area of lung. Consequently, the proposed framework outperforms various state-of-the-art methods by accomplishing superior infection segmentation results with an average Dice similarity coefficient of 75.7%, Intersection over Union of 67.22%, Average Precision of 69.92%, Sensitivity of 72.78%, Specificity of 99.79%, Enhance-Alignment Measure of 91.11%, and Mean Absolute Error of 0.026.
Muhammad Owais, Na Rae Baek, Kang Ryoung Park
Expert Syst. Appl.3
2022 Deep Features Aggregation-Based Joint Segmentation of Cytoplasm and Nuclei in White Blood Cells
abstract
White blood cells (WBCs), also known as leukocytes, are one of the valuable parts of the blood and immune system. Typically, pathologists use microscope for the manual inspection of blood smears which is a time-consuming, error-prone, and labor-intensive procedure. To address these issues, we present two novel shallow networks: a leukocyte deep segmentation network (LDS-Net) and leukocyte deep aggregation segmentation network (LDAS-Net) for the joint segmentation of cytoplasm and nuclei in WBC images. LDS-Net is a shallow architecture with three downsampling stages and seven convolution layers. LDAS-Net is an extended version of LDS-Net that utilizes a novel pool-less low-level information transfer bridge to transfer low-level information to the deep layers of the network. This information is aggregated with deep features in a dense feature concatenation block to achieve accurate cytoplasm and nuclei joint segmentation. We evaluated our developed architectures on four WBC publicly available datasets. For cytoplasmic segmentation in WBCs, the proposed method achieved the dice coefficients of 98.97%, 99.0%, 96.05%, and 98.79% on Datasets 1, 2, 3, and 4, respectively. For nuclei segmentation, the dice coefficients of 96.35% and 98.09% are achieved for Datasets 1 and 2, respectively. Proposed method outperforms state-of-the-art methods with superior computational efficiency and requires only 6.5 million trainable parameters.
Adnan Haider, Young Won Lee, Kang Ryoung Park
IEEE J. Biomed. Health Informatics4
2021 Multilevel Deep-Aggregated Boosted Network to Recognize COVID-19 Infection from Large-Scale Heterogeneous Radiographic Data
abstract
In the present epidemic of the coronavirus disease 2019 (COVID-19), radiological imaging modalities, such as X-ray and computed tomography (CT), have been identified as effective diagnostic tools. However, the subjective assessment of radiographic examination is a time-consuming task and demands expert radiologists. Recent advancements in artificial intelligence have enhanced the diagnostic power of computer-aided diagnosis (CAD) tools and assisted medical specialists in making efficient diagnostic decisions. In this work, we propose an optimal multilevel deep-aggregated boosted network to recognize COVID-19 infection from heterogeneous radiographic data, including X-ray and CT images. Our method leverages multilevel deep-aggregated features and multistage training via a mutually beneficial approach to maximize the overall CAD performance. To improve the interpretation of CAD predictions, these multilevel deep features are visualized as additional outputs that can assist radiologists in validating the CAD results. A total of six publicly available datasets were fused to build a single large-scale heterogeneous radiographic collection that was used to analyze the performance of the proposed technique and other baseline methods. To preserve generality of our method, we selected different patient data for training, validation, and testing, and consequently, the data of same patient were not included in training, validation, and testing subsets. In addition, fivefold cross-validation was performed in all the experiments for a fair evaluation. Our method exhibits promising performance values of 95.38%, 95.57%, 92.53%, 98.14%, 93.16%, and 98.55% in terms of average accuracy, F-measure, specificity, sensitivity, precision, and area under the curve, respectively and outperforms various state-of-the-art methods.
Muhammad Owais, Young Won Lee, Tahir Mahmood 0003, Adnan Haider, Haseeb Sultan, Kang Ryoung Park
IEEE J. Biomed. Health Informatics6
2020 OR-Skip-Net: Outer residual skip network for skin segmentation in non-ideal situations
Dong Seop Kim, Muhammad Owais, Kang Ryoung Park
Expert Syst. Appl.4
2019 FRED-Net: Fully residual encoder-decoder network for accurate iris segmentation
Dong Seop Kim, Min Beom Lee, Muhammad Owais, Kang Ryoung Park
Expert Syst. Appl.5
2019 Driver's eye-based gaze tracking system by one-point calibration
Hyo Sik Yoon, Hyung Gil Hong, Dong Eun Lee, Kang Ryoung Park
Multim. Tools Appl.4
2018 Body-movement-based human identification using convolutional neural network
Ganbayar Batchuluun, Rizwan Ali Naqvi, Wan Kim, Kang Ryoung Park
Expert Syst. Appl.4
2018 Pedestrian detection based on faster R-CNN in nighttime by fusing deep convolutional features of successive images
Ganbayar Batchuluun, Kang Ryoung Park
Expert Syst. Appl.3
2018 Fuzzy-based estimation of continuous Z-distances and discrete directions of home appliances for NIR camera-based gaze tracking system
Jae Woong Jang, Hwan Heo, Jae Won Bang, Hyung Gil Hong, Rizwan Ali Naqvi, Phong Nguyen 0001, Tien Dat Nguyen, Min Beom Lee, Kang Ryoung Park
Multim. Tools Appl.9
2017 Fuzzy system based human behavior recognition by combining behavior prediction and recognition
Ganbayar Batchuluun, Hyung Gil Hong, Jin Kyu Kang, Kang Ryoung Park
Expert Syst. Appl.5
2017 Periocular-based biometrics robust to eye rotation based on polar coordinates
So Ra Cho, Gi Pyo Nam, Kwang Yong Shin, Tien Dat Nguyen, Tuyen Danh Pham, Eui Chul Lee, Kang Ryoung Park
Multim. Tools Appl.7
2017 Banknote recognition based on optimization of discriminative regions by genetic algorithm with one-dimensional visible-light line sensor
Tuyen Danh Pham, Ki-Wan Kim, Jeonggoo Kang, Kang Ryoung Park
Pattern Recognit.4
2016 Enhanced age estimation by considering the areas of non-skin and the non-uniform illumination of visible light camera sensor
Tien Dat Nguyen, Kang Ryoung Park
Expert Syst. Appl.2
2014 Detecting driver drowsiness using feature-level fusion and user-specific classification
Jaeik Jo, Sung Joo Lee, Kang Ryoung Park, Ig-Jae Kim, Jaihie Kim
Expert Syst. Appl.3
2012 New iris recognition method for noisy iris images
Kwang Yong Shin, Gi Pyo Nam, Dae Sik Jeong, Dal Ho Cho, Byung Jun Kang, Kang Ryoung Park, Jaihie Kim
Pattern Recognit. Lett.6
2011 Age estimation using a hierarchical classifier based on global and local facial features
Sung Eun Choi, Youn Joo Lee, Sung Joo Lee, Kang Ryoung Park, Jaihie Kim
Pattern Recognit.4
2011 A SfM-based 3D face reconstruction method robust to self-occlusion by using a shape conversion matrix
Sung Joo Lee, Kang Ryoung Park, Jaihie Kim
Pattern Recognit.2
2011 Real-Time Gaze Estimator Based on Driver's Head Orientation for Forward Collision Warning System
abstract
This paper presents a vision-based real-time gaze zone estimator based on a driver's head orientation composed of yaw and pitch. Generally, vision-based methods are vulnerable to the wearing of eyeglasses and image variations between day and night. The proposed method is novel in the following four ways: First, the proposed method can work under both day and night conditions and is robust to facial image variation caused by eyeglasses because it only requires simple facial features and not specific features such as eyes, lip corners, and facial contours. Second, an ellipsoidal face model is proposed instead of a cylindrical face model to exactly determine a driver's yaw. Third, we propose new features—the normalized mean and the standard deviation of the horizontal edge projection histogram—to reliably and rapidly estimate a driver's pitch. Fourth, the proposed method obtains an accurate gaze zone by using a support vector machine. Experimental results from 200 000 images showed that the root mean square errors of the estimated yaw and pitch angles are below 7 under both daylight and nighttime conditions. Equivalent results were obtained for drivers with glasses or sunglasses, and 18 gaze zones were accurately estimated using the proposed gaze estimation method.
Sung Joo Lee, Jaeik Jo, Ho Gi Jung, Kang Ryoung Park, Jaihie Kim
IEEE Trans. Intell. Transp. Syst.4
2010 A comparative study of local feature extraction for age estimation
abstract
Many age estimation methods have been proposed for various applications such as Age Specific Human Computer Interaction (ASHCI) system, age simulation system and so on. Because the performance of the age estimation is greatly affected by the aging feature, the aging feature extraction from facial images is very important. The aging features used in previous works can be divided into global and local features. As global features, Active Appearance Models (AAM) was mainly used for age estimation in previous works. However, AAM is not enough to represent local features such as wrinkle and skin. Therefore, the research about local features is required. In previous works, local features were generally used to determine age group rather than detailed age, and the comparative studies about various local features extraction methods were not conducted. In this paper, the performances of sobel filter, difference image between original and smoothed image, ideal high pass filter (IHPF), gaussian high pass filter (GHPF), Haar and Daubechies discrete wavelet transform (DWT) are compared for extracting local features and detailed age estimation is performed by Support Vector Regression (SVR) on BERC and PAL aging database. The experimental results show that local features can be used for detailed age estimation and GHPF gives a better performance than other methods.
Sung Eun Choi, Youn Joo Lee, Sung Joo Lee, Kang Ryoung Park, Jaihie Kim
ICARCV4
2010 A new iris segmentation method for non-ideal iris images
Dae Sik Jeong, Jae Won Hwang, Byung Jun Kang, Kang Ryoung Park, Chee Sun Won, Dong Kwon Park, Jaihie Kim
Image Vis. Comput.4
2010 Finger vein recognition using weighted local binary pattern code based on a support vector machine
abstract
Finger vein recognition is a biometric technique which identifies individuals using their unique finger vein patterns. It is reported to have a high accuracy and rapid processing speed. In addition, it is impossible to steal a vein pattern located inside the finger. We propose a new identification method of finger vascular patterns using a weighted local binary pattern (LBP) and support vector machine (SVM). This research is novel in the following three ways. First, holistic codes are extracted through the LBP method without using a vein detection procedure. This reduces the processing time and the complexities in detecting finger vein patterns. Second, we classify the local areas from which the LBP codes are extracted into three categories based on the SVM classifier: local areas that include a large amount (LA), a medium amount (MA), and a small amount (SA) of vein patterns. Third, different weights are assigned to the extracted LBP code according to the local area type (LA, MA, and SA) from which the LBP codes were extracted. The optimal weights are determined empirically in terms of the accuracy of the finger vein recognition. Experimental results show that our equal error rate (EER) is significantly lower compared to that without the proposed method or using a conventional method.
Hyeon Chang Lee, Byung Jun Kang, Eui Chul Lee, Kang Ryoung Park
J. Zhejiang Univ. Sci. C4
2010 A new multi-unit iris authentication based on quality assessment and score level fusion for mobile phones
Byung Jun Kang, Kang Ryoung Park
Mach. Vis. Appl.2
2009 A robust eye gaze tracking method based on a virtual eyeball model
Eui Chul Lee, Kang Ryoung Park
Mach. Vis. Appl.2
2009 A comparative study of facial appearance modeling methods for active appearance models
Sung Joo Lee, Kang Ryoung Park, Jaihie Kim
Pattern Recognit. Lett.2
2008 A study on eyelid localization considering image focus for iris recognition
Young Kyoon Jang, Byung Jun Kang, Kang Ryoung Park
Pattern Recognit. Lett.3
2008 New focus assessment method for iris recognition systems
Jain Jang, Kang Ryoung Park, Jaihie Kim, Yillbyung Lee
Pattern Recognit. Lett.2
2008 A robust gaze detection method by compensating for facial movements based on corneal specularities
You Jin Ko, Eui Chul Lee, Kang Ryoung Park
Pattern Recognit. Lett.3
2008 A New Method for Generating an Invariant Iris Private Key Based on the Fuzzy Vault System
abstract
Cryptographic systems have been widely used in many information security applications. One main challenge that these systems have faced has been how to protect private keys from attackers. Recently, biometric cryptosystems have been introduced as a reliable way of concealing private keys by using biometric data. A fuzzy vault refers to a biometric cryptosystem that can be used to effectively protect private keys and to release them only when legitimate users enter their biometric data. In biometric systems, a critical problem is storing biometric templates in a database. However, fuzzy vault systems do not need to directly store these templates since they are combined with private keys by using cryptography. Previous fuzzy vault systems were designed by using fingerprint, face, and so on. However, there has been no attempt to implement a fuzzy vault system that used an iris. In biometric applications, it is widely known that an iris can discriminate between persons better than other biometric modalities. In this paper, we propose a reliable fuzzy vault system based on local iris features. We extracted multiple iris features from multiple local regions in a given iris image, and the exact values of the unordered set were then produced using the clustering method. To align the iris templates with the new input iris data, a shift-matching technique was applied. Experimental results showed that 128-bit private keys were securely and robustly generated by using any given iris data without requiring prealignment.
Youn Joo Lee, Kang Ryoung Park, Sung Joo Lee, Kwanghyuk Bae, Jaihie Kim
IEEE Trans. Syst. Man Cybern. Part B2
2007 A robust eyelash detection based on iris focus assessment
Byung Jun Kang, Kang Ryoung Park
Pattern Recognit. Lett.2
2007 Iris recognition based on score level fusion by using SVM
Hyun-Ae Park, Kang Ryoung Park
Pattern Recognit. Lett.2
2007 Real-Time Image Restoration for Iris Recognition Systems
abstract
In the field of biometrics, it has been reported that iris recognition techniques have shown high levels of accuracy because unique patterns of the human iris, which has very many degrees of freedom, are used. However, because conventional iris cameras have small depth-of-field (DOF) areas, input iris images can easily be blurred, which can lead to lower recognition performance, since iris patterns are transformed by the blurring caused by optical defocusing. To overcome these problems, an autofocusing camera can be used. However, this inevitably increases the cost, size, and complexity of the system. Therefore, we propose a new real-time iris image-restoration method, which can increase the camera's DOF without requiring any additional hardware. This paper presents five novelties as compared to previous works: 1) by excluding eyelash and eyelid regions, it is possible to obtain more accurate focus scores from input iris images; 2) the parameter of the point spread function (PSF) can be estimated in terms of camera optics and measured focus scores; therefore, parameter estimation is more accurate than it has been in previous research; 3) because the PSF parameter can be obtained by using a predetermined equation, iris image restoration can be done in real-time; 4) by using a constrained least square (CLS) restoration filter that considers noise, performance can be greatly enhanced; and 5) restoration accuracy can also be enhanced by estimating the weight value of the noise-regularization term of the CLS filter according to the amount of image blurring. Experimental results showed that iris recognition errors when using the proposed restoration method were greatly reduced as compared to those results achieved without restoration or those achieved using previous iris-restoration methods.
Byung Jun Kang, Kang Ryoung Park
IEEE Trans. Syst. Man Cybern. Part B2
2007 A Real-Time Gaze Position Estimation Method Based on a 3-D Eye Model
abstract
This paper proposes a new gaze-detection method based on a 3-D eye position and the gaze vector of the human eyeball. Seven new developments compared to previous works are presented. First, a method of using three camera systems, i.e., one wide-view camera and two narrow-view cameras, is proposed. The narrow-view cameras use autozooming, focusing, panning, and tilting procedures (based on the detected 3-D eye feature position) for gaze detection. This allows for natural head and eye movement by users. Second, in previous conventional gaze-detection research, one or multiple illuminators were used. These studies did not consider specular reflection (SR) problems, which were caused by the illuminators when working with users who wore glasses. To solve this problem, a method based on dual illuminators is proposed in this paper. Third, the proposed method does not require user-dependent calibration, so all procedures for detecting gaze position operate automatically without human intervention. Fourth, the intrinsic characteristics of the human eye, such as the disparity between the pupillary and the visual axes in order to obtain accurate gaze positions, are considered. Fifth, all the coordinates obtained by the left and right narrow-view cameras, as well as the wide-view camera coordinates and the monitor coordinates, are unified. This simplifies the complex 3-D converting calculation and allows for calculation of the 3-D feature position and gaze position on the monitor. Sixth, to upgrade eye-detection performance when using a wide-view camera, the adaptive-selection method is used. This involves an IR-LED on/off scheme, an AdaBoost classifier, and a principle component analysis method based on the number of SR elements. Finally, the proposed method uses an eigenvector matrix (instead of simply averaging six gaze vectors) in order to obtain a more accurate final gaze vector that can compensate for noise. Experimental results show that the root mean square error of gaze detection was about 0.627 cm on a 19-in monitor. The processing speed of the proposed method (used to obtain the gaze position on the monitor) was 32 ms (using a Pentium IV 1.8-GHz PC). It was possible to detect the user's gaze position at real-time speed.
Kang Ryoung Park
IEEE Trans. Syst. Man Cybern. Part B1
2006 Design and Implementation of a Fast Integral Image Rendering Method
Bin-Na-Ra Lee, Yongjoo Cho, Kyoung Shin Park, Sung-Wook Min, Joa Sang Lim, Kang Ryoung Park
ICEC7
2006 Robust Gaze Estimation for Human Computer Interaction
Kang Ryoung Park
PRICAI1
2006 Pupil and Iris Localization for Iris Recognition in Mobile Phones
abstract
Until now, iris recognition has been used in many fields. Recently, there have been attempts to adopt iris recognition technology for the security of mobile phones. For example, in case of bank transaction service by using a mobile phone, using a mobile phone can use high level of security based on iris recognition. In this paper, we propose a new pupil & iris segmentation method apt for the mobile environment. We find the pupil & iris at the same time, using both information of the pupil and iris. And we also use characteristic of the eye image. Experimental result shows that our algorithm has good performance in various images, which include motion or optical blurring, ghost, specular refection and etc. from various environments for iris recognition system.
Dal Ho Cho, Kang Ryoung Park, Dae Woong Rhee, Jonghoon Yang
SNPD2
2005 A Study on Non-intrusive Facial and Eye Gaze Detection
Kang Ryoung Park, Joa Sang Lim
ACIVS1
2005 A Real-Time Iris Image Acquisition Algorithm Based on Specular Reflection and Eye Model
Kang Ryoung Park, Jang-Hee Yoo
ACIVS1
2005 A Study on Fast Iris Image Acquisition Method
Kang Ryoung Park
CAIP1
2005 Real-Time Iris Localization for Iris Recognition in Cellular Phone
abstract
With the increasing need of guaranteeing the security in case of using bank transaction service by using cellular phone, it is required to apply biometrics for the security of cellular phone. Especially, iris recognition is good for cellular phone security because of its reliability and accuracy compared to other biometrics such as face, fingerprint and voice recognition. In this paper, we propose a new pupil and iris localization algorithm, which is apt for cellular phone platform based on detecting dark pupil and corneal specular reflection by changing brightness and contrast value. In addition, we lessen the processing time by excluding floating point operation in our algorithm, which is not apt for ARM CPU of CDMA cellular phone. Results show that our algorithm can be used for real-time iris localization for iris recognition in cellular phone.
Dal Ho Cho, Kang Ryoung Park, Dae Woong Rhee
SNPD2
2005 A real-time focusing algorithm for iris recognition camera
abstract
For fast iris recognition, it is very important to capture the user's focused eye image at fast speed. Previous researchers have used the focusing method which has been applied to general landscape scenes without considering the characteristics of the iris image. So, they take much focusing time, especially in the case of the user with glasses. To overcome such problems, we propose a new iris image acquisition method to capture focused eye images at very fast speed based on corneal specular reflection. Experimental results show that the focusing time for both users with and without glasses averages 480 ms, and we conclude that our method can be used for the real-time iris recognition camera.
Kang Ryoung Park, Jaihie Kim
IEEE Trans. Syst. Man Cybern. Part C1
2004 Gaze Detection by Wide and Narrow View Stereo Camera
Kang Ryoung Park
CIARP1
2004 A study on multi-unit iris recognition
abstract
Iris recognition system has achieved good performance, but it is affected by the quality of input data. In this paper, we propose a multi-unit iris recognition system, which can select the good quality data between multi-unit eye images of the same person. The system is composed of four stages. First, both iris data are captured at the same time. After that the eye image check algorithm rejects noisy and counterfeit data. At the third stage, features are extracted by Daubechies' wavelet. Finally, features are classified by support vector machines (SVM) and Euclidian distance. We select the better accuracy rate between results of two methods. Experiment results involve 1694 eye images of 111 different people and the best accuracy rate is 99.1%.
Jain Jang, Kang Ryoung Park, Jinho Son, Yillbyung Lee
ICARCV2
2004 Real-Time Gaze Detection via Neural Network
Kang Ryoung Park
ICONIP1
2004 CLOVES: A Virtual World Builder for Constructing Virtual Environments for Science Inquiry Learning
Yongjoo Cho, Kyoung Shin Park, Thomas G. Moher, Andrew E. Johnson 0001, Juno Chang, Joa Sang Lim, Dae Woong Rhee, Kang Ryoung Park, Hung Kook Park
ICEC9
2002 Gaze position detection by computing the three dimensional facial positions and motions
Kang Ryoung Park, Jeong Jun Lee, Jaihie Kim
Pattern Recognit.1
1998 Substroke matching by segmenting and merging for online Korean cursive character recognition
abstract
The Korean character is composed of several alphabets in two-dimensional formation and the total number of Korean characters exceeds eleven thousand. Therefore, the previous approaches to Korean cursive characters pay most of their attention to segmenting a character into alphabets accurately. However, it is difficult because the boundaries of alphabets are not apparent in most cases. We propose an alphabet-based method without assuming accurate alphabet segmentation. In the proposed method, a cursive character is segmented into substrokes by a set of segmenting conditions. Then it is matched with the reference substrokes generated from alphabet models and ligatures by segmenting and merging in the process of recognition. Among substrokes, a certain substroke can be either an alphabet itself a part of alphabet or a composite of the alphabet and ligature. We applied the proposed method to 5000 Korean characters and got the result of 83.4% for the first rank and 89.2% for the top 5 result candidates with the speed of 0.17 seconds on average per character on a PC which uses Intel Pentium 90 Mhz CPU.
Kang Ryoung Park, Byung Hwan Jun, Jaihie Kim
ICPR2