Gi Pyo Nam

dblp:23/8267 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-3383-7806ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Navigating Label Ambiguity for Facial Expression Recognition in the Wild
abstract
Facial expression recognition (FER) remains a challenging task due to label ambiguity caused by the subjective nature of facial expressions and noisy samples. Additionally, class imbalance, which is common in real-world datasets, further complicates FER. Although many studies have shown impressive improvements, they typically address only one of these issues, leading to suboptimal results. To tackle both challenges simultaneously, we propose a novel framework called Navigating Label Ambiguity (NLA), which is robust under real-world conditions. The motivation behind NLA is that dynamically estimating and emphasizing ambiguous samples at each iteration helps mitigate noise and class imbalance by reducing the model's bias toward majority classes. To achieve this, NLA consists of two main components: Noise-aware Adaptive Weighting (NAW) and consistency regularization. Specifically, NAW adaptively assigns higher importance to ambiguous samples and lower importance to noisy ones, based on the correlation between the intermediate prediction scores for the ground truth and the nearest negative. Moreover, we incorporate a regularization term to ensure consistent latent distributions. Consequently, NLA enables the model to progressively focus on more challenging ambiguous samples, which primarily belong to the minority class, in the later stages of training. Extensive experiments demonstrate that NLA outperforms existing methods in both overall and mean accuracy, confirming its robustness against noise and class imbalance. To the best of our knowledge, this is the first framework to address both problems simultaneously.
JunGyu Lee 0003, Yeji Choi, Haksub Kim, Ig-Jae Kim, Gi Pyo Nam
AAAI5
2025 VIGFace: Virtual Identity Generation for Privacy-Free Face Recognition Dataset
Min-Cheol Sagong, Gi Pyo Nam, Junghyun Cho, Ig-Jae Kim
ICCV3
2025 Where and What: Contextual Dynamics-Aware Anomaly Detection in Surveillance Videos
abstract
In surveillance environments, detecting anomalies requires understanding the contextual dynamics of the environment, human behaviors, and movements within a scene. Effective anomaly detection must address both the where and what of events, but existing approaches such as unimodal action-based methods or LLM-integrated multimodal frameworks have limitations. These methods either rely on implicit scene information, making it difficult to localize where anomalies occur, or fail to adapt to surveillance specific challenges such as view changes, subtle actions, low light conditions, and crowded scenes. As a result, these challenges hinder accurate detection of what occurs. To overcome these limitations, our system takes advantage of features from a lightweight scene classification model to discern where an event occurs, acquiring explicit location-based context. To identify what events occur, it focuses on atomic actions, which remain underexplored in this field and are better suited to interpreting intricate abnormal behaviors than conventional abstract action features. To achieve robust anomaly detection, the proposed Temporal-Semantic Relationship Network (TSRN) models spatio-temporal relationships among multimodal features and employs a Segment-selective Focal Margin loss (SFML) to effectively address class imbalance, outperforming conventional MIL-based methods. Experimental results on public datasets demonstrate that the proposed system effectively reduces false alarms while maintaining robustness and practicality for real-world surveillance applications.
Deok-Hyun Ahn, Yong-Jin Jo, Dong-Bum Kim, Gi Pyo Nam, Jae-Ho Han, Haksub Kim
IEEE Trans. Image Process.4
2023 Detection of Road Accidents Using Synthetically Generated Multi-Perspective Accident Videos
abstract
Road accidents are often caused by short abnormal events, including traffic violations, abrupt change in vehicular motion, driver fatigue, etc. Observing an accident event from the right camera perspective plays a crucial role while detecting accidents. However, it may not be possible to capture such abnormal events from a limited camera perspective. We present a deep learning framework to analyze the accident events recorded from multiple perspectives. First, we estimate feature similarity in videos recorded from multiple perspectives. We then divided the video samples into high and low feature similarity groups. Next, we extract spatio-temporal features from each group using two-branch DCNNs and fuse them using a rank-based weighted average pooling strategy followed by classification. We present a new road accident video dataset (MP-RAD), where each accident event is synthetically generated and captured from five independent camera perspectives using a computer gaming platform. Most of the existing road accident datasets use egocentric views or they are captured in fixed camera setups. However, our dataset is large and multi-perspective that can be used to validate ITS-related tasks such as accident detection, accident localization, traffic monitoring, etc. The dataset contains 400 accident events with a total of 2000 videos. We provide temporal annotations of all videos. The proposed framework and the dataset have been cross-validated with latest accident detection baselines trained on real-world road accident videos and vice-versa. The sub-optimal detection accuracy obtained using the baselines indicates that the proposed framework and the dataset can be useful for ITS related research. Code and dataset is available at: https://github.com/draxler1/MP-RAD-Dataset-ITS-
Thakare Kamalakar Vijay, Debi Prosad Dogra, Heeseung Choi, Gi Pyo Nam, Ig-Jae Kim
IEEE Trans. Intell. Transp. Syst.4
2022 PPL: Pairwise Prototype Learning for Masked Face Recognition
Gi Pyo Nam, Yu-Jin Hong, Ig-Jae Kim
BMVC2
2021 A 3d Model-Based Approach For Fitting Masks To Faces In The Wild
abstract
Face recognition now requires a large number of labelled masked face images in the era of this unprecedented COVID19 pandemic. Unfortunately, the rapid spread of the virus has left us little time to prepare for such dataset in the wild. To circumvent this issue, we present a 3D model-based approach called WearMask3D for augmenting face images of various poses to the masked face counterparts. Our method proceeds by first fitting a 3D morphable model on the input image, second overlaying the mask surface onto the face model and warping the respective mask texture, and last projecting the 3D mask back to 2D. The mask texture is adapted based on the brightness and resolution of the input image. By working in 3D, our method can produce more natural masked faces of diverse poses from a single mask texture. To compare precisely between different augmentation approaches, we have constructed a dataset comprising masked and unmasked faces with labels called MFW-mini. Experimental results demonstrate WearMask3D1produces more realistic masked faces, and utilizing these images for training leads to state-of-the-art recognition accuracy for masked faces.
Je Hyeong Hong, Hanjo Kim, Gi Pyo Nam, Junghyun Cho, Hyeong-Seok Ko, Ig-Jae Kim
ICIP4
2020 Query-Based Video Synopsis for Intelligent Traffic Monitoring Applications
abstract
Synopsis of a long-duration video has many applications in intelligent transportation systems. It can help to monitor traffic with lesser manpower. However, generating meaningful synopsis of a long-duration video recording can be challenging. Often summarized outputs include redundant contents or activities that may not be helpful to the observer. Moving object trajectories are possible sources of information that can be used to generate the synopsis of long-duration videos. The synopsis generation faces challenges due to object tracking, grouping of the trajectories with respect to activity type, object category, and contextual information, and generating smooth synopsis according to a query. In this paper, we propose a method to generate meaningful and smooth synopsis of long-duration videos according to the users' query. We have tracked moving objects and adopted deep learning to classify the objects into known categories (e.g., car, bike, and pedestrians). We then identify regions in the surveillance scene with the help of unsupervised clustering. Each tube (spatiotemporal object trajectory) is represented by the source and the destination. In the final stage, we take a query from the user and generate the synopsis video by smoothly blending the appropriate tubes over the background frame through energy minimization. The proposed method has been evaluated on two publicly available datasets and our own surveillance datasets. We have compared the method with popular state-of-the-art techniques. The experiments reveal that the proposed method is superior to the existing techniques and it produces visually seamless video synopsis.
Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Renuka Patnaik, Seung-Cheol Lee, Heeseung Choi, Gi Pyo Nam, Ig-Jae Kim
IEEE Trans. Intell. Transp. Syst.7
2017 Periocular-based biometrics robust to eye rotation based on polar coordinates
So Ra Cho, Gi Pyo Nam, Kwang Yong Shin, Tien Dat Nguyen, Tuyen Danh Pham, Eui Chul Lee, Kang Ryoung Park
Multim. Tools Appl.2
2012 New iris recognition method for noisy iris images
Kwang Yong Shin, Gi Pyo Nam, Dae Sik Jeong, Dal Ho Cho, Byung Jun Kang, Kang Ryoung Park, Jaihie Kim
Pattern Recognit. Lett.2