EDBT 2026 Demo / reviewers in the wild / expert
Soo-Hyung Kim
dblp:15/5759
· DBLP profile ↗
88ranked-venue papers
8as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 7 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 19 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 10 · 4 first-authorHuman-computer interaction and ubiquitous computing · 6 · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Computer networks · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring the Effects of Support Type and Task Difficulty for Virtual Assistants as Social CompanionsabstractVirtual assistants (VAs) are increasingly positioned not just as tools, but as potential social companions—capable of offering either emotional or informational support. Yet, how these forms of support should adapt to varying task difficulties and embodiment styles remains underexplored. We conducted two user studies with cognitive and physical tasks to investigate how support type (emotional vs. informational) shapes user perceptions across variations in task difficulty (easy vs. hard) and embodiment (non-embodied vs. embodied). In Study 1, emotional support positively influenced users’ impressions of VA in easy tasks, while informational support was more effective in difficult tasks. In Study 2, participants also preferred emotional support for easy tasks, but differences between support types were less pronounced for difficult tasks. Notably, embodiment exerted no significant influence in either study. These findings underscore the role of context in shaping effective support strategies, offering design insights for VAs as social companions. Sei Kang, Yunsu Lee, Gun A. Lee, Soo-Hyung Kim, Ji-Eun Shin, Seungwon Kim |
CHI | 5 |
| 2026 | Multi-modal adaptive empathy assessment in online dyadic interaction using bi-directional multi-layer perceptron-mixer and dynamic weights fusionabstractEmpathy modeling in online interactions presents a significant challenge due to its dynamic, context-sensitive, and bidirectional nature. To address these complexities, we propose the Bi-directional MLP-Mixer (Bi-Mixer) and dynamic weights fusion model—a novel neural architecture that captures temporal dependencies in both forward and reverse directions and performs dynamic multi-modal fusion using a cross-attention mechanism. The model adaptively reweights input from visual, audio, text, and biological modalities based on contextual cues, enabling more accurate real-time empathy prediction. To support the training and evaluation of our proposed model, we build the Multi-modal online interaction EMPathy (Multi-EMP) dataset, consisting of unscripted dyadic conversations recorded through online video conferencing. The dataset includes four synchronized modalities: video, audio, text, and bio-signals such as electrodermal activity, blood volume pulse, temperature, and metabolic equivalent of task. It enables dual empathy assessment based on both explicit self-reports and computed emotional alignment between speaker and listener. Our contributions include: (1) the Bi-Mixer and dynamic weights fusion model for bidirectional and adaptive multi-modal representation learning, (2) the release of a comprehensive multi-modal dataset for naturalistic empathy research, and (3) a dual-perspective framework for empathy evaluation. Eunchae Lim, Hwaryung Lee, Ji-Eun Shin, Soo-Hyung Kim, Seungwon Kim, Aera Kim |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Multi-stage wavelet-attention deep learning for brain MRI classification
Tuan-Khoi Tran, Soo-Hyung Kim, Myungeun Lee |
Multim. Syst. | 2 |
| 2026 | Leveraging deep visual geometry group network for facial emotion recognition through RGB and thermal image fusion
Tuan-Khoi Tran, Soo-Hyung Kim, Seung-Won Kim, Ji-Eun Shin |
Multim. Tools Appl. | 2 |
| 2026 | AttCo: Attention-based co-Learning fusion of deep feature representation for medical image segmentation using multimodalityabstractAccurate tissue segmentation is crucial for advancing healthcare, particularly in disease prediction and treatment planning. Precisely identifying abnormal tissue locations is a critical step for clinical analysis. While medical image segmentation increasingly utilizes multimodal and three-dimensional (3D) information to capture spatial relationships, current methods often struggle to effectively learn complementary information from multiple inputs, especially with complex 3D structures. In this study, we introduce AttCo, a novel multimodal semantic segmentation network built upon an attention-based co-learning fusion of deep feature representations. AttCo first employs multiple encoder branches to extract unimodal 3D representations from each imaging modality. These unimodal representations are then processed by a co-learning fusion module, which integrates both intra-modality (using SEAT) and inter-modality (using OSCAT) feature learning components. This dual approach ensures the capture of intricate interactions within each modality and across different modalities. Finally, these fused features pass through an up-sampling module to generate 3D segmented tumor maps. Our end-to-end deep network effectively addresses two key aspects: (i) extracting robust unimodal 3D representations and (ii) exploiting comprehensive inter- and intra-modality feature interactions. Experimental results demonstrate that AttCo significantly outperforms competing methods in terms of Dice score across various datasets. The source code can be found here https://github.com/duyphuongcri/AttCo. Duy-Phuong Dao, Soo-Hyung Kim, Sae-Ryung Kang |
Neural Networks | 3 |
| 2025 | ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise DiffusionabstractAudio-driven talking head generation requires precise synchronization between facial animations and audio signals. This paper introduces ATL-Diff, a novel approach addressing synchronization limitations while reducing noise and computational costs. Our framework features three key components: a Landmark Generation Module converting audio to facial landmarks, a Landmarks-Guide Noise approach that decouples audio by distributing noise according to landmarks, and a 3D Identity Diffusion network preserving identity characteristics. Experiments on MEAD and CREMA-D datasets demonstrate that ATL-Diff outperforms state-of-the-art methods across all metrics. Our approach achieves near real-time processing with high-quality animations, computational efficiency, and exceptional preservation of facial nuances. This advancement offers promising applications for virtual assistants, education, medical communication, and digital platforms. The source code is available at: https://github.com/sonvth/ATL-Diff Thanh Hoang Son Vo, Quang-Vinh Nguyen, Seungwon Kim, Soonja Yeom, Soo-Hyung Kim |
AVSS | 6 |
| 2025 | A Time-Aware Mental State Space for Multimodal Depression Detection on Social Media
Dong Thanh Nguyen, Duc Duy Nguyen, Doan Khai Ta, Hai Binh Nguyen, Ji-Eun Shin, Seungwon Kim, Soo-Hyung Kim |
CogSci | 9 |
| 2025 | Anatomical Attention Alignment Representation for Radiology Report GenerationabstractAutomated Radiology report generation (RRG) aims at producing detailed descriptions of medical images, reducing radiologists’ workload and improving access to high-quality diagnostic services. Existing encoder-decoder models only rely on visual features extracted from raw input images, which can limit the understanding of spatial structures and semantic relationships, often resulting in suboptimal text generation. To address this, we propose Anatomical Attention Alignment Network (A3Net), a framework that enhance visual-textual understanding by constructing hyper-visual representations. Our approach integrates a knowledge dictionary of anatomical structures with patch-level visual features, enabling the model to effectively associate image regions with their corresponding anatomical entities. This structured representation improves semantic reasoning, interpretability, and cross-modal alignment, ultimately enhancing the accuracy and clinical relevance of generated reports. Experimental results on IU X-Ray and MIMIC-CXR datasets demonstrate that A3Net significantly improves both visual perception and text generation quality. Our code is available at GitHub. Thanh Hoang Son Vo, Soo-Hyung Kim |
ICIP | 5 |
| 2025 | Three Techniques for Enhancing Emotional Expression on Embodied Avatar Face in VRabstractPeople often attempt to mask their true emotions through deliberate facial expressions, but such efforts are not always successful. In contrast, emotional concealment could be more easily achieved in Virtual Reality (VR) when appropriate functionalities are available. This study introduces three techniques in VR that enable users to manually adjust emotional facial expression while still reflecting real-time facial tracking results. The Ekman (Ek) technique allows users to select six discrete emotions via button interaction, while the Scrollable-Ekman (SEk) technique extends this by allowing users to scale the intensity of the selected emotion. The Arousal-Valence (AV) technique offers nuanced control within a two-dimensional arousal-valence space. We evaluated these techniques against a baseline condition that synchronizes users' natural facial expressions, focusing on the expression of happiness, sadness, and anger. In most measurements, the Ek and SEk showed better results compared to the baseline and the AV techniques. Notably, the SEk technique was particularly effective in enhancing hedonic quality. During the free-flowing conversation, the most critical factor was the timely and well-synchronized coordination between speech and controlled facial expressions. Participant satisfaction also varied by usage style: those who tried to use the techniques continuously and naturally to mimic real-life communication reported lower satisfaction, while those who used them occasionally for playful or exaggerated expressions tended to report higher satisfaction. Jaejoon Jeong, Gun A. Lee, Soo-Hyung Kim, Ji-Eun Shin, Gayun Suh, Sei Kang, Seungwon Kim |
ISMAR | 4 |
| 2025 | Design and Evaluation of a Virtual Agent for Interpersonal Emotion Regulation in VRabstractManaging negative emotions through emotion regulation (ER) is key to mental well-being. While virtual reality (VR) shows promise for supporting ER, prior work has primarily focused on selfregulation. This paper introduces a virtual agent that helps users manage emotions through conversation-based ER strategies. We compared three conditions: no agent, an agent with non-supportive responses, and an agent with ER-supportive responses. Results showed that the ER-supportive agent significantly improved users' emotional states and overall experience. Building on this, we conducted a second experiment to examine how the agent's appearance (realistic vs. cartoon) and voice tone (emotional vs. neutral) affect ER. Results indicated that an emotional voice tone improved users' ability to regulate emotions. Although a realistic appearance did not directly improve ER, it increased users' trust and sense of social presence. This paper contributes to VR and human-agent interaction by demonstrating the potential of virtual agents to support ER and offering design implications for future ER-supportive agents. Sei Kang, Gun A. Lee, Soo-Hyung Kim, Ji-Eun Shin, Jaejoon Jeong, Myungho Lee, Seungwon Kim |
ISMAR | 4 |
| 2025 | Effects of Co-speech Gesture Size of Virtual Agents on Persuasive CommunicationabstractCo-speech gestures are crucial for enriching both human-human and human-agent communications. Yet, the specific impacts of gesture size—especially when being generated by advanced data-driven techniques—remain underexplored. This study investigates how varying gesture sizes affect human-agent interactions across two distinct persuasive contexts (informational and emotional), with a focus on social outcomes such as persuasion and empathy. We conducted two controlled experiments, each involving 36 participants, comparing three gesture conditions: Minimal gesture, Small gesture, and Large gesture conditions. Experiment 1, set in an informational sales context, showed that small and large gestures significantly enhanced persuasive effectiveness, social presence, and communication quality compared to the minimal gesture condition, although no meaningful differences emerged between small and large gestures. In contrast, Experiment 2, situated in an emotionally charged context, revealed that larger gestures progressively amplified both persuasive impact and perceived empathy. These findings highlight that gesture size matters in emotionally intensive communications and the substantial social benefits of deep-learning techniques for gesture generation. Gayun Suh, Gun A. Lee, Soo-Hyung Kim, Ji-Eun Shin, Jaejoon Jeong, Sei Kang, Seungwon Kim |
VRST | 4 |
| 2025 | Multilevel spatial-temporal feature analysis for generic event boundary detection in videos
Van Thong Huynh, Seungwon Kim, Soo-Hyung Kim |
Comput. Vis. Image Underst. | 4 |
| 2024 | Polyp-SES: Automatic Polyp Segmentation with Self-enriched Semantic Model
Quang Vinh Nguyen 0003, Thanh Hoang Son Vo, Sae-Ryung Kang, Soo-Hyung Kim |
ACCV (2) | 4 |
| 2024 | The RayHand Navigation: A Virtual Navigation Method with Relative Position between Hand and Gaze-RayabstractIn this paper, we introduce a novel Virtual Reality (VR) navigation method using gaze ray and hand, named RayHand navigation. It supports controlling navigation speed and direction by quickly indicating the initial direction using gaze and then using dexterous hand movement for controlling the speed and direction based on the relative position between the gaze ray and user's hand. We conducted a user study comparing our approach to the head-hand and torso-leaning-based navigation methods, and also evaluated their learning effect. The results showed that the RayHand and head-hand navigations were less physically demanding than the torso-leaning navigation, and the RayHand supported rich navigation experience with high hedonic quality and solved the issue of the user unintentionally stepping out from the designated interaction area. In addition, our approach showed a significant improvement over time with a learning effect. Sei Kang, Jaejoon Jeong, Gun A. Lee, Soo-Hyung Kim, Seungwon Kim |
CHI | 4 |
| 2024 | Multiple Facial Reaction Generation Using Gaussian Mixture of Models and Multimodal Bottleneck TransformerabstractFacial reaction generation has gained prominence in recent years. However, while there has been extensive research on synthesizing facial expressions from the perspective of the speaker, the generation of reactions from the listener's standpoint remains relatively unexplored. Predicting the facial reactions of the listener in a conversational setting presents a challenge due to the diverse range of reactions that can be elicited by the behavior of a single speaker. In this study, we introduce a Multimodal Transformer-based Variational Autoencoder designed to learn the distribution of listener facial reactions based on speaker audiovisual cues. Our proposed approach incorporates the Multimodal Bottleneck Token mechanism to capture interactions between acoustic and visual speaker features and utilizes the Variational Au-to encoder framework to generate latent representations of multiple listener reactions. Additionally, we employ Gaussian Mixture Models to enhance the generative capabilities of the Autoencoder. Experimental results demonstrate that our method surpasses baseline models and previous approaches on the REACT24 benchmark dataset. Dang-Khanh Nguyen, Prabesh Paudel, Seung-Won Kim, Ji-Eun Shin, Soo-Hyung Kim |
FG | 5 |
| 2024 | Vector Quantized Diffusion Models for Multiple Appropriate Reactions GenerationabstractIn the realm of dyadic interactions, the ability to generate appropriate facial reactions is paramount for the conveyance of empathy and understanding. This paper introduces a novel framework that leverages the strengths of a diffusion model architecture, underpinned by a vector quantized variational autoencoder (VQ-VAE) to synthesize facial reactions that are contextually apt. We rigorously evaluate our model on the IEEE FG REACT2024 dataset, where it demonstrates superior performance, outshining baseline methods in terms of effectiveness. The results underscore the potential of our framework to enhance the fidelity of digital human interactions, paving the way for more nuanced and emotionally intelligent systems. Ngoc-Huynh Ho, Soo-Hyung Kim, Seungwon Kim, Ji-Eun Shin |
FG | 4 |
| 2024 | Latent Behavior Diffusion for Sequential Reaction Generation in Dyadic Setting
Soo-Hyung Kim, Ji-Eun Shin, Seung-Won Kim |
ICPR (25) | 3 |
| 2024 | 4T-Net: Multitask deep learning for nuclear analysis from pathology images
Vi Thi-Tuong Vo, Myung-Giun Noh, Soo-Hyung Kim |
Multim. Tools Appl. | 3 |
| 2023 | Adaptation of Distinct Semantics for Uncertain Areas in Polyp Segmentation
Van Thong Huynh, Soo-Hyung Kim |
BMVC | 3 |
| 2023 | 3D-DDA: 3D Dual-Domain Attention for Brain Tumor SegmentationabstractAccurate brain tumor segmentation plays an essential role in the diagnosis process. However, there are challenges due to the variety of tumors in low contrast, morphology, location, annotation bias, and imbalance among tumor regions. This work proposes a novel 3D dual-domain attention module to learn local and global information in spatial and context domains from encoding feature maps in Unet. Our attention module generates refined feature maps from the enlarged reception field at every stage by attention mechanisms and residual learning to focus on complex tumor regions. Our experiments on BraTS 2018 have demonstrated superior performance compared to existing state-of-the-art methods. Nhu-Tai Do, Hoang-Son Vo-Thanh, Tram-Tran Nguyen-Quynh, Soo-Hyung Kim |
ICIP | 4 |
| 2023 | DCTM: Dilated Convolutional Transformer Model for Multimodal Engagement Estimation in ConversationabstractConversational engagement estimation is posed as a regression problem, entailing the identification of the favorable attention and involvement of the participants in the conversation. This task arises as a crucial pursuit to gain insights into human's interaction dynamics and behavior patterns within a conversation. In this research, we introduce a dilated convolutional Transformer for modeling and estimating human engagement in the MULTIMEDIATE 2023 competition. Our proposed system surpasses the baseline models, exhibiting a noteworthy 7% improvement on test set and 4% on validation set. Moreover, we employ different modality fusion mechanism and show that for this type of data, a simple concatenated method with self-attention fusion gains the best performance. Vu Ngoc Tu, Van Thong Huynh, Soo-Hyung Kim, Shah Nawaz, Karthik Nandakumar, Muhammad Zaigham Zaheer |
ACM Multimedia | 4 |
| 2023 | Prediction of evoked expression from videos with temporal position fusion
Van Thong Huynh, Soo-Hyung Kim |
Pattern Recognit. Lett. | 4 |
| 2022 | Survival Analysis based on Lung Tumor Segmentation using Global Context-aware Transformer in MultimodalityabstractLung cancer is the most common type of cancer and is on the rise. Recently, most young people were diagnosed with the disease, accounting for approximately 12% of all cancer patients worldwide. Early diagnosis combined with immediate treatment can increase the chance of successful recovery. We observed that tumor information is essential for diagnosis. However, it is expensive and time-consuming to provide tumor location in practice. We therefore first proposed a tumor segmentation model, Multi-scale Aggregation-based Parallel Transformer Network (MAPTransNet), to segment where the tumor cells are and where the normal tissues are in 3D PET/CT images. In this segmentation model, we employed parallel transformer mechanism to capture global context of multi-scale encoder feature maps and then concatenated them to obtain global context at multi-scale maps. Here, we integrate external attention into original vision transformer (ViT) mechanism to learn position properties of each small patches in the entire dataset. After that, we use output of MAPTransNet model as the region of interest (RoI) image for survival analysis task. Since there is a statistically significant relationship between intensity and size of tumors and disease stage that reflects to individual’s survival time, we proposed Multimodality-based Survival Network (MSNet) for predicting hazard rate of non- small cell lung cancer (NSCLC) patients using PET/CT scans with tumor information (RoI image) and clinical data (disease stages, age, histology, etc). The experimental results prove that our proposal achieves competitive performance compared to competing methods in terms of Dice score metric for tumor segmentation task. For survival analysis task, we emphasize that our proposed method outperforms other competing methods in terms of C-index metric. From the results, we conclude that the use of multimodality (clinical and image features) provides rich information related to survival analysis task in NSCLC. Duy-Phuong Dao, Ngoc-Huynh Ho, Sudarshan Pant, Soo-Hyung Kim, In-Jae Oh, Sae-Ryung Kang |
ICPR | 5 |
| 2022 | Positional Multi-Cross-Attention for Bone Age Estimation Using Deep Multiple Instance LearningabstractThe bone age estimation is clinically important as it can be used by physicians to read and report medial images such as bone scintigraphy considering age-related bone changes. Some recent research has applied deep learning techniques to assess bone age and have achieved positive results. Whole-body bone scintigraphy usually contains millions of pixels, which is computationally infeasible. This problem can be solved effectively with multiple instance learning by treating a whole image as a bag of instances. However, many approaches usually assume that there are no dependencies or ordering among instances. In the present study, we propose a multiple instance learning based method for bone age prediction using whole-body bone scan images. Our network architecture combines attention-based multiple instances learning with a multi-cross attention module to handle large input images and dependencies among instances. The experiments conducted on the Chonnam National University Hwasun Hospital data set have shown the potential performance of our model. The proposed framework can be used as a robust support tool for clinicians to analyze and prognosticate bone aging and age-related diseases. Thanh Cong Do, Sae-Ryung Kang, Soo-Hyung Kim, Jung-Joon Min |
ICPR | 4 |
| 2022 | Facial Expression Recognition Using a Temporal Ensemble of Multi-Level Convolutional Neural NetworksabstractEmotion recognition is indispensable in human-machine interaction systems. It comprises locating facial regions of interest in images and classifying them into one of seven classes: angry, disgust, fear, happy, neutral, sad, and surprise. Despite several breakthroughs in image classification, particularly in facial expression recognition, this research area is still challenging, as sampling in the wild is a demanding task. In this study, a two-stage method is proposed for recognizing facial expressions given a sequence of images. At the first stage, all face regions are extracted in each frame, and essential information that would be helpful and related to human emotion is obtained. Then, the extracted features from the previous step are considered temporal data and are assigned to one of the seven basic emotions. In addition, a study of multi-level features is conducted in a convolutional neural network for facial expression recognition. Moreover, various network connections are introduced to improve the classification task. By combining the proposed network connections, superior results are obtained compared to state-of-the-art methods on the FER2013 dataset. Furthermore, the performance of our temporal model is better than that of the single architecture of the 2017 EmotiW challenge winner on the AFEW 7.0 dataset. Hai Duong Nguyen, Sun-Hee Kim, In Seop Na, Soo-Hyung Kim |
IEEE Trans. Affect. Comput. | 6 |
| 2021 | Survival time prediction by integrating cox proportional hazards network and distribution function networkabstractBACKGROUND: The Cox proportional hazards model is commonly used to predict hazard ratio, which is the risk or probability of occurrence of an event of interest. However, the Cox proportional hazard model cannot directly generate an individual survival time. To do this, the survival analysis in the Cox model converts the hazard ratio to survival times through distributions such as the exponential, Weibull, Gompertz or log-normal distributions. In other words, to generate the survival time, the Cox model has to select a specific distribution over time. RESULTS: This study presents a method to predict the survival time by integrating hazard network and a distribution function network. The Cox proportional hazards network is adapted in DeepSurv for the prediction of the hazard ratio and a distribution function network applied to generate the survival time. To evaluate the performance of the proposed method, a new evaluation metric that calculates the intersection over union between the predicted curve and ground truth was proposed. To further understand significant prognostic factors, we use the 1D gradient-weighted class activation mapping method to highlight the network activations as a heat map visualization over an input data. The performance of the proposed method was experimentally verified and the results compared to other existing methods. CONCLUSIONS: Our results confirmed that the combination of the two networks, Cox proportional hazards network and distribution function network, can effectively generate accurate survival time. Eu-Tteum Baek, Soo-Hyung Kim, In-Jae Oh, Sae-Ryung Kang, Jung-Joon Min |
BMC Bioinform. | 3 |
| 2021 | Korean video dataset for emotion recognition in the wildabstractAbstract Emotion recognition is one of the hottest fields in affective computing research. Recognizing emotions is an important task for facilitating communication between machines and humans. However, it is a very challenging task based on a lack of ethnically diverse databases. In particular, emotional expressions tend to be very dissimilar between Western and Eastern people. Therefore, diverse emotion databases are required for studying emotional expression. However, majority of the well-known emotion databases focus on Western people, which exhibit different characteristics compared to Eastern people. In this study, we constructed a novel emotion dataset containing more than 1200 video clips collected from Korean movies, called Korean Video Dataset for Emotion Recognition in the Wild (KVDERW). Which are similar to real-world conditions, with the goal of studying the emotions of Eastern people, particularly Korean people. Additionally, we developed a semi-automatic video emotion labelling tool that could be used to generate video clips and annotate the emotions in clips. Trinh Le Ba Khanh, Soo-Hyung Kim, Eu-Tteum Baek |
Multim. Tools Appl. | 2 |
| 2021 | Real-time virtual mouse system using RGB-D images and fingertip detectionabstractAbstract A real-time fingertip-gesture-based interface is still challenging for human–computer interactions, due to sensor noise, changing light levels, and the complexity of tracking a fingertip across a variety of subjects. Using fingertip tracking as a virtual mouse is a popular method of interacting with computers without a mouse device. In this work, we propose a novel virtual-mouse method using RGB-D images and fingertip detection. The hand region of interest and the center of the palm are first extracted using in-depth skeleton-joint information images from a Microsoft Kinect Sensor version 2, and then converted into a binary image. Then, the contours of the hands are extracted and described by a border-tracing algorithm. The K-cosine algorithm is used to detect the fingertip location, based on the hand-contour coordinates. Finally, the fingertip location is mapped to RGB images to control the mouse cursor based on a virtual screen. The system tracks fingertips in real-time at 30 FPS on a desktop computer using a single CPU and Kinect V2. The experimental results showed a high accuracy level; the system can work well in real-world environments with a single CPU. This fingertip-gesture-based interface allows humans to easily interact with computers by hand. Dinh-Son Tran, Ngoc-Huynh Ho, Soo-Hyung Kim |
Multim. Tools Appl. | 4 |
| 2021 | Deep neural network-based fusion model for emotion recognition using visual dataabstractAbstract In this study, we present a fusion model for emotion recognition based on visual data. The proposed model uses video information as its input and generates emotion labels for each video sample. Based on the video data, we first choose the most significant face regions with the use of a face detection and selection step. Subsequently, we employ three CNN-based architectures to extract the high-level features of the face image sequence. Furthermore, we adjusted one additional module for each CNN-based architecture to capture the sequential information of the entire video dataset. The combination of the three CNN-based models in a late-fusion-based approach yields a competitive result when compared to the baseline approach while using two public datasets: AFEW 2016 and SAVEE. Luu Ngoc Do, Hai Duong Nguyen, Soo-Hyung Kim, In Seop Na |
J. Supercomput. | 4 |
| 2020 | Affective Expression Analysis in-the-wild using Multi-Task Temporal Statistical Deep Learning ModelabstractAffective behavior analysis plays an important role in human-computer interaction, customer marketing, health monitoring. ABAW Challenge and Aff-Wild2 dataset raise the new challenge for classifying basic emotions and regression valence-arousal value under in-the-wild environments. In this paper, we present an affective expression analysis model that deals with the above challenges. Our approach includes STAT and Temporal Module for fine-tuning again face feature model. We experimented on Aff-Wild2 dataset, a large-scale dataset for ABAW Challenge with the annotations for both the categorical and valence-arousal emotion. We achieved the expression score 0.543 and valence-arousal score 0.534 on the validation set. Nhu-Tai Do, Tram-Tran Nguyen-Quynh, Soo-Hyung Kim |
FG | 3 |
| 2020 | Multimodality Pain and related Behaviors Recognition based on Attention LearningabstractOur work aimed to study facial data as well as movement data for recognition of pain and related behaviors in the context of everyday physical activities, which was provided as three tasks in EmoPain 2020 challenge. We explored deep visual representation and geometric features, which included head pose, facial landmarks, and action units in facial data with a combination of fully connected layers for estimating pain from facial data. In tasks with movement data, we employed long short-term memory layers to learn temporal information in each segment of 180 frames. We examined attention mechanism to investigate the relationship and gather data from multiple sources together. Experiments on EmoPain dataset showed that our methods significantly outperformed baseline results on pain recognition tasks. Van Thong Huynh, Soo-Hyung Kim |
FG | 4 |
| 2020 | Representing Temporal Attributes for Schema MatchingabstractTemporal data are prevalent, where one or several time attributes present. It is challenging to identify the temporal attributes from heterogeneous sources. The reason is that the same attribute could contain distinct values in different time spans, whereas different attributes may have highly similar timestamps and alike values. Existing studies on schema matching seldom explore the temporal information for matching attributes. In this paper, we argue to order the values in an attribute A by some time attribute T as a time series. To learn deep temporal features in the attribute pair (T, A), we devise an auto-encoder to embed the transitions of values in the time series into a vector. The temporal attribute matching (TAM) is thus to evaluate matching distance of two temporal attribute pairs by comparing their transition vectors. We show that computing the optimal matching distance is NP-hard, and present an approximation algorithm. Experiments on real datasets demonstrate the superiority of our proposal in matching temporal attributes compared to the generic schema matching approaches. Yinan Mei, Shaoxu Song, Yunsu Lee, Jungho Park, Soo-Hyung Kim, Sungmin Yi |
KDD | 5 |
| 2020 | Effective and Efficient Retrieval of Structured EntitiesabstractStructured entities are commonly abstracted, such as from XML, RDF or hidden-web databases. Direct retrieval of various structured entities is highly demanded in data lakes, e.g., given a JSON object, to find the XML entities that denote the same real-world object. Existing approaches on evaluating structured entity similarity emphasize too much the structural inconsistency. Indeed, entities from heterogeneous sources could have very distinct structures, owing to various information representation conventions. We argue that the retrieval could be more tolerant to structural differences and focus more on the contents of the entities. In this paper, we first identify the unique challenge of parent-child (containment) relationships among structured entities, which unfortunately prevent the retrieval of proper entities (returning parents or children). To solve the problem, a novel hierarchy smooth function is proposed to combine the term scores in different nodes of a structured entity. Entities sharing the same structure, namely an entity family, are employed to learn the coefficient in aggregating the scores, and thus distinguish/prune the parent or child entities. Remarkably, the proposed method could cooperate with both the bag-of-words (BOW) and word embedding models, successful in retrieving unstructured documents, for querying structured entities. Extensive experiments on real datasets demonstrate that our proposal is effective and efficient. Ruihong Huang, Shaoxu Song, Yunsu Lee, Jungho Park, Soo-Hyung Kim, Sungmin Yi |
Proc. VLDB Endow. | 5 |
| 2019 | Continuous Hand Gesture Spotting and Classification Using 3D Finger Joints InformationabstractThis paper presents a novel approach for continuous dynamic hand gesture recognition for RGB video input. Our approach contains two main modules. Firstly, in the gesture spotting module, the video sequence with continuous gestures are pre-segmented into isolated gestures. Secondly, the gesture classification module classifies the segmented gestures. In the gesture spotting module, the motion of the hand palm and finger movements are fed into Bidirectional Long Short-Term Memory (Bi-LSTM) network for gesture spotting purpose. In the gesture classification module, two 3D Convolution Neural Networks (3D_CNN) and one Long Short-Term Memory (LSTM) network are combined to efficiently utilize the combination of multiple data channels such as RGB, Optical Flow, and 3D key pose positions. The promising performance of our approach is obtained by experiments conducted on three publish datasets - Chalearn LAP ConGD dataset, 20BN-Jester, and NVIDIA Dynamic Hand gesture Dataset. Our approach achieves mean Jaccard Index of 0.5535, which outperforms the state-of-the-art methods on Chalearn LAP ConGD dataset. Nguyen Ngoc Hoang, Soo-Hyung Kim |
ICIP | 3 |
| 2019 | A 3d Face Modeling Approach for in-The-Wild Facial Expression Recognition on Image DatasetsabstractThis paper explores the benefits of 3D face modeling for in-the-wild facial expression recognition (FER). Since there is limited in-the-wild 3D FER dataset, we first construct 3D facial data from available 2D dataset using recent advances in 3D face reconstruction. The 3D facial geometry representation is then extracted by deep learning technique. In addition, we also take advantage of manipulating the 3D face, such as using 2D projected images of 3D face as additional input for FER. These features are then fused with that of 2D FER typical network. By doing so, despite using common approaches, we achieve a competent recognition accuracy on Real-World Affective Faces (RAF) database and Static Facial Expressions in the Wild (SFEW 2.0) compared with the state-of-the-art reports. To the best of our knowledge, this is the first time such a deep learning combination of 3D and 2D facial modalities is presented in the context of in-the-wild FER. Son Thai Ly, Nhu-Tai Do, Soo-Hyung Kim |
ICIP | 4 |
| 2019 | Group-level Cohesion Prediction using Deep Learning Models with A Multi-stream Hybrid NetworkabstractIn this paper, we propose a hybrid deep learning network for predicting group cohesion in images. It is a kind of regression problem and its objective is to predict the Group Cohesion Score (GCS), which is in the range of [0,3]. In order to solve this issue, we exploit four types of visual cues, such as scene, skeleton, UV coordinates and face image, along with state-of-the-art convolutional neural networks (CNNs). We use not only fusion but also ensemble methods to combine these approaches. Our proposed hybrid network achieves 0.517 and 0.416 mean square errors (MSEs) on validation and testing sets, respectively. We finally achieved the first place on the Group-level Cohesion Sub-challenge (GC) in the EmotiW 2019. Dang Xuan Tien, Soo-Hyung Kim, Thanh Hung Vo |
ICMI | 2 |
| 2019 | Engagement Intensity Prediction withFacial Behavior FeaturesabstractThis paper describes an approach for the engagement prediction task, a sub-challenge of the 7th Emotion Recognition in the Wild Challenge (EmotiW 2019). Our method involves three fundamental steps: feature extraction, regression and model ensemble. In the first step, an input video is divided into multiple overlapped segments (instances) and the features extracted for each instance. The combinations of Long short-term memory (LSTM) and Fully connected layers deployed to capture the temporal information and regress the engagement intensity for the features in previous step. In the last step, we performed fusions to achieve better performance. Finally, our approach achieved a mean square error of 0.0597, which is 4.63% lower than the best results last year. Van Thong Huynh, Soo-Hyung Kim |
ICMI | 2 |
| 2019 | Facial Emotion Recognition Using an Ensemble of Multi-Level Convolutional Neural NetworksabstractEmotion recognition plays an indispensable role in human–machine interaction system. The process includes finding interesting facial regions in images and classifying them into one of seven classes: angry, disgust, fear, happy, neutral, sad, and surprise. Although many breakthroughs have been made in image classification, especially in facial expression recognition, this research area is still challenging in terms of wild sampling environment. In this paper, we used multi-level features in a convolutional neural network for facial expression recognition. Based on our observations, we introduced various network connections to improve the classification task. By combining the proposed network connections, our method achieved competitive results compared to state-of-the-art methods on the FER2013 dataset. Hai Duong Nguyen, Soonja Yeom, In Seop Na, Soo-Hyung Kim |
Int. J. Pattern Recognit. Artif. Intell. | 6 |
| 2019 | A novel 2D and 3D multimodal approach for in-the-wild facial expression recognition
Son Thai Ly, Nhu-Tai Do, Soo-Hyung Kim |
Image Vis. Comput. | 3 |
| 2019 | Multiple human tracking in drone image
Hai Duong Nguyen, In Seop Na, Soo-Hyung Kim, Jun Ho Choi |
Multim. Tools Appl. | 3 |
| 2018 | Text line segmentation using a fully convolutional network in handwritten document imagesabstractLine detection in handwritten documents is an important problem for processing of scanned documents. While existing approaches mainly use hand‐designed features or heuristic rules to estimate the location of text lines, the authors present a novel approach that trains a fully convolutional network (FCN) to predict text line structure in document images. A rough estimation of text line, or a line map, is obtained by using FCN, from which text strings that pass through characters in each text line are constructed. Finally, the touching characters should be separated and assigned to different text lines to complete the segmentation, for which line adjacency graph is used. Experimental results on ICDAR2013 Handwritten Segmentation Contest data set show high performance together with the robustness of the system with different types of languages and multi‐skewed text lines. Quang Nhat Vo, Soo-Hyung Kim |
IET Image Process. | 2 |
| 2018 | Recognition of Music Scores with Non-Linear Distortions in Mobile Devices
Quang Nhat Vo, Soo-Hyung Kim |
Multim. Tools Appl. | 3 |
| 2018 | Binarization of degraded document images based on hierarchical deep supervised network
Quang Nhat Vo, Soo-Hyung Kim |
Pattern Recognit. | 2 |
| 2017 | Quality of Private Information (QoPI) model for effective representation and prediction of privacy controls in mobile computing
In-Young Ko, Soo-Hyung Kim |
Comput. Secur. | 3 |
| 2017 | A robust system for document layout analysis using multilevel homogeneity structure
Tuan Anh Tran 0001, Kang Han Oh, In Seop Na, Soo-Hyung Kim |
Expert Syst. Appl. | 6 |
| 2017 | Salient object detection using recursive regional feature clustering
Kang Han Oh, Myungeun Lee, Yura Lee, Soo-Hyung Kim |
Inf. Sci. | 4 |
| 2017 | Music symbol recognition by a LAG-based combination model
In Seop Na, Soo-Hyung Kim |
Multim. Tools Appl. | 2 |
| 2017 | Multi-query processing of XML data streams on multicore
Soo-Hyung Kim, Kyong-Ha Lee, Yoon-Joon Lee |
J. Supercomput. | 1 |
| 2016 | An entropy-based analytic model for the privacy-preserving in open dataabstractIn a Big Data era, a lot of open data set is published and shared with the public. That creates new services and business. However, the publication may cause a leakage problem of private information. In general, de-identification techniques are applied to the data before publication. The problem, however, has not been solved completely. Personal data can be obtained from the several sources such as Internet service and social media. In this situation, a de-identified open data may be simply joined with the leaked external data and it may result in a re-identification issue. We propose a new analytic model to measure the personal information leakage risk in the open data before publishing. The proposed model formulates the entropy-based re-identification risk to measure the privacy leakage risk. We also try to find the data utility measure by using the entropy while preserving the privacy. Based on both the risk and the utility measure, we propose the guideline for data open to the public. We show the guideline including the risk and utility measurement can be applicable with the empirical experiments. Soo-Hyung Kim, Changwook Jung, Yoon-Joon Lee |
IEEE BigData | 1 |
| 2016 | Recovery of drawing order from multi-stroke English handwritten images based on graph models and ambiguous zone analysis
Minh Dinh, Soo-Hyung Kim, Luu Ngoc Do |
Expert Syst. Appl. | 4 |
| 2016 | Page segmentation using minimum homogeneity algorithm and adaptive mathematical morphology
Tuan Anh Tran 0001, In Seop Na, Soo-Hyung Kim |
Int. J. Document Anal. Recognit. | 3 |
| 2016 | Detection of multiple salient objects through the integration of estimated foreground clues
Kang Han Oh, Myungeun Lee, Gwang Bok Kim, Soo-Hyung Kim |
Image Vis. Comput. | 4 |
| 2016 | A mixture model using Random Rotation Bounding Box to detect table region in document image
Tuan Anh Tran 0001, Hong Tai Tran, In Seop Na, Soo-Hyung Kim |
J. Vis. Commun. Image Represent. | 6 |
| 2016 | An MRF model for binarization of music scores with complex background
Vo Quang Nhat, Soo-Hyung Kim |
Pattern Recognit. Lett. | 2 |
| 2015 | Automatic extraction of text regions from document images by multilevel thresholding and k-means clusteringabstractTextual data plays an important role in a number of applications such as image database indexing, document understanding, and image-based web searching. The target of automatic real-life text extracting in document images without character recognition module is to identify image regions that contain only text. These textual regions can then be either input of optical character recognition application or highlighted for user focusing. In this paper we propose a method which consists of three stages-preprocessing which improves contrast of grayscale image, multi-level thresholding for separating textual region from non-textual object such as graphics, pictures, and complex background, and heuristic filter, recursive filter for text localizing in textual region. In many of these applications, it is not necessary to identify all the text regions, therefor we emphasize on identifying important text region with relatively large size and high contrast. Experimental results on real-life dataset images demonstrate that the proposed method is effective in identifying textual region with various illuminations, size and font from various types of background. Tuan Anh Tran 0001, In Seop Na, Soo-Hyung Kim |
ICIS | 4 |
| 2015 | Extraction of salient objects based on image clustering and saliency
In Seop Na, Ha Le, Soo-Hyung Kim |
Pattern Anal. Appl. | 3 |
| 2014 | A novel and effective method for specular detection and removal by tensor votingabstractMost specular detection methods assumed that dominant highlight regions should be uniform for the detection of highlights, which may not be the case in real images. Even when non-uniformity is allowed in the detection, the specular removal can still suffer from non-converged artifacts due to discontinuities in surface colors, especially in highly textured and multicolor images. In this paper, we propose a novel and effective resolution to separate and remove specular components from a single image by adopting tensor voting to obtain reflectance distribution of an input image. Specular and noise pixels denoted as small tensors are isolated and removed. Diffuse reflectance distribution is achieved by analyzing salient and orientation information of tensors around the specular region. The proposed method is non-iterative and does not require any predefined constraints in the input image. We evaluate our proposed method on a dataset consisting of highly textured and multicolor images. Experimental results showed that our result is outstanding compared to other state-of-the-art techniques. Vo Quang Nhat, Soo-Hyung Kim |
ICIP | 3 |
| 2014 | Staff Line Removal Using Line Adjacency Graph and Staff Line Skeleton for Camera-Based Printed Music ScoresabstractOn camera-based music scores, curved and uneven staff-lines tend to incur more frequently, and with the loss in performance of binarization methods, line thickness variation and space variation between lines are inevitable. We propose a novel and effective staff-line removal method based on following 3 main ideas. First, the state-of-the-art staff-line detection method, Stable Path, is used to extract staff-line skeletons of the music score. Second, a line adjacency graph (LAG) model is exploited in a different manner of over segmentation to cluster pixel runs generated from the run-length encoding (RLE) of the image. Third, a two-pass staff-line removal pipeline called filament filtering is applied to remove clusters lying on the staff-line. Our method shows impressive results on music score images captured from cameras, and gives high performance when applied to the ICDAR/GREC 2013 database. Hoang-Nam Bui, In Seop Na, Soo-Hyung Kim |
ICPR | 3 |
| 2014 | Boosted Stable Path for Staff-Line Detection Using Order Statistic Downscaling and Coarse-to-Fine TechniqueabstractStaff-line detection is the key component in any Optical Music Recognition (OMR) system. The state-of-the-art Stable Path method has the powerful capability on skewed and distorted music sheets. However, the naive cost function calculation and graph-traversing for shortest paths is time consuming. In this paper we present a novel method to overcome this challenge. A coarse-to-fine technique is applied for accelerating the speed of staff-line detection. First, coarse-level staff-line detection is performed on a 2D order-statistic-based scaled binary image to estimate staff-line positions. Second, we estimate staff-line boundaries by interpolation and translation of coarse detection results. Finally, fine-level staff-line detection is applied for refining the result from the first step. Experiments show that our Boosted Stable Path technique can impressively speed-up the naive method, hence a user-friendly mobile OMR application is possible. Hoang-Nam Bui, In Seop Na, Soo-Hyung Kim |
ICPR | 5 |
| 2014 | Distorted Music Score Recognition without Staffline RemovalabstractThis paper proposes a new approach for recognizing the primitive musical symbols in distorted music scores without the staff line removal. We try to overcome two main issues. The first problem is the difficult and unreliable removal of staff lines required as a pre-processing step for most of recognition systems. The second problem is the non-linear distortion of the music score images captured by digital cameras. At the beginning, we detect the locations of bar-lines on each staff and segment it into sub-areas which can be rectified into undistorted shapes by biquadratic transformation. Then, musical rules, template matching, run length coding and projection methods are employed to extract the musical note information without the application of staff removal. The proposed method is implemented on smart phones and shows promising results. Vo Quang Nhat, Soo-Hyung Kim |
ICPR | 3 |
| 2013 | Retrieval of Effective Leaf Area Index in Heterogeneous Forests With Terrestrial Laser ScanningabstractTerrestrial laser scanner (TLS)-based leaf area index (LAI) retrieval is an appealing concept, due to the ability to capture structural information of canopies as 3-D point cloud data (PCD). TLS-based LAI estimation methods promise a nondestructive tool for spatially explicit calibration of LAI estimated by aerial or satellite remote sensing techniques. These methods also overcome the sky condition restrictions of on-ground optical instruments such as hemispherical photography frequently used for LAI estimation. This paper presents a new method for estimating the effective LAI (LAIe) directly from PCD generated by TLS in heterogeneous forests. We converted the 3-D PCD into 2-D raster images, similar to hemispherical photographs, using two geometrical projection techniques in order to estimate gap fraction and LAIe using a linear least squares method. Our results indicated that the TLS-based algorithm was able to capture the variability in LAIe of forest stands with a range of densities. The TLS-based LAIe estimation method explained 89.1% (rmse= 0.01 ;p<; 0.001) of the variation in results from digital hemispherical photographs taken of the same stands and used for validation. The Breusch-Pagan test score confirmed that the stereographic-projection-based TLS LAIe model was more robust compared to the Lambert azimuthal equal-area projection TLS LAIe model. Finally, we explore and show significant relationships between airborne-laser-scanner (ALS)-based and TLS-based LAIe estimates, showing promise for further exploration of utilizing TLS as a calibration tool for ALS. Guang Zheng, L. Monika Moskal, Soo-Hyung Kim |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2012 | Fast automatic saliency map driven geometric active contour model for color object segmentation
Nguyen Tran Lan Anh, Vo Quang Nhat, Elyor Kodirov, Soo-Hyung Kim |
ICPR | 4 |
| 2011 | Segmentation of the Pectoral Muscle Boundary in Breast MR ImagesabstractRecently, breast MR images have been used more frequently for diagnosis of that area. Several methods have been proposed for segmenting these MR images, however, these methods are not dedicated to segmenting the pectoral muscle. Therefore, in this paper, we propose a practical and general-purpose approach for detecting segmenting in the pectoral muscle boundary based on the structure tensor. The segmentation work flow comprises four key steps: preprocessing, detection of the region of interest (ROI) within the breast region, segmenting the pectoral muscle and finally extracting and refining the pectoral muscle boundary. From experimental results we show that the proposed method can segment the pectoral muscle exactly. In addition, the proposed method will allow the application of the CAD and registration for various breast images. Myungeun Lee, Yanjuan Chen, Soo-Hyung Kim, Jong Hyo Kim |
BIBE | 3 |
| 2011 | Automatically improving image quality using tensor voting
Toan Nguyen Dinh, Soo-Hyung Kim, Hyuk Ro Park |
Neural Comput. Appl. | 3 |
| 2010 | Segmentation of 3D object in volume dataset using active deformable modelabstractThe level set approach can be used as powerful tool for volume segmentation of a region-of-interest (ROI), to achieve an accurate estimation of tumor or soft tissue in medical images. A major challenge of such algorithms is required to set the equation parameters, especially in the speed function. In this paper, we introduce a geometric active surface scheme that uses level set approach for tumor segmentation in volume datasets by the surface evolution framework based on the geometric variation principle. In this scheme, the level set speed function is designed using hybrid information of geodesic active region and geodesic active contour. Our method handles topological changes of the deformable surface using geometric integral measures and the level set theory. These integral measures contain the robust alignment term, the active region term and the minimal surface term. The proposed algorithm is tested on medical images of the head for tumor segmentation and its performance is evaluated visually and quantitatively. The experimental results confirm the effectiveness of the proposed method and its superior performance when compared with traditional approaches. Wan Hyun Cho, Soon-Young Park, Sun-Worl Kim, Soo-Hyung Kim, Gukdong Ahn, Myung-Eun Lee |
ICIP | 5 |
| 2010 | Level-Set Segmentation of Brain Tumors Using a New Hybrid Speed FunctionabstractThis paper presents a new hybrid speed function needed to perform image segmentation within the level-set framework. This speed function provides a general form that incorporates the alignment term as a part of the driving force for the proper edge direction of an active contour by using the probability term derived from the region partition scheme and, for regularization, the geodesics contour term. First, we use an external force for active contours as the Gradient Vector Flow field. This is computed as the diffusion of gradient vectors of a gray level edge map derived from an image. Second, we partition the image domain by progressively fitting statistical models to the intensity of each region. Here we adopt two Gaussian distributions to model the intensity distribution of the inside and outside of the evolving curve partitioning the image domain. Third, we use the active contour model that has the computation of geodesics or minimal distance curves, which allows stable boundary detection when the model's gradients suffer from large variations including gaps or noise. Finally, we test the accuracy and robustness of the proposed method for various medical images. Experimental results show that our method can properly segment low contrast, complex images. Wan Hyun Cho, Soon-Young Park, Soo-Hyung Kim, Sun-Worl Kim, Gukdong Ahn, Myung-Eun Lee |
ICPR | 4 |
| 2010 | Automatic detection and recognition of Korean text in outdoor signboard images
Euichul Kim, Junsik Lim, Soo-Hyung Kim, Myunghun Lee, Seongtaek Hwang |
Pattern Recognit. Lett. | 5 |
| 2009 | Segmentation of Brain MR Images Using an Ant Colony Optimization AlgorithmabstractIn this paper, we describe a segmentation method for brain MR images using an ant colony optimization (ACO) algorithm. This is a relatively new meta-heuristic algorithm and a successful paradigm of all the algorithms which take advantage of the insectpsilas behavior. It has been applied to solve many optimization problems with good discretion, parallel, robustness and positive feedback. As an advanced optimization algorithm, only recently, researchers began to apply ACO to image processing tasks. Hence, we segment the MR brain image using ant colony optimization algorithm. Compared to traditional meta-heuristic segmentation methods, the proposed method has advantages that it can effectively segment the fine details. Myung-Eun Lee, Soo-Hyung Kim, Wan Hyun Cho, Soon-Young Park, Junsik Lim |
BIBE | 2 |
| 2009 | A Generic Framework of Integrating Segmentation and RegistrationabstractWe propose an integrated framework that can simultaneously segment and register a set of medical images using a geometric active model and a pseudo-likelihood method. First, we segment given medical volume data using a fast matching and geometric deformable model and extract a surface of an object from the segmented volume data. Second, we use the hidden Markov random field model and the pseudo-likelihood method to statistically model the intensity distribution of each voxel at the surface region. We adopt the Bernoulli probability model to formulate a prior distribution of the labeling variable for the transformed voxels. The Gaussian mixture model is taken as a probability distribution function for the intensity of the transformed voxel. We use the deterministic annealing EM (DAEM) algorithm to get the proper estimators for the parameters of the complete-data log likelihood function. Then, we define a new registration measure with the maximization function, called Q-function, obtained by the DAEM algorithm. We evaluate the precision of the proposed approach by comparing the registration traces of our measure with other measures such as Mutual Information or Cross Correlation for the original image and its transformed image with respect to translation and rotation. The experimental results show that our method has great potential power to segment and register various medical images given by different modalities. Wan Hyun Cho, Soon-Young Park, Junsik Lim, Soo-Hyung Kim |
BIBE | 5 |
| 2009 | Automatic Image Restoration Based on Tensor Voting
Toan Nguyen Dinh, Soo-Hyung Kim, Hyuk Ro Park |
ICONIP (1) | 3 |
| 2008 | Segmentation of medical image based on mean shift and deterministic annealing EM algorithmabstractIn this paper, we use the mean shift procedure to determine the number of components in a mixture model and to detect their modes of each mixture component. Next, we have adopted the Gaussian mixture model to represent the probability distribution of feature vectors. A deterministic annealing expectation maximization algorithm is used to estimate the parameters of the GMM. The experimental results show that the mean shift part of the proposed algorithm is efficient to determine the number of components and modes of each component in mixture models. And it shows that the DAEM part provides a global optimal solution for the parameter estimation in a mixture model. Myung-Eun Lee, Soo-Hyung Kim, Wan Hyun Cho |
AICCSA | 2 |
| 2008 | Improved image segmentation method based on optimized threshold using Genetic AlgorithmabstractIn image segmentation, threshold segmentation is becoming more and more widely used because of its simplicity and efficiency. In this paper, an improved image segmentation method based on optimized threshold using genetic algorithm is proposed. Compared with the traditional threshold segmentation methods, this method has advantages that it can nicely segment the thin and it can efficiently reduce calculation time and it has good capability and stabilization nature. The results show that using this proposed method can obtain satisfactory segmentation effect. Myung-Eun Lee, Soo-Hyung Kim |
AICCSA | 3 |
| 2008 | Multimodality image registration using ordinary procrustes analysis and entropy of bivariate normal kernel densityabstractWe present a registration method for medical images based on shape information and voxel intensities. First, we segment volume images using the Markov random field and the Gibbs distribution. We extract the 3D feature points of the shape from the surface of the segmented object. Then, we conduct first registration using ordinary Procrustes analysis for two sets of 3D feature points. For the second registration, we define the new optimization measure of registration as the entropy of the bivariate normal kernel density for pairs of intensities given from the extracted feature points as well as the transformed feature points. The final registration for two volume images is carried out by finding the appropriate transformation parameter yielding the minimum value of this optimization measure. To evaluate the performance of the proposed registration method, we conduct various experiments comparing our method with existing ones such as the Mutual Information measure. Wan Hyun Cho, Sun-Worl Kim, Myung-Eun Lee, Soo-Hyung Kim, Soon-Young Park, Chang Bu Jeong |
BIBE | 4 |
| 2008 | Text image matching without language model using a Hausdorff distance
Hwa Jeong Son, Soo-Hyung Kim, Ji Soo Kim |
Inf. Process. Manag. | 2 |
| 2007 | Multi-modal Data Integration Using Graph for Collaborative Assembly Design Information Sharing and Reuse
Hyung-Jae Lee, Soo-Hyung Kim, Sook-Young Choi |
IEA/AIE | 4 |
| 2006 | Delivery and Storage Architecture for Sensed Information Using SNMP
Deokjai Choi 0001, Hongseok Jang, Kugsang Jeong, Punghyeok Kim, Soo-Hyung Kim |
APNOMS | 5 |
| 2006 | A Middleware Architecture Determining Application Context Using Shared Ontology
Kugsang Jeong, Deokjai Choi 0001, Soo-Hyung Kim |
ICCSA (4) | 3 |
| 2005 | Text Locating from Natural Scene Images Using Image IntensitieabstractIn this paper, we propose three text extraction methods based on intensity information for natural scene images. The first method is composed of gray value stretching and binarization by an average intensity of the image. This method is appropriate to extract texts from complex backgrounds. The second method is a split and merge approach which is one of well-known algorithms for image segmentation. The third one is a combination of the two. Experimental results show that the proposed approaches are superior to conventional methods both in simple and complex images. Ji Soo Kim, Sang-Cheol Park, Soo-Hyung Kim |
ICDAR | 3 |
| 2005 | Keyword Spotting on Hangul Document Images Using Two-Level Image-to-Image Matching
Sang-Cheol Park, Hwa Jeong Son, Chang Bu Jeong, Soo-Hyung Kim |
IEA/AIE | 4 |
| 2004 | Fast Half Pixel Motion Estimation Based on Spatio-temporal Correlations
Hyo Sun Yoon, Soo-Hyung Kim, Deokjai Choi 0001 |
ICONIP | 3 |
| 2004 | Word-Level Optical Font Recognition Using Typographical FeaturesabstractPrevious research efforts on optical font recognition have mostly limited applications since they deal with only a few types of font attributes and estimate them from a line or block of text. This paper proposes a word-level optical font recognition system for printed Korean and English documents. At the word-level, it has the advantages of obtaining more detailed font attributes including the following: script (Korean and English), font style (regular, bold, italic, and underlined), typeface (Myung-jo and Gothic), point size (10, 12, 14 pts), and word length (2, 3, 4, 5 for Korean, and 4 to 10 for English). A hierarchical classifier and several typographical features have been devised for the system, and their effectiveness are proven by an experiment with a database of 100 sets of 264 font categories. Soo-Hyung Kim, Hee K. Kwag, Ching Y. Suen |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2003 | Analysis and Recognition of Asian Scripts - the State of the ArtabstractThis paper summarizes the research activities of the pastdecade on the recognition of handwritten scripts used inChina, Japan, and Korea. It presents the recognitionmethodologies, features explored, databases used, andclassification schemes investigated. In addition, it includes adescription of the performance of numerous recognitionsystems found in both academic and industrial researchlaboratories. Recent achievements and applications are alsopresented. A list of relevant references is attached togetherwith our remarks on this subject. Ching Y. Suen, Shunji Mori, Soo-Hyung Kim, Cheung Hoi Leung |
ICDAR | 3 |
| 2001 | Word Segmentation in Handwritten Korean Text Lines Based on Gap Clustering TechniquesabstractWe propose a word segmentation method for handwritten Korean text lines. It uses gap information to separate a text line into word units, where the gap is defined as a white-run obtained after a vertical projection of the line image. Each gap is classified into a between-word gap or a within-word gap using a clustering technique. We take up three gap metrics - the bounding box (BB), run-length/Euclidean (RLE) and convex hull (CH) distances - which are known to have superior performance in Roman-style word segmentation, and three clustering techniques - the average linkage method, the modified MAX method and sequential clustering. An experiment with 498 text-line images extracted from live mail pieces has shown that the best performance is obtained by the sequential clustering technique using all three gap metrics. Soo-Hyung Kim, Ching Y. Suen, S. Jeong |
ICDAR | 1 |
| 2001 | A lexicon-driven approach for optimal segment combination in off-line recognition of unconstrained handwritten Korean words
Soo-Hyung Kim, S. Jeong, Ching Y. Suen |
Pattern Recognit. | 1 |
| 2000 | Mean Field Annealing EM for Image SegmentationabstractWe present a statistical model-based approach to the color image segmentation. A novel deterministic annealing expectation-maximization (EM) and mean field theory are used to estimate the posterior probability of each pixel and the parameters of the Gaussian mixture model which represents the multi-colored objects statistically. Image segmentation is carried out by clustering each pixel into the most probable component Gaussian. The experimental results show that the mean field annealing EM provides a global optimal solution for the maximum likelihood parameter estimation and the real images are segmented efficiently using the estimates computed by the maximum entropy principle and mean field theory. Wan Hyun Cho, Soo-Hyung Kim, Soon-Young Park |
ICIP | 2 |
| 1999 | A Lexicon Driven Approach for Off-line Recognition of Unconstrained Handwritten Korean WordsabstractWe propose a new method for the recognition of unconstrained handwritten words consisting of Korean and numeric characters. To overcome the difficulty of separating touching characters, we adopt an over-segmentation technique and we find the optimal segment combination using a lexicon-driven word scoring technique and a nearest neighbor classifier. The optimal combination gives the final segmentation positions for individual characters with the best matching word in the lexicon. The proposed system has yielded an accuracy of 90.64% for 908 word images on live mail pieces. Soo-Hyung Kim, S. Jeong, Ching Y. Suen |
ICDAR | 1 |
| 1995 | Off-line Recognition of Korean Scripts Using Distance Matching and Neural Network ClassifiersabstractAn off-line recognition engine is proposed for handwritten Korean characters based on a distance matching and the neural network technique. The distance matching selects a set of several candidates from the large set of character classes, and the neural network performs a detailed classification on the candidates. As an approach for combining the two methodologies, a clustering method based on sample distributions has been devised. Recognition accuracy of the engine on a public database, PE92, is 84.1%. About four character patterns can be processed in a second on PC. Soo-Hyung Kim, Jeong-In Doh |
ICDAR | 1 |
| 1994 | Automatic Input of Logic Diagrams by Recognizing Loop-Symbols and Rectilinear ConnectionsabstractAn automatic system that reads an image of a logic diagram, digitized by a document scanner, and generates a description of the diagram in terms of logic symbols and their interconnections is proposed. The digitized image is converted into a set of line segments by way of a sequence of picture processing operations, Then symbols and connections are recognized by identifying loops and rectilinear polylines. Graphical description of a set of symbol models is provided as prior knowledge about input diagrams. Experiments show that the system can recognize more than 96% of logic symbols and connections and that an A4 size diagram (210×297 mm ) of average complexity can be processed within 10 seconds on a workstation. Soo-Hyung Kim, Jin Hyung Kim |
Int. J. Pattern Recognit. Artif. Intell. | 1 |