Ha Young Kim

dblp:191/2588 · DBLP profile ↗
← Back
27ranked-venue papers
3as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Towards Temporal Fusion Beyond the Field of View for Camera-based Semantic Scene Completion
abstract
Recent camera-based 3D semantic scene completion (SSC) methods have increasingly explored leveraging temporal cues to enrich the features of the current frame. However, while these approaches primarily focus on enhancing in-frame regions, they often struggle to reconstruct critical out-of-frame areas near the sides of the ego-vehicle, although previous frames commonly contain valuable contextual information about these unseen regions. To address this limitation, we propose the Current-Centric Contextual 3D Fusion (C3DFusion) module, which generates hidden region-aware 3D feature geometry by explicitly aligning 3D-lifted point features from both current and historical frames. C3DFusion performs enhanced temporal fusion through two complementary techniques—historical context blurring and current-centric feature densification—which suppress noise from inaccurately warped historical point features by attenuating their scale, and enhance current point features by increasing their volumetric contribution. Simply integrated into standard SSC architectures, C3DFusion demonstrates strong effectiveness, significantly outperforming state-of-the-art methods on the SemanticKITTI and SSCBench-KITTI-360 datasets. Furthermore, it exhibits robust generalization, achieving notable performance gains when applied to other baseline models.
Jongseong Bae, Junwoo Ha, Jinnyeong Heo, Yeongin Lee, Ha Young Kim
AAAI5
2025 Three Cars Approaching within 100m! Enhancing Distant Geometry by Tri-Axis Voxel Scanning for Camera-based Semantic Scene Completion
abstract
Camera-based Semantic Scene Completion (SSC) is gaining attentions in the 3D perception field. However, properties such as perspective and occlusion lead to the underestimation of the geometry in distant regions, posing a critical issue for safety-focused autonomous driving systems. To tackle this, we propose ScanSSC, a novel camera-based SSC model composed of a Scan Module and Scan Loss, both designed to enhance distant scenes by leveraging context from near-viewpoint scenes. The Scan Module uses axis-wise masked attention, where each axis employing a near-to-far cascade masking that enables distant voxels to capture relationships with preceding voxels. In addition, the Scan Loss computes the cross-entropy along each axis between cumulative logits and corresponding class distributions in a near-to-far direction, thereby propagating rich context-aware signals to distant voxels. Leveraging the synergy between these components, ScanSSC achieves state-of-the-art performance, with IoUs of 44.54 and 48.29, and mIoUs of 17.40 and 20.14 on the SemanticKITTI and SSCBench-KITTI-360 benchmarks.
Jongseong Bae, Junwoo Ha, Ha Young Kim
CVPR3
2025 Leveraging the Power of MLLMs for Gloss-Free Sign Language Translation
abstract
Sign language translation (SLT) is a challenging task that involves translating sign language images into spoken language. For SLT models to perform this task successfully, they must bridge the modality gap and identify subtle variations in sign language components to understand their meanings accurately. To address these challenges, we propose a novel gloss-free SLT framework called Multimodal Sign Language Translation (MMSLT), which leverages the representational capabilities of off-the-shelf multimodal large language models (MLLMs). Specifically, we use MLLMs to generate detailed textual descriptions of sign language components. Then, through our proposed multimodal-language pre-training module, we integrate these description features with sign video features to align them within the spoken sentence space. Our approach achieves state-of-the-art performance on benchmark datasets PHOENIX14T and CSL-Daily, highlighting the potential of MLLMs to be utilized effectively in SLT. Code is available at https://github.com/hwjeon98/MMSLT.
Jungeun Kim, Hyeongwoo Jeon, Jongseong Bae, Ha Young Kim
ICCV4
2025 Multi-View Slot Attention using Paraphrased Texts for Face Anti-Spoofing
abstract
Recent face anti-spoofing (FAS) methods have shown remarkable cross-domain performance by employing vision-language models like CLIP. However, existing CLIP-based FAS models do not fully exploit CLIP's patch embedding tokens, failing to detect critical spoofing clues. Moreover, these models rely on a single text prompt per class (e.g., 'live' or 'fake'), which limits generalization. To address these issues, we propose MVP-FAS, a novel framework incorporating two key modules: Multi-View Slot attention (MVS) and Multi-Text Patch Alignment (MTPA). Both modules utilize multiple paraphrased texts to generate generalized features and reduce dependence on domain-specific text. MVS extracts local detailed spatial features and global context from patch embeddings by leveraging diverse texts with multiple perspectives. MTPA aligns patches with multiple text representations to improve semantic robustness. Extensive experiments demonstrate that MVP-FAS achieves superior generalization performance, outperforming previous state-of-the-art methods on cross-domain datasets. Code: https://github.com/Elune001/MVP-FAS.
Jeongmin Yu, Susang Kim, Kisu Lee, Taekyoung Kwon 0002, Won-Yong Shin, Ha Young Kim
ICCV6
2025 Looping In: Exploring Feedback Strategies to Motivate Human Engagement in Interactive Machine Learning
abstract
This study investigates effective feedback mechanisms to maintain human engagement in interactive machine learning (IML) systems, focusing on social media platforms. We developed “Loop,” an IML system based on human-in-the-loop (HITL) principles that recommends content while encouraging users to report inaccuracies for model refinement. Loop implements three types of artificial intelligence (AI) feedback on user reports: (a) machine learning (ML)-centric, (b) personal-centric, and (c) community-centric feedback. In addition, we evaluated the relative effectiveness of these feedback types under two different task criticality scenarios: high and low. A user study with 30 participants was conducted to evaluate Loop through questionnaires and interviews. Results showed that participants preferred algorithmic improvements for personal benefit over altruistic contributions to the community, especially for low-criticality tasks. Furthermore, personal-centric feedback had a significant impact on user engagement and satisfaction. Our findings provide insights into the effectiveness of machine feedback in HITL-ML systems, contributing to the design of more engaging and effective IML interfaces. We discuss implications and strategies for encouraging proactive user engagement in HITL-ML-based systems, emphasizing the importance of tailored feedback mechanisms.
Hyorim Shin, Jeongeun Park 0003, Jeongmin Yu, Jungeun Kim, Ha Young Kim, Changhoon Oh
Int. J. Hum. Comput. Interact.5
2025 HeLPNet: Facial structure-aware landmark coordinate regression with heatmap-guided local patch embedding
Jeongmin Yu, Jongseong Bae, Ha Young Kim
Neurocomputing3
2025 MVFormer: Diversifying feature normalization and token mixing for efficient vision transformers
Jongseong Bae, Susang Kim, Minsu Cho, Ha Young Kim
Pattern Recognit. Lett.4
2025 DiffSLT: Enhancing diversity in sign language translation via diffusion model
JiHwan Moon, Jungeun Kim, Jongseong Bae, Hyeongwoo Jeon, Ha Young Kim
Pattern Recognit. Lett.6
2024 "Is Text-Based Music Search Enough to Satisfy Your Needs?" A New Way to Discover Music with Images
abstract
Music is intrinsically connected to human experience, yet the plethora of choices often renders the search for the ideal piece perplexing, especially when the search terms are ambiguous. This study questions the viability of employing visual data, specifically images, in innovative queries for music search, and it aims to better align search results with users’ moods and situational context. We designed and evaluated three prototype systems for music search—TTTune (text-based), VisTune (image-based), and VTTune (hybrid)—to comparatively assess user experience and system usability. In a comprehensive user study involving 236 participants, each participant interacted with one of the systems and subsequently completed post-experimental surveys. A subset of participants also participated in in-depth interviews to further elucidate the potential and the advantages of image-based music retrieval (IMR) systems. Our findings reveal a marked preference for the user experience and usability offered by the IMR approach, as compared with the traditional text-based method. This underscores the potential of the image in an effective search query. Based on these findings, we discuss interface design guidelines tailored for IMR systems and factors affecting system performance, contributing to the evolving landscape of music search methods.
Jeongeun Park 0003, Hyorim Shin, Changhoon Oh, Ha Young Kim
CHI4
2024 Do Topological Characteristics Help in Knowledge Distillation?
abstract
Knowledge distillation (KD) aims to transfer knowledge from larger (teacher) to smaller (student) networks. Previous studies focus on point-to-point or pairwise relationships in embedding features as knowledge and struggle to efficiently transfer relationships of complex latent spaces. To tackle this issue, we propose a novel KD method called TopKD, which considers the global topology of the latent spaces. We define global topology knowledge using the persistence diagram (PD) that captures comprehensive geometric structures such as shape of distribution, multiscale structure and connectivity, and the topology distillation loss for teaching this knowledge. To make the PD transferable within reasonable computational time, we employ approximated persistence images of PDs. Through experiments, we support the benefits of using global topology as knowledge and demonstrate the potential of TopKD. Code is available at https://github.com/jekim5418/TopKD
Jungeun Kim, Junwon You, Ha Young Kim, Jae-Hun Jung
ICML4
2024 TSpred: a robust prediction framework for TCR-epitope interactions using paired chain TCR sequence data
abstract
MOTIVATION: Prediction of T-cell receptor (TCR)-epitope interactions is important for many applications in biomedical research, such as cancer immunotherapy and vaccine design. The prediction of TCR-epitope interactions remains challenging especially for novel epitopes, due to the scarcity of available data. RESULTS: We propose TSpred, a new deep learning approach for the pan-specific prediction of TCR binding specificity based on paired chain TCR data. We develop a robust model that generalizes well to unseen epitopes by combining the predictive power of CNN and the attention mechanism. In particular, we design a reciprocal attention mechanism which focuses on extracting the patterns underlying TCR-epitope interactions. Upon a comprehensive evaluation of our model, we find that TSpred achieves state-of-the-art performances in both seen and unseen epitope specificity prediction tasks. Also, compared to other predictors, TSpred is more robust to bias related to peptide imbalance in the dataset. In addition, the reciprocal attention component of our model allows for model interpretability by capturing structurally important binding regions. Results indicate that TSpred is a robust and reliable method for the task of TCR-epitope binding prediction. AVAILABILITY AND IMPLEMENTATION: Source code is available at https://github.com/ha01994/TSpred.
Ha Young Kim, Sungsik Kim, Woong-Yang Park, Dongsup Kim
Bioinform.1
2024 Image Is All for Music Retrieval: Interactive Music Retrieval System Using Images with Mood and Theme Attributes
abstract
We propose an intuitive image-to-music retrieval (IMR) framework to improve the user experience on these platforms. The proposed method extracts mood and theme tags by searching for images from a pre-built database that are similar to a query image and then retrieves music with matching tag information. We investigated the system’s effectiveness by comparing participants’ satisfaction, intention to use, and valence between those who interacted with the system and those who did not. We also examined whether using mood or theme attributes affected the user-perceived suitability of the retrieved music. Results showed that all three variables of the interaction group were significantly higher than that of the non-interaction group and that there was no difference in the perceived suitability of music between the mood and theme attributes. Our study concludes that image attributes are effective in successful music retrieval and that interaction is a crucial factor in designing IMR systems.
Jeongeun Park 0003, Minchae Kim, Ha Young Kim
Int. J. Hum. Comput. Interact.3
2024 Human, Do You Think This Painting is the Work of a Real Artist?
abstract
Artificial intelligence (AI) is beginning to be applied in the field of art, which had hitherto been an area exclusively reserved for human creativity. Using online AI tools, lay people can easily create artworks that imitate the style of famous artists. Consequently, human judgment on the authenticity of artworks has become critical. While many studies have focused on copyright or value of AI-created artworks, we examine whether human beings can distinguish between paintings drawn by artists and fake paintings created using AI tools. We selected the AI’s recommendations for each artwork and prior information about the artists as factors that can affect human judgment and investigated how the two factors affect people’s discriminative abilities. We found that people have difficulty distinguishing authentic from fake artwork and that additional information about artists and artworks can affect people’s criteria for judging paintings. Furthermore, AI recommendations can help discriminate fake paintings, suggesting that AI-assisted decision-making could play an assistive role in human identification of digitized fake paintings.
Jeongeun Park 0003, Hyunmin Kang, Ha Young Kim
Int. J. Hum. Comput. Interact.3
2024 Integrated Recognition Assistant Framework Based on Deep Learning for Autonomous Driving: Human-Like Restoring Damaged Road Sign Information
abstract
Unpredictable situations frequently occur in real driving environments, and it is often difficult to recognize road signs. In this case, autonomous vehicles (AVs) have a limited ability to predict areas that cannot be detected, making it difficult to judge objects accurately when some information is lost. Therefore, we propose a framework that helps AVs infer proper information under limited conditions. The entire process consists of three steps. First, the missing part of the road sign is restored using the image generative pre-trained transformer model. Next, the sample image with the highest classification accuracy and restored quality is selected among several sample images. Finally, the selected image is provided to users through the designed user interface. The proposed framework improved recognition accuracy compared with unrestored accuracy, indicating the possibility of application as a driving assistance system, and is meaningful in that it is a system that mimics human reasoning ability.
Jeongeun Park 0003, Kisu Lee, Ha Young Kim
Int. J. Hum. Comput. Interact.3
2023 Decompose, Adjust, Compose: Effective Normalization by Playing with Frequency for Domain Generalization
abstract
Domain generalization (DG) is a principal task to evaluate the robustness of computer vision models. Many previous studies have used normalization for DG. In normalization, statistics and normalized features are regarded as style and content, respectively. However, it has a content variation problem when removing style because the boundary between content and style is unclear. This study addresses this problem from the frequency domain perspective, where amplitude and phase are considered as style and content, respectively. First, we verify the quantitative phase variation of normalization through the mathematical derivation of the Fourier transform formula. Then, based on this, we propose a novel normalization method, PC Norm, which eliminates style only as the preserving content through spectral decomposition. Furthermore, we propose advanced PCNorm variants, CCNorm and SCNorm, which adjust the degrees of variations in content and style, respectively. Thus, they can learn domain-agnostic representations for DG. With the normalization methods, we propose ResNet-variant models, DAC-P and DAC-SC, which are robust to the domain gap. The proposed models outperform other recent DG methods. The DAC-SC achieves an average state-of-the-art performance of 65.6% on five datasets: PACS, VLCS, Office-Home, DomainNet, and TerraIncognita.
Sangrok Lee, Jongseong Bae, Ha Young Kim
CVPR3
2023 S3M: Scalable Statistical Shape Modeling Through Unsupervised Correspondences
Lennart Bastian, Alex Baumann, Emily Hoppe, Vincent Bürgin, Ha Young Kim, Mahdi Saleh, Benjamin Busam, Nassir Navab
MICCAI (10)5
2023 CSLT-AK: Convolutional-embedded transformer with an action tokenizer and keypoint emphasizer for sign language translation
Jungeun Kim, Ha Young Kim
Pattern Recognit. Lett.2
2022 ADEL: Adaptive Distribution Effective-Matching Method for Guiding Generators of GANs
Jungeun Kim, Jeongeun Park 0003, Ha Young Kim
ACCV (7)3
2022 Ultra-lightweight face activation for dynamic vision sensor with convolutional filter-level fusion using facial landmarks
Jeongeun Park 0003, Donguk Yang, Dongyup Shin, Jungyeon Kim, Hyunsurk Ryu, Ha Young Kim
Expert Syst. Appl.7
2021 Deep-learning-based automated terminology mapping in OMOP-CDM
abstract
OBJECTIVE: Accessing medical data from multiple institutions is difficult owing to the interinstitutional diversity of vocabularies. Standardization schemes, such as the common data model, have been proposed as solutions to this problem, but such schemes require expensive human supervision. This study aims to construct a trainable system that can automate the process of semantic interinstitutional code mapping. MATERIALS AND METHODS: To automate mapping between source and target codes, we compute the embedding-based semantic similarity between corresponding descriptive sentences. We also implement a systematic approach for preparing training data for similarity computation. Experimental results are compared to traditional word-based mappings. RESULTS: The proposed model is compared against the state-of-the-art automated matching system, which is called Usagi, of the Observational Medical Outcomes Partnership common data model. By incorporating multiple negative training samples per positive sample, our semantic matching method significantly outperforms Usagi. Its matching accuracy is at least 10% greater than that of Usagi, and this trend is consistent across various top-k measurements. DISCUSSION: The proposed deep learning-based mapping approach outperforms previous simple word-level matching algorithms because it can account for contextual and semantic information. Additionally, we demonstrate that the manner in which negative training samples are selected significantly affects the overall performance of the system. CONCLUSION: Incorporating the semantics of code descriptions more significantly increases matching accuracy compared to traditional text co-occurrence-based approaches. The negative training sample collection methodology is also an important component of the proposed trainable system that can be adopted in both present and future related systems.
Byungkon Kang, Jisang Yoon, Ha Young Kim, Sung Jin Jo, Yourim Lee, Hye Jin Kam
J. Am. Medical Informatics Assoc.3
2020 Prediction of mutation effects using a deep temporal convolutional network
abstract
MOTIVATION: Accurate prediction of the effects of genetic variation is a major goal in biological research. Towards this goal, numerous machine learning models have been developed to learn information from evolutionary sequence data. The most effective method so far is a deep generative model based on the variational autoencoder (VAE) that models the distributions using a latent variable. In this study, we propose a deep autoregressive generative model named mutationTCN, which employs dilated causal convolutions and attention mechanism for the modeling of inter-residue correlations in a biological sequence. RESULTS: We show that this model is competitive with the VAE model when tested against a set of 42 high-throughput mutation scan experiments, with the mean improvement in Spearman rank correlation ∼0.023. In particular, our model can more efficiently capture information from multiple sequence alignments with lower effective number of sequences, such as in viral sequence families, compared with the latent variable model. Also, we extend this architecture to a semi-supervised learning framework, which shows high prediction accuracy. We show that our model enables a direct optimization of the data likelihood and allows for a simple and stable training process. AVAILABILITY AND IMPLEMENTATION: Source code is available at https://github.com/ha01994/mutationTCN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ha Young Kim, Dongsup Kim
Bioinform.1
2020 Stock market forecasting with super-high dimensional time-series data using ConvLSTM, trend sampling, and specialized data augmentation
Si Woon Lee, Ha Young Kim
Expert Syst. Appl.2
2020 BshapeNet: Object detection and instance segmentation with bounding shape masks
Ba Rom Kang, Hyunku Lee, Keunju Park, Hyunsurk Ryu, Ha Young Kim
Pattern Recognit. Lett.5
2019 Improving financial trading decisions using deep Q-learning: Predicting the number of shares, action strategies, and transfer learning
Gyeeun Jeong, Ha Young Kim
Expert Syst. Appl.2
2018 ModAugNet: A new forecasting framework for stock market index value with an overfitting prevention LSTM module and a prediction LSTM module
Yujin Baek, Ha Young Kim
Expert Syst. Appl.2
2018 Forecasting the volatility of stock price index: A hybrid model integrating LSTM with multiple GARCH-type models
Ha Young Kim, Chang Hyun Won
Expert Syst. Appl.1
2016 Performance improvement of deep learning based gesture recognition using spatiotemporal demosaicing technique
abstract
We propose a novel method for the demosaicing of event-based images that offers substantial performance improvement of far-distance gesture recognition based on deep Convolutional Neural Network. Unlike the conventional demosaicing technique using the spatial color interpolation of Bayer patterns, our new approach utilizes spatiotemporal correlation between pixel arrays, whereby timestamps of high-resolution pixels are efficiently generated in real-time from the event data. In this paper, we describe this new method and evaluate its performance with a hand motion recognition task.
Paul K. J. Park, Baek Hwan Cho, Jin Man Park, Kyoobin Lee, Ha Young Kim, Hyo Ah Kang, Hyun Goo Lee, Jooyeon Woo, Yohan Roh, Won Jo Lee, Chang-Woo Shin, Qiang Wang 0023, Hyunsurk Ryu
ICIP5