Jaekoo Lee

dblp:08/5973 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-5947-5487ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Training-Free Few-Shot Segmentation via Vision-Language Guided Prompting
abstract
Object segmentation relies heavily on costly pixel-level annotations and struggles to generalize to unseen domains. The recent introduction of the Segment Anything Model (SAM), a foundation model for segmentation, offers a prompt-driven, zero-shot capability that has been applied in various domains (e.g., autonomous driving, satellite imagery, medical imaging) and extended to Few-Shot Segmentation (FSS) tasks. However, existing SAM-based FSS methods typically generate prompts by using a vision encoder to measure support–query image similarity, which often biases towards the support images and fails when there are significant support–query context shifts. To address this limitation, we propose a training-free FSS approach that combines visual and textual cues to generate effective prompts for the target class. By leveraging both vision and language information, our approach bridges the support–query gap and guides SAM to segment novel objects more reliably. Without any additional training, our method outperforms previous state-of-the-art FSS methods on established benchmarks (COCO-20i, Pascal-5i), demonstrating its effectiveness and robust generalization. Our code is publicly available on GitHub.
Euihyun Yoon, Taejin Park, Jaekoo Lee
WACV3
2026 SAMIRO: Spatial Attention Mutual Information Regularization with a pre-trained model as Oracle for lane detection
abstract
Lane detection is an important topic in the future mobility solutions. Real-world environmental challenges such as background clutter, varying illumination, and occlusions pose significant obstacles to effective lane detection, particularly when relying on data-driven approaches that require substantial effort and cost for data collection and annotation. To address these issues, lane detection methods must leverage contextual and global information from surrounding lanes and objects. In this paper, we propose a Spatial Attention Mutual Information Regularization with a pre-trained model as an Oracle , called SAMIRO . SAMIRO enhances lane detection performance by transferring knowledge from a pre-trained model while preserving domain-agnostic spatial information. Leveraging SAMIRO’s plug-and-play characteristic, we integrate it into various state-of-the-art lane detection approaches and conduct extensive experiments on major benchmarks such as CULane, Tusimple, and LLAMAS. The results demonstrate that SAMIRO consistently improves performance across different models and datasets. The code will be made available upon publication. • We propose SAMIRO, a novel regularization framework that improves generalization in lane detection by transferring domain-agnostic knowledge from a pre-trained model. • SAMIRO is plug-and-play and scalable with state-of-the-art lane detection methods. • We validate SAMIRO on public benchmarks (CULane, Tusimple, and LLAMAS), showing that it achieves performance comparable to or surpassing SOTA models.
Jaekoo Lee
Pattern Recognit. Lett.3
2025 GUIDE-CoT: Goal-driven and User-Informed Dynamic Estimation for Pedestrian Trajectory using Chain-of-Thought
Sungsik Kim, Janghyun Baek, Jaekoo Lee
AAMAS4
2025 Leveraging Text-Driven Semantic Variation for Robust OOD Segmentation
abstract
In autonomous driving and robotics, ensuring road safety and reliable decision-making critically depends on out-of-distribution (OOD) segmentation. While numerous methods have been proposed to detect anomalous objects on the road, leveraging the vision-language space–which provides rich linguistic knowledge–remains an underexplored field. We hypothesize that incorporating these linguistic cues can be especially beneficial in the complex contexts found in real-world autonomous driving scenarios.To this end, we present a novel approach that trains a Text-Driven OOD Segmentation model to learn a semantically diverse set of objects in the vision-language space. Concretely, our approach combines a vision-language model’s encoder with a transformer decoder, employs Distance-Based OOD prompts located at varying semantic distances from in-distribution (ID) classes, and utilizes OOD Semantic Augmentation for OOD representaitons. By aligning visual and textual information, our approach effectively generalizes to unseen objects and provides robust OOD segmentation in diverse driving environments.We conduct extensive experiments on publicly available OOD segmentation datasets such as Fishyscapes, Segment-Me-If-You-Can, and Road Anomaly datasets, demonstrating that our approach achieves state-of-the-art performance across both pixel-level and object-level evaluations. This result underscores the potential of vision-language–based OOD segmentation to bolster the safety and reliability of future autonomous driving systems.
Seungheon Song, Jaekoo Lee
IROS2
2025 Deep Sequential Feature Learning for Phase Noise Compensation in Sub-THz Systems
abstract
For data rates exceeding several gigabits per second, which are achievable in fifth-generation wireless networks, the integration of subterahertz (sub-THz) communications has become increasingly important. However, hardware limitations, such as phase noise arising from local oscillator imperfections, can hinder the performance of sub-THz systems. This paper addresses the challenge of phase noise compensation for data decoding in orthogonal frequency division multiplexing systems. To address this issue, we propose a new sequential feature learning mechanism designed to capture both the subcarrier-shift invariance feature of phase noise and the subcarrier-dependent feature of intercarrier interference. The proposed approach involves empirical performance evaluations under various conditions and a theoretical analysis that considers computational complexity and the signal-to-interference-plus-noise power ratio. By performing comprehensive comparisons with existing analytical and deep learning-based methods, we demonstrate the effectiveness of our proposed method, even with an extremely low pilot overhead.
Kyungsik Seo, Jeongwon Jeon, Gilwon Lee, Jaekoo Lee, Song Noh
IEEE Trans. Commun.4
2024 Leveraging Inductive Bias in ViT for Medical Image Diagnosis
Jungmin Ha, Euihyun Yoon, Sungsik Kim, Jaekoo Lee
BMVC5
2024 Controllable 3D Object Generation with Single Image Prompt
Jaeseok Lee, Jaekoo Lee
ICPR (6)2
2024 Learning Temporal Cues by Predicting Objects Move for Multi-camera 3D Object Detection
abstract
In autonomous driving and robotics, there is a growing interest in utilizing short-term historical data to enhance multi-camera 3D object detection, leveraging the continuous and correlated nature of input video streams. Recent work has focused on spatially aligning BEV-based features over timesteps. However, this is often limited as its gain does not scale well with long-term past observations. To address this, we advocate for supervising a model to predict objects’ poses given past observations, thus explicitly guiding to learn objects’ temporal cues. To this end, we propose a model called DAP (Detection After Prediction), consisting of a two-branch network: (i) a branch responsible for forecasting the current objects’ poses given past observations and (ii) another branch that detects objects based on the current and past observations. The features predicting the current objects from branch (i) is fused into branch (ii) to transfer predictive knowledge. We conduct extensive experiments with the large-scale nuScenes datasets, and we observe that utilizing such predictive information significantly improves the overall detection performance. Our model can be used plug-and-play, showing consistent performance gain.
Seokha Moon, Hongbeen Park, Jaekoo Lee, Jinkyu Kim 0001
ICRA3
2023 Distillation for High-Quality Knowledge Extraction via Explainable Oracle Approach
MyungHak Lee, Wooseong Syz Cho, Sungsik Kim, Jaekoo Lee
BMVC5
2022 Bridging the Domain Gap Towards Generalization in Automatic Colorization
Hyejin Lee, Daehee Kim 0003, Daeun Lee 0001, Jinkyu Kim 0001, Jaekoo Lee
ECCV (17)5
2022 Fall Detection using Biometric Information Based on Multi-Horizon Forecasting
abstract
With the steady growth of the aging population, there has been increasing interest in the health of elderly people. In particular, a physical fall can be fatal for the elderly or result in complications that cause more serious physical and mental damage than the cause of the falls. For this reason, it is important to accurately predict the occurrence of abnormal behavior in advance. Although existing methods can be used to analyze general behavioral characteristics, personalized and customized behavioral prediction remains challenging. In this paper, we propose a customized pre-prediction method using sensor wearers’ biometric information as well as behavioral data reflecting the wearers’ unique behavioral characteristics. The technique is based on a temporal fusion transformers method [1] that can be effectively applied to multi-horizon forecasting. A performance evaluation was made to compare the effects of including and excluding biometric information to quantitatively evaluate the enhancement achieved by extending the prediction model. The use of additional biometric information enabled superior performance with 0.9702, 0.9702, 0.9702, 0.9883 in F-1 score, precision, recall respectively, which is all evaluation indicators. The code for the experiments can be found at the following GitHub address: https://github.com/IKKIM00/Fall_Detection_using_multihorizon_forecasting.
Inkyung Kim, Daehee Kim 0003, Sunyoung Kwon, Sheayun Lee, Jaekoo Lee
ICPR5
2021 SelfReg: Self-supervised Contrastive Regularization for Domain Generalization
abstract
In general, an experimental environment for deep learning assumes that the training and the test dataset are sampled from the same distribution. However, in real-world situations, a difference in the distribution between two datasets, i.e. domain shift, may occur, which becomes a major factor impeding the generalization performance of the model. The research field to solve this problem is called domain generalization, and it alleviates the domain shift problem by extracting domain-invariant features explicitly or implicitly. In recent studies, contrastive learning-based domain generalization approaches have been proposed and achieved high performance. These approaches require sampling of the negative data pair. However, the performance of contrastive learning fundamentally depends on quality and quantity of negative data pairs. To address this issue, we propose a new regularization method for domain generalization based on contrastive learning, called self-supervised contrastive regularization (SelfReg). The proposed approach use only positive data pairs, thus it resolves various problems caused by negative pair sampling. Moreover, we propose a class-specific domain perturbation layer (CDPL), which makes it possible to effectively apply mixup augmentation even when only positive data pairs are used. The experimental results show that the techniques incorporated by SelfReg contributed to the performance in a compatible manner. In the recent benchmark, DomainBed, the proposed method shows comparable performance to the conventional state-of-the-art alternatives.
Daehee Kim 0003, Youngjun Yoo, Seunghyun Park 0001, Jinkyu Kim 0001, Jaekoo Lee
ICCV5
2017 Transfer Learning for Deep Learning on Graph-Structured Data
abstract
Graphs provide a powerful means for representing complex interactions between entities. Recently, new deep learning approaches have emerged for representing and modeling graph-structured data while the conventional deep learning methods, such as convolutional neural networks and recurrent neural networks, have mainly focused on the grid-structured inputs of image and audio. Leveraged by representation learning capabilities, deep learning-based techniques can detect structural characteristics of graphs, giving promising results for graph applications. In this paper, we attempt to advance deep learning for graph-structured data by incorporating another component: transfer learning. By transferring the intrinsic geometric information learned in the source domain, our approach can construct a model for a new but related task in the target domain without collecting new data and without training a new model from scratch. We thoroughly tested our approach with large-scale real-world text data and confirmed the effectiveness of the proposed transfer learning framework for deep learning on graphs. According to our experiments, transfer learning is most effective when the source and target domains bear a high level of structural similarity in their graph representations.
Jaekoo Lee, Hyunjae Kim, Jongsun Lee, Sungroh Yoon
AAAI1
2015 Measuring Large-Scale Dynamic Graph Similarity by RICom: RWR with Intergraph Compression
abstract
By how much is a large-scale graph transformed over time or by a significant event?' or 'how structurally similar are two large-scale graphs?' are the two questions that this paper attempts to address. The proposed method efficiently calculates and accurately produces graph similarity. Our approach is based on the well-known random walk with restart (RWR) algorithm, which quantifies relevance between nodes to express the structural and connection characteristics of graphs. Intergraph compression, which is inspired by interframe compression, merges two input graphs and reorders their nodes contributing to improved process-data storage efficiency and processing convenience. This is a boon to the RWR algorithm for large-scale graphs. The representation of a graph transformed via intergraph compression can be used to accurately show similarity because sub-matrix blocks are reordered to concentrate nonzero elements. In performing the RWR algorithm, which quantifies inter-node relevance, transformed representation of graph with Intergraph compression is efficient in space requirement and produces results more quickly and accurately over conventional graph transformation schemes. We demonstrate the validity of our method through experiments and apply it to the usage data of public transportation SmartCard in a large metropolitan area to suggest usefulness of the proposed algorithm.
Jaekoo Lee, Gunn Kim, Sungroh Yoon
ICDM1