VLDB 2026 Research / reviewers in the wild / expert
Seong Jong Ha
dblp:42/3078
· DBLP profile ↗
12ranked-venue papers
5as first author
5since 2021 · last 2024
0009-0006-7231-5206ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | GLAD: Global-Local View Alignment and Background Debiasing for Unsupervised Video Domain Adaptation with Large Domain GapabstractIn this work, we tackle the challenging problem of unsupervised video domain adaptation (UVDA) for action recognition. We specifically focus on scenarios with a substantial domain gap, in contrast to existing works primarily deal with small domain gaps between labeled source domains and unlabeled target domains. To establish a more realistic setting, we introduce a novel UVDA scenario, denoted as Kinetics→BABEL, with a more considerable domain gap in terms of both temporal dynamics and background shifts. To tackle the temporal shift, i.e., action duration difference between the source and target domains, we propose a global-local view alignment approach. To mitigate the background shift, we propose to learn temporal order sensitive representations by temporal order learning and background invariant representations by background augmentation. We empirically validate that the proposed method shows significant improvement over the existing methods on the Kinetics→BABEL dataset with a large domain gap. The code is available at https://github.com/KHU-VLL/GLAD. Hyogun Lee, Kyungho Bae, Seong Jong Ha, Yumin Ko, Gyeong-Moon Park, Jinwoo Choi 0001 |
WACV | 3 |
| 2022 | Future Transformer for Long-term Action AnticipationabstractThe task of predicting future actions from a video is crucial for a real-world agent interacting with others. When anticipating actions in the distant future, we humans typically consider long-term relations over the whole sequence of actions, i.e., not only observed actions in the past but also potential actions in the future. In a similar spirit, we propose an end-to-end attention model for action anticipation, dubbed Future Transformer (FUTR), that leverages global attention over all input frames and output tokens to predict a minutes-long sequence of future actions. Unlike the previous autoregressive models, the proposed method learns to predict the whole sequence of future actions in parallel decoding, enabling more accurate and fast inference for long-term anticipation. We evaluate our method on two standard benchmarks for long-term action anticipation, Breakfast and 50 Salads, achieving state-of-the-art results. Dayoung Gong, Joonseok Lee, Manjin Kim, Seong Jong Ha, Minsu Cho |
CVPR | 4 |
| 2021 | Video Question Answering Using Language-Guided Deep Compressed-Domain Video FeatureabstractVideo Question Answering (Video QA) aims to give an answer to the question through semantic reasoning between visual and linguistic information. Recently, handling large amounts of multi-modal video and language information of a video is considered important in the industry. However, the current video QA models use deep features, suffered from significant computational complexity and insufficient representation capability both in training and testing. Existing features are extracted using pre-trained networks after all the frames are decoded, which is not always suitable for video QA tasks. In this paper, we develop a novel deep neural network to provide video QA features obtained from coded video bit-stream to reduce the complexity. The proposed network includes several dedicated deep modules to both the video QA and the video compression system, which is the first attempt at the video QA task. The proposed network is predominantly model-agnostic. It is integrated into the state-of-the-art networks for improved performance without any computationally expensive motion-related deep models. The experimental results demonstrate that the proposed network outperforms the previous studies at lower complexity. https://github.com/Nayoung-Kim-ICP/VQAC Seong Jong Ha, Je-Won Kang |
ICCV | 2 |
| 2021 | Zero-shot Natural Language Video LocalizationabstractUnderstanding videos to localize moments with natural language often requires large expensive annotated video regions paired with language queries. To eliminate the annotation costs, we make a first attempt to train a natural language video localization model in zero-shot manner. Inspired by unsupervised image captioning setup, we merely require random text corpora, unlabeled video collections, and an off-the-shelf object detector to train a model. With the unpaired data, we propose to generate pseudo-supervision of candidate temporal regions and corresponding query sentences, and develop a simple NLVL model to train with the pseudo-supervision. Our empirical validations show that the proposed pseudo-supervised method outperforms several baseline approaches and a number of methods using stronger supervision on Charades-STA and ActivityNet-Captions. Jinwoo Nam, Daechul Ahn, Dongyeop Kang, Seong Jong Ha |
ICCV | 4 |
| 2021 | Semantic-Preserving Metric Learning for Video-Text RetrievalabstractVideo-text retrieval requires finding an optimal space for comparing the similarity of two different modalities. Most approaches adopt ranking loss as a primary training objective to find the space. The loss is only interested in bringing the samples annotated as pairs closer to each other without considering the semantic relevance of different samples. This rather causes even semantically similar pairs not to get close. To deal with the problem, we propose semantic-preserving metric learning. The proposed method entails the metric space where the similarity ratio between samples is proportional to semantic relevance between annotations. In the extensive experiments on video-text datasets, the proposed method presents a close alignment between the learned metric space and the semantic space. It also demonstrates state-of-the-art retrieval performance. Sungkwon Choo, Seong Jong Ha, Joonsoo Lee |
ICIP | 2 |
| 2012 | Reduction of ghost effect in exposure fusion by detecting the ghost pixels in saturated and non-saturated regionsabstractThis paper proposes a multiple exposure image fusion algorithm with reduced ghost. The basic idea is to adjust the weight map in the conventional fusion method in such a way that the ghost pixels are excluded. For this, pixels that cause ghost effect are detected in both of the saturated and non-saturated regions. In order to detect ghost pixels in the non-saturated region, we use the photometric relation and Gaussian mixture modeling (GMM) of a zero mean normalized cross correlation (ZNCC) map between a given exposure image and the reference. From this, we can also obtain static region where we construct an intensity mapping function (IMF) to detect the ghost pixels in saturated regions. Experimental results show that the proposed method generates high quality image without noticeable ghost effect, and yields less artifacts than the conventional methods. Jaehyun An, Seong Jong Ha, Nam Ik Cho |
ICASSP | 2 |
| 2012 | Fast text line extraction in document imagesabstractThis paper proposes an algorithm for fast text line extraction in document image. Instead of binarization or multi-oriented Gaussian blurring of an image as in the conventional methods, we use integral image and design filters that are proper to detect text regions on the integral image. After the filtering, the center points in the regions are discovered by cascade text region verification followed by non-maximum suppression. Finally, text lines are extracted by grouping the points on the same line. The proposed method is tested with document images taken in various environments, and it is shown to be faster than the conventional ones while its performance is comparable. Seong Jong Ha, Bora Jin, Nam Ik Cho |
ICIP | 1 |
| 2012 | Image registration by using a descriptor for repetitive patternsabstractThis paper proposes a new feature-based image registration method based on the description of feature clusters. This method can find larger number of correspondences than the conventional methods using singleton feature descriptors, which often fail in repetitive patterns. The reason for the failure of conventional methods in a repeating pattern is due to the existence of too many similar features, which in turn gives geometrically inconsistent matching or do not survive ratio test. Hence the proposed method follows the strategy that first separate the similar features from the repetitive patterns from the others. Then the similar features in a pattern are grouped into a set that is described by a support vector descriptor in terms of the cluster's center and radius. Once the same pattern in different images are matched, the geometric cue is added to find many geometrically consistent correspondences of the features. In the experiments, it has been demonstrated that the larger number of geometrically consistent correspondences from the repetitive pattern give more accurate registration, and thus more pleasing results in image stitching and panoramic image generation. Seong Jong Ha, Seyun Kim, Nam Ik Cho |
VCIP | 1 |
| 2012 | Discrimination and description of repetitive patterns for enhancing the performance of feature-based recognition
Seong Jong Ha, Sang Hwa Lee, Nam Ik Cho |
Image Vis. Comput. | 1 |
| 2011 | Discrimination and description of repetitive patterns for enhancing object recognition performanceabstractObjects with repetitive patterns are not well recognized by SIFT/SURF based matching because the features from those patterns are too similar and thus it is often difficult to find the homography of matched pairs. In this paper, we propose a new feature matching strategy to alleviate this problem by differentiating repetitive patterns from the other salient ones and also by developing a way of utilizing the patterns for robust feature matching. Specifically, we develop a classifier that tells whether the features are from repetitive patterns or salient features, based on mean shift clustering followed by support vector data description. Then the homography is found over the salient features by excluding the repetitive features at first, which is then validated and refined by the patterns. The proposed method is tested with the examples of matching the buildings with repeating patterns, and it is shown to be robuster and more reliable than the conventional ones. Seong Jong Ha, Sang Hwa Lee, Nam Ik Cho |
ICIP | 1 |
| 2009 | Stereo matching using hierarchical belief propagation along ambiguity gradientabstractThis paper proposes a stereo matching algorithm based on hierarchical belief propagation and occlusion handling. We define a new order for message passing in belief propagation instead of the scanline approach. The primary assumption is that a pixel with a well-defined minimum in its likelihood field is more likely to contain a correct disparity, when compared to a pixel having an ill-defined minimum with several local minima. The order for message passing is determined by the variance of likelihood field at each pixel. The variances evaluate the ambiguity of likelihood fields, and the messages are hierarchically updated along the gradient of ambiguity. The experimental results show that the proposed method estimates the disparities correctly in the hard regions such as large occlusions and textureless regions. The proposed algorithm is currently tied with the best performing algorithm on the Middlebury stereo site. Sumit Srivastava, Seong Jong Ha, Sang Hwa Lee, Nam Ik Cho, Sang Uk Lee |
ICIP | 2 |
| 2008 | Pnoramic mosaic system for mobile devicesabstractThis paper deals with a panoramic mosaic system for mobile devices. The proposed system is optimized for mobile devices by integer-programmable algorithms and auto-shot user interface. The proposed system consists of auto-shot interface, transform onto cylindrical surface, color compensation, local alignment, image stitching, and blending. The auto-shot interface senses user's camera motion and takes pictures when the camera motion is matched to the predefined motion model. The captured images are projected onto a cylindrical mosaic surface by modified projection equation. Exposure difference is removed by color compensation, where the luminance differences between images are adjusted and color saturation is eliminated. The projected images are locally aligned by fast hierarchical hexagonal search technique. Then, the optimal boundary of overlapped images are determined using dynamic programming and synthesized seamlessly. According to the experiments using mobile devices, the proposed system shows good performance compared with other PC-based mosaic algorithms. Seong Jong Ha, Sang Hwa Lee, Yu Ri Ahn, Nam Ik Cho |
ICIP | 1 |