VLDB 2026 Research / reviewers in the wild / expert
Sungyeon Kim
dblp:69/8198
· DBLP profile ↗
17ranked-venue papers
10as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-author · 12 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Audio-guided Video Representation with Gated Attention for Video-Text RetrievalabstractVideo-text retrieval, the task of retrieving videos based on a textual query or vice versa, is of paramount importance for video understanding and multimodal information retrieval. Recent methods in this area rely primarily on visual and textual features and often ignore audio, although it helps enhance overall comprehension of video content. Moreover, traditional models that incorporate audio blindly utilize the audio input regardless of whether it is useful or not, resulting in suboptimal video representation. To address these limitations, we propose a novel video-text retrieval framework, Audio-guided VIdeo representation learning with GATEd attention (AVIGATE), that effectively leverages audio cues through a gated attention mechanism that selectively filters out uninformative audio signals. In addition, we propose an adaptive margin-based contrastive loss to deal with the inherently unclear positive-negative relationship between video and text, which facilitates learning better video-text alignment. Our extensive experiments demonstrate that AVIGATE achieves state-of-the-art performance on all the public benchmarks. Boseung Jeong, Jicheol Park, Sungyeon Kim, Suha Kwak |
CVPR | 3 |
| 2025 | GENIUS: A Generative Framework for Universal Multimodal SearchabstractGenerative retrieval is an emerging approach in information retrieval that generates identifiers (IDs) of target data based on a query, providing an efficient alternative to traditional embedding-based retrieval methods. However, existing models are task-specific and fall short of embedding-based retrieval in performance. This paper proposes GENIUS, a universal generative retrieval framework supporting diverse tasks across multiple modalities and domains. At its core, GENIUS introduces modality-decoupled semantic quantization, transforming multimodal data into discrete IDs encoding both modality and semantics. Moreover, to enhance generalization, we propose a query augmentation that interpolates between a query and its target, allowing GENIUS to adapt to varied query forms. Evaluated on the M-BEIR benchmark, it surpasses prior generative methods by a clear margin. Unlike embedding-based retrieval, GENIUS consistently maintains high retrieval speed across database size, with competitive performance across multiple benchmarks. With additional re-ranking, GENIUS often achieves results close to those of embedding-based methods while preserving efficiency. Sungyeon Kim, Xinliang Zhu, Muhammet Bastan, Douglas Gray 0001, Suha Kwak |
CVPR | 1 |
| 2025 | Learning Unified Distance Metric Across Diverse Data Distributions with Parameter-Efficient Transfer LearningabstractA common practice in metric learning is to train and test an embedding model for each dataset. This dataset-specific approach fails to simulate real-world scenarios that involve multiple heterogeneous distributions of data. In this regard, we explore a new metric learning paradigm, called Uni-fied Metric Learning (UML), which learns a unified dis-tance metric capable of capturing relations across multi-ple data distributions. UML presents new challenges, such as imbalanced data distribution and bias towards dom-inant distributions. These issues cause standard metric learning methods to fail in learning a unified metric. To address these challenges, we propose Parameter-efficient Unified Metric leArning (PUMA), which consists of a pre-trained frozen model and two additional modules, stochas-tic adapter and prompt pool. These modules enable to capture dataset-specific knowledge while avoiding bias to-wards dominant distributions. Additionally, we compile a new unified metric learning benchmark with a total of 8 different datasets. PUMA outperforms the state-of-the-art dataset-specific models while using about 69 times fewer trainable parameters. Sungyeon Kim, Donghyun Kim 0006, Suha Kwak |
WACV | 1 |
| 2024 | Efficient and Versatile Robust Fine-Tuning of Zero-Shot Models
Sungyeon Kim, Boseung Jeong, Donghyun Kim 0006, Suha Kwak |
ECCV (17) | 1 |
| 2024 | FREST: Feature RESToration for Semantic Segmentation Under Multiple Adverse Conditions
Sohyun Lee, Namyup Kim, Sungyeon Kim, Suha Kwak |
ECCV (29) | 3 |
| 2023 | HIER: Metric Learning Beyond Class Labels via Hierarchical RegularizationabstractSupervision for metric learning has long been given in the form of equivalence between human-labeled classes. Although this type of supervision has been a basis of metric learning for decades, we argue that it hinders further advances in the field. In this regard, we propose a new regularization method, dubbed HIER, to discover the latent semantic hierarchy of training data, and to deploy the hierarchy to provide richer and more fine-grained supervision than inter-class separability induced by common metric learning losses. HIER achieves this goal with no annotation for the semantic hierarchy but by learning hierarchical proxies in hyperbolic spaces. The hierarchical proxies are learnable parameters, and each of them is trained to serve as an ancestor of a group of data or other proxies to approximate the semantic hierarchy among them. HIER deals with the proxies along with data in hyperbolic space since the geometric properties of the space are well-suited to represent their hierarchical structure. The efficacy of HIER is evaluated on four standard benchmarks, where it consistently improved the performance of conventional methods when integrated with them, and consequently achieved the best records, surpassing even the existing hyperbolic metric learning technique, in almost all settings. Sungyeon Kim, Boseung Jeong, Suha Kwak |
CVPR | 1 |
| 2023 | PromptStyler: Prompt-driven Style Generation for Source-free Domain GeneralizationabstractIn a joint vision-language space, a text feature (e.g., from "a photo of a dog") could effectively represent its relevant image features (e.g., from dog photos). Also, a recent study has demonstrated the cross-modal transferability phenomenon of this joint space. From these observations, we propose PromptStyler which simulates various distribution shifts in the joint space by synthesizing diverse styles via prompts without using any images to deal with source-free domain generalization. The proposed method learns to generate a variety of style features (from "a S∗style of a") via learnable style word vectors for pseudo-words S∗. To ensure that learned styles do not distort content information, we force style-content features (from "a S∗style of a [class]") to be located nearby their corresponding content features (from "[class]") in the joint vision-language space. After learning style word vectors, we train a linear classifier using synthesized style-content features. PromptStyler achieves the state of the art on PACS, VLCS, OfficeHome and DomainNet, even though it does not require any images for training. Junhyeong Cho, Gilhyun Nam, Sungyeon Kim, Hunmin Yang, Suha Kwak |
ICCV | 3 |
| 2022 | Self-Taught Metric Learning without LabelsabstractWe present a novel self-taught framework for unsuper-vised metric learning, which alternates between predicting class-equivalence relations between data through a moving average of an embedding model and learning the model with the predicted relations as pseudo labels. At the heart of our framework lies an algorithm that investigates contexts of data on the embedding space to predict their class-equivalence relations as pseudo labels. The algorithm enables efficient end-to-end training since it demands no off-the-shelf module for pseudo labeling. Also, the class-equivalence relations provide rich supervisory signals for learning an embedding space. On standard benchmarks for metric learning, it clearly outperforms existing unsupervised learning methods and sometimes even beats supervised learning models using the same backbone network. It is also applied to semi-supervised metric learning as a way of exploiting additional unlabeled data, and achieves the state of the art by boosting performance of supervised learning substantially. Sungyeon Kim, Minsu Cho, Suha Kwak |
CVPR | 1 |
| 2022 | Combating Label Distribution Shift for Active Domain Adaptation
Sehyun Hwang, Sohyun Lee, Sungyeon Kim, Jungseul Ok, Suha Kwak |
ECCV (33) | 3 |
| 2022 | Cross-domain Ensemble Distillation for Domain Generalization
Kyungmoon Lee, Sungyeon Kim, Suha Kwak |
ECCV (25) | 2 |
| 2021 | Learning to Generate Novel Classes for Deep Metric Learning
Kyungmoon Lee, Sungyeon Kim, Seunghoon Hong, Suha Kwak |
BMVC | 2 |
| 2021 | Embedding Transfer With Label Relaxation for Improved Metric LearningabstractThis paper presents a novel method for embedding transfer, a task of transferring knowledge of a learned embedding model to another. Our method exploits pairwise similarities between samples in the source embedding space as the knowledge, and transfers them through a loss used for learning target embedding models. To this end, we design a new loss called relaxed contrastive loss, which employs the pairwise similarities as relaxed labels for intersample relations. Our loss provides a rich supervisory signal beyond class equivalence, enables more important pairs to contribute more to training, and imposes no restriction on manifolds of target embedding spaces. Experiments on metric learning benchmarks demonstrate that our method largely improves performance, or reduces sizes and output dimensions of target models effectively. We further show that it can be also used to enhance quality of self-supervised representation and performance of classification models. In all the experiments, our method clearly outperforms existing embedding transfer techniques. Sungyeon Kim, Minsu Cho, Suha Kwak |
CVPR | 1 |
| 2020 | Proxy Anchor Loss for Deep Metric LearningabstractExisting metric learning losses can be categorized into two classes: pair-based and proxy-based losses. The former class can leverage fine-grained semantic relations between data points, but slows convergence in general due to its high training complexity. In contrast, the latter class enables fast and reliable convergence, but cannot consider the rich data-to-data relations. This paper presents a new proxy-based loss that takes advantages of both pair- and proxy-based methods and overcomes their limitations. Thanks to the use of proxies, our loss boosts the speed of convergence and is robust against noisy labels and outliers. At the same time, it allows embedding vectors of data to interact with each other in its gradients to exploit data-to-data relations. Our method is evaluated on four public benchmarks, where a standard network trained with our loss achieves state-of-the-art performance and most quickly converges. Sungyeon Kim, Minsu Cho, Suha Kwak |
CVPR | 1 |
| 2019 | Deep Metric Learning Beyond Binary SupervisionabstractMetric Learning for visual similarity has mostly adopted binary supervision indicating whether a pair of images are of the same class or not. Such a binary indicator covers only a limited subset of image relations, and is not sufficient to represent semantic similarity between images described by continuous and/or structured labels such as object poses, image captions, and scene graphs. Motivated by this, we present a novel method for deep metric learning using continuous labels. First, we propose a new triplet loss that allows distance ratios in the label space to be preserved in the learned metric space. The proposed loss thus enables our model to learn the degree of similarity rather than just the order. Furthermore, we design a triplet mining strategy adapted to metric learning with continuous labels. We address three different image retrieval tasks with continuous labels in terms of human poses, room layouts and image captions, and demonstrate the superior performance of our approach compared to previous methods. Sungyeon Kim, Minkyo Seo, Ivan Laptev, Minsu Cho, Suha Kwak |
CVPR | 1 |
| 2012 | To Cooperate or Not to Cooperate: System Throughput and Fairness PerspectiveabstractThe cooperative transmission, in which some nodes help the transmission of other nodes, has been actively studied to overcome the channel fading effects that deteriorate the communication quality. Thus far, most researches on the cooperative transmission have been studied from the reliability point of view focusing on a single transmission and showing that the cooperative transmission can increase transmission reliability. In this paper, we study the effects of the cooperative transmission from the system throughput and fairness point of view, considering the following fundamental questions: Is the cooperative transmission always helpful to increase the system throughput and improve the degree of fairness among nodes? If not, when is it helpful to increase the system throughput and improve the degree of fairness among nodes? We provide the answers to the above questions with a simple system in which two source nodes are capable of being cooperative with each other by using the decode-and-forward cooperative scheme to transmit their data to a single destination node. Sungyeon Kim, Jang-Won Lee 0001 |
IEEE J. Sel. Areas Commun. | 1 |
| 2011 | Process estimation and optimized recipes of ZnO: Ga thin film characteristics for transparent electrode applications
Chang Eun Kim, Pyung Moon, Ilgu Yun, Sungyeon Kim, Jae-Min Myoung, Hyeon Woo Jang, Jungsik Bang |
Expert Syst. Appl. | 4 |
| 2009 | Joint Resource Allocation for Uplink and Downlink in Wireless Networks: A Case Study with User-Level Utility FunctionsabstractIn most of researches in resource allocation for wireless networks, uplink and downlink problems are considered separately, especially when resources for uplink and downlink are statically partitioned, as in FDD and static TDD systems. However, even in those systems, joint resource allocation for uplink and downlink can improve system efficiency and we study this issue in this paper with the concept of the user-level utility function. In most cases, a user has a two-way communication that consists of two sessions: uplink and downlink sessions and its overall satisfaction to its communication depends on its satisfaction to each of its sessions. To model user's overall satisfaction to its communication, we define a user-level utility function, which is defined as a function of its session-level utility functions. We then formulate and solve the optimization problem with user-level utility functions for cell-level resource scheduling that jointly considers uplink and downlink resource allocation. Simulation results show that our cell-level scheduling in which resource allocation in both uplink and downlink is done jointly outperforms link-level scheduling, in which resource allocation in each of uplink and downlink is done separately in most cases, especially when the asymmetry between uplink and downlink is large. Sungyeon Kim, Jang-Won Lee 0001 |
VTC Spring | 1 |