Jongsuk Kim

dblp:330/3774 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Representation and self-supervised learning · 47% Language models and text generation · 22% Autonomous driving · 13%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Autonomous driving
end-to-end driving
0.912025
SynAD: Enhancing Real-World End-to-End Autonomous Driving Models through Synthetic Data Integration · ICCV 2025
Machine learning › Representation and self-supervised learning › contrastive learning › multimodal contrastive learning
audio-visual contrastive learning
0.812024
EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning · ICML 2024
Natural language and speech › Language models and text generation › prompting › prompt engineering › prompt optimization
automatic prompt optimization
0.812024
StablePrompt : Automatic Prompt Tuning using Reinforcement Learning for Large Language Model · EMNLP 2024
Machine learning › Representation and self-supervised learning › equivariance
equivariant representation learning
0.812024
EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning · ICML 2024
Natural language and speech › Language models and text generation
prompt tuning
0.812024
StablePrompt : Automatic Prompt Tuning using Reinforcement Learning for Large Language Model · EMNLP 2024
Machine learning › Representation and self-supervised learning
contrastive learning
0.612022
UniCLIP: Unified Framework for Contrastive Language-Image Pre-training · NeurIPS 2022
Machine learning › Representation and self-supervised learning › contrastive learning
contrastive loss
0.612022
UniCLIP: Unified Framework for Contrastive Language-Image Pre-training · NeurIPS 2022
Machine learning › Representation and self-supervised learning › multimodal representation learning
cross-modal representation learning
0.612022
UniCLIP: Unified Framework for Contrastive Language-Image Pre-training · NeurIPS 2022
Computer vision › Vision and language
vision-language pretraining
0.612022
UniCLIP: Unified Framework for Contrastive Language-Image Pre-training · NeurIPS 2022
Robotics › Robot navigation and mapping › mobile robot perception
driving perception
0.312025
SynAD: Enhancing Real-World End-to-End Autonomous Driving Models through Synthetic Data Integration · ICCV 2025
Machine learning › Deep learning architectures and training
data augmentation
0.212024
EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning · ICML 2024
Machine learning › Reinforcement learning
policy optimization
0.212024
StablePrompt : Automatic Prompt Tuning using Reinforcement Learning for Large Language Model · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

synthetic data generation · 0.9reinforcement learning · 0.8proximal policy optimization · 0.8equivariance · 0.8contrastive learning · 0.8attention-based transformation predictor · 0.8augmentation-aware feature embedding · 0.6MP-NCE loss · 0.6
YearPublicationVenuePosition
2025 SynAD: Enhancing Real-World End-to-End Autonomous Driving Models through Synthetic Data Integration
Jongsuk Kim, Gyojin Han, Minki Jeong
ICCV1
2025 FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
Jongsuk Kim, Jaemyung Yu, Minchan Kwon, Junmo Kim 0002
INTERSPEECH1
2024 StablePrompt : Automatic Prompt Tuning using Reinforcement Learning for Large Language Model
abstract
Finding appropriate prompts for the specific task has become an important issue as the usage of Large Language Models (LLM) has expanded.Reinforcement Learning (RL) is widely used for prompt tuning, but its inherent instability and environmental dependency make it difficult to use in practice.In this paper, we propose StablePrompt, which strikes a balance between training stability and search space, mitigating the instability of RL and producing high-performance prompts.We formulate prompt tuning as an online RL problem between the agent and target LLM and introduce Adaptive Proximal Policy Optimization (APPO).APPO introduces an LLM anchor model to adaptively adjust the rate of policy updates.This allows for flexible prompt search while preserving the linguistic ability of the pre-trained LLM.StablePrompt outperforms previous methods on various tasks including text classification, question answering, and text generation.Our code can be found in github.
Minchan Kwon, Gaeun Kim, Jongsuk Kim, Haeil Lee, Junmo Kim 0002
EMNLP3
2024 EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning
abstract
Recent advancements in self-supervised audio-visual representation learning have demonstrated its potential to capture rich and comprehensive representations. However, despite the advantages of data augmentation verified in many learning methods, audio-visual learning has struggled to fully harness these benefits, as augmentations can easily disrupt the correspondence between input pairs. To address this limitation, we introduce EquiAV, a novel framework that leverages equivariance for audio-visual contrastive learning. Our approach begins with extending equivariance to audio-visual learning, facilitated by a shared attention-based transformation predictor. It enables the aggregation of features from diverse augmentations into a representative embedding, providing robust supervision. Notably, this is achieved with minimal computational overhead. Extensive ablation studies and qualitative results verify the effectiveness of our method. EquiAV outperforms previous works across various audio-visual benchmarks. The code is available on https://github.com/JongSuk1/EquiAV
Jongsuk Kim, Hyeongkeun Lee, Kyeongha Rho, Junmo Kim 0002, Joon Son Chung
ICML1
2024 AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning
Jongsuk Kim, Jiwon Shin, Junmo Kim 0002
INTERSPEECH1
2024 Deep dependence in hydroclimatological variables
Taesam Lee, Jongsuk Kim
Appl. Intell.2
2022 UniCLIP: Unified Framework for Contrastive Language-Image Pre-training
abstract
Pre-training vision-language models with contrastive objectives has shown promising results that are both scalable to large uncurated datasets and transferable to many downstream applications. Some following works have targeted to improve data efficiency by adding self-supervision terms, but inter-domain (image-text) contrastive loss and intra-domain (image-image) contrastive loss are defined on individual spaces in those works, so many feasible combinations of supervision are overlooked. To overcome this issue, we propose UniCLIP, a Unified framework for Contrastive Language-Image Pre-training. UniCLIP integrates the contrastive loss of both inter-domain pairs and intra-domain pairs into a single universal space. The discrepancies that occur when integrating contrastive loss between different domains are resolved by the three key components of UniCLIP: (1) augmentation-aware feature embedding, (2) MP-NCE loss, and (3) domain dependent similarity measure. UniCLIP outperforms previous vision-language pre-training methods on various single- and multi-modality downstream tasks. In our experiments, we show that each component that comprises UniCLIP contributes well to the final performance.
Janghyeon Lee 0001, Jongsuk Kim, Hyounguk Shon, Bumsoo Kim 0005, Honglak Lee, Junmo Kim 0002
NeurIPS2