EDBT 2026 Demo / reviewers in the wild / expert
Wonil Song
dblp:193/5955
· DBLP profile ↗
6ranked-venue papers
4as first author
5since 2021 · last 2024
0000-0002-8332-5294ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Reinforcement learning · 74% Representation and self-supervised learning · 17% Deep learning architectures and training · 6% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning |
1.4 | 2 | 2024 | A Simple Framework for Generalization in Visual RL under Dynamic Scene Perturbations · NeurIPS 2024 Local-Guided Global: Paired Similarity Representation for Visual Reinforcement Learning · CVPR 2023 |
Machine learning › Reinforcement learning › deep reinforcement learning
visual reinforcement learning |
1.4 | 2 | 2024 | A Simple Framework for Generalization in Visual RL under Dynamic Scene Perturbations · NeurIPS 2024 Local-Guided Global: Paired Similarity Representation for Visual Reinforcement Learning · CVPR 2023 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
0.7 | 1 | 2023 | Local-Guided Global: Paired Similarity Representation for Visual Reinforcement Learning · CVPR 2023 |
Image and video processing › super-resolution
image super-resolution |
0.4 | 1 | 2020 | Stereoscopic Image Super-Resolution with Stereo Consistent Feature · AAAI 2020 |
Image and video processing › super-resolution › image super-resolution
stereo image super-resolution |
0.4 | 1 | 2020 | Stereoscopic Image Super-Resolution with Stereo Consistent Feature · AAAI 2020 |
Machine learning › Deep learning architectures and training
data augmentation |
0.2 | 1 | 2024 | A Simple Framework for Generalization in Visual RL under Dynamic Scene Perturbations · NeurIPS 2024 |
Computer vision › 3D vision
stereo vision |
0.1 | 1 | 2020 | Stereoscopic Image Super-Resolution with Stereo Consistent Feature · AAAI 2020 |
Methods — techniques the papers use, named apart from their topics
stereo-consistency loss · 0.9self and parallax attention mechanism · 0.9shifted random overlay augmentation · 0.8frame stacking · 0.8self-supervised learning · 0.7contrastive similarity constraints · 0.7action-aware transform · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Improving Self-Supervised Vision Transformers for Visual ControlabstractDespite the tremendous success of vision transformer (ViT) architectures in a broad range of computer vision tasks, the potential of ViT for vision-based deep reinforcement learning (RL) has not been fully explored yet. To improve the performance of the ViT model in visual RL, we propose a simple yet effective approach for self-supervised learning by utilizing the structural capability of a single ViT model, which can learn multiple, distinct representations through extra learnable token embeddings. To this end, in addition to an RL token used for RL input, which corresponds to the classification token in computer vision, we introduce additional extra tokens that are tailored to two auxiliary self-supervised tasks specialized to learn visual and environmental dynamics representations. By interacting with embeddings of these extra tokens through self-attention, our approach provides additional learning signals to the ViT encoder, enabling it to learn more comprehensive representations that are beneficial to RL tasks. In experiments on benchmarks including the DeepMind Control Suite (DMControl) and Atari games, we demonstrate that the proposed approach outperforms the baselines that utilize ViT encoders, particularly achieving state-of-the-art performance in 4 out of 5 tasks in DMControl. Wonil Song, Kwanghoon Sohn, Dongbo Min |
ICIP | 1 |
| 2024 | A Simple Framework for Generalization in Visual RL under Dynamic Scene PerturbationsabstractIn the rapidly evolving domain of vision-based deep reinforcement learning (RL), a pivotal challenge is to achieve generalization capability to dynamic environmental changes reflected in visual observations.
Our work delves into the intricacies of this problem, identifying two key issues that appear in previous approaches for visual RL generalization: (i) imbalanced saliency and (ii) observational overfitting.
Imbalanced saliency is a phenomenon where an RL agent disproportionately identifies salient features across consecutive frames in a frame stack.
Observational overfitting occurs when the agent focuses on certain background regions rather than task-relevant objects.
To address these challenges, we present a simple yet effective framework for generalization in visual RL (SimGRL) under dynamic scene perturbations.
First, to mitigate the imbalanced saliency problem, we introduce an architectural modification to the image encoder to stack frames at the feature level rather than the image level.
Simultaneously, to alleviate the observational overfitting problem, we propose a novel technique called shifted random overlay augmentation, which is specifically designed to learn robust representations capable of effectively handling dynamic visual scenes.
Extensive experiments demonstrate the superior generalization capability of SimGRL, achieving state-of-the-art performance in benchmarks including the DeepMind Control Suite. Wonil Song, Hyesong Choi, Kwanghoon Sohn, Dongbo Min |
NeurIPS | 1 |
| 2023 | Local-Guided Global: Paired Similarity Representation for Visual Reinforcement LearningabstractRecent vision-based reinforcement learning (RL) methods have found extracting high-level features from raw pixels with self-supervised learning to be effective in learning policies. However, these methods focus on learning global representations of images, and disregard local spatial structures present in the consecutively stacked frames. In this paper, we propose a novel approach, termed self-supervised Paired Similarity Representation Learning (PSRL) for effectively encoding spatial structures in an unsupervised manner. Given the input frames, the latent volumes are first generated individually using an encoder, and they are used to capture the variance in terms of local spatial structures, i.e., correspondence maps among multiple frames. This enables for providing plenty of fine-grained samples for training the encoder of deep RL. We further attempt to learn the global semantic representations in the action aware transform module that predicts future state representations using action vectors as a medium. The proposed method imposes similarity constraints on the three latent volumes; transformed query representations by estimated pixel-wise correspondence, predicted query representations from the action aware transform model, and target representations of future state, guiding action aware transform with locality-inherent volume. Experimental results on complex tasks in Atari Games and DeepMind Control Suite demonstrate that the RL methods are significantly boosted by the proposed self-supervised learning of paired similarity representations. Hyesong Choi, Hunsang Lee, Wonil Song, Sangryul Jeon, Kwanghoon Sohn, Dongbo Min |
CVPR | 3 |
| 2023 | Learning disentangled skills for hierarchical reinforcement learning through trajectory autoencoder with weak labels
Wonil Song, Sangryul Jeon, Hyesong Choi, Kwanghoon Sohn, Dongbo Min |
Expert Syst. Appl. | 1 |
| 2021 | Wide and Narrow: Video Prediction from Context and Motion
Jaehoon Cho, Jiyoung Lee 0005, Changjae Oh, Wonil Song, Kwanghoon Sohn |
BMVC | 4 |
| 2020 | Stereoscopic Image Super-Resolution with Stereo Consistent FeatureabstractWe present a first attempt for stereoscopic image super-resolution (SR) for recovering high-resolution details while preserving stereo-consistency between stereoscopic image pair. The most challenging issue in the stereoscopic SR is that the texture details should be consistent for corresponding pixels in stereoscopic SR image pair. However, existing stereo SR methods cannot maintain the stereo-consistency, thus causing 3D fatigue to the viewers. To address this issue, in this paper, we propose a self and parallax attention mechanism (SPAM) to aggregate the information from its own image and the counterpart stereo image simultaneously, thus reconstructing high-quality stereoscopic SR image pairs. Moreover, we design an efficient network architecture and effective loss functions to enforce stereo-consistency constraint. Finally, experimental results demonstrate the superiority of our method over state-of-the-art SR methods in terms of both quantitative metrics and qualitative visual quality while maintaining stereo-consistency between stereoscopic image pair. Wonil Song, Sungil Choi, Somi Jeong, Kwanghoon Sohn |
AAAI | 1 |