EDBT 2026 Demo / reviewers in the wild / expert
Jaejun Yoo 0001
dblp:141/8878-1
· DBLP profile ↗
32ranked-venue papers
5as first author
24since 2021 · last 2025
0000-0001-5252-9668ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 2 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Singular Value Scaling: Efficient Generative Model Compression via Pruned Weights RefinementabstractWhile pruning methods effectively maintain model performance without extra training costs, they often focus solely on preserving crucial connections, overlooking the impact of pruned weights on subsequent fine-tuning or distillation, leading to inefficiencies. Moreover, most compression techniques for generative models have been developed primarily for GANs, tailored to specific architectures like StyleGAN, and research into compressing Diffusion models has just begun. Even more, these methods are often applicable only to GANs or Diffusion models, highlighting the need for approaches that work across both model types. In this paper, we introduce Singular Value Scaling (SVS), a versatile technique for refining pruned weights, applicable to both model types. Our analysis reveals that pruned weights often exhibit dominant singular vectors, hindering fine-tuning efficiency and leading to suboptimal performance compared to random initialization. Our method enhances weight initialization by minimizing the disparities between singular values of pruned weights, thereby improving the fine-tuning process. This approach not only guides the compressed model toward superior solutions but also significantly speeds up fine-tuning. Extensive experiments on StyleGAN2, StyleGAN3 and DDPM demonstrate that SVS improves compression performance across model types without additional training costs. Jaejun Yoo 0001 |
AAAI | 2 |
| 2025 | BF-STVSR: B-Splines and Fourier - Best Friends for High Fidelity Spatial-Temporal Video Super-ResolutionabstractWhile prior methods in Continuous Spatial-Temporal Video Super-Resolution (C-STVSR) employ Implicit Neural Representation (INR) for continuous encoding, they often struggle to capture the complexity of video data, relying on simple coordinate concatenation and pre-trained optical flow networks for motion representation. Interestingly, we find that adding position encoding, contrary to common observations, does not improve—and even degrades—performance. This issue becomes particularly pronounced when combined with pre-trained optical flow networks, which can limit the model’s flexibility. To address these issues, we propose BF-STVSR, a C-STVSR framework with two key modules tailored to better represent spatial and temporal characteristics of video: 1) B-spline Mapper for smooth temporal interpolation, and 2) Fourier Mapper for capturing dominant spatial frequencies. Our approach achieves state-of-the-art in various metrics, including PSNR and SSIM, showing enhanced spatial details and natural temporal consistency. Our code is available ${\color{Cyan}\text{here}}$. Kyong Hwan Jin, Jaejun Yoo 0001 |
CVPR | 4 |
| 2025 | Beyond Spatial Frequency: Pixel-Wise Temporal Frequency-Based Deepfake Video Detection
Taehoon Kim 0004, Jongwook Choi 0001, Yonghyun Jeong, Haeun Noh, Jaejun Yoo 0001, Seungryul Baek |
ICCV | 5 |
| 2025 | Understanding Flatness in Generative Models: Its Role and BenefitsabstractFlat minima, known to enhance generalization and robustness in supervised learning, remain largely unexplored in generative models. In this work, we systematically investigate the role of loss surface flatness in generative models, both theoretically and empirically, with a particular focus on diffusion models. We establish a theoretical claim that flatter minima improve robustness against perturbations in target prior distributions, leading to benefits such as reduced exposure bias -- where errors in noise estimation accumulate over iterations -- and significantly improved resilience to model quantization, preserving generative performance even under strong quantization constraints. We further observe that Sharpness-Aware Minimization (SAM), which explicitly controls the degree of flatness, effectively enhances flatness in diffusion models even surpassing the indirectly promoting flatness methods -- Input Perturbation (IP) which enforces the Lipschitz condition, ensembling-based approach like Stochastic Weight Averaging (SWA) and Exponential Moving Average (EMA) -- are less effective. Through extensive experiments on CIFAR-10, LSUN Tower, and FFHQ, we demonstrate that flat minima in diffusion models indeed improve not only generative performance but also robustness. Taehwan Lee, Kyeongkook Seo, Jaejun Yoo 0001, Sung Whan Yoon |
ICCV | 3 |
| 2025 | PRISM: Privacy-Preserving Improved Stochastic Masking for Federated Generative ModelsabstractDespite recent advancements in federated learning (FL), the integration of generative models into FL has been limited due to challenges such as high communication costs and unstable training in heterogeneous data environments. To address these issues, we propose PRISM, a FL framework tailored for generative models that ensures (i) stable performance in heterogeneous data distributions and (ii) resource efficiency in terms of communication cost and final model size. The key of our method is to search for an optimal stochastic binary mask for a random network rather than updating the model weights, identifying a sparse subnetwork with high generative performance; i.e., a ``strong lottery ticket''. By communicating binary masks in a stochastic manner, PRISM minimizes communication overhead. This approach, combined with the utilization of maximum mean discrepancy (MMD) loss and a mask-aware dynamic moving average aggregation method (MADA) on the server side, facilitates stable and strong generative capabilities by mitigating local divergence in FL scenarios. Moreover, thanks to its sparsifying characteristic, PRISM yields a lightweight model without extra pruning or quantization, making it ideal for environments such as edge devices. Experiments on MNIST, FMNIST, CelebA, and CIFAR10 demonstrate that PRISM outperforms existing methods, while maintaining privacy with minimal communication costs. PRISM is the first to successfully generate images under challenging non-IID and privacy-preserving FL environments on complex datasets, where previous methods have struggled. Kyeongkook Seo, Dong-Jun Han, Jaejun Yoo 0001 |
ICLR | 3 |
| 2025 | MultiDreamer3D: Multi-concept 3D Customization with Concept-Aware Diffusion GuidanceabstractWhile single-concept customization has been studied in 3D, multi-concept customization remains largely unexplored. To address this, we propose MultiDreamer3D that can generate coherent multi-concept 3D content in a divide-and-conquer manner. First, we generate 3D bounding boxes using an LLM-based layout controller. Next, a selective point cloud generator creates coarse point clouds for each concept. These point clouds are placed in the 3D bounding boxes and initialized into 3D Gaussian Splatting with concept labels, enabling precise identification of concept attributions in 2D projections. Finally, we refine 3D Gaussians via concept-aware interval score matching, guided by concept-aware diffusion. Our experimental results show that MultiDreamer3D not only ensures object presence and preserves the distinct identities of each concept but also successfully handles complex cases such as property change or interaction. To the best of our knowledge, we are the first to address the multi-concept customization in 3D. Wooseok Song, Seunggyu Chang, Jaejun Yoo 0001 |
IJCAI | 3 |
| 2025 | Dynamic-Aware Spatio-Temporal Representation Learning for Dynamic MRI Reconstruction
Dayoung Baik, Jaejun Yoo 0001 |
MICCAI (4) | 2 |
| 2025 | Spherical Diffusion Process for Score-Guided Cortical Correspondence via Spectral Attention
Seungeun Lee, Sergey Pyatkovskiy, Jaejun Yoo 0001, Ilwoo Lyu |
MICCAI (16) | 3 |
| 2024 | MOVES: Motion-Oriented VidEo Sampling for Natural Language-Based Vehicle RetrievalabstractRetrieving the target vehicle through natural language descriptions plays a crucial role in intelligent transportation systems. Existing methods tackle this task by employing models that leverage the correlation between textual and visual representations, such as CLIP. However, these models struggle to capture the temporal characteristics of video data, and researchers enhance temporal understanding performance through various data augmentation and video encoders. Yet, conventional approaches in previous studies often overlook the detailed temporal characteristics of vehicles. To overcome this limitation, we introduce a MOVES: Motion-Oriented VidEo Sampling method to effectively utilize the motion information of the target vehicle. Furthermore, we construct a robust model by implementing a re-ranking algorithm to address a variety of vehicle attributes. As a result, our proposed model achieves state-of-the-art performance on the public vehicle retrieval dataset. Dongyoung Kim, Kyoungoh Lee, In-Su Jang, Kwang-Ju Kim, Pyong-Kun Kim, Jaejun Yoo 0001 |
AVSS | 6 |
| 2024 | Hybrid Video Diffusion Models with 2D Triplane and 3D Wavelet Representation
Haneol Lee, Kwanghee Lee, Seungryong Kim, Jaejun Yoo 0001 |
ECCV (52) | 7 |
| 2024 | PosterLlama: Bridging Design Ability of Language Model to Content-Aware Layout Generation
Jaejung Seol, Seojun Kim, Jaejun Yoo 0001 |
ECCV (82) | 3 |
| 2024 | Nickel and Diming Your GAN: A Dual-Method Approach to Enhancing GAN Efficiency via Knowledge Distillation
Sangyeop Yeo, Yoojin Jang 0001, Jaejun Yoo 0001 |
ECCV (88) | 3 |
| 2024 | STREAM: Spatio-TempoRal Evaluation and Analysis Metric for Video Generative ModelsabstractImage generative models have made significant progress in generating realistic and diverse images, supported by comprehensive guidance from various evaluation metrics. However, current video generative models struggle to generate even
short video clips, with limited tools that provide insights for improvements. Current video evaluation metrics are simple adaptations of image metrics by switching the embeddings with video embedding networks, which may underestimate the unique characteristics of video. Our analysis reveals that the widely used Frechet Video Distance (FVD) has a stronger emphasis on the spatial aspect than the temporal naturalness of video and is inherently constrained by the input size of the embedding networks used, limiting it to 16 frames. Additionally, it demonstrates considerable instability and diverges from human evaluations. To address the limitations, we propose STREAM, a new video evaluation metric uniquely designed to independently evaluate spatial and temporal aspects. This feature allows comprehensive analysis and evaluation of video generative models from various perspectives, unconstrained by video length. We provide analytical and experimental evidence demonstrating that STREAM provides an effective evaluation tool for both visual and temporal quality of videos, offering insights into area of improvement for video generative models. To the best of our knowledge, STREAM is the first evaluation metric that can separately assess the temporal and spatial aspects of videos. Our code is available at https://github.com/pro2nit/STREAM. Pum Jun Kim, Seojun Kim, Jaejun Yoo 0001 |
ICLR | 3 |
| 2024 | RADIO: Reference-Agnostic Dubbing Video SynthesisabstractOne of the most challenging problems in audio-driven talking head generation is achieving high-fidelity detail while ensuring precise synchronization. Given only a single reference image, extracting meaningful identity attributes becomes even more challenging, often causing the network to mirror the facial and lip structures too closely. To address these issues, we introduce RADIO, a framework engineered to yield high-quality dubbed videos regardless of the pose or expression in reference images. The key is to modulate the decoder layers using latent space composed of audio and reference features. Additionally, we incorporate ViT blocks into the decoder to emphasize high-fidelity details, especially in the lip region. Our experimental results demonstrate that RADIO displays high synchronization without the loss of fidelity. Especially in harsh scenarios where the reference frame deviates significantly from the ground truth, our method outperforms state-of-the-art methods, highlighting its robustness. Dongyeun Lee, Sangjoon Yu, Jaejun Yoo 0001, Gyeong-Moon Park |
WACV | 4 |
| 2024 | Data Augmentation for Low-Level Vision: CutBlur and Mixture-of-Augmentation
Namhyuk Ahn, Jaejun Yoo 0001, Kyung-Ah Sohn 0001 |
Int. J. Comput. Vis. | 2 |
| 2024 | Bridging the Domain Gap: A Simple Domain Matching Method for Reference-Based Image Super-Resolution in Remote SensingabstractRecently, reference-based image super-resolution (RefSR) has shown excellent performance in image super-resolution (SR) tasks. The main idea of RefSR is to utilize additional information from the reference (Ref) image to recover the high-frequency components in low-resolution (LR) images. By transferring relevant textures through feature matching, RefSR models outperform existing single-image SR (SISR) models. However, their performance significantly declines when a domain gap between Ref and LR images exists, which often occurs in real-world scenarios, such as satellite imaging. In this letter, we introduce a domain matching (DM) module that can be seamlessly integrated with existing RefSR models to enhance their performance in a plug-and-play manner. To the best of our knowledge, we are the first to explore DM-based RefSR in remote sensing image processing. Our analysis reveals that their domain gaps often occur in different satellites, and our model effectively addresses these challenges, whereas existing models struggle. Our experiments demonstrate that the proposed DM module improves SR performance both qualitatively and quantitatively for remote sensing SR tasks. Jeongho Min, Yejun Lee, Dongyoung Kim, Jaejun Yoo 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Universal Dehazing via Haze Style TransferabstractSingle image dehazing has been actively studied to overcome the quality degradation of hazy images. Most of the existing methods take model-based approaches and the existing learning-based methods usually target specific haze styles only, e.g., daytime, varicolored, and nighttime haze. Therefore, they suffer from the limited performance on arbitrary hazy images with diverse characteristics due to the lack of universal training dataset. In this paper, we first propose a fully data-driven learning-based framework for universal dehazing based on the haze style transfer (HST). We define multiple domains of haze styles by applying the K-means clustering to the background light of diverse real hazy images. We design the haze style modulator to extract the scene radiance features and the haze-related features, respectively. We employ the unpaired image-to-image translation methodology to transfer a source hazy image into different hazy images with diverse styles while preserving the scene radiance. The generated diverse hazy images are used to train the universal dehazing network in a semi-supervised manner, where we implement the dehazing as a special instance of HST into no haze style. The experimental results show that the proposed framework reliably generates realistic and diverse hazy images, and achieves better performance of universal dehazing regardless of the haze styles compared with the existing state-of-the art dehazing methods. Eunpil Park, Jaejun Yoo 0001, Jae-Young Sim |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Can We Find Strong Lottery Tickets in Generative Models?abstractYes. In this paper, we investigate strong lottery tickets in generative models, the subnetworks that achieve good generative performance without any weight update. Neural network pruning is considered the main cornerstone of model compression for reducing the costs of computation and memory. Unfortunately, pruning a generative model has not been extensively explored, and all existing pruning algorithms suffer from excessive weight-training costs, performance degradation, limited generalizability, or complicated training. To address these problems, we propose to find a strong lottery ticket via moment-matching scores. Our experimental results show that the discovered subnetwork can perform similarly or better than the trained dense model even when only 10% of the weights remain. To the best of our knowledge, we are the first to show the existence of strong lottery tickets in generative models and provide an algorithm to find it stably. Our code and supplementary materials are publicly available at https://lait-cvlab.github.io/SLT-in-Generative-Models/. Sangyeop Yeo, Yoojin Jang 0001, Jy-yong Sohn, Dongyoon Han, Jaejun Yoo 0001 |
AAAI | 5 |
| 2023 | Fix the Noise: Disentangling Source Feature for Controllable Domain TranslationabstractRecent studies show strong generative performance in domain translation especially by using transfer learning techniques on the unconditional generator. However, the control between different domain features using a single model is still challenging. Existing methods often require additional models, which is computationally demanding and leads to unsatisfactory visual quality. In addition, they have restricted control steps, which prevents a smooth transition. In this paper, we propose a new approach for high-quality domain translation with better controllability. The key idea is to preserve source features within a disentangled subspace of a target feature space. This allows our method to smoothly control the degree to which it preserves source features while generating images from an entirely new domain using only a single model. Our extensive experiments show that the proposed method can produce more consistent and realistic images than previous works and maintain precise controllability over different levels of transformation. The code is available at LeeDongYeun/FixNoise. Dongyeun Lee, Jae Young Lee 0002, Jaehyun Choi, Jaejun Yoo 0001, Junmo Kim 0002 |
CVPR | 5 |
| 2023 | LANIT: Language-Driven Image-to-Image Translation for Unlabeled DataabstractExisting techniques for image-to-image translation commonly have suffered from two critical problems: heavy reliance on per-sample domain annotation and/or inability to handle multiple attributes per image. Recent truly-unsupervised methods adopt clustering approaches to easily provide per-sample one-hot domain labels. However, they cannot account for the real-world setting: one sample may have multiple attributes. In addition, the semantics of the clusters are not easily coupled to human understanding. To overcome these, we present LANguage-driven Image-to-image Translation model, dubbed LANIT. We leverage easy-to-obtain candidate attributes given in texts for a dataset: the similarity between images and attributes indicates per-sample domain labels. This formulation naturally enables multi-hot labels so that users can specify the target domain with a set of attributes in language. To account for the case that the initial prompts are inaccurate, we also present prompt learning. We further present domain regularization loss that enforces translated images to be mapped to the corresponding domain. Experiments on several standard benchmarks demonstrate that LANIT achieves comparable or superior performance to existing models. The code is available at github.com/KU-CVLAB/LANIT. Seokju Cho, Jaejun Yoo 0001, Youngjung Uh, Seungryong Kim |
CVPR | 5 |
| 2023 | Towards Robust Contrail Detection by Mitigating Label Bias via a Probabilistic Deep Learning Model: A Preliminary StudyabstractContrails, formed by jet flights, alter Earth's energy balance, prompting research into monitoring contrails and developing satellite-based automated contrail detection. This demand has advanced deep learning (DL)-based techniques. However, training DL algorithms to detect contrails has limitations: class imbalance, labeling difficulty, and a lack of reliable labeled datasets. We propose a probabilistic DL approach using P-UNet to alleviate label bias in contrail detection. By observing model outputs based on two labeled datasets, OpenContrails and MIT-Contrails, we found the probabilistic approach robust against potentially biased labels. Yejun Lee, Eun-Kyeong Kim, Jaejun Yoo 0001 |
SIGSPATIAL/GIS | 3 |
| 2023 | TopP&R: Robust Support Estimation Approach for Evaluating Fidelity and Diversity in Generative ModelsabstractWe propose a robust and reliable evaluation metric for generative models called Topological Precision and Recall (TopP&R, pronounced “topper”), which systematically estimates supports by retaining only topologically and statistically significant features with a certain level of confidence. Existing metrics, such as Inception Score (IS), Frechet Inception Distance (FID), and various Precision and Recall (P&R) variants, rely heavily on support estimates derived from sample features. However, the reliability of these estimates has been overlooked, even though the quality of the evaluation hinges entirely on their accuracy. In this paper, we demonstrate that current methods not only fail to accurately assess sample quality when support estimation is unreliable, but also yield inconsistent results. In contrast, TopP&R reliably evaluates the sample quality and ensures statistical consistency in its results. Our theoretical and experimental findings reveal that TopP&R provides a robust evaluation, accurately capturing the true trend of change in samples, even in the presence of outliers and non-independent and identically distributed (Non-IID) perturbations where other methods result in inaccurate support estimations. To our knowledge, TopP&R is the first evaluation metric specifically focused on the robust estimation of supports, offering statistical consistency under noise conditions. Pum Jun Kim, Yoojin Jang 0001, Jaejun Yoo 0001 |
NeurIPS | 4 |
| 2021 | Rethinking the Truly Unsupervised Image-to-Image TranslationabstractEvery recent image-to-image translation model inherently requires either image-level (i.e. input-output pairs) or set-level (i.e. domain labels) supervision. However, even set-level supervision can be a severe bottleneck for data collection in practice. In this paper, we tackle image-to-image translation in a fully unsupervised setting, i.e., neither paired images nor domain labels. To this end, we propose a truly unsupervised image-to-image translation model (TUNIT) that simultaneously learns to separate image domains and translates input images into the estimated domains. Experimental results show that our model achieves comparable or even better performance than the set-level supervised model trained with full labels, generalizes well on various datasets, and is robust against the choice of hyperparameters (e.g. the preset number of pseudo domains). Furthermore, TUNIT can be easily extended to semi-supervised learning with a few labeled data. Kyungjune Baek, Yunjey Choi, Youngjung Uh, Jaejun Yoo 0001, Hyunjung Shim |
ICCV | 4 |
| 2021 | Time-Dependent Deep Image Prior for Dynamic MRIabstractWe propose a novel unsupervised deep-learning-based algorithm for dynamic magnetic resonance imaging (MRI) reconstruction. Dynamic MRI requires rapid data acquisition for the study of moving organs such as the heart. We introduce a generalized version of the deep-image-prior approach, which optimizes the weights of a reconstruction network to fit a sequence of sparsely acquired dynamic MRI measurements. Our method needs neither prior training nor additional data. In particular, for cardiac images, it does not require the marking of heartbeats or the reordering of spokes. The key ingredients of our method are threefold: 1) a fixed low-dimensional manifold that encodes the temporal variations of images; 2) a network that maps the manifold into a more expressive latent space; and 3) a convolutional neural network that generates a dynamic series of MRI images from the latent variables and that favors their consistency with the measurements in k -space. Our method outperforms the state-of-the-art methods quantitatively and qualitatively in both retrospective and real fetal cardiac datasets. To the best of our knowledge, this is the first unsupervised deep-learning-based method that can reconstruct the continuous variation of dynamic MRI sequences with high spatial resolution. Jaejun Yoo 0001, Kyong Hwan Jin, Jérôme Yerly, Matthias Stuber, Michael Unser |
IEEE Trans. Medical Imaging | 1 |
| 2020 | StarGAN v2: Diverse Image Synthesis for Multiple DomainsabstractA good image-to-image translation model should learn a mapping between different visual domains while satisfying the following properties: 1) diversity of generated images and 2) scalability over multiple domains. Existing methods address either of the issues, having limited diversity or multiple models for all domains. We propose StarGAN v2, a single framework that tackles both and shows significantly improved results over the baselines. Experiments on CelebA-HQ and a new animal faces dataset (AFHQ) validate our superiority in terms of visual quality, diversity, and scalability. To better assess image-to-image translation models, we release AFHQ, high-quality animal faces with large inter- and intra-domain differences. The code, pretrained models, and dataset are available at https://github.com/clovaai/stargan-v2. Yunjey Choi, Youngjung Uh, Jaejun Yoo 0001, Jung-Woo Ha 0001 |
CVPR | 3 |
| 2020 | Rethinking Data Augmentation for Image Super-resolution: A Comprehensive Analysis and a New StrategyabstractData augmentation is an effective way to improve the performance of deep networks. Unfortunately, current methods are mostly developed for high-level vision tasks (e.g., classification) and few are studied for low-level vision tasks (e.g., image restoration). In this paper, we provide a comprehensive analysis of the existing augmentation methods applied to the super-resolution task. We find that the methods discarding or manipulating the pixels or features too much hamper the image restoration, where the spatial relationship is very important. Based on our analyses, we propose CutBlur that cuts a low-resolution patch and pastes it to the corresponding high-resolution image region and vice versa. The key intuition of CutBlur is to enable a model to learn not only "how" but also "where" to super-resolve an image. By doing so, the model can understand "how much", instead of blindly learning to apply super-resolution to every given pixel. Our method consistently and significantly improves the performance across various scenarios, especially when the model size is big and the data is collected under real-world environments. We also show that our method improves other low-level vision tasks, such as denoising and compression artifact removal. Jaejun Yoo 0001, Namhyuk Ahn, Kyung-Ah Sohn 0001 |
CVPR | 1 |
| 2020 | Reliable Fidelity and Diversity Metrics for Generative ModelsabstractDevising indicative evaluation metrics for the image generation task remains an open problem. The most widely used metric for measuring the similarity between real and generated images has been the Frechet Inception Distance (FID) score. Since it does not differentiate the fidelity and diversity aspects of the generated images, recent papers have introduced variants of precision and recall metrics to diagnose those properties separately. In this paper, we show that even the latest version of the precision and recall metrics are not reliable yet. For example, they fail to detect the match between two identical distributions, they are not robust against outliers, and the evaluation hyperparameters are selected arbitrarily. We propose density and coverage metrics that solve the above issues. We analytically and experimentally show that density and coverage provide more interpretable and reliable signals for practitioners than the existing metrics. Muhammad Ferjad Naeem, Seong Joon Oh, Youngjung Uh, Yunjey Choi, Jaejun Yoo 0001 |
ICML | 5 |
| 2020 | Deep Learning Diffuse Optical TomographyabstractDiffuse optical tomography (DOT) has been investigated as an alternative imaging modality for breast cancer detection thanks to its excellent contrast to hemoglobin oxidization level. However, due to the complicated non-linear photon scattering physics and ill-posedness, the conventional reconstruction algorithms are sensitive to imaging parameters such as boundary conditions. To address this, here we propose a novel deep learning approach that learns non-linear photon scattering physics and obtains an accurate three dimensional (3D) distribution of optical anomalies. In contrast to the traditional black-box deep learning approaches, our deep network is designed to invert the Lippman-Schwinger integral equation using the recent mathematical theory of deep convolutional framelets. As an example of clinical relevance, we applied the method to our prototype DOT system. We show that our deep neural network, trained with only simulation data, can accurately recover the location of anomalies within biomimetic phantoms and live animals without the use of an exogenous contrast agent. Jaejun Yoo 0001, Sohail Sabir, Duchang Heo, Kee Hyun Kim, Abdul Wahab 0002, Yoonseok Choi, Seul-I Lee, Eun Young Chae, Hak Hee Kim, Young Min Bae, Young-Wook Choi, Seungryong Cho, Jong Chul Ye |
IEEE Trans. Medical Imaging | 1 |
| 2019 | Photorealistic Style Transfer via Wavelet TransformsabstractRecent style transfer models have provided promising artistic results. However, given a photograph as a reference style, existing methods are limited by spatial distortions or unrealistic artifacts, which should not happen in real photographs. We introduce a theoretically sound correction to the network architecture that remarkably enhances photorealism and faithfully transfers the style. The key ingredient of our method is wavelet transforms that naturally fits in deep networks. We propose a wavelet corrected transfer based on whitening and coloring transforms (WCT2) that allows features to preserve their structural information and statistical properties of VGG feature space during stylization. This is the first and the only end-to-end model that can stylize a 1024x1024 resolution image in 4.7 seconds, giving a pleasing and photorealistic quality without any post-processing. Last but not least, our model provides a stable video stylization without temporal constraints. Our code, generated images, pre-trained models and supplementary documents are all available at https://github.com/ClovaAI/WCT2. Jaejun Yoo 0001, Youngjung Uh, Sanghyuk Chun, Byeongkyu Kang, Jung-Woo Ha 0001 |
ICCV | 1 |
| 2019 | Large-Scale Answerer in Questioner's Mind for Visual Dialog Question Generation
Sang-Woo Lee 0001, Sohee Yang, Jaejun Yoo 0001, Jung-Woo Ha 0001 |
ICLR (Poster) | 4 |
| 2018 | Deep Convolutional Framelet Denosing for Low-Dose CT via Wavelet Residual NetworkabstractModel-based iterative reconstruction algorithms for low-dose X-ray computed tomography (CT) are computationally expensive. To address this problem, we recently proposed a deep convolutional neural network (CNN) for low-dose X-ray CT and won the second place in 2016 AAPM Low-Dose CT Grand Challenge. However, some of the textures were not fully recovered. To address this problem, here we propose a novel framelet-based denoising algorithm using wavelet residual network which synergistically combines the expressive power of deep learning and the performance guarantee from the framelet-based denoising algorithms. The new algorithms were inspired by the recent interpretation of the deep CNN as a cascaded convolution framelet signal representation. Extensive experimental results confirm that the proposed networks have significantly improved performance and preserve the detail texture of the original images. Eunhee Kang, Won Chang, Jaejun Yoo 0001, Jong Chul Ye |
IEEE Trans. Medical Imaging | 3 |
| 2017 | A Joint Sparse Recovery Framework for Accurate Reconstruction of Inclusions in Elastic MediaabstractA robust algorithm is proposed to reconstruct the spatial support and the Lamé parameters of multiple inclusions in a homogeneous background elastic material using a few measurements of the displacement field over a finite collection of boundary points. The algorithm does not require any linearization or iterative update of Green's function but still allows very accurate reconstruction. The breakthrough comes from a novel interpretation of Lippmann--Schwinger type integral representation of the displacement field in terms of unknown densities having common sparse support on the location of inclusions. Accordingly, the proposed algorithm consists of a two-step approach. First, the localization problem is recast as a joint sparse recovery problem that renders the densities and the inclusion support simultaneously. Then, a noise robust constrained optimization problem is formulated for the reconstruction of elastic parameters. An efficient algorithm is designed for numerical implementation using the Multiple Sparse Bayesian Learning (M-SBL) for joint sparse recovery problem and the Constrained Split Augmented Lagrangian Shrinkage Algorithm (C-SALSA) for the constrained optimization problem. The efficacy of the proposed framework is manifested through extensive numerical simulations. To the best of our knowledge, this is the first algorithm tailored for parameter reconstruction problems in elastic media using highly under-sampled data in the sense of Nyquist rate. Jaejun Yoo 0001, Younghoon Jung, Mikyoung Lim, Jong Chul Ye, Abdul Wahab 0002 |
SIAM J. Imaging Sci. | 1 |