EDBT 2026 Demo / reviewers in the wild / expert
Juan Luis Gonzalez 0001
dblp:228/6824 · also Juan Luis Gonzalez Bello
· DBLP profile ↗
14ranked-venue papers
10as first author
10since 2021 · last 2026
0000-0003-3963-4433ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 9 · 6 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoBGS: Motion Deblurring Dynamic 3D Gaussian Splatting for Blurry Monocular VideoabstractWe present MoBGS, a novel motion deblurring 3D Gaussian Splatting (3DGS) framework capable of reconstructing sharp and high-quality novel spatio-temporal views from blurry monocular videos in an end-to-end manner. Existing dynamic novel view synthesis (NVS) methods are highly sensitive to motion blur in casually captured videos, resulting in significant degradation of rendering quality. While recent approaches address motion-blurred inputs for NVS, they primarily focus on static scene reconstruction and lack dedicated motion modeling for dynamic objects. To overcome these limitations, our MoBGS introduces a novel Blur-adaptive Latent Camera Estimation (BLCE) method using a proposed Blur-adaptive Neural Ordinary Differential Equation (ODE) solver for effective latent camera trajectory estimation, improving global camera motion deblurring. In addition, we propose a Latent Camera-induced Exposure Estimation (LCEE) method to ensure consistent deblurring of both a global camera and local object motions. Extensive experiments on the Stereo Blur dataset and real-world blurry videos show that our MoBGS significantly outperforms the very recent methods, achieving state-of-the-art performance for dynamic NVS under motion blur. Minh-Quan Viet Bui, Jongmin Park 0001, Juan Luis Gonzalez 0001, Jaeho Moon, Jihyong Oh, Munchurl Kim |
AAAI | 3 |
| 2026 | Zero-Shot Video Deraining with Video Diffusion ModelsabstractExisting video deraining methods are often trained on paired datasets, either synthetic, which limits their ability to generalize to real-world rain, or captured by static cameras, which restricts their effectiveness in dynamic scenes with background and camera motion. Furthermore, recent works in fine-tuning diffusion models have shown promising results, but the fine-tuning tends to weaken the generative prior, limiting generalization to unseen cases. In this paper, we introduce the first zero-shot video deraining method for complex dynamic scenes that does not require synthetic data nor model fine-tuning, by leveraging a pretrained text-to-video diffusion model that demonstrates strong generalization capabilities. By inverting an input video into the latent space of diffusion models, its reconstruction process can be intervened and pushed away from the model’s concept of rain using negative prompting. At the core of our approach is an attention switching mechanism that we found is crucial for maintaining dynamic backgrounds as well as structural consistency between the input and the derained video, mitigating artifacts introduced by naive negative prompting. Our approach is validated through extensive experiments on real-world rain datasets, demonstrating substantial improvements over prior methods and showcasing robust generalization without the need for supervised training. Tuomas Varanka, Juan Luis Gonzalez 0001, Hyeongwoo Kim, Pablo Garrido 0001 |
WACV | 2 |
| 2025 | VideoSPatS: Video SPatiotemporal Splines for Disentangled Occlusion, Appearance and Motion Modeling and EditingabstractWe present an implicit video representation for occlusions, appearance, and motion disentanglement from monocular videos, which we call Video SPatiotemporal Splines (VideoSPatS). Unlike previous methods that map time and coordinates to deformation and canonical colors, our VideoSPatS maps input coordinates into Spatial and Color Spline deformation fields ${\mathcal{D}_{\text{s}}}$ and ${\mathcal{D}_c}$, which disentangle motion and appearance in videos. With spline-based parametrization, our method naturally generates temporally consistent flow and guarantees long-term temporal consistency, which is crucial for convincing video editing. Using multiple prediction branches, our VideoSPatS model also performs layer separation between the latent video and the selected occluder. By disentangling occlusions, appearance, and motion, our method enables better spatiotemporal modeling and editing of diverse videos, including in-the-wild talking head videos with challenging occlusions, shadows, and specularities while maintaining an appropriate canonical space for editing. We also present general video modeling results on the DAVIS and CoDeF datasets, as well as our own talking head video dataset collected from open-source web videos. Extensive ablations show the combination of ${\mathcal{D}_{\text{s}}}$ and ${\mathcal{D}_c}$ under neural splines can overcome motion and appearance ambiguities, paving the way for more advanced video editing models. Visit our project site1. Juan Luis Gonzalez 0001, Alex Whelan, Kyle Olszewski, Hyeongwoo Kim, Pablo Garrido 0001 |
CVPR | 1 |
| 2025 | SplineGS: Robust Motion-Adaptive Spline for Real-Time Dynamic 3D Gaussians from Monocular VideoabstractSynthesizing novel views from in-the-wild monocular videos is challenging due to scene dynamics and the lack of multi-view cues. To address this, we propose SplineGS, a COLMAP-free dynamic 3D Gaussian Splatting (3DGS) framework for high-quality reconstruction and fast rendering from monocular videos. At its core is a novel Motion-Adaptive Spline (MAS) method, which represents continuous dynamic 3D Gaussian trajectories using cubic Her-mite splines with a small number of control points. For MAS, we introduce a Motion-Adaptive Control points Pruning (MACP) method to model the deformation of each dynamic 3D Gaussian across varying motions, progressively pruning control points while maintaining dynamic modeling integrity. Additionally, we present a joint optimization strategy for camera parameter estimation and 3D Gaussian attributes, leveraging photometric and geometric consistency. This eliminates the need for Structure-from-Motion preprocessing and enhances SplineGS’s robustness in real-world conditions. Experiments show that SplineGS significantly outperforms state-of-the-art methods in novel view synthesis quality for dynamic scenes from monocular videos, achieving thousands times faster rendering speed. Jongmin Park 0001, Minh-Quan Viet Bui, Juan Luis Gonzalez 0001, Jaeho Moon, Jihyong Oh, Munchurl Kim |
CVPR | 3 |
| 2024 | Novel View Synthesis with View-Dependent Effects from a Single ImageabstractIn this paper, we address single image-based novel view synthesis (NVS) by firstly integrating view-dependent effects (VDE) into the process. Our approach leverages camera motion priors to model VDE, treating negative disparity as the representation of these effects in the scene. By identifying that specularities align with camera motion, we infuse VDEs into input images by aggregating pixel colors along the negative depth region of epipolar lines. Additionally, we introduce a ‘relaxed volumetric rendering’ approximation, enhancing efficiency by computing densities in a single pass for NVS from single images. Notably, our method learns single-image NVS from image sequences alone, making it a fully self-supervised learning approach that requires no depth or camera pose annotations. We present extensive experimental results and show that our proposed method can learn NVS with VDEs, outperforming the SOTA single-view NVS methods on the RealEstate10k and MannequinChallenge datasets. Visit our project site11https://kaist-viclab.github.io/monovde-site. Juan Luis Gonzalez 0001, Munchurl Kim |
CVPR | 1 |
| 2024 | From-Ground-To-Objects: Coarse-to-Fine Self-supervised Monocular Depth Estimation of Dynamic Objects with Ground Contact PriorabstractSelf-supervised monocular depth estimation (DE) is an approach to learning depth without costly depth ground truths. However, it often struggles with moving objects that violate the static scene assumption during training. To address this issue, we introduce a coarse-to-fine training strategy leveraging the ground contacting prior based on the observation that most moving objects in outdoor scenes contact the ground. In the coarse training stage, we exclude the objects in dynamic classes from the reprojection loss calculation to avoid inaccurate depth learning. To provide precise supervision on the depth of the objects, we present a novel Ground-contacting-prior Disparity Smoothness Loss (GDS-Loss) that encourages a DE network to align the depth of the objects with their ground-contacting points. Subsequently, in the fine training stage, we refine the DE network to learn the detailed depth of the objects from the reprojection loss, while ensuring accurate DE on the moving object regions by employing our regularization loss with a cost-volume-based weighting factor. Our overall coarse-to-fine training strategy can easily be integrated with existing DE methods without any modifications, significantly enhancing DE performance on challenging Cityscapes and KITTI datasets, especially in the moving object regions. Jaeho Moon, Juan Luis Gonzalez 0001, Byeongjun Kwon, Munchurl Kim |
CVPR | 2 |
| 2024 | Self-Supervised Monocular Depth Estimation With Positional Shift Depth Variance and Adaptive Disparity QuantizationabstractRecently, attempts to learn the underlying 3D structures of a scene from monocular videos in a fully self-supervised fashion have drawn much attention. One of the most challenging aspects of this task is to handle independently moving objects as they break the rigid-scene assumption. In this paper, we show for the first time that pixel positional information can be exploited to learn SVDE (Single View Depth Estimation) from videos. The proposed moving object (MO) masks, which are induced by the depth variance to shifted positional information (SPI) and are referred to as 'SPIMO' masks, are highly robust and consistently remove independently moving objects from the scenes, allowing for robust and consistent learning of SVDE from videos. Additionally, we introduce a new adaptive quantization scheme that assigns the best per-pixel quantization curve for depth discretization, improving the fine granularity and accuracy of the final aggregated depth maps. Finally, we employ existing boosting techniques in a new way that self-supervises moving object depths further. With these features, our pipeline is robust against moving objects and generalizes well to high-resolution images, even when trained with small patches, yielding state-of-the-art (SOTA) results with four- to eight-fold fewer parameters than the previous SOTA techniques that learn from videos. We present extensive experiments on KITTI and CityScapes that show the effectiveness of our method. Juan Luis Gonzalez 0001, Jaeho Moon, Munchurl Kim |
IEEE Trans. Image Process. | 1 |
| 2023 | Optical-Flow-Aided Self-Supervised Single View Depth Estimation from Monocular VideosabstractRecent advances in deep-learning-based accurate optical flow (OF) estimation have sparked interest in building hardware dedicated to it, making OF a real-time commodity for downstream computer vision tasks. Relative camera pose, and depth estimations (DE) are closely related to OF, but obtaining one from the others is not a trivial task. While previous works have attempted simultaneous learning of OF and DE, or utilized OF-nets only for camera pose prediction and sparse depth map supervision, we propose OF-aided self-supervised DE, that is, pre-computed OF is used as a network input. One of the main challenges in OF-aided DE is preventing the DE network from learning an incorrect OF magnitude prior, which is valid for rigid regions but breaks for scenes with dynamic objects or static camera sequences. We achieve OF-aided DE by instead transforming the input OF into a 3D surface normals space, which provides a well-informed geometrical input while being invariant to the OF magnitude. In combination with a new training strategy, we show that our models with OF-aided DE achieve state-of-the-art (SOTA) results on the KITTI dataset. Juan Luis Gonzalez 0001, Munchurl Kim |
VCIP | 1 |
| 2022 | Self-Supervised Deep Monocular Depth Estimation With Ambiguity BoostingabstractWe propose a novel two-stage training strategy with ambiguity boosting for the self-supervised learning of single view depths from stereo images. Our proposed two-stage learning strategy first aims to obtain a coarse depth prior by training an auto-encoder network for a stereoscopic view synthesis task. This prior knowledge is then boosted and used to self-supervise the model in the second stage of training in our novel ambiguity boosting loss. Our ambiguity boosting loss is a confidence-guided type of data augmentation loss that improves the accuracy and consistency of generated depth maps under several transformations of the single-image input. To show the benefits of the proposed two-stage training strategy with boosting, our two previous depth estimation (DE) networks, one with t-shaped adaptive kernels and the other with exponential disparity volumes, are extended with our new learning strategy, referred to as DBoosterNet-t and DBoosterNet-e, respectively. Our self-supervised DBoosterNets are competitive, and in some cases even better, compared to the most recent supervised SOTA methods, and are remarkably superior to the previous self-supervised methods for monocular DE on the challenging KITTI dataset. We present intensive experimental results, showing the efficacy of our method for the self-supervised monocular DE task. Juan Luis Gonzalez 0001, Munchurl Kim |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | PLADE-Net: Towards Pixel-Level Accuracy for Self-Supervised Single-View Depth Estimation With Neural Positional Encoding and Distilled Matting LossabstractIn this paper, we propose a self-supervised singleview pixel-level accurate depth estimation network, called PLADE-Net. The PLADE-Net is the first work that shows remarkable accuracy levels, exceeding 95% in terms of the δ1metric on the challenging KITTI dataset. Our PLADENet is based on a new network architecture with neural positional encoding and a novel loss function that borrows from the closed-form solution of the matting Laplacian to learn pixel-level accurate depth estimation from stereo images. Neural positional encoding allows our PLADENet to obtain more consistent depth estimates by letting the network reason about location-specific image properties such as projection (and potentially lens) distortions. Our novel distilled matting Laplacian loss allows our network to predict sharp depths at object boundaries and more consistent depths in highly homogeneous regions. Our proposed method outperforms all previous self-supervised single-view depth estimation methods by a large margin on the challenging KITTI dataset, with unparalleled levels of accuracy. Furthermore, our PLADE-Net, naively extended for stereo inputs, outperforms the most recent self-supervised stereo methods, even without any advanced blocks like 1D correlations, 3D convolutions, or spatial pyramid pooling. We present extensive ablation studies and experiments that support our method’s effectiveness on the KITTI, CityScapes, and Make3D datasets. Juan Luis Gonzalez 0001, Munchurl Kim |
CVPR | 1 |
| 2020 | Pan-Sharpening With Color-Aware Perceptual Loss And Guided Re-ColorizationabstractIn remote sensing, “pan-sharpening” is the task of enhancing the spatial resolution of a multi-spectral (MS) image by exploiting the high-frequency information in a panchromatic (PAN) reference image. We present a novel color-aware perceptual (CAP) loss for learning the task of pan-sharpening. Our CAP loss is designed to focus on the deep features of a pre-trained VGG network that are more sensitive to spatial details and ignore color information to allow the network to extract the structural information from the PAN image while keeping the color from the lower resolution MS image. Additionally, we propose “guided re-colorization”, which generates a pan-sharpened image with real colors from the MS input by “picking” the closest MS pixel color for each pan-sharpened pixel, as a human operator would do in manual colorization. Such a re-colorized (RC) image is completely aligned with the pan-sharpened (PS) network output and can be used as a self-supervision signal during training, or to enhance the colors in the PS image during test. We present several experiments where our network trained with our CAP loss generates naturally looking pan-sharpened images with fewer artifacts and outperforms the state-of-the-arts on the WorldView3 dataset in terms of ERGAS, SCC, and QNR metrics. Juan Luis Gonzalez 0001, Soomin Seo, Munchurl Kim |
ICIP | 1 |
| 2020 | Deep 3D Pan via local adaptive "t-shaped" convolutions with global and local adaptive dilations
Juan Luis Gonzalez 0001, Munchurl Kim |
ICLR | 1 |
| 2020 | Forget About the LiDAR: Self-Supervised Depth Estimators with MED Probability VolumesabstractSelf-supervised depth estimators have recently shown results comparable to the supervised methods on the challenging single image depth estimation (SIDE) task, by exploiting the geometrical relations between target and reference views in the training data. However, previous methods usually learn forward or backward image synthesis, but not depth estimation, as they cannot effectively neglect occlusions between the target and the reference images. Previous works rely on rigid photometric assumptions or on the SIDE network to infer depth and occlusions, resulting in limited performance. On the other hand, we propose a method to "Forget About the LiDAR" (FAL), with Mirrored Exponential Disparity (MED) probability volumes for the training of monocular depth estimators from stereo images. Our MED representation allows us to obtain geometrically inspired occlusion maps with our novel Mirrored Occlusion Module (MOM), which does not impose a learning burden on our FAL-net. Contrary to the previous methods that learn SIDE from stereo pairs by regressing disparity in the linear space, our FAL-net regresses disparity by binning it into the exponential space, which allows for better detection of distant and nearby objects. We define a two-step training strategy for our FAL-net: It is first trained for view synthesis and then fine-tuned for depth estimation with our MOM. Our FAL-net is remarkably light-weight and outperforms the previous state-of-the-art methods with 8$\times$ fewer parameters and 3$\times$ faster inference speeds on the challenging KITTI dataset. We present extensive experimental results on the KITTI, CityScapes, and Make3D datasets to verify our method's effectiveness. To the authors' best knowledge, the presented method performs the best among all the previous self-supervised methods until now. Juan Luis Gonzalez 0001, Munchurl Kim |
NeurIPS | 1 |
| 2019 | A Novel Monocular Disparity Estimation Network with Domain Transformation and Ambiguity LearningabstractConvolutional neural networks (CNN) have shown state-of-the-art results for low-level computer vision problems such as stereo and monocular disparity estimations, but still, have much room to further improve their performance in terms of accuracy, numbers of parameters, etc. Recent works have uncovered the advantages of using an unsupervised scheme to train CNN's to estimate monocular disparity, where only the relatively-easy-to-obtain stereo images are needed for training. We propose a novel encoder-decoder architecture that outperforms previous unsupervised monocular depth estimation networks by (i) taking into account ambiguities, (ii) efficient fusion between encoder and decoder features with rectangular convolutions and (iii) domain transformations between encoder and decoder. Our architecture outperforms the Monodepth baseline in all metrics, even with a considerable reduction of parameters. Furthermore, our architecture is capable of estimating a full disparity map in a single forward pass, whereas the baseline requires two passes. We perform extensive experiments to verify the effectiveness of our method on the KITTI dataset. Juan Luis Gonzalez 0001, Munchurl Kim |
ICIP | 1 |