VLDB 2026 Research / reviewers in the wild / expert
Yu-Lun Liu 0001
dblp:142/0282-1 · also Yu-Lun (Alex) Liu
· DBLP profile ↗
42ranked-venue papers
6as first author
35since 2021 · last 2026
0000-0002-7561-6884ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 6 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 5 first-author · 25 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Splannequin: Freezing Monocular Mannequin-Challenge Footage with Dual-Detection SplattingabstractSynthesizing high-fidelity frozen 3D scenes from monocular Mannequin-Challenge (MC) videos is a unique problem distinct from standard dynamic scene reconstruction. Instead of focusing on modeling motion, our goal is to create a frozen scene while strategically preserving subtle dynamics to enable user-controlled instant selection. To achieve this, we introduce a novel application of dynamic Gaussian splatting: the scene is modeled dynamically, which retains nearby temporal variation, and a static scene is rendered by fixing the model’s time parameter. However, under this usage, monocular capture with sparse temporal supervision introduces artifacts like ghosting and blur for Gaussians that become unobserved or occluded at weakly supervised timestamps. We propose Splannequin, an architecture-agnostic regularization that detects two states of Gaussian primitives, hidden and defective, and applies temporal anchoring. Under predominantly forward camera motion, hidden states are anchored to their well-observed past states, while defective states are anchored to future states with stronger supervision. Our method integrates into existing dynamic Gaussian pipelines via simple loss terms, requires no architectural changes, and adds zero inference overhead. This results in markedly improved visual quality, enabling high-fidelity, user-selectable frozen-time renderings, validated by a 96% user preference. Project page: https://chien90190.github.io/splannequin/ Hao-Jen Chien, Yi-Chuan Huang, Chung-Ho Wu, Wei-Lun Chao, Yu-Lun Liu 0001 |
WACV | 5 |
| 2026 | TED-4DGS: Temporally Activated and Embedding-based Deformation for 4DGS CompressionabstractBuilding on the success of 3D Gaussian Splatting (3DGS) in static 3D scene representation, its extension to dynamic scenes-commonly referred to as 4DGS or dynamic 3DGS- has attracted increasing attention. However, designing more compact, efficient deformation schemes together with rate-distortion-optimized compression strategies for dynamic 3DGS representations remains an underexplored area. Prior methods either rely on space-time 4DGS with overspecified, short-lived Gaussian primitives or on canonical 3DGS with deformation that lacks explicit temporal control. To address this, we present TED-4DGS, a temporally activated and embedding-based deformation scheme for rate-distortion- optimized 4DGS compression that unifies the strengths of both families. TED-4DGS is built on a sparse anchor-based 3DGS representation. Each canonical anchor is assigned with learnable temporal-activation parameters to specify its appearance and disappearance transitions over time, while a lightweight per-anchor temporal embedding queries a shared deformation bank to produce anchor-specific deformation. For rate-distortion compression, we incorporate an implicit neural representation (INR)-based hyperprior to model anchor attribute distributions, along with a channelwise autoregressive model to capture intra-anchor correlations. With these novel elements, our scheme achieves the state-of-the-art rate-distortion performance on several commonly used real-world datasets. To the best of our knowledge, this work represents one of the first attempts to pursue a rate-distortion-optimized compression framework for dynamic 3DGS representations. Cheng-Yuan Ho, Hebi Yang, Jui-Chiu Chiang, Yu-Lun Liu 0001, Wen-Hsiao Peng |
WACV | 4 |
| 2026 | PS3: Part level instance segmentation in 3DabstractOpen-vocabulary 3D segmentation allows exploration of 3D environments using unrestricted natural language queries. Current approaches to open-vocabulary 3D instance segmentation largely concentrate on recognizing object-level instances but face difficulties when dealing with more fine-grained elements of a scene, such as object parts. Some previous work constructs hierarchical open-vocabulary 3D scene representations by geometric over-segmentation, which can’t identify parts with similar geometry. In this work, we introduce PS3, an approach to generate 3D part proposals from multi-view 2D masks. PS3 outperforms baselines that rely on geometric over-segmentation in scene-scale open-vocabulary 3D part segmentation. Hong-Xuan Yen, Chiamin Chen, Yu-Lun Liu 0001, Min Sun 0001 |
WACV | 4 |
| 2025 | GCC: Generative Color Constancy via Diffusing a Color CheckerabstractColor constancy methods often struggle to generalize across different camera sensors due to varying spectral sensitivities. We present GCC, which leverages diffusion models to inpaint color checkers into images for illumination estimation. Our key innovations include (1) a single-step deterministic inference approach that inpaints color checkers reflecting scene illumination, (2) a Laplacian decomposition technique that preserves checker structure while allowing illumination-dependent color adaptation, and (3) a mask-based data augmentation strategy for handling imprecise color checker annotations. By harnessing rich priors from pre-trained diffusion models, GCC demonstrates strong robustness in challenging cross-camera scenarios. These results highlight our method’s effective generalization capability across different camera characteristics without requiring sensor-specific training, making it a versatile and practical solution for real-world applications. Chen-Wei Chang, Cheng-De Fan, Chia-Che Chang, Yi-Chen Lo, Yu-Chee Tseng, Jiun-Long Huang, Yu-Lun Liu 0001 |
CVPR | 7 |
| 2025 | SpectroMotion: Dynamic 3D Reconstruction of Specular ScenesabstractWe present SpectroMotion, a novel approach that combines 3D Gaussian Splatting (3DGS) with physically-based rendering (PBR) and deformation fields to reconstruct dynamic specular scenes. Previous methods extending 3DGS to model dynamic scenes have struggled to represent specular surfaces accurately. Our method addresses this limitation by introducing a residual correction technique for accurate surface normal computation during deformation, complemented by a deformable environment map that adapts to time-varying lighting conditions. We implement a coarse-to-fine training strategy significantly enhancing scene geometry and specular color prediction. It is the only existing 3DGS method capable of synthesizing photorealistic real-world dynamic specular scenes, outperforming state-of-the-art methods in rendering complex, dynamic, and specular scenes. Please see our project page at cdfan0627.github.io/spectromotion. Cheng-De Fan, Chen-Wei Chang, Yi-Ruei Liu, Jie-Ying Lee, Jiun-Long Huang, Yu-Chee Tseng, Yu-Lun Liu 0001 |
CVPR | 7 |
| 2025 | FrugalNeRF: Fast Convergence for Extreme Few-shot Novel View Synthesis without Learned PriorsabstractNeural Radiance Fields (NeRF) face significant challenges in extreme few-shot scenarios, primarily due to overfitting and long training times. Existing methods, such as FreeNeRF and SparseNeRF, use frequency regularization or pre-trained priors but struggle with complex scheduling and bias. We introduce FrugalNeRF, a novel few-shot NeRF framework that leverages weight-sharing voxels across multiple scales to efficiently represent scene details. Our key contribution is a cross-scale geometric adaptation scheme that selects pseudo ground truth depth based on reprojection errors across scales. This guides training without relying on externally learned priors, enabling full utilization of the training data. It can also integrate pre-trained priors, enhancing quality without slowing convergence. Experiments on LLFF, DTU, and RealEstate-10K show that FrugalNeRF outperforms other few-shot NeRF methods while significantly reducing training time, making it a practical solution for efficient and accurate 3D scene reconstruction. Chin-Yang Lin, Chung-Ho Wu, Changhan Yeh, Shih-Han Yen, Cheng Sun 0004, Yu-Lun Liu 0001 |
CVPR | 6 |
| 2025 | AuraFusion360: Augmented Unseen Region Alignment for Reference-based 360deg Unbounded Scene InpaintingabstractThree-dimensional scene inpainting is crucial for applications from virtual reality to architectural visualization, yet existing methods struggle with view consistency and geometric accuracy in 360° unbounded scenes. We present AuraFusion360, a novel reference-based method that enables high-quality object removal and hole filling in 3D scenes represented by Gaussian Splatting. Our approach introduces (1) depth-aware unseen mask generation for accurate occlusion identification, (2) Adaptive Guided Depth Diffusion, a zero-shot method for accurate initial point placement without requiring additional training, and (3) SDEdit-based detail enhancement for multi-view coherence. We also introduce 360-USID, the first comprehensive dataset for 360° unbounded scene inpainting with ground truth. Extensive experiments demonstrate that AuraFusion360 significantly outperforms existing methods, achieving superior perceptual quality while maintaining geometric accuracy across dramatic viewpoint changes. Chung-Ho Wu, Yang-Jung Chen, Ying-Huan Chen, Jie-Ying Lee, Bo-Hsu Ke, Chun-Wei Tuan Mu, Yi-Chuan Huang, Chin-Yang Lin, Min-Hung Chen, Yen-Yu Lin, Yu-Lun Liu 0001 |
CVPR | 11 |
| 2025 | DeNVeR: Deformable Neural Vessel Representations for Unsupervised Video Vessel SegmentationabstractThis paper presents Deformable Neural Vessel Representations (DeNVeR), an unsupervised approach for vessel segmentation in X-ray angiography videos without annotated ground truth. DeNVeR utilizes optical flow and layer separation techniques, enhancing segmentation accuracy and adaptability through test-time training. Key contributions include a novel layer separation bootstrapping technique, a parallel vessel motion loss, and the integration of Eulerian motion fields for modeling complex vessel dynamics. A significant component of this research is the introduction of the XACV dataset, the first X-ray angiography coronary video dataset with high-quality, manually labeled segmentation ground truth. Extensive evaluations on both XACV and CADICA datasets demonstrate that DeNVeR outperforms current state-of-the-art methods in vessel segmentation accuracy and generalization capability while maintaining temporal coherency. Please see our project page at kirito878.github.io/DeNVeR. Chun-Hung Wu, Shih-Hong Chen, Chih-Yao Hu, Hsin-Yu Wu, Kai-Hsin Chen, Yu-You Chen, Chih-Hai Su, Chih-Kuo Lee, Yu-Lun Liu 0001 |
CVPR | 9 |
| 2025 | 3D Gaussian Splatting with Grouped Uncertainty for Unconstrained Imagesabstract3D Gaussian Splatting (3DGS) [1] is a promising method for 3D reconstruction and novel view synthesis. However, training it with unconstrained images presents challenges due to transient objects that cause undesired floaters and ghosting artifacts. Although related works using Neural Radiance Fields (NeRF) [2] have attempted to address these issues, those techniques proved ineffective when directly applied to 3DGS. In this study, we propose an uncertainty estimation approach to assist 3DGS in reconstructing 3D scenes and removing transients. Specifically, we introduce a grouped uncertainty map using the Segment Anything Model [3] for per-area uncertainty estimation, combined with appearance embeddings to handle diverse lighting conditions. Experimental results on tourism photo collections [4] demonstrate that our method improves transient separation and rendering clarity. Furthermore, it facilitates effective color training and enables 3DGS to reconstruct target scenes from unconstrained images with fewer floaters or artifacts. Hao-Yu Hou, Chia-Chi Hsu, Yu-Chen Huang, Mu-Yi Shen, Wei-Fang Sun, Cheng Sun 0004, Chia-Che Chang, Yu-Lun Liu 0001, Chun-Yi Lee |
ICASSP | 8 |
| 2025 | OpenM3D: Open Vocabulary Multi-View Indoor 3D Object Detection without Human Annotations
Peng-Hao Hsu, Ke Zhang 0028, Fu-En Wang, Tao Tu 0002, Ming-Feng Li, Yu-Lun Liu 0001, Albert Chen 0001, Min Sun 0001, Cheng-Hao Kuo |
ICCV | 6 |
| 2025 | StealthAttack: Robust 3D Gaussian Splatting Poisoning via Density-Guided Illusionsabstract3D scene representation methods like Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have significantly advanced novel view synthesis. As these methods become prevalent, addressing their vulnerabilities becomes critical. We analyze 3DGS robustness against image-level poisoning attacks and propose a novel density-guided poisoning method. Our method strategically injects Gaussian points into low-density regions identified via Kernel Density Estimation (KDE), embedding viewpoint-dependent illusory objects clearly visible from poisoned views while minimally affecting innocent views. Additionally, we introduce an adaptive noise strategy to disrupt multi-view consistency, further enhancing attack effectiveness. We propose a KDE-based evaluation protocol to assess attack difficulty systematically, enabling objective benchmarking for future research. Extensive experiments demonstrate our method's superior performance compared to state-of-the-art techniques. Project page: https://hentci.github.io/stealthattack/ Bo-Hsu Ke, You-Zhe Xie, Yu-Lun Liu 0001, Walon Wei-Chen Chiu |
ICCV | 3 |
| 2025 | LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long VideosabstractLongSplat addresses critical challenges in novel view synthesis (NVS) from casually captured long videos characterized by irregular camera motion, unknown camera poses, and expansive scenes. Current methods often suffer from pose drift, inaccurate geometry initialization, and severe memory limitations. To address these issues, we introduce LongSplat, a robust unposed 3D Gaussian Splatting framework featuring: (1) Incremental Joint Optimization that concurrently optimizes camera poses and 3D Gaussians to avoid local minima and ensure global consistency; (2) a robust Pose Estimation Module leveraging learned 3D priors; and (3) an efficient Octree Anchor Formation mechanism that converts dense point clouds into anchors based on spatial density. Extensive experiments on challenging benchmarks demonstrate that LongSplat achieves state-of-the-art results, substantially improving rendering quality, pose accuracy, and computational efficiency compared to prior approaches. Project page: https://linjohnss.github.io/longsplat/ Chin-Yang Lin, Cheng Sun 0004, Fu-En Yang, Min-Hung Chen, Yen-Yu Lin, Yu-Lun Liu 0001 |
ICCV | 6 |
| 2025 | CAT-3DGS: A Context-Adaptive Triplane Approach to Rate-Distortion-Optimized 3DGS Compressionabstract3D Gaussian Splatting (3DGS) has recently emerged as a promising 3D representation. Much research has been focused on reducing its storage requirements and memory footprint. However, the needs to compress and transmit the 3DGS representation to the remote side are overlooked. This new application calls for rate-distortion-optimized 3DGS compression. How to quantize and entropy encode sparse Gaussian primitives in the 3D space remains largely unexplored. Few early attempts resort to the hyperprior framework from learned image compression. But, they fail to utilize fully the inter and intra correlation inherent in Gaussian primitives. Built on ScaffoldGS, this work, termed CAT-3DGS, introduces a context-adaptive triplane approach to their rate-distortion-optimized coding. It features multi-scale triplanes, oriented according to the principal axes of Gaussian primitives in the 3D space, to capture their inter correlation (i.e. spatial correlation) for spatial autoregressive coding in the projected 2D planes. With these triplanes serving as the hyperprior, we further perform channel-wise autoregressive coding to leverage the intra correlation within each individual Gaussian primitive. Our CAT-3DGS incorporates a view frequency-aware masking mechanism. It actively skips from coding those Gaussian primitives that potentially have little impact on the rendering quality. When trained end-to-end to strike a good rate-distortion trade-off, our CAT-3DGS achieves the state-of-the-art compression performance on the commonly used real-world datasets. Yu-Ting Zhan, Cheng-Yuan Ho, Hebi Yang, Yi-Hsin Chen, Jui-Chiu Chiang, Yu-Lun Liu 0001, Wen-Hsiao Peng |
ICLR | 6 |
| 2025 | FIPER: Factorized Features for Robust Image Super-Resolution and CompressionabstractIn this work, we propose using a unified representation, termed **Factorized Features**, for low-level vision tasks, where we test on **Single Image Super-Resolution (SISR)** and **Image Compression**. Motivated by the shared principles between these tasks, they require recovering and preserving fine image details, whether by enhancing resolution for SISR or reconstructing compressed data for Image Compression. Unlike previous methods that mainly focus on network architecture, our proposed approach utilizes a basis-coefficient decomposition as well as an explicit formulation of frequencies to capture structural components and multi-scale visual features in images, which addresses the core challenges of both tasks.
We replace the representation of prior models from simple feature maps with Factorized Features to validate the potential for broad generalizability.
In addition, we further optimize the compression pipeline by leveraging the mergeable-basis property of our Factorized Features, which consolidates shared structures on multi-frame compression.
Extensive experiments show that our unified representation delivers state-of-the-art performance, achieving an average relative improvement of 204.4\% in PSNR over the baseline in Super-Resolution (SR) and 9.35\% BD-rate reduction in Image Compression compared to the previous SOTA. Yang-Che Sun, Cheng Yu Yeo, Ernie Chu, Jun-Cheng Chen, Yu-Lun Liu 0001 |
NeurIPS | 5 |
| 2025 | Prior-Enhanced Gaussian Splatting for Dynamic Scene Reconstruction from Casual VideoabstractWe introduce a fully automatic pipeline for dynamic scene reconstruction from casually captured monocular RGB videos. Rather than designing a new scene representation, we enhance the priors that drive Dynamic Gaussian Splatting. Video segmentation combined with epipolar-error maps yields object-level masks that closely follow thin structures; these masks (i) guide an object-depth loss that sharpens the consistent video depth, and (ii) support skeleton-based sampling plus mask-guided re-identification to produce reliable, comprehensive 2-D tracks. Two additional objectives embed the refined priors in the reconstruction stage: a virtual-view depth loss removes floaters, and a scaffold-projection loss ties motion nodes to the tracks, preserving fine geometry and coherent motion. The resulting system surpasses previous monocular dynamic scene reconstruction methods and delivers visibly superior renderings. Project page: https://priorenhancedgaussian.github.io/ Meng-Li Shih, Ying-Huan Chen, Yu-Lun Liu 0001, Brian Curless |
SIGGRAPH Asia | 3 |
| 2025 | ORFormer: Occlusion-Robust Transformer for Accurate Facial Landmark DetectionabstractAlthough facial landmark detection (FLD) has gained significant progress, existing FLD methods still suffer from performance drops on partially non-visible faces, such as faces with occlusions or under extreme lighting conditions or poses. To address this issue, we introduce ORFormer, a novel transformer-based method that can detect non-visible regions and recover their missing features from vis-ible parts. Specifically, ORFormer associates each image patch token with one additional learnable token called the messenger token. The messenger token aggregates features from all but its patch. This way, the consensus between a patch and other patches can be assessed by referring to the similarity between its regular and messenger embeddings, enabling non-visible region identification. Our method then recovers occluded patches with features aggregated by the messenger tokens. Leveraging the recovered features, OR-Former compiles high-quality heatmaps for the downstream FLD task. Extensive experiments show that our method generates heatmaps resilient to partial occlusions. By inte-grating the resultant heatmaps into existing FLD methods, our method performs favorably against the state of the arts on challenging datasets such as WFLWand COFW. Jui-Che Chiang, Hou-Ning Hu, Bo-Syuan Hou, Chia-Yu Tseng, Yu-Lun Liu 0001, Min-Hung Chen, Yen-Yu Lin |
WACV | 5 |
| 2025 | CorrFill: Enhancing Faithfulness in Reference-Based Inpainting with Correspondence Guidance in Diffusion ModelsabstractIn the task of reference-based image inpainting, an additional reference image is provided to restore a damaged target image to its original state. The advancement of diffusion models, particularly Stable Diffusion, allows for simple formulations in this task. However, existing diffusion-based methods often lack explicit constraints on the correlation between the reference and damaged images, resulting in lower faithfulness to the reference images in the inpainting results. In this work, we propose CorrFill, a training-free module designed to enhance the awareness of geometric correlations between the reference and target images. This enhancement is achieved by guiding the inpainting process with correspondence constraints estimated during inpainting, utilizing attention masking in self-attention layers and an objective function to update the input tensor according to the constraints. Experimental results demonstrate that CorrFill significantly enhances the performance of multiple baseline diffusion-based methods, including state-of-the-art approaches, by emphasizing faithfulness to the reference images. Kuan-Hung Liu, Cheng-Kun Yang, Min-Hung Chen, Yu-Lun Liu 0001, Yen-Yu Lin |
WACV | 4 |
| 2024 | Improving Robustness for Joint Optimization of Camera Pose and Decomposed Low-Rank Tensorial Radiance FieldsabstractIn this paper, we propose an algorithm that allows joint refinement of camera pose and scene geometry represented by decomposed low-rank tensor, using only 2D images as supervision. First, we conduct a pilot study based on a 1D signal and relate our findings to 3D scenarios, where the naive joint pose optimization on voxel-based NeRFs can easily lead to sub-optimal solutions. Moreover, based on the analysis of the frequency spectrum, we propose to apply convolutional Gaussian filters on 2D and 3D radiance fields for a coarse-to-fine training schedule that enables joint camera pose optimization. Leveraging the decomposition property in decomposed low-rank tensor, our method achieves an equivalent effect to brute-force 3D convolution with only incurring little computational overhead. To further improve the robustness and stability of joint optimization, we also propose techniques of smoothed 2D supervision, randomly scaled kernel parameters, and edge-guided loss mask. Extensive quantitative and qualitative evaluations demonstrate that our proposed framework achieves superior performance in novel view synthesis as well as rapid convergence for optimization. The source code is available at https://github.com/Nemo1999/Joint-TensoRF. Bo-Yu Chen, Walon Wei-Chen Chiu, Yu-Lun Liu 0001 |
AAAI | 3 |
| 2024 | HumanNeRF-SE: A Simple yet Effective Approach to Animate HumanNeRF with Diverse PosesabstractWe present HumanNeRF-SE, a simple yet effective method that synthesizes diverse novel pose images with sim-ple input. Previous HumanNeRF works require a large number of optimizable parameters to fit the human images. Instead, we reload these approaches by combining explicit and implicit human representations to design both general-ized rigid deformation and specific non-rigid deformation. Our key insight is that explicit shape can reduce the sam-pling points used to fit implicit representation, and frozen blending weights from SMPL constructing a generalized rigid deformation can effectively avoid overfitting and im-prove pose generalization performance. Our architecture involving both explicit and implicit representation is sim-ple yet effective. Experiments demonstrate our model can synthesize images under arbitrary poses with few-shot input and increase the speed of synthesizing images by 15 times through a reduction in computational complexity without using any existing acceleration modules. Compared to the state-of-the-art HumanNeRF studies, HumanNeRF-SE achieves better performance with fewer learnable parame-ters and less training time. Caoyuan Ma, Yu-Lun Liu 0001, Zhixiang Wang 0001, Wu Liu 0005, Xinchen Liu, Zheng Wang 0007 |
CVPR | 2 |
| 2024 | Image-Text Co-Decomposition for Text-Supervised Semantic SegmentationabstractThis paper addresses text-supervised semantic segmentation, aiming to learn a model capable of segmenting arbitrary visual concepts within images by using only image-text pairs without dense annotations. Existing methods have demonstrated that contrastive learning on image-text pairs effectively aligns visual segments with the meanings of texts. We notice that there is a discrepancy between text alignment and semantic segmentation: A text often consists of multiple semantic concepts, whereas semantic segmentation strives to create semantically homogeneous segments. To address this issue, we propose a novel framework, Image-Text Co-Decomposition (CoDe), where the paired image and text are jointly decomposed into a set of image regions and a set of word segments, respectively, and contrastive learning is developed to enforce region-word alignment. To work with a vision-language model, we present a prompt learning mechanism that derives an extra representation to highlight an image segment or a word segment of interest, with which more effective features can be extracted from that segment. Comprehensive experimental results demonstrate that our method performs favorably against existing text-supervised semantic segmentation methods on six benchmark datasets. The code is available at https://github.com/072jiajia/image-text-co-decomposition. Ji-Jia Wu, Andy Chia-Hao Chang, Chieh-Yu Chuang, Chun-Pei Chen, Yu-Lun Liu 0001, Min-Hung Chen, Hou-Ning Hu, Yung-Yu Chuang, Yen-Yu Lin |
CVPR | 5 |
| 2024 | GenRC: Generative 3D Room Completion from Sparse Image Collections
Ming-Feng Li, Yueh-Feng Ku, Hong-Xuan Yen, Yu-Lun Liu 0001, Albert Chen 0001, Cheng-Hao Kuo, Min Sun 0001 |
ECCV (37) | 5 |
| 2024 | Dual Associated Encoder for Face RestorationabstractRestoring facial details from low-quality (LQ) images has remained challenging due to the nature of the problem caused by various degradations in the wild.
The codebook prior has been proposed to address the ill-posed problems by leveraging an autoencoder and learned codebook of high-quality (HQ) features, achieving remarkable quality.
However, existing approaches in this paradigm frequently depend on a single encoder pre-trained on HQ data for restoring HQ images, disregarding the domain gap and distinct feature representations between LQ and HQ images.
As a result, encoding LQ inputs with the same encoder could be insufficient, resulting in imprecise feature representation and leading to suboptimal performance.
To tackle this problem, we propose a novel dual-branch framework named $\textit{DAEFR}$. Our method introduces an auxiliary LQ branch that extracts domain-specific information from the LQ inputs.
Additionally, we incorporate association training to promote effective synergy between the two branches, enhancing code prediction and restoration quality.
We evaluate the effectiveness of DAEFR on both synthetic and real-world datasets, demonstrating its superior performance in restoring facial details.
Project page: https://liagm.github.io/DAEFR/ Yu-Ju Tsai, Yu-Lun Liu 0001, Lu Qi 0001, Kelvin C. K. Chan, Ming-Hsuan Yang 0001 |
ICLR | 2 |
| 2024 | Precise Pick-and-Place using Score-Based Diffusion NetworksabstractIn this paper, we propose a novel coarse-to-fine continuous pose diffusion method to enhance the precision of pick-and-place operations within robotic manipulation tasks. Leveraging the capabilities of diffusion networks, we facilitate the accurate perception of object poses. This accurate perception enhances both pick-and-place success rates and overall manipulation precision. Our methodology utilizes a top-down RGB image projected from an RGB-D camera and adopts a coarse-to-fine architecture. This architecture enables efficient learning of coarse and fine models. A distinguishing feature of our approach is its focus on continuous pose estimation, which enables more precise object manipulation, particularly concerning rotational angles. In addition, we employ pose and color augmentation techniques to enable effective training with limited data. Through extensive experiments in simulated and real-world scenarios, as well as an ablation study, we comprehensively evaluate our proposed methodology. Taken together, the findings validate its effectiveness in achieving high-precision pick-and-place tasks. Shih-Wei Guo, Tsu-Ching Hsiao, Yu-Lun Liu 0001, Chun-Yi Lee |
IROS | 3 |
| 2024 | NaRCan: Natural Refined Canonical Image with Integration of Diffusion Prior for Video EditingabstractWe propose a video editing framework, NaRCan, which integrates a hybrid deformation field and diffusion prior to generate high-quality natural canonical images to represent the input video. Our approach utilizes homography to model global motion and employs multi-layer perceptrons (MLPs) to capture local residual deformations, enhancing the model’s ability to handle complex video dynamics. By introducing a diffusion prior from the early stages of training, our model ensures that the generated images retain a high-quality natural appearance, making the produced canonical images suitable for various downstream tasks in video editing, a capability not achieved by current canonical-based methods. Furthermore, we incorporate low-rank adaptation (LoRA) fine-tuning and introduce a noise and diffusion prior update scheduling technique that accelerates the training process by 14 times. Extensive experimental results show that our method outperforms existing approaches in various video editing tasks and produces coherent and high-quality edited video sequences. See our project page for video results: [koi953215.github.io/NaRCan_page](https://koi953215.github.io/NaRCan_page/). Ting-Hsuan Chen, Jiewen Chan, Hau-Shiang Shiu, Shih-Han Yen, Changhan Yeh, Yu-Lun Liu 0001 |
NeurIPS | 6 |
| 2024 | ReF-LDM: A Latent Diffusion Model for Reference-based Face Image RestorationabstractWhile recent works on blind face image restoration have successfully produced impressive high-quality (HQ) images with abundant details from low-quality (LQ) input images, the generated content may not accurately reflect the real appearance of a person. To address this problem, incorporating well-shot personal images as additional reference inputs may be a promising strategy. Inspired by the recent success of the Latent Diffusion Model (LDM) in image generation, we propose ReF-LDM—an adaptation of LDM designed to generate HQ face images conditioned on one LQ image and multiple HQ reference images. Our LDM-based model incorporates an effective and efficient mechanism, CacheKV, for conditioning on reference images. Additionally, we design a timestep-scaled identity loss, enabling LDM to focus on learning the discriminating features of human faces. Lastly, we construct FFHQ-ref, a dataset consisting of 20,406 high-quality (HQ) face images with corresponding reference images, which can serve as both training and evaluation data for reference-based face restoration models. Chi-Wei Hsiao, Yu-Lun Liu 0001, Cheng-Kun Yang, Sheng-Po Kuo, Kevin Jou, Chia-Ping Chen |
NeurIPS | 2 |
| 2024 | Depth Anywhere: Enhancing 360 Monocular Depth Estimation via Perspective Distillation and Unlabeled Data AugmentationabstractAccurately estimating depth in 360-degree imagery is crucial for virtual reality, autonomous navigation, and immersive media applications. Existing depth estimation methods designed for perspective-view imagery fail when applied to 360-degree images due to different camera projections and distortions. We propose a new depth estimation framework that uses unlabeled 360-degree data effectively. Our approach uses state-of-the-art perspective depth estimation models as teacher models to generate pseudo labels through a six-face cube projection technique, enabling efficient labeling of depth in 360-degree images. This method leverages the increasing availability of large datasets. It includes two main stages: offline mask generation for invalid regions and an online semi-supervised joint training regime. We tested our approach on benchmark datasets such as Matterport3D and Stanford2D3D, showing significant improvements in depth estimation accuracy, particularly in zero-shot scenarios. Our proposed training pipeline can enhance any 360 monocular depth estimator and demonstrate effective knowledge transfer across different camera projections and data types. Ning-Hsu Wang, Yu-Lun Liu 0001 |
NeurIPS | 2 |
| 2024 | DisCO: Portrait Distortion Correction with Perspective-Aware 3D GANs
Zhixiang Wang 0001, Yu-Lun Liu 0001, Jia-Bin Huang 0001, Shin'ichi Satoh 0001, Sizhuo Ma, Gurunandan Krishnan, Jian Wang 0100 |
Int. J. Comput. Vis. | 2 |
| 2023 | Robust Dynamic Radiance FieldsabstractDynamic radiance field reconstruction methods aim to model the time-varying structure and appearance of a dynamic scene. Existing methods, however, assume that accurate camera poses can be reliably estimated by Structure from Motion (SfM) algorithms. These methods, thus, are unreliable as SfM algorithms often fail or produce erroneous poses on challenging videos with highly dynamic objects, poorly textured surfaces, and rotating camera motion. We address this robustness issue by jointly estimating the static and dynamic radiance fields along with the camera parameters (poses and focal length). We demonstrate the robustness of our approach via extensive quantitative and qualitative experiments. Our results show favorable performance over the state-of-the-art dynamic view synthesis methods. Yu-Lun Liu 0001, Chen Gao 0003, Andreas Meuleman, Hung-Yu Tseng, Ayush Saraf, Changil Kim 0001, Yung-Yu Chuang, Johannes Kopf 0001, Jia-Bin Huang 0001 |
CVPR | 1 |
| 2023 | Progressively Optimized Local Radiance Fields for Robust View SynthesisabstractWe present an algorithm for reconstructing the radiance field of a large-scale scene from a single casually captured video. The task poses two core challenges. First, most existing radiance field reconstruction approaches rely on accurate pre-estimated camera poses from Structure-from-Motion algorithms, which frequently fail on in-the-wild videos. Second, using a single, global radiance field with finite representational capacity does not scale to longer trajectories in an unbounded scene. For handling unknown poses, we jointly estimate the camera poses with radiance field in a progressive manner. We show that progressive optimization significantly improves the robustness of the reconstruction. For handling large unbounded scenes, we dynamically allocate new local radiance fields trained with frames within a temporal window. This further improves robustness (e.g., performs well even under moderate pose drifts) and allows us to scale to large scenes. Our extensive evaluation on the TANKS AND TEMPLES dataset and our collected outdoor dataset, STATIC HIKES, show that our approach compares favorably with the state-of-the-art. Andreas Meuleman, Yu-Lun Liu 0001, Chen Gao 0003, Jia-Bin Huang 0001, Changil Kim 0001, Min H. Kim 0001, Johannes Kopf 0001 |
CVPR | 2 |
| 2023 | Learning Continuous Exposure Value Representations for Single-Image HDR ReconstructionabstractDeep learning is commonly used to reconstruct HDR images from LDR images. LDR stack-based methods are used for single-image HDR reconstruction, generating an HDR image from a deep learning-generated LDR stack. However, current methods generate the stack with predetermined exposure values (EVs), which may limit the quality of HDR reconstruction. To address this, we propose the continuous exposure value representation (CEVR), which uses an implicit function to generate LDR images with arbitrary EVs, including those unseen during training. Our approach generates a continuous stack with more images containing diverse EVs, significantly improving HDR reconstruction. We use a cycle training strategy to supervise the model in generating continuous EV LDR images without corresponding ground truths. Our CEVR model outperforms existing methods, as demonstrated by experimental results. Su-Kai Chen, Hung-Lin Yen, Yu-Lun Liu 0001, Min-Hung Chen, Hou-Ning Hu, Wen-Hsiao Peng, Yen-Yu Lin |
ICCV | 3 |
| 2023 | ImGeoNet: Image-induced Geometry-aware Voxel Representation for Multi-view 3D Object DetectionabstractWe propose ImGeoNet, a multi-view image-based 3D object detection framework that models a 3D space by an image-induced geometry-aware voxel representation. Unlike previous methods which aggregate 2D features into 3D voxels without considering geometry, ImGeoNet learns to induce geometry from multi-view images to alleviate the confusion arising from voxels of free space, and during the inference phase, only images from multiple views are required. Besides, a powerful pre-trained 2D feature extractor can be leveraged by our representation, leading to a more robust performance. To evaluate the effectiveness of ImGeoNet, we conduct quantitative and qualitative experiments on three indoor datasets, namely ARKitScenes, ScanNetV2, and ScanNet200. The results demonstrate that ImGeoNet outperforms the current state-of-the-art multiview image-based method, ImVoxelNet, on all three datasets in terms of detection accuracy. In addition, ImGeoNet shows great data efficiency by achieving results comparable to ImVoxelNet with 100 views while utilizing only 40 views. Furthermore, our studies indicate that our proposed image-induced geometry-aware representation can enable image-based methods to attain superior detection accuracy than the seminal point cloud-based method, VoteNet, in two practical scenarios: (1) scenarios where point clouds are sparse and noisy, such as in ARKitScenes, and (2) scenarios involve diverse object classes, particularly classes of small objects, as in the case in ScanNet200. Project page: https://ttaoretw.github.io/imgeonet. Tao Tu 0002, Shun-Po Chuang, Yu-Lun Liu 0001, Cheng Sun 0004, Ke Zhang 0028, Donna Roy, Cheng-Hao Kuo, Min Sun 0001 |
ICCV | 3 |
| 2022 | Denoising Likelihood Score Matching for Conditional Score-based Data Generation
Chen-Hao Chao, Wei-Fang Sun, Bo-Wun Cheng, Yi-Chen Lo, Chia-Che Chang, Yu-Lun Liu 0001, Yu-Lin Chang, Chia-Ping Chen, Chun-Yi Lee |
ICLR | 6 |
| 2022 | Learning to See Through Obstructions With Layered DecompositionabstractWe present a learning-based approach for removing unwanted obstructions, such as window reflections, fence occlusions, or adherent raindrops, from a short sequence of images captured by a moving camera. Our method leverages motion differences between the background and obstructing elements to recover both layers. Specifically, we alternate between estimating dense optical flow fields of the two layers and reconstructing each layer from the flow-warped images via a deep convolutional neural network. This learning-based layer reconstruction module facilitates accommodating potential errors in the flow estimation and brittle assumptions, such as brightness consistency. We show that the proposed approach learned from synthetically generated data performs well to real images. Experimental results on numerous challenging scenarios of reflection and fence removal demonstrate the effectiveness of the proposed method. Yu-Lun Liu 0001, Wei-Sheng Lai, Ming-Hsuan Yang 0001, Yung-Yu Chuang, Jia-Bin Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Hybrid Neural Fusion for Full-frame Video StabilizationabstractExisting video stabilization methods often generate visible distortion or require aggressive cropping of frame boundaries, resulting in smaller field of views. In this work, we present a frame synthesis algorithm to achieve full-frame video stabilization. We first estimate dense warp fields from neighboring frames and then synthesize the stabilized frame by fusing the warped contents. Our core technical novelty lies in the learning-based hybrid-space fusion that alleviates artifacts caused by optical flow inaccuracy and fast-moving objects. We validate the effectiveness of our method on the NUS, selfie, and DeepStab video datasets. Extensive experiment results demonstrate the merits of our approach over prior video stabilization methods. Yu-Lun Liu 0001, Wei-Sheng Lai, Ming-Hsuan Yang 0001, Yung-Yu Chuang, Jia-Bin Huang 0001 |
ICCV | 1 |
| 2021 | Bridging Unsupervised and Supervised Depth from Focus via All-in-Focus SupervisionabstractDepth estimation is a long-lasting yet important task in computer vision. Most of the previous works try to estimate depth from input images and assume images are all-in-focus (AiF), which is less common in real-world applications. On the other hand, a few works take defocus blur into account and consider it as another cue for depth estimation. In this paper, we propose a method to estimate not only a depth map but an AiF image from a set of images with different focus positions (known as a focal stack). We design a shared architecture to exploit the relationship between depth and AiF estimation. As a result, the proposed method can be trained either supervisedly with ground truth depth, or unsupervisedly with AiF images as supervisory signals. We show in various experiments that our method outperforms the state-of-the-art methods both quantitatively and qualitatively, and also has higher efficiency in inference time. Ning-Hsu Wang, Ren Wang 0014, Yu-Lun Liu 0001, Yu-Lin Chang, Chia-Ping Chen, Kevin Jou |
ICCV | 3 |
| 2020 | Attention-Based View Selection Networks for Light-Field Disparity EstimationabstractThis paper introduces a novel deep network for estimating depth maps from a light field image. For utilizing the views more effectively and reducing redundancy within views, we propose a view selection module that generates an attention map indicating the importance of each view and its potential for contributing to accurate depth estimation. By exploring the symmetric property of light field views, we enforce symmetry in the attention map and further improve accuracy. With the attention map, our architecture utilizes all views more effectively and efficiently. Experiments show that the proposed method achieves state-of-the-art performance in terms of accuracy and ranks the first on a popular benchmark for disparity estimation for light field images. Yu-Ju Tsai, Yu-Lun Liu 0001, Ouhyoung Ming, Yung-Yu Chuang |
AAAI | 2 |
| 2020 | Learning to See Through ObstructionsabstractWe present a learning-based approach for removing unwanted obstructions, such as window reflections, fence occlusions or raindrops, from a short sequence of images captured by a moving camera. Our method leverages the motion differences between the background and the obstructing elements to recover both layers. Specifically, we alternate between estimating dense optical flow fields of the two layers and reconstructing each layer from the flow-warped images via a deep convolutional neural network. The learning-based layer reconstruction allows us to accommodate potential errors in the flow estimation and brittle assumptions such as brightness consistency. We show that training on synthetically generated data transfers well to real images. Our results on numerous challenging scenarios of reflection and fence removal demonstrate the effectiveness of the proposed method. Yu-Lun Liu 0001, Wei-Sheng Lai, Ming-Hsuan Yang 0001, Yung-Yu Chuang, Jia-Bin Huang 0001 |
CVPR | 1 |
| 2020 | Single-Image HDR Reconstruction by Learning to Reverse the Camera PipelineabstractRecovering a high dynamic range (HDR) image from a single low dynamic range (LDR) input image is challenging due to missing details in under-/over-exposed regions caused by quantization and saturation of camera sensors. In contrast to existing learning-based methods, our core idea is to incorporate the domain knowledge of the LDR image formation pipeline into our model. We model the HDR-to-LDR image formation pipeline as the (1) dynamic range clipping, (2) non-linear mapping from a camera response function, and (3) quantization. We then propose to learn three specialized CNNs to reverse these steps. By decomposing the problem into specific sub-tasks, we impose effective physical constraints to facilitate the training of individual sub-networks. Finally, we jointly fine-tune the entire model end-to-end to reduce error accumulation. With extensive quantitative and qualitative experiments on diverse image datasets, we demonstrate that the proposed method performs favorably against state-of-the-art single-image HDR reconstruction algorithms. Yu-Lun Liu 0001, Wei-Sheng Lai, Yu-Sheng Chen, Yi-Lung Kao, Ming-Hsuan Yang 0001, Yung-Yu Chuang, Jia-Bin Huang 0001 |
CVPR | 1 |
| 2020 | Learning Camera-Aware Noise Models
Ke-Chi Chang, Ren Wang 0014, Hung-Jin Lin, Yu-Lun Liu 0001, Chia-Ping Chen, Yu-Lin Chang, Hwann-Tzong Chen |
ECCV (24) | 4 |
| 2020 | Explorable Tone Mapping OperatorsabstractTone-mapping plays an essential role in high dynamic range (HDR) imaging. It aims to preserve visual information of HDR images in a medium with a limited dynamic range. Although many works have been proposed to provide tone-mapped results from HDR images, most of them can only perform tone-mapping in a single pre-designed way. However, the subjectivity of tone-mapping quality varies from person to person, and the preference of tone-mapping style also differs from application to application. In this paper, a learning-based multimodal tone-mapping method is proposed, which not only achieves excellent visual quality but also explores the style diversity. Based on the framework of BicycleGAN [1], the proposed method can provide a variety of expert-level tone-mapped results by manipulating different latent codes. Finally, we show that the proposed method performs favorably against state-of-the-art tone-mapping algorithms both quantitatively and qualitatively. Chien-Chuan Su, Ren Wang 0014, Hung-Jin Lin, Yu-Lun Liu 0001, Chia-Ping Chen, Yu-Lin Chang, Soo-Chang Pei |
ICPR | 4 |
| 2019 | Deep Video Frame Interpolation Using Cyclic Frame GenerationabstractVideo frame interpolation algorithms predict intermediate frames to produce videos with higher frame rates and smooth view transitions given two consecutive frames as inputs. We propose that: synthesized frames are more reliable if they can be used to reconstruct the input frames with high quality. Based on this idea, we introduce a new loss term, the cycle consistency loss. The cycle consistency loss can better utilize the training data to not only enhance the interpolation results, but also maintain the performance better with less training data. It can be integrated into any frame interpolation network and trained in an end-to-end manner. In addition to the cycle consistency loss, we propose two extensions: motion linearity loss and edge-guided training. The motion linearity loss approximates the motion between two input frames to be linear and regularizes the training. By applying edge-guided training, we further improve results by integrating edge information into training. Both qualitative and quantitative experiments demonstrate that our model outperforms the state-of-the-art methods. The source codes of the proposed method and more experimental results will be available at https://github.com/alex04072000/CyclicGen. Yu-Lun Liu 0001, Yi-Tung Liao, Yen-Yu Lin, Yung-Yu Chuang |
AAAI | 1 |
| 2013 | Virtual view synthesis using backward depth warping algorithmabstractThe virtual view synthesis reference software offered by the MPEG standard committee adopts the forward warping technique in projecting the depth map from the reference view to the target (virtual) view location. Often, this warping process results in many artifacts, holes and cracks, due to quantization errors and occlusion. In this study, we propose a backward warping process to replace the forward warping process, and the artifacts (particularly the ones produced by quantization) are significantly reduced. The subjective quality of the synthesized virtual view images is thus much improved. Du-Hsiu Li, Hsueh-Ming Hang, Yu-Lun Liu 0001 |
PCS | 3 |