EDBT 2026 Demo / reviewers in the wild / expert
Kang Liao
dblp:231/8721
· DBLP profile ↗
52ranked-venue papers
11as first author
47since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 8 first-author · 33 since 2021Artificial intelligence and machine learning · 26 · 6 first-author · 26 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Wide-Angle Images: Structure-to-Detail Video Portrait Correction via Unsupervised Spatiotemporal AdaptationabstractWide-angle cameras, despite their popularity for content creation, suffer from distortion-induced facial stretching—especially at the edge of the lens—which degrades visual appeal. To address this issue, we propose a structure-to-detail portrait correction model named ImagePC. It integrates the long-range awareness of the transformer and multi-step denoising of diffusion models into a unified framework, achieving global structural robustness and local detail refinement. Besides, considering the high cost of obtaining video labels, we then repurpose ImagePC for unlabeled wide-angle videos (termed VideoPC), by spatiotemporal diffusion adaption with spatial consistency and temporal smoothness constraints. For the former, we encourage the denoised image to approximate pseudo labels following the wide-angle distortion distribution pattern, while for the latter, we derive rectification trajectories with backward optical flows and smooth them. Compared with ImagePC, VideoPC maintains high-quality facial corrections in space and mitigates the potential temporal shakes sequentially in blind scenarios. Finally, to establish an evaluation benchmark and train the framework, we establish a video portrait dataset with a large diversity in the number of people, lighting conditions, and background. Experiments demonstrate that the proposed methods outperform existing solutions quantitatively and qualitatively, contributing to high-fidelity wide-angle videos with stable and natural portraits. Wenbo Nie, Lang Nie, Chunyu Lin, Jiyuan Wang 0001, Kang Liao |
AAAI | 7 |
| 2026 | Revisiting 360 Depth Estimation With PanoGabor: A New Fusion PerspectiveabstractDepth estimation from a monocular 360 image is important to the perception of the entire 3D environment. However, the inherent distortion and large field of view (FoV) in 360 images pose great challenges for this task. To this end, existing mainstream solutions typically introduce additional perspective-based 360 representations (e.g., Cubemap) to achieve effective feature extraction. Nevertheless, regardless of the introduced representations, they eventually need to be unified into the equirectangular projection (ERP) format for the subsequent depth estimation, which inevitably reintroduces additional distortions. In this work, we propose an oriented-distortion-aware Gabor Fusion framework (PGFuse) to address the above challenges. First, we introduce Gabor filters that analyze texture in the frequency domain, extending the receptive fields and enhancing depth cues. To address the reintroduced distortions, we design a latitude-aware distortion representation to generate customized, distortion-aware Gabor filters (PanoGabor filters). Furthermore, we design a channel-wise and spatial-wise unidirectional fusion module (CS-UFM) that integrates the proposed PanoGabor filters to unify other representations into the ERP format, delivering effective and distortion-aware features. Considering the orientation sensitivity of the Gabor transform, we further introduce a spherical gradient constraint to stabilize this sensitivity. Experimental results on three popular indoor 360 benchmarks demonstrate the superiority of the proposed PGFuse to existing state-of-the-art solutions. Code and models will be available at https://github.com/zhijieshen-bjtu/PGFuse. Zhijie Shen, Chunyu Lin, Lang Nie, Kang Liao, Weisi Lin, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Robust Image Stitching With Optimal PlaneabstractWe present RopStitch, an unsupervised deep image stitching framework with both robustness and naturalness. To ensure the robustness of RopStitch, we propose to incorporate the universal prior of content perception into the image stitching model by a dual-branch architecture. It separately captures coarse and fine features and integrates them to achieve highly generalizable performance across diverse unseen real-world scenes. Concretely, the dual-branch model consists of a pretrained branch to capture semantically invariant representations and a learnable branch to extract fine-grained discriminative features, which are then merged into a whole by a controllable factor at the correlation level. Besides, considering that content alignment and structural preservation are often contradictory to each other, we propose a concept of virtual optimal planes to relieve this conflict. To this end, we model this problem as a process of estimating homography decomposition coefficients, and design an iterative coefficient predictor and minimal semantic distortion constraint to identify the optimal plane. This scheme is finally incorporated into RopStitch by warping both views onto the optimal plane bidirectionally. Extensive experiments across various datasets demonstrate that RopStitch significantly outperforms existing methods, particularly in scene robustness and content naturalness. Lang Nie, Kang Liao, Yunqiu Xu, Chunyu Lin, Bin Xiao 0002 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | Arbitrary-steps Image Super-resolution via Diffusion InversionabstractThis study presents a new image super-resolution (SR) technique based on diffusion inversion, aiming at harnessing the rich image priors encapsulated in large pre-trained diffusion models to improve SR performance. We design a Partial noise Prediction strategy to construct an intermediate state of the diffusion model, which serves as the starting sampling point. Central to our approach is a deep noise predictor to estimate the optimal noise maps for the forward diffusion process. Once trained, this noise predictor can be used to initialize the sampling process partially along the diffusion trajectory, generating the desirable high-resolution result. Compared to existing approaches, our method offers a flexible and efficient sampling mechanism that supports an arbitrary number of sampling steps, ranging from one to five. Even with a single sampling step, our method demonstrates superior or comparable performance to recent state-of-the-art approaches. The code and model are publicly available at https://github.com/zsyOAOA/InvSR. Zongsheng Yue, Kang Liao, Chen Change Loy |
CVPR | 2 |
| 2025 | Lifting the Structural Morphing for Wide-Angle Images Rectification: Unified Content and Boundary Modeling
Wenting Luan, Siqi Lu, Yongbin Zheng, Wanying Xu, Lang Nie, Zongtan Zhou, Kang Liao |
ICCV | 7 |
| 2025 | Denoising as Adaptation: Noise-Space Domain Adaptation for Image RestorationabstractAlthough learning-based image restoration methods have made significant progress, they still struggle with limited generalization to real-world scenarios due to the substantial domain gap caused by training on synthetic data. Existing methods address this issue by improving data synthesis pipelines, estimating degradation kernels, employing deep internal learning, and performing domain adaptation and regularization. Previous domain adaptation methods have sought to bridge the domain gap by learning domain-invariant knowledge in either feature or pixel space. However, these techniques often struggle to extend to low-level vision tasks within a stable and compact framework. In this paper, we show that it is possible to perform domain adaptation via the noise space using diffusion models. In particular, by leveraging the unique property of how auxiliary conditional inputs influence the multi-step denoising process, we derive a meaningful *diffusion loss* that guides the restoration model in progressively aligning both restored synthetic and real-world outputs with a target clean distribution. We refer to this method as *denoising as adaptation*. To prevent shortcuts during joint training, we present crucial strategies such as channel-shuffling layer and residual-swapping contrastive learning in the diffusion model. They implicitly blur the boundaries between conditioned synthetic and real data and prevent the reliance of the model on easily distinguishable features. Experimental results on three classical image restoration tasks, namely denoising, deblurring, and deraining, demonstrate the effectiveness of the proposed method. Kang Liao, Zongsheng Yue, Zhouxia Wang, Chen Change Loy |
ICLR | 1 |
| 2025 | Jasmine: Harnessing Diffusion Prior for Self-supervised Depth EstimationabstractIn this paper, we propose \textbf{Jasmine}, the first Stable Diffusion (SD)-based self-supervised framework for monocular depth estimation, which effectively harnesses SD’s visual priors to enhance the sharpness and generalization of unsupervised prediction. Previous SD-based methods are all supervised since adapting diffusion models for dense prediction requires high-precision supervision. In contrast, self-supervised reprojection suffers from inherent challenges (\textit{e.g.}, occlusions, texture-less regions, illumination variance), and the predictions exhibit blurs and artifacts that severely compromise SD's latent priors. To resolve this, we construct a novel surrogate task of mix-batch image reconstruction. Without any additional supervision, it preserves the detail priors of SD models by reconstructing the images themselves while preventing depth estimation from degradation. Furthermore, to address the inherent misalignment between SD's scale and shift invariant estimation and self-supervised scale-invariant depth estimation, we build the Scale-Shift GRU. It not only bridges this distribution gap but also isolates the fine-grained texture of SD output against the interference of reprojection loss. Extensive experiments demonstrate that Jasmine achieves SoTA performance on the KITTI benchmark and exhibits superior zero-shot generalization across multiple datasets. Jiyuan Wang 0001, Chunyu Lin, Cheng Guan, Lang Nie, Kang Liao, Yao Zhao 0001 |
NeurIPS | 7 |
| 2025 | MOWA: Multiple-in-One Image Warping ModelabstractWhile recent image warping approaches achieved remarkable success on existing benchmarks, they still require training separate models for each specific task and cannot generalize well to different camera models or customized manipulations. To address diverse types of warping in practice, we propose a Multiple-in-One image WArping model (named MOWA) in this work. Specifically, we mitigate the difficulty of multi-task learning by disentangling the motion estimation at both the region level and pixel level. To further enable dynamic task-aware image warping, we introduce a lightweight point-based classifier that predicts the task type, serving as prompts to modulate the feature maps for more accurate estimation. To our knowledge, this is the first work that solves multiple practical warping tasks in one single model. Extensive experiments demonstrate that our MOWA, which is trained on six tasks for multiple-in-one single image warping, outperforms state-of-the-art task-specific models across most tasks. Moreover, MOWA also exhibits promising potential to generalize into unseen scenes, as evidenced by cross-domain and zero-shot evaluations. Kang Liao, Zongsheng Yue, Chen Change Loy |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | StabStitch++: Unsupervised Online Video Stitching With Spatiotemporal Bidirectional WarpsabstractWe retarget video stitching to an emerging issue, named warping shake, which unveils the temporal content shakes induced by sequentially unsmooth warps when extending image stitching to video stitching. Even if the input videos are stable, the stitched video can inevitably cause undesired warping shakes and affect the visual experience. To address this issue, we propose StabStitch++, a novel video stitching framework to realize spatial stitching and temporal stabilization with unsupervised learning simultaneously. First, different from existing learning-based image stitching solutions that typically warp one image to align with another, we suppose a virtual midplane between original image planes and project them onto it. Concretely, we design a differentiable bidirectional decomposition module to disentangle the homography transformation and incorporate it into our spatial warp, evenly spreading alignment burdens and projective distortions across two views. Then, inspired by camera paths in video stabilization, we derive the mathematical expression of stitching trajectories in video stitching by elaborately integrating spatial and temporal warps. Finally, a warp smoothing model is presented to produce stable stitched videos with a hybrid loss to simultaneously encourage content alignment, trajectory smoothness, and online collaboration. Compared with StabStitch that sacrifices alignment for stabilization, StabStitch++ makes no compromise and optimizes both of them simultaneously, especially in the online mode. To establish an evaluation benchmark and train the learning framework, we build a video stitching dataset with a rich diversity in camera motions and scenes. Experiments exhibit that StabStitch++ surpasses current solutions in stitching performance, robustness, and efficiency, offering compelling advancements in this field by building a real-time online video stitching system. Lang Nie, Chunyu Lin, Kang Liao, Yun Zhang 0024, Shuaicheng Liu, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | SGFormer: Spherical Geometry Transformer for 360° Depth EstimationabstractPanoramic distortion poses a significant challenge in 360° depth estimation, particularly pronounced at the north and south poles. Existing methods either adopt a bi-projection fusion strategy to remove distortions or model long-range dependencies to capture global structures, resulting in either unclear structure or insufficient local perception. In this paper, we propose a spherical geometry transformer, named SGFormer, to address the above issues, with an innovative step to integrate spherical geometric priors into vision transformers. To this end, we retarget the transformer decoder to a spherical prior decoder (termed SPDecoder), which endeavors to uphold the integrity of spherical structures during decoding. Concretely, we leverage bipolar reprojection, circular rotation, and curve local embedding to preserve the spherical characteristics of equidistortion, continuity, and surface distance, respectively. Furthermore, we present a query-based global conditional position embedding to compensate for spatial structure at varying resolutions. It not only boosts the global perception of spatial position but also sharpens the depth structure across different patches. Finally, we conduct extensive experiments on popular benchmarks, demonstrating our superiority over state-of-the-art solutions. Our code will be made publicly athttps://github.com/iuiuJaon/SGFormer. Junsong Zhang, Zisong Chen, Chunyu Lin, Zhijie Shen, Lang Nie, Kang Liao, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | FishFormer: Annulus Slicing-based Transformer for Fisheye RectificationabstractNumerous significant progress on fisheye image rectification has been achieved through CNN. Nevertheless, constrained by a fixed receptive field, the global distribution and the local symmetry of the distortion have not been fully exploited. To leverage these two characteristics, we introduce FishFormer that processes the fisheye image as a sequence to enhance global and local perception. We tune the Transformer according to the structural properties of fisheye images. First, the uneven distortion distribution in patches generated by the existing square slicing method hinders the understanding of the global structure. Therefore, we propose an annulus slicing method to maintain the consistency of the distortion in each patch; thus, the applicability of the Transformer is expanded to perceive the distortion distribution efficiently. Second, the distortion of adjacent patches is progressive. Such explicit correlations in local regions need to be rapidly constructed and maintained, but Transformer has a weakness in local area perception. Hence, a novel layer attention mechanism is introduced to enhance the local perception and feature interaction. Our network simultaneously implements global perception and focused local perception. Extensive experiments demonstrate that our method provides superior performance compared with state-of-the-art methods. Shangrong Yang, Chunyu Lin, Kang Liao, Yao Zhao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Eliminating Warping Shakes for Unsupervised Online Video Stitching
Lang Nie, Chunyu Lin, Kang Liao, Yun Zhang 0024, Shuaicheng Liu, Rui Ai 0001, Yao Zhao 0001 |
ECCV (4) | 3 |
| 2024 | Digging into Contrastive Learning for Robust Depth Estimation with Diffusion ModelsabstractRecently, diffusion-based depth estimation methods have drawn widespread attention due to their elegant denoising patterns and promising performance. However, they are typically unreliable under adverse conditions prevalent in real-world scenarios, such as rainy, snowy, etc. In this paper, we propose a novel robust depth estimation method called D4RD, featuring a custom contrastive learning mode tailored for diffusion models to mitigate performance degradation in complex environments. Concretely, we integrate the strength of knowledge distillation into contrastive learning, building the `trinity' contrastive scheme. This scheme utilizes the sampled noise of the forward diffusion process as a natural reference, guiding the predicted noise in diverse scenes toward a more stable and precise optimum. Moreover, we extend noise-level trinity to encompass more generic feature and image levels, establishing a multi-level contrast to distribute the burden of robust perception across the overall network. Before addressing complex scenarios, we enhance the stability of the baseline diffusion model with three straightforward yet effective improvements, which facilitate convergence and remove depth outliers. Extensive experiments demonstrate that D4RD surpasses existing state-of-the-art solutions on synthetic corruption datasets and real-world weather conditions. Source code and data are available at \url{https://github.com/wangjiyuan9/D4RD}. Jiyuan Wang 0001, Chunyu Lin, Lang Nie, Kang Liao, Shuwei Shao, Yao Zhao 0001 |
ACM Multimedia | 4 |
| 2024 | Semi-Supervised Coupled Thin-Plate Spline Model for Rotation Correction and BeyondabstractThin-plate spline (TPS) is a principal warp that allows for representing elastic, nonlinear transformation with control point motions. With the increase of control points, the warp becomes increasingly flexible but usually encounters a bottleneck caused by undesired issues, e.g., content distortion. In this paper, we explore generic applications of TPS in single-image-based warping tasks, such as rotation correction, rectangling, and portrait correction. To break this bottleneck, we propose the coupled thin-plate spline model (CoupledTPS), which iteratively couples multiple TPS with limited control points into a more flexible and powerful transformation. Concretely, we first design an iterative search to predict new control points according to the current latent condition. Then, we present the warping flow as a bridge for the coupling of different TPS transformations, effectively eliminating interpolation errors caused by multiple warps. Besides, in light of the laborious annotation cost, we develop a semi-supervised learning scheme to improve warping quality by exploiting unlabeled data. It is formulated through dual transformation between the searched control points of unlabeled data and its graphic augmentation, yielding an implicit correction consistency constraint. Finally, we collect massive unlabeled data to exhibit the benefit of our semi-supervised scheme in rotation correction. Extensive experiments demonstrate the superiority and universality of CoupledTPS over the existing State-of-the-Art (SoTA) solutions for rotation correction and beyond. Lang Nie, Chunyu Lin, Kang Liao, Shuaicheng Liu, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | 360 Layout Estimation via Orthogonal Planes Disentanglement and Multi-View Geometric Consistency PerceptionabstractExisting panoramic layout estimation solutions tend to recover room boundaries from a vertically compressed sequence, yielding imprecise results as the compression process often muddles the semantics between various planes. Besides, these data-driven approaches impose an urgent demand for massive data annotations, which are laborious and time-consuming. For the first problem, we propose an orthogonal plane disentanglement network (termed DOPNet) to distinguish ambiguous semantics. DOPNet consists of three modules that are integrated to deliver distortion-free, semantics-clean, and detail-sharp disentangled representations, which benefit the subsequent layout recovery. For the second problem, we present an unsupervised adaptation technique tailored for horizon-depth and ratio representations. Concretely, we introduce an optimization strategy for decision-level layout analysis and a 1D cost volume construction method for feature-level multi-view aggregation, both of which are designed to fully exploit the geometric consistency across multiple perspectives. The optimizer provides a reliable set of pseudo-labels for network training, while the 1D cost volume enriches each view with comprehensive scene information derived from other perspectives. Extensive experiments demonstrate that our solution outperforms other SoTA models on both monocular layout estimation and multi-view layout estimation tasks. Zhijie Shen, Chunyu Lin, Junsong Zhang, Lang Nie, Kang Liao, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Cylin-Painting: Seamless 360° Panoramic Image Outpainting and BeyondabstractImage outpainting gains increasing attention since it can generate the complete scene from a partial view, providing a valuable solution to construct 360° panoramic images. As image outpainting suffers from the intrinsic issue of unidirectional completion flow, previous methods convert the original problem into inpainting, which allows a bidirectional flow. However, we find that inpainting has its own limitations and is inferior to outpainting in certain situations. The question of how they may be combined for the best of both has as yet remained under-explored. In this paper, we provide a deep analysis of the differences between inpainting and outpainting, which essentially depends on how the source pixels contribute to the unknown regions under different spatial arrangements. Motivated by this analysis, we present a Cylin-Painting framework that involves meaningful collaborations between inpainting and outpainting and efficiently fuses the different arrangements, with a view to leveraging their complementary benefits on a seamless cylinder. Nevertheless, straightforwardly applying the cylinder-style convolution often generates visually unpleasing results as it discards important positional information. To address this issue, we further present a learnable positional embedding strategy to incorporate the missing component of positional encoding into the cylinder convolution, which significantly improves the panoramic results. It is noted that while developed for image outpainting, the proposed algorithm can be effectively extended to other panoramic vision tasks, such as object detection, depth estimation, and image super-resolution. Code will be made available at https://github.com/KangLiao929/Cylin-Painting. Kang Liao, Xiangyu Xu 0002, Chunyu Lin, Wenqi Ren, Yunchao Wei, Yao Zhao 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Spatiotemporal Deformation Perception for Fisheye Video RectificationabstractAlthough the distortion correction of fisheye images has been extensively studied, the correction of fisheye videos is still an elusive challenge. For different frames of the fisheye video, the existing image correction methods ignore the correlation of sequences, resulting in temporal jitter in the corrected video. To solve this problem, we propose a temporal weighting scheme to get a plausible global optical flow, which mitigates the jitter effect by progressively reducing the weight of frames. Subsequently, we observe that the inter-frame optical flow of the video is facilitated to perceive the local spatial deformation of the fisheye video. Therefore, we derive the spatial deformation through the flows of fisheye and distorted-free videos, thereby enhancing the local accuracy of the predicted result. However, the independent correction for each frame disrupts the temporal correlation. Due to the property of fisheye video, a distorted moving object may be able to find its distorted-free pattern at another moment. To this end, a temporal deformation aggregator is designed to reconstruct the deformation correlation between frames and provide a reliable global feature. Our method achieves an end-to-end correction and demonstrates superiority in correction quality and stability compared with the SOTA correction methods. Shangrong Yang, Chunyu Lin, Kang Liao, Yao Zhao 0001 |
AAAI | 3 |
| 2023 | Disentangling Orthogonal Planes for Indoor Panoramic Room Layout Estimation with Cross-Scale Distortion AwarenessabstractBased on the Manhattan World assumption, most existing indoor layout estimation schemes focus on recovering layouts from vertically compressed 1D sequences. However, the compression procedure confuses the semantics of different planes, yielding inferior performance with ambiguous interpretability. To address this issue, we propose to disentangle this 1D representation by pre-segmenting orthogonal (vertical and horizontal) planes from a complex scene, explicitly capturing the geometric cues for indoor layout estimation. Considering the symmetry between the floor boundary and ceiling boundary, we also design a soft-flipping fusion strategy to assist the pre-segmentation. Besides, we present a feature assembling mechanism to effectively integrate shallow and deep features with distortion distribution awareness. To compensate for the potential errors in pre-segmentation, we further leverage triple attention to reconstruct the disentangled sequences for better performance. Experiments on four popular benchmarks demonstrate our superiority over existing SoTA solutions, especially on the 3DIoU metric. The code is available at https://github.com/zhijieshen-bjtu/DOPNet. Zhijie Shen, Zishuo Zheng, Chunyu Lin, Lang Nie, Kang Liao, Shuai Zheng 0005, Yao Zhao 0001 |
CVPR | 5 |
| 2023 | Towards Reliable Image Outpainting: Learning Structure-Aware Multimodal Fusion with Depth GuidanceabstractImage outpainting technology generates visually plausible content regardless of authenticity, making it unreliable to be applied in practice. Thus, we propose a reliable image outpainting task, introducing the sparse depth from LiDARs (Light Detection And Ranging devices) to extrapolate authentic RGB scenes. The large field view of LiDARs allows it to serve for data enhancement and further multimodal tasks. Concretely, we propose a Depth-Guided Outpainting Network to model different feature representations of two modalities and learn the structure-aware cross-modal fusion. And two components are designed: 1) The Multimodal Learning Module produces unique depth and RGB feature representations from the perspectives of different modal characteristics. 2) The Depth Guidance Fusion Module leverages the complete depth modality to guide the establishment of RGB contents by progressive multimodal feature fusion. Furthermore, we specially design an additional constraint strategy consisting of Cross-modal Loss and Edge Loss to enhance ambiguous contours and expedite reliable content generation. Extensive experiments on KITTI and Waymo datasets demonstrate our superiority over the state-of-the-art method, quantitatively and qualitatively. Lei Zhang 0116, Chunyu Lin, Kang Liao, Yao Zhao 0001 |
ICASSP | 3 |
| 2023 | RecRecNet: Rectangling Rectified Wide-Angle Images by Thin-Plate Spline Model and DoF-based Curriculum LearningabstractThe wide-angle lens shows appealing applications in VR technologies, but it introduces severe radial distortion into its captured image. To recover the realistic scene, previous works devote to rectifying the content of the wide-angle image. However, such a rectification solution inevitably distorts the image boundary, which changes related geometric distributions and misleads the current vision perception models. In this work, we explore constructing a win-win representation on both content and boundary by contributing a new learning model, i.e., Rectangling Rectification Network (RecRecNet). In particular, we propose a thin-plate spline (TPS) module to formulate the nonlinear and non-rigid transformation for rectangling images. By learning the control points on the rectified image, our model can flexibly warp the source structure to the target domain and achieves an end-to-end unsupervised deformation. To relieve the complexity of structure approximation, we then inspire our RecRecNet to learn the gradual deformation rules with a DoF (Degree of Freedom)-based curriculum learning. By increasing the DoF in each curriculum stage, namely, from similarity transformation (4-DoF) to homography transformation (8-DoF), the network is capable of investigating more detailed deformations, offering fast convergence on the final rectangling task. Experiments show the superiority of our solution over the compared methods on both quantitative and qualitative evaluations. The code and dataset are available at https://github.com/KangLiao929/RecRecNet. Kang Liao, Lang Nie, Chunyu Lin, Zishuo Zheng, Yao Zhao 0001 |
ICCV | 1 |
| 2023 | Parallax-Tolerant Unsupervised Deep Image StitchingabstractTraditional image stitching approaches tend to leverage increasingly complex geometric features (e.g., point, line, edge, etc.) for better performance. However, these hand-crafted features are only suitable for specific natural scenes with adequate geometric structures. In contrast, deep stitching schemes overcome adverse conditions by adaptively learning robust semantic features, but they cannot handle large-parallax cases.To solve these issues, we propose a parallax-tolerant unsupervised deep image stitching technique. First, we propose a robust and flexible warp to model the image registration from global homography to local thin-plate spline motion. It provides accurate alignment for overlapping regions and shape preservation for non-overlapping regions by joint optimization concerning alignment and distortion. Subsequently, to improve the generalization capability, we design a simple but effective iterative strategy to enhance the warp adaption in cross-dataset and cross-resolution applications. Finally, to further eliminate the parallax artifacts, we propose to composite the stitched image seamlessly by unsupervised learning for seam-driven composition masks. Compared with existing methods, our solution is parallax-tolerant and free from laborious designs of complicated geometric features for specific scenes. Extensive experiments show our superiority over the SoTA methods, both quantitatively and qualitatively. The code is available at https://github.com/nie-lang/UDIS2. Lang Nie, Chunyu Lin, Kang Liao, Shuaicheng Liu, Yao Zhao 0001 |
ICCV | 3 |
| 2023 | Innovating Real Fisheye Image Correction with Dual Diffusion ArchitectureabstractFisheye image rectification is hindered by synthetic models producing poor results for real-world correction. To address this, we propose a Dual Diffusion Architecture (DDA) for fisheye rectification that offers better practicality. The DDA leverages Denoising Diffusion Probabilistic Models (DDPMs) to gradually introduce bidirectional noise, allowing the synthesized and real images to develop into a consistent noise distribution. As a result, our network can perceive the distribution of unlabelled real fisheye images without relying on a transfer network, thus improving the performance of real fisheye correction. Additionally, we design an unsupervised one-pass network that generates a plausible new condition to strengthen guidance and address the non-negligible indeterminacy between the prior condition and the target. It can significantly affect the rectification task, especially in cases where radial distortion causes significant artifacts. This network can be regarded as an alternate scheme for fast producing reliable results without iterative inference. Compared to the state-of-the-art methods, our approach achieves superior performance in both synthetic and real fisheye image corrections. Shangrong Yang, Chunyu Lin, Kang Liao, Yao Zhao 0001 |
ICCV | 3 |
| 2023 | Learning Deposition Policies for Fused Multi-Material 3D Printingabstract3D printing based on continuous deposition of materials, such as filament-based 3D printing, has seen widespread adoption thanks to its versatility in working with a wide range of materials. An important shortcoming of this type of technology is its limited multi-material capabilities. While there are simple hardware designs that enable multi-material printing in principle, the required software is heavily underdeveloped. A typical hardware design fuses together individual materials fed into a single chamber from multiple inlets before they are deposited. This design, however, introduces a time delay between the intended material mixture and its actual deposition. In this work, inspired by diverse path planning research in robotics, we show that this mechanical challenge can be addressed via improved printer control. We propose to formulate the search for optimal multi-material printing policies in a reinforcement learning setup. We put forward a simple numerical deposition model that takes into account the non-linear material mixing and delayed material deposition. To validate our system we focus on color fabrication, a problem known for its strict requirements for varying material mixtures at a high spatial frequency. We demonstrate that our learned control policy outperforms state-of-the-art hand-crafted algorithms. Kang Liao, Thibault Tricard, Michal Piovarci, Hans-Peter Seidel, Vahid Babaei |
ICRA | 1 |
| 2023 | Unsupervised OmniMVS: Efficient Omnidirectional Depth Inference via Establishing Pseudo-Stereo SupervisionabstractOmnidirectional multi-view stereo (MVS) vision is attractive for its ultra-wide field-of-view (FoV), enabling machines to perceive 360°3D surroundings. However, the existing solutions require expensive dense depth labels for supervision, making them impractical in real-world applications. In this paper, we propose the first unsupervised omnidirectional MVS framework based on multiple fisheye images. To this end, we project all images to a virtual view center and composite two panoramic images with spherical geometry from two pairs of back-to-back fisheye images. The two 360° images formulate a stereo pair with a special pose, and the photometric consistency is leveraged to establish the unsupervised constraint, which we term “Pseudo-Stereo Supervision”. In addition, we propose Un-OmniMVS, an efficient unsupervised omnidirectional MVS network, to facilitate the inference speed with two efficient components. First, a novel feature extractor with frequency attention is proposed to simultaneously capture the non-local Fourier features and local spatial features, explicitly facilitating the feature representation. Then, a variance-based light cost volume is put forward to reduce the computational complexity. Experiments exhibit that the performance of our unsupervised solution is competitive to that of the state-of-the-art (SoTA) supervised methods with better generalization in real-world data. The code will be available at https://github.com/Chen-z-s/Un-OmniMVS. Zisong Chen, Chunyu Lin, Lang Nie, Kang Liao, Yao Zhao 0001 |
IROS | 4 |
| 2023 | S-OmniMVS: Incorporating Sphere Geometry into Omnidirectional Stereo MatchingabstractMulti-fisheye stereo matching is a promising task that employs the traditional multi-view stereo (MVS) pipeline with spherical sweeping to acquire omnidirectional depth. However, the existing omnidirectional MVS technologies neglect fisheye and omnidirectional distortions, yielding inferior performance. In this paper, we revisit omnidirectional MVS by incorporating three sphere geometry priors: spherical projection, spherical continuity, and spherical position. To deal with fisheye distortion, we propose a new distortion-adaptive fusion module to convert fisheye inputs into distortion-free spherical tangent representations by constructing a spherical projection space. Then these multi-scale features are adaptively aggregated with additional learnable offsets to enhance content perception. To handle omnidirectional distortion, we present a new spherical cost aggregation module with a comprehensive consideration of the spherical continuity and position. Concretely, we first design a rotation continuity compensation mechanism to ensure omnidirectional depth consistency of left-right boundaries without introducing extra computation. On the other hand, we encode the geometry-aware spherical position and push them into the cost aggregation to relieve panoramic distortion and perceive the 3D structure. Furthermore, to avoid the excessive concentration of depth hypothesis caused by inverse depth linear sampling, we develop a segmented sampling strategy that combines linear and exponential spaces to create S-OmniMVS, along with three sphere priors. Extensive experiments demonstrate the proposed method outperforms the state-of-the-art (SoTA) solutions by a large margin on various datasets both quantitatively and qualitatively. Zisong Chen, Chunyu Lin, Lang Nie, Zhijie Shen, Kang Liao, Yuanzhouhan Cao, Yao Zhao 0001 |
ACM Multimedia | 5 |
| 2023 | Complementary Bi-directional Feature Compression for Indoor 360° Semantic Segmentation with Self-distillationabstractSemantic segmentation on 360° images is a vital component of scene understanding due to the rich surrounding information. Recently, horizontal representation-based approaches outperform projection-based solutions, because the distortions can be effectively removed by compressing the spherical data in the vertical direction. However, these methods ignore the distortion distribution prior and are limited to unbalanced receptive fields, e.g., the receptive fields are sufficient in the vertical direction and insufficient in the horizontal direction. Differently, a vertical representation compressed in another direction can offer implicit distortion prior and enlarge horizontal receptive fields. In this paper, we combine the two different representations and propose a novel 360° semantic segmentation solution from a complementary perspective. Our network comprises three modules: a feature extraction module, a bi-directional compression module, and an ensemble decoding module. First, we extract multi-scale features from a panorama. Then, a bi-directional compression module is designed to compress features into two complementary low-dimensional representations, which provide content perception and distortion prior. Furthermore, to facilitate the fusion of bi-directional features, we design a unique self distillation strategy in the ensemble decoding module to enhance the interaction of different features and further improve the performance. Experimental results show that our approach outperforms the state-of-the-art solutions on quantitative evaluations while displaying the best performance on visual appearance. Zishuo Zheng, Chunyu Lin, Lang Nie, Kang Liao, Zhijie Shen, Yao Zhao 0001 |
WACV | 4 |
| 2023 | As-Deformable-As-Possible Single-Image-Based View Synthesis Without Depth PriorabstractDepth-image-based rendering (DIBR) technologies have been widely employed to synthesize novel realistic views from a single image in 3D video applications. However, DIBR-oriented approaches heavily rely on the accuracy of depth maps, usually requiring the depth GT as a prior. Despite that, there might exist extensive float precision losses and invalid holes in the synthesized view due to warping error and occlusion. In this paper, we propose an end-to-end as-deformable-as-possible (ADAP) single-image-based view synthesis solution without depth prior. It addresses the above issues through two stages: alignment and reconstruction, where we first transform the input image to the latent feature space and then reconstruct the novel view in the image domain. In the first stage, the input image is deformed to align with the synthesized view at feature level. To this end, we propose an ADAP alignment mechanism through pixel-level warping to error-level quantization to feature-level alignment, progressively improving the deformable capability in handling challenging motion conditions in real-world scenes. In the second stage, we exploit an occlusion-aware reconstruction module to recover the content details from the deformed feature at pixel level. Extensive experiments demonstrate that our alignment-reconstruction approach is robust to the depth map. Even with a coarsely estimated depth map, our solution outperforms other SoTA schemes in the popular benchmarks. Chunlan Zhang, Chunyu Lin, Kang Liao, Lang Nie, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Deep Rotation Correction Without Angle PriorabstractNot everybody can be equipped with professional photography skills and sufficient shooting time, and there can be some tilts in the captured images occasionally. In this paper, we propose a new and practical task, named Rotation Correction, to automatically correct the tilt with high content fidelity in the condition that the rotated angle is unknown. This task can be easily integrated into image editing applications, allowing users to correct the rotated images without any manual operations. To this end, we leverage a neural network to predict the optical flows that can warp the tilted images to be perceptually horizontal. Nevertheless, the pixel-wise optical flow estimation from a single image is severely unstable, especially in large-angle tilted images. To enhance its robustness, we propose a simple but effective prediction strategy to form a robust elastic warp. Particularly, we first regress the mesh deformation that can be transformed into robust initial optical flows. Then we estimate residual optical flows to facilitate our network the flexibility of pixel-wise deformation, further correcting the details of the tilted images. To establish an evaluation benchmark and train the learning framework, a comprehensive rotation correction dataset is presented with a large diversity in scenes and rotated angles. Extensive experiments demonstrate that even in the absence of the angle prior, our algorithm can outperform other state-of-the-art solutions requiring this prior. The code and dataset are available at https://github.com/nie-lang/RotationCorrection. Lang Nie, Chunyu Lin, Kang Liao, Shuaicheng Liu, Yao Zhao 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | Monocular Pseudo-LiDAR Point Cloud Extrapolation Based on Iterative Hybrid RenderingabstractRecently, a pseudo-LiDAR point cloud extrapolation algorithm equipped with stereo cameras has been introduced, bridging the gap between the expensive 3D sensor LiDAR and relatively cheap 2D sensor camera in autonomous driving. In this paper, we explore an approach to further bridge this gap using only a monocular camera and extrapolate a wide field of view 3D point cloud from a limited 2D view. However, this task is extremely challenging as it requires inferring the occluded contents in the scene. To this end, we propose a ‘render-refine-iterate-fuse’ framework that takes advantage of both image view synthesis and image inpainting techniques, guiding the neural network to learn the potential spatial distribution. In addition, we design a hybrid rendering scheme to ensure that the visible content moves in a geometrically correct manner and fills the pixels caused by occlusion. Benefitting from the proposed framework, our approach achieves significant improvements on the pseudo-LiDAR point cloud extrapolation task. The gap between LiDAR and cameras is further bridged, showing an economical and practical application in the environment perception module of autonomous driving. The experimental results evaluated on the KITTI dataset demonstrate that our approach achieves superior quantitative and qualitative performance. Chunlan Zhang, Chunyu Lin, Kang Liao, Yao Zhao 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Deep Rectangling for Image Stitching: A Learning BaselineabstractStitched images provide a wide field-of-view (FoV) but suffer from unpleasant irregular boundaries. To deal with this problem, existing image rectangling methods devote to searching an initial mesh and optimizing a target mesh to form the mesh deformation in two stages. Then rectangu-lar images can be generated by warping stitched images. However, these solutions only work for images with rich linear structures, leading to noticeable distortions for por-traits and landscapes with non-linear objects. In this paper, we address these issues by proposing the first deep learning solution to image rectangling. Con-cretely, we predefine a rigid target mesh and only estimate an initial mesh to form the mesh deformation, contributing to a compact one-stage solution. The initial mesh is predicted using a fully convolutional network with a resid-ual progressive regression strategy. To obtain results with high content fidelity, a comprehensive objective function is proposed to simultaneously encourage the boundary rect-angular, mesh shape-preserving, and content perceptually natural. Besides, we build the first image stitching rectan-gling dataset with a large diversity in irregular boundaries and scenes. Experiments demonstrate our superiority over traditional methods both quantitatively and qualitatively. Lang Nie, Chunyu Lin, Kang Liao, Shuaicheng Liu, Yao Zhao 0001 |
CVPR | 3 |
| 2022 | PanoFormer: Panorama Transformer for Indoor 360$^{\circ }$ Depth Estimation
Zhijie Shen, Chunyu Lin, Kang Liao, Lang Nie, Zishuo Zheng, Yao Zhao 0001 |
ECCV (1) | 3 |
| 2022 | SivsFormer: Parallax-Aware Transformers for Single-image-based View SynthesisabstractSingle-image-based view synthesis is significant for generating a 3D scene and gains increasing attention in recent years. However, this task is challenging as it requires inferring contents beyond what is immediately visible. Previous methods directly predict the unknown views using the convolutional neural networks, but the generated views suffer from visually unpleasant holes, deformations, and artifacts. In this paper, we propose a Single-image-based view synthesis transformer (named SivsFormer) for high-quality and realistic view synthesis. In particular, a warping and occlusion handing module is designed to reduce the influence of parallax on the network. Subsequently, a disparity alignment module captures the long-range information over the scene and ensures that pixels move in a geometrically correct manner with soft probabilistic disparity maps. Moreover, we present a parallax-aware loss function to improve the quality of the synthetic images, which explicitly quantifies the magnitude of parallaxes. We conduct extensive experiments on popular KITTI and Cityscapes datasets. Benefitting from the proposed parallax-aware transformer, our approach achieves superior performance in both quantitative and qualitative evaluations. Chunlan Zhang, Chunyu Lin, Kang Liao, Lang Nie, Yao Zhao 0001 |
VR | 3 |
| 2022 | Learning edge-preserved image stitching from multi-scale deep homography
Lang Nie, Chunyu Lin, Kang Liao, Yao Zhao 0001 |
Neurocomputing | 3 |
| 2022 | Bi-projection for 360°image object detection bridged by RoI Searcher
Zishuo Zheng, Chunyu Lin, Lang Nie, Kang Liao, Yao Zhao 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2022 | Depth-Aware Multi-Grid Deep Homography Estimation With Contextual CorrelationabstractHomography estimation is an important task in computer vision applications, such as image stitching, video stabilization, and camera calibration. Traditional homography estimation methods heavily depend on the quantity and distribution of feature correspondences, leading to poor robustness in low-texture scenes. The learning solutions, on the contrary, try to learn robust deep features but demonstrate unsatisfying performance in the scenes with low overlap rates. In this paper, we address these two problems simultaneously by designing a contextual correlation layer (CCL). The CCL can efficiently capture the long-range correlation within feature maps and can be flexibly used in a learning framework. In addition, considering that a single homography can not represent the complex spatial transformation in depth-varying images with parallax, we propose to predict multi-grid homography from global to local. Moreover, we equip our network with a depth perception capability, by introducing a novel depth-aware shape-preserved loss. Extensive experiments demonstrate the superiority of our method over state-of-the-art solutions in the synthetic benchmark dataset and real-world dataset. The codes and models will be available athttps://github.com/nie-lang/Multi-Grid-Deep-Homography. Lang Nie, Chunyu Lin, Kang Liao, Shuaicheng Liu, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Neural Contourlet Network for Monocular 360° Depth EstimationabstractFor a monocular 360° image, depth estimation is a challenging because the distortion increases along the latitude. To perceive the distortion, existing methods devote to designing a deep and complex network architecture. In this paper, we provide a new perspective that constructs an interpretable and sparse representation for a 360° image. Considering the importance of the geometric structure in depth estimation, we utilize the contourlet transform to capture an explicit geometric cue in the spectral domain and integrate it with an implicit cue in the spatial domain. Specifically, we propose a neural contourlet network consisting of a convolutional neural network and a contourlet transform branch. In the encoder stage, we design a spatial–spectral fusion module to effectively fuse two types of cues. Contrary to the encoder, we employ the inverse contourlet transform with learned low-pass subbands and band-pass directional subbands to compose the depth in the decoder. Experiments on the three popular 360° panoramic image datasets demonstrate that the proposed approach outperforms the state-of-the-art schemes with faster convergence. Code is available athttps://github.com/zhijieshen-bjtu/Neural-Contourlet-Network-for-MODE. Zhijie Shen, Chunyu Lin, Lang Nie, Kang Liao, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Revisiting Radial Distortion Rectification in Polar-Coordinates: A New and Efficient Learning PerspectiveabstractFisheye cameras can capture a large field-of-view (Fov) scene but it introduces severe radial distortion in images. Thus, distortion rectification is a crucial step for subsequent computer vision tasks using fisheye cameras. A prevalent type of method predicts the displacement field between the input and output to rectify the distorted images. However, it is challenging to estimate the accurate flow in Cartesian coordinates (both$x$and$y$directions), in which the sampling strategy of the convolution kernel ignores the radial symmetry of distortion. In general, the pixel’s distortion at the same radius from the center is the same, while the radius corresponds to one parameter in polar coordinates. Motivated by this fact, we exploit the radial symmetry of distortion to predict a more straightforward one-dimensional flow, transforming the distorted image into the polar coordinates domain instead of predicting two-dimensional flow in$x$and$y$directions. Specifically, we propose a Polar coordinates Distortion Rectification Network (PCDRN), whose sampling strategy corresponds to the radial distortion characteristic so that a more accurate flow can be predicted. To eliminate the blurs and ring artifacts induced by the coordinates transformation, a Polar-To-Cartesian Appearance Enhancement Network is designed to enhance the local appearance of rectified images. Experimental results on the synthesized dataset and real-world dataset demonstrate the superiority of our approach in both quantitative and qualitative evaluations. Keyao Zhao, Chunyu Lin, Kang Liao, Shangrong Yang, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Pseudo-LiDAR Point Cloud Interpolation Based on 3D Motion Representation and Spatial SupervisionabstractPseudo-LiDAR point cloud interpolation is a novel and challenging task in autonomous driving, which aims to address the frequency mismatching problem between a camera and a LiDAR. Previous works represent the 3D spatial motion relationship with a coarse 2D optical flow, and the quality of interpolated point clouds only depends on the supervision of depth maps. As a result, the generated point clouds suffer from inferior global distributions and local appearances. To solve the above problems, we propose a Pseudo-LiDAR point cloud interpolation network to generate temporally and spatially high-quality point cloud sequences. By exploiting the scene flow from point clouds, the proposed network is able to learn a more accurate representation of the 3D spatial motion relationship. For a more comprehensive perception of the distribution of a point cloud, we design a novel reconstruction loss function with the chamfer distance to supervise the generation of Pseudo-LiDAR point clouds in 3D space. In addition, we introduce a multi-modal deep aggregation module to facilitate the efficient fusion of texture and depth features. As the benefits of the improved motion representation, training loss function, and model structure, our approach gains significant improvements on the Pseudo-LiDAR point cloud interpolation task. The experimental results evaluated on KITTI dataset demonstrate the state-of-the-art quantitative and qualitative performance of the proposed network. Kang Liao, Chunyu Lin, Yao Zhao 0001, Yulan Guo |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Progressively Complementary Network for Fisheye Image Rectification Using Appearance FlowabstractDistortion rectification is often required for fisheye images. The generation-based method is one mainstream solution due to its label-free property, but its naive skip-connection and overburdened decoder will cause blur and incomplete correction. First, the skip-connection directly transfers the image features, which may introduce distortion and cause incomplete correction. Second, the decoder is overburdened during simultaneously reconstructing the content and structure of the image, resulting in vague performance. To solve these two problems, in this paper, we focus on the interpretable correction mechanism of the distortion rectification network and propose a feature-level correction scheme. We embed a correction layer in skip-connection and leverage the appearance flows in different layers to pre-correct the image features. Consequently, the decoder can easily reconstruct a plausible result with the remaining distortion-less information. In addition, we propose a parallel complementary structure. It effectively reduces the burden of the decoder by separating content reconstruction and structure correction. Subjective and objective experiment results on different datasets demonstrate the superiority of our method. Shangrong Yang, Chunyu Lin, Kang Liao, Chunjie Zhang 0001, Yao Zhao 0001 |
CVPR | 3 |
| 2021 | Multi-Level Curriculum for Training A Distortion-Aware Barrel Distortion Rectification ModelabstractBarrel distortion rectification aims at removing the radial distortion in a distorted image captured by a wide-angle lens. Previous deep learning methods mainly solve this problem by learning the implicit distortion parameters or the nonlinear rectified mapping function in a direct manner. However, this type of manner results in an indistinct learning process of rectification and thus limits the deep perception of distortion. In this paper, inspired by the curriculum learning, we analyze the barrel distortion rectification task in a progressive and meaningful manner. By considering the relationship among different construction levels in an image, we design a multi-level curriculum that disassembles the rectification task into three levels, structure recovery, semantics embedding, and texture rendering. With the guidance of the curriculum that corresponds to the construction of images, the proposed hierarchical architecture enables a progressive rectification and achieves more accurate results. Moreover, we present a novel distortion-aware pre-training strategy to facilitate the initial learning of neural networks, promoting the model to converge faster and better. Experimental results on the synthesized and real-world distorted image datasets show that the proposed approach significantly outperforms other learning methods, both qualitatively and quantitatively. Kang Liao, Chunyu Lin, Lixin Liao, Yao Zhao 0001, Weiyao Lin |
ICCV | 1 |
| 2021 | Towards Complete Scene and Regular Shape for Distortion Rectification by Curve-Aware ExtrapolationabstractThe wide-angle lens gains increasing attention since it can capture a wide field-of-view (FoV) scene. However, the obtained image is contaminated with radial distortion, making the scene not realistic. Previous distortion rectification methods rectify the image in a rectangle or invagination, failing to display the complete content and regular shape simultaneously. In this paper, we rethink the representation of rectification results and present a Rectification OutPainting (ROP) method, aiming to extrapolate the coherent semantics to the blank area and create a wider FoV beyond the original wide-angle lens. To address the specific challenges such as the variable painting region and curve boundary, a rectification module is designed to rectify the image with geometry supervision, and the extrapolated results are generated using a dual conditional expansion strategy. In terms of the spatially discounted correlation, a curve-aware correlation measurement is proposed to focus on the generated region to enforce the local consistency. To our knowledge, we are the first to tackle the challenging rectification via outpainting, and our curve-aware strategy can reach a rectification construction with complete content and regular shape. Extensive experiments well demonstrate the superiority of our ROP over other state-of-the-art solutions. Kang Liao, Chunyu Lin, Yunchao Wei, Feng Li 0037, Shangrong Yang, Yao Zhao 0001 |
ICCV | 1 |
| 2021 | Distortion-Tolerant Monocular Depth Estimation on Omnidirectional Images Using Dual-CubemapabstractEstimating the depth of omnidirectional images is more challenging than that of normal field-of-view (NFoV) images because the varying distortion can significantly twist an object’s shape. The existing methods suffer from troublesome distortion while estimating the depth of omnidirectional images, leading to inferior performance. To reduce the negative impact of the distortion influence, we propose a distortion-tolerant omnidirectional depth estimation algorithm using a dual-cubemap. It comprises two modules: Dual-Cubemap Depth Estimation (DCDE) module and Boundary Revision (BR) module. In DCDE module, we present a rotation-based dual-cubemap model to estimate the accurate NFoV depth, reducing the distortion at the cost of boundary discontinuity on omnidirectional depths. Then a boundary revision module is designed to smooth the discontinuous boundaries, which contributes to the precise and visually continuous omnidirectional depths. Extensive experiments demonstrate the superiority of our method over other state-of-the-art solutions. Zhijie Shen, Chunyu Lin, Lang Nie, Kang Liao, Yao Zhao 0001 |
ICME | 4 |
| 2021 | Image Outpainting with Depth Assistance
Lei Zhang 0116, Kang Liao, Chunyu Lin, Meiqin Liu 0002, Yao Zhao 0001 |
PRCV (3) | 2 |
| 2021 | Pseudo-LiDAR point cloud magnification
Chunlan Zhang, Kang Liao, Chunyu Lin, Yao Zhao 0001 |
Neurocomputing | 2 |
| 2021 | Joint distortion rectification and super-resolution for self-driving scene perception
Keyao Zhao, Kang Liao, Chunyu Lin, Meiqin Liu 0002, Yao Zhao 0001 |
Neurocomputing | 2 |
| 2021 | A Deep Ordinal Distortion Estimation Approach for Distortion RectificationabstractRadial distortion has widely existed in the images captured by popular wide-angle cameras and fisheye cameras. Despite the long history of distortion rectification, accurately estimating the distortion parameters from a single distorted image is still challenging. The main reason is that these parameters are implicit to image features, influencing the networks to learn the distortion information fully. In this work, we propose a novel distortion rectification approach that can obtain more accurate parameters with higher efficiency. Our key insight is that distortion rectification can be cast as a problem of learning an ordinal distortion from a single distorted image. To solve this problem, we design a local-global associated estimation network that learns the ordinal distortion to approximate the realistic distortion distribution. In contrast to the implicit distortion parameters, the proposed ordinal distortion has a more explicit relationship with image features, and significantly boosts the distortion perception of neural networks. Considering the redundancy of distortion information, our approach only uses a patch of the distorted image for the ordinal distortion estimation, showing promising applications in efficient distortion rectification. In the distortion rectification field, we are the first to unify the heterogeneous distortion parameters into a learning-friendly intermediate representation through ordinal distortion, bridging the gap between image feature and distortion rectification. The experimental results demonstrate that our approach outperforms the state-of-the-art methods by a significant margin, with approximately 23% improvement on the quantitative evaluation while displaying the best performance on visual appearance. Kang Liao, Chunyu Lin, Yao Zhao 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Unsupervised Deep Image Stitching: Reconstructing Stitched Features to ImagesabstractTraditional feature-based image stitching technologies rely heavily on feature detection quality, often failing to stitch images with few features or low resolution. The learning-based image stitching solutions are rarely studied due to the lack of labeled data, making the supervised methods unreliable. To address the above limitations, we propose an unsupervised deep image stitching framework consisting of two stages: unsupervised coarse image alignment and unsupervised image reconstruction. In the first stage, we design an ablation-based loss to constrain an unsupervised homography network, which is more suitable for large-baseline scenes. Moreover, a transformer layer is introduced to warp the input images in the stitching-domain space. In the second stage, motivated by the insight that the misalignments in pixel-level can be eliminated to a certain extent in feature-level, we design an unsupervised image reconstruction network to eliminate the artifacts from features to pixels. Specifically, the reconstruction network can be implemented by a low-resolution deformation branch and a high-resolution refined branch, learning the deformation rules of image stitching and enhancing the resolution simultaneously. To establish an evaluation benchmark and train the learning framework, a comprehensive real-world image dataset for unsupervised deep image stitching is presented and released. Extensive experiments well demonstrate the superiority of our method over other state-of-the-art solutions. Even compared with the supervised solutions, our image stitching quality is still preferred by users. Lang Nie, Chunyu Lin, Kang Liao, Shuaicheng Liu, Yao Zhao 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | A view-free image stitching network based on global homography
Lang Nie, Chunyu Lin, Kang Liao, Meiqin Liu 0002, Yao Zhao 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2020 | Unsupervised fisheye image correction through bidirectional loss with geometric prior
Shangrong Yang, Chunyu Lin, Kang Liao, Yao Zhao 0001, Meiqin Liu 0002 |
J. Vis. Commun. Image Represent. | 3 |
| 2020 | DR-GAN: Automatic Radial Distortion Rectification Using Conditional GAN in Real-TimeabstractRadial distortion, which severely hinders object detection and semantic recognition, frequently exists in images captured using a wide-angle lens. Correction of this distortion of images is crucial in many computer vision applications. In this paper, we present distortion rectification generative adversarial network (DR-GAN), a conditional generative adversarial network (GAN) for automatic radial DR. To the best of our knowledge, this is the first end-to-end trainable adversarial framework for radial distortion rectification. The DR-GAN trained using the proposed low-to-high perceptual loss learns the mapping relation between different structural images rather than estimating multifarious distortion parameters, while also realizing label-free training and one-stage rectification. As a benefit of one-stage rectification, the proposed method is extremely fast with the completion of rectification in real time. This is approximately 22 times faster than the state-of-the-art methods. The experimental results show that the DR-GAN achieves an excellent performance in both quantitative measure (PSNR and SSIM) and visual qualitative appearance. Kang Liao, Chunyu Lin, Yao Zhao 0001, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Distortion Rectification From Static to Dynamic: A Distortion Sequence Construction PerspectiveabstractDistortion rectification is a fundamental task in the field of computer vision and image processing. Nevertheless, previous methods have regarded distortion rectification as a static problem that learns a mapping function and corrects the distorted image to a unique state. However, this state is generally not the optimal solution, as it would result in an under-rectified or over-rectified structure. In this study, we revisit the classical distortion rectification task with a new perspective and redesign the algorithm, inspired by video processing techniques. Specifically, we regard distortion rectification as a dynamic problem that can be extended to a sequence of different distortion states: the input distorted image (t), under-rectified image (t+1), ideal-rectified image (t+2), and over-rectified image (t+3). We first estimate the residual distortion map (RDM) between the input distorted image and the coarse-rectified (t+1 or t+3) image. Here, RDM indicates the motion difference between two distorted images. Subsequently, the RDM is used to guide the refinement rectification process, aiming to convert the coarse-rectified state into the ideal-rectified state. In addition, the flexible implementation of the proposed refinement process with RDM to improve the rectification results of any method is appealing. The experimental results demonstrate that our method outperforms the state-of-the-art schemes by a significant margin, revealing approximately 40% improvement through quantitative evaluation. Kang Liao, Chunyu Lin, Yao Zhao 0001, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Model-Free Distortion Rectification Framework Bridged by Distortion Distribution MapabstractRecently, learning-based distortion rectification schemes have shown high efficiency. However, most of these methods only focus on a specific camera model with fixed parameters, thus failing to be extended to other models. To avoid such a disadvantage, we propose a model-free distortion rectification framework for the single-shot case, bridged by the distortion distribution map (DDM). Our framework is based on an observation that the pixel-wise distortion information is mathematically regular in a distorted image, despite different models having different types and numbers of distortion parameters. Motivated by this observation, instead of estimating the heterogeneous distortion parameters, we construct a proposed distortion distribution map that intuitively indicates the global distortion features of a distorted image. In addition, we develop a dual-stream feature learning module, benefitting from both the advantages of traditional methods that leverage the local handcrafted feature and learning-based methods that focus on the global semantic feature perception. Due to the sparsity of handcrafted features, we discrete the features into a 2D point map and learn the structure inspired by PointNet. Finally, a multimodal attention fusion module is designed to attentively fuse the local structural and global semantic features, providing the hybrid features for the more reasonable scene recovery. The experimental results demonstrate the excellent generalization ability and more significant performance of our method in both quantitative and qualitative evaluations, compared with the stateof- the-art methods. Kang Liao, Chunyu Lin, Yao Zhao 0001, Mai Xu |
IEEE Trans. Image Process. | 1 |