EDBT 2026 Demo / reviewers in the wild / expert
Chunyu Lin
dblp:04/7485
· DBLP profile ↗
112ranked-venue papers
7as first author
70since 2021 · last 2026
0000-0003-2847-0349ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 87 · 7 first-author · 51 since 2021Artificial intelligence and machine learning · 34 · 32 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Systems, architecture and hardware · 3 · 2 since 2021Computer networks · 3 · 2 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Wide-Angle Images: Structure-to-Detail Video Portrait Correction via Unsupervised Spatiotemporal AdaptationabstractWide-angle cameras, despite their popularity for content creation, suffer from distortion-induced facial stretching—especially at the edge of the lens—which degrades visual appeal. To address this issue, we propose a structure-to-detail portrait correction model named ImagePC. It integrates the long-range awareness of the transformer and multi-step denoising of diffusion models into a unified framework, achieving global structural robustness and local detail refinement. Besides, considering the high cost of obtaining video labels, we then repurpose ImagePC for unlabeled wide-angle videos (termed VideoPC), by spatiotemporal diffusion adaption with spatial consistency and temporal smoothness constraints. For the former, we encourage the denoised image to approximate pseudo labels following the wide-angle distortion distribution pattern, while for the latter, we derive rectification trajectories with backward optical flows and smooth them. Compared with ImagePC, VideoPC maintains high-quality facial corrections in space and mitigates the potential temporal shakes sequentially in blind scenarios. Finally, to establish an evaluation benchmark and train the framework, we establish a video portrait dataset with a large diversity in the number of people, lighting conditions, and background. Experiments demonstrate that the proposed methods outperform existing solutions quantitatively and qualitatively, contributing to high-fidelity wide-angle videos with stable and natural portraits. Wenbo Nie, Lang Nie, Chunyu Lin, Jiyuan Wang 0001, Kang Liao |
AAAI | 3 |
| 2026 | Unsupervised multi-modal domain adaptation for RGB-T Semantic Segmentation
Zeyang Chen, Chunyu Lin, Yao Zhao 0001, Tammam Tillo |
Comput. Vis. Image Underst. | 2 |
| 2026 | Point cloud accumulation via multi-dimensional pseudo label and progressive instance association
Chunyu Lin, Lang Nie, Meiqin Liu 0002, Yao Zhao 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2026 | Revisiting 360 Depth Estimation With PanoGabor: A New Fusion PerspectiveabstractDepth estimation from a monocular 360 image is important to the perception of the entire 3D environment. However, the inherent distortion and large field of view (FoV) in 360 images pose great challenges for this task. To this end, existing mainstream solutions typically introduce additional perspective-based 360 representations (e.g., Cubemap) to achieve effective feature extraction. Nevertheless, regardless of the introduced representations, they eventually need to be unified into the equirectangular projection (ERP) format for the subsequent depth estimation, which inevitably reintroduces additional distortions. In this work, we propose an oriented-distortion-aware Gabor Fusion framework (PGFuse) to address the above challenges. First, we introduce Gabor filters that analyze texture in the frequency domain, extending the receptive fields and enhancing depth cues. To address the reintroduced distortions, we design a latitude-aware distortion representation to generate customized, distortion-aware Gabor filters (PanoGabor filters). Furthermore, we design a channel-wise and spatial-wise unidirectional fusion module (CS-UFM) that integrates the proposed PanoGabor filters to unify other representations into the ERP format, delivering effective and distortion-aware features. Considering the orientation sensitivity of the Gabor transform, we further introduce a spherical gradient constraint to stabilize this sensitivity. Experimental results on three popular indoor 360 benchmarks demonstrate the superiority of the proposed PGFuse to existing state-of-the-art solutions. Code and models will be available at https://github.com/zhijieshen-bjtu/PGFuse. Zhijie Shen, Chunyu Lin, Lang Nie, Kang Liao, Weisi Lin, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | TransVFC: A transformable video feature compression framework for machines
Yuxiao Sun, Yao Zhao 0001, Meiqin Liu 0002, Huihui Bai 0001, Chunyu Lin, Weisi Lin |
Pattern Recognit. | 6 |
| 2026 | Toward Oriented Multi-Object Tracking for Fisheye Images: Dataset and FrameworkabstractMulti-object tracking (MOT) in fisheye images becomes particularly challenging due to significant radial distortion. In MOT, Camera Motion Compensation (CMC) is crucial for mitigating inter-frame errors caused by camera movement. However, conventional CMC methods degrade in fisheye images, as their rigid motion models cannot characterize the non-uniform motion fields introduced by severe distortion. In this paper, we first establish this limitation through theoretical derivation and then propose a plug-and-play ROI-Centric Motion Compensation (RCMC) mechanism, which serves as a new CMC solution that leverages instance-aware ROI for motion estimation. Moreover, fisheye distortion introduces significant target tilt, distorting object geometry and thereby hindering accurate detection and separation of adjacent instances. To address this issue, we integrate RCMC with tailored adjustments for oriented bounding boxes (OBB) to form FO-MOT, a MOT framework for fisheye images. In addition, to overcome the lack of dedicated evaluation benchmarks, we present FisheyeMOT, a comprehensively annotated dataset with OBB-based labels, designed to support the development and validation of fisheye MOT systems. Finally, we demonstrate the superiority of our framework by qualitative and quantitative experiments. The dataset and code will be released at https://github.com/lukanightfever/FO-MOT. Chunyu Lin, Lang Nie, Yepeng Tang, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Revisiting Monocular 3D Object Detection With Depth Thickness FieldabstractMonocular 3D object detection is challenging due to the lack of accurate depth. However, existing depth-assisted solutions still exhibit inferior performance, whose reason is universally acknowledged as the unsatisfactory accuracy of monocular depth estimation models. In this paper, we revisit monocular 3D object detection from the depth perspective and formulate an additional issue as the limited 3D structure-aware capability of existing depth representations (e.g., depth one-hot encoding or depth distribution). To address this issue, we introduce a novel Depth Thickness Field approach to embed clear 3D structures of the scenes. Specifically, we present MonoDTF, a scene-to-instance depth-adapted network comprising a Scene-Level Depth Retargeting (SDR) module and an Instance-Level Spatial Refinement (ISR) module. The former retargets traditional depth representations to the proposed depth thickness field, incorporating the scene-level perception of 3D structures. The latter refines the voxel space with the guidance of instances, enhancing the 3D instance-aware capability of the depth thickness field and thus improving detection accuracy. Extensive experiments on the KITTI and Waymo datasets demonstrate our superiority to existing state-of-the-art (SoTA) methods and the universality when equipped with different depth estimation models. The source codes are available at https://github.com/QiuDeZhang/MonoDTF. Qiude Zhang, Chunyu Lin, Zhijie Shen, Lang Nie, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Robust Image Stitching With Optimal PlaneabstractWe present RopStitch, an unsupervised deep image stitching framework with both robustness and naturalness. To ensure the robustness of RopStitch, we propose to incorporate the universal prior of content perception into the image stitching model by a dual-branch architecture. It separately captures coarse and fine features and integrates them to achieve highly generalizable performance across diverse unseen real-world scenes. Concretely, the dual-branch model consists of a pretrained branch to capture semantically invariant representations and a learnable branch to extract fine-grained discriminative features, which are then merged into a whole by a controllable factor at the correlation level. Besides, considering that content alignment and structural preservation are often contradictory to each other, we propose a concept of virtual optimal planes to relieve this conflict. To this end, we model this problem as a process of estimating homography decomposition coefficients, and design an iterative coefficient predictor and minimal semantic distortion constraint to identify the optimal plane. This scheme is finally incorporated into RopStitch by warping both views onto the optimal plane bidirectionally. Extensive experiments across various datasets demonstrate that RopStitch significantly outperforms existing methods, particularly in scene robustness and content naturalness. Lang Nie, Kang Liao, Yunqiu Xu, Chunyu Lin, Bin Xiao 0002 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | PixelStitch: Structure-Preserving Pixel-Wise Bidirectional Warps for Unsupervised Image Stitching
Hengzhe Jin, Lang Nie, Chunyu Lin, Xiaomei Feng, Yao Zhao 0001 |
ICCV | 3 |
| 2025 | What Really Matters for Learning-based LiDAR-Camera CalibrationabstractCalibration is an essential prerequisite for the accurate data fusion of LiDAR and camera sensors. Traditional calibration techniques often require specific targets or suitable scenes to obtain reliable 2D-3D correspondences. To tackle the challenge of target-less and online calibration, deep neural networks have been introduced to solve the problem in a data-driven manner. While previous learning-based methods have achieved impressive performance on specific datasets, they still struggle in complex real-world scenarios. Most existing works focus on improving calibration accuracy but overlook the underlying mechanisms. In this paper, we revisit the development of learning-based LiDAR-Camera calibration and encourage the community to pay more attention to the underlying principles to advance practical applications. We systematically analyze the paradigm of mainstream learning-based methods, and identify the critical limitations of regression-based methods with the widely used data generation pipeline. Our findings reveal that most learning-based methods inadvertently operate as retrieval networks, focusing more on single-modality distributions rather than cross-modality correspondences. We also investigate how the input data format and preprocessing operations impact network performance and summarize the regression clues to inform further improvements. Chunyu Lin, Yao Zhao 0001 |
MMAsia | 2 |
| 2025 | Jasmine: Harnessing Diffusion Prior for Self-supervised Depth EstimationabstractIn this paper, we propose \textbf{Jasmine}, the first Stable Diffusion (SD)-based self-supervised framework for monocular depth estimation, which effectively harnesses SD’s visual priors to enhance the sharpness and generalization of unsupervised prediction. Previous SD-based methods are all supervised since adapting diffusion models for dense prediction requires high-precision supervision. In contrast, self-supervised reprojection suffers from inherent challenges (\textit{e.g.}, occlusions, texture-less regions, illumination variance), and the predictions exhibit blurs and artifacts that severely compromise SD's latent priors. To resolve this, we construct a novel surrogate task of mix-batch image reconstruction. Without any additional supervision, it preserves the detail priors of SD models by reconstructing the images themselves while preventing depth estimation from degradation. Furthermore, to address the inherent misalignment between SD's scale and shift invariant estimation and self-supervised scale-invariant depth estimation, we build the Scale-Shift GRU. It not only bridges this distribution gap but also isolates the fine-grained texture of SD output against the interference of reprojection loss. Extensive experiments demonstrate that Jasmine achieves SoTA performance on the KITTI benchmark and exhibits superior zero-shot generalization across multiple datasets. Jiyuan Wang 0001, Chunyu Lin, Cheng Guan, Lang Nie, Kang Liao, Yao Zhao 0001 |
NeurIPS | 2 |
| 2025 | Unsupervised Global and Local Homography Estimation With Coplanarity-Aware GANabstractUnsupervised methods have received increasing attention in homography learning due to their promising performance and label-free training. However, existing methods do not explicitly consider the plane-induced parallax, making the prediction compromised on multiple planes. In this work, we propose a novel method HomoGAN to guide unsupervised homography estimation to focus on the dominant plane. First, a multi-scale transformer is designed to predict homography from the feature pyramids of input images in a coarse-to-fine fashion. Moreover, we propose an unsupervised GAN to impose coplanarity constraint on the predicted homography, which is realized by using a generator to predict a mask of aligned regions, and then a discriminator to check if two masked feature maps are induced by a single homography. Based on the global homography framework, we extend it to the local mesh-grid homography estimation, namely, MeshHomoGAN, where plane constraints can be enforced on each mesh cell to go beyond a single dominant plane, such that scenes with multiple depth planes can be better aligned. To validate the effectiveness of our method and its components, we conduct extensive experiments on large-scale datasets. Results show that our matching error is 22% lower than previous SOTA methods. Code is available at https://github.com/megvii-research/HomoGAN. Shuaicheng Liu, Mingbo Hong, Nianjin Ye, Chunyu Lin, Bing Zeng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | StabStitch++: Unsupervised Online Video Stitching With Spatiotemporal Bidirectional WarpsabstractWe retarget video stitching to an emerging issue, named warping shake, which unveils the temporal content shakes induced by sequentially unsmooth warps when extending image stitching to video stitching. Even if the input videos are stable, the stitched video can inevitably cause undesired warping shakes and affect the visual experience. To address this issue, we propose StabStitch++, a novel video stitching framework to realize spatial stitching and temporal stabilization with unsupervised learning simultaneously. First, different from existing learning-based image stitching solutions that typically warp one image to align with another, we suppose a virtual midplane between original image planes and project them onto it. Concretely, we design a differentiable bidirectional decomposition module to disentangle the homography transformation and incorporate it into our spatial warp, evenly spreading alignment burdens and projective distortions across two views. Then, inspired by camera paths in video stabilization, we derive the mathematical expression of stitching trajectories in video stitching by elaborately integrating spatial and temporal warps. Finally, a warp smoothing model is presented to produce stable stitched videos with a hybrid loss to simultaneously encourage content alignment, trajectory smoothness, and online collaboration. Compared with StabStitch that sacrifices alignment for stabilization, StabStitch++ makes no compromise and optimizes both of them simultaneously, especially in the online mode. To establish an evaluation benchmark and train the learning framework, we build a video stitching dataset with a rich diversity in camera motions and scenes. Experiments exhibit that StabStitch++ surpasses current solutions in stitching performance, robustness, and efficiency, offering compelling advancements in this field by building a real-time online video stitching system. Lang Nie, Chunyu Lin, Kang Liao, Yun Zhang 0024, Shuaicheng Liu, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | SGFormer: Spherical Geometry Transformer for 360° Depth EstimationabstractPanoramic distortion poses a significant challenge in 360° depth estimation, particularly pronounced at the north and south poles. Existing methods either adopt a bi-projection fusion strategy to remove distortions or model long-range dependencies to capture global structures, resulting in either unclear structure or insufficient local perception. In this paper, we propose a spherical geometry transformer, named SGFormer, to address the above issues, with an innovative step to integrate spherical geometric priors into vision transformers. To this end, we retarget the transformer decoder to a spherical prior decoder (termed SPDecoder), which endeavors to uphold the integrity of spherical structures during decoding. Concretely, we leverage bipolar reprojection, circular rotation, and curve local embedding to preserve the spherical characteristics of equidistortion, continuity, and surface distance, respectively. Furthermore, we present a query-based global conditional position embedding to compensate for spatial structure at varying resolutions. It not only boosts the global perception of spatial position but also sharpens the depth structure across different patches. Finally, we conduct extensive experiments on popular benchmarks, demonstrating our superiority over state-of-the-art solutions. Our code will be made publicly athttps://github.com/iuiuJaon/SGFormer. Junsong Zhang, Zisong Chen, Chunyu Lin, Zhijie Shen, Lang Nie, Kang Liao, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Ada3DLane: Adaptive 3D Lane Detection From Monocular Imagesabstract3D lane detection is a fundamental task in autonomous driving, and detecting 3D lanes from monocular images has been widely adopted due to the low computational cost and the property of lanes. Recent progresses have been made based on surrogate representations such as bird’s eye view (BEV) features. However, monocular BEV construction strictly relies on flat groud assumption, and the misalignment between perspective view and BEV is inevitable. In this paper, we propose Ada3DLane, a BEV-free, query based 3D lane detector, which adaptively generates queries with rich semantic and geometric information as well as lane interactions; adaptively samples perspective view image features in spatial and temporal domain; adaptively decode the sampled features for fast and accurate 3D lane detection. We conduct extensive experiments on benchmark lane detection datasets and outperforms previous state-of-the-art methods. Zhiming Hou, Yuanzhouhan Cao, Naiyue Chen, Chao Ren 0002, Chunyu Lin, Congyan Lang, Yidong Li |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Advancing Real-World Parking Slot Detection With Large-Scale Dataset and Semi-Supervised Baseline
Chunyu Lin, Lang Nie, Jiyuan Wang 0001, Yao Zhao 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | FishFormer: Annulus Slicing-based Transformer for Fisheye RectificationabstractNumerous significant progress on fisheye image rectification has been achieved through CNN. Nevertheless, constrained by a fixed receptive field, the global distribution and the local symmetry of the distortion have not been fully exploited. To leverage these two characteristics, we introduce FishFormer that processes the fisheye image as a sequence to enhance global and local perception. We tune the Transformer according to the structural properties of fisheye images. First, the uneven distortion distribution in patches generated by the existing square slicing method hinders the understanding of the global structure. Therefore, we propose an annulus slicing method to maintain the consistency of the distortion in each patch; thus, the applicability of the Transformer is expanded to perceive the distortion distribution efficiently. Second, the distortion of adjacent patches is progressive. Such explicit correlations in local regions need to be rapidly constructed and maintained, but Transformer has a weakness in local area perception. Hence, a novel layer attention mechanism is introduced to enhance the local perception and feature interaction. Our network simultaneously implements global perception and focused local perception. Extensive experiments demonstrate that our method provides superior performance compared with state-of-the-art methods. Shangrong Yang, Chunyu Lin, Kang Liao, Yao Zhao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Toward Oriented Fisheye Object Detection: Dataset and BaselineabstractFisheye object detection is challenging due to the fisheye distortion, which inclines objects to different extents and pushes extensive irrelevant pixels into the predicted horizontal bounding box (HBB). To address the problems above, we establish a new fisheye object detection dataset (named FishOBB) with compact oriented bounding box (OBB) annotations, as well as an OBB-customized mosaic augmentation technology. To our knowledge, there are very few fisheye datasets labeled by OBB, especially the open source forword view dataset like ours. Besides, we provide a fisheye object detection baseline (named FDA-YOLO) with two fisheye adaption units. Concretely, we first design a distortion orientation aggregation (DOA) unit guided by polar sampling to capture distortion-aware fisheye features. On the other hand, to transfer HBB-based detection models to OBB-based counterparts, we propose an oriented anchor attention unit. It automatically weights the unbalanced positive/negative samples and facilitates convergence for multi-anchor models. Finally, we demonstrate that the two adaption units can be easily integrated into various anchor-based YOLO methods, e.g., ScaledYOLOv4 and YOLOv7, contributing to superior performance to existing state-of-the-art (SoTA) solutions in the proposed dataset. Meanwhile, our method has also achieved SoTA performance on other popular datasets like WEPDTOF. The dataset and code are released at https://github.com/lukanightfever/FishOBB . Chunyu Lin, Lang Nie, Zisen Kong, Jiapeng Wang 0002, Yao Zhao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | CurrI2P: inter- and intra-modality similarity curriculum learning for image-to-point cloud registration
Chunyu Lin, Lang Nie, Yao Zhao 0001 |
Vis. Comput. | 2 |
| 2024 | Efficient Meshflow and Optical Flow Estimation from Event CamerasabstractIn this paper, we explore the problem of event-based meshflow estimation, a novel task that involves predicting a spatially smooth sparse motion field from event cameras. To start, we generate a large-scale High-Resolution Event Meshflow (HREM) dataset, which showcases its superiority by encompassing the merits of high resolution at 1280×720, handling dynamic objects and complex motion patterns, and offering both optical flow and meshflow labels. These aspects have not been fully explored in previous works. Besides, we propose Efficient Event-based MeshFlow (EEMFlow) network, a lightweight model featuring a specially crafted encoder-decoder architecture to facilitate swift and accurate meshflow estimation. Furthermore, we upgrade EEMFlow network to support dense event optical flow, in which a Confidence-induced Detail Completion (CDC) module is proposed to preserve sharp motion boundaries. We conduct comprehensive experiments to show the exceptional performance and runtime efficiency (39× faster) of our EEMFlow model compared to recent state-of-the-art flow methods. Our code is available at https://github.com/boomluo02/EEMFlow. Xinglong Luo, Ao Luo, Zhengning Wang, Chunyu Lin, Bing Zeng 0001, Shuaicheng Liu |
CVPR | 4 |
| 2024 | Eliminating Warping Shakes for Unsupervised Online Video Stitching
Lang Nie, Chunyu Lin, Kang Liao, Yun Zhang 0024, Shuaicheng Liu, Rui Ai 0001, Yao Zhao 0001 |
ECCV (4) | 2 |
| 2024 | WeatherDepth: Curriculum Contrastive Learning for Self-Supervised Depth Estimation under Adverse Weather ConditionsabstractDepth estimation models have shown promising performance on clear scenes but fail to generalize to adverse weather conditions due to illumination variations, weather particles, etc. In this paper, we propose WeatherDepth, a self-supervised robust depth estimation model with curriculum contrastive learning, to tackle performance degradation in complex weather conditions. Concretely, we first present a progressive curriculum learning scheme with three simple-to-complex curricula to gradually adapt the model from clear to relative adverse, and then to adverse weather scenes. It encourages the model to gradually grasp beneficial depth cues against the weather effect, yielding smoother and better domain adaption. Meanwhile, to prevent the model from forgetting previous curricula, we integrate contrastive learning into different curricula. By drawing reference knowledge from the previous course, our strategy establishes a depth consistency constraint between different courses toward robust depth estimation in diverse weather. Besides, to reduce manual intervention and better adapt to different models, we designed an adaptive curriculum scheduler to automatically search for the best timing for course switching. In the experiment, the proposed solution is proven to be easily incorporated into various architectures and demonstrates state-of-the-art (SoTA) performance on both synthetic and real weather datasets. Source code and data are available at https://github.com/wangjiyuan9/WeatherDepth. Jiyuan Wang 0001, Chunyu Lin, Lang Nie, Shujun Huang, Yao Zhao 0001, Xing Pan, Rui Ai 0001 |
ICRA | 2 |
| 2024 | Digging into Contrastive Learning for Robust Depth Estimation with Diffusion ModelsabstractRecently, diffusion-based depth estimation methods have drawn widespread attention due to their elegant denoising patterns and promising performance. However, they are typically unreliable under adverse conditions prevalent in real-world scenarios, such as rainy, snowy, etc. In this paper, we propose a novel robust depth estimation method called D4RD, featuring a custom contrastive learning mode tailored for diffusion models to mitigate performance degradation in complex environments. Concretely, we integrate the strength of knowledge distillation into contrastive learning, building the `trinity' contrastive scheme. This scheme utilizes the sampled noise of the forward diffusion process as a natural reference, guiding the predicted noise in diverse scenes toward a more stable and precise optimum. Moreover, we extend noise-level trinity to encompass more generic feature and image levels, establishing a multi-level contrast to distribute the burden of robust perception across the overall network. Before addressing complex scenarios, we enhance the stability of the baseline diffusion model with three straightforward yet effective improvements, which facilitate convergence and remove depth outliers. Extensive experiments demonstrate that D4RD surpasses existing state-of-the-art solutions on synthetic corruption datasets and real-world weather conditions. Source code and data are available at \url{https://github.com/wangjiyuan9/D4RD}. Jiyuan Wang 0001, Chunyu Lin, Lang Nie, Kang Liao, Shuwei Shao, Yao Zhao 0001 |
ACM Multimedia | 2 |
| 2024 | Multimodal spatiotemporal aggregation for point cloud accumulation
Chunyu Lin, Lang Nie, Meiqin Liu 0002, Yao Zhao 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | Semi-Supervised Coupled Thin-Plate Spline Model for Rotation Correction and BeyondabstractThin-plate spline (TPS) is a principal warp that allows for representing elastic, nonlinear transformation with control point motions. With the increase of control points, the warp becomes increasingly flexible but usually encounters a bottleneck caused by undesired issues, e.g., content distortion. In this paper, we explore generic applications of TPS in single-image-based warping tasks, such as rotation correction, rectangling, and portrait correction. To break this bottleneck, we propose the coupled thin-plate spline model (CoupledTPS), which iteratively couples multiple TPS with limited control points into a more flexible and powerful transformation. Concretely, we first design an iterative search to predict new control points according to the current latent condition. Then, we present the warping flow as a bridge for the coupling of different TPS transformations, effectively eliminating interpolation errors caused by multiple warps. Besides, in light of the laborious annotation cost, we develop a semi-supervised learning scheme to improve warping quality by exploiting unlabeled data. It is formulated through dual transformation between the searched control points of unlabeled data and its graphic augmentation, yielding an implicit correction consistency constraint. Finally, we collect massive unlabeled data to exhibit the benefit of our semi-supervised scheme in rotation correction. Extensive experiments demonstrate the superiority and universality of CoupledTPS over the existing State-of-the-Art (SoTA) solutions for rotation correction and beyond. Lang Nie, Chunyu Lin, Kang Liao, Shuaicheng Liu, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | 360 Layout Estimation via Orthogonal Planes Disentanglement and Multi-View Geometric Consistency PerceptionabstractExisting panoramic layout estimation solutions tend to recover room boundaries from a vertically compressed sequence, yielding imprecise results as the compression process often muddles the semantics between various planes. Besides, these data-driven approaches impose an urgent demand for massive data annotations, which are laborious and time-consuming. For the first problem, we propose an orthogonal plane disentanglement network (termed DOPNet) to distinguish ambiguous semantics. DOPNet consists of three modules that are integrated to deliver distortion-free, semantics-clean, and detail-sharp disentangled representations, which benefit the subsequent layout recovery. For the second problem, we present an unsupervised adaptation technique tailored for horizon-depth and ratio representations. Concretely, we introduce an optimization strategy for decision-level layout analysis and a 1D cost volume construction method for feature-level multi-view aggregation, both of which are designed to fully exploit the geometric consistency across multiple perspectives. The optimizer provides a reliable set of pseudo-labels for network training, while the 1D cost volume enriches each view with comprehensive scene information derived from other perspectives. Extensive experiments demonstrate that our solution outperforms other SoTA models on both monocular layout estimation and multi-view layout estimation tasks. Zhijie Shen, Chunyu Lin, Junsong Zhang, Lang Nie, Kang Liao, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Cylin-Painting: Seamless 360° Panoramic Image Outpainting and BeyondabstractImage outpainting gains increasing attention since it can generate the complete scene from a partial view, providing a valuable solution to construct 360° panoramic images. As image outpainting suffers from the intrinsic issue of unidirectional completion flow, previous methods convert the original problem into inpainting, which allows a bidirectional flow. However, we find that inpainting has its own limitations and is inferior to outpainting in certain situations. The question of how they may be combined for the best of both has as yet remained under-explored. In this paper, we provide a deep analysis of the differences between inpainting and outpainting, which essentially depends on how the source pixels contribute to the unknown regions under different spatial arrangements. Motivated by this analysis, we present a Cylin-Painting framework that involves meaningful collaborations between inpainting and outpainting and efficiently fuses the different arrangements, with a view to leveraging their complementary benefits on a seamless cylinder. Nevertheless, straightforwardly applying the cylinder-style convolution often generates visually unpleasing results as it discards important positional information. To address this issue, we further present a learnable positional embedding strategy to incorporate the missing component of positional encoding into the cylinder convolution, which significantly improves the panoramic results. It is noted that while developed for image outpainting, the proposed algorithm can be effectively extended to other panoramic vision tasks, such as object detection, depth estimation, and image super-resolution. Code will be made available at https://github.com/KangLiao929/Cylin-Painting. Kang Liao, Xiangyu Xu 0002, Chunyu Lin, Wenqi Ren, Yunchao Wei, Yao Zhao 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Camera calibration for the surround-view system: a benchmark and dataset
Leidong Qin, Chunyu Lin, Shangrong Yang, Yao Zhao 0001 |
Vis. Comput. | 2 |
| 2023 | Spatiotemporal Deformation Perception for Fisheye Video RectificationabstractAlthough the distortion correction of fisheye images has been extensively studied, the correction of fisheye videos is still an elusive challenge. For different frames of the fisheye video, the existing image correction methods ignore the correlation of sequences, resulting in temporal jitter in the corrected video. To solve this problem, we propose a temporal weighting scheme to get a plausible global optical flow, which mitigates the jitter effect by progressively reducing the weight of frames. Subsequently, we observe that the inter-frame optical flow of the video is facilitated to perceive the local spatial deformation of the fisheye video. Therefore, we derive the spatial deformation through the flows of fisheye and distorted-free videos, thereby enhancing the local accuracy of the predicted result. However, the independent correction for each frame disrupts the temporal correlation. Due to the property of fisheye video, a distorted moving object may be able to find its distorted-free pattern at another moment. To this end, a temporal deformation aggregator is designed to reconstruct the deformation correlation between frames and provide a reliable global feature. Our method achieves an end-to-end correction and demonstrates superiority in correction quality and stability compared with the SOTA correction methods. Shangrong Yang, Chunyu Lin, Kang Liao, Yao Zhao 0001 |
AAAI | 2 |
| 2023 | Disentangling Orthogonal Planes for Indoor Panoramic Room Layout Estimation with Cross-Scale Distortion AwarenessabstractBased on the Manhattan World assumption, most existing indoor layout estimation schemes focus on recovering layouts from vertically compressed 1D sequences. However, the compression procedure confuses the semantics of different planes, yielding inferior performance with ambiguous interpretability. To address this issue, we propose to disentangle this 1D representation by pre-segmenting orthogonal (vertical and horizontal) planes from a complex scene, explicitly capturing the geometric cues for indoor layout estimation. Considering the symmetry between the floor boundary and ceiling boundary, we also design a soft-flipping fusion strategy to assist the pre-segmentation. Besides, we present a feature assembling mechanism to effectively integrate shallow and deep features with distortion distribution awareness. To compensate for the potential errors in pre-segmentation, we further leverage triple attention to reconstruct the disentangled sequences for better performance. Experiments on four popular benchmarks demonstrate our superiority over existing SoTA solutions, especially on the 3DIoU metric. The code is available at https://github.com/zhijieshen-bjtu/DOPNet. Zhijie Shen, Zishuo Zheng, Chunyu Lin, Lang Nie, Kang Liao, Shuai Zheng 0005, Yao Zhao 0001 |
CVPR | 3 |
| 2023 | SIGVIC: Spatial Importance Guided Variable-Rate Image CompressionabstractVariable-rate mechanism has improved the flexibility and efficiency of learning-based image compression that trains multiple models for different rate-distortion tradeoffs. One of the most common approaches for variable-rate is to channel- wisely or spatial-uniformly scale the internal features. However, the diversity of spatial importance is instructive for bit allocation of image compression. In this paper, we introduce a Spatial Importance Guided Variable-rate Image Compression (SigVIC), in which a spatial gating unit (SGU) is designed for adaptively learning a spatial importance mask. Then, a spatial scaling network (SSN) takes the spatial importance mask to guide the feature scaling and bit allocation for variablerate. Moreover, to improve the quality of decoded image, Top-K shallow features are selected to refine the decoded features through a shallow feature fusion module (SFFM). Experiments show that our method outperforms other learning- based methods (whether variable-rate or not) and traditional codecs, with storage saving and high flexibility. Meiqin Liu 0002, Chunyu Lin, Yao Zhao 0001 |
ICASSP | 4 |
| 2023 | Towards Reliable Image Outpainting: Learning Structure-Aware Multimodal Fusion with Depth GuidanceabstractImage outpainting technology generates visually plausible content regardless of authenticity, making it unreliable to be applied in practice. Thus, we propose a reliable image outpainting task, introducing the sparse depth from LiDARs (Light Detection And Ranging devices) to extrapolate authentic RGB scenes. The large field view of LiDARs allows it to serve for data enhancement and further multimodal tasks. Concretely, we propose a Depth-Guided Outpainting Network to model different feature representations of two modalities and learn the structure-aware cross-modal fusion. And two components are designed: 1) The Multimodal Learning Module produces unique depth and RGB feature representations from the perspectives of different modal characteristics. 2) The Depth Guidance Fusion Module leverages the complete depth modality to guide the establishment of RGB contents by progressive multimodal feature fusion. Furthermore, we specially design an additional constraint strategy consisting of Cross-modal Loss and Edge Loss to enhance ambiguous contours and expedite reliable content generation. Extensive experiments on KITTI and Waymo datasets demonstrate our superiority over the state-of-the-art method, quantitatively and qualitatively. Lei Zhang 0116, Chunyu Lin, Kang Liao, Yao Zhao 0001 |
ICASSP | 2 |
| 2023 | RecRecNet: Rectangling Rectified Wide-Angle Images by Thin-Plate Spline Model and DoF-based Curriculum LearningabstractThe wide-angle lens shows appealing applications in VR technologies, but it introduces severe radial distortion into its captured image. To recover the realistic scene, previous works devote to rectifying the content of the wide-angle image. However, such a rectification solution inevitably distorts the image boundary, which changes related geometric distributions and misleads the current vision perception models. In this work, we explore constructing a win-win representation on both content and boundary by contributing a new learning model, i.e., Rectangling Rectification Network (RecRecNet). In particular, we propose a thin-plate spline (TPS) module to formulate the nonlinear and non-rigid transformation for rectangling images. By learning the control points on the rectified image, our model can flexibly warp the source structure to the target domain and achieves an end-to-end unsupervised deformation. To relieve the complexity of structure approximation, we then inspire our RecRecNet to learn the gradual deformation rules with a DoF (Degree of Freedom)-based curriculum learning. By increasing the DoF in each curriculum stage, namely, from similarity transformation (4-DoF) to homography transformation (8-DoF), the network is capable of investigating more detailed deformations, offering fast convergence on the final rectangling task. Experiments show the superiority of our solution over the compared methods on both quantitative and qualitative evaluations. The code and dataset are available at https://github.com/KangLiao929/RecRecNet. Kang Liao, Lang Nie, Chunyu Lin, Zishuo Zheng, Yao Zhao 0001 |
ICCV | 3 |
| 2023 | GAFlow: Incorporating Gaussian Attention into Optical FlowabstractOptical flow, or the estimation of motion fields from image sequences, is one of the fundamental problems in computer vision. Unlike most pixel-wise tasks that aim at achieving consistent representations of the same category, optical flow raises extra demands for obtaining local discrimination and smoothness, which yet is not fully explored by existing approaches. In this paper, we push Gaussian Attention (GA) into the optical flow models to accentuate local properties during representation learning and enforce the motion affinity during matching. Specifically, we introduce a novel Gaussian-Constrained Layer (GCL) which can be easily plugged into existing Transformer blocks to highlight the local neighborhood that contains fine-grained structural information. Moreover, for reliable motion analysis, we provide a new Gaussian-Guided Attention Module (GGAM) which not only inherits properties from Gaussian distribution to instinctively revolve around the neighbor fields of each point but also is empowered to put the emphasis on contextually related regions during matching. Our fully-equipped model, namely Gaussian Attention Flow network (GAFlow), naturally incorporates a series of novel Gaussian-based modules into the conventional optical flow framework for reliable motion analysis. Extensive experiments on standard optical flow datasets consistently demonstrate the exceptional performance of the proposed approach in terms of both generalization ability evaluation and online benchmark testing. Code is available at https://github.com/LA30/GAFlow. Ao Luo, Fan Yang 0054, Xin Li 0005, Lang Nie, Chunyu Lin, Haoqiang Fan, Shuaicheng Liu |
ICCV | 5 |
| 2023 | Parallax-Tolerant Unsupervised Deep Image StitchingabstractTraditional image stitching approaches tend to leverage increasingly complex geometric features (e.g., point, line, edge, etc.) for better performance. However, these hand-crafted features are only suitable for specific natural scenes with adequate geometric structures. In contrast, deep stitching schemes overcome adverse conditions by adaptively learning robust semantic features, but they cannot handle large-parallax cases.To solve these issues, we propose a parallax-tolerant unsupervised deep image stitching technique. First, we propose a robust and flexible warp to model the image registration from global homography to local thin-plate spline motion. It provides accurate alignment for overlapping regions and shape preservation for non-overlapping regions by joint optimization concerning alignment and distortion. Subsequently, to improve the generalization capability, we design a simple but effective iterative strategy to enhance the warp adaption in cross-dataset and cross-resolution applications. Finally, to further eliminate the parallax artifacts, we propose to composite the stitched image seamlessly by unsupervised learning for seam-driven composition masks. Compared with existing methods, our solution is parallax-tolerant and free from laborious designs of complicated geometric features for specific scenes. Extensive experiments show our superiority over the SoTA methods, both quantitatively and qualitatively. The code is available at https://github.com/nie-lang/UDIS2. Lang Nie, Chunyu Lin, Kang Liao, Shuaicheng Liu, Yao Zhao 0001 |
ICCV | 2 |
| 2023 | Innovating Real Fisheye Image Correction with Dual Diffusion ArchitectureabstractFisheye image rectification is hindered by synthetic models producing poor results for real-world correction. To address this, we propose a Dual Diffusion Architecture (DDA) for fisheye rectification that offers better practicality. The DDA leverages Denoising Diffusion Probabilistic Models (DDPMs) to gradually introduce bidirectional noise, allowing the synthesized and real images to develop into a consistent noise distribution. As a result, our network can perceive the distribution of unlabelled real fisheye images without relying on a transfer network, thus improving the performance of real fisheye correction. Additionally, we design an unsupervised one-pass network that generates a plausible new condition to strengthen guidance and address the non-negligible indeterminacy between the prior condition and the target. It can significantly affect the rectification task, especially in cases where radial distortion causes significant artifacts. This network can be regarded as an alternate scheme for fast producing reliable results without iterative inference. Compared to the state-of-the-art methods, our approach achieves superior performance in both synthetic and real fisheye image corrections. Shangrong Yang, Chunyu Lin, Kang Liao, Yao Zhao 0001 |
ICCV | 2 |
| 2023 | Unsupervised OmniMVS: Efficient Omnidirectional Depth Inference via Establishing Pseudo-Stereo SupervisionabstractOmnidirectional multi-view stereo (MVS) vision is attractive for its ultra-wide field-of-view (FoV), enabling machines to perceive 360°3D surroundings. However, the existing solutions require expensive dense depth labels for supervision, making them impractical in real-world applications. In this paper, we propose the first unsupervised omnidirectional MVS framework based on multiple fisheye images. To this end, we project all images to a virtual view center and composite two panoramic images with spherical geometry from two pairs of back-to-back fisheye images. The two 360° images formulate a stereo pair with a special pose, and the photometric consistency is leveraged to establish the unsupervised constraint, which we term “Pseudo-Stereo Supervision”. In addition, we propose Un-OmniMVS, an efficient unsupervised omnidirectional MVS network, to facilitate the inference speed with two efficient components. First, a novel feature extractor with frequency attention is proposed to simultaneously capture the non-local Fourier features and local spatial features, explicitly facilitating the feature representation. Then, a variance-based light cost volume is put forward to reduce the computational complexity. Experiments exhibit that the performance of our unsupervised solution is competitive to that of the state-of-the-art (SoTA) supervised methods with better generalization in real-world data. The code will be available at https://github.com/Chen-z-s/Un-OmniMVS. Zisong Chen, Chunyu Lin, Lang Nie, Kang Liao, Yao Zhao 0001 |
IROS | 2 |
| 2023 | S-OmniMVS: Incorporating Sphere Geometry into Omnidirectional Stereo MatchingabstractMulti-fisheye stereo matching is a promising task that employs the traditional multi-view stereo (MVS) pipeline with spherical sweeping to acquire omnidirectional depth. However, the existing omnidirectional MVS technologies neglect fisheye and omnidirectional distortions, yielding inferior performance. In this paper, we revisit omnidirectional MVS by incorporating three sphere geometry priors: spherical projection, spherical continuity, and spherical position. To deal with fisheye distortion, we propose a new distortion-adaptive fusion module to convert fisheye inputs into distortion-free spherical tangent representations by constructing a spherical projection space. Then these multi-scale features are adaptively aggregated with additional learnable offsets to enhance content perception. To handle omnidirectional distortion, we present a new spherical cost aggregation module with a comprehensive consideration of the spherical continuity and position. Concretely, we first design a rotation continuity compensation mechanism to ensure omnidirectional depth consistency of left-right boundaries without introducing extra computation. On the other hand, we encode the geometry-aware spherical position and push them into the cost aggregation to relieve panoramic distortion and perceive the 3D structure. Furthermore, to avoid the excessive concentration of depth hypothesis caused by inverse depth linear sampling, we develop a segmented sampling strategy that combines linear and exponential spaces to create S-OmniMVS, along with three sphere priors. Extensive experiments demonstrate the proposed method outperforms the state-of-the-art (SoTA) solutions by a large margin on various datasets both quantitatively and qualitatively. Zisong Chen, Chunyu Lin, Lang Nie, Zhijie Shen, Kang Liao, Yuanzhouhan Cao, Yao Zhao 0001 |
ACM Multimedia | 2 |
| 2023 | Kernel Dimension Matters: To Activate Available Kernels for Real-time Video Super-ResolutionabstractReal-time video super-resolution requires low latency with high-quality reconstruction. Existing methods mostly use pruning schemes or neglect complicated modules to reduce the calculation complexity. However, the video contains large amounts of temporal redundancies due to the inter-frame correlation, which is rarely investigated in existing methods. The static and dynamic information lies in feature maps and represents the redundant complements and temporal offsets respectively. It is crucial to split channels with dynamic and static information for efficient processing. Thus, this paper proposes a kernel-split strategy to activate available kernels for real-time inference. This strategy focuses on the dimensions of convolutional kernels, including the channel and depth dimensions. Available kernel dimensions are activated according to the split of high-value and low-value channels. Specifically, a multi-channel selection unit is designed to discriminate the importance of channels and filter the high-value channels hierarchically. At each hierarchy, low-dimensional convolutional kernels are activated to reuse the low-value channel and re-parameterized convolutional kernels are employed on the high-value channel to merge the depth dimension. In addition, we design a multiple flow deformable alignment module for a sufficient temporal representation with affordable calculation cost. Experimental results demonstrate that our method outperforms other state-of-the-art (SOTA) ones in terms of reconstruction quality and runtime. Codes will be available at https://github.com/Kimsure/KSNet. Meiqin Liu 0002, Chunyu Lin, Yao Zhao 0001 |
ACM Multimedia | 4 |
| 2023 | Complementary Bi-directional Feature Compression for Indoor 360° Semantic Segmentation with Self-distillationabstractSemantic segmentation on 360° images is a vital component of scene understanding due to the rich surrounding information. Recently, horizontal representation-based approaches outperform projection-based solutions, because the distortions can be effectively removed by compressing the spherical data in the vertical direction. However, these methods ignore the distortion distribution prior and are limited to unbalanced receptive fields, e.g., the receptive fields are sufficient in the vertical direction and insufficient in the horizontal direction. Differently, a vertical representation compressed in another direction can offer implicit distortion prior and enlarge horizontal receptive fields. In this paper, we combine the two different representations and propose a novel 360° semantic segmentation solution from a complementary perspective. Our network comprises three modules: a feature extraction module, a bi-directional compression module, and an ensemble decoding module. First, we extract multi-scale features from a panorama. Then, a bi-directional compression module is designed to compress features into two complementary low-dimensional representations, which provide content perception and distortion prior. Furthermore, to facilitate the fusion of bi-directional features, we design a unique self distillation strategy in the ensemble decoding module to enhance the interaction of different features and further improve the performance. Experimental results show that our approach outperforms the state-of-the-art solutions on quantitative evaluations while displaying the best performance on visual appearance. Zishuo Zheng, Chunyu Lin, Lang Nie, Kang Liao, Zhijie Shen, Yao Zhao 0001 |
WACV | 2 |
| 2023 | Camouflaged object detection based on context-aware and boundary refinement
Caijuan Shi, Bijuan Ren, Houru Chen, Chunyu Lin, Yao Zhao 0001 |
Appl. Intell. | 5 |
| 2023 | MSPNet: Multi-stage progressive network for image denoising
Meiqin Liu 0002, Chunyu Lin, Yao Zhao 0001 |
Neurocomputing | 4 |
| 2023 | Temporal Consistency Learning of Inter-Frames for Video Super-ResolutionabstractVideo super-resolution (VSR) is a task that aims to reconstruct high-resolution (HR) frames from the low-resolution (LR) reference frame and multiple neighboring frames. The vital operation is to utilize the relative misaligned frames for the current frame reconstruction and preserve the consistency of the results. Existing methods generally explore information propagation and frame alignment to improve the performance of VSR. However, few studies focus on the temporal consistency of inter-frames. In this paper, we propose a Temporal Consistency learning Network (TCNet) for VSR in an end-to-end manner, to enhance the consistency of the reconstructed videos. A spatio-temporal stability module is designed to learn the self-alignment from inter-frames. Especially, the correlative matching is employed to exploit the spatial dependency from each frame to maintain structural stability. Moreover, a self-attention mechanism is utilized to learn the temporal correspondence to implement an adaptive warping operation for temporal consistency among multi-frames. Besides, a hybrid recurrent architecture is designed to leverage short-term and long-term information. We further present a progressive fusion module to perform a multistage fusion of spatio-temporal features. And the final reconstructed frames are refined by these fused features. Objective and subjective results of various experiments demonstrate that TCNet has superior performance on different benchmark datasets, compared to several state-of-the-art methods. Meiqin Liu 0002, Chunyu Lin, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | As-Deformable-As-Possible Single-Image-Based View Synthesis Without Depth PriorabstractDepth-image-based rendering (DIBR) technologies have been widely employed to synthesize novel realistic views from a single image in 3D video applications. However, DIBR-oriented approaches heavily rely on the accuracy of depth maps, usually requiring the depth GT as a prior. Despite that, there might exist extensive float precision losses and invalid holes in the synthesized view due to warping error and occlusion. In this paper, we propose an end-to-end as-deformable-as-possible (ADAP) single-image-based view synthesis solution without depth prior. It addresses the above issues through two stages: alignment and reconstruction, where we first transform the input image to the latent feature space and then reconstruct the novel view in the image domain. In the first stage, the input image is deformed to align with the synthesized view at feature level. To this end, we propose an ADAP alignment mechanism through pixel-level warping to error-level quantization to feature-level alignment, progressively improving the deformable capability in handling challenging motion conditions in real-world scenes. In the second stage, we exploit an occlusion-aware reconstruction module to recover the content details from the deformed feature at pixel level. Extensive experiments demonstrate that our alignment-reconstruction approach is robust to the depth map. Even with a coarsely estimated depth map, our solution outperforms other SoTA schemes in the popular benchmarks. Chunlan Zhang, Chunyu Lin, Kang Liao, Lang Nie, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | JNMR: Joint Non-Linear Motion Regression for Video Frame InterpolationabstractVideo frame interpolation (VFI) aims to generate predictive frames by motion-warping from bidirectional references. Most examples of VFI utilize spatiotemporal semantic information to realize motion estimation and interpolation. However, due to variable acceleration, irregular movement trajectories, and camera movement in real-world cases, they can not be sufficient to deal with non-linear middle frame estimation. In this paper, we present a reformulation of the VFI as a joint non-linear motion regression (JNMR) strategy to model the complicated inter-frame motions. Specifically, the motion trajectory between the target frame and multiple reference frames is regressed by a temporal concatenation of multi-stage quadratic models. Then, a comprehensive joint distribution is constructed to connect all temporal motions. Moreover, to reserve more contextual details for joint regression, the feature learning network is devised to explore clarified feature expressions with dense skip-connection. Later, a coarse-to-fine synthesis enhancement module is utilized to learn visual dynamics at different resolutions with multi-scale textures. The experimental VFI results show the effectiveness and significant improvement of joint motion regression over the state-of-the-art methods. The code is available at https://github.com/ruhig6/JNMR. Meiqin Liu 0002, Chunyu Lin, Yao Zhao 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | Deep Rotation Correction Without Angle PriorabstractNot everybody can be equipped with professional photography skills and sufficient shooting time, and there can be some tilts in the captured images occasionally. In this paper, we propose a new and practical task, named Rotation Correction, to automatically correct the tilt with high content fidelity in the condition that the rotated angle is unknown. This task can be easily integrated into image editing applications, allowing users to correct the rotated images without any manual operations. To this end, we leverage a neural network to predict the optical flows that can warp the tilted images to be perceptually horizontal. Nevertheless, the pixel-wise optical flow estimation from a single image is severely unstable, especially in large-angle tilted images. To enhance its robustness, we propose a simple but effective prediction strategy to form a robust elastic warp. Particularly, we first regress the mesh deformation that can be transformed into robust initial optical flows. Then we estimate residual optical flows to facilitate our network the flexibility of pixel-wise deformation, further correcting the details of the tilted images. To establish an evaluation benchmark and train the learning framework, a comprehensive rotation correction dataset is presented with a large diversity in scenes and rotated angles. Extensive experiments demonstrate that even in the absence of the angle prior, our algorithm can outperform other state-of-the-art solutions requiring this prior. The code and dataset are available at https://github.com/nie-lang/RotationCorrection. Lang Nie, Chunyu Lin, Kang Liao, Shuaicheng Liu, Yao Zhao 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Monocular Pseudo-LiDAR Point Cloud Extrapolation Based on Iterative Hybrid RenderingabstractRecently, a pseudo-LiDAR point cloud extrapolation algorithm equipped with stereo cameras has been introduced, bridging the gap between the expensive 3D sensor LiDAR and relatively cheap 2D sensor camera in autonomous driving. In this paper, we explore an approach to further bridge this gap using only a monocular camera and extrapolate a wide field of view 3D point cloud from a limited 2D view. However, this task is extremely challenging as it requires inferring the occluded contents in the scene. To this end, we propose a ‘render-refine-iterate-fuse’ framework that takes advantage of both image view synthesis and image inpainting techniques, guiding the neural network to learn the potential spatial distribution. In addition, we design a hybrid rendering scheme to ensure that the visible content moves in a geometrically correct manner and fills the pixels caused by occlusion. Benefitting from the proposed framework, our approach achieves significant improvements on the pseudo-LiDAR point cloud extrapolation task. The gap between LiDAR and cameras is further bridged, showing an economical and practical application in the environment perception module of autonomous driving. The experimental results evaluated on the KITTI dataset demonstrate that our approach achieves superior quantitative and qualitative performance. Chunlan Zhang, Chunyu Lin, Kang Liao, Yao Zhao 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Unsupervised Homography Estimation with Coplanarity-Aware GANabstractEstimating homography from an image pair is a fundamental problem in image alignment. Unsupervised learning methods have received increasing attention in this field due to their promising performance and label-free training. However, existing methods do not explicitly consider the problem of plane-induced parallax, which will make the predicted homography compromised on multiple planes. In this work, we propose a novel method HomoGAN to guide unsupervised homography estimation to focus on the dominant plane. First, a multi-scale transformer network is designed to predict homography from the feature pyramids of input images in a coarse-to-fine fashion. Moreover, we propose an unsupervised GAN to impose coplanarity constraint on the predicted homography, which is realized by using a generator to predict a mask of aligned regions, and then a discriminator to check if two masked feature maps are induced by a single homography. To validate the effectiveness of HomoGAN and its components, we conduct extensive experiments on a large-scale dataset, and results show that our matching error is 22% lower than the previous SOTA method. Code is available at https://github.com/megvii-research/HomoGAN Mingbo Hong, Nianjin Ye, Chunyu Lin, Qijun Zhao, Shuaicheng Liu |
CVPR | 4 |
| 2022 | Deep Rectangling for Image Stitching: A Learning BaselineabstractStitched images provide a wide field-of-view (FoV) but suffer from unpleasant irregular boundaries. To deal with this problem, existing image rectangling methods devote to searching an initial mesh and optimizing a target mesh to form the mesh deformation in two stages. Then rectangu-lar images can be generated by warping stitched images. However, these solutions only work for images with rich linear structures, leading to noticeable distortions for por-traits and landscapes with non-linear objects. In this paper, we address these issues by proposing the first deep learning solution to image rectangling. Con-cretely, we predefine a rigid target mesh and only estimate an initial mesh to form the mesh deformation, contributing to a compact one-stage solution. The initial mesh is predicted using a fully convolutional network with a resid-ual progressive regression strategy. To obtain results with high content fidelity, a comprehensive objective function is proposed to simultaneously encourage the boundary rect-angular, mesh shape-preserving, and content perceptually natural. Besides, we build the first image stitching rectan-gling dataset with a large diversity in irregular boundaries and scenes. Experiments demonstrate our superiority over traditional methods both quantitatively and qualitatively. Lang Nie, Chunyu Lin, Kang Liao, Shuaicheng Liu, Yao Zhao 0001 |
CVPR | 2 |
| 2022 | PanoFormer: Panorama Transformer for Indoor 360$^{\circ }$ Depth Estimation
Zhijie Shen, Chunyu Lin, Kang Liao, Lang Nie, Zishuo Zheng, Yao Zhao 0001 |
ECCV (1) | 2 |
| 2022 | SivsFormer: Parallax-Aware Transformers for Single-image-based View SynthesisabstractSingle-image-based view synthesis is significant for generating a 3D scene and gains increasing attention in recent years. However, this task is challenging as it requires inferring contents beyond what is immediately visible. Previous methods directly predict the unknown views using the convolutional neural networks, but the generated views suffer from visually unpleasant holes, deformations, and artifacts. In this paper, we propose a Single-image-based view synthesis transformer (named SivsFormer) for high-quality and realistic view synthesis. In particular, a warping and occlusion handing module is designed to reduce the influence of parallax on the network. Subsequently, a disparity alignment module captures the long-range information over the scene and ensures that pixels move in a geometrically correct manner with soft probabilistic disparity maps. Moreover, we present a parallax-aware loss function to improve the quality of the synthetic images, which explicitly quantifies the magnitude of parallaxes. We conduct extensive experiments on popular KITTI and Cityscapes datasets. Benefitting from the proposed parallax-aware transformer, our approach achieves superior performance in both quantitative and qualitative evaluations. Chunlan Zhang, Chunyu Lin, Kang Liao, Lang Nie, Yao Zhao 0001 |
VR | 2 |
| 2022 | Learning edge-preserved image stitching from multi-scale deep homography
Lang Nie, Chunyu Lin, Kang Liao, Yao Zhao 0001 |
Neurocomputing | 2 |
| 2022 | Bi-projection for 360°image object detection bridged by RoI Searcher
Zishuo Zheng, Chunyu Lin, Lang Nie, Kang Liao, Yao Zhao 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2022 | Future pseudo-LiDAR frame prediction for autonomous driving
Chunyu Lin, Lang Nie, Yao Zhao 0001 |
Multim. Syst. | 2 |
| 2022 | Depth-Aware Multi-Grid Deep Homography Estimation With Contextual CorrelationabstractHomography estimation is an important task in computer vision applications, such as image stitching, video stabilization, and camera calibration. Traditional homography estimation methods heavily depend on the quantity and distribution of feature correspondences, leading to poor robustness in low-texture scenes. The learning solutions, on the contrary, try to learn robust deep features but demonstrate unsatisfying performance in the scenes with low overlap rates. In this paper, we address these two problems simultaneously by designing a contextual correlation layer (CCL). The CCL can efficiently capture the long-range correlation within feature maps and can be flexibly used in a learning framework. In addition, considering that a single homography can not represent the complex spatial transformation in depth-varying images with parallax, we propose to predict multi-grid homography from global to local. Moreover, we equip our network with a depth perception capability, by introducing a novel depth-aware shape-preserved loss. Extensive experiments demonstrate the superiority of our method over state-of-the-art solutions in the synthetic benchmark dataset and real-world dataset. The codes and models will be available athttps://github.com/nie-lang/Multi-Grid-Deep-Homography. Lang Nie, Chunyu Lin, Kang Liao, Shuaicheng Liu, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Neural Contourlet Network for Monocular 360° Depth EstimationabstractFor a monocular 360° image, depth estimation is a challenging because the distortion increases along the latitude. To perceive the distortion, existing methods devote to designing a deep and complex network architecture. In this paper, we provide a new perspective that constructs an interpretable and sparse representation for a 360° image. Considering the importance of the geometric structure in depth estimation, we utilize the contourlet transform to capture an explicit geometric cue in the spectral domain and integrate it with an implicit cue in the spatial domain. Specifically, we propose a neural contourlet network consisting of a convolutional neural network and a contourlet transform branch. In the encoder stage, we design a spatial–spectral fusion module to effectively fuse two types of cues. Contrary to the encoder, we employ the inverse contourlet transform with learned low-pass subbands and band-pass directional subbands to compose the depth in the decoder. Experiments on the three popular 360° panoramic image datasets demonstrate that the proposed approach outperforms the state-of-the-art schemes with faster convergence. Code is available athttps://github.com/zhijieshen-bjtu/Neural-Contourlet-Network-for-MODE. Zhijie Shen, Chunyu Lin, Lang Nie, Kang Liao, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Revisiting Radial Distortion Rectification in Polar-Coordinates: A New and Efficient Learning PerspectiveabstractFisheye cameras can capture a large field-of-view (Fov) scene but it introduces severe radial distortion in images. Thus, distortion rectification is a crucial step for subsequent computer vision tasks using fisheye cameras. A prevalent type of method predicts the displacement field between the input and output to rectify the distorted images. However, it is challenging to estimate the accurate flow in Cartesian coordinates (both$x$and$y$directions), in which the sampling strategy of the convolution kernel ignores the radial symmetry of distortion. In general, the pixel’s distortion at the same radius from the center is the same, while the radius corresponds to one parameter in polar coordinates. Motivated by this fact, we exploit the radial symmetry of distortion to predict a more straightforward one-dimensional flow, transforming the distorted image into the polar coordinates domain instead of predicting two-dimensional flow in$x$and$y$directions. Specifically, we propose a Polar coordinates Distortion Rectification Network (PCDRN), whose sampling strategy corresponds to the radial distortion characteristic so that a more accurate flow can be predicted. To eliminate the blurs and ring artifacts induced by the coordinates transformation, a Polar-To-Cartesian Appearance Enhancement Network is designed to enhance the local appearance of rectified images. Experimental results on the synthesized dataset and real-world dataset demonstrate the superiority of our approach in both quantitative and qualitative evaluations. Keyao Zhao, Chunyu Lin, Kang Liao, Shangrong Yang, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Pseudo-LiDAR Point Cloud Interpolation Based on 3D Motion Representation and Spatial SupervisionabstractPseudo-LiDAR point cloud interpolation is a novel and challenging task in autonomous driving, which aims to address the frequency mismatching problem between a camera and a LiDAR. Previous works represent the 3D spatial motion relationship with a coarse 2D optical flow, and the quality of interpolated point clouds only depends on the supervision of depth maps. As a result, the generated point clouds suffer from inferior global distributions and local appearances. To solve the above problems, we propose a Pseudo-LiDAR point cloud interpolation network to generate temporally and spatially high-quality point cloud sequences. By exploiting the scene flow from point clouds, the proposed network is able to learn a more accurate representation of the 3D spatial motion relationship. For a more comprehensive perception of the distribution of a point cloud, we design a novel reconstruction loss function with the chamfer distance to supervise the generation of Pseudo-LiDAR point clouds in 3D space. In addition, we introduce a multi-modal deep aggregation module to facilitate the efficient fusion of texture and depth features. As the benefits of the improved motion representation, training loss function, and model structure, our approach gains significant improvements on the Pseudo-LiDAR point cloud interpolation task. The experimental results evaluated on KITTI dataset demonstrate the state-of-the-art quantitative and qualitative performance of the proposed network. Kang Liao, Chunyu Lin, Yao Zhao 0001, Yulan Guo |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Towards Fast and Accurate Real-World Depth Super-Resolution: Benchmark Dataset and BaselineabstractDepth maps obtained by commercial depth sensors are always in low-resolution, making it difficult to be used in various computer vision tasks. Thus, depth map super-resolution (SR) is a practical and valuable task, which up-scales the depth map into high-resolution (HR) space. However, limited by the lack of real-world paired low-resolution (LR) and HR depth maps, most existing methods use down-sampling to obtain paired training samples. To this end, we first construct a large-scale dataset named "RGB-D-D", which can greatly promote the study of depth map SR and even more depth-related real-world tasks. The "D-D" in our dataset represents the paired LR and HR depth maps captured from mobile phone and Lucid Helios respectively ranging from indoor scenes to challenging outdoor scenes. Besides, we provide a fast depth map super-resolution (FDSR) baseline, in which the high-frequency component adaptively decomposed from RGB image to guide the depth map SR. Extensive experiments on existing public datasets demonstrate the effectiveness and efficiency of our network compared with the state-of-the-art methods. Moreover, for the real-world LR depth maps, our algorithm can produce more accurate HR depth maps with clearer boundaries and to some extent correct the depth value errors. Lingzhi He, Hongguang Zhu, Feng Li 0037, Huihui Bai 0001, Runmin Cong, Chunjie Zhang 0001, Chunyu Lin, Meiqin Liu 0002, Yao Zhao 0001 |
CVPR | 7 |
| 2021 | Progressively Complementary Network for Fisheye Image Rectification Using Appearance FlowabstractDistortion rectification is often required for fisheye images. The generation-based method is one mainstream solution due to its label-free property, but its naive skip-connection and overburdened decoder will cause blur and incomplete correction. First, the skip-connection directly transfers the image features, which may introduce distortion and cause incomplete correction. Second, the decoder is overburdened during simultaneously reconstructing the content and structure of the image, resulting in vague performance. To solve these two problems, in this paper, we focus on the interpretable correction mechanism of the distortion rectification network and propose a feature-level correction scheme. We embed a correction layer in skip-connection and leverage the appearance flows in different layers to pre-correct the image features. Consequently, the decoder can easily reconstruct a plausible result with the remaining distortion-less information. In addition, we propose a parallel complementary structure. It effectively reduces the burden of the decoder by separating content reconstruction and structure correction. Subjective and objective experiment results on different datasets demonstrate the superiority of our method. Shangrong Yang, Chunyu Lin, Kang Liao, Chunjie Zhang 0001, Yao Zhao 0001 |
CVPR | 2 |
| 2021 | Multi-Level Curriculum for Training A Distortion-Aware Barrel Distortion Rectification ModelabstractBarrel distortion rectification aims at removing the radial distortion in a distorted image captured by a wide-angle lens. Previous deep learning methods mainly solve this problem by learning the implicit distortion parameters or the nonlinear rectified mapping function in a direct manner. However, this type of manner results in an indistinct learning process of rectification and thus limits the deep perception of distortion. In this paper, inspired by the curriculum learning, we analyze the barrel distortion rectification task in a progressive and meaningful manner. By considering the relationship among different construction levels in an image, we design a multi-level curriculum that disassembles the rectification task into three levels, structure recovery, semantics embedding, and texture rendering. With the guidance of the curriculum that corresponds to the construction of images, the proposed hierarchical architecture enables a progressive rectification and achieves more accurate results. Moreover, we present a novel distortion-aware pre-training strategy to facilitate the initial learning of neural networks, promoting the model to converge faster and better. Experimental results on the synthesized and real-world distorted image datasets show that the proposed approach significantly outperforms other learning methods, both qualitatively and quantitatively. Kang Liao, Chunyu Lin, Lixin Liao, Yao Zhao 0001, Weiyao Lin |
ICCV | 2 |
| 2021 | Towards Complete Scene and Regular Shape for Distortion Rectification by Curve-Aware ExtrapolationabstractThe wide-angle lens gains increasing attention since it can capture a wide field-of-view (FoV) scene. However, the obtained image is contaminated with radial distortion, making the scene not realistic. Previous distortion rectification methods rectify the image in a rectangle or invagination, failing to display the complete content and regular shape simultaneously. In this paper, we rethink the representation of rectification results and present a Rectification OutPainting (ROP) method, aiming to extrapolate the coherent semantics to the blank area and create a wider FoV beyond the original wide-angle lens. To address the specific challenges such as the variable painting region and curve boundary, a rectification module is designed to rectify the image with geometry supervision, and the extrapolated results are generated using a dual conditional expansion strategy. In terms of the spatially discounted correlation, a curve-aware correlation measurement is proposed to focus on the generated region to enforce the local consistency. To our knowledge, we are the first to tackle the challenging rectification via outpainting, and our curve-aware strategy can reach a rectification construction with complete content and regular shape. Extensive experiments well demonstrate the superiority of our ROP over other state-of-the-art solutions. Kang Liao, Chunyu Lin, Yunchao Wei, Feng Li 0037, Shangrong Yang, Yao Zhao 0001 |
ICCV | 2 |
| 2021 | Distortion-Tolerant Monocular Depth Estimation on Omnidirectional Images Using Dual-CubemapabstractEstimating the depth of omnidirectional images is more challenging than that of normal field-of-view (NFoV) images because the varying distortion can significantly twist an object’s shape. The existing methods suffer from troublesome distortion while estimating the depth of omnidirectional images, leading to inferior performance. To reduce the negative impact of the distortion influence, we propose a distortion-tolerant omnidirectional depth estimation algorithm using a dual-cubemap. It comprises two modules: Dual-Cubemap Depth Estimation (DCDE) module and Boundary Revision (BR) module. In DCDE module, we present a rotation-based dual-cubemap model to estimate the accurate NFoV depth, reducing the distortion at the cost of boundary discontinuity on omnidirectional depths. Then a boundary revision module is designed to smooth the discontinuous boundaries, which contributes to the precise and visually continuous omnidirectional depths. Extensive experiments demonstrate the superiority of our method over other state-of-the-art solutions. Zhijie Shen, Chunyu Lin, Lang Nie, Kang Liao, Yao Zhao 0001 |
ICME | 2 |
| 2021 | ODE-Inspired Image Denoiser: An End-to-End Dynamical Denoising Network
Meiqin Liu 0002, Chunyu Lin, Yao Zhao 0001 |
PRCV (3) | 4 |
| 2021 | Image Outpainting with Depth Assistance
Lei Zhang 0116, Kang Liao, Chunyu Lin, Meiqin Liu 0002, Yao Zhao 0001 |
PRCV (3) | 3 |
| 2021 | Pseudo-LiDAR point cloud magnification
Chunlan Zhang, Kang Liao, Chunyu Lin, Yao Zhao 0001 |
Neurocomputing | 3 |
| 2021 | Joint distortion rectification and super-resolution for self-driving scene perception
Keyao Zhao, Kang Liao, Chunyu Lin, Meiqin Liu 0002, Yao Zhao 0001 |
Neurocomputing | 3 |
| 2021 | A Deep Ordinal Distortion Estimation Approach for Distortion RectificationabstractRadial distortion has widely existed in the images captured by popular wide-angle cameras and fisheye cameras. Despite the long history of distortion rectification, accurately estimating the distortion parameters from a single distorted image is still challenging. The main reason is that these parameters are implicit to image features, influencing the networks to learn the distortion information fully. In this work, we propose a novel distortion rectification approach that can obtain more accurate parameters with higher efficiency. Our key insight is that distortion rectification can be cast as a problem of learning an ordinal distortion from a single distorted image. To solve this problem, we design a local-global associated estimation network that learns the ordinal distortion to approximate the realistic distortion distribution. In contrast to the implicit distortion parameters, the proposed ordinal distortion has a more explicit relationship with image features, and significantly boosts the distortion perception of neural networks. Considering the redundancy of distortion information, our approach only uses a patch of the distorted image for the ordinal distortion estimation, showing promising applications in efficient distortion rectification. In the distortion rectification field, we are the first to unify the heterogeneous distortion parameters into a learning-friendly intermediate representation through ordinal distortion, bridging the gap between image feature and distortion rectification. The experimental results demonstrate that our approach outperforms the state-of-the-art methods by a significant margin, with approximately 23% improvement on the quantitative evaluation while displaying the best performance on visual appearance. Kang Liao, Chunyu Lin, Yao Zhao 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Unsupervised Deep Image Stitching: Reconstructing Stitched Features to ImagesabstractTraditional feature-based image stitching technologies rely heavily on feature detection quality, often failing to stitch images with few features or low resolution. The learning-based image stitching solutions are rarely studied due to the lack of labeled data, making the supervised methods unreliable. To address the above limitations, we propose an unsupervised deep image stitching framework consisting of two stages: unsupervised coarse image alignment and unsupervised image reconstruction. In the first stage, we design an ablation-based loss to constrain an unsupervised homography network, which is more suitable for large-baseline scenes. Moreover, a transformer layer is introduced to warp the input images in the stitching-domain space. In the second stage, motivated by the insight that the misalignments in pixel-level can be eliminated to a certain extent in feature-level, we design an unsupervised image reconstruction network to eliminate the artifacts from features to pixels. Specifically, the reconstruction network can be implemented by a low-resolution deformation branch and a high-resolution refined branch, learning the deformation rules of image stitching and enhancing the resolution simultaneously. To establish an evaluation benchmark and train the learning framework, a comprehensive real-world image dataset for unsupervised deep image stitching is presented and released. Extensive experiments well demonstrate the superiority of our method over other state-of-the-art solutions. Even compared with the supervised solutions, our image stitching quality is still preferred by users. Lang Nie, Chunyu Lin, Kang Liao, Shuaicheng Liu, Yao Zhao 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Graph Learning Based Head Movement Prediction for Interactive 360 Video StreamingabstractUltra-high definition (UHD) 360 videos encoded in fine quality are typically too large to stream in its entirety over bandwidth (BW)-constrained networks. One popular approach is to interactively extract and send a spatial sub-region corresponding to a viewer's current field-of-view (FoV) in a head-mounted display (HMD) for more BW-efficient streaming. Due to the non-negligible round-trip-time (RTT) delay between server and client, accurate head movement prediction foretelling a viewer's future FoVs is essential. In this paper, we cast the head movement prediction task as a sparse directed graph learning problem: three sources of relevant information-collected viewers' head movement traces, a 360 image saliency map, and a biological human head model-are distilled into a view transition Markov model. Specifically, we formulate a constrained maximum a posteriori (MAP) problem with likelihood and prior terms defined using the three information sources. We solve the MAP problem alternately using a hybrid iterative reweighted least square (IRLS) and Frank-Wolfe (FW) optimization strategy. In each FW iteration, a linear program (LP) is solved, whose runtime is reduced thanks to warm start initialization. Having estimated a Markov model from data, we employ it to optimize a tile-based 360 video streaming system. Extensive experiments show that our head movement prediction scheme noticeably outperformed existing proposals, and our optimized tile-based streaming scheme outperformed competitors in rate-distortion performance. Xue Zhang 0008, Gene Cheung, Yao Zhao 0001, Patrick Le Callet, Chunyu Lin, Jack Z. G. Tan |
IEEE Trans. Image Process. | 5 |
| 2020 | IET Image Processingabstract360 video is very popular due to its 360 views of a scene. Although 360 videos are also compressed by a hybrid coding framework like 2D video, its high resolution and serious shape deformation affect coding efficiency. In equirectangular projection (ERP) format of 360 videos, if an object moves from equator regions to pole regions or vice versa, large deformation will be introduced and motion estimation cannot find the best‐matched part. To solve the above problem, the authors propose to generate a better reference frame for the current to be encoded frame. First, they project the frame prior to the current one from ERP to the sphere and rotate it at an appropriate angle depending on motion vectors. Subsequently, they insert this generated frame to the rear of the reference queue and let the encoder work as usual. The advantage is that the inserted frame has a more similar shape deformation as the current frame, which greatly helps motion estimation and makes full use of 360 video characters. Their method is simple and friendly compatible with the existing compression standard. Experiments prove that their method achieves 1.57% Bjøntegaard Delta (BD)‐gain compared with standard high efficiency video coding. Chunyu Lin, Yao Zhao 0001, Meiqin Liu 0002, Xue Zhang 0008 |
IET Image Process. | 2 |
| 2020 | A view-free image stitching network based on global homography
Lang Nie, Chunyu Lin, Kang Liao, Meiqin Liu 0002, Yao Zhao 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2020 | Unsupervised fisheye image correction through bidirectional loss with geometric prior
Shangrong Yang, Chunyu Lin, Kang Liao, Yao Zhao 0001, Meiqin Liu 0002 |
J. Vis. Commun. Image Represent. | 2 |
| 2020 | Preface
Yao Zhao 0001, Chunyu Lin |
Pattern Recognit. Lett. | 2 |
| 2020 | Pixel-Level View Synthesis Distortion Estimation for 3D Video CodingabstractRecently, region-based 3D video coding has been proposed. However, existing view synthesis distortion estimation (VSDE) methods are performed at the frame level. To guide the rate-distortion optimization process of region-based 3D video coding schemes, this paper proposes the first pixel-level VSDE (PL-VSDE) method. We first give the definition of the pixel-level view synthesis distortion. To estimate it, a backward prediction method is then developed, which starts from the pixels of interest (POIs) in the virtual view and finds their corresponding pixels in the reference view via a coarse-to-fine approach, denoted as coarse-to-fine backward prediction (CFBP) method. Additionally, the CFBP fully considers the details of 3D warping, the rounding operation and the warping competition in view synthesis, leading to improve accuracy of the prediction. Besides, a table-lookup method and a warping property are introduced to speed up the CFBP. After integrating the CFBP into the PL-VSDE, we can estimate the view synthesis distortion at the pixel level. Our method is carried out pixel-by-pixel independently, which is friendly for parallel processing. The experimental results demonstrate that our proposed method has significant advantages in both accuracy and efficiency compared with the state-of-the-art frame-level VSDE methods. Jie Liang 0001, Yao Zhao 0001, Chunyu Lin, Lili Meng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | DR-GAN: Automatic Radial Distortion Rectification Using Conditional GAN in Real-TimeabstractRadial distortion, which severely hinders object detection and semantic recognition, frequently exists in images captured using a wide-angle lens. Correction of this distortion of images is crucial in many computer vision applications. In this paper, we present distortion rectification generative adversarial network (DR-GAN), a conditional generative adversarial network (GAN) for automatic radial DR. To the best of our knowledge, this is the first end-to-end trainable adversarial framework for radial distortion rectification. The DR-GAN trained using the proposed low-to-high perceptual loss learns the mapping relation between different structural images rather than estimating multifarious distortion parameters, while also realizing label-free training and one-stage rectification. As a benefit of one-stage rectification, the proposed method is extremely fast with the completion of rectification in real time. This is approximately 22 times faster than the state-of-the-art methods. The experimental results show that the DR-GAN achieves an excellent performance in both quantitative measure (PSNR and SSIM) and visual qualitative appearance. Kang Liao, Chunyu Lin, Yao Zhao 0001, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Distortion Rectification From Static to Dynamic: A Distortion Sequence Construction PerspectiveabstractDistortion rectification is a fundamental task in the field of computer vision and image processing. Nevertheless, previous methods have regarded distortion rectification as a static problem that learns a mapping function and corrects the distorted image to a unique state. However, this state is generally not the optimal solution, as it would result in an under-rectified or over-rectified structure. In this study, we revisit the classical distortion rectification task with a new perspective and redesign the algorithm, inspired by video processing techniques. Specifically, we regard distortion rectification as a dynamic problem that can be extended to a sequence of different distortion states: the input distorted image (t), under-rectified image (t+1), ideal-rectified image (t+2), and over-rectified image (t+3). We first estimate the residual distortion map (RDM) between the input distorted image and the coarse-rectified (t+1 or t+3) image. Here, RDM indicates the motion difference between two distorted images. Subsequently, the RDM is used to guide the refinement rectification process, aiming to convert the coarse-rectified state into the ideal-rectified state. In addition, the flexible implementation of the proposed refinement process with RDM to improve the rectification results of any method is appealing. The experimental results demonstrate that our method outperforms the state-of-the-art schemes by a significant margin, revealing approximately 40% improvement through quantitative evaluation. Kang Liao, Chunyu Lin, Yao Zhao 0001, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Model-Free Distortion Rectification Framework Bridged by Distortion Distribution MapabstractRecently, learning-based distortion rectification schemes have shown high efficiency. However, most of these methods only focus on a specific camera model with fixed parameters, thus failing to be extended to other models. To avoid such a disadvantage, we propose a model-free distortion rectification framework for the single-shot case, bridged by the distortion distribution map (DDM). Our framework is based on an observation that the pixel-wise distortion information is mathematically regular in a distorted image, despite different models having different types and numbers of distortion parameters. Motivated by this observation, instead of estimating the heterogeneous distortion parameters, we construct a proposed distortion distribution map that intuitively indicates the global distortion features of a distorted image. In addition, we develop a dual-stream feature learning module, benefitting from both the advantages of traditional methods that leverage the local handcrafted feature and learning-based methods that focus on the global semantic feature perception. Due to the sparsity of handcrafted features, we discrete the features into a 2D point map and learn the structure inspired by PointNet. Finally, a multimodal attention fusion module is designed to attentively fuse the local structural and global semantic features, providing the hybrid features for the more reasonable scene recovery. The experimental results demonstrate the excellent generalization ability and more significant performance of our method in both quantitative and qualitative evaluations, compared with the stateof- the-art methods. Kang Liao, Chunyu Lin, Yao Zhao 0001, Mai Xu |
IEEE Trans. Image Process. | 2 |
| 2019 | Improving Cube-to-ERP Conversion Performance with Geometry Features of 360 Video Structureabstract360 videos provide an omnidirectional view of the scene with extremely large data. Therefore, representing 360 videos with less data has become more and more important. Cube format is such a popular representation of 360 videos. However, we have to convert cube to Equirectangula(ERP) for displaying convenience. In this paper, we enhance Cube-to-ERP conversion performance by joint using Convolutional Neural Network(CNN) and classical interpolation method. The optimal threshold of boundary is derived according to geometry features of the cube-to-ERP format. This threshold is the guidance of how to combine CNN and classical interpolation method. Our experiment results prove that the derived threshold has a certain degree of guiding significance. Furthermore, we propose a new evaluation criterion with the help of Marsaglia model. It is much easier and more accurate to evaluate geometry conversion process. Chunyu Lin, Huihui Bai 0001, Meiqin Liu 0002, Yao Zhao 0001 |
DCC | 1 |
| 2019 | Block Partitioning Decision Based on Content Complexity for Future Video Coding
Yanhong Zhang, Yao Zhao 0001, Chunyu Lin, Meiqin Liu 0002 |
ICIG (3) | 3 |
| 2019 | A Depth-Bin-Based Graphical Model for Fast View Synthesis Distortion EstimationabstractDuring 3-D video communication, transmission errors, such as packet loss, could happen to the texture and depth sequences. View synthesis distortion will be generated when these sequences are used to synthesize virtual views according to the depth-image-based rendering method. A depth-value-based graphical model (DVGM) has been employed to achieve the accurate packet-loss-caused view synthesis distortion estimation (VSDE). However, the DVGM models the complicated view synthesis processes at depth-value level, which costs too much computation and is difficult to be applied in practice. In this paper, a depth-bin-based graphical model (DBGM) is developed, in which the complicated view synthesis processes are modeled at depth-bin level so that it can be used for the fast VSDE with 1-D parallel camera configuration. To this end, several depth values are fused into one depth bin, and a depth-bin-oriented rule is developed to handle the warping competition process. Then, the properties of the depth bin are analyzed and utilized to form the DBGM. Finally, a conversion algorithm is developed to convert the per-pixel input depth value probability distribution into the depth-bin format. Experimental results verify that our proposed method is 8-$32\times $ faster and requires 17%-60% less memory than the DVGM, with exactly the same accuracy. Jie Liang 0001, Yao Zhao 0001, Chunyu Lin, Anhong Wang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Adaptive Streaming in Interactive Multiview Video SystemsabstractMultiview applications endow final users with the possibility to freely navigate within 3D scenes with minimum-delay. A real feeling of scene navigation is enabled by transmitting multiple high-quality camera views, which can be used to synthesize additional virtual views to offer a smooth navigation. However, when network resources are limited, not all camera views can be sent at high quality. It is therefore important, yet challenging, to find the right tradeoff between coding artifacts (reducing the quality of camera views) and virtual synthesis artifacts (reducing the number of camera views sent to users). To this aim, we propose an optimal transmission strategy for interactive multiview HTTP adaptive streaming. We propose a problem formulation to select the optimal set of camera views that the client requests for downloading, such that the navigation quality experienced by the user is optimized while the bandwidth constraints are satisfied. We show that our optimization problem is NP-hard, and we therefore develop an optimal solution based on the dynamic programming algorithm with polynomial time complexity. To further simplify the deployment, we present a suboptimal greedy algorithm with effective performance and lower complexity. The proposed controller is evaluated in theoretical and realistic settings characterized by realistic network statistics estimation, buffer management, and server-side representation optimization. Simulation results show significant improvement in terms of navigation quality compared with alternative baseline multiview adaptation logic solutions. Xue Zhang 0008, Laura Toni, Pascal Frossard, Yao Zhao 0001, Chunyu Lin |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | A packetization strategy for interactive multiview video streaming over lossy networks
Xue Zhang 0008, Yao Zhao 0001, Tammam Tillo, Chunyu Lin |
Signal Process. | 4 |
| 2018 | 3-D Surround View for Advanced Driver Assistance SystemsabstractAs the primary means of transportations in modern society, the automobile is developing toward the trend of intelligence, automation, and comfort. In this paper, we propose a more immersive 3-D surround view covering the automobiles around for advanced driver assistance systems. The 3-D surround view helps drivers to become aware of the driving environment and eliminates visual blind spots. The system first uses four fish-eye lenses mounted around a vehicle to capture images. Then, according to the pattern of image acquisition, camera calibration, image stitching, and scene generation, the 3-D surround driving environment is created. To achieve the real-time and easy-to-handle performance, we only use one image to finish the camera calibration through a special designed checkerboard. Furthermore, in the process of image stitching, a 3-D ship model is built to be the supporter, where texture mapping and image fusion algorithms are utilized to preserve the real texture information. The algorithms used in this system can reduce the computational complexity and improve the stitching efficiency. The fidelity of the surround view is also improved, thereby optimizing the immersion experience of the system under the premise of preserving the information of the surroundings. Chunyu Lin, Yao Zhao 0001, Xin Wang 0046, Shikui Wei |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | Region-Based Multiple Description Coding for Multiview Video Plus Depth VideoabstractInterframe and interview predictions are widely employed in multiview video coding. This technique improves the coding efficiency, but it also increases the vulnerability of the coded bitstream. Thus, one packet loss will affect many subsequent frames in the same view and probably in other referenced views. To address this problem, a region-based multiple description coding scheme is proposed for robust 3-D video communication in this paper, in which two descriptions are formed by setting the left and right view as dominant in the first and second description, respectively. This approach exploits the fact that most regions in the reference view could be synthesized from the base view. Hence, these regions could be skipped or only coarsely encoded. In our work, the disoccluded regions, illumination-affected regions, and remaining regions are first determined and extracted. By assigning different quantization parameters for these three different regions according to the network status, an efficient multiple description scheme is formed. Experimental results demonstrate that the proposed scheme achieves considerably better performance compared with the traditional approach. Chunyu Lin, Yao Zhao 0001, Jimin Xiao, Tammam Tillo |
IEEE Trans. Multim. | 1 |
| 2017 | Optimized receiver control in interactive multiview video streaming systemsabstractMultiview applications endow final users with the possibility to freely navigate within 3D scenes with minimum-delay. High-quality rendering of the scene is enabled by transmitting multiple high-quality camera views, which can be used to synthesize additional virtual views to offer a smooth navigation in the scene. When network resources are limited, the set of camera views needs to be properly selected by the client. The right tradeoff between coding artifacts (reducing the quality of camera views) and virtual synthesis artifacts (reducing the number of camera views sent to users) has to be optimized. Existing client adaptation logic strategies usually fail to properly consider the content characteristics and the client navigation properties in the view selection problem. We therefore propose an optimal representation selection for interactive multiview HTTP adaptive streaming (HAS), with a complete problem formulation to select the optimal set of camera views that optimize the navigation quality experienced by the user while satisfying the bandwidth constraints. We show that our optimization problem is NP-hard and develop an effective solution based on a dynamic programming algorithm with polynomial time complexity. Simulation results show significant navigation quality improvement compared to two baseline multiview adaptation logic solutions. This confirms that adaptation logics have to consider both video content and interactivity level of the user in the representation selection strategy. Xue Zhang 0008, Laura Toni, Pascal Frossard, Yao Zhao 0001, Chunyu Lin |
ICC | 5 |
| 2017 | A Vehicle-Mounted Multi-camera 3D Panoramic Imaging Algorithm Based on Ship-Shaped Model
Xin Wang 0046, Chunyu Lin, Shikui Wei, Yao Zhao 0001 |
ICIG (3) | 2 |
| 2017 | Depth map up-sampling with fractal dimension and texture-depth boundary consistencies
Meiqin Liu 0002, Yao Zhao 0001, Jie Liang 0001, Chunyu Lin, Huihui Bai 0001 |
Neurocomputing | 4 |
| 2017 | 3D saliency detection based on background detection
Hongyun Lin, Chunyu Lin, Yao Zhao 0001, Anhong Wang |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | Just Noticeable Difference Based Fast Coding Unit Partition in 3D-HEVC Intra CodingabstractSummary form only given. This paper mainly studies currently developing 3D video coding based on HEVC. HEVC-based 3D video coding mainly focuses on 3DTV and auto-stereoscopic video compression system. A variety of new encoding tools, such as inter-view motion prediction and depth modeling modes, have been added in 3D-HEVC. Although 3D-HEVC provides greater bit rate saving, it also brings the enormous encoding complexity increase. The coding time is increased correspondingly. It is necessary to reduce the encoding time. In this paper, a fast CU-sized partition algorithm is proposed for 3D-HEVC intra coding. The key point of this algorithm is to find the relationship between the texture characteristic and the sub-partition in each CU. It needs to determine whether the LCU can be subdivided to smaller CU according to the relationship. In order to reduce the redundancy of the human eye, just noticeable difference (JND) is a high efficiency model in the base of psychology and physiology. Instead of the time-consuming rate distortion optimization for coding mode decision, the variance of JND in each CU can be exploited to partition the coding unit according to human visual system characteristics. In other words, the larger blocks with higher JND variance will be subdivided to smaller blocks with lower JND variance. Consequently, the rules of CU preliminary partition are decided as follows: (a) For a 64×64 CU, if the variance of JND is larger than 0.25, the CU will be sub-divided into four 32×32 sub-blocks. (b) For a 32×32 CU, if the variance of JND is larger than 0.15, the CU will be sub-divided into four 16×16 sub-blocks. (c) For a 16×16 CU, if the variance of JND is larger than 0.10, the CU will be sub-divided into four 8×8 sub-blocks. The proposed algorithm is implemented based on HTM-13.1 reference software. The experiment condition is set up as "All Intra-Main" (AI-Main) configuration [1]. The quantization parameter (QP) values of texture are set to 25, 30, 35 and 40, respectively and the corresponding QPs of depth can be set to 34,39,42,45. The experimental results show that the fast intra mode decision algorithm provides over 29.25% encoding time saving on average with comparable rate distortion performance. Hai Ren, Huihui Bai 0001, Chunyu Lin, Mengmeng Zhang 0008, Yao Zhao 0001 |
DCC | 3 |
| 2016 | Depth map up-sampling with texture edge feature via sparse representationabstractDepth information has been an important part in various computer vision applications. However, the low resolution depth map captured by TOF cameras and other depth sensors cannot be directly used in visual depth perception and 3D reconstruction. Thus, this paper proposes an effective depth map up-sampling algorithm which reconstructs a high resolution depth map from a low depth map, by integrating the registered color image. Different from the other depth map up-sampling methods, we establish the correspondence between the high resolution and low resolution image pairs via image sparse representation, then extract the geometric structure component and texture component from high resolution depth map and the relevant color image, respectively. Specially, we employ a guided filter to preserve regions smooth and edges sharp. The experimental results show that the proposed depth map up-sampling method obtains quality improvement on depth map and the synthesized view compared with state-of-the-art methods. Chunyu Lin, Yao Zhao 0001, Jingxuan Hou |
VCIP | 2 |
| 2016 | Feature-based depth refinement for view synthesisabstractThe accuracy of depth map limited the performance of Depth-Image-Based-Rendering (DIBR) system, especially the misalignment of objects between the wrapped image and the target image which locates on the real viewpoint would seriously decreases the quality of the wrapped image. In practice, the misalignment is a major cause that the objective quality is not enough, comparing the wrapped image and the target image. In this paper, we propose a new depth refinement method which aims to align the wrapped positions and the real target positions. More specifically, to align the wrapped image, some features extracted from texture images are used. Then, based on the matched features, an energy function is minimized the depth errors. This energy function is efficient enough to provide high quality rendered image, even comparing with the real target image. Experiments show that the proposed method provides visually pleasing results and high objective quality. Yao Zhao 0001, Chunyu Lin |
VCIP | 3 |
| 2016 | Packetization strategies for MVD-based 3D video transmissionabstractIn multi-view video plus depth (MVD) format, virtual views are synthesized by the compressed texture videos and their associated depth through depth-image-based rendering. In this paper, we consider the setup where both the encoded texture and depth bitstreams experience packet losses during transmission. Different packetization strategies are investigated and a novel strategy is developed to improve error resilience of MVD-based video transmission, where texture data and its corresponding depth are put into the same packet. The size of texture plus associated depth data included in each packet needs to be less than the Maximum Transfer Unit (MTU). Experimental results demonstrate that our proposed packetization scheme yields a significant improvement in terms of both texture views and synthesized virtual views quality when fit in H.264/AVC. Xue Zhang 0008, Yao Zhao 0001, Tammam Tillo, Chunyu Lin, Jimin Xiao, Anhong Wang |
VCIP | 4 |
| 2016 | Region-Aware 3-D Warping for DIBRabstractIn 3-D video (3DV) applications, depth-image-based rendering (DIBR) has been widely employed to synthesize virtual views. However, this approach is performed in a frame-based way, meaning each whole frame is dealt with and the characteristics of different regions in the frame are ignored. As a result, redundant pixels in some regions are abused during the subsequent warping and blending stage. This paper proposes a region-aware 3-D warping approach for DIBR in which warped frames are reasonably divided beforehand so that only the indispensable regions are used. With the proposed scheme, it is possible to avoid noneffective and repeated pixels during the warping stage. In addition, the blending process is also saved. The experimental results show that compared to the state-of-the-art VSRS3.5 and VSRS-1D-fast algorithms, our approach can achieve significant computation savings without sacrificing synthesis quality. Anhong Wang, Yao Zhao 0001, Chunyu Lin, Bing Zeng 0001 |
IEEE Trans. Multim. | 4 |
| 2016 | Depth Map Down-Sampling and Coding Based on Synthesized View DistortionabstractIn this paper, we propose a depth map down-sampling and coding scheme that minimizes the view synthesis distortion. Moreover, a solution for the optimal depth map down-sampling problem that minimizes the depth-caused distortion in the virtual view by exploiting the depth map and the associated texture information along with the up-sampling method to be used in the decoder side is derived. Furthermore, to enhance compression performance, the synthesized view distortion, which is evaluated by emulating the interpolation and the virtual view synthesis process, is used in the optimization objective function for coding mode selection in the video encoder. Experimental results show that both the proposed depth map down-sampling and encoding methods lead to good performance, and the average bit rate reduction is 2.62% compared with 3D-AVC. Jimin Xiao, Tammam Tillo, Yao Zhao 0001, Chunyu Lin, Huihui Bai 0001 |
IEEE Trans. Multim. | 5 |
| 2015 | Intra-/inter-View Correlation Based Multiple Description Coding for Multiview TransmissionabstractWith the development of 3D video technology, many studies have paid attention to compression efficiency and rate distortion performance. When 3D videos are transmitted over error-prone channels, they may suffer significant quality degradation. In this paper, we combine multiview video coding (MVC) with multiple description coding (MDC) for robust transmission. The proposed scheme can give full consideration of both intra-view and inter-view correlation for better estimation. Furthermore, an adaptive mode decision is designed to generate a label as redundant information. The experiments show that the redundant information occupies just a few bits while the PSNR values of the reconstructed videos demonstrate a significant improvement. Jiansheng Guo, Huihui Bai 0001, Chunyu Lin, Mengmeng Zhang 0008, Yao Zhao 0001 |
DCC | 3 |
| 2015 | Texture Characteristics Based Fast Coding Unit Partition in HEVC Intra CodingabstractHigh efficiency video coding (HEVC) is an emerging video compression standard, developed by the Joint Collaborative Team on Video Coding (JCT-VC). The aim of HEVC standardization effort is to save about 50% bit rate for equal perceptual video quality relative to H.264/AVC. Although HEVC provides greater bit rate saving, it also brings the enormous encoding complexity increase. In this paper, we propose a fast intra CU decision algorithm based on the texture characteristics of video. Furthermore, we also consider the coding bits of each CU as auxiliary information to refine the partition results. Experimental results show that the fast intra mode decision algorithm provides over 33% complexity reduction in terms of encoding time with negligible quality loss, compared with the original HEVC test model version HM-12.0+RExt-4.0rc2. Huihui Bai 0001, Chunyu Lin, Mengmeng Zhang 0008, Yao Zhao 0001 |
DCC | 3 |
| 2015 | A fast region-level 3D-warping method for depth-image-based renderingabstractIn 3D video, depth-image-based rendering (DIBR) is widely employed in view synthesis to generate virtual views. However, the processing of this algorithm is based on the frame-level, and the characteristics in different regions cannot be fully taken into account before rendering. This drawback will lead to the unnecessary and redundant information in some regions being abused, which increases extra computation. This paper proposes a region-level 3D-warping method for DIBR, where regions are divided according to their characteristics. Then, only the necessary information in some important regions is utilized warping so that the redundant information could be avoided in the computation. Experimental results show that our approach is almost 4 times faster than VSRS-1D-fast, while declines 0.12 dB PSNR in the performance of synthesis views averagely. Hence, our method can achieve a good trade-off between the computation and view synthesis and will be especially useful for applications where the computation is the concern. Anhong Wang, Yao Zhao 0001, Chunyu Lin |
MMSP | 4 |
| 2015 | Quantized dictionary for sparse representationabstractDictionary learning for sparse representation has drawn considerable attention in recent years. In particular, the K-SVD algorithm is an efficient approach, and various modifications of the K-SVD have been developed for applications such as face recognition. However, the efficient storage of the dictionary has not been studied. Currently, the dictionary is simply normalized and saved as floating-point numbers, which could be quite large and lead to excessive cost and delay if the dictionary needs to be transmitted, e.g., to mobile users. In this paper, we develop a quantized K-SVD (Q-KSVD) to reduce the storage of the dictionary. We compress each basis image in the dictionary by the conventional image coding method. Moreover, we integrate the image compression step into various modified K-SVD optimization schemes, and develop an algorithm to find the optimal dictionary when there is a constraint on the total bits of the compressed dictionary. Our algorithm selects dictionary bases by ranking the contribution-rate slopes of all bases. This method also serves as an efficient approach to find the optimal number of bases of the dictionary at each rate constraint. Face recognition experiments using four K-SVD-based methods show that our method can achieve different tradeoffs between the dictionary storage space and the recognition accuracy. It can achieve comparable performance with as little as 3% of the original storage space. It can even yield higher accuracy than the uncompressed dictionary in some cases. Jie Liang 0001, Yao Zhao 0001, Chunyu Lin, Huihui Bai 0001 |
MMSP | 4 |
| 2015 | Multiple Description Coding for Stereoscopic Videos With Stagger Frame OrderabstractDue to the prediction structures employed in video coding, the loss of one packet will affect many following frames. In this paper, a multiple description coding scheme with stagger frame order is proposed for stereoscopic 3-D videos. First, the reference and auxiliary views in stereoscopic sequences will be asymmetrically encoded into one description, whereas the other description will be formed in the same way with one dumb frame delay. Because of the stagger frame order, the coarsely encoded B frames will be inserted into different positions of the two descriptions. If a certain frame encoded with I/P mode is lost, then its corresponding B-frame version will be employed to compensate for the loss. In each description, the quantization steps of B frames are tuned based on a closed-form solution that considers the video contents, network status, frame positions in the group of picture, and the layer of the views. For further improvement, a fusing scheme is provided. The experimental results demonstrate that the proposed scheme outperforms state-of-the-art schemes. Specifically, up to 1.3-dB gain is achieved in the case of packet loss, and 2-dB gain is obtained for the side/central performance. Chunyu Lin, Yao Zhao 0001, Tammam Tillo, Jimin Xiao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Local stereo matching algorithm using rotation-skeleton-based regionabstractThis paper proposes a local stereo matching algorithm for accurate disparity estimation by using 45° rotation-skeleton-based region (RSBR). For local stereo matching, an adaptive local region is important for the performance of the disparity estimation. In order to generate more accurate regions, we use a skeleton with 45° rotation to divide an initial window that achieved by mean-shift segmentation into four parts. All the pixels in the same height level are judged simultaneously, and the valid pixels in the four parts construct the 45° RSBR. Compared with the common local support region based on orthogonal skeleton, 45° RSBR enables to maintain the consistency in images and have the advantage of error tolerance. The local stereo matching algorithm with RSBR is improved in two aspects. First, the hybrid cost aggregation using the RSBR helps to remove some noise caused by outliers and improve the subjective performance. Second, the candidate values in the refinement step are selected from the RSBR, which ensures the validity of candidates. The experiment results demonstrate a good performance in Middlebury Stereo datasets, both in objective and subjective performance. Yao Zhao 0001, Chunyu Lin |
MMSP | 3 |
| 2014 | Optimizing the deadzone width to improve the polyphase-based multiple description coding
Chunyu Lin, Tammam Tillo, Jimin Xiao, Yao Zhao 0001 |
Multim. Tools Appl. | 1 |
| 2013 | M-channel multiple description coding based on uniformly offset quantizers with optimal deadzoneabstractThis paper proposes an improved source-splitting-based two-rate M-channel multiple description coding scheme, where the source is split into M subsets. In each description, one subset is coded at a high rate, and others are predictively coded at a low rate. Uniform offsets among low-rate quantizers of different descriptions are achieved by employing unequal deadzones and by quantizing the predictions. When several descriptions are received, the optimal reconstruction of each subset is achieved by finding the intersection of all received quantization bins. The closed-form expression of the expected distortion is obtained. The proposed scheme is applied to lapped transform-based multiple description image coding and achieves improved performance. The optimal deadzone selection and its impact are also given in this paper. Lili Meng, Jie Liang 0001, Yao Zhao 0001, Huihui Bai 0001, Chunyu Lin, André Kaup |
ICASSP | 5 |
| 2013 | Multiple description video coding based on forward error correction within expanding windowsabstractIn this paper, an MDC scheme based on forward error correction(FEC) within expanding windows is proposed. Firstly, the video sequence will be coded into source packets with/without slice group enabled. Secondly, the appropriate FEC packets are inserted according to the packet loss rate. Since the previous frames in a GOP is generally more important than the following frames in the GOP, an expanding window is exploited so that the FEC packets for the current frame will also protect the previous frames in the window. After this, the source packets with the inserted FEC packets will be divided into two descriptions and transmitted into two independent channels. When some packets in one description are lost, FEC decoding will try to recover the lost packets. Through this scheme, the source packets can get appropriate protection while the compression efficiency will not be degraded too much. The experimental results show that the proposed scheme outperforms the compared schemes up to 3 dB. Chunyu Lin, Yao Zhao 0001, Jimin Xiao, Tammam Tillo |
ICIP | 1 |
| 2013 | Directional block compressed sensing for image codingabstractCompared with traditional Nyquist sampling, compressed sensing (CS) enables a highly precise reconstruction of the signal from fewer measurements, suggesting great potential for efficient and simplistic data acquisition. In this paper, we propose a directional block-based compressed sensing (DBCS) scheme for image coding, where the directionalities inherently exhibited within image blocks are exploited as the “a priori” information. The image block is first directionally scanned following the dominating direction of its edges/textures. Then the vectorized image block is sampled by a block-based compressed sensing (BCS) method. At the decoder, each image block is recovered and then rearranged by the corresponding inverse-scan to obtain the recovered image. Experimental results show that the proposed DBCS scheme outperforms BCS due to the exploitation of the directional information within image blocks. Anhong Wang, Kongfen Zhu, Chunyu Lin, Yao Zhao 0001 |
ISCAS | 4 |
| 2013 | Fast bottom-up pruning for HEVC intraframe codingabstractIn intraframe coding of the High Efficiency Video Coding (HEVC) standard, up to 35 modes are defined for intra prediction and the quadtree structure is used for adaptive block partition. While such flexibility leads to more efficient compression, it also dramatically increases the encoder complexity. In this paper, a simple yet effective fast bottom-up pruning algorithm is proposed to reduce the computational cost. Mode decision at a large coding unit (CU) is selectively skipped based on the block structures of its sub-CUs. Our experimental results show that the proposed scheme can effectively reduce the encoder complexity without compromising the compression efficiency. Han Huang 0001, Yao Zhao 0001, Chunyu Lin, Huihui Bai 0001 |
VCIP | 3 |
| 2012 | Real-time video streaming exploiting the late-arrival packetsabstractFor real-time video applications, such as video telephony service, the allowed maximum end-to-end transmission delay is usually fixed. The packets arriving at the destination out of the maximum end-to-end delay are treated as late-arrival packets, and these packets are discarded in traditional video transmission systems. In this paper, in order to improve the system performance, we propose to exploit these packets to update the decoder reference buffer. Two schemes are proposed to exploit the late-arrival packets, one scheme is to use sliding-window updating, where the updating window is moving; another scheme is to use fixed-window update together with systematic Reed-Solomon code. The effectiveness of the two schemes are validated by simulation results without adding extra delay. It is found that in both schemes, the updating window size plays an important role on the system performance. Jimin Xiao, Tammam Tillo, Chunyu Lin, Yao Zhao 0001 |
PCS | 3 |
| 2012 | Dynamic Sub-GOP Forward Error Correction Code for Real-Time Video ApplicationsabstractReed-Solomon erasure codes are commonly studied as a method to protect the video streams when transmitted over unreliable networks. As a block-based error correcting code, on one hand, enlarging the block size can enhance the performance of the Reed-Solomon codes; on the other hand, large block size leads to long delay which is not tolerable for real-time video applications. In this paper a novel Dynamic Sub-GOP FEC (DSGF) approach is proposed to improve the performance of Reed-Solomon codes for video applications. With the proposed approach, the Sub-GOP, which contains more than one video frame, is dynamically tuned and used as the RS coding block, yet no delay is introduced. For a fixed number of extra introduced packets, for protection, the length of the Sub-GOP and the redundancy devoted to each Sub-GOP becomes a constrained optimization problem. To solve this problem, a fast greedy algorithm is proposed. Experimental results show that the proposed ap proach outperforms other real-time error resilient video coding technologies. Jimin Xiao, Tammam Tillo, Chunyu Lin, Yao Zhao 0001 |
IEEE Trans. Multim. | 3 |
| 2011 | Real-time forward error correction for video transmissionabstractWhen the video streams are transmitted over the unreliable networks, forward error correction (FEC) codes are usually used to protect them. Reed-Solomon codes are block-based FEC codes. On one hand, enlarging the block size can enhance the performance of the Reed- Solomon codes. On the other hand, large Reed-Solomon block size leads to long delay which is not tolerable for real-time video applications. In this paper a novel approach is proposed to improve the performance of Reed-Solomon codes. With the proposed approach, more than one video frame are encompassed in the Reed-Solomon coding block yet no delay is introduced. Experimental results show that the proposed approach outperforms other real-time error resilient video coding technologies. Jimin Xiao, Tammam Tillo, Chunyu Lin, Yao Zhao 0001 |
VCIP | 3 |
| 2011 | Multiple Description Coding for H.264/AVC With Redundancy Allocation at Macro Block LevelabstractIn this paper, a novel multiple description video coding scheme is proposed to insert and control the redundancy at macro block (MB) level. By analyzing the error propagation paths, the relative importance of each MB is determined. The paths, in practice, depend on both the video content and the adopted video coder. Considering the relative importance of the MB and the network status, an unequal protection for the video data can be realized to exploit the redundancy effectively. In addition, a simple and effective approach is introduced to tune the quantization parameter for the variable rate coding case. The whole scheme is implemented in H.264/AVC by employing its coding options, thus generating descriptions that are compatible with the baseline profile and extended profile of H.264/AVC. Due to its general property, the proposed approach can be employed for other hybrid video codecs. The results demonstrate the advantage of the proposed approach over other H.264/AVC multiple description schemes. Chunyu Lin, Tammam Tillo, Yao Zhao 0001, Byeungwoo Jeon |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | Distributed Video Coding Based on Multiple DescriptionabstractIn our paper, we propose a novel distributed video coding (DVC) scheme using the theory of multiple description (MD), in which key frame is encoded by MD codec and transmitted over the corresponding channel. This scheme combines the advantage of DVC as well as robustness of MD, and exploits three different methods to generate multiple descriptions for the key frames that are essential to side information. Experiments demonstrate that it can get better performance than some general DVC methods. Besides, it demonstrates higher robustness in packet-loss channel than general DVC due to the MD algorithm. Hongxia Ma, Yao Zhao 0001, Chunyu Lin, Anhong Wang |
IAS | 3 |
| 2008 | Two-Stage Diversity-Based Multiple Description Image CodingabstractIn this letter, a diversity-based two-description image coding scheme is firstly presented and analyzed, for which a two-stage coding is introduced to facilitate the tuning of central/side distortion tradeoff. By subsampling the central decoded errors from the first stage to constitute the second part for each description, respectively, we show that not only the central distortion but also the side distortion can be reduced with the second stage information. We further generalize the proposed scheme to 4-description coding. Experiment results demonstrate that the proposed scheme outperforms the state-of-the-art techniques in terms of both central and side coding performance. Chunyu Lin, Yao Zhao 0001, Ce Zhu |
IEEE Signal Process. Lett. | 1 |