EDBT 2026 Demo / reviewers in the wild / expert
Shuaicheng Liu
dblp:49/8652
· DBLP profile ↗
144ranked-venue papers
15as first author
104since 2021 · last 2026
0000-0002-8815-5335ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 120 · 11 first-author · 83 since 2021Artificial intelligence and machine learning · 88 · 8 first-author · 76 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RAW-Flow: Advancing RGB-to-RAW Image Reconstruction with Deterministic Latent Flow MatchingabstractRGB-to-RAW reconstruction, or the reverse modeling of a camera Image Signal Processing (ISP) pipeline, aims to recover high-fidelity RAW data from RGB images. Despite notable progress, existing learning-based methods typically treat this task as a direct regression objective and still struggle with detail inconsistency and color deviation, due to the ill-posed nature of inverse ISP and the inherent information loss in quantized RGB images. To address these limitations, we pioneer a generative perspective by reformulating RGB-to-RAW reconstruction as a deterministic latent transport problem and introduce a novel framework named RAW-Flow, which leverages flow matching to learn a deterministic vector field in latent space, to effectively bridge the gap between RGB and RAW representations and enable accurate reconstruction of structural details and color information. To further enhance latent transport, we introduce a cross-scale context guidance module that injects hierarchical RGB features into the flow estimation process. Moreover, we design a Dual-domain Latent Autoencoder (DLAE) with a feature alignment constraint to support the proposed latent transport framework, which jointly encodes RGB and RAW inputs while promoting stable training and high-fidelity reconstruction. Extensive experiments demonstrate that RAW-Flow outperforms state-of-the-art approaches both quantitatively and visually. Zhen Liu 0022, Diedong Feng, Hai Jiang 0006, Liaoyuan Zeng, Hao Wang 0073, Chaoyu Feng, Bing Zeng 0001, Shuaicheng Liu |
AAAI | 9 |
| 2026 | Robust Partial-to-Partial Point Cloud Registration with Overlapping Mask Learning
Hao Xu 0018, Guanghui Liu 0001, Bing Zeng 0001, Shuaicheng Liu |
Int. J. Comput. Vis. | 4 |
| 2026 | Supervised Small-Baseline and Large-Baseline Homography Learning With Diffusion-Based Data GenerationabstractIn this paper, we propose an iterative framework, which consists of two phases: a generation phase and a training phase, to generate realistic training data for supervised small-baseline and large-baseline homography learning and yield a state-of-the-art homography estimation network. In the generation phase, given an unlabeled image pair, we utilize the pre-estimated dominant plane masks and homography of the pair, along with another sampled homography that serves as ground truth to generate a new labeled training pair with realistic motion. In the training phase, the generated data is used to train the supervised homography network, in which the training data is refined via a content refinement diffusion model. Once an iteration is finished, the trained network is used in the next data generation phase to update the pre-estimated homography. Through such an iterative strategy, the quality of the dataset and the performance of the network can be gradually and simultaneously improved. Experimental results show that our method outperforms existing competitors and previous supervised methods can also be improved based on the generated dataset. Hai Jiang 0006, Haipeng Li 0001, Songchen Han, Bing Zeng 0001, Shuaicheng Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | SS-NeRF: Physically Based Sparse Spectral Rendering With Neural Radiance FieldabstractIn this paper, we propose SS-NeRF, the end-to-end Neural Radiance Field (NeRF)-based architectures for high-quality physically based rendering with sparse inputs. We modify the classical spectral rendering into two main steps, 1) the generation of a series of spectrum maps spanning different wavelengths, 2) the combination of these spectrum maps for the RGB output. The proposed architecture follows these two steps through the proposed multi-layer perceptron (MLP)-based architecture (SpectralMLP) and spectrum attention UNet (SAUNet). Given the ray origin and the ray direction, the SpectralMLP constructs the spectral radiance field to obtain spectrum maps of novel views, which are then sent to the SAUNet to produce RGB images of white-light illumination. Applying NeRF to build up the spectral rendering is a more physically-based way from the perspective of ray-tracing. Further, the spectral radiance fields decompose difficult scenes and improve the performance of NeRF-based methods. Previous baseline, such as SpectralNeRF, outperforms recent methods in synthesizing novel views but requires relatively dense viewpoints for accurate scene reconstruction. To tackle this, we propose SS-NeRF to enhance the detail of scene representation with sparse inputs. In SS-NeRF, we first design the depth-aware continuity to optimize the reconstruction based on single-view depth predictions. Then, the geometric-projected consistency is introduced to optimize the multi-view geometry alignment. Additionally, we introduce a superpixel-aligned consistency to ensure that the average color within each superpixel region remains consistent. Comprehensive experimental results demonstrate that the proposed method is superior to recent state-of-the-art methods when synthesizing new views on both synthetic and real-world datasets. Ru Li 0002, Guanghui Liu 0001, Shengping Zhang, Bing Zeng 0001, Shuaicheng Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Learning Efficient Meshflow and Optical Flow From Event CamerasabstractIn this paper, we explore the problem of event-based meshflow estimation, a novel task that involves predicting a spatially smooth sparse motion field from event cameras. To start, we review the state-of-the-art in event-based flow estimation, highlighting two key areas for further research: i) the lack of meshflow-specific event datasets and methods, and ii) the underexplored challenge of event data density. First, we generate a large-scale High-Resolution Event Meshflow (HREM) dataset, which showcases its superiority by encompassing the merits of high resolution at 1280 × 720, handling dynamic objects and complex motion patterns, and offering both optical flow and meshflow labels. These aspects have not been fully explored in previous works. Besides, we propose Efficient Event-based MeshFlow (EEMFlow) network, a lightweight model featuring a specially crafted encoder-decoder architecture to facilitate swift and accurate meshflow estimation. Furthermore, we upgrade EEMFlow network to support dense event optical flow, in which a Confidence-induced Detail Completion (CDC) module is proposed to preserve sharp motion boundaries. We conduct comprehensive experiments to show the exceptional performance and runtime efficiency (30×faster) of our EEMFlow model compared to the recent state-of-the-art flow method. As an extension, we expand HREM into HREM+, a multi-density event dataset contributing to a thorough study of the robustness of existing methods across data with varying densities, and propose an Adaptive Density Module (ADM) to adjust the density of input event data to a more optimal range, enhancing the model's generalization ability. We empirically demonstrate that ADM helps to significantly improve the performance of EEMFlow and EEMFlow+ by 8% and 10%, respectively. Xinglong Luo, Ao Luo, Kunming Luo, Zhengning Wang, Ping Tan 0002, Bing Zeng 0001, Shuaicheng Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Solving ILL-posed Regions in High Dynamic Range Reconstruction With Uncertainty-Aware Diffusion ModelsabstractLearning-based approaches have achieved promising progress in High Dynamic Range (HDR) image reconstruction, particularly in ghost removal. However, they often struggle in ill-posed regions, such as areas with occlusion or saturation, where insufficient or unreliable information leads to persistent residual ghosting artifacts and structural distortions. In this paper, we present UA-Diff, an uncertainty-aware diffusion framework designed to generate visually coherent, ghost-free HDR images. Specifically, our approach introduces an Uncertainty Generation Module (UGM) that estimates pixel-wise reconstruction confidence via a probabilistic Laplacian loss, producing an uncertainty map that explicitly highlights challenging ill-posed regions. To address these regions effectively, we develop an Uncertainty-Aware Diffusion Module (UADM) that operates selectively on the average-coefficient component of a 2D discrete wavelet transform, where dominant artifacts tend to concentrate. This enables reduced computational overhead while preserving high-quality details. Moreover, we propose an Uncertainty-Guided Sampling (UGS) strategy that leverages the uncertainty map to guide the denoising process, ensuring faithful reconstruction in reliable regions and targeted refinement in uncertain areas. Extensive experiments on three public HDR benchmarks demonstrate that UA-Diff surpasses state-of-the-art methods both quantitatively and perceptually, especially in challenging ill-posed scenarios. Zhen Liu 0022, Hai Jiang 0006, Haipeng Li 0001, Shuaicheng Liu, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | HybridReg: Robust 3D Point Cloud Registration with Hybrid MotionsabstractScene-level point cloud registration is very challenging when considering dynamic foregrounds. Existing indoor datasets mostly assume rigid motions, so the trained models cannot robustly handle scenes with non-rigid motions. On the other hand, non-rigid datasets are mainly object-level, so the trained models cannot generalize well to complex scenes. This paper presents HybridReg, a new approach to 3D point cloud registration, learning uncertainty mask to account for hybrid motions: rigid for backgrounds and non-rigid/rigid for instance-level foregrounds. First, we build a scene-level 3D registration dataset, namely HybridMatch, designed specifically with strategies to arrange diverse deforming foregrounds in a controllable manner. Second, we account for different motion types and formulate a mask-learning module to alleviate the interference of deforming outliers. Third, we exploit a simple yet effective negative log-likelihood loss to adopt uncertainty to guide the feature extraction and correlation computation. To our best knowledge, HybridReg is the first work that exploits hybrid motions for robust point cloud registration. Extensive experiments show HybridReg's strengths, leading it to achieve state-of-the-art performance on both widely-used indoor and outdoor datasets. Keyu Du, Hao Xu 0018, Haipeng Li 0001, Chi-Wing Fu, Shuaicheng Liu |
AAAI | 6 |
| 2025 | Diff-Shadow: Global-guided Diffusion Model for Shadow RemovalabstractWe propose Diff-Shadow, a global-guided diffusion model for high-quality shadow removal. Previous transformer-based approaches can utilize global information to relate shadow and non-shadow regions but are limited in their synthesis ability and recover images with obvious boundaries. In contrast, diffusion-based methods can generate better content but they are not exempt from issues related to inconsistent illumination. In this work, we combine the advantages of diffusion models and global guidance to realize shadow-free restoration. Specifically, we propose a parallel UNets architecture: 1) the local branch performs the patch-based noise estimation in the diffusion process, and 2) the global branch recovers the low-resolution shadow-free images. A Reweight Cross Attention (RCA) module is designed to integrate global contextual information of non-shadow regions into the local branch. We further design a Global-guided Sampling Strategy (GSS) that mitigates patch boundary issues and ensures consistent illumination across shaded and unshaded regions in the recovered image. Comprehensive experiments on three publicly standard datasets ISTD, ISTD+, and SRD have demonstrated the effectiveness of Diff-Shadow. Compared to state-of-the-art methods, our method achieves a significant improvement in terms of PSNR, increasing from 32.33dB to 33.69dB on the ISTD dataset. Jinting Luo, Ru Li 0002, Chengzhi Jiang, Xiaoming Zhang 0008, Mingyan Han, Ting Jiang 0005, Haoqiang Fan, Shuaicheng Liu |
AAAI | 8 |
| 2025 | ISPDiffuser: Learning RAW-to-sRGB Mappings with Texture-Aware Diffusion Models and Histogram-Guided Color ConsistencyabstractRAW-to-sRGB mapping, or the simulation of the traditional camera image signal processor (ISP), aims to generate DSLR-quality sRGB images from raw data captured by smartphone sensors. Despite achieving comparable results to sophisticated handcrafted camera ISP solutions, existing learning-based methods still struggle with detail disparity and color distortion. In this paper, we present ISPDiffuser, a diffusion-based decoupled framework that separates the RAW-to-sRGB mapping into detail reconstruction in grayscale space and color consistency mapping from grayscale to sRGB. Specifically, we propose a texture-aware diffusion model that leverages the generative ability of diffusion models to focus on local detail recovery, in which a texture enrichment loss is further proposed to prompt the diffusion model to generate more intricate texture details. Subsequently, we introduce a histogram-guided color consistency module that utilizes color histogram as guidance to learn precise color information for grayscale to sRGB color consistency mapping, with a color consistency loss designed to constrain the learned color information. Extensive experimental results show that the proposed ISPDiffuser outperforms state-of-the-art competitors both quantitatively and visually. Yang Ren 0001, Hai Jiang 0006, Menglong Yang, Wei Li 0075, Shuaicheng Liu |
AAAI | 5 |
| 2025 | Realistic Noise Synthesis with Diffusion ModelsabstractDeep denoising models require extensive real-world training data, which is challenging to acquire. Current noise synthesis techniques struggle to accurately model complex noise distributions. We propose a novel Realistic Noise Synthesis Diffusor (RNSD) method using diffusion models to address these challenges. By encoding camera settings into a time-aware camera-conditioned affine modulation (TCCAM), RNSD generates more realistic noise distributions under various camera conditions. Additionally, RNSD integrates a multi-scale content-aware module (MCAM), enabling the generation of structured noise with spatial correlations across multiple frequencies. We also introduce Deep Image Prior Sampling (DIPS), a learnable sampling sequence based on depth image prior, which significantly accelerates the sampling process while maintaining the high quality of synthesized noise. Extensive experiments demonstrate that our RNSD method significantly outperforms existing techniques in synthesizing realistic noise under multiple metrics and improving image denoising performance. Qi Wu 0017, Mingyan Han, Ting Jiang 0005, Chengzhi Jiang, Jinting Luo, Man Jiang, Haoqiang Fan, Shuaicheng Liu |
AAAI | 8 |
| 2025 | Single Image Rolling Shutter Removal with Diffusion ModelsabstractWe present RS-Diffusion, the first Diffusion Models-based method for single-frame Rolling Shutter (RS) correction. RS artifacts compromise visual quality of frames due to the row-wise exposure of CMOS sensors. Most previous methods have focused on multi-frame approaches, using temporal information from consecutive frames for the motion rectification. However, few approaches address the more challenging but important single frame RS correction. In this work, we present an ``image-to-motion" framework via diffusion techniques, with a designed patch-attention module. In addition, we present the RS-Real dataset, comprised of captured RS frames alongside their corresponding Global Shutter (GS) ground-truth pairs. The GS frames are corrected from the RS ones, guided by the corresponding Inertial Measurement Unit (IMU) gyroscope data acquired during capture. Experiments show that RS-Diffusion surpasses previous single-frame RS methods, demonstrates the potential of diffusion-based approaches, and provides a valuable dataset for further research. Zhanglei Yang, Haipeng Li 0001, Mingbo Hong, Chen-Lin Zhang, Shuaicheng Liu |
AAAI | 6 |
| 2025 | FlowPolicy: Enabling Fast and Robust 3D Flow-Based Policy via Consistency Flow Matching for Robot ManipulationabstractRobots can acquire complex manipulation skills by learning policies from expert demonstrations, which is often known as vision-based imitation learning. Generating policies based on diffusion and flow matching models has been shown to be effective, particularly in robotic manipulation tasks. However, recursion-based approaches are inference inefficient in working from noise distributions to policy distributions, posing a challenging trade-off between efficiency and quality. This motivates us to propose FlowPolicy, a novel framework for fast policy generation based on consistency flow matching and 3D vision. Our approach refines the flow dynamics by normalizing the self-consistency of the velocity field, enabling the model to derive task execution policies in a single inference step. Specifically, FlowPolicy conditions on the observed 3D point cloud, where consistency flow matching directly defines straight-line flows from different time states to the same action space, while simultaneously constraining their velocity values, that is, we approximate the trajectories from noise to robot actions by normalizing the self-consistency of the velocity field within the action space, thus improving the inference efficiency. We validate the effectiveness of FlowPolicy in Adroit and Metaworld, demonstrating a 7× increase in inference speed while maintaining competitive average success rates compared to state-of-the-art methods. Qinglun Zhang, Zhen Liu 0022, Haoqiang Fan, Guanghui Liu 0001, Bing Zeng 0001, Shuaicheng Liu |
AAAI | 6 |
| 2025 | Learning Hazing to Dehazing: Towards Realistic Haze Generation for Real-World Image DehazingabstractExisting real-world image dehazing methods primarily attempt to fine-tune pre-trained models or adapt their inference procedures, thus heavily relying on the pre-trained models and associated training data. Moreover, restoring heavily distorted information under dense haze requires generative diffusion models, whose potential in de-hazing remains underutilized partly due to their lengthy sampling processes. To address these limitations, we introduce a novel hazing-dehazing pipeline consisting of a Realistic Hazy Image Generation framework (HazeGen) and a Diffusion-based Dehazing framework (DiffDehaze). Specifically, HazeGen harnesses robust generative diffusion priors of real-world hazy images embedded in a pre-trained text-to-image diffusion model. By employing specialized hybrid training and blended sampling strategies, HazeGen produces realistic and diverse hazy images as high-quality training data for DiffDehaze. To alleviate the inefficiency and fidelity concerns associated with diffusion-based methods, DiffDehaze adopts an Accelerated Fidelity-Preserving Sampling process (AccSamp). The core of AccSamp is the Tiled Statistical Alignment Operation (AlignOp), which can provide a clean and faithful dehazing estimate within a small fraction of sampling steps to reduce complexity and enable effective fidelity guidance. Extensive experiments demonstrate the superior dehazing performance and visual quality of our approach over existing methods. The code is available at https://github.com/ruiyi-w/Learning-Hazing-to-Dehazing. Ruiyi Wang, Yushuo Zheng, Chunyi Li 0001, Shuaicheng Liu, Guangtao Zhai, Xiaohong Liu 0001 |
CVPR | 5 |
| 2025 | Learning to See in the Extremely Dark
Hai Jiang 0006, Binhao Guan, Zhen Liu 0022, Xiaohong Liu 0001, Songchen Han, Shuaicheng Liu |
ICCV | 8 |
| 2025 | Estimating 2D Camera Motion with Hybrid Motion Basis
Haipeng Li 0001, Tianhao Zhou, Zhanglei Yang, Yan Chen 0007, Zijing Mao, Shen Cheng, Bing Zeng 0001, Shuaicheng Liu |
ICCV | 9 |
| 2025 | Ultra High-Resolution Image Inpainting with Patch-Based Content Consistency AdapterabstractIn this work, we present Patch-Adapter, an effective framework for high-resolution text-guided image inpainting. Unlike existing methods limited to lower resolutions, our approach achieves 4K+ resolution while maintaining precise content consistency and prompt alignment, two critical challenges in image inpainting that intensify with increasing resolution and texture complexity. Patch-Adapter leverages a two-stage adapter architecture to scale the diffusion model's resolution from 1K to 4K+ without requiring structural overhauls: (1) Dual Context Adapter learns coherence between masked and unmasked regions at reduced resolutions to establish global structural consistency; and (2) Reference Patch Adapter implements a patch-level attention mechanism for full-resolution inpainting, preserving local detail fidelity through adaptive feature fusion. This dual-stage architecture uniquely addresses the scalability gap in high-resolution inpainting by decoupling global semantics from localized refinement. Experiments demonstrate that Patch-Adapter not only resolves artifacts common in large-scale inpainting but also achieves state-of-the-art performance on the OpenImages and Photo-Concept-Bucket datasets, outperforming existing methods in both perceptual quality and text-prompt adherence. Qirui Sun, Wang Luyang, Chaoyu Feng, Jue Wang 0001, Shuaicheng Liu |
ICCV | 10 |
| 2025 | The Parallel Pneumatic Artificial Muscle Platform Based on RBF Neural Network CompensationabstractA two-degree-of-freedom parallel mechanism control system based on an adaptive learning rate and radial basis function (RBF) neural network controller is studied in this paper. The mechanism is composed of four pneumatic artificial muscles(PAM), forming two pairs of antagonistic single-degree-of-freedom joints, which enable two-degree-of-freedom motion along the X and Y axes. The core objective of the system is to automatically output the air pressure values for the X and Y axes based on the input desired angle, driving the joints to precisely reach the specified angle. In this research, dynamic modeling of the two pairs of driving joints composed of four pneumatic muscles was conducted, analyzing the motion characteristics of the system. Subsequently, an RBF neural network was employed to approximate system modeling errors and external disturbances, combined with a PID controller to optimize the driving performance of the pneumatic muscles. The stability of the controller was proven by designing the Lyapunov function, ensuring that the system remains stable during dynamic changes. Finally, simulation experiments were conducted using MATLAB/Simulink to verify the effectiveness of the proposed algorithm. The experimental results demonstrate that the control algorithm enables the actual angle to track the desired angle in real-time, with high control accuracy and stability. This research provides a new solution for the precise control of pneumatic muscle-driven parallel joint systems, with broad application prospects, effectively addressing the limitations of traditional PAM control methods that require precise modeling and suffer from poor robustness. Yuanquan Dai, Ruidong Yu, Mingkang Zi, Shuaicheng Liu, Yinhui Xie |
IROS | 6 |
| 2025 | Coding-Prior Guided Diffusion Network for Video DeblurringabstractWhile recent video deblurring methods have advanced significantly, they often overlook two valuable prior information: (1) motion vectors (MVs) and coding residuals (CRs) from video codecs, which provide efficient inter-frame alignment cues, and (2) the rich real-world knowledge embedded in pre-trained diffusion generative models. We present CPGD-Net, a novel two-stage framework that effectively leverages both coding priors and generative diffusion priors for high-quality deblurring. First, our coding-prior feature propagation (CPFP) module utilizes MVs for efficient frame alignment and CRs to generate attention masks, addressing motion inaccuracies and texture variations. Second, a coding-prior controlled generation (CPC) module network integrates coding priors into a pre-trained diffusion model, guiding it to enhance critical regions and synthesize realistic details. Experiments demonstrate our method achieves state-of-the-art perceptual quality with up to 30% improvement in IQA metrics. The code and the coding-prior-augmented dataset are available at: https://github.com/liuyike422/CPGD-Net. Haipeng Li 0001, Shuaicheng Liu, Bing Zeng 0001 |
ACM Multimedia | 4 |
| 2025 | Learning Arbitrary-Scale RAW Image Downscaling with Wavelet-based Recurrent ReconstructionabstractImage downscaling is critical for efficient storage and transmission of high-resolution (HR) images. Existing learning-based methods focus on performing downscaling within the sRGB domain, which typically suffers from blurred details and unexpected artifacts. RAW images, with their unprocessed photonic information, offer greater flexibility but lack specialized downscaling frameworks. In this paper, we propose a wavelet-based recurrent reconstruction framework that leverages the information lossless attribute of wavelet transformation to fulfill the arbitrary-scale RAW image downscaling in a coarse-to-fine manner, in which the Low-Frequency Arbitrary-Scale Downscaling Module (LASDM) and the High-Frequency Prediction Module (HFPM) are proposed to preserve structural and textural integrity of the reconstructed low-resolution (LR) RAW images, alongside an energy-maximization loss to align high-frequency energy between HR and LR domain. Furthermore, we introduce the Realistic Non-Integer RAW Downscaling (Real-NIRD) dataset, featuring a non-integer downscaling factor of 1.3×, and incorporate it with publicly available datasets with integer factors (2×, 3×, 4×) for comprehensive benchmarking arbitrary-scale image downscaling purposes. Extensive experiments demonstrate that our method outperforms existing state-of-the-art competitors both quantitatively and visually. The code and dataset will be released at https://github.com/RenYangSCU/ASRD. Yang Ren 0001, Hai Jiang 0006, Wei Li 0075, Menglong Yang, Heng Zhang 0042, Zehua Sheng, Qingsheng Ye, Shuaicheng Liu |
ACM Multimedia | 8 |
| 2025 | Hierarchical Neural Semantic Representation for 3D Semantic CorrespondenceabstractThis paper presents a new approach to estimate accurate and robust 3D semantic correspondence with the hierarchical neural semantic representation. Our work has three key contributions. First, we design the hierarchical neural semantic representation (HNSR), which consists of a global semantic feature to capture high-level structure and multi-resolution local geometric features to preserve fine details, by carefully harnessing 3D priors from pre-trained 3D generative models. Second, we design a progressive global-to-local matching strategy, which establishes coarse semantic correspondence using the global semantic feature, then iteratively refines it with local geometric features, yielding accurate and semantically-consistent mappings. Third, our framework is training-free and broadly compatible with various pre-trained 3D generative backbones, demonstrating strong generalization across diverse shape categories. Our method also supports various applications, such as shape co-segmentation, keypoint matching, and texture transfer, and generalizes well to structurally diverse shapes, with promising results even in cross-category scenarios. Both qualitative and quantitative evaluations show that our method outperforms previous state-of-the-art techniques. Keyu Du, Jingyu Hu 0001, Haipeng Li 0001, Hao Xu 0018, Chi-Wing Fu, Shuaicheng Liu |
SIGGRAPH Asia | 7 |
| 2025 | ISPFormer: Learning RAW-to-sRGB mappings with wavelet-based self-attention
Yang Ren 0001, Xinhan Niu, Hai Jiang 0006, Zhen Liu 0022, Ting Jiang 0005, Guanghui Liu 0001, Shuaicheng Liu |
Neurocomputing | 7 |
| 2025 | Unsupervised Global and Local Homography Estimation With Coplanarity-Aware GANabstractUnsupervised methods have received increasing attention in homography learning due to their promising performance and label-free training. However, existing methods do not explicitly consider the plane-induced parallax, making the prediction compromised on multiple planes. In this work, we propose a novel method HomoGAN to guide unsupervised homography estimation to focus on the dominant plane. First, a multi-scale transformer is designed to predict homography from the feature pyramids of input images in a coarse-to-fine fashion. Moreover, we propose an unsupervised GAN to impose coplanarity constraint on the predicted homography, which is realized by using a generator to predict a mask of aligned regions, and then a discriminator to check if two masked feature maps are induced by a single homography. Based on the global homography framework, we extend it to the local mesh-grid homography estimation, namely, MeshHomoGAN, where plane constraints can be enforced on each mesh cell to go beyond a single dominant plane, such that scenes with multiple depth planes can be better aligned. To validate the effectiveness of our method and its components, we conduct extensive experiments on large-scale datasets. Results show that our matching error is 22% lower than previous SOTA methods. Code is available at https://github.com/megvii-research/HomoGAN. Shuaicheng Liu, Mingbo Hong, Nianjin Ye, Chunyu Lin, Bing Zeng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Minimum Latency Deep Online Video Stabilization and Its ExtensionsabstractWe present a novel deep camera path optimization framework for minimum latency online video stabilization. Typically, a stabilization pipeline consists of three steps: motion estimation, path smoothing, and novel view synthesis. Most previous methods concentrate on motion estimation while path optimization receives less attention, particularly in the crucial online setting where future frames are inaccessible. In this work, we adopt off-the-shelf high-quality deep motion models for motion estimation and focus only on the path optimization. Specifically, our camera path smoothing network takes a short 2D camera path in a sliding window as input and outputs the stabilizing warp field of the last frame, which warps the coming frame to its stabilized position. We explore three motion densities: a global single camera path, local mesh-based bundled paths, and dense flow paths. A hybrid loss and an efficient motion smoothing attention (EMSA) module are proposed for spatially and temporally consistent path smoothing. Moreover, we build a motion dataset that contains stable and unstable motion pairs for training. Extensive experiments demonstrate that our method surpasses state-of-the-art online stabilization methods and rivals the performance of offline methods, offering compelling advancements in the field of video stabilization. Shuaicheng Liu, Zhuofan Zhang, Zhen Liu 0022, Ping Tan 0002, Bing Zeng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | StabStitch++: Unsupervised Online Video Stitching With Spatiotemporal Bidirectional WarpsabstractWe retarget video stitching to an emerging issue, named warping shake, which unveils the temporal content shakes induced by sequentially unsmooth warps when extending image stitching to video stitching. Even if the input videos are stable, the stitched video can inevitably cause undesired warping shakes and affect the visual experience. To address this issue, we propose StabStitch++, a novel video stitching framework to realize spatial stitching and temporal stabilization with unsupervised learning simultaneously. First, different from existing learning-based image stitching solutions that typically warp one image to align with another, we suppose a virtual midplane between original image planes and project them onto it. Concretely, we design a differentiable bidirectional decomposition module to disentangle the homography transformation and incorporate it into our spatial warp, evenly spreading alignment burdens and projective distortions across two views. Then, inspired by camera paths in video stabilization, we derive the mathematical expression of stitching trajectories in video stitching by elaborately integrating spatial and temporal warps. Finally, a warp smoothing model is presented to produce stable stitched videos with a hybrid loss to simultaneously encourage content alignment, trajectory smoothness, and online collaboration. Compared with StabStitch that sacrifices alignment for stabilization, StabStitch++ makes no compromise and optimizes both of them simultaneously, especially in the online mode. To establish an evaluation benchmark and train the learning framework, we build a video stitching dataset with a rich diversity in camera motions and scenes. Experiments exhibit that StabStitch++ surpasses current solutions in stitching performance, robustness, and efficiency, offering compelling advancements in this field by building a real-time online video stitching system. Lang Nie, Chunyu Lin, Kang Liao, Yun Zhang 0024, Shuaicheng Liu, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | HandBooster+: Boosting 3D Hand-Mesh Reconstruction From Data Synthesis to Progressive Multi-Hypothesis AggregationabstractRobustly reconstructing 3D hand mesh from a single image is very challenging, due to (i) the lack of diversity in existing real-world datasets and (ii) the ambiguity in occluded hand regions. While data synthesis helps relieve issue (i), the syn-to-real gap still hinders its usage. For issue (ii), most previous works produce deterministic results while other probabilistic methods rely on ground truths to choose the best hypothesis. In this work, we explore the diffusion model to alleviate these problems by collectively considering two perspectives: (i) conditional synthesis and sampling approach for realistic data generation and (ii) probabilistic modeling with progressive multi-hypothesis aggregation. First, we present HandBooster, a new approach to uplift the data diversity by training a conditional generative space on hand-object interactions and sampling the space to synthesize effective data with reliable 3D annotations and diverse hand appearances, poses, views, and backgrounds. Second, we design HandBooster+, a probabilistic diffusion-based model to further boost the 3D hand-mesh reconstruction performance by progressively aggregating the multiple hypotheses. Extensive experimental results show that our method significantly improves several baselines and achieves SOTA on the HO3D and DexYCB benchmarks. Hao Xu 0018, Haipeng Li 0001, Yinqiao Wang, Shuaicheng Liu, Chi-Wing Fu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Kernel Reformulation With Deep Constrained Least Squares for Blind Image Super-ResolutionabstractThis work proposes to learn blind image super-resolution (SR) using deep constrained least squares deconvolution with low-resolution (LR) space kernels. Our method recovers the high-resolution (HR) image with a kernel estimation step and a kernel-based image restoration process. Specifically, we first reformulate the classical degradation model to transfer the deblurring kernel estimation into the LR space. We show that the LR space kernel has a closed-form solution given a pair of LR-HR images, which can be learned without ground truth kernels. Next, we introduce a dynamic deep linear filter module, which can generate deblurring kernel weights adaptively. Subsequently, the estimated kernel is integrated with a deep constrained least square filtering module to produce clean features. For reconstruction, we adopt a dual-path structured SR network that inputs both the deblurred feature and the original feature to suppress deconvolution artifacts. Finally, we learn discriminative features for deblurring and then restore the HR image in a single branch, producing a lighter weight network that can achieve comparable performance while only using 56% parameters and 60% inference time. Extensive experiments on both synthetic and real-world datasets demonstrate that our method achieves better accuracy and visual improvements against state-of-the-art approaches. Ziwei Luo 0002, Bing Zeng 0001, Shuaicheng Liu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Multi-Frame Rolling Shutter Correction With Diffusion Models
Zhanglei Yang, Haipeng Li 0001, Shen Cheng, Mingbo Hong, Bing Zeng 0001, Shuaicheng Liu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Hand-Shadow PoserabstractHand shadow art is a captivating art form, creatively using hand shadows to reproduce expressive shapes on the wall. In this work, we study an inverse problem: given a target shape, find the poses of left and right hands that together best produce a shadow resembling the input. This problem is nontrivial, since the design space of 3D hand poses is huge while being restrictive due to anatomical constraints. Also, we need to attend to the input's shape and crucial features, though the input is colorless and textureless. To meet these challenges, we design Hand-Shadow Poser, a three-stage pipeline, to decouple the anatomical constraints (by hand) and semantic constraints (by shadow shape): (i) a generative hand assignment module to explore diverse but reasonable left/right-hand shape hypotheses; (ii) a generalized hand-shadow alignment module to infer coarse hand poses with a similarity-driven strategy for selecting hypotheses; and (iii) a shadow-feature-aware refinement module to optimize the hand poses for physical plausibility and shadow feature preservation. Further, we design our pipeline to be trainable on generic public hand data, thus avoiding the need for any specialized training dataset. For method validation, we build a benchmark of 210 diverse shadow shapes of varying complexity and a comprehensive set of metrics, including a novel DINOv2-based evaluation metric. Through extensive comparisons with multiple baselines and user studies, our approach is demonstrated to effectively generate bimanual hand poses for a large variety of hand shapes for over 85% of the benchmark cases. Hao Xu 0018, Yinqiao Wang, Niloy J. Mitra, Shuaicheng Liu, Pheng-Ann Heng, Chi-Wing Fu |
ACM Trans. Graph. | 4 |
| 2024 | SpectralNeRF: Physically Based Spectral Rendering with Neural Radiance FieldabstractIn this paper, we propose SpectralNeRF, an end-to-end Neural Radiance Field (NeRF)-based architecture for high-quality physically based rendering from a novel spectral perspective. We modify the classical spectral rendering into two main steps, 1) the generation of a series of spectrum maps spanning different wavelengths, 2) the combination of these spectrum maps for the RGB output. Our SpectralNeRF follows these two steps through the proposed multi-layer perceptron (MLP)-based architecture (SpectralMLP) and Spectrum Attention UNet (SAUNet). Given the ray origin and the ray direction, the SpectralMLP constructs the spectral radiance field to obtain spectrum maps of novel views, which are then sent to the SAUNet to produce RGB images of white-light illumination. Applying NeRF to build up the spectral rendering is a more physically-based way from the perspective of ray-tracing. Further, the spectral radiance fields decompose difficult scenes and improve the performance of NeRF-based methods. Comprehensive experimental results demonstrate the proposed SpectralNeRF is superior to recent NeRF-based methods when synthesizing new views on synthetic and real datasets. The codes and datasets are available at https://github.com/liru0126/SpectralNeRF. Ru Li 0002, Guanghui Liu 0001, Shengping Zhang, Bing Zeng 0001, Shuaicheng Liu |
AAAI | 6 |
| 2024 | Efficient Meshflow and Optical Flow Estimation from Event CamerasabstractIn this paper, we explore the problem of event-based meshflow estimation, a novel task that involves predicting a spatially smooth sparse motion field from event cameras. To start, we generate a large-scale High-Resolution Event Meshflow (HREM) dataset, which showcases its superiority by encompassing the merits of high resolution at 1280×720, handling dynamic objects and complex motion patterns, and offering both optical flow and meshflow labels. These aspects have not been fully explored in previous works. Besides, we propose Efficient Event-based MeshFlow (EEMFlow) network, a lightweight model featuring a specially crafted encoder-decoder architecture to facilitate swift and accurate meshflow estimation. Furthermore, we upgrade EEMFlow network to support dense event optical flow, in which a Confidence-induced Detail Completion (CDC) module is proposed to preserve sharp motion boundaries. We conduct comprehensive experiments to show the exceptional performance and runtime efficiency (39× faster) of our EEMFlow model compared to recent state-of-the-art flow methods. Our code is available at https://github.com/boomluo02/EEMFlow. Xinglong Luo, Ao Luo, Zhengning Wang, Chunyu Lin, Bing Zeng 0001, Shuaicheng Liu |
CVPR | 6 |
| 2024 | FlowDiffuser: Advancing Optical Flow Estimation with Diffusion ModelsabstractOptical flow estimation, a process of predicting pixel-wise displacement between consecutive frames, has commonly been approached as a regression task in the age of deep learning. Despite notable advancements, this de facto paradigm unfortunately falls short in generalization performance when trained on synthetic or constrained data. Pioneering a paradigm shift, we reformulate optical flow estimation as a conditional flow generation challenge, unveiling FlowDiffuser — a new family of optical flow models that could have stronger learning and generalization capabilities. FlowDiffuser estimates optical flow through a ‘noise-to-flow’ strategy, progressively eliminating noise from randomly generated flows conditioned on the provided pairs. To optimize accuracy and efficiency, our FlowDiffuser incorporates a novel Conditional Recurrent Denoising Decoder (Conditional-RDD), streamlining the flow estimation process. It incorporates a unique Hidden State Denoising (HSD) paradigm, effectively leveraging the information from previous time steps. Moreover, FlowDiffuser can be easily integrated into existing flow networks, leading to significant improvements in performance metrics compared to conventional implementations. Experiments on challenging benchmarks, including Sintel and KITTI, demonstrate the effectiveness of our FlowDiffuser with superior performance to existing state-of-the-art models. Code is available at https://github.com/LA30/FlowDiffuser. Ao Luo, Fan Yang 0054, Jiangyu Liu, Haoqiang Fan, Shuaicheng Liu |
CVPR | 6 |
| 2024 | HandBooster: Boosting 3D Hand-Mesh Reconstruction by Conditional Synthesis and Sampling of Hand-Object InteractionsabstractReconstructing 3D hand mesh robustly from a single image is very challenging, due to the lack of diversity in existing real-world datasets. While data synthesis helps relieve the issue, the syn-to-real gap still hinders its usage. In this work, we present HandBooster, a new approach to uplift the data diversity and boost the 3D hand-mesh reconstruction performance by training a conditional generative space on hand-object interactions and purposely sampling the space to synthesize effective data samples. First, we construct versatile content-aware conditions to guide a diffusion model to produce realistic images with diverse hand appearances, poses, views, and backgrounds; favorably, accurate 3D an-notations are obtained for free. Then, we design a novel condition creator based on our similarity-aware distribution sampling strategies to deliberately find novel and realistic interaction poses that are distinctive from the training set. Equipped with our method, several baselines can be significantly improved beyond the SOTA on the HO3D and DexYCB benchmarks. Our code will be released on https://github.com/hxwork/HandBooster_Pytorch. Hao Xu 0018, Haipeng Li 0001, Yinqiao Wang, Shuaicheng Liu, Chi-Wing Fu |
CVPR | 4 |
| 2024 | RecDiffusion: Rectangling for Image Stitching with Diffusion ModelsabstractImage stitching from different captures often results in non-rectangular boundaries, which is often considered un-appealing. To solve non-rectangular boundaries, current solutions involve cropping, which discards image content, inpainting, which can introduce unrelated content, or warping, which can distort non-linear features and introduce artifacts. To overcome these issues, we introduce a novel diffusion-based learning framework, RecDiffusion, for image stitching rectangling. This framework combines Motion Diffusion Models (MDM) to generate motion fields, ef-fectively transitioning from the stitched image's irregular borders to a geometrically corrected intermediary. Fol-lowed by Content Diffusion Models (CDM) for image de-tail refinement. Notably, our sampling process utilizes a weighted map to identify regions needing correction during each iteration of CDM. Our RecDiffusion ensures geomet-ric accuracy and overall visual appeal, surpassing all pre-vious methods in both quantitative and qualitative measures when evaluated on public benchmarks. Code is released at https://github.com/haippp/RecDiffusion. Tianhao Zhou, Haipeng Li 0001, Ao Luo, Chen-Lin Zhang, Bing Zeng 0001, Shuaicheng Liu |
CVPR | 8 |
| 2024 | PointRegGPT: Boosting 3D Point Cloud Registration Using Generative Point-Cloud Pairs for Training
Suyi Chen, Hao Xu 0018, Haipeng Li 0001, Kunming Luo, Guanghui Liu 0001, Chi-Wing Fu, Ping Tan 0002, Shuaicheng Liu |
ECCV (51) | 8 |
| 2024 | LightenDiffusion: Unsupervised Low-Light Image Enhancement with Latent-Retinex Diffusion Models
Hai Jiang 0006, Ao Luo, Xiaohong Liu 0001, Songchen Han, Shuaicheng Liu |
ECCV (48) | 5 |
| 2024 | Eliminating Warping Shakes for Unsupervised Online Video Stitching
Lang Nie, Chunyu Lin, Kang Liao, Yun Zhang 0024, Shuaicheng Liu, Rui Ai 0001, Yao Zhao 0001 |
ECCV (4) | 5 |
| 2024 | Neural Spectral Decomposition for Dataset Distillation
Shaolei Yang, Shen Cheng, Mingbo Hong, Haoqiang Fan, Shuaicheng Liu |
ECCV (52) | 6 |
| 2024 | GLARE: Low Light Image Enhancement via Generative Latent Feature Based Codebook Retrieval
Han Zhou 0003, Wei Dong 0011, Xiaohong Liu 0001, Shuaicheng Liu, Xiongkuo Min, Guangtao Zhai, Jun Chen 0005 |
ECCV (48) | 4 |
| 2024 | Efficient Single Image Super-Resolution with Entropy Attention and Receptive Field AugmentationabstractTransformer-based deep models for single image super-resolution (SISR) have greatly improved the performance of lightweight SISR tasks in recent years. However, they often suffer from heavy computational burden and slow inference due to the complex calculation of multi-head self-attention (MSA), seriously hindering their practical application and deployment. In this work, we present an efficient SR model to mitigate the dilemma between model efficiency and SR performance, which is dubbed Entropy Attention and Receptive Field Augmentation network (EARFA), and composed of a novel entropy attention (EA) and a shifting large kernel attention (SLKA). From the perspective of information theory, EA increases the entropy of intermediate features conditioned on a Gaussian distribution, providing more informative input for subsequent reasoning. On the other hand, SLKA extends the receptive field of SR models with the assistance of channel shifting, which also favors to boost the diversity of hierarchical features. Since the implementation of EA and SLKA does not involve complex computations (such as extensive matrix multiplications), the proposed method can achieve faster nonlinear inference than Transformer-based SR models while maintaining better SR performance. Extensive experiments show that the proposed model can significantly reduce the delay of model inference while achieving the SR performance comparable with other advanced models. Xiaole Zhao, Linze Li 0001, Chengxing Xie, Xiaoming Zhang 0008, Ting Jiang 0005, Shuaicheng Liu, Tianrui Li 0001 |
ACM Multimedia | 7 |
| 2024 | You Only Look Around: Learning Illumination-Invariant Feature for Low-light Object DetectionabstractIn this paper, we introduce YOLA, a novel framework for object detection in low-light scenarios. Unlike previous works, we propose to tackle this challenging problem from the perspective of feature learning. Specifically, we propose to learn illumination-invariant features through the Lambertian image formation model. We observe that, under the Lambertian assumption, it is feasible to approximate illumination-invariant feature maps by exploiting the interrelationships between neighboring color channels and spatially adjacent pixels. By incorporating additional constraints, these relationships can be characterized in the form of convolutional kernels, which can be trained in a detection-driven manner within a network. Towards this end, we introduce a novel module dedicated to the extraction of illumination-invariant features from low-light images, which can be easily integrated into existing object detection frameworks. Our empirical findings reveal significant improvements in low-light object detection tasks, as well as promising results in both well-lit and over-lit scenarios. Mingbo Hong, Shen Cheng, Haoqiang Fan, Shuaicheng Liu |
NeurIPS | 5 |
| 2024 | GyroFlow+: Gyroscope-Guided Unsupervised Deep Homography and Optical Flow Learning
Haipeng Li 0001, Kunming Luo, Bing Zeng 0001, Shuaicheng Liu |
Int. J. Comput. Vis. | 4 |
| 2024 | Semi-Supervised Coupled Thin-Plate Spline Model for Rotation Correction and BeyondabstractThin-plate spline (TPS) is a principal warp that allows for representing elastic, nonlinear transformation with control point motions. With the increase of control points, the warp becomes increasingly flexible but usually encounters a bottleneck caused by undesired issues, e.g., content distortion. In this paper, we explore generic applications of TPS in single-image-based warping tasks, such as rotation correction, rectangling, and portrait correction. To break this bottleneck, we propose the coupled thin-plate spline model (CoupledTPS), which iteratively couples multiple TPS with limited control points into a more flexible and powerful transformation. Concretely, we first design an iterative search to predict new control points according to the current latent condition. Then, we present the warping flow as a bridge for the coupling of different TPS transformations, effectively eliminating interpolation errors caused by multiple warps. Besides, in light of the laborious annotation cost, we develop a semi-supervised learning scheme to improve warping quality by exploiting unlabeled data. It is formulated through dual transformation between the searched control points of unlabeled data and its graphic augmentation, yielding an implicit correction consistency constraint. Finally, we collect massive unlabeled data to exhibit the benefit of our semi-supervised scheme in rotation correction. Extensive experiments demonstrate the superiority and universality of CoupledTPS over the existing State-of-the-Art (SoTA) solutions for rotation correction and beyond. Lang Nie, Chunyu Lin, Kang Liao, Shuaicheng Liu, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Reconstruction flow recurrent network for compressed video quality enhancement
Zhengning Wang, Xuhang Liu, Chuan Wang 0001, Ting Jiang 0005, Tianjiao Zeng, Zhenni Zeng, Guoqing Wang 0001, Shuaicheng Liu |
Pattern Recognit. | 8 |
| 2024 | GLOCAL: A self-supervised learning framework for global and local motion estimation
Yihao Zheng 0002, Kunming Luo, Shuaicheng Liu, Zun Li 0001, Ye Xiang, Lifang Wu, Bing Zeng 0001, Chang Wen Chen |
Pattern Recognit. Lett. | 3 |
| 2024 | PBR-GAN: Imitating Physically-Based Rendering With Generative Adversarial NetworksabstractWe propose a Generative Adversarial Network (GAN)-based architecture for achieving high-quality physically based rendering (PBR). Conventional PBR relies heavily on ray tracing, which is computationally expensive in complicated environments. Some recent deep learning-based methods can improve efficiency but cannot deal with illumination variation well. In this paper, we propose PBR-GAN, an end-to-end GAN-based network that solves these problems while generating natural photo-realistic images. Two encoders (the shading encoder and albedo encoder) and two decoders (the image decoder and light decoder) are introduced to achieve our target. The two encoders and the image decoder constitute the generator that learns the mapping between the generated domain and the real domain. The light decoder produces light maps that pay more attention to the highlight and shadow regions. The discriminator aims to optimize the generator by distinguishing target images from the generated ones. Three novel loss items, concentrating on domain translation, overall shading preservation, and light map estimation, are proposed to optimize the photo-realistic outputs. Furthermore, a real dataset is collected to provide realistic information for training GAN architecture. Extensive experiments indicate that PBR-GAN can preserve the illumination variation and improve the image perceptual quality. Ru Li 0002, Peng Dai 0003, Guanghui Liu 0001, Shengping Zhang, Bing Zeng 0001, Shuaicheng Liu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | CodingHomo: Bootstrapping Deep Homography With Video CodingabstractHomography estimation is a fundamental task in computer vision with applications in diverse fields. Recent advances in deep learning have improved homography estimation, particularly with unsupervised learning approaches, offering increased robustness and generalizability. However, accurately predicting homography, especially in complex motions, remains a challenge. In response, this work introduces a novel method leveraging video coding, particularly by harnessing inherent motion vectors (MVs) present in videos. We present CodingHomo, an unsupervised framework for homography estimation. Our framework features a Mask-Guided Fusion (MGF) module that identifies and utilizes beneficial features among the MVs, thereby enhancing the accuracy of homography prediction. Additionally, the Mask-Guided Homography Estimation (MGHE) module is presented for eliminating undesired features in the coarse-to-fine homography refinement process. CodingHomo outperforms existing state-of-the-art unsupervised methods, delivering good robustness and generalizability. The code and dataset are available at:https://github.com/liuyike422/CodingHomo. Haipeng Li 0001, Shuaicheng Liu, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Single-Image-Based Deep Learning for Segmentation of Early Esophageal Cancer LesionsabstractAccurate segmentation of lesions is crucial for diagnosis and treatment of early esophageal cancer (EEC). However, neither traditional nor deep learning-based methods up to today can meet the clinical requirements, with the mean Dice score - the most important metric in medical image analysis - hardly exceeding 0.75. In this paper, we present a novel deep learning approach for segmenting EEC lesions. Our method stands out for its uniqueness, as it relies solely on a single input image from a patient, forming the so-called "You-Only-Have-One" (YOHO) framework. On one hand, this "one-image-one-network" learning ensures complete patient privacy as it does not use any images from other patients as the training data. On the other hand, it avoids nearly all generalization-related problems since each trained network is applied only to the same input image itself. In particular, we can push the training to "over-fitting" as much as possible to increase the segmentation accuracy. Our technical details include an interaction with clinical doctors to utilize their expertise, a geometry-based data augmentation over a single lesion image to generate the training dataset (the biggest novelty), and an edge-enhanced UNet. We have evaluated YOHO over an EEC dataset collected by ourselves and achieved a mean Dice score of 0.888, which is much higher as compared to the existing deep-learning methods, thus representing a significant advance toward clinical applications. The code and dataset are available at: https://github.com/lhaippp/YOHO. Haipeng Li 0001, Dingrui Liu, Shuaicheng Liu, Tao Gan, Nini Rao, Jinlin Yang, Bing Zeng 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | DMHomo: Learning Homography with Diffusion ModelsabstractSupervised homography estimation methods face a challenge due to the lack of adequate labeled training data. To address this issue, we propose DMHomo , a diffusion model-based framework for supervised homography learning. This framework generates image pairs with accurate labels, realistic image content, and realistic interval motion, ensuring that they satisfy adequate pairs. We utilize unlabeled image pairs with pseudo labels such as homography and dominant plane masks, computed from existing methods, to train a diffusion model that generates a supervised training dataset. To further enhance performance, we introduce a new probabilistic mask loss, which identifies outlier regions through supervised training, and an iterative mechanism to optimize the generative and homography models successively. Our experimental results demonstrate that DMHomo effectively overcomes the scarcity of qualified datasets in supervised homography learning and improves generalization to real-world scenes. The code and dataset are available at GitHub ( https://github.com/lhaippp/DMHomo ). Haipeng Li 0001, Hai Jiang 0006, Ao Luo, Ping Tan 0002, Haoqiang Fan, Bing Zeng 0001, Shuaicheng Liu |
ACM Trans. Graph. | 7 |
| 2023 | Semi-supervised Deep Large-Baseline Homography Estimation with Progressive Equivalence ConstraintabstractHomography estimation is erroneous in the case of large-baseline due to the low image overlay and limited receptive field. To address it, we propose a progressive estimation strategy by converting large-baseline homography into multiple intermediate ones, cumulatively multiplying these intermediate items can reconstruct the initial homography. Meanwhile, a semi-supervised homography identity loss, which consists of two components: a supervised objective and an unsupervised objective, is introduced. The first supervised loss is acting to optimize intermediate homographies, while the second unsupervised one helps to estimate a large-baseline homography without photometric losses. To validate our method, we propose a large-scale dataset that covers regular and challenging scenes. Experiments show that our method achieves state-of-the-art performance in large-baseline scenes while keeping competitive performance in small-baseline scenes. Code and dataset are available at https://github.com/megvii-research/LBHomo. Hai Jiang 0006, Haipeng Li 0001, Songchen Han, Shuaicheng Liu |
AAAI | 5 |
| 2023 | Low-Light Image Enhancement with Illumination-Aware Gamma Correction and Complete Image Modelling NetworkabstractThis paper presents a novel network structure with illumination-aware gamma correction and complete image modelling to solve the low-light image enhancement problem. Low-light environments usually lead to less informative large-scale dark areas, directly learning deep representations from low-light images is insensitive to recovering normal illumination. We propose to integrate the effectiveness of gamma correction with the strong modelling capacities of deep networks, which enables the correction factor gamma to be learned in a coarse to elaborate manner via adaptively perceiving the deviated illumination. Because exponential operation introduces high computational complexity, we propose to use Taylor Series to approximate gamma correction, accelerating the training and inference speed. Dark areas usually occupy large scales in low-light images, common local modelling structures, e.g., CNN, SwinIR, are thus insufficient to recover accurate illumination across whole low-light images. We propose a novel Transformer block to completely simulate the dependencies of all pixels across images via a local-to-global hierarchical attention mechanism, so that dark areas could be inferred by borrowing the information from far informative regions in a highly effective manner. Extensive experiments on several benchmark datasets demonstrate that our approach outperforms state-of-the-art methods. Yinglong Wang 0002, Zhen Liu 0022, Jianzhuang Liu, Songcen Xu, Shuaicheng Liu |
ICCV | 5 |
| 2023 | Supervised Homography Learning with Realistic Dataset GenerationabstractIn this paper, we propose an iterative framework, which consists of two phases: a generation phase and a training phase, to generate realistic training data and yield a supervised homography network. In the generation phase, given an unlabeled image pair, we utilize the pre-estimated dominant plane masks and homography of the pair, along with another sampled homography that serves as ground truth to generate a new labeled training pair with realistic motion. In the training phase, the generated data is used to train the supervised homography network, in which the training data is refined via a content consistency module and a quality assessment module. Once an iteration is finished, the trained network is used in the next data generation phase to update the pre-estimated homography. Through such an iterative strategy, the quality of the dataset and the performance of the network can be gradually and simultaneously improved. Experimental results show that our method achieves state-of-the-art performance and existing supervised methods can be also improved based on the generated dataset. Code and dataset are available at https://github.com/JianghaiSCU/RealSH. Hai Jiang 0006, Haipeng Li 0001, Songchen Han, Haoqiang Fan, Bing Zeng 0001, Shuaicheng Liu |
ICCV | 6 |
| 2023 | SIRA-PCR: Sim-to-Real Adaptation for 3D Point Cloud RegistrationabstractPoint cloud registration is essential for many applications. However, existing real datasets require extremely tedious and costly annotations, yet may not provide accurate camera poses. For the synthetic datasets, they are mainly object-level, so the trained models may not generalize well to real scenes. We design SIRA-PCR, a new approach to 3D point cloud registration. First, we build a synthetic scene-level 3D registration dataset, specifically designed with physically-based and random strategies to arrange diverse objects. Second, we account for variations in different sensing mechanisms and layout placements, then formulate a sim-to-real adaptation framework with an adaptive re-sample module to simulate patterns in real point clouds. To our best knowledge, this is the first work that explores sim-to-real adaptation for point cloud registration. Extensive experiments show the SOTA performance of SIRA-PCR on widely-used indoor and out-door datasets. The code and dataset will be released on https://github.com/Chen-Suyi/SIRA_Pytorc.h Suyi Chen, Hao Xu 0018, Ru Li 0002, Guanghui Liu 0001, Chi-Wing Fu, Shuaicheng Liu |
ICCV | 6 |
| 2023 | Explicit Motion Disentangling for Efficient Optical Flow EstimationabstractIn this paper, we propose a novel framework for optical flow estimation that achieves a good balance between performance and efficiency. Our approach involves disentangling global motion learning from local flow estimation, treating global matching and local refinement as separate stages. We offer two key insights: First, the multi-scale 4D cost-volume based recurrent flow decoder is computationally expensive and unnecessary for handling small displacement. With the separation, we can utilize lightweight methods for both parts and maintain similar performance. Second, a dense and robust global matching is essential for both flow initialization as well as stable and fast convergence for the refinement stage. Towards this end, we introduce EMD-Flow, a framework that explicitly separates global motion estimation from the recurrent refinement stage. We propose two novel modules: Multi-scale Motion Aggregation (MMA) and Confidence-induced Flow Propagation (CFP). These modules leverage cross-scale matching prior and self-contained confidence maps to handle the ambiguities of dense matching in a global manner, generating a dense initial flow. Additionally, a lightweight decoding module is followed to handle small displacements, resulting in an efficient yet robust flow estimation framework. We further conduct comprehensive experiments on standard optical flow benchmarks with the proposed framework, and the experimental results demonstrate its superior balance between performance and runtime. Code is available at https://github.com/gddcx/EMD-Flow. Changxing Deng, Ao Luo, Shaodan Ma, Jiangyu Liu, Shuaicheng Liu |
ICCV | 6 |
| 2023 | MEFLUT: Unsupervised 1D Lookup Tables for Multi-exposure Image FusionabstractIn this paper, we introduce a new approach for high-quality multi-exposure image fusion (MEF). We show that the fusion weights of an exposure can be encoded into a 1D lookup table (LUT), which takes pixel intensity value as input and produces fusion weight as output. We learn one 1D LUT for each exposure, then all the pixels from different exposures can query 1D LUT of that exposure independently for high-quality and efficient fusion. Specifically, to learn these 1D LUTs, we involve attention mechanism in various dimensions including frame, channel and spatial ones into the MEF task so as to bring us significant quality improvement over the state-of-the-art (SOTA). In addition, we collect a new MEF dataset consisting of 960 samples, 155 of which are manually tuned by professionals as ground-truth for evaluation. Our network is trained by this dataset in an unsupervised manner. Extensive experiments are conducted to demonstrate the effectiveness of all the newly proposed components, and results show that our approach outperforms the SOTA in our and another representative dataset SICE, both qualitatively and quantitatively. Moreover, our 1D LUT approach takes less than 4ms to run a 4K image on a PC GPU. Given its high quality, efficiency and robustness, our method has been shipped into millions of Android mobiles across multiple brands world-wide. Code is available at: https://github.com/Hedlen/MEFLUT. Ting Jiang 0005, Chuan Wang 0001, Xinpeng Li 0002, Ru Li 0002, Haoqiang Fan, Shuaicheng Liu |
ICCV | 6 |
| 2023 | Uncertainty Guided Adaptive Warping for Robust and Efficient Stereo MatchingabstractCorrelation based stereo matching has achieved outstanding performance, which pursues cost volume between two feature maps. Unfortunately, current methods with a fixed model do not work uniformly well across various datasets, greatly limiting their real-world applicability. To tackle this issue, this paper proposes a new perspective to dynamically calculate correlation for robust stereo matching. A novel Uncertainty Guided Adaptive Correlation (UGAC) module is introduced to robustly adapt the same model for different scenarios. Specifically, a variance-based uncertainty estimation is employed to adaptively adjust the sampling area during warping operation. Additionally, we improve the traditional non-parametric warping with learnable parameters, such that the position-specific weights can be learned. We show that by empowering the recurrent network with the UGAC module, stereo matching can be exploited more robustly and effectively. Extensive experiments demonstrate that our method achieves state-of-the-art performance over the ETH3D, KITTI, and Middlebury datasets when employing the same fixed model over these datasets without any retraining procedure. To target real-time applications, we further design a lightweight model based on UGAC, which also outperforms other methods over KITTI benchmarks with only 0.6 M parameters. Junpeng Jing, Jiankun Li, Pengfei Xiong, Jiangyu Liu, Shuaicheng Liu, Xin Deng 0002, Mai Xu, Lai Jiang 0004, Leonid Sigal |
ICCV | 5 |
| 2023 | Learning Optical Flow from Event Camera with Rendered DatasetabstractWe study the problem of estimating optical flow from event cameras. One important issue is how to build a high-quality event-flow dataset with accurate event values and flow labels. Previous datasets are created by either capturing real scenes by event cameras or synthesizing from images with pasted foreground objects. The former case can produce real event values but with calculated flow labels, which are sparse and inaccurate. The latter case can generate dense flow labels but the interpolated events are prone to errors. In this work, we propose to render a physically correct event-flow dataset using computer graphics models. In particular, we first create indoor and outdoor 3D scenes by Blender with rich scene content variations. Second, diverse camera motions are included for the virtual capturing, producing images and accurate flow labels. Third, we render high-framerate videos between images for accurate events. The rendered dataset can adjust the density of events, based on which we further introduce an adaptive density module (ADM). Experiments show that our proposed dataset can facilitate event-flow learning, whereas previous approaches when trained on our dataset can improve their performances constantly by a relatively large margin. In addition, event-flow pipelines when equipped with our ADM can further improve performances. Our code is available at https://github.com/boomluo02/ADMFlow. Xinglong Luo, Kunming Luo, Ao Luo, Zhengning Wang, Ping Tan 0002, Shuaicheng Liu |
ICCV | 6 |
| 2023 | GAFlow: Incorporating Gaussian Attention into Optical FlowabstractOptical flow, or the estimation of motion fields from image sequences, is one of the fundamental problems in computer vision. Unlike most pixel-wise tasks that aim at achieving consistent representations of the same category, optical flow raises extra demands for obtaining local discrimination and smoothness, which yet is not fully explored by existing approaches. In this paper, we push Gaussian Attention (GA) into the optical flow models to accentuate local properties during representation learning and enforce the motion affinity during matching. Specifically, we introduce a novel Gaussian-Constrained Layer (GCL) which can be easily plugged into existing Transformer blocks to highlight the local neighborhood that contains fine-grained structural information. Moreover, for reliable motion analysis, we provide a new Gaussian-Guided Attention Module (GGAM) which not only inherits properties from Gaussian distribution to instinctively revolve around the neighbor fields of each point but also is empowered to put the emphasis on contextually related regions during matching. Our fully-equipped model, namely Gaussian Attention Flow network (GAFlow), naturally incorporates a series of novel Gaussian-based modules into the conventional optical flow framework for reliable motion analysis. Extensive experiments on standard optical flow datasets consistently demonstrate the exceptional performance of the proposed approach in terms of both generalization ability evaluation and online benchmark testing. Code is available at https://github.com/LA30/GAFlow. Ao Luo, Fan Yang 0054, Xin Li 0005, Lang Nie, Chunyu Lin, Haoqiang Fan, Shuaicheng Liu |
ICCV | 7 |
| 2023 | Parallax-Tolerant Unsupervised Deep Image StitchingabstractTraditional image stitching approaches tend to leverage increasingly complex geometric features (e.g., point, line, edge, etc.) for better performance. However, these hand-crafted features are only suitable for specific natural scenes with adequate geometric structures. In contrast, deep stitching schemes overcome adverse conditions by adaptively learning robust semantic features, but they cannot handle large-parallax cases.To solve these issues, we propose a parallax-tolerant unsupervised deep image stitching technique. First, we propose a robust and flexible warp to model the image registration from global homography to local thin-plate spline motion. It provides accurate alignment for overlapping regions and shape preservation for non-overlapping regions by joint optimization concerning alignment and distortion. Subsequently, to improve the generalization capability, we design a simple but effective iterative strategy to enhance the warp adaption in cross-dataset and cross-resolution applications. Finally, to further eliminate the parallax artifacts, we propose to composite the stitched image seamlessly by unsupervised learning for seam-driven composition masks. Compared with existing methods, our solution is parallax-tolerant and free from laborious designs of complicated geometric features for specific scenes. Extensive experiments show our superiority over the SoTA methods, both quantitatively and qualitatively. The code is available at https://github.com/nie-lang/UDIS2. Lang Nie, Chunyu Lin, Kang Liao, Shuaicheng Liu, Yao Zhao 0001 |
ICCV | 4 |
| 2023 | AccFlow: Backward Accumulation for Long-Range Optical FlowabstractRecent deep learning-based optical flow estimators have exhibited impressive performance in generating local flows between consecutive frames. However, the estimation of long-range flows between distant frames, particularly under complex object deformation and large motion occlusion, remains a challenging task. One promising solution is to accumulate local flows explicitly or implicitly to obtain the desired long-range flow. Nevertheless, the accumulation errors and flow misalignment can hinder the effectiveness of this approach. This paper proposes a novel recurrent framework called AccFlow, which recursively backward accumulates local flows using a deformable module called as AccPlus. In addition, an adaptive blending module is designed along with AccPlus to alleviate the occlusion effect by backward accumulation and rectify the accumulation error. Notably, we demonstrate the superiority of backward accumulation over conventional forward accumulation, which to the best of our knowledge has not been explicitly established before. To train and evaluate the proposed AccFlow, we have constructed a large-scale high-quality dataset named CVO, which provides ground-truth optical flow labels between adjacent and distant frames. Extensive experiments validate the effectiveness of AccFlow in handling long-range optical flow estimation. Codes are available at https://github.com/mulns/AccFlow. Guangyang Wu, Xiaohong Liu 0001, Kunming Luo, Qingqing Zheng, Shuaicheng Liu, Xinyang Jiang, Guangtao Zhai, Wenyi Wang 0005 |
ICCV | 6 |
| 2023 | Deep Homography Mixture for Single Image Rolling Shutter CorrectionabstractWe present a deep homography mixture motion model for single image rolling shutter correction. Rolling shutter (RS) effects are often caused by row-wise exposure delay in the widely adopted CMOS sensor. Previous methods often require more than one frame for the correction, leading to data quality requirements. Few approaches address the more challenging task of single image RS correction, which often adopt designs like trajectory estimation or long rectangular kernels, to learn the camera motion parameters of an RS image, to restore the global shutter (GS) image. In this work, we adopt a more straightforward method to learn deep homography mixture motion between an RS image and its corresponding GS image, without large solution space or strict restrictions on image features. We show that dividing an image into blocks with a Gaussian weight of block scanlines fits well for the RS setting. Moreover, instead of directly learning the motion mapping, we learn coefficients that assemble several motion bases to produce the correction motion, where these bases are learned from the consecutive frames of natural videos beforehand. Experiments show that our method outperforms existing single RS methods statistically and visually, in both synthesized and real RS images. Our code and dataset are available at https://github.com/DavidYan2001/Deep_HM. Weilong Yan, Robby T. Tan, Bing Zeng 0001, Shuaicheng Liu |
ICCV | 4 |
| 2023 | Minimum Latency Deep Online Video StabilizationabstractWe present a novel camera path optimization framework for the task of online video stabilization. Typically, a stabilization pipeline consists of three steps: motion estimating, path smoothing, and novel view rendering. Most previous methods concentrate on motion estimation, proposing various global or local motion models. In contrast, path optimization receives relatively less attention, especially in the important online setting, where no future frames are available. In this work, we adopt recent off-the-shelf high-quality deep motion models for motion estimation to recover the camera trajectory and focus on the latter two steps. Our network takes a short 2D camera path in a sliding window as input and outputs the stabilizing warp field of the last frame in the window, which warps the coming frame to its stabilized position. A hybrid loss is well-defined to constrain the spatial and temporal consistency. In addition, we build a motion dataset that contains stable and unstable motion pairs for the training. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art online methods both qualitatively and quantitatively and achieves comparable performance to offline methods. Our code and dataset are available at https://github.com/liuzhen03/NNDVS. Zhuofan Zhang, Zhen Liu 0022, Ping Tan 0002, Bing Zeng 0001, Shuaicheng Liu |
ICCV | 5 |
| 2023 | Dynamic updating self-training for semi-weakly supervised object detection
Shuaicheng Liu, Bing Zeng 0001 |
Neurocomputing | 2 |
| 2023 | Unsupervised Global and Local Homography Estimation With Motion Basis LearningabstractIn this paper, we introduce a new framework for unsupervised deep homography estimation. Our contributions are 3 folds. First, unlike previous methods that regress 4 offsets for a homography, we propose a homography flow representation, which can be estimated by a weighted sum of 8 pre-defined homography flow bases. Second, considering a homography contains 8 Degree-of-Freedoms (DOFs) that is much less than the rank of the network features, we propose a Low Rank Representation (LRR) block that reduces the feature rank, so that features corresponding to the dominant motions are retained while others are rejected. Last, we propose a Feature Identity Loss (FIL) to enforce the learned image feature warp-equivariant, meaning that the result should be identical if the order of warp operation and feature extraction is swapped. With this constraint, the unsupervised optimization can be more effective and the learned features are more stable. With global-to-local homography flow refinement, we also naturally generalize the proposed method to local mesh-grid homography estimation, which can go beyond the constraint of a single homography. Extensive experiments are conducted to demonstrate the effectiveness of all the newly proposed components, and results show that our approach outperforms the state-of-the-art on the homography benchmark dataset both qualitatively and quantitatively. Code is available at https://github.com/megvii-research/BasesHomo. Shuaicheng Liu, Hai Jiang 0006, Nianjin Ye, Chuan Wang 0001, Bing Zeng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Content-Aware Unsupervised Deep Homography Estimation and its ExtensionsabstractHomography estimation is a basic image alignment method in many applications. It is usually done by extracting and matching sparse feature points, which are error-prone in low-light and low-texture images. On the other hand, previous deep homography approaches use either synthetic images for supervised learning or aerial images for unsupervised learning, both ignoring the importance of handling depth disparities and moving objects in real-world applications. To overcome these problems, in this work, we propose an unsupervised deep homography method with a new architecture design. In the spirit of the RANSAC procedure in traditional methods, we specifically learn an outlier mask to only select reliable regions for homography estimation. We calculate loss with respect to our learned deep features instead of directly comparing image content as did previously. To achieve the unsupervised training, we also formulate a novel triplet loss customized for our network. We verify our method by conducting comprehensive comparisons on a new dataset that covers a wide range of scenes with varying degrees of difficulties for the task. Experimental results reveal that our method outperforms the state-of-the-art, including deep solutions and feature-based solutions. Shuaicheng Liu, Nianjin Ye, Chuan Wang 0001, Jirong Zhang, Lanpeng Jia, Kunming Luo, Jue Wang 0001, Jian Sun 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Deep Rotation Correction Without Angle PriorabstractNot everybody can be equipped with professional photography skills and sufficient shooting time, and there can be some tilts in the captured images occasionally. In this paper, we propose a new and practical task, named Rotation Correction, to automatically correct the tilt with high content fidelity in the condition that the rotated angle is unknown. This task can be easily integrated into image editing applications, allowing users to correct the rotated images without any manual operations. To this end, we leverage a neural network to predict the optical flows that can warp the tilted images to be perceptually horizontal. Nevertheless, the pixel-wise optical flow estimation from a single image is severely unstable, especially in large-angle tilted images. To enhance its robustness, we propose a simple but effective prediction strategy to form a robust elastic warp. Particularly, we first regress the mesh deformation that can be transformed into robust initial optical flows. Then we estimate residual optical flows to facilitate our network the flexibility of pixel-wise deformation, further correcting the details of the tilted images. To establish an evaluation benchmark and train the learning framework, a comprehensive rotation correction dataset is presented with a large diversity in scenes and rotated angles. Extensive experiments demonstrate that even in the absence of the angle prior, our algorithm can outperform other state-of-the-art solutions requiring this prior. The code and dataset are available at https://github.com/nie-lang/RotationCorrection. Lang Nie, Chunyu Lin, Kang Liao, Shuaicheng Liu, Yao Zhao 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | Stereo RGB and Deeper LIDAR-Based Network for 3D Object Detection in Autonomous Drivingabstract3D object detection has become an emerging task in autonomous driving scenarios. Most of previous works process 3D point clouds using either projection-based or voxel-based models. However, both approaches contain some drawbacks. The voxel-based methods lack semantic information, while the projection-based methods suffer from numerous spatial information loss when projected to different views. In this paper, we propose the Stereo RGB and Deeper LIDAR (SRDL) framework which can utilize semantic and spatial information simultaneously such that the performance of network for 3D object detection can be improved naturally. Specifically, the network generates candidate boxes from stereo pairs and combines different region-wise features using a deep fusion scheme. The stereo strategy offers more information for prediction compared with prior works. Then, several local and global feature extractors are stacked in the segmentation module to capture richer deep semantic geometric features from point clouds. After aligning the interior points with fused features, the proposed network refines the prediction in a more accurate manner and encodes the whole box in a novel compact method. The decent experimental results on the challenging KITTI detection benchmark demonstrate the effectiveness of utilizing both stereo images and point clouds for 3D object detection. Qingdong He, Zhengning Wang, Yijun Liu 0012, Shuaicheng Liu, Bing Zeng 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | Low-Light Image Enhancement with Wavelet-Based Diffusion ModelsabstractDiffusion models have achieved promising results in image restoration tasks, yet suffer from time-consuming, excessive computational resource consumption, and unstable restoration. To address these issues, we propose a robust and efficient Diffusion-based Low-Light image enhancement approach, dubbed DiffLL. Specifically, we present a wavelet-based conditional diffusion model (WCDM) that leverages the generative power of diffusion models to produce results with satisfactory perceptual fidelity. Additionally, it also takes advantage of the strengths of wavelet transformation to greatly accelerate inference and reduce computational resource usage without sacrificing information. To avoid chaotic content and diversity, we perform both forward diffusion and denoising in the training phase of WCDM, enabling the model to achieve stable denoising and reduce randomness during inference. Moreover, we further design a high-frequency restoration module (HFRM) that utilizes the vertical and horizontal details of the image to complement the diagonal information for better fine-grained restoration. Extensive experiments on publicly available real-world benchmarks demonstrate that our method outperforms the existing state-of-the-art methods both quantitatively and visually, and it achieves remarkable improvements in efficiency compared to previous diffusion-based methods. In addition, we empirically show that the application for low-light face detection also reveals the latent practical values of our method. Code is available at https://github.com/JianghaiSCU/Diffusion-Low-Light. Hai Jiang 0006, Ao Luo, Haoqiang Fan, Songchen Han, Shuaicheng Liu |
ACM Trans. Graph. | 5 |
| 2022 | Learning Optical Flow with Adaptive Graph ReasoningabstractEstimating per-pixel motion between video frames, known as optical flow, is a long-standing problem in video understanding and analysis. Most contemporary optical flow techniques largely focus on addressing the cross-image matching with feature similarity, with few methods considering how to explicitly reason over the given scene for achieving a holistic motion understanding. In this work, taking a fresh perspective, we introduce a novel graph-based approach, called adaptive graph reasoning for optical flow (AGFlow), to emphasize the value of scene/context information in optical flow. Our key idea is to decouple the context reasoning from the matching procedure, and exploit scene information to effectively assist motion estimation by learning to reason over the adaptive graph. The proposed AGFlow can effectively exploit the context information and incorporate it within the matching procedure, producing more robust and accurate results. On both Sintel clean and final passes, our AGFlow achieves the best accuracy with EPE of 1.43 and 2.47 pixels, outperforming state-of-the-art approaches by 11.2% and 13.6%, respectively. Code is publicly available at https://github.com/megvii-research/AGFlow. Ao Luo, Fan Yang 0054, Kunming Luo, Xin Li 0079, Haoqiang Fan, Shuaicheng Liu |
AAAI | 6 |
| 2022 | FINet: Dual Branches Feature Interaction for Partial-to-Partial Point Cloud RegistrationabstractData association is important in the point cloud registration. In this work, we propose to solve the partial-to-partial registration from a new perspective, by introducing multi-level feature interactions between the source and the reference clouds at the feature extraction stage, such that the registration can be realized without the attentions or explicit mask estimation for the overlapping detection as adopted previously. Specifically, we present FINet, a feature interactionbased structure with the capability to enable and strengthen the information associating between the inputs at multiple stages. To achieve this, we first split the features into two components, one for rotation and one for translation, based on the fact that they belong to different solution spaces, yielding a dual branches structure. Second, we insert several interaction modules at the feature extractor for the data association. Third, we propose a transformation sensitivity loss to obtain rotation-attentive and translation-attentive features. Experiments demonstrate that our method performs higher precision and robustness compared to the state-of-the-art traditional and learning-based methods. Code is available at https://github.com/megvii-research/FINet. Hao Xu 0018, Nianjin Ye, Guanghui Liu 0001, Bing Zeng 0001, Shuaicheng Liu |
AAAI | 5 |
| 2022 | Unsupervised Homography Estimation with Coplanarity-Aware GANabstractEstimating homography from an image pair is a fundamental problem in image alignment. Unsupervised learning methods have received increasing attention in this field due to their promising performance and label-free training. However, existing methods do not explicitly consider the problem of plane-induced parallax, which will make the predicted homography compromised on multiple planes. In this work, we propose a novel method HomoGAN to guide unsupervised homography estimation to focus on the dominant plane. First, a multi-scale transformer network is designed to predict homography from the feature pyramids of input images in a coarse-to-fine fashion. Moreover, we propose an unsupervised GAN to impose coplanarity constraint on the predicted homography, which is realized by using a generator to predict a mask of aligned regions, and then a discriminator to check if two masked feature maps are induced by a single homography. To validate the effectiveness of HomoGAN and its components, we conduct extensive experiments on a large-scale dataset, and results show that our matching error is 22% lower than the previous SOTA method. Code is available at https://github.com/megvii-research/HomoGAN Mingbo Hong, Nianjin Ye, Chunyu Lin, Qijun Zhao, Shuaicheng Liu |
CVPR | 6 |
| 2022 | RAGO: Recurrent Graph Optimizer For Multiple Rotation AveragingabstractThis paper proposes a deep recurrent Rotation Averaging Graph Optimizer (RAGO) for Multiple Rotation Averaging (MRA). Conventional optimization-based methods usually fail to produce accurate results due to corrupted and noisy relative measurements. Recent learning-based approaches regard MRA as a regression problem, while these methods are sensitive to initialization due to the gauge freedom problem. To handle these problems, we propose a learnable iterative graph optimizer minimizing a gauge- invariant cost function with an edge rectification strategy to mitigate the effect of inaccurate measurements. Our graph optimizer iteratively refines the global camera rotations by minimizing each node's single rotation objective function. Besides, our approach iteratively rectifies relative rotations to make them more consistent with the current camera orientations and observed relative rotations. Furthermore,$we$employ a gated recurrent unit to improve the result by tracing the temporal information of the cost graph. Our framework is a real-time learning-to-optimize rotation averaging graph optimizer with a tiny size deployed for real-world applications. RAGO outperforms previous traditional and deep methods on real-world and synthetic datasets. The code is available at github.com/sfu-gruvi-3dv/RAGO. Heng Li 0009, Zhaopeng Cui, Shuaicheng Liu, Ping Tan 0002 |
CVPR | 3 |
| 2022 | Practical Stereo Matching via Cascaded Recurrent Network with Adaptive CorrelationabstractWith the advent of convolutional neural networks, stereo matching algorithms have recently gained tremendous progress. However, it remains a great challenge to accurately extract disparities from real-world image pairs taken by consumer-level devices like smartphones, due to practical complicating factors such as thin structures, non-ideal rectification, camera module inconsistencies and various hard-case scenes. In this paper, we propose a set of innovative designs to tackle the problem of practical stereo matching: 1) to better recover fine depth details, we design a hierarchical network with recurrent refinement to update disparities in a coarse-to-fine manner, as well as a stacked cascaded architecture for inference; 2) we propose an adaptive group correlation layer to mitigate the impact of erroneous rectification; 3) we introduce a new synthetic dataset with special attention to difficult cases for better generalizing to real-world scenes. Our results not only rank 1ston both Middlebury and ETH3D benchmarks, outperforming existing state-of-the-art methods by a notable margin, but also exhibit high-quality details for real-life photos, which clearly demonstrates the efficacy of our contributions. Jiankun Li, Peisen Wang, Pengfei Xiong, Ziwei Yan, Jiangyu Liu, Haoqiang Fan, Shuaicheng Liu |
CVPR | 9 |
| 2022 | Deep Constrained Least Squares for Blind Image Super-ResolutionabstractIn this paper, we tackle the problem of blind image super-resolution(SR) with a reformulated degradation model and two novel modules. Following the common practices of blind SR, our method proposes to improve both the kernel estimation as well as the kernel based high resolution image restoration. To be more specific, we first reformulate the degradation model such that the deblurring kernel estimation can be transferred into the low resolution space. On top of this, we introduce a dynamic deep linear filter module. Instead of learning a fixed kernel for all images, it can adaptively generate deblurring kernel weights conditional on the input and yields more robust kernel estimation. Subsequently, a deep constrained least square filtering module is applied to generate clean features based on the reformulation and estimated kernel. The deblurred feature and the low input image feature are then fed into a dual-path structured SR network and restore the final high resolution result. To evaluate our method, we further conduct evaluations on several benchmarks, including Gaussian8 and DIV2KRK. Our experiments demonstrate that the proposed method achieves better accuracy and visual improvements against state-of-the-art methods. Codes and models are available at https://github.com/megvii-research/DCLS-SR. Ziwei Luo 0002, Haoqiang Fan, Shuaicheng Liu |
CVPR | 6 |
| 2022 | Learning Optical Flow with Kernel Patch AttentionabstractOptical flow is a fundamental method used for quantitative motion estimation on the image plane. In the deep learning era, most works treat it as a task of ‘matching of features’, learning to pull matched pixels as close as possible in feature space and vice versa. However, spatial affinity (smoothness constraint), another important component for motion understanding, has been largely overlooked. In this paper, we introduce a novel approach, called kernel patch attention (KPA), to better resolve the ambiguity in dense matching by explicitly taking the local context relations into consideration. Our KPA operates on each local patch, and learns to mine the context affinities for better inferring the flow fields. It can be plugged into contemporary optical flow architecture and empower the model to conduct comprehensive motion analysis with both feature similarities and spatial relations. On Sintel dataset, the proposed KPA-Flow achieves the best performance with EPE of 1.35 on clean pass and 2.36 on final pass, and it sets a new record of 4.60% in F1-all on KITTI-15 benchmark. Code is publicly available at https://github.com/megvii-research/KPAFlow. Ao Luo, Fan Yang 0054, Xin Li 0079, Shuaicheng Liu |
CVPR | 4 |
| 2022 | Deep Rectangling for Image Stitching: A Learning BaselineabstractStitched images provide a wide field-of-view (FoV) but suffer from unpleasant irregular boundaries. To deal with this problem, existing image rectangling methods devote to searching an initial mesh and optimizing a target mesh to form the mesh deformation in two stages. Then rectangu-lar images can be generated by warping stitched images. However, these solutions only work for images with rich linear structures, leading to noticeable distortions for por-traits and landscapes with non-linear objects. In this paper, we address these issues by proposing the first deep learning solution to image rectangling. Con-cretely, we predefine a rigid target mesh and only estimate an initial mesh to form the mesh deformation, contributing to a compact one-stage solution. The initial mesh is predicted using a fully convolutional network with a resid-ual progressive regression strategy. To obtain results with high content fidelity, a comprehensive objective function is proposed to simultaneously encourage the boundary rect-angular, mesh shape-preserving, and content perceptually natural. Besides, we build the first image stitching rectan-gling dataset with a large diversity in irregular boundaries and scenes. Experiments demonstrate our superiority over traditional methods both quantitatively and qualitatively. Lang Nie, Chunyu Lin, Kang Liao, Shuaicheng Liu, Yao Zhao 0001 |
CVPR | 4 |
| 2022 | Learning to Zoom Inside Camera Imaging PipelineabstractExisting single image super-resolution methods are either designed for synthetic data, or for real data but in the RGB-to-RGB or the RAW-to-RGB domain. This paper proposes to zoom an image from RAW to RAW inside the camera imaging pipeline. The RAW-to-RAW domain closes the gap between the ideal and the real degradation models. It also excludes the image signal processing pipeline, which refocuses the model learning onto the super-resolution. To these ends, we design a method that receives a low-resolution RAW as the input and estimates the desired higher-resolution RAW jointly with the degradation model. In our method, two convolutional neural networks are learned to constrain the high-resolution image and the degradation model in lower-dimensional subspaces. This subspace constraint converts the ill-posed SISR problem to a well-posed one. To demonstrate the superiority of the proposed method and the RAW-to-RAW domain, we conduct evaluations on the RealSR and the SR-RAW datasets. The results show that our method performs superiorly over the state-of-the-arts both qualitatively and quantitatively, and it also generalizes well and enables zero-shot transfer across different sensors. Chengzhou Tang, Yuqiang Yang, Bing Zeng 0001, Ping Tan 0002, Shuaicheng Liu |
CVPR | 5 |
| 2022 | SceneSqueezer: Learning to Compress Scene for Camera RelocalizationabstractStandard visual localization methods build a priori 3D model of a scene which is used to establish correspondences against the 2D keypoints in a query image. Storing these pre-built 3D scene models can be prohibitively expensive for large-scale environments, especially on mobile devices with limited storage and communication bandwidth. We design a novel framework that compresses a scene while still maintaining localization accuracy. The scene is compressed in three stages: first, the database frames are clustered using pairwise co-visibility information. Then, a learned point selection module prunes the points in each cluster taking into account the final pose estimation accuracy. In the final stage, the features of the selected points are further compressed using learned quantization. Query image registration is done using only the compressed scene points. To the best of our knowledge, we are the first to propose learned scene compression for visual localization. We also demonstrate the effectiveness and efficiency of our method on various outdoor datasets where it can perform accurate localization with low memory consumption. Luwei Yang, Rakesh Shrestha, Shuaicheng Liu, Guofeng Zhang 0001, Zhaopeng Cui, Ping Tan 0002 |
CVPR | 4 |
| 2022 | Semi-Supervised Wide-Angle Portraits Correction by Multi-Scale TransformerabstractWe propose a semi-supervised network for wide-angle portraits correction. Wide-angle images often suffer from skew and distortion affected by perspective distortion, especially noticeable at the face regions. Previous deep learning based approaches need the ground-truth correction flow maps for training guidance. However, such labels are expensive, which can only be obtained manually. In this work, we design a semi-supervised scheme and build a high-quality unlabeled dataset with rich scenarios, allowing us to simultaneously use labeled and unlabeled data to improve performance. Specifically, our semi-supervised scheme takes advantage of the consistency mechanism, with several novel components such as direction and range consistency (DRC) and regression consistency (RC). Furthermore, different from the existing methods, we propose the Multi-Scale Swin-Unet (MS-Unet) based on the multi-scale swin transformer block (MSTB), which can simultaneously learn short-distance and long-distance information to avoid artifacts. Extensive experiments demonstrate that the proposed method is superior to the state-of-the-art methods and other representative baselines. The source code and dataset are available at https://github.corn/megvii-research/PortraitsCorrection Fushun Zhu, Shan Zhao 0010, Hao Wang 0073, Shuaicheng Liu |
CVPR | 6 |
| 2022 | RealFlow: EM-Based Realistic Optical Flow Dataset Generation from Videos
Yunhui Han, Kunming Luo, Ao Luo, Jiangyu Liu, Haoqiang Fan, Guiming Luo, Shuaicheng Liu |
ECCV (19) | 7 |
| 2022 | D2C-SR: A Divergence to Convergence Approach for Real-World Image Super-Resolution
Lanpeng Jia, Haoqiang Fan, Shuaicheng Liu |
ECCV (19) | 5 |
| 2022 | Ghost-free High Dynamic Range Imaging with Context-Aware Transformer
Zhen Liu 0022, Yinglong Wang 0002, Bing Zeng 0001, Shuaicheng Liu |
ECCV (19) | 4 |
| 2022 | UPHDR-GAN: Generative Adversarial Network for High Dynamic Range Imaging With Unpaired DataabstractThe paper proposes a method to effectively fuse multi-exposure inputs and generate high-quality high dynamic range (HDR) images with unpaired datasets. Deep learning-based HDR image generation methods rely heavily on paired datasets. The ground truth images play a leading role in generating reasonable HDR images. Datasets without ground truth are hard to be applied to train deep neural networks. Recently, Generative Adversarial Networks (GAN) have demonstrated their potentials of translating images from source domain$X$to target domain$Y$in the absence of paired examples. In this paper, we propose a GAN-based network for solving such problems while generating enjoyable HDR results, named UPHDR-GAN. The proposed method relaxes the constraint of the paired dataset and learns the mapping from the LDR domain to the HDR domain. Although the pair data are missing, UPHDR-GAN can properly handle the ghosting artifacts caused by moving objects or misalignments with the help of the modified GAN loss, the improved discriminator network and the useful initialization phase. The proposed method preserves the details of important regions and improves the total image perceptual quality. Qualitative and quantitative comparisons against the representative methods demonstrate the superiority of the proposed UPHDR-GAN. Ru Li 0002, Chuan Wang 0001, Jue Wang 0001, Guanghui Liu 0001, Heng-Yu Zhang, Bing Zeng 0001, Shuaicheng Liu |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2022 | ASFlow: Unsupervised Optical Flow Learning With Adaptive Pyramid SamplingabstractWe present an unsupervised optical flow estimation method by proposing an adaptive pyramid sampling in the deep pyramid network. Specifically, in the pyramid downsampling, we propose a Content-Aware Pooling (CAP) module, which promotes local feature gathering by avoiding cross region pooling, so that the learned features become more representative. In the pyramid upsampling, we propose an Adaptive Flow Upsampling (AFU) module, where cross edge interpolation can be avoided, producing sharp motion boundaries. Equipped with these two modules, our method achieves the best performance for unsupervised optical flow estimation on multiple leading benchmarks, including MPI-Sintel, KITTI 2012 and KITTI 2015. Particularly, we achieve EPE=1.5 on KITTI 2012 and F1=9.67% KITTI 2015, which outperform the previous state-of-the-art methods by 16.7% and 13.1%, respectively. Shuaicheng Liu, Kunming Luo, Ao Luo, Chuan Wang 0001, Fanman Meng, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | DeepOIS: Gyroscope-Guided Deep Optical Image Stabilizer CompensationabstractMobile captured images can be aligned using their gyroscope sensors. Optical image stabilizer (OIS) terminates this possibility by adjusting the images during the capturing. In this work, we propose a deep network that compensates for the motions caused by the OIS, such that the gyroscopes can be used for image alignment on the OIS cameras. To achieve this, we first record both videos and gyroscope readings with an OIS camera as training data. Then, we convert gyroscope readings into motion fields. Second, we propose an Essential Mixtures motion model for rolling shutter cameras, where an array of rotations within a frame are extracted as the ground-truth guidance. Third, we train a convolutional neural network with gyroscope motions as input to compensate for the OIS motion. Once finished, the compensation network can be applied for other scenes, where the image alignment is purely based on gyroscopes with no need for images contents, delivering strong robustness. Experiments show that our results are comparable with that of non-OIS cameras, and outperform image-based alignment results with a relatively large margin. Code and dataset is available at:https://github.com/lhaippp/DeepOIS. Shuaicheng Liu, Haipeng Li 0001, Zhengning Wang, Jue Wang 0001, Shuyuan Zhu, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Depth-Aware Multi-Grid Deep Homography Estimation With Contextual CorrelationabstractHomography estimation is an important task in computer vision applications, such as image stitching, video stabilization, and camera calibration. Traditional homography estimation methods heavily depend on the quantity and distribution of feature correspondences, leading to poor robustness in low-texture scenes. The learning solutions, on the contrary, try to learn robust deep features but demonstrate unsatisfying performance in the scenes with low overlap rates. In this paper, we address these two problems simultaneously by designing a contextual correlation layer (CCL). The CCL can efficiently capture the long-range correlation within feature maps and can be flexibly used in a learning framework. In addition, considering that a single homography can not represent the complex spatial transformation in depth-varying images with parallax, we propose to predict multi-grid homography from global to local. Moreover, we equip our network with a depth perception capability, by introducing a novel depth-aware shape-preserved loss. Extensive experiments demonstrate the superiority of our method over state-of-the-art solutions in the synthetic benchmark dataset and real-world dataset. The codes and models will be available athttps://github.com/nie-lang/Multi-Grid-Deep-Homography. Lang Nie, Chunyu Lin, Kang Liao, Shuaicheng Liu, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Quadratic Terms Based Point-to-Surface 3D Representation for Deep Learning of Point CloudabstractIn this paper, we introduce a novel point-to-surface representation for 3D point cloud learning. Unlike the previous methods that mainly adopt voxel, mesh, or point coordinates, we propose to tackle this problem from a new perspective: learn a set of quadratic terms based static and global reference surfaces to describe 3D shapes, such that the coordinates of a 3D point (x, y, z) can be extended to quadratic terms (xy, xz, yz,$\ldots $) and transformed to the relationship between the local point and the global reference surfaces. Then, the static surfaces are changed into dynamic surfaces by adaptive contribution weighting to improve the descriptive capability. Towards this end, we propose our point-to-surface representation, a new representation for 3D point cloud learning that has not been attempted before, which can assemble local and global geometric information effectively by building connections between the point cloud and the learned reference surfaces. Given 3D points, we show how the reference surfaces are constructed, and how they are inserted into the 3D learning pipeline for different tasks. The experimental results confirm the effectiveness of our new representation, which has outperformed the state-of-the-art methods on the tasks of 3D classification and segmentation. Tiecheng Sun, Guanghui Liu 0001, Ru Li 0002, Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | JigsawGAN: Auxiliary Learning for Solving Jigsaw Puzzles With Generative Adversarial NetworksabstractThe paper proposes a solution based on Generative Adversarial Network (GAN) for solving jigsaw puzzles. The problem assumes that an image is divided into equal square pieces, and asks to recover the image according to information provided by the pieces. Conventional jigsaw puzzle solvers often determine the relationships based on the boundaries of pieces, which ignore the important semantic information. In this paper, we propose JigsawGAN, a GAN-based auxiliary learning method for solving jigsaw puzzles with unpaired images (with no prior knowledge of the initial images). We design a multi-task pipeline that includes, (1) a classification branch to classify jigsaw permutations, and (2) a GAN branch to recover features to images in correct orders. The classification branch is constrained by the pseudo-labels generated according to the shuffled pieces. The GAN branch concentrates on the image semantic information, where the generator produces the natural images to fool the discriminator, while the discriminator distinguishes whether a given image belongs to the synthesized or the real target domain. These two branches are connected by a flow-based warp module that is applied to warp features to correct the order according to the classification results. The proposed method can solve jigsaw puzzles more efficiently by utilizing both semantic information and boundary information simultaneously. Qualitative and quantitative comparisons against several representative jigsaw puzzle solvers demonstrate the superiority of our method. Ru Li 0002, Shuaicheng Liu, Guangfu Wang, Guanghui Liu 0001, Bing Zeng 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | NBNet: Noise Basis Learning for Image Denoising With Subspace ProjectionabstractIn this paper, we introduce NBNet, a novel framework for image denoising. Unlike previous works, we propose to tackle this challenging problem from a new perspective: noise reduction by image-adaptive projection. Specifically, we propose to train a network that can separate signal and noise by learning a set of reconstruction basis in the feature space. Subsequently, image denosing can be achieved by selecting corresponding basis of the signal subspace and projecting the input into such space. Our key insight is that projection can naturally maintain the local structure of input signal, especially for areas with low light or weak textures. Towards this end, we propose SSA, a non-local attention module we design to explicitly learn the basis generation as well as subspace projection. We further incorporate SSA with NBNet, a UNet structured network designed for end-to-end image denosing based. We conduct evaluations on benchmarks, including SIDD and DND, and NBNet achieves state-of-the-art performance on PSNR and SSIM with significantly less computational cost. Shen Cheng, Yuzhi Wang, Donghao Liu, Haoqiang Fan, Shuaicheng Liu |
CVPR | 6 |
| 2021 | UPFlow: Upsampling Pyramid for Unsupervised Optical Flow LearningabstractWe present an unsupervised learning approach for optical flow estimation by improving the upsampling and learning of pyramid network. We design a self-guided upsample module to tackle the interpolation blur problem caused by bilinear upsampling between pyramid levels. Moreover, we propose a pyramid distillation loss to add supervision for intermediate levels via distilling the finest flow as pseudo labels. By integrating these two components together, our method achieves the best performance for unsupervised optical flow learning on multiple leading benchmarks, including MPI-SIntel, KITTI 2012 and KITTI 2015. In particular, we achieve EPE=1.4 on KITTI 2012 and F1=9.38% on KITTI 2015, which outperform the previous state-of-the-art methods by 22.2% and 15.7%, respectively. Kunming Luo, Chuan Wang 0001, Shuaicheng Liu, Haoqiang Fan, Jue Wang 0001, Jian Sun 0001 |
CVPR | 3 |
| 2021 | Practical Wide-Angle Portraits Correction With Deep Structured ModelsabstractWide-angle portraits often enjoy expanded views. However, they contain perspective distortions, especially noticeable when capturing group portrait photos, where the background is skewed and faces are stretched. This paper introduces the first deep learning based approach to remove such artifacts from freely-shot photos. Specifically, given a wide-angle portrait as input, we build a cascaded network consisting of a LineNet, a ShapeNet, and a transition module (TM), which corrects perspective distortions on the background, adapts to the stereographic projection on facial regions, and achieves smooth transitions between these two projections, accordingly. To train our network, we build the first perspective portrait dataset with a large diversity in identities, scenes and camera modules. For the quantitative evaluation, we introduce two novel metrics, line consistency and face congruence. Compared to the previous state-of-the-art approach, our method does not require camera distortion parameters. We demonstrate that our approach significantly outperforms the previous state-of-the-art approach both qualitatively and quantitatively. Shan Zhao 0010, Pengfei Xiong, Jiangyu Liu, Haoqiang Fan, Shuaicheng Liu |
CVPR | 6 |
| 2021 | Holistic 3D Scene Understanding From a Single Image With Implicit RepresentationabstractWe present a new pipeline for holistic 3D scene understanding from a single image, which could predict object shapes, object poses, and scene layout. As it is a highly ill-posed problem, existing methods usually suffer from inaccurate estimation of both shapes and layout especially for the cluttered scene due to the heavy occlusion between objects. We propose to utilize the latest deep implicit representation to solve this challenge. We not only propose an image-based local structured implicit network to improve the object shape estimation, but also refine the 3D object pose and scene layout via a novel implicit scene graph neural network that exploits the implicit local object features. A novel physical violation loss is also proposed to avoid incorrect context between objects. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods in terms of object shape, scene layout estimation, and 3D object detection. Zhaopeng Cui, Yinda Zhang 0001, Bing Zeng 0001, Marc Pollefeys, Shuaicheng Liu |
CVPR | 6 |
| 2021 | GyroFlow: Gyroscope-Guided Unsupervised Optical Flow LearningabstractExisting optical flow methods are erroneous in challenging scenes, such as fog, rain, and night because the basic optical flow assumptions such as brightness and gradient constancy are broken. To address this problem, we present an unsupervised learning approach that fuses gyroscope into optical flow learning. Specifically, we first convert gyroscope readings into motion fields named gyro field. Second, we design a self-guided fusion module to fuse the background motion extracted from the gyro field with the optical flow and guide the network to focus on motion details. To the best of our knowledge, this is the first deep learning-based framework that fuses gyroscope data and image content for optical flow learning. To validate our method, we propose a new dataset that covers regular and challenging scenes. Experiments show that our method outperforms the state-of-art methods in both regular and challenging scenes. Code and dataset are available at https://github.com/megvii-research/GyroFlow. Haipeng Li 0001, Kunming Luo, Shuaicheng Liu |
ICCV | 3 |
| 2021 | OMNet: Learning Overlapping Mask for Partial-to-Partial Point Cloud RegistrationabstractPoint cloud registration is a key task in many computational fields. Previous correspondence matching based methods require the inputs to have distinctive geometric structures to fit a 3D rigid transformation according to point-wise sparse feature matches. However, the accuracy of transformation heavily relies on the quality of extracted features, which are prone to errors with respect to partiality and noise. In addition, they can not utilize the geometric knowledge of all the overlapping regions. On the other hand, previous global feature based approaches can utilize the entire point cloud for the registration, however they ignore the negative effect of non-overlapping points when aggregating global features. In this paper, we present OM-Net, a global feature based iterative network for partial-to-partial point cloud registration. We learn overlapping masks to reject non-overlapping regions, which converts the partial-to-partial registration to the registration of the same shape. Moreover, the previously used data is sampled only once from the CAD models for each object, resulting in the same point clouds for the source and reference. We propose a more practical manner of data generation where a CAD model is sampled twice for the source and reference, avoiding the previously prevalent over-fitting issue. Experimental results show that our method achieves state-of-the-art performance compared to traditional and deep learning based methods. Code is available at https://github.com/megvii-research/OMNet. Hao Xu 0018, Shuaicheng Liu, Guangfu Wang, Guanghui Liu 0001, Bing Zeng 0001 |
ICCV | 2 |
| 2021 | Motion Basis Learning for Unsupervised Deep Homography Estimation with Subspace ProjectionabstractIn this paper, we introduce a new framework for unsupervised deep homography estimation. Our contributions are 3 folds. First, unlike previous methods that regress 4 offsets for a homography, we propose a homography flow representation, which can be estimated by a weighted sum of 8 pre-defined homography flow bases. Second, considering a homography contains 8 Degree-of-Freedoms (DOFs) that is much less than the rank of the network features, we propose a Low Rank Representation (LRR) block that reduces the feature rank, so that features corresponding to the dominant motions are retained while others are rejected. Last, we propose a Feature Identity Loss (FIL) to enforce the learned image feature warp-equivariant, meaning that the result should be identical if the order of warp operation and feature extraction is swapped. With this constraint, the unsupervised optimization is achieved more effectively and more stable features are learned. Extensive experiments are conducted to demonstrate the effectiveness of all the newly proposed components, and results show that our approach outperforms the state-of-the-art on the homography benchmark datasets both qualitatively and quantitatively. Code is available at https://github.com/megvii-research/BasesHomo Nianjin Ye, Chuan Wang 0001, Haoqiang Fan, Shuaicheng Liu |
ICCV | 4 |
| 2021 | DeepPanoContext: Panoramic 3D Scene Understanding with Holistic Scene Context Graph and Relation-based OptimizationabstractPanorama images have a much larger field-of-view thus naturally encode enriched scene context information compared to standard perspective images, which however is not well exploited in the previous scene understanding methods. In this paper, we propose a novel method for panoramic 3D scene understanding which recovers the 3D room layout and the shape, pose, position, and semantic category for each object from a single full-view panorama image. In order to fully utilize the rich context information, we design a novel graph neural network based context model to predict the relationship among objects and room layout, and a differentiable relationship-based optimization module to optimize object arrangement with well-designed objective functions on-the-fly. Realizing the existing data are either with incomplete ground truth or overly-simplified scene, we present a new synthetic dataset with good diversity in room layout and furniture placement, and realistic image quality for total panoramic 3D scene understanding. Experiments demonstrate that our method outperforms existing methods on panoramic scene understanding in terms of both geometry accuracy and object arrangement. Code is available at https://chengzhag.github.io/publication/dpc. Zhaopeng Cui, Cai Chen 0002, Shuaicheng Liu, Bing Zeng 0001, Hujun Bao, Yinda Zhang 0001 |
ICCV | 4 |
| 2021 | Hierarchical Region Proposal Refinement Network for Weakly Supervised Object DetectionabstractWeakly supervised object detection (WSOD) has attracted more attention because it only requires image-level annotations to indicate whether a certain class exists. Most WSOD methods utilize multiple instance learning (MIL) to train an object detector where an image is treated as a bag of candidate proposals. Unlike fully supervised object detection (FSOD) that uses the object-aware region proposal network (RPN) to generate effective candidate proposals, WSOD only utilizes region proposal methods (e.g., selective search or edge boxes) due to the lack of instance-level annotations (i.e., bounding boxes). However, the quality of proposals can influence the training of the detector. To solve this problem, we propose a hierarchical region proposal refinement network (HRPRN) to refine these proposals gradually. Specifically, our network contains multiple weakly supervised detectors that are trained stage by stage. In addition, we propose an instance regression refinement model to generate object-aware coordinate offsets to refine proposals at each stage. In order to demonstrate the effectiveness of our method, we conduct experiments on PASCAL VOC 2007 dataset that is the widely used benchmark. Compared with our baseline method, online instance classifier refinement (OICR), our method achieves 9% and 5.6% improvements in terms of mAP and CorLoc, respectively. Shuaicheng Liu, Bing Zeng 0001 |
ICIP | 2 |
| 2021 | GLM-Net: Global and Local Motion Estimation via Task-Oriented Encoder-Decoder StructureabstractIn this work, we study the problem of separating the global camera motion and the local dynamic motion from an optical flow. Previous methods either estimate global motions by a parametric model, such as a homography, or estimate both of them by an optical flow field. However, none of these methods can directly estimate global and local motions through an end-to-end manner. In addition, separating the two motions accurately from a hybrid flow field is challenging. Because one motion can easily confuse the estimate of the other one when they are compounded together. To this end, we propose an end-to-end global and local motion estimation network GLM-Net. We design two encoder-decoder structures for the motion separation in the optical flow based on different task orientations. One structure adopts a mask autoencoder to extract the global motion, while the other one uses attention U-net for the local motion refinement. We further designed two effective training methods to overcome the problem of lacking supervisions. We apply our method on the action recognition datasets NCAA and UCF-101 to verify the accuracy of the local motion, and the homography estimation dataset DHE for the accuracy of the global motion. Experimental results show that our method can achieve competitive performance in both tasks at the same time, validating the effectiveness of the motion separation. Ye Xiang, Shuaicheng Liu, Lifang Wu, Boxuan Zhao, Bing Zeng 0001 |
ACM Multimedia | 3 |
| 2021 | OAENet: Oriented attention ensemble for accurate facial expression recognition
Zhengning Wang, Fanwei Zeng, Shuaicheng Liu, Bing Zeng 0001 |
Pattern Recognit. | 3 |
| 2021 | SDP-GAN: Saliency Detail Preservation Generative Adversarial Networks for High Perceptual Quality Style TransferabstractThe paper proposes a solution to effectively handle salient regions for style transfer between unpaired datasets. Recently, Generative Adversarial Networks (GAN) have demonstrated their potentials of translating images from source domain X to target domain Y in the absence of paired examples. However, such a translation cannot guarantee to generate high perceptual quality results. Existing style transfer methods work well with relatively uniform content, they often fail to capture geometric or structural patterns that always belong to salient regions. Detail losses in structured regions and undesired artifacts in smooth regions are unavoidable even if each individual region is correctly transferred into the target style. In this paper, we propose SDP-GAN, a GAN-based network for solving such problems while generating enjoyable style transfer results. We introduce a saliency network, which is trained with the generator simultaneously. The saliency network has two functions: (1) providing constraints for content loss to increase punishment for salient regions, and (2) supplying saliency features to generator to produce coherent results. Moreover, two novel losses are proposed to optimize the generator and saliency networks. The proposed method preserves the details on important salient regions and improves the total image perceptual quality. Qualitative and quantitative comparisons against several leading prior methods demonstrates the superiority of our method. Ru Li 0002, Chihao Wu 0001, Shuaicheng Liu, Jue Wang 0001, Guangfu Wang, Guanghui Liu 0001, Bing Zeng 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | OIFlow: Occlusion-Inpainting Optical Flow Estimation by Unsupervised LearningabstractOcclusion is an inevitable and critical problem in unsupervised optical flow learning. Existing methods either treat occlusions equally as non-occluded regions or simply remove them to avoid incorrectness. However, the occlusion regions can provide effective information for optical flow learning. In this paper, we present OIFlow, an occlusion-inpainting framework to make full use of occlusion regions. Specifically, a new appearance-flow network is proposed to inpaint occluded flows based on the image content. Moreover, a boundary dilated warp is proposed to deal with occlusions caused by displacement beyond the image border. We conduct experiments on multiple leading flow benchmark datasets such as Flying Chairs, KITTI and MPI-Sintel, which demonstrate that the performance is significantly improved by our proposed occlusion handling framework. Shuaicheng Liu, Kunming Luo, Nianjin Ye, Chuan Wang 0001, Jue Wang 0001, Bing Zeng 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Unsupervised Deep Image Stitching: Reconstructing Stitched Features to ImagesabstractTraditional feature-based image stitching technologies rely heavily on feature detection quality, often failing to stitch images with few features or low resolution. The learning-based image stitching solutions are rarely studied due to the lack of labeled data, making the supervised methods unreliable. To address the above limitations, we propose an unsupervised deep image stitching framework consisting of two stages: unsupervised coarse image alignment and unsupervised image reconstruction. In the first stage, we design an ablation-based loss to constrain an unsupervised homography network, which is more suitable for large-baseline scenes. Moreover, a transformer layer is introduced to warp the input images in the stitching-domain space. In the second stage, motivated by the insight that the misalignments in pixel-level can be eliminated to a certain extent in feature-level, we design an unsupervised image reconstruction network to eliminate the artifacts from features to pixels. Specifically, the reconstruction network can be implemented by a low-resolution deformation branch and a high-resolution refined branch, learning the deformation rules of image stitching and enhancing the resolution simultaneously. To establish an evaluation benchmark and train the learning framework, a comprehensive real-world image dataset for unsupervised deep image stitching is presented and released. Extensive experiments well demonstrate the superiority of our method over other state-of-the-art solutions. Even compared with the supervised solutions, our image stitching quality is still preferred by users. Lang Nie, Chunyu Lin, Kang Liao, Shuaicheng Liu, Yao Zhao 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | SlimConv: Reducing Channel Redundancy in Convolutional Neural Networks by Features RecombiningabstractThe channel redundancy of convolutional neural networks (CNNs) results in the large consumption of memories and computational resources. In this work, we design a novel Slim Convolution (SlimConv) module to boost the performance of CNNs by reducing channel redundancies. Our SlimConv consists of three main steps: Reconstruct, Transform, and Fuse. It aims to reorganize and fuse the learned features more efficiently, such that the method can compress the model effectively. Our SlimConv is a plug-and-play architectural unit that can be used to replace convolutional layers in CNNs directly. We validate the effectiveness of SlimConv by conducting comprehensive experiments on various leading benchmarks, such as ImageNet, MS COCO2014, Pascal VOC2012 segmentation, and Pascal VOC2007 detection datasets. The experiments show that SlimConv-equipped models can achieve better performances consistently, less consumption of memory and computation resources than non-equipped counterparts. For example, the ResNet-101 fitted with SlimConv achieves 77.84% top-1 classification accuracy with 4.87 GFLOPs and 27.96M parameters on ImageNet, which shows almost 0.5% better performance with about 3 GFLOPs and 38% parameters reduced. Jiaxiong Qiu, Cai Chen 0002, Shuaicheng Liu, Heng-Yu Zhang, Bing Zeng 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Semi-Supervised Pixel-Level Scene Text Segmentation by Mutually Guided NetworkabstractIn this paper we present a new data-driven method for pixel-level scene text segmentation from a single natural image. Although scene text detection, i.e. producing a text region mask, has been well studied in the past decade, pixel-level text segmentation is still an open problem due to the lack of massive pixel-level labeled data for supervised training. To tackle this issue, we incorporate text region mask as an auxiliary data into this task, considering acquiring large-scale of labeled text region mask is commonly less expensive and time-consuming. To be specific, we propose a mutually guided network which produces a polygon-level mask in one branch and a pixel-level text mask in the other. The two branches' outputs serve as guidance for each other and the whole network is trained via a semi-supervised learning strategy. Extensive experiments are conducted to demonstrate the effectiveness of our mutually guided network, and experimental results show our network outperforms the state-of-the-art in pixel-level scene text segmentation. We also demonstrate the mask produced by our network could improve the text recognition performance besides the trivial image editing application. Chuan Wang 0001, Shan Zhao 0010, Li Zhu 0003, Kunming Luo, Yanwen Guo 0001, Jue Wang 0001, Shuaicheng Liu |
IEEE Trans. Image Process. | 7 |
| 2021 | Rich Visual Knowledge-Based Augmentation Network for Visual Question AnsweringabstractVisual question answering (VQA) that involves understanding an image and paired questions develops very quickly with the boost of deep learning in relevant research fields, such as natural language processing and computer vision. Existing works highly rely on the knowledge of the data set. However, some questions require more professional cues other than the data set knowledge to answer questions correctly. To address such an issue, we propose a novel framework named a knowledge-based augmentation network (KAN) for VQA. We introduce object-related open-domain knowledge to assist the question answering. Concretely, we extract more visual information from images and introduce a knowledge graph to provide the necessary common sense or experience for the reasoning process. For these two augmented inputs, we design an attention module that can adjust itself according to the specific questions, such that the importance of external knowledge against detected objects can be balanced adaptively. Extensive experiments show that our KAN achieves state-of-the-art performance on three challenging VQA data sets, i.e., VQA v2, VQA-CP v2, and FVQA. In addition, our open-domain knowledge is also beneficial to VQA baselines. Code is available at https://github.com/yyyanglz/KAN. Shuaicheng Liu, Donghao Liu, Pengpeng Zeng, Jingkuan Song, Lianli Gao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Neural Point Cloud Rendering via Multi-Plane ProjectionabstractWe present a new deep point cloud rendering pipeline through multi-plane projections. The input to the network is the raw point cloud of a scene and the output are image or image sequences from a novel view or along a novel camera trajectory. Unlike previous approaches that directly project features from 3D points onto 2D image domain, we propose to project these features into a layered volume of camera frustum. In this way, the visibility of 3D points can be automatically learnt by the network, such that ghosting effects due to false visibility check as well as occlusions caused by noise interferences are both avoided successfully. Next, the 3D feature volume is fed into a 3D CNN to produce multiple planes of images w.r.t. the space division in the depth directions. The multi-plane images are then blended based on learned weights to produce the final rendering results. Experiments show that our network produces more stable renderings compared to previous methods, especially near the object boundaries. Moreover, our pipeline is robust to noisy and relatively sparse point cloud for a variety of challenging scenes. Peng Dai 0003, Yinda Zhang 0001, Zhuwen Li, Shuaicheng Liu, Bing Zeng 0001 |
CVPR | 4 |
| 2020 | DaST: Data-Free Substitute Training for Adversarial AttacksabstractMachine learning models are vulnerable to adversarial examples. For the black-box setting, current substitute attacks need pre-trained models to generate adversarial examples. However, pre-trained models are hard to obtain in real-world tasks. In this paper, we propose a data-free substitute training method (DaST) to obtain substitute models for adversarial black-box attacks without the requirement of any real data. To achieve this, DaST utilizes specially designed generative adversarial networks (GANs) to train the substitute models. In particular, we design a multi-branch architecture and label-control loss for the generative model to deal with the uneven distribution of synthetic samples. The substitute model is then trained by the synthetic samples generated by the generative model, which are labeled by the attacked model subsequently. The experiments demonstrate the substitute models produced by DaST can achieve competitive performance compared with the baseline models which are trained by the same train set with attacked models. Additionally, to evaluate the practicability of the proposed method on the real-world task, we attack an online machine learning model on the Microsoft Azure platform. The remote model misclassifies 98.35% of the adversarial examples crafted by our method. To the best of our knowledge, we are the first to train a substitute model for adversarial attacks without any real data. Mingyi Zhou, Jing Wu 0021, Yipeng Liu 0001, Shuaicheng Liu, Ce Zhu |
CVPR | 4 |
| 2020 | Flow-Guided Temporal-Spatial Network for HEVC Compressed Video Quality EnhancementabstractIn this paper, a flow-guided temporal-spatial network (FGTSN) is proposed to enhance the quality of HEVC compressed video. Specifically, we first employ a robust motion estimation subnet via trainable optical flow module to estimate the motion flow between the target frame and its adjacent frames, and these adjacent frames are pre-warped guided by the predicted motion flow. Then, a temporal encoder is proposed to fuse the related information between the target frame and its pre-warped frames. Finally, a quality enhancement subnet with multi-scale encoder-decoder structure is designed to generate high quality frame by training the network in a multi-supervised fashion. Experimental results show the superior performance of our proposed FGTSN method for the reconstruction quality of HEVC compressed frames, much better than the state-of-the-art quality enhancement methods. In addition, our FGTSN method can also effectively mitigate the quality fluctuation of adjacent frames. Xiandong Meng, Shuyuan Zhu, Shuaicheng Liu, Bing Zeng 0001 |
DCC | 4 |
| 2020 | Content-Aware Unsupervised Deep Homography Estimation
Jirong Zhang, Chuan Wang 0001, Shuaicheng Liu, Lanpeng Jia, Nianjin Ye, Jue Wang 0001, Ji Zhou 0001, Jian Sun 0001 |
ECCV (1) | 3 |
| 2020 | Multi-exposure photomontage with hand-held cameras
Ru Li 0002, Shuaicheng Liu, Guanghui Liu 0001, Tiecheng Sun, Jishun Guo |
Comput. Vis. Image Underst. | 2 |
| 2020 | An efficient and compact 3D local descriptor based on the weighted height image
Tiecheng Sun, Guanghui Liu 0001, Shuaicheng Liu, Fanman Meng, Liaoyuan Zeng, Ru Li 0002 |
Inf. Sci. | 3 |
| 2020 | PBR-Net: Imitating Physically Based Rendering Using Deep Neural NetworkabstractPhysically based rendering has been widely used to generate photo-realistic images, which greatly impacts industry by providing appealing rendering, such as for entertainment and augmented reality, and academia by serving large scale high-fidelity synthetic training data for data hungry methods like deep learning. However, physically based rendering heavily relies on ray-tracing, which can be computational expensive in complicated environment and hard to parallelize. In this paper, we propose an end-to-end deep learning based approach to generate physically based rendering efficiently. Our system consists of two stacked neural networks, which effectively simulates the physical behavior of the rendering process and produces photo-realistic images. The first network, namely shading network, is designed to predict the optimal shading image from surface normal, depth and illumination; the second network, namely composition network, learns to combine the predicted shading image with the reflectance to generate the final result. Our approach is inspired by intrinsic image decomposition, and thus it is more physically reasonable to have shading as intermediate supervision. Extensive experiments show that our approach is robust to noise thanks to a modified perceptual loss and even outperforms the physically based rendering systems in complex scenes given a reasonable time budget. Peng Dai 0003, Zhuwen Li, Yinda Zhang 0001, Shuaicheng Liu, Bing Zeng 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | DeepLiDAR: Deep Surface Normal Guided Depth Prediction for Outdoor Scene From Sparse LiDAR Data and Single Color ImageabstractIn this paper, we propose a deep learning architecture that produces accurate dense depth for the outdoor scene from a single color image and a sparse depth. Inspired by the indoor depth completion, our network estimates surface normals as the intermediate representation to produce dense depth, and can be trained end-to-end. With a modified encoder-decoder structure, our network effectively fuses the dense color image and the sparse LiDAR depth. To address outdoor specific challenges, our network predicts a confidence mask to handle mixed LiDAR signals near foreground boundaries due to occlusion, and combines estimates from the color image and surface normals with learned attention maps to improve the depth accuracy especially for distant areas. Extensive experiments demonstrate that our model improves upon the state-of-the-art performance on KITTI depth completion benchmark. Ablation study shows the positive impact of each model components to the final performance, and comprehensive analysis shows that our model generalizes well to the input with higher sparsity or from indoor scenes. Jiaxiong Qiu, Zhaopeng Cui, Yinda Zhang 0001, Xingdi Zhang, Shuaicheng Liu, Bing Zeng 0001, Marc Pollefeys |
CVPR | 5 |
| 2019 | C3AE: Exploring the Limits of Compact Model for Age EstimationabstractAge estimation is a classic learning problem in computer vision. Many larger and deeper CNNs have been proposed with promising performance, such as AlexNet, VggNet, GoogLeNet and ResNet. However, these models are not practical for the embedded/mobile devices. Recently, MobileNets and ShuffleNets have been proposed to reduce the number of parameters, yielding lightweight models. However, their representation has been weakened because of the adoption of depth-wise separable convolution. In this work, we investigate the limits of compact model for small-scale image and propose an extremely Compact yet efficient Cascade Context-based Age Estimation model(C3AE). This model possesses only 1/9 and 1/2000 parameters compared with MobileNets/ShuffleNets and VggNet, while achieves competitive performance. In particular, we re-define age estimation problem by two-points representation, which is implemented by a cascade model. Moreover, to fully utilize the facial context information, multi-branch CNN network is proposed to aggregate multi-scale context. Experiments are carried out on three age estimation datasets. The state-of-the-art performance on compact model has been achieved with a relatively large margin. Chao Zhang 0072, Shuaicheng Liu, Xun Xu 0002, Ce Zhu |
CVPR | 2 |
| 2019 | Hybrid Synthesis for Exposure Fusion from Hand-Held Camera InputsabstractThe paper proposes a hybrid synthesis method for multi-exposure image fusion taken by hand-held cameras. Motions either due to the shaky cameras or caused by dynamic scenes should be compensated before any content fusion. The misalignment will cause blurring/ghosting artifacts in the fused result. The proposed method can deal with such motions and maintain the exposure information of each input effectively. In particular, the proposed method first applies optical flow for a coarse registration, which performs well with complex non-rigid motion but produces deformations at regions with missing correspondences. To correct such error registration, we segment images into superpixels and identify problematic alignments based on each superpixel, which is further aligned by PatchMatch. After that, the proposed method obtains a fully aligned image stack which facilitates a high-quality fusion that is free from blurring/ghosting artifacts. We compare our method with existing fusion algorithms on various challenging examples, including the static/dynamic, the indoor/outdoor and the daytime/nighttime scenes. Experiment results demonstrate the effectiveness and robustness. Ru Li 0002, Shuaicheng Liu, Guanghui Liu 0001, Bing Zeng 0001 |
ICIP | 2 |
| 2019 | High-Quality Color Image Compression by Quantization Crossing Color SpacesabstractCoding of a color image usually happens in the YCbCr space so that the rate-distortion optimization is conducted in this space. Due to the use of a non-unitary matrix in the RGB-to-YCbCr conversion, an optimal coding performance achieved in the YCbCr space does not guarantee an optimal quality in the RGB space, which would impact most display devices that need RGB signals as the inputs. In this paper, we first study the relationship between the coding distortions of the compressed RGB signals and the quantization errors occurred in the coded YCbCr signals. Then, we design a new quantization scheme crossing the RGB and YCbCr spaces to achieve a high-quality color image compression with the YCbCr 4:4:4 format. Although our proposed quantization takes place in the YCbCr space, it aims at reducing the coding distortion in the RGB space as much as possible. Experimental results demonstrate that our proposed method offers a significant quality gain over the existing block-based coding methods for various images. Shuyuan Zhu, Zhiying He, Chen Chen 0015, Shuaicheng Liu, Jiantao Zhou 0001, Yuanfang Guo, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Photomontage for Robust HDR Imaging with Hand-Held CamerasabstractThis paper studies the image fusion from multiple images taken by hand-held cameras with different exposures. The existing methods often generate unsatisfactory results, such as the blurring/ghosting artifacts due to the problematic handling of camera motions, dynamic contents, and inappropriate fusion of local regions (e.g., over or under exposed). They often require high quality image registration before fusion. However, the accurate alignment is hard to obtain in many scenarios, such as scenes with large depth variations and dynamic textures. Besides, high quality alignment is also time consuming. In this paper, we only enable a rough registration by a single homography and combine the inputs seamlessly to hide any possible misalignment. Specifically, we propose to use a Markov Random Filed (MRF) function for the labelling of all pixels, which assigns different labels to different aligned input images. During the labelling, we choose well-exposured regions and skip moving objects simultaneously. Then, we combine a Laplace image according to the labels and construct the fusion result by solving the Poisson equation. We present various challenging examples to demonstrate the effectiveness and practicability of our approach. Ru Li 0002, Xiaowu He, Shuaicheng Liu, Guanghui Liu 0001, Bing Zeng 0001 |
ICIP | 3 |
| 2018 | Coding Trajectory: Enable Video Coding for Video DenoisingabstractWe introduce a novel video denoising approach which can produce a clean video by utilizing redundant image patches existed in the video frames. Previous multi-frame video denosing approaches either require image registration or employ Patch Match algorithms for the discovery of the patch redundancy. However, these computations are time-consuming and prone to errors. On the other hand, nearly all captured videos have been compressed. Such a compression can produce a rich set of block-based motion vectors that can be utilized for the redundant patch extraction, leading to the efficient video denosing. To be specific, the motion vectors and frame references can be obtained from the video coding. Given a noised frame block, we follow its motion vectors from the coding to form a trajectory and gather a set of block candidates along the routes from its nearby frames. The trajectory is referred to as Coding Trajectory. Then, the corresponding denoised block is generated by weighted fusing the block candidates with outlier rejections. A denoised frame is consisted of all the denoised blocks. We compare our method with several state-of-the-art approaches, such as VBM3D and VB-M4D, in terms of PSNR and SSIM. The experiments show that our method can achieve high quality results while runs much faster then the other approaches. Zhihang Ren, Peng Dai 0003, Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001 |
ICIP | 3 |
| 2018 | A 3D Descriptor based on Local Height ImageabstractThis paper proposes a novel 3D local descriptor, which seeks a good balance between the efficiency and the accuracy. We use the Local Reference Frame (LRF) to estimate a robust coordinate system to describe the local 3D shape. A novel Local Height Image (LHI) is defined by projecting the 3D points in the support region onto the tangent plane of the basis point. The Local Height Image Descriptor (LHID) is then defined by calculating the averaged projection distances. We further smooth the LHID to resist various kinds of interferences. We setup several experiments to assess the performance of our descriptor by comparison with the state-of-the-art algorithms. The experimental results demonstrate the effectiveness of the proposed method, which not only achieves the high accuracy as well as the robustness, but also possesses low complexity for the efficiency. Tiecheng Sun, Shuaicheng Liu, Guanghui Liu 0001, Shuyuan Zhu, Zhipeng Zhu |
ISCAS | 2 |
| 2018 | Block-based Image Coding by Compression-Constrained Transform Domain Down-ScalingabstractTransform domain down-scaling (TDDS) is traditionally implemented by dropping most of high-frequency components of the transformed block. Applying it to image compression can improve the compression efficiency by saving considerable bit-cost. Due to losing some necessary high-frequency information, the resulted image compressed by using the traditional TDDS-based coding often suffers a serious quality degradation. In this paper, we propose a compression-constrained TDDS and perform it on each N × N block to produce an N/2 × N/2 coefficient block for the compression. Our proposed TDDS not only guarantees a high reconstruction quality but also makes a low bit-cost for compression. We integrate it in practical image coding to build up our proposed compression scheme. Experimental results show that our proposed method demonstrates excellent coding performance when used to compress image signals. Chang Cui, Shuyuan Zhu, Xiandong Meng, Shuaicheng Liu, Bing Zeng 0001 |
VCIP | 4 |
| 2018 | Multi-exposure Fusion With JPEG Compression GuidanceabstractConstruct a High Dynamic Range (HDR) image is the primary method to solve the information loss caused by insufficient dynamic range of cameras. We propose a technique for fusing a bracketed low dynamic range (LDR) image sequence of varying exposures into an HDR image, skipping the physically-based HDR assembly step. Traditionally approaches often rely on complicated algorithms to select good regions from the input LDR images for the fusion. However, we found that the selection strategy can purely base on JPEG compression bits, bypassing the calculations of image low-level features, such as image gradients, local saturations, over/under exposure evaluations, as long as the input LDR image is compressed by the JPEG formats. In this way, lots of computations can be saved. In particular, we extract the coding bits from the intermediate product of the JPEG. The coding bits of blocks can be modified as the weights for the exposure fusion. Well-exposed regions often require higher bits for the compression while overexposure or saturated regions often correspond to lower bits. The objective and subjective evaluations demonstrate the effectiveness of our method. Xingdi Zhang, Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001 |
VCIP | 2 |
| 2018 | Cross-Space Distortion Directed Color Image CompressionabstractTraditional color image compression is usually conducted in the YCbCr space but many color displayers only accept RGB signals as inputs. Due to the use of a non-unitary matrix in the YCbCr-RGB conversion, low distortion achieved in the YCbCr space cannot guarantee low distortion for the RGB signals. To solve this problem, we propose a novel compression scheme for color images through defining a cross-space distortion so as to reduce as much as possible the distortion in the RGB space. To this end, we first derive the relationship between the distortions in the YCbCr space and RGB space. Then, we develop two solutions to implement color image compression for the most popular 4:2:0 chroma format. The first solution focuses on the design of a new spatial downsampling method to generate the 4:2:0 YCbCr image for a high-efficiency compression. The second one provides a novel way to reduce the distortion of the compressed color image by controlling the quantization error of the 4:2:0 YCbCr image, especially the one generated by using the traditional spatial downsampling. Experimental results show that both proposed solutions offer a remarkable quality gain over some state-of-the-art approaches when tested on various textured color images. Shuyuan Zhu, Chen Chen 0015, Shuaicheng Liu, Bing Zeng 0001 |
IEEE Trans. Multim. | 4 |
| 2017 | Direct Photometric Alignment by Mesh DeformationabstractThe choice of motion models is vital in applications like image/video stitching and video stabilization. Conventional methods explored different approaches ranging from simple global parametric models to complex per-pixel optical flow. Mesh-based warping methods achieve a good balance between computational complexity and model flexibility. However, they typically require high quality feature correspondences and suffer from mismatches and low-textured image content. In this paper, we propose a mesh-based photometric alignment method that minimizes pixel intensity difference instead of Euclidean distance of known feature correspondences. The proposed method combines the superior performance of dense photometric alignment with the efficiency of mesh-based image warping. It achieves better global alignment quality than the feature-based counterpart in textured images, and more importantly, it is also robust to low-textured image content. Abundant experiments show that our method can handle a variety of images and videos, and outperforms representative state-of-the-art methods in both image stitching and video stabilization tasks. Kaimo Lin, Nianjuan Jiang, Shuaicheng Liu, Loong Fah Cheong, Minh N. Do, Jiangbo Lu |
CVPR | 3 |
| 2017 | Shape Recovery of Endoscopic Videos by Shape from Shading Using Mesh Regularization
Zhihang Ren, Lingbing Peng, Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001 |
ICIG (3) | 4 |
| 2017 | Long-Distance/Environment Face Image Enhancement Method for Recognition
Zhengning Wang, Shanshan Ma, Mingyan Han, Shuaicheng Liu |
ICIG (1) | 5 |
| 2017 | Uncertain Region Identification for Stereoscopic Foreground Cutout
Taotao Yang, Shuaicheng Liu, Zhengning Wang, Bing Zeng 0001 |
ICIG (3) | 2 |
| 2017 | Meshflow video denoisingabstractWe propose an efficient video denoising approach that produces clean videos by utilizing the recently proposed meshflow motion model for the camera motion compensation. The meshflow is a spatially-smooth sparse motion field with motion vectors located at the mesh vertexes. The model is very effective and efficient for the purpose of the multiframes denoising due to its internal characteristics such as the lightweight, the nonparametric form, and the spatially-variant motion representation. Specifically, the meshflows are estimated between adjacent frames, which are used to align frames within a sliding time window. A denoised frame is generated by fusing of several registered frames in a spatial and temporal manner with outlier rejections. Various challenging examples demonstrate the effectiveness and practicability of the proposed approach. Zhihang Ren, Shuaicheng Liu, Bing Zeng 0001 |
ICIP | 3 |
| 2017 | Endoscopic video deblurring via synthesisabstractEndoscopic videos have been widely used for stomach diagnoses. However, endoscopic devices often capture videos with motion blurs, due to the dimly-lit environment and the camera shakiness during the capturing, which severely disturbs the diagnoses. In this paper, we present a framework that can restore blurry frames by synthesizing image details from the nearby sharp frames. Specifically, the blurry frame and their corresponding nearby sharp frames are identified according to the image gradient sharpness. To restore one blurry frame, a non-parametric mesh-based motion model is proposed to align the sharp frame to the blurry frame. The motion model leverages motions from image feature matches and optical flows, which yields high quality alignments to overcome challenges such as noisy, blurry, reflective and textureless interferences. After the alignment, the deblurred frame is synthesized by matching patches locally between the blurry frame and the aligned sharp frame. Without the estimation of blur kernels, we show that it is possible to directly compare a blurry patch against the sharp patches for the nearest neighbor matches in endoscopic images. The experiments demonstrate the effectiveness of our algorithm. Lingbing Peng, Shuaicheng Liu, Dehua Xie, Shuyuan Zhu, Bing Zeng 0001 |
VCIP | 2 |
| 2017 | MMSE-Directed Linear Image Interpolation Based on Nonlocal Geometric SimilarityabstractIn this letter, we propose a minimum mean square error (MMSE) directed linear interpolation to compose the high-resolution image from a single low-resolution image. We build up our interpolation model by using some similar image patches selected according to the nonlocal geometric similarity. First, we use a two-stage search scheme to collect the matched patches inside the whole image. Second, a similarity scaling factor is used in the second search to refine the collected patches so as to help find a robust solution to the MMSE-directed interpolation. Third, our MMSE-directed interpolation is regularized by the involved reference patches to make the solved interpolation coefficients more reliable. Experimental results show that our proposed method outperforms the state-of-the-art MMSE-directed linear interpolation schemes and works competitively with the state-of-the-art learning-based ones. Shuyuan Zhu, Zhiying He, Shuaicheng Liu, Bing Zeng 0001 |
IEEE Signal Process. Lett. | 3 |
| 2017 | A Hybrid Approach for Near-Range Video StabilizationabstractNear-range videos contain objects that are close to the camera. These videos often contain discontinuous depth variation (DDV), which is the main challenge to the existing video stabilization methods. Traditionally, 2D methods are robust to various camera motions (e.g., quick rotation and zooming) under scenes with continuous depth variation (CDV). However, in the presence of DDV, they often generate wobbled results due to the limited ability of their 2D motion models. Alternatively, 3D methods are more robust in handling near-range videos. We show that, by compensating rotational motions and ignoring translational motions, near-range videos can be successfully stabilized by 3D methods without sacrificing the stability too much. However, it is time-consuming to reconstruct the 3D structures for the entire video and sometimes even impossible due to rapid camera motions. In this paper, we combine the advantages of 2D and 3D methods, yielding a hybrid approach that is robust to various camera motions and can handle the near-range scenarios well. To this end, we automatically partition the input video into CDV and DDV segments. Then, the 2D and 3D approaches are adopted for CDV and DDV clips, respectively. Finally, these segments are stitched seamlessly via a constrained optimization. We validate our method on a large variety of consumer videos. Shuaicheng Liu, Binhan Xu, Chuang Deng, Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | CodingFlow: Enable Video Coding for Video StabilizationabstractVideo coding focuses on reducing the data size of videos. Video stabilization targets at removing shaky camera motions. In this paper, we enable video coding for video stabilization by constructing the camera motions based on the motion vectors employed in the video coding. The existing stabilization methods rely heavily on image features for the recovery of camera motions. However, feature tracking is time-consuming and prone to errors. On the other hand, nearly all captured videos have been compressed before any further processing and such a compression has produced a rich set of block-based motion vectors that can be utilized for estimating the camera motion. More specifically, video stabilization requires camera motions between two adjacent frames. However, motion vectors extracted from video coding may refer to non-adjacent frames. We first show that these non-adjacent motions can be transformed into adjacent motions such that each coding block within a frame contains a motion vector referring to its adjacent previous frame. Then, we regularize these motion vectors to yield a spatially-smoothed motion field at each frame, named as CodingFlow, which is optimized for a spatially-variant motion compensation. Based on CodingFlow, we finally design a grid-based 2D method to accomplish the video stabilization. Our method is evaluated in terms of efficiency and stabilization quality, both quantitatively and qualitatively, which shows that our method can achieve high-quality results compared with the state-of-the-art methods (feature-based). Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | A Hierarchical Approach for Rain or Snow Removing in a Single Color ImageabstractIn this paper, we propose an efficient algorithm to remove rain or snow from a single color image. Our algorithm takes advantage of two popular techniques employed in image processing, namely, image decomposition and dictionary learning. At first, a combination of rain/snow detection and a guided filter is used to decompose the input image into a complementary pair: 1) the low-frequency part that is free of rain or snow almost completely and 2) the high-frequency part that contains not only the rain/snow component but also some or even many details of the image. Then, we focus on the extraction of image's details from the high-frequency part. To this end, we design a 3-layer hierarchical scheme. In the first layer, an overcomplete dictionary is trained and three classifications are carried out to classify the high-frequency part into rain/snow and non-rain/snow components in which some common characteristics of rain/snow have been utilized. In the second layer, another combination of rain/snow detection and guided filtering is performed on the rain/snow component obtained in the first layer. In the third layer, the sensitivity of variance across color channels is computed to enhance the visual quality of rain/snow-removed image. The effectiveness of our algorithm is verified through both subjective (the visual quality) and objective (through rendering rain/snow on some ground-truth images) approaches, which shows a superiority over several state-of-the-art works. Yinglong Wang 0002, Shuaicheng Liu, Chen Chen 0015, Bing Zeng 0001 |
IEEE Trans. Image Process. | 2 |
| 2016 | MeshFlow: Minimum Latency Online Video Stabilization
Shuaicheng Liu, Ping Tan 0002, Lu Yuan 0001, Jian Sun 0001, Bing Zeng 0001 |
ECCV (6) | 1 |
| 2016 | Joint bundled camera paths for stereoscopic video stabilizationabstractThis paper presents a method to stabilize shaky stereoscopic videos captured by hand-held devices. Directly applying traditional monocular video stabilization techniques to two views independently is problematic as it often brings undesirable vertical disparities and produces inaccurate horizontal disparities, which violate original stereoscopic disparity constraints, leading to erroneous depth perception. In this paper, we show that monocular video stabilization methods, such as the bundled camera paths stabilization, can be extended for stereoscopic videos by taking additional disparity constraints during the stabilization. In particular, we first estimate disparities between two views. Then, we compute camera motions as meshes of bundled paths for each view. Next, we smooth paths of two views separately and iteratively. During each iteration, we adjust the meshes of one view by our proposed `Joint Disparity and Stability mesh Warp (JDSW)'. The final result is generated after several iterations of paths smoothing and meshes adjusting, in which temporal stability and correct depth perception are achieved simultaneously. We evaluate our method by various challenging stereoscopic videos with different camera motions and scene types. The experiments demonstrate the effectiveness of our method. Heng Guo 0003, Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001 |
ICIP | 2 |
| 2016 | Intrinsic decomposition for stereoscopic imagesabstractIntrinsic image decomposition is an important technique that decomposes an image into reflectance and shading components. In this paper, we enable intrinsic decomposition for stereoscopic images. Traditional approaches cannot be directly applied to decompose stereoscopic images, yielding inconsistent reflectance and 3D artifacts after recoloring. To solve this problem, we propose a straight yet effective method for stereoscopic intrinsic decomposition, which consists of classical retinex constraint as well as disparity constraint. The former encodes the shading smoothness prior while the latter controls the reflectance similarity between two views. To further reduce ambiguity, we employ local and non-local texture cues by using superpixels within and across two views. The experiments show that our method can effectively decompose stereoscopic images with high quality and offer a comfortable 3D viewing experience. Dehua Xie, Shuaicheng Liu, Kaimo Lin, Shuyuan Zhu, Bing Zeng 0001 |
ICIP | 2 |
| 2016 | Automatic Reflection Removal using Gradient Intensity and Motion CuesabstractWe present a method to separate the background image and reflection from two photos that are taken in front of a transparent glass under slightly different viewpoints. In our method, the SIFT-flow between two images is first calculated and a motion hierarchy is constructed from the SIFT-flow at multiple levels of spatial smoothness. To distinguish background edges and reflection edges, we calculate a motion score for each edge pixel by its variance along the motion hierarchy. Alternatively, we make use of the so-called superpixels to group edge pixels into edge segments and calculate the motion scores by averaging over each segment. In the meantime, we also calculate an intensity score for each edge pixel by its gradient magnitude. We combine both motion and intensity scores to get a combination score. A binary labelling (for separation) can be obtained by thresholding the combination scores. The background image is finally reconstructed from the separated gradients. Compared to the existing approaches that require a sequence of images or a video clip for the separation, we only need two images, which largely improves its feasibility. Various challenging examples are tested to validate the effectiveness of our method. Shuaicheng Liu, Taotao Yang, Bing Zeng 0001, Zhengning Wang, Guanghui Liu 0001 |
ACM Multimedia | 2 |
| 2016 | Geometry-based PSF estimation and deblurring of defocused images with depth informationabstractWe propose in this paper an algorithm to recover the blurred details of an image caused by defocusing during the photo-taking. Our algorithm takes one RGB image as well as its depth map as the input. We build up a model in which each captured image pixel is regarded as a light-emitting source that goes through a synthetic camera system. Thanks to the depth map, we have the geometrical information of the scene so that the point spread function (PSF) can be derived for each pixel more accurately as compared to conventional approaches where only RGB images are involved. Then, we make use of the derived PSFs to solve an optimization so as to reconstruct an all-in-focus image. The reconstructed results are evaluated by comparison with the original all-in-focus images. Compared to other methods for the deblurring of defocused images, our method shows a better recovery of image details. Yiqun Wu 0003, Bing Zeng 0001, Dehua Xie, Shuaicheng Liu |
VCIP | 5 |
| 2016 | Seamless Video Stitching from Hand-held Camera InputsabstractAbstract Images/videos captured by portable devices (e.g., cellphones, DV cameras) often have limited fields of view. Image stitching, also referred to as mosaics or panorama, can produce a wide angle image by compositing several photographs together. Although various methods have been developed for image stitching in recent years, few works address the video stitching problem. In this paper, we present the first system to stitch videos captured by hand‐held cameras. We first recover the 3D camera paths and a sparse set of 3D scene points using CoSLAM system, and densely reconstruct the 3D scene in the overlapping regions. Then, we generate a smooth virtual camera path, which stays in the middle of the original paths. Finally, the stitched video is synthesized along the virtual path as if it was taken from this new trajectory. The warping required for the stitching is obtained by optimizing over both temporal stability and alignment quality, while leveraging on 3D information at our disposal. The experiments show that our method can produce high quality stitching results for various challenging scenarios. Kaimo Lin, Shuaicheng Liu, Loong Fah Cheong, Bing Zeng 0001 |
Comput. Graph. Forum | 2 |
| 2016 | Joint Video Stitching and Stabilization From Moving CamerasabstractIn this paper, we extend image stitching to video stitching for videos that are captured for the same scene simultaneously by multiple moving cameras. In practice, videos captured under this circumstance often appear shaky. Directly applying image stitching methods for shaking videos often suffers from strong spatial and temporal artifacts. To solve this problem, we propose a unified framework in which video stitching and stabilization are performed jointly. Specifically, our system takes several overlapping videos as inputs. We estimate both inter motions (between different videos) and intra motions (between neighboring frames within a video). Then, we solve an optimal virtual 2D camera path from all original paths. An enlarged field of view along the virtual path is finally obtained by a space-temporal optimization that takes both inter and intra motions into consideration. Two important components of this optimization are that: 1) a grid-based tracking method is designed for an improved robustness, which produces features that are distributed evenly within and across multiple views and 2) a mesh-based motion model is adopted for the handling of the scene parallax. Some experimental results are provided to demonstrate the effectiveness of our approach on various consumer-level videos and a Plugin, named "Video Stitcher" is developed at Adobe After Effects CC2015 to show the processed videos. Heng Guo 0003, Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj |
IEEE Trans. Image Process. | 2 |
| 2014 | SteadyFlow: Spatially Smooth Optical Flow for Video StabilizationabstractWe propose a novel motion model, SteadyFlow, to represent the motion between neighboring video frames for stabilization. A SteadyFlow is a specific optical flow by enforcing strong spatial coherence, such that smoothing feature trajectories can be replaced by smoothing pixel profiles, which are motion vectors collected at the same pixel location in the SteadyFlow over time. In this way, we can avoid brittle feature tracking in a video stabilization system. Besides, SteadyFlow is a more general 2D motion model which can deal with spatially-variant motion. We initialize the SteadyFlow by optical flow and then discard discontinuous motions by a spatial-temporal analysis and fill in missing regions by motion completion. Our experiments demonstrate the effectiveness of our stabilization on real-world challenging videos. Shuaicheng Liu, Lu Yuan 0001, Ping Tan 0002, Jian Sun 0001 |
CVPR | 1 |
| 2014 | TrackCam: 3D-aware tracking shots from consumer videoabstractPanning and tracking shots are popular photography techniques in which the camera tracks a moving object and keeps it at the same position, resulting in an image where the moving foreground is sharp but the background is blurred accordingly, creating an artistic illustration of the foreground motion. Such shots however are hard to capture even for professionals, especially when the foreground motion is complex (e.g., non-linear motion trajectories). In this work we propose a system to generate realistic, 3D-aware tracking shots from consumer videos. We show how computer vision techniques such as segmentation and structure-from-motion can be used to lower the barrier and help novice users create high quality tracking shots that are physically plausible. We also introduce a pseudo 3D approach for relative depth estimation to avoid expensive 3D reconstruction for improved robustness and a wider application range. We validate our system through extensive quantitative and qualitative evaluations. Shuaicheng Liu, Jue Wang 0001, Sunghyun Cho, Ping Tan 0002 |
ACM Trans. Graph. | 1 |
| 2013 | Bundled camera paths for video stabilizationabstractWe present a novel video stabilization method which models camera motion with a bundle of (multiple) camera paths. The proposed model is based on a mesh-based, spatially-variant motion representation and an adaptive, space-time path optimization. Our motion representation allows us to fundamentally handle parallax and rolling shutter effects while it does not require long feature trajectories or sparse 3D reconstruction. We introduce the 'as-similar-as-possible' idea to make motion estimation more robust. Our space-time path smoothing adaptively adjusts smoothness strength by considering discontinuities, cropping size and geometrical distortion in a unified optimization framework. The evaluation on a large variety of consumer videos demonstrates the merits of our method. Shuaicheng Liu, Lu Yuan 0001, Ping Tan 0002, Jian Sun 0001 |
ACM Trans. Graph. | 1 |
| 2012 | Video stabilization with a depth cameraabstractPrevious video stabilization methods often employ homographies to model transitions between consecutive frames, or require robust long feature tracks. However, the homography model is invalid for scenes with significant depth variations, and feature point tracking is fragile in videos with textureless objects, severe occlusion or camera rotation. To address these challenging cases, we propose to solve video stabilization with an additional depth sensor such as the Kinect camera. Though the depth image is noisy, incomplete and low resolution, it facilitates both camera motion estimation and frame warping, which make the video stabilization a much well posed problem. The experiments demonstrate the effectiveness of our algorithm. Shuaicheng Liu, Yinting Wang, Lu Yuan 0001, Jiajun Bu, Ping Tan 0002, Jian Sun 0001 |
CVPR | 1 |
| 2010 | Super resolution using edge prior and single image detail synthesisabstractEdge-directed image super resolution (SR) focuses on ways to remove edge artifacts in upsampled images. Under large magnification, however, textured regions become blurred and appear homogenous, resulting in a super-resolution image that looks unnatural. Alternatively, learning-based SR approaches use a large database of exemplar images for “hallucinating” detail. The quality of the upsampled image, especially about edges, is dependent on the suitability of the training images. This paper aims to combine the benefits of edge-directed SR with those of learning-based SR. In particular, we propose an approach to extend edge-directed super-resolution to include detail from an image/texture example provided by the user (e.g., from the Internet). A significant benefit of our approach is that only a single exemplar image is required to supply the missing detail - strong edges are obtained in the SR image even if they are not present in the example image due to the combination of the edge-directed approach. In addition, we can achieve quality results at very large magnification, which is often problematic for both edge-directed and learning-based approaches. Yu-Wing Tai, Shuaicheng Liu, Michael S. Brown, Stephen Lin 0001 |
CVPR | 2 |
| 2010 | Colorization for Single Image Super Resolution
Shuaicheng Liu, Michael S. Brown, Seon Joo Kim, Yu-Wing Tai |
ECCV (6) | 1 |