Jiangyu Liu

dblp:73/7188 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
8since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2024 FlowDiffuser: Advancing Optical Flow Estimation with Diffusion Models
abstract
Optical flow estimation, a process of predicting pixel-wise displacement between consecutive frames, has commonly been approached as a regression task in the age of deep learning. Despite notable advancements, this de facto paradigm unfortunately falls short in generalization performance when trained on synthetic or constrained data. Pioneering a paradigm shift, we reformulate optical flow estimation as a conditional flow generation challenge, unveiling FlowDiffuser — a new family of optical flow models that could have stronger learning and generalization capabilities. FlowDiffuser estimates optical flow through a ‘noise-to-flow’ strategy, progressively eliminating noise from randomly generated flows conditioned on the provided pairs. To optimize accuracy and efficiency, our FlowDiffuser incorporates a novel Conditional Recurrent Denoising Decoder (Conditional-RDD), streamlining the flow estimation process. It incorporates a unique Hidden State Denoising (HSD) paradigm, effectively leveraging the information from previous time steps. Moreover, FlowDiffuser can be easily integrated into existing flow networks, leading to significant improvements in performance metrics compared to conventional implementations. Experiments on challenging benchmarks, including Sintel and KITTI, demonstrate the effectiveness of our FlowDiffuser with superior performance to existing state-of-the-art models. Code is available at https://github.com/LA30/FlowDiffuser.
Ao Luo, Fan Yang 0054, Jiangyu Liu, Haoqiang Fan, Shuaicheng Liu
CVPR4
2023 Explicit Motion Disentangling for Efficient Optical Flow Estimation
abstract
In this paper, we propose a novel framework for optical flow estimation that achieves a good balance between performance and efficiency. Our approach involves disentangling global motion learning from local flow estimation, treating global matching and local refinement as separate stages. We offer two key insights: First, the multi-scale 4D cost-volume based recurrent flow decoder is computationally expensive and unnecessary for handling small displacement. With the separation, we can utilize lightweight methods for both parts and maintain similar performance. Second, a dense and robust global matching is essential for both flow initialization as well as stable and fast convergence for the refinement stage. Towards this end, we introduce EMD-Flow, a framework that explicitly separates global motion estimation from the recurrent refinement stage. We propose two novel modules: Multi-scale Motion Aggregation (MMA) and Confidence-induced Flow Propagation (CFP). These modules leverage cross-scale matching prior and self-contained confidence maps to handle the ambiguities of dense matching in a global manner, generating a dense initial flow. Additionally, a lightweight decoding module is followed to handle small displacements, resulting in an efficient yet robust flow estimation framework. We further conduct comprehensive experiments on standard optical flow benchmarks with the proposed framework, and the experimental results demonstrate its superior balance between performance and runtime. Code is available at https://github.com/gddcx/EMD-Flow.
Changxing Deng, Ao Luo, Shaodan Ma, Jiangyu Liu, Shuaicheng Liu
ICCV5
2023 Uncertainty Guided Adaptive Warping for Robust and Efficient Stereo Matching
abstract
Correlation based stereo matching has achieved outstanding performance, which pursues cost volume between two feature maps. Unfortunately, current methods with a fixed model do not work uniformly well across various datasets, greatly limiting their real-world applicability. To tackle this issue, this paper proposes a new perspective to dynamically calculate correlation for robust stereo matching. A novel Uncertainty Guided Adaptive Correlation (UGAC) module is introduced to robustly adapt the same model for different scenarios. Specifically, a variance-based uncertainty estimation is employed to adaptively adjust the sampling area during warping operation. Additionally, we improve the traditional non-parametric warping with learnable parameters, such that the position-specific weights can be learned. We show that by empowering the recurrent network with the UGAC module, stereo matching can be exploited more robustly and effectively. Extensive experiments demonstrate that our method achieves state-of-the-art performance over the ETH3D, KITTI, and Middlebury datasets when employing the same fixed model over these datasets without any retraining procedure. To target real-time applications, we further design a lightweight model based on UGAC, which also outperforms other methods over KITTI benchmarks with only 0.6 M parameters.
Junpeng Jing, Jiankun Li, Pengfei Xiong, Jiangyu Liu, Shuaicheng Liu, Xin Deng 0002, Mai Xu, Lai Jiang 0004, Leonid Sigal
ICCV4
2023 Promoting Monocular Depth Estimation by Multi-Scale Residual Laplacian Pyramid Fusion
abstract
Deep learning approach has achieved great success in monocular depth estimation. However, the learned deep network may produce a depth map with fewer details and incorrect global depth layout, especially when the learned network is applied to a high-resolution image. In order to generate a high-quality depth map with better global structure and richer details, we propose a multi-scale residual Laplacian pyramid fusion net (MS-RLap-FNet), to fuse the multi-scale depth maps estimated by the existing depth estimation models, for depth refinement. Our approach relies on a proposed multi-scale residual Laplacian pyramid decomposition of the multi-scale depth maps, and the fusion network modules to gradually refine the depth maps based on the decomposition from low to high resolution. Comprehensive experiments show that our method, by refining the depth maps based on three popular monocular depth estimation models (DPT, MiDas, SGR), outperforms the existing state-of-the-art methods both in quantity and quality on three public datasets with different image resolutions. The depth map refined by our method has better global depth layout with richer fine details.
Anmei Zhang, Yunchao Ma, Jiangyu Liu, Jian Sun 0009
IEEE Signal Process. Lett.3
2022 Practical Stereo Matching via Cascaded Recurrent Network with Adaptive Correlation
abstract
With the advent of convolutional neural networks, stereo matching algorithms have recently gained tremendous progress. However, it remains a great challenge to accurately extract disparities from real-world image pairs taken by consumer-level devices like smartphones, due to practical complicating factors such as thin structures, non-ideal rectification, camera module inconsistencies and various hard-case scenes. In this paper, we propose a set of innovative designs to tackle the problem of practical stereo matching: 1) to better recover fine depth details, we design a hierarchical network with recurrent refinement to update disparities in a coarse-to-fine manner, as well as a stacked cascaded architecture for inference; 2) we propose an adaptive group correlation layer to mitigate the impact of erroneous rectification; 3) we introduce a new synthetic dataset with special attention to difficult cases for better generalizing to real-world scenes. Our results not only rank 1ston both Middlebury and ETH3D benchmarks, outperforming existing state-of-the-art methods by a notable margin, but also exhibit high-quality details for real-life photos, which clearly demonstrates the efficacy of our contributions.
Jiankun Li, Peisen Wang, Pengfei Xiong, Ziwei Yan, Jiangyu Liu, Haoqiang Fan, Shuaicheng Liu
CVPR7
2022 DIP: Deep Inverse Patchmatch for High-Resolution Optical Flow
abstract
Recently, the dense correlation volume method achieves state-of-the-art performance in optical flow. However, the correlation volume computation requires a lot of memory, which makes prediction difficult on high-resolution images. In this paper, we propose a novel Patchmatch-based framework to work on high-resolution optical flow estimation. Specifically, we introduce the first end-to-end Patchmatch based deep learning optical flow. It can get high-precision results with lower memory benefiting from propagation and local search of Patchmatch. Furthermore, a new inverse propagation is proposed to decouple the complex operations of propagation, which can significantly reduce calculations in multiple iterations. At the time of submission, our method ranks 1st on all the metrics on the popular KITTI2015 [28] benchmark, and ranks 2ndon EPE on the Sintel [7] clean benchmark among published optical flow methods. Experiment shows our method has a strong cross-dataset generalization ability that the F1-all achieves 13.73%, reducing 21% from the best published result 17.4% on KITTI2015. What's more, our method shows a good details preserving result on the high-resolution dataset DAVIS [1] and consumes 2× less memory than RAFT [36]. Code will be available at github.com/zihuarheng/DIP
Zihua Zheng, Ni Nie, Pengfei Xiong, Jiangyu Liu, Jiankun Li
CVPR5
2022 RealFlow: EM-Based Realistic Optical Flow Dataset Generation from Videos
Yunhui Han, Kunming Luo, Ao Luo, Jiangyu Liu, Haoqiang Fan, Guiming Luo, Shuaicheng Liu
ECCV (19)4
2021 Practical Wide-Angle Portraits Correction With Deep Structured Models
abstract
Wide-angle portraits often enjoy expanded views. However, they contain perspective distortions, especially noticeable when capturing group portrait photos, where the background is skewed and faces are stretched. This paper introduces the first deep learning based approach to remove such artifacts from freely-shot photos. Specifically, given a wide-angle portrait as input, we build a cascaded network consisting of a LineNet, a ShapeNet, and a transition module (TM), which corrects perspective distortions on the background, adapts to the stereographic projection on facial regions, and achieves smooth transitions between these two projections, accordingly. To train our network, we build the first perspective portrait dataset with a large diversity in identities, scenes and camera modules. For the quantitative evaluation, we introduce two novel metrics, line consistency and face congruence. Compared to the previous state-of-the-art approach, our method does not require camera distortion parameters. We demonstrate that our approach significantly outperforms the previous state-of-the-art approach both qualitatively and quantitatively.
Shan Zhao 0010, Pengfei Xiong, Jiangyu Liu, Haoqiang Fan, Shuaicheng Liu
CVPR4
2019 Disentangled Image Matting
abstract
Most previous image matting methods require a roughly-specificed trimap as input, and estimate fractional alpha values for all pixels that are in the unknown region of the trimap. In this paper, we argue that directly estimating the alpha matte from a coarse trimap is a major limitation of previous methods, as this practice tries to address two difficult and inherently different problems at the same time: identifying true blending pixels inside the trimap region, and estimate accurate alpha values for them. We propose AdaMatting, a new end-to-end matting framework that disentangles this problem into two sub-tasks: trimap adaptation and alpha estimation. Trimap adaptation is a pixel-wise classification problem that infers the global structure of the input image by identifying definite foreground, background, and semi-transparent image regions. Alpha estimation is a regression problem that calculates the opacity value of each blended pixel. Our method separately handles these two sub-tasks within a single deep convolutional neural network (CNN). Extensive experiments show that AdaMatting has additional structure awareness and trimap fault-tolerance. Our method achieves the state-of-the-art performance on Adobe Composition-1k dataset both qualitatively and quantitatively. It is also the current best-performing method on the alphamatting.com online evaluation for all commonly-used metrics.
Shaofan Cai, Xiaoshuai Zhang, Haoqiang Fan, Jiangyu Liu, Jiaying Liu 0001, Jue Wang 0001, Jian Sun 0001
ICCV5
2017 Deep Learning Based Intelligent Basketball Arena with Energy Image
Wu Liu 0005, Jiangyu Liu, Xiaoyan Gu 0001, Kun Liu 0016, Xiaowei Dai, Huadong Ma
MMM (1)2
2017 Deep learning based basketball video analysis for intelligent arena application
Wu Liu 0005, Chenggang Yan 0001, Jiangyu Liu, Huadong Ma
Multim. Tools Appl.3
2010 Parallel graph-cuts by adaptive bottom-up merging
abstract
Graph-cuts optimization is prevalent in vision and graphics problems. It is thus of great practical importance to parallelize the graph-cuts optimization using today's ubiquitous multi-core machines. However, the current best serial algorithm by Boykov and Kolmogorov (called the BK algorithm) still has the superior empirical performance. It is non-trivial to parallelize as expensive synchronization overhead easily offsets the advantage of parallelism. In this paper, we propose a novel adaptive bottom-up approach to parallelize the BK algorithm. We first uniformly partition the graph into a number of regularly-shaped disjoint subgraphs and process them in parallel, then we incrementally merge the subgraphs in an adaptive way to obtain the global optimum. The new algorithm has three benefits: 1) it is more cache-friendly within smaller subgraphs; 2) it keeps balanced workloads among computing cores; 3) it causes little overhead and is adaptable to the number of available cores. Extensive experiments in common applications such as 2D/3D image segmentations and 3D surface fitting demonstrate the effectiveness of our approach.
Jiangyu Liu, Jian Sun 0001
CVPR1
2009 Paint selection
abstract
In this paper, we present Paint Selection, a progressive painting-based tool for local selection in images. Paint Selection facilitates users to progressively make a selection by roughly painting the object of interest using a brush. More importantly, Paint Selection is efficient enough that instant feedback can be provided to users as they drag the mouse. We demonstrate that high quality selections can be quickly and effectively "painted" on a variety of multi-megapixel images.
Jiangyu Liu, Jian Sun 0001, Harry Shum
ACM Trans. Graph.1