Jiangbo Lu

dblp:77/6697 · DBLP profile ↗
← Back
84ranked-venue papers
15as first author
28since 2021 · last 2025
0000-0002-0048-3140ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 72 · 14 first-author · 21 since 2021Artificial intelligence and machine learning · 43 · 3 first-author · 21 since 2021Systems, architecture and hardware · 3 · 2 since 2021
YearPublicationVenuePosition
2025 FlexUOD: The Answer to Real-world Unsupervised Image Outlier Detection
abstract
How many outliers are within an unlabeled and contaminated dataset? Despite a series of unsupervised outlier detection (UOD) approaches have been proposed, they cannot correctly answer this critical question, resulting in their performance instability across various real-world (varying contamination factor) scenarios. To address this problem, we propose FlexUOD, with a novel contamination factor estimation perspective. FlexUOD not only achieves its remarkable robustness but also is a general and plug-and-play framework, which can significantly improve the performance of existing UOD methods. Extensive experiments demonstrate that FlexUOD achieves state-of-the-art results as well as high efficacy on diverse evaluation benchmarks.
Zhonghang Liu, Kun Zhou 0001, Changshuo Wang 0001, Wen-Yan Lin, Jiangbo Lu
CVPR5
2025 TSP-Mamba: The Travelling Salesman Problem Meets Mamba for Image Super-resolution and Beyond
abstract
Recently, Mamba-based frameworks have achieved substantial advancements across diverse computer vision and NLP tasks, particularly in their capacity for reasoning over long-range information with linear complexity. However, the fixed 2D-to-1D scanning pattern overlooks the local structures of an image, limiting its effectiveness in aggregating 2D spatial information. While stacking additional Mamba layers can partially address this issue, it increases parameter intensity and constrains real-time application. In this work, we reconsider the local optimal scanning path in Mamba, enhancing the rigid and uniform 1D scan through the local shortest path theory, thus creating a structure-aware Mamba suited for lightweight single-image super-resolution. Specifically, we draw inspiration from the Traveling Salesman Problem (TSP) to establish a local optimal scanning path for improved structural 2D information utilization. Here, local patch aggregation occurs in a content-adaptive manner with minimal propagation cost. TSP-Mamba demonstrates substantial improvements over existing Mamba-based and Transformer-based architectures. For example, TSP-Mamba surpasses MambaIR by up to 0.7dB in lightweight SISR, with comparable parameters and very slightly extra computational demands (1-2 GFlops for 720P images).
Kun Zhou 0001, Jiangbo Lu
CVPR3
2024 NeRF-HuGS: Improved Neural Radiance Fields in Non-static Scenes Using Heuristics-Guided Segmentation
abstract
Neural Radiance Field (NeRF) has been widely recognized for its excellence in novel view synthesis and 3D scene reconstruction. However, their effectiveness is in-herently tied to the assumption of static scenes, rendering them susceptible to undesirable artifacts when confronted with transient distractors such as moving objects or shad-ows. In this work, we propose a novel paradigm, namely “Heuristics-Guided Segmentation” (HuGS), which signifi-cantly enhances the separation of static scenes from tran-sient distractors by harmoniously combining the strengths of hand-crafted heuristics and state-of-the-art segmentation models, thus significantly transcending the limitations of previous solutions. Furthermore, we delve into the metic-ulous design of heuristics, introducing a seamless fusion of Structure-from-Motion (SfM)-based heuristics and color residual heuristics, catering to a diverse range of texture profiles. Extensive experiments demonstrate the superiority and robustness of our method in mitigating transient dis-tractors for NeRFs trained in non-static scenes. Project page: https://cnhaox.github.io/NeRF-HuGS/
Yipeng Qin, Lingjie Liu, Jiangbo Lu, Guanbin Li
CVPR4
2024 Unveiling Advanced Frequency Disentanglement Paradigm for Low-Light Image Enhancement
Kun Zhou 0001, Wenbo Li 0002, Xiaogang Xu 0002, Yuanhao Cai, Zhonghang Liu, Xiaoguang Han 0001, Jiangbo Lu
ECCV (7)8
2024 Emphasizing Semantic Consistency of Salient Posture for Speech-Driven Gesture Generation
abstract
Speech-driven gesture generation aims at synthesizing a gesture sequence synchronized with the input speech signal. Previous methods leverage neural networks to directly map a compact audio representation to the gesture sequence, ignoring the semantic association of different modalities and failing to deal with salient gestures. In this paper, we propose a novel speech-driven gesture generation method by emphasizing the semantic consistency of salient posture. Specifically, we first learn a joint manifold space for the individual representation of audio and body pose to exploit the inherent semantic association between two modalities, and propose to enforce semantic consistency via a consistency loss. Furthermore, we emphasize the semantic consistency of salient postures by introducing a weakly-supervised detector to identify salient postures, and reweighting the consistency loss to focus more on learning the correspondence between salient postures and the high-level semantics of speech content. In addition, we propose to extract audio features dedicated to facial expression and body gesture separately, and design separate branches for face and body gesture synthesis. Extensive experimental results demonstrate the superiority of our method over the state-of-the-art approaches.
Fengqi Liu, Jingyu Gong, Ran Yi 0002, Qianyu Zhou 0001, Xuequan Lu, Jiangbo Lu, Lizhuang Ma
ACM Multimedia7
2024 UPS: Unified Projection Sharing for Lightweight Single-Image Super-resolution and Beyond
abstract
To date, transformer-based frameworks have demonstrated impressive results in single-image super-resolution (SISR). However, under practical lightweight scenarios, the complex interaction of deep image feature extraction and similarity modeling limits the performance of these methods, since they require simultaneous layer-specific optimization of both two tasks. In this work, we introduce a novel Unified Projection Sharing algorithm(UPS) to decouple the feature extraction and similarity modeling, achieving notable performance. To do this, we establish a unified projection space defined by a learnable projection matrix, for similarity calculation across all self-attention layers. As a result, deep image feature extraction remains a per-layer optimization manner, while similarity modeling is carried out by projecting these image features onto the shared projection space. Extensive experiments demonstrate that our proposed UPS achieves state-of-the-art performance relative to leading lightweight SISR methods, as verified by various popular benchmarks. Moreover, our unified optimized projection space exhibits encouraging robustness performance for unseen data (degraded and depth images). Finally, UPS also demonstrates promising results across various image restoration tasks, including real-world and classic SISR, image denoising, and image deblocking.
Kun Zhou 0001, Zhonghang Liu, Xiaoguang Han 0001, Jiangbo Lu
NeurIPS5
2024 HAWK: Learning to Understand Open-World Video Anomalies
abstract
Video Anomaly Detection (VAD) systems can autonomously monitor and identify disturbances, reducing the need for manual labor and associated costs. However, current VAD systems are often limited by their superficial semantic understanding of scenes and minimal user interaction. Additionally, the prevalent data scarcity in existing datasets restricts their applicability in open-world scenarios. In this paper, we introduce HAWK, a novel framework that leverages interactive large Visual Language Models (VLM) to interpret video anomalies precisely. Recognizing the difference in motion information between abnormal and normal videos, HAWK explicitly integrates motion modality to enhance anomaly identification. To reinforce motion attention, we construct an auxiliary consistency loss within the motion and video space, guiding the video branch to focus on the motion modality. Moreover, to improve the interpretation of motion-to-language, we establish a clear supervisory relationship between motion and its linguistic representation. Furthermore, we have annotated over 8,000 anomaly videos with language descriptions, enabling effective training across diverse open-world scenarios, and also created 8,000 question-answering pairs for users' open-world questions. The final results demonstrate that HAWK achieves SOTA performance, surpassing existing baselines in both video description generation and question-answering. Our codes/dataset/demo will be released at https://github.com/jqtangust/hawk.
Jiaqi Tang 0005, Hao Lu 0009, Ruizheng Wu, Xiaogang Xu 0002, Bin Guo 0001, Jiangbo Lu, Qifeng Chen 0001, Ying-Cong Chen
NeurIPS8
2024 From NeRFLiX to NeRFLiX++: A General NeRF-Agnostic Restorer Paradigm
abstract
Neural radiance fields (NeRF) have shown great success in novel view synthesis. However, recovering high-quality details from real-world scenes is still challenging for the existing NeRF-based approaches, due to the potential imperfect calibration information and scene representation inaccuracy. Even with high-quality training frames, the synthetic novel views produced by NeRF models still suffer from notable rendering artifacts, such as noise and blur. To address this, we propose NeRFLiX, a general NeRF-agnostic restorer paradigm that learns a degradation-driven inter-viewpoint mixer. Specially, we design a NeRF-style degradation modeling approach and construct large-scale training data, enabling the possibility of effectively removing NeRF-native rendering artifacts for deep neural networks. Moreover, beyond the degradation removal, we propose an inter-viewpoint aggregation framework that fuses highly related high-quality training images, pushing the performance of cutting-edge NeRF models to entirely new levels and producing highly photo-realistic synthetic views. Based on this paradigm, we further present NeRFLiX++ with a stronger two-stage NeRF degradation simulator and a faster inter-viewpoint mixer, achieving superior performance with significantly improved computational efficiency. Notably, NeRFLiX++ is capable of restoring photo-realistic ultra-high-resolution outputs from noisy low-resolution NeRF-rendered views. Extensive experiments demonstrate the excellent restoration ability of NeRFLiX++ on various novel view synthesis benchmarks.
Kun Zhou 0001, Wenbo Li 0002, Nianjuan Jiang, Xiaoguang Han 0001, Jiangbo Lu
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 A High-Performance Accelerator for Real-Time Super-Resolution on Edge FPGAs
abstract
In the digital era, the prevalence of low-quality images contrasts with the widespread use of high-definition displays, primarily due to low-resolution cameras and compression technologies. Image super-resolution (SR) techniques, particularly those leveraging deep learning, aim to enhance these images for high-definition presentation. However, real-time execution of deep neural network (DNN)-based SR methods at the edge poses challenges due to their high computational and storage requirements. To address this, field-programmable gate arrays (FPGAs) have emerged as a promising platform, offering flexibility, programmability, and adaptability to evolving models. Previous FPGA-based SR solutions have focused on reducing computational and memory costs through aggressive simplification techniques, often sacrificing the quality of the reconstructed images. This paper introduces a novel SR network specifically designed for edge applications, which maintains reconstruction performance while managing computation costs effectively. Additionally, we propose an architectural design that enables the real-time and end-to-end inference of the proposed SR network on embedded FPGAs. Our key contributions include a tailored SR algorithm optimized for embedded FPGAs, a DSP-enhanced design that achieves a significant four-fold speedup, a novel scalable cache strategy for handling large feature maps, optimization of DSP cascade consumption, and a constraint optimization approach for resource allocation. Experimental results demonstrate that our FPGA-specific accelerator surpasses existing solutions, delivering superior throughput, energy efficiency, and image quality.
Hongduo Liu, Yijian Qian, Youqiang Liang, Zhaohan Liu, Wenqian Zhao 0002, Jiangbo Lu, Bei Yu 0001
ACM Trans. Design Autom. Electr. Syst.8
2023 CRIN: Rotation-Invariant Point Cloud Analysis and Rotation Estimation via Centrifugal Reference Frame
abstract
Various recent methods attempt to implement rotation-invariant 3D deep learning by replacing the input coordinates of points with relative distances and angles. Due to the incompleteness of these low-level features, they have to undertake the expense of losing global information. In this paper, we propose the CRIN, namely Centrifugal Rotation-Invariant Network. CRIN directly takes the coordinates of points as input and transforms local points into rotation-invariant representations via centrifugal reference frames. Aided by centrifugal reference frames, each point corresponds to a discrete rotation so that the information of rotations can be implicitly stored in point features. Unfortunately, discrete points are far from describing the whole rotation space. We further introduce a continuous distribution for 3D rotations based on points. Furthermore, we propose an attention-based down-sampling strategy to sample points invariant to rotations. A relation module is adopted at last for reinforcing the long-range dependencies between sampled points and predicts the anchor point for unsupervised rotation estimation. Extensive experiments show that our method achieves rotation invariance, accurately estimates the object rotation, and obtains state-of-the-art results on rotation-augmented classification and part segmentation. Ablation studies validate the effectiveness of the network design.
Yujing Lou, Zelin Ye, Yang You 0004, Nianjuan Jiang, Jiangbo Lu, Lizhuang Ma, Cewu Lu
AAAI5
2023 Low-Light Image Enhancement via Structure Modeling and Guidance
abstract
This paper proposes a new framework for low-light image enhancement by simultaneously conducting the appearance as well as structure modeling. It employs the structural feature to guide the appearance enhancement, leading to sharp and realistic results. The structure modeling in our framework is implemented as the edge detection in low-light images. It is achieved with a modified generative model via designing a structure-aware feature extractor and generator. The detected edge maps can accurately emphasize the essential structural information, and the edge prediction is robust towards the noises in dark areas. Moreover, to improve the appearance modeling, which is implemented with a simple U-Net, a novel structure-guided enhancement module is proposed with structure-guided feature synthesis layers. The appearance modeling, edge detector, and enhancement module can be trained end-to-end. The experiments are conducted on representative datasets (sRGB and RAW domains), showing that our model consistently achieves SOTA performance on all datasets with the same architecture. The code is available at https://github.com/xiaogangOO/SMG-LLIE.
Xiaogang Xu 0002, Ruixing Wang, Jiangbo Lu
CVPR3
2023 Exploring Motion Ambiguity and Alignment for High-Quality Video Frame Interpolation
abstract
For video frame interpolation (VFI), existing deep-learning-based approaches strongly rely on the ground-truth (GT) intermediate frames, which sometimes ignore the non-unique nature of motion judging from the given adjacent frames. As a result, these methods tend to produce averaged solutions that are not clear enough. To alleviate this issue, we propose to relax the requirement of reconstructing an intermediate frame as close to the GT as possible. Towards this end, we develop a texture consistency loss (TCL) upon the assumption that the interpolated content should maintain similar structures with their counterparts in the given frames. Predictions satisfying this constraint are encouraged, though they may differ from the predefined GT. Without the bells and whistles, our plug-and-play TCL is capable of improving the performance of existing VFI frameworks consistently. On the other hand, previous methods usually adopt the cost volume or correlation map to achieve more accurate image or feature warping. However, the O (N2) (N refers to the pixel count) computational complexity makes it infeasible for high-resolution cases. In this work, we design a simple, efficient O (N) yet powerful guided cross-scale pyramid alignment (GCSPA) module, where multi-scale information is highly exploited. Extensive experiments justify the efficiency and effectiveness of the proposed strategy.
Kun Zhou 0001, Wenbo Li 0002, Xiaoguang Han 0001, Jiangbo Lu
CVPR4
2023 NeRFLiX: High-Quality Neural View Synthesis by Learning a Degradation-Driven Inter-viewpoint MiXer
abstract
Neural radiance fields (NeRF) show great success in novel view synthesis. However, in real-world scenes, recovering high-quality details from the source images is still challenging for the existing NeRF-based approaches, due to the potential imperfect calibration information and scene representation inaccuracy. Even with high-quality training frames, the synthetic novel views produced by NeRF models still suffer from notable rendering artifacts, such as noise, blur, etc. Towards to improve the synthesis quality of NeRF-based approaches, we propose NeRFLiX, a general NeRF-agnostic restorer paradigm by learning a degradation-driven inter-viewpoint mixer. Specially, we design a NeRF-style degradation modeling approach and construct large-scale training data, enabling the possibility of effectively removing NeRF-native rendering artifacts for existing deep neural networks. Moreover, beyond the degradation removal, we propose an inter-viewpoint aggregation framework that is able to fuse highly related high-quality training images, pushing the performance of cutting-edge NeRF models to entirely new levels and producing highly photo-realistic synthetic views.
Kun Zhou 0001, Wenbo Li 0001, Yi Wang 0074, Tao Hu 0011, Nianjuan Jiang, Xiaoguang Han 0001, Jiangbo Lu
CVPR7
2023 On Efficient Transformer-Based Image Pre-training for Low-Level Vision
abstract
Pre-training has marked numerous state of the arts in high-level computer vision, while few attempts have ever been made to investigate how pre-training acts in image processing systems. In this paper, we tailor transformer-based pre-training regimes that boost various low-level tasks. To comprehensively diagnose the influence of pre-training, we design a whole set of principled evaluation tools that uncover its effects on internal representations. The observations demonstrate that pre-training plays strikingly different roles in low-level tasks. For example, pre-training introduces more local information to intermediate layers in super-resolution (SR), yielding significant performance gains, while pre-training hardly affects internal feature representations in denoising, resulting in limited gains. Further, we explore different methods of pre-training, revealing that multi-related-task pre-training is more effective and data-efficient than other alternatives. Finally, we extend our study to varying data scales and model sizes, as well as comparisons between transformers and CNNs. Based on the study, we successfully develop state-of-the-art models for multiple low-level tasks.
Wenbo Li 0002, Xin Lu 0006, Shengju Qian, Jiangbo Lu
IJCAI4
2023 A High-Performance Accelerator for Super-Resolution Processing on Embedded GPU
abstract
Over the past few years, super-resolution (SR) processing has achieved astonishing progress along with the development of deep learning. Nevertheless, the rigorous requirement for real-time inference, especially for video tasks, leaves a harsh challenge for both the model architecture design and the hardware-level implementation. In this article, we propose a hardware-aware acceleration on embedded GPU devices as a full-stack SR deployment framework. The most critical stage with dictionary learning applied in SR flow was analyzed in details and optimized with a tailored dictionary slimming strategy. Moreover, we also delve into the programming architecture of hardware while analyzing the model structure to optimize the computation kernels to reduce inference latency and maximize the throughput given restricted computing power. In addition, we further accelerate the model with 8-bit integer inference by quantizing the weights in the compressed model. An adaptive 8-bit quantization flow for SR task enables the quantized model to achieve a comparable result with the full-precision baselines. With the help of our approaches, the computation and communication bottlenecks in the deep dictionary learning-based SR models can be overcome effectively. The experiments on both edge embedded device NVIDIA NX and 2080Ti prove that our framework exceeds the performance of state-of-the-art NVIDIA TensorRT significantly and can achieve real-time performance.
Wenqian Zhao 0002, Qi Sun 0002, Wenbo Li 0002, Haisheng Zheng, Nianjuan Jiang, Jiangbo Lu, Bei Yu 0001, Martin D. F. Wong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2022 Best-Buddy GANs for Highly Detailed Image Super-resolution
abstract
We consider the single image super-resolution (SISR) problem, where a high-resolution (HR) image is generated based on a low-resolution (LR) input. Recently, generative adversarial networks (GANs) become popular to hallucinate details. Most methods along this line rely on a predefined single-LR-single-HR mapping, which is not flexible enough for the ill-posed SISR task. Also, GAN-generated fake details may often undermine the realism of the whole image. We address these issues by proposing best-buddy GANs (Beby-GAN) for rich-detail SISR. Relaxing the rigid one-to-one constraint, we allow the estimated patches to dynamically seek trustworthy surrogates of supervision during training, which is beneficial to producing more reasonable details. Besides, we propose a region-aware adversarial learning strategy that directs our model to focus on generating details for textured areas adaptively. Extensive experiments justify the effectiveness of our method. An ultra-high-resolution 4K dataset is also constructed to facilitate future super-resolution research.
Wenbo Li 0002, Kun Zhou 0001, Lu Qi 0001, Liying Lu, Jiangbo Lu
AAAI5
2022 Video Frame Interpolation with Transformer
abstract
Video frame interpolation (VFI), which aims to synthesize intermediate frames of a video, has made remarkable progress with development of deep convolutional networks over past years. Existing methods built upon convolutional networks generally face challenges of handling large motion due to the locality of convolution operations. To overcome this limitation, we introduce a novel framework, which takes advantage of Transformer to model long-range pixel correlation among video frames. Further, our network is equipped with a novel cross-scale window-based attention mechanism, where cross-scale windows interact with each other. This design effectively enlarges the receptive field and aggregates multi-scale information. Extensive quantitative and qualitative experiments demonstrate that our method achieves new state-of-the-art results on various benchmarks.
Liying Lu, Ruizheng Wu, Huaijia Lin, Jiangbo Lu, Jiaya Jia
CVPR4
2022 Revisiting Temporal Alignment for Video Restoration
abstract
Long-range temporal alignment is critical yet challenging for video restoration tasks. Recently, some works attempt to divide the long-range alignment into several sub-alignments and handle them progressively. Although this operation is helpful in modeling distant correspondences, error accumulation is inevitable due to the propagation mechanism. In this work, we present a novel, generic iterative alignment module which employs a gradual refinement scheme for sub-alignments, yielding more accurate motion compensation. To further enhance the alignment accuracy and temporal consistency, we develop a non-parametric re-weighting method, where the importance of each neighboring frame is adaptively evaluated in a spatial-wise way for aggregation. By virtue of the proposed strategies, our model achieves state-of-the-art performance on multiple benchmarks across a range of video restoration tasks including video super-resolution, denoising and deblurring.
Kun Zhou 0001, Wenbo Li 0002, Liying Lu, Xiaoguang Han 0001, Jiangbo Lu
CVPR5
2022 ScatterNet: Point Cloud Learning via Scatters
abstract
Design of point cloud shape descriptors is a challenging problem in practical applications due to the sparsity and the inscrutable distribution of the point clouds. In this paper, we propose ScatterNet, a novel 3D local feature learning approach for exploring and aggregating hypothetical scatters of the point clouds. Scatters of relational points are first organized in point cloud via guided explorations, and then propagated back to extend the capacity in representing the point-wise characteristics. We provide an practical implementation of the ScatterNet, which involves an unique scatter exploration operator and a scatter convolution operator. Our method achieves the state-of-the-art performance on several point cloud analysis tasks like classification, part segmentation and normal estimation. The source code of ScatterNet is available in supplementary materials.
Nianjuan Jiang, Jiangbo Lu, Mingang Chen, Ran Yi 0002, Lizhuang Ma
ACM Multimedia3
2022 HEMlets PoSh: Learning Part-Centric Heatmap Triplets for 3D Human Pose and Shape Estimation
abstract
Estimating 3D human pose from a single image is a challenging task. This work attempts to address the uncertainty of lifting the detected 2D joints to the 3D space by introducing an intermediate state - Part-Centric Heatmap Triplets (HEMlets), which shortens the gap between the 2D observation and the 3D interpretation. The HEMlets utilize three joint-heatmaps to represent the relative depth information of the end-joints for each skeletal body part. In our approach, a Convolutional Network (ConvNet) is first trained to predict HEMlets from the input image, followed by a volumetric joint-heatmap regression. We leverage on the integral operation to extract the joint locations from the volumetric heatmaps, guaranteeing end-to-end learning. Despite the simplicity of the network design, the quantitative comparisons show a significant performance improvement over the best-of-grade methods (e.g., 20 percent on Human3.6M). The proposed method naturally supports training with "in-the-wild" images, where only weakly-annotated relative depth information of skeletal joints is available. This further improves the generalization ability of our model, as validated by qualitative comparisons on outdoor images. Leveraging the strength of the HEMlets pose estimation, we further design and append a shallow yet effective network module to regress the SMPL parameters of the body pose and shape. We term the entire HEMlets-based human pose and shape recovery pipeline HEMlets PoSh. Extensive quantitative and qualitative experiments on the existing human body recovery benchmarks justify the state-of-the-art results obtained with our HEMlets PoSh approach.
Kun Zhou 0001, Xiaoguang Han 0001, Nianjuan Jiang, Kui Jia, Jiangbo Lu
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Group Sparsity Mixture Model and Its Application on Image Denoising
abstract
Prior learning is a fundamental problem in the field of image processing. In this paper, we conduct a detailed study on (1) how to model and learn the prior of the image patch group, which consists of a group of non-local similar image patches, and (2) how to apply the learned prior to the whole image denoising task. To tackle the first problem, we propose a new prior model named Group Sparsity Mixture Model (GSMM). With the bilateral matrix multiplication, the GSMM can model both the local feature of a single patch and the relation among non-local similar patches, and thus it is very suitable for patch group based prior learning. This is supported by the parameter analysis which demonstrates that the learned GSMM successfully captures the inherent strong sparsity embodied in the image patch group. Besides, as a mixture model, GSMM can be used for patch group classification. This makes the image denoising method based on GSMM capable of processing patch groups flexibly. To tackle the second problem, we propose an efficient and effective patch group based image denoising framework, which is plug-and-play and compatible with any patch group prior model. Using this framework, we construct two versions of GSMM based image denoising methods, both of which outperform the competing methods based on other prior models, e.g., Field of Experts (FoE) and Gaussian Mixture Model (GMM). Also, the better version is competitive with the state-of-the-art model based method WNNM with about ×8 faster average running speed.
Haosen Liu 0001, Laquan Li, Jiangbo Lu, Tan Shan
IEEE Trans. Image Process.3
2021 MASA-SR: Matching Acceleration and Spatial Adaptation for Reference-Based Image Super-Resolution
abstract
Reference-based image super-resolution (RefSR) has shown promising success in recovering high-frequency details by utilizing an external reference image (Ref). In this task, texture details are transferred from the Ref image to the low-resolution (LR) image according to their point- or patch-wise correspondence. Therefore, high-quality correspondence matching is critical. It is also desired to be computationally efficient. Besides, existing RefSR methods tend to ignore the potential large disparity in distributions between the LR and Ref images, which hurts the effectiveness of the information utilization. In this paper, we propose the MASA network for RefSR, where two novel modules are designed to address these problems. The proposed Match & Extraction Module significantly reduces the computational cost by a coarse-to-fine correspondence matching scheme. The Spatial Adaptation Module learns the difference of distribution between the LR and Ref images, and remaps the distribution of Ref features to that of LR features in a spatially adaptive way. This scheme makes the network robust to handle different reference images. Extensive quantitative and qualitative experiments validate the effectiveness of our proposed model.
Liying Lu, Wenbo Li 0002, Xin Tao 0001, Jiangbo Lu, Jiaya Jia
CVPR4
2021 Video Instance Segmentation with a Propose-Reduce Paradigm
abstract
Video instance segmentation (VIS) aims to segment and associate all instances of predefined classes for each frame in videos. Prior methods usually obtain segmentation for a frame or clip first, and merge the incomplete results by tracking or matching. These methods may cause error accumulation in the merging step. Contrarily, we propose a new paradigm – Propose-Reduce, to generate complete sequences for input videos by a single step. We further build a sequence propagation head on the existing image-level instance segmentation network for long-term propagation. To ensure robustness and high recall of our proposed framework, multiple sequences are proposed where redundant sequences of the same instance are reduced. We achieve state-of-the-art performance on two representative benchmark datasets – we obtain 47.6% in terms of AP on YouTube-VIS validation set and 70.4 % for J&F on DAVIS-UVOS validation set.
Huaijia Lin, Ruizheng Wu, Shu Liu 0005, Jiangbo Lu, Jiaya Jia
ICCV4
2021 Self-Supervised Image Prior Learning with GMM from a Single Noisy Image
abstract
The lack of clean images undermines the practicability of supervised image prior learning methods, of which the training schemes require a large number of clean images. To free image prior learning from the image collection burden, a novel Self-Supervised learning method for Gaussian Mixture Model (SS-GMM) is proposed in this paper. It can simultaneously achieve the noise level estimation and the image prior learning directly from only a single noisy image. This work is derived from our study on eigenvalues of the GMM’s covariance matrix. Through statistical experiments and theoretical analysis, we conclude that (1) covariance eigenvalues for clean images hold the sparsity; and that (2) those for noisy images contain sufficient information for noise estimation. The first conclusion inspires us to impose a sparsity constraint on covariance eigenvalues during the learning process to suppress the influence of noise. The second conclusion leads to a self-contained noise estimation module of high accuracy in our proposed method. This module serves to estimate the noise level and automatically determine the specific level of the sparsity constraint. Our final derived method requires only minor modifications to the standard expectation-maximization algorithm. This makes it easy to implement. Very interestingly, the GMM learned via our proposed self-supervised learning method can even achieve better image denoising performance than its supervised counterpart, i.e., the EPLL. Also, it is on par with the state-of-the-art self-supervised deep learning method, i.e., the Self2Self. Code is available at https://github.com/HUST-Tan/SS-GMM.
Haosen Liu 0001, Jiangbo Lu, Tan Shan
ICCV3
2021 Seeing Dynamic Scene in the Dark: A High-Quality Video Dataset with Mechatronic Alignment
abstract
Low-light video enhancement is an important task. Previous work is mostly trained on paired static images or videos. We compile a new dataset formed by our new strategy that contains high-quality spatially-aligned video pairs from dynamic scenes in low- and normal-light conditions. We built it using a mechatronic system to precisely control the dynamics during the video capture process, and further align the video pairs, both spatially and temporally, by identifying the system’s uniform motion stage. Besides the dataset, we propose an end-to-end framework, in which we design a self-supervised strategy to reduce noise, while enhancing the illumination based on the Retinex theory. Extensive experiments based on various metrics and large-scale user study demonstrate the value of our dataset and effectiveness of our method. The dataset and code are available at https://github.com/dvlab-research/SDSD.
Ruixing Wang, Xiaogang Xu 0002, Chi-Wing Fu, Jiangbo Lu, Bei Yu 0001, Jiaya Jia
ICCV4
2021 Sparse Steerable Convolutions: An Efficient Learning of SE(3)-Equivariant Features for Estimation and Tracking of Object Poses in 3D Space
abstract
As a basic component of SE(3)-equivariant deep feature learning, steerable convolution has recently demonstrated its advantages for 3D semantic analysis. The advantages are, however, brought by expensive computations on dense, volumetric data, which prevent its practical use for efficient processing of 3D data that are inherently sparse. In this paper, we propose a novel design of Sparse Steerable Convolution (SS-Conv) to address the shortcoming; SS-Conv greatly accelerates steerable convolution with sparse tensors, while strictly preserving the property of SE(3)-equivariance. Based on SS-Conv, we propose a general pipeline for precise estimation of object poses, wherein a key design is a Feature-Steering module that takes the full advantage of SE(3)-equivariance and is able to conduct an efficient pose refinement. To verify our designs, we conduct thorough experiments on three tasks of 3D object semantic analysis, including instance-level 6D pose estimation, category-level 6D pose and size estimation, and category-level 6D pose tracking. Our proposed pipeline based on SS-Conv outperforms existing methods on almost all the metrics evaluated by the three tasks. Ablation studies also show the superiority of our SS-Conv over alternative convolutions in terms of both accuracy and efficiency. Our code is released publicly at https://github.com/Gorilla-Lab-SCUT/SS-Conv.
Jiehong Lin, Ke Chen 0004, Jiangbo Lu, Kui Jia
NeurIPS4
2021 JOLO-GCN: Mining Joint-Centered Light-Weight Information for Skeleton-Based Action Recognition
abstract
Skeleton-based action recognition has attracted research attentions in recent years. One common drawback in currently popular skeleton-based human action recognition methods is that the sparse skeleton information alone is not sufficient to fully characterize human motion. This limitation makes several existing methods incapable of correctly classifying action categories which exhibit only subtle motion differences. In this paper, we propose a novel framework for employing human pose skeleton and joint-centered light-weight information jointly in a two-stream graph convolutional network, namely, JOLO-GCN. Specifically, we use Joint-aligned optical Flow Patches (JFP) to capture the local subtle motion around each joint as the pivotal joint-centered visual information. Compared to the pure skeleton-based baseline, this hybrid scheme effectively boosts performance, while keeping the computational and memory overheads low. Experiments on the NTU RGB+D, NTU RGB+D 120, and the Kinetics-Skeleton dataset demonstrate clear accuracy improvements attained by the proposed method over the state-of-the-art skeleton-based methods.
Jinmiao Cai, Nianjuan Jiang, Xiaoguang Han 0001, Kui Jia, Jiangbo Lu
WACV5
2021 Image Co-Skeletonization via Co-Segmentation
abstract
Recent advances in the joint processing of a set of images have shown its advantages over individual processing. Unlike the existing works geared towards co-segmentation or co-localization, in this article, we explore a new joint processing topic: image co-skeletonization, which is defined as joint skeleton extraction of the foreground objects in an image collection. It is well known that object skeletonization in a single natural image is challenging, because there is hardly any prior knowledge available about the object present in the image. Therefore, we resort to the idea of image co-skeletonization, hoping that the commonness prior that exists across the semantically similar images can be leveraged to have such knowledge, similar to other joint processing problems such as co-segmentation. Moreover, earlier research has found that augmenting a skeletonization process with the object's shape information is highly beneficial in capturing the image context. Having made these two observations, we propose a coupled framework for co-skeletonization and co-segmentation tasks to facilitate shape information discovery for our co-skeletonization process through the co-segmentation process. While image co-skeletonization is our primary goal, the co-segmentation process might also benefit, in turn, from exploiting skeleton outputs of the co-skeletonization process as central object seeds through such a coupled framework. As a result, both can benefit from each other synergistically. For evaluating image co-skeletonization results, we also construct a novel benchmark dataset by annotating nearly 1.8 K images and dividing them into 38 semantic categories. Although the proposed idea is essentially a weakly supervised method, it can also be employed in supervised and unsupervised scenarios. Extensive experiments demonstrate that the proposed method achieves promising results in all three scenarios.
Koteswar Rao Jerripothula, Jianfei Cai 0001, Jiangbo Lu, Junsong Yuan 0001
IEEE Trans. Image Process.3
2020 MuCAN: Multi-correspondence Aggregation Network for Video Super-Resolution
Wenbo Li 0002, Xin Tao 0001, Taian Guo, Lu Qi 0001, Jiangbo Lu, Jiaya Jia
ECCV (10)5
2020 LAPAR: Linearly-Assembled Pixel-Adaptive Regression Network for Single Image Super-resolution and Beyond
abstract
Single image super-resolution (SISR) deals with a fundamental problem of upsampling a low-resolution (LR) image to its high-resolution (HR) version. Last few years have witnessed impressive progress propelled by deep learning methods. However, one critical challenge faced by existing methods is to strike a sweet spot of deep model complexity and resulting SISR quality. This paper addresses this pain point by proposing a linearly-assembled pixel-adaptive regression network (LAPAR), which casts the direct LR to HR mapping learning into a linear coefficient regression task over a dictionary of multiple predefined filter bases. Such a parametric representation renders our model highly lightweight and easy to optimize while achieving state-of-the-art results on SISR benchmarks. Moreover, based on the same idea, LAPAR is extended to tackle other restoration tasks, e.g., image denoising and JPEG image deblocking, and again, yields strong performance.
Wenbo Li 0002, Kun Zhou 0001, Lu Qi 0001, Nianjuan Jiang, Jiangbo Lu, Jiaya Jia
NeurIPS5
2019 HEMlets Pose: Learning Part-Centric Heatmap Triplets for Accurate 3D Human Pose Estimation
abstract
Estimating 3D human pose from a single image is a challenging task. This work attempts to address the uncertainty of lifting the detected 2D joints to the 3D space by introducing an intermediate state - Part-Centric Heatmap Triplets (HEMlets), which shortens the gap between the 2D observation and the 3D interpretation. The HEMlets utilize three joint-heatmaps to represent the relative depth information of the end-joints for each skeletal body part. In our approach, a Convolutional Network(ConvNet) is first trained to predict HEMlests from the input image, followed by a volumetric joint-heatmap regression. We leverage on the integral operation to extract the joint locations from the volumetric heatmaps, guaranteeing end-to-end learning. Despite the simplicity of the network design, the quantitative comparisons show a significant performance improvement over the best-of-grade method (by 20% on Human3.6M). The proposed method naturally supports training with "in-the-wild'' images, where only weakly-annotated relative depth information of skeletal joints is available. This further improves the generalization ability of our model, as validated by qualitative comparisons on outdoor images.
Kun Zhou 0001, Xiaoguang Han 0001, Nianjuan Jiang, Kui Jia, Jiangbo Lu
ICCV5
2018 Robust Video Background Identification by Dominant Rigid Motion Estimation
Kaimo Lin, Nianjuan Jiang, Loong Fah Cheong, Jiangbo Lu, Xun Xu 0002
ACCV (2)4
2018 CODE: Coherence Based Decision Boundaries for Feature Correspondence
abstract
A key challenge in feature correspondence is the difficulty in differentiating true and false matches at a local descriptor level. This forces adoption of strict similarity thresholds that discard many true matches. However, if analyzed at a global level, false matches are usually randomly scattered while true matches tend to be coherent (clustered around a few dominant motions), thus creating a coherence based separability constraint. This paper proposes a non-linear regression technique that can discover such a coherence based separability constraint from highly noisy matches and embed it into a correspondence likelihood model. Once computed, the model can filter the entire set of nearest neighbor matches (which typically contains over 90 percent false matches) for true matches. We integrate our technique into a full feature correspondence system which reliably generates large numbers of good quality correspondences over wide baselines where previous techniques provide few or no matches.
Wen-Yan Lin, Fan Wang 0010, Ming-Ming Cheng, Sai-Kit Yeung, Philip Torr 0001, Minh N. Do, Jiangbo Lu
IEEE Trans. Pattern Anal. Mach. Intell.7
2018 Locating 3D Object Proposals: A Depth-Based Online Approach
abstract
2D object proposals, quickly detected regions in an image that likely contain an object of interest, are an effective approach for improving the computational efficiency and accuracy of object detection in color images. In this paper, we propose a novel online method that generates 3D object proposals in an RGB-D video sequence. Our main observation is that depth images provide important information about the geometry of the scene. Diverging from the traditional goal of 2D object proposals to provide a high recall, we aim for precise 3D proposals. We leverage on depth information per frame and multiview scene information to obtain accurate 3D object proposals. Using efficient but robust registration enables us to combine multiple frames of a scene in near real time and generate 3D bounding boxes for potential 3D regions of interest. Using standard metrics, such as precision-recall (P-R) curves and F-measure, we show that the proposed approach is significantly more accurate than the current state-of-the-art techniques. Our online approach can be integrated into simultaneous localization and mapping-based video processing for quick 3D object localization. Our method takes less than a second in MATLAB on the UW-RGBD scene data set on a single thread CPU and, thus, has potential to be used in low-power chips in unmanned aerial vehicles, quadcopters, and drones.
Ramanpreet Singh Pahwa, Jiangbo Lu, Nianjuan Jiang, Tian-Tsong Ng, Minh N. Do
IEEE Trans. Circuits Syst. Video Technol.2
2017 Object Co-skeletonization with Co-segmentation
abstract
Recent advances in the joint processing of images have certainly shown its advantages over the individual processing. Different from the existing works geared towards co-segmentation or co-localization, in this paper, we explore a new joint processing topic: co-skeletonization, which is defined as joint skeleton extraction of common objects in a set of semantically similar images. Object skeletonization in real world images is a challenging problem, because there is no prior knowledge of the objects shape if we consider only a single image. This motivates us to resort to the idea of object co-skeletonization hoping that the commonness prior existing across the similar images may help, just as it does for other joint processing problems such as co-segmentation. Noting that skeleton can provide good scribbles for segmentation, and skeletonization, in turn, needs good segmentation, we propose a coupled framework for co-skeletonization and co-segmentation tasks so that they are well informed by each other, and benefit each other synergistically. Since it is a new problem, we also construct a benchmark dataset for the co-skeletonization task. Extensive experiments demonstrate that proposed method achieves very competitive results.
Koteswar Rao Jerripothula, Jianfei Cai 0001, Jiangbo Lu, Junsong Yuan 0001
CVPR3
2017 Direct Photometric Alignment by Mesh Deformation
abstract
The choice of motion models is vital in applications like image/video stitching and video stabilization. Conventional methods explored different approaches ranging from simple global parametric models to complex per-pixel optical flow. Mesh-based warping methods achieve a good balance between computational complexity and model flexibility. However, they typically require high quality feature correspondences and suffer from mismatches and low-textured image content. In this paper, we propose a mesh-based photometric alignment method that minimizes pixel intensity difference instead of Euclidean distance of known feature correspondences. The proposed method combines the superior performance of dense photometric alignment with the efficiency of mesh-based image warping. It achieves better global alignment quality than the feature-based counterpart in textured images, and more importantly, it is also robust to low-textured image content. Abundant experiments show that our method can handle a variety of images and videos, and outperforms representative state-of-the-art methods in both image stitching and video stabilization tasks.
Kaimo Lin, Nianjuan Jiang, Shuaicheng Liu, Loong Fah Cheong, Minh N. Do, Jiangbo Lu
CVPR6
2017 PatchMatch Filter: Edge-Aware Filtering Meets Randomized Search for Visual Correspondence
abstract
Though many tasks in computer vision can be formulated elegantly as pixel-labeling problems, a typical challenge discouraging such a discrete formulation is often due to computational efficiency. Recent studies on fast cost volume filtering based on efficient edge-aware filters provide a fast alternative to solve discrete labeling problems, with the complexity independent of the support window size. However, these methods still have to step through the entire cost volume exhaustively, which makes the solution speed scale linearly with the label space size. When the label space is huge or even infinite, which is often the case for (subpixel-accurate) stereo and optical flow estimation, their computational complexity becomes quickly unacceptable. Developed to search approximate nearest neighbors rapidly, the PatchMatch method can significantly reduce the complexity dependency on the search space size. But, its pixel-wise randomized search and fragmented data access within the 3D cost volume seriously hinder the application of efficient cost slice filtering. This paper presents a generic and fast computational framework for general multi-labeling problems called PatchMatch Filter (PMF). We explore effective and efficient strategies to weave together these two fundamental techniques developed in isolation, i.e., PatchMatch-based randomized search and efficient edge-aware image filtering. By decompositing an image into compact superpixels, we also propose superpixel-based novel search strategies that generalize and improve the original PatchMatch method. Further motivated to improve the regularization strength, we propose a simple yet effective cross-scale consistency constraint, which handles labeling estimation for large low-textured regions more reliably than a single-scale PMF algorithm. Focusing on dense correspondence field estimation in this paper, we demonstrate PMF's applications in stereo and optical flow. Our PMF methods achieve top-tier correspondence accuracy but run much faster than other related competing methods, often giving over 10-100 times speedup.
Jiangbo Lu, Yu Li 0003, Hongsheng Yang, Dongbo Min, Wei Yong Eng, Minh N. Do
IEEE Trans. Pattern Anal. Mach. Intell.1
2017 Single Image Rain Streak Decomposition Using Layer Priors
abstract
Rain streaks impair visibility of an image and introduce undesirable interference that can severely affect the performance of computer vision and image analysis systems. Rain streak removal algorithms try to recover a rain streak free background scene. In this paper, we address the problem of rain streak removal from a single image by formulating it as a layer decomposition problem, with a rain streak layer superimposed on a background layer containing the true scene content. Existing decomposition methods that address this problem employ either sparse dictionary learning methods or impose a low rank structure on the appearance of the rain streaks. While these methods can improve the overall visibility, their performance can often be unsatisfactory, for they tend to either over-smooth the background images or generate -images that still contain noticeable rain streaks. To address the problems, we propose a method that imposes priors for both the background and rain streak layers. These priors are based on Gaussian mixture models learned on small patches that can accommodate a variety of background appearances as well as the appearance of the rain streaks. Moreover, we introduce a structure residue recovery step to further separate the background residues and improve the decomposition quality. Quantitative evaluation shows our method outperforms existing methods by a large margin. We overview our method and demonstrate its effectiveness over prior work on a number of examples.
Yu Li 0003, Robby T. Tan, Xiaojie Guo 0001, Jiangbo Lu, Michael S. Brown
IEEE Trans. Image Process.4
2016 Rain Streak Removal Using Layer Priors
abstract
This paper addresses the problem of rain streak removal from a single image. Rain streaks impair visibility of an image and introduce undesirable interference that can severely affect the performance of computer vision algorithms. Rain streak removal can be formulated as a layer decomposition problem, with a rain streak layer superimposed on a background layer containing the true scene content. Existing decomposition methods that address this problem employ either dictionary learning methods or impose a low rank structure on the appearance of the rain streaks. While these methods can improve the overall visibility, they tend to leave too many rain streaks in the background image or over-smooth the background image. In this paper, we propose an effective method that uses simple patch-based priors for both the background and rain layers. These priors are based on Gaussian mixture models and can accommodate multiple orientations and scales of the rain streaks. This simple approach removes rain streaks better than the existing methods qualitatively and quantitatively. We overview our method and demonstrate its effectiveness over prior work on a number of examples.
Yu Li 0003, Robby T. Tan, Xiaojie Guo 0001, Jiangbo Lu, Michael S. Brown
CVPR4
2016 Fast Guided Global Interpolation for Depth and Motion
Yu Li 0003, Dongbo Min, Minh N. Do, Jiangbo Lu
ECCV (3)4
2016 SEAGULL: Seam-Guided Local Alignment for Parallax-Tolerant Image Stitching
Kaimo Lin, Nianjuan Jiang, Loong Fah Cheong, Minh N. Do, Jiangbo Lu
ECCV (3)5
2016 RepMatch: Robust Feature Matching and Pose for Reconstructing Modern Cities
Wen-Yan Lin, Nianjuan Jiang, Minh N. Do, Jiangbo Lu
ECCV (1)6
2016 Action Recognition in Still Images With Minimum Annotation Efforts
abstract
We focus on the problem of still image-based human action recognition, which essentially involves making prediction by analyzing human poses and their interaction with objects in the scene. Besides image-level action labels (e.g., riding, phoning), during both training and testing stages, existing works usually require additional input of human bounding boxes to facilitate the characterization of the underlying human-object interactions. We argue that this additional input requirement might severely discourage potential applications and is not very necessary. To this end, a systematic approach was developed in this paper to address this challenging problem of minimum annotation efforts, i.e., to perform recognition in the presence of only image-level action labels in the training stage. Experimental results on three benchmark data sets demonstrate that compared with the state-of-the-art methods that have privileged access to additional human bounding-box annotations, our approach achieves comparable or even superior recognition accuracy using only action annotations in training. Interestingly, as a by-product in many cases, our approach is able to segment out the precise regions of underlying human-object interactions.
Yu Zhang 0004, Li Cheng 0001, Jianxin Wu 0001, Jianfei Cai 0001, Minh N. Do, Jiangbo Lu
IEEE Trans. Image Process.6
2016 Weakly Supervised Fine-Grained Categorization With Part-Based Image Representation
abstract
In this paper, we propose a fine-grained image categorization system with easy deployment. We do not use any object/part annotation (weakly supervised) in the training or in the testing stage, but only class labels for training images. Fine-grained image categorization aims to classify objects with only subtle distinctions (e.g., two breeds of dogs that look alike). Most existing works heavily rely on object/part detectors to build the correspondence between object parts, which require accurate object or object part annotations at least for training images. The need for expensive object annotations prevents the wide usage of these methods. Instead, we propose to generate multi-scale part proposals from object proposals, select useful part proposals, and use them to compute a global image representation for categorization. This is specially designed for the weakly supervised fine-grained categorization task, because useful parts have been shown to play a critical role in existing annotation-dependent works, but accurate part detectors are hard to acquire. With the proposed image representation, we can further detect and visualize the key (most discriminative) parts in objects of different classes. In the experiments, the proposed weakly supervised method achieves comparable or better accuracy than the state-of-the-art weakly supervised methods and most existing annotation-dependent methods on three challenging datasets. Its success suggests that it is not always necessary to learn expensive object/part detectors in fine-grained image categorization.
Yu Zhang 0004, Xiu-Shen Wei, Jianxin Wu 0001, Jianfei Cai 0001, Jiangbo Lu, Minh N. Do
IEEE Trans. Image Process.5
2016 Multiple Human Identification and Cosegmentation: A Human-Oriented CRF Approach With Poselets
abstract
Localizing, identifying, and extracting humans with consistent appearance jointly from a personal photo stream is an important problem and has wide applications. The strong variations in foreground and background and irregularly occurring foreground humans make this realistic problem challenging. Inspired by advancements in object detection, scene understanding, and image cosegmentation, we explore explicit constraints to label and segment human objects rather than other nonhuman objects and “stuff.” We refer to such a problem as multiple human identification and cosegmentation (MHIC). To identify specific human subjects, we propose an efficient human instance detector by combining an extended color line model with a poselet-based human detector. Moreover, to capture high-level human shape information, a novel soft shape cue is proposed. It is initialized by the human detector, then further enhanced through a generalized geodesic distance transform, and finally refined with a joint bilateral filter. We also propose to capture the rich feature context around each pixel by using an adaptive cross-region data structure, which gives a higher discriminative power than a single pixel-based estimation. The high-level object cues from the detector and the shape are then integrated with the low-level pixel cues and midlevel contour cues into a principled conditional random field (CRF) framework, which can be efficiently solved by using fast graph cut algorithms. We evaluate our method over a newly created NTU-MHIC human dataset, which contains 351 images with manually annotated groundtruth segmentation. Both visual and quantitative results demonstrate that our method achieves state-of-the-art performance for the MHIC task.
Hongyuan Zhu 0002, Jiangbo Lu, Jianfei Cai 0001, Jianmin Zheng, Shijian Lu, Nadia Magnenat-Thalmann
IEEE Trans. Multim.2
2015 Direct structure estimation for 3D reconstruction
abstract
Most conventional structure-from-motion (SFM) techniques require camera pose estimation before computing any scene structure. In this work we show that when combined with single/multiple homography estimation, the general Euclidean rigidity constraint provides a simple formulation for scene structure recovery without explicit camera pose computation. This direct structure estimation (DSE) opens a new way to design a SFM system that reverses the order of structure and motion estimation. We show that this alternative approach works well for recovering scene structure and camera poses from sideway motion given planar or general man-made scenes.
Nianjuan Jiang, Wen-Yan Lin, Minh N. Do, Jiangbo Lu
CVPR4
2015 SPM-BP: Sped-Up PatchMatch Belief Propagation for Continuous MRFs
abstract
Markov random fields are widely used to model many computer vision problems that can be cast in an energy minimization framework composed of unary and pairwise potentials. While computationally tractable discrete optimizers such as Graph Cuts and belief propagation (BP) exist for multi-label discrete problems, they still face prohibitively high computational challenges when the labels reside in a huge or very densely sampled space. Integrating key ideas from PatchMatch of effective particle propagation and resampling, PatchMatch belief propagation (PMBP) has been demonstrated to have good performance in addressing continuous labeling problems and runs orders of magnitude faster than Particle BP (PBP). However, the quality of the PMBP solution is tightly coupled with the local window size, over which the raw data cost is aggregated to mitigate ambiguity in the data constraint. This dependency heavily influences the overall complexity, increasing linearly with the window size. This paper proposes a novel algorithm called sped-up PMBP (SPM-BP) to tackle this critical computational bottleneck and speeds up PMBP by 50-100 times. The crux of SPM-BP is on unifying efficient filter-based cost aggregation and message passing with PatchMatch-based particle generation in a highly effective way. Though simple in its formulation, SPM-BP achieves superior performance for sub-pixel accurate stereo and optical-flow on benchmark datasets when compared with more complex and task-specific approaches.
Yu Li 0003, Dongbo Min, Michael S. Brown, Minh N. Do, Jiangbo Lu
ICCV5
2015 PISA: Pixelwise Image Saliency by Aggregating Complementary Appearance Contrast Measures With Edge-Preserving Coherence
abstract
Driven by recent vision and graphics applications such as image segmentation and object recognition, computing pixel-accurate saliency values to uniformly highlight foreground objects becomes increasingly important. In this paper, we propose a unified framework called pixelwise image saliency aggregating (PISA) various bottom-up cues and priors. It generates spatially coherent yet detail-preserving, pixel-accurate, and fine-grained saliency, and overcomes the limitations of previous methods, which use homogeneous superpixel based and color only treatment. PISA aggregates multiple saliency cues in a global context, such as complementary color and structure contrast measures, with their spatial priors in the image domain. The saliency confidence is further jointly modeled with a neighborhood consistence constraint into an energy minimization formulation, in which each pixel will be evaluated with multiple hypothetical saliency levels. Instead of using global discrete optimization methods, we employ the cost-volume filtering technique to solve our formulation, assigning the saliency levels smoothly while preserving the edge-aware structure details. In addition, a faster version of PISA is developed using a gradient-driven image subsampling strategy to greatly improve the runtime efficiency while keeping comparable detection accuracy. Extensive experiments on a number of public data sets suggest that PISA convincingly outperforms other state-of-the-art approaches. In addition, with this work, we also create a new data set containing 800 commodity images for evaluating saliency detection.
Keze Wang, Liang Lin 0004, Jiangbo Lu, Chenglong Li 0002, Keyang Shi
IEEE Trans. Image Process.3
2014 DAISY Filter Flow: A Generalized Discrete Approach to Dense Correspondences
abstract
Establishing dense correspondences reliably between a pair of images is an important vision task with many applications. Though significant advance has been made towards estimating dense stereo and optical flow fields for two images adjacent in viewpoint or in time, building reliable dense correspondence fields for two general images still remains largely unsolved. For instance, two given images sharing some content exhibit dramatic photometric and geometric variations, or they depict different 3D scenes of similar scene characteristics. Fundamental challenges to such an image or scene alignment task are often multifold, which render many existing techniques fall short of producing dense correspondences robustly and efficiently. This paper presents a novel approach called DAISY filter flow (DFF) to address this challenging task. Inspired by the recent PatchMatch Filter technique, we leverage and extend a few established methods: DAISY descriptors, filter-based efficient flow inference, and the PatchMatch fast search. Coupling and optimizing these modules seamlessly with image segments as the bridge, the proposed DFF approach enables efficiently performing dense descriptor-based correspondence field estimation in a generalized high-dimensional label space, which is augmented by scales and rotations. Experiments on a variety of challenging scenes show that our DFF approach estimates spatially coherent yet discontinuity-preserving image alignment results both robustly and efficiently.
Hongsheng Yang, Wen-Yan Lin, Jiangbo Lu
CVPR3
2014 Bilateral Functions for Global Motion Modeling
Wen-Yan Lin, Ming-Ming Cheng, Jiangbo Lu, Hongsheng Yang, Minh N. Do, Philip Torr 0001
ECCV (4)3
2014 Poselet-based multiple human identification and cosegmentation
abstract
Localizing, identifying and extracting human groups with consistent appearance jointly from a personal photo stream is an important problem and has wide applications. Inspired by recent advances in object detection, scene understanding and image cosegmentation, in this paper we explore explicit constraints to label and segment human objects rather than other non-human objects and “stuff”. We propose a novel soft human shape cue, which is initialized by color line poselet-based human part detection, further processed through a generalized geodesic distance transform, and refined finally with a joint bilateral filter. Such a high-level object cue is then integrated with other low-level unary and pairwise terms into a principled conditional random field framework, which can be efficiently solved by fast graph cut algorithms. We evaluate our algorithm over the FlickrMFC human dataset, and show that it achieves state-of-the-art performance for this challenging task.
Hongyuan Zhu 0002, Jiangbo Lu, Jianfei Cai 0001, Jianmin Zheng, Nadia Magnenat-Thalmann
ICIP2
2014 Multiple foreground recognition and cosegmentation: An object-oriented CRF model with robust higher-order potentials
abstract
Localizing, recognizing, and segmenting multiple foreground objects jointly from a general user's photo stream that records a specific event is an important task with many useful applications. As argued in recent Multiple Foreground Cosegmentation (MFC) work by Kim and Xing, this task is very challenging in that it contrasts substantially from the classical cosegmentation problem, and aims to parse a set of realistic event photos but each containing irregularly occurring multiple foregrounds with high appearance and scene configuration variations. Inspired by the impressive advance in scene understanding and object recognition, this paper casts the multiple foreground recognition and cosegmentation (MFRC) problem within a conditional random fields (CRFs) framework in a principled manner. We capitalize centrally on the key objective that MFRC is to segment out and annotate foreground objects or “things” rather than “stuff”. To this end, we exploit a few complementary objectness cues (e.g. contours, object detectors and layout) and propose novel and efficient methods to capture object-level information. Integrating object potentials as soft constraints (e.g. robust higher-order potentials defined over detected object regions) with low-level unary and pairwise terms holistically, we solve the MFRC task with a probabilistic CRF model. The inference for such a CRF model is performed efficiently with graph cut based move making algorithms. With a minimal amount of user annotations on just a few example photos, the proposed approach produces spatially coherent, boundary-aligned segmentation results with correct and consistent object labeling. Experiments on the FlickrMFC dataset justify that our method achieves state-of-the-art performance.
Hongyuan Zhu 0002, Jiangbo Lu, Jianfei Cai 0001, Jianmin Zheng, Nadia Magnenat-Thalmann
WACV2
2014 Fast Global Image Smoothing Based on Weighted Least Squares
abstract
This paper presents an efficient technique for performing a spatially inhomogeneous edge-preserving image smoothing, called fast global smoother. Focusing on sparse Laplacian matrices consisting of a data term and a prior term (typically defined using four or eight neighbors for 2D image), our approach efficiently solves such global objective functions. In particular, we approximate the solution of the memory-and computation-intensive large linear system, defined over a d-dimensional spatial domain, by solving a sequence of 1D subsystems. Our separable implementation enables applying a linear-time tridiagonal matrix algorithm to solve d three-point Laplacian matrices iteratively. Our approach combines the best of two paradigms, i.e., efficient edge-preserving filters and optimization-based smoothing. Our method has a comparable runtime to the fast edge-preserving filters, but its global optimization formulation overcomes many limitations of the local filtering approaches. Our method also achieves high-quality results as the state-of-the-art optimization-based techniques, but runs ∼10-30 times faster. Besides, considering the flexibility in defining an objective function, we further propose generalized fast algorithms that perform Lγ norm smoothing (0 < γ < 2) and support an aggregated (robust) data term for handling imprecise data constraints. We demonstrate the effectiveness and efficiency of our techniques in a range of image processing and computer graphics applications.
Dongbo Min, Sunghwan Choi, Jiangbo Lu, Bumsub Ham, Kwanghoon Sohn, Minh N. Do
IEEE Trans. Image Process.3
2014 Efficient Hybrid Tree-Based Stereo Matching With Applications to Postcapture Image Refocusing
abstract
Estimating dense correspondence or depth information from a pair of stereoscopic images is a fundamental problem in computer vision, which finds a range of important applications. Despite intensive past research efforts in this topic, it still remains challenging to recover the depth information both reliably and efficiently, especially when the input images contain weakly textured regions or are captured under uncontrolled, real-life conditions. Striking a desired balance between computational efficiency and estimation quality, a hybrid minimum spanning tree-based stereo matching method is proposed in this paper. Our method performs efficient nonlocal cost aggregation at pixel-level and region-level, and then adaptively fuses the resulting costs together to leverage their respective strength in handling large textureless regions and fine depth discontinuities. Experiments on the standard Middlebury stereo benchmark show that the proposed stereo method outperforms all prior local and nonlocal aggregation-based methods, achieving particularly noticeable improvements for low texture regions. To further demonstrate the effectiveness of the proposed stereo method, also motivated by the increasing desire to generate expressive depth-induced photo effects, this paper is tasked next to address the emerging application of interactive depth-of-field rendering given a real-world stereo image pair. To this end, we propose an accurate thin-lens model for synthetic depth-of-field rendering, which considers the user-stroke placement and camera-specific parameters and performs the pixel-adapted Gaussian blurring in a principled way. Taking ~1.5 s to process a pair of 640×360 images in the off-line step, our system named Scribble2focus allows users to interactively select in-focus regions by simple strokes using the touch screen and returns the synthetically refocused images instantly to the user.
Dung T. Vu, Benjamin Chidester, Hongsheng Yang, Minh N. Do, Jiangbo Lu
IEEE Trans. Image Process.5
2013 Patch Match Filter: Efficient Edge-Aware Filtering Meets Randomized Search for Fast Correspondence Field Estimation
abstract
Though many tasks in computer vision can be formulated elegantly as pixel-labeling problems, a typical challenge discouraging such a discrete formulation is often due to computational efficiency. Recent studies on fast cost volume filtering based on efficient edge-aware filters have provided a fast alternative to solve discrete labeling problems, with the complexity independent of the support window size. However, these methods still have to step through the entire cost volume exhaustively, which makes the solution speed scale linearly with the label space size. When the label space is huge, which is often the case for (sub pixel-accurate) stereo and optical flow estimation, their computational complexity becomes quickly unacceptable. Developed to search approximate nearest neighbors rapidly, the Patch Match method can significantly reduce the complexity dependency on the search space size. But, its pixel-wise randomized search and fragmented data access within the 3D cost volume seriously hinder the application of efficient cost slice filtering. This paper presents a generic and fast computational framework for general multi-labeling problems called Patch Match Filter (PMF). For the very first time, we explore effective and efficient strategies to weave together these two fundamental techniques developed in isolation, i.e., Based-based randomized search and efficient edge-aware image filtering. By decompositing an image into compact super pixels, we also propose super pixel-based novel search strategies that generalize and improve the original Patch Match method. Focusing on dense correspondence field estimation in this paper, we demonstrate PMF's applications in stereo and optical flow. Our PMF methods achieve state-of-the-art correspondence accuracy but run much faster than other competing methods, often giving over 10-times speedup for large label space cases.
Jiangbo Lu, Hongsheng Yang, Dongbo Min, Minh N. Do
CVPR1
2013 PISA: Pixelwise Image Saliency by Aggregating Complementary Appearance Contrast Measures with Spatial Priors
abstract
Driven by recent vision and graphics applications such as image segmentation and object recognition, assigning pixel-accurate saliency values to uniformly highlight foreground objects becomes increasingly critical. More often, such fine-grained saliency detection is also desired to have a fast runtime. Motivated by these, we propose a generic and fast computational framework called PISA - Pixel wise Image Saliency Aggregating complementary saliency cues based on color and structure contrasts with spatial priors holistically. Overcoming the limitations of previous methods often using homogeneous super pixel-based and color contrast-only treatment, our PISA approach directly performs saliency modeling for each individual pixel and makes use of densely overlapping, feature-adaptive observations for saliency measure computation. We further impose a spatial prior term on each of the two contrast measures, which constrains pixels rendered salient to be compact and also centered in image domain. By fusing complementary contrast measures in such a pixel wise adaptive manner, the detection effectiveness is significantly boosted. Without requiring reliable region segmentation or post-relaxation, PISA exploits an efficient edge-aware image representation and filtering technique and produces spatially coherent yet detail-preserving saliency maps. Extensive experiments on three public datasets demonstrate PISA's superior detection accuracy and competitive runtime speed over the state-of-the-arts approaches.
Keyang Shi, Keze Wang, Jiangbo Lu, Liang Lin 0004
CVPR3
2013 Robust Non-parametric Data Fitting for Correspondence Modeling
abstract
We propose a generic method for obtaining nonparametric image warps from noisy point correspondences. Our formulation integrates a huber function into a motion coherence framework. This makes our fitting function especially robust to piecewise correspondence noise (where an image section is consistently mismatched). By utilizing over parameterized curves, we can generate realistic nonparametric image warps from very noisy correspondence. We also demonstrate how our algorithm can be used to help stitch images taken from a panning camera by warping the images onto a virtual push-broom camera imaging plane.
Wen-Yan Lin, Ming-Ming Cheng, Shuai Zheng 0001, Jiangbo Lu, Nigel T. Crook
ICCV4
2013 Joint Histogram-Based Cost Aggregation for Stereo Matching
abstract
This paper presents a novel method for performing efficient cost aggregation in stereo matching. The cost aggregation problem is reformulated from the perspective of a histogram, giving us the potential to reduce the complexity of the cost aggregation in stereo matching significantly. Differently from previous methods which have tried to reduce the complexity in terms of the size of an image and a matching window, our approach focuses on reducing the computational redundancy that exists among the search range, caused by a repeated filtering for all the hypotheses. Moreover, we also reduce the complexity of the window-based filtering through an efficient sampling scheme inside the matching window. The tradeoff between accuracy and complexity is extensively investigated by varying the parameters used in the proposed method. Experimental results show that the proposed method provides high-quality disparity maps with low complexity and outperforms existing local methods. This paper also provides new insights into complexity-constrained stereo-matching algorithm design.
Dongbo Min, Jiangbo Lu, Minh N. Do
IEEE Trans. Pattern Anal. Mach. Intell.2
2012 Cross-based local multipoint filtering
abstract
This paper presents a cross-based framework of performing local multipoint filtering efficiently. We formulate the filtering process as a local multipoint regression problem, consisting of two main steps: 1) multipoint estimation, calculating the estimates for a set of points within a shape-adaptive local support, and 2) aggregation, fusing a number of multipoint estimates available for each point. Compared with the guided filter that applies the linear regression to all pixels covered by a fixed-sized square window non-adaptively, the proposed filtering framework is a more generalized form. Two specific filtering methods are instantiated from this framework, based on piecewise constant and piecewise linear modeling, respectively. Leveraging a cross-based local support representation and integration technique, the proposed filtering methods achieve theoretically strong results in an efficient manner, with the two main steps' complexity independent of the filtering kernel size. We demonstrate the strength of the proposed filters in various applications including stereo matching, depth map enhancement, edge-preserving smoothing, color image denoising, detail enhancement, and flash/no-flash denoising.
Jiangbo Lu, Keyang Shi, Dongbo Min, Liang Lin 0004, Minh N. Do
CVPR1
2012 Weighted mode filtering and its applications to depth video enhancement and coding
abstract
This paper presents a novel approach for improving the quality of depth video. Given a high-quality color image and its corresponding low-quality depth image, we handle various artifacts which may exist on the depth video by applying a weighted mode filtering method based on a joint histogram. When the histogram is generated, the weight based on color similarity between reference and neighboring pixels on the color image is computed and then used for counting each bin on the joint histogram of the depth map. A final solution is determined by seeking a global mode on the histogram. Experimental results show that the proposed method has outstanding performance and is very efficient in various applications such as depth video enhancement and compression.
Dongbo Min, Jiangbo Lu, Minh N. Do
ICASSP2
2012 Efficient video compression methods for a lightweight tele-immersive video chat system
abstract
A lightweight tele-immersive (TI) video chat system named CuteChat has been developed recently to provide a radically new video chat experience by merging each participant in the same shared space, allowing them to interact more naturally in an integrated manner. This paper presents an insight of the coding component in the system. Specifically, we present an efficient standard-compliant method to compress and deliver video contents in a more semantic manner. Taking into account the characteristics of real-life video chat sequences, we also propose a low-complexity coding method to significantly reduce the encoder complexity while retaining acceptable visual quality. Experimental results have shown the effectiveness of the proposed methods in comparison with the existing methods.
Jiangbo Lu, Minh N. Do
ISCAS2
2012 ITEM: immersive telepresence for entertainment and meetings with commodity setup
abstract
This paper presents an Immersive Telepresence system for Entertainment and Meetings (ITEM). The system aims to provide a radically new video communication experience by seamlessly merging participants into the same virtual space to allow a natural interaction among them and shared collaborative contents. With the goal to make a scalable, flexible system for various business solutions as well as easily accessible by massive consumers, we address the challenges in the whole pipeline of media processing, communication, and displaying in our design and realization of such a system. Extensive experiments show the developed system runs reliably and comfortably in real time with a minimal setup requirement (e.g., a webcam, a laptop/desktop connected to the public Internet) for tele-immersive video communication. With such a really minimal deployment requirement, we present a variety of interesting applications and user experiences created by ITEM.
Tien Dung Vu, Hongsheng Yang, Jiangbo Lu, Minh N. Do
ACM Multimedia4
2012 Quicktoon: a real-time video stylization and sharing system on general processors
abstract
We present a video stylization and sharing system named QuickToon, which supports generating a variety of useful and pleasing visual effects and allows easy use in video chat or sharing via social networking services (SNS). Based on our highly efficient edge-preserving smoothing filter, non-photo-realistic video rendering effects can be generated on the fly such as skin beautification, cartoon-like rendition, object outline, pencil sketch, and color stroke. Without requiring any GPUs that would otherwise seriously limits portability, the QuickToon system runs comfortably in real time for VGA- or HD-sized videos on general processors such as CPUs. The system can take in video frames either from a live webcam feed or any photos/videos from the local store, while the transformed imagery can be used in a live Skype video call, or saved locally, or uploaded and shared to SNS with one click. We demonstrate with concrete examples the functionality of this system, and underline its utility in video communication and photo/video sharing.
Hongsheng Yang, Huanliang Sun, Jiangbo Lu
ACM Multimedia3
2012 Depth Video Enhancement Based on Weighted Mode Filtering
abstract
This paper presents a novel approach for depth video enhancement. Given a high-resolution color video and its corresponding low-quality depth video, we improve the quality of the depth video by increasing its resolution and suppressing noise. For that, a weighted mode filtering method is proposed based on a joint histogram. When the histogram is generated, the weight based on color similarity between reference and neighboring pixels on the color image is computed and then used for counting each bin on the joint histogram of the depth map. A final solution is determined by seeking a global mode on the histogram. We show that the proposed method provides the optimal solution with respect to L(1) norm minimization. For temporally consistent estimate on depth video, we extend this method into temporally neighboring frames. Simple optical flow estimation and patch similarity measure are used for obtaining the high-quality depth video in an efficient manner. Experimental results show that the proposed method has outstanding performance and is very efficient, compared with existing methods. We also show that the temporally consistent enhancement of depth video addresses a flickering problem and improves the accuracy of depth video.
Dongbo Min, Jiangbo Lu, Minh N. Do
IEEE Trans. Image Process.2
2011 A revisit to MRF-based depth map super-resolution and enhancement
abstract
This paper presents a Markov Random Field (MRF)-based approach for depth map super-resolution and enhancement. Given a low-resolution or moderate quality depth map, we study the problem of enhancing its resolution or quality with a registered high-resolution color image. Different from the previous methods, this MRF-based approach is based on a novel data term formulation that fits well to the unique characteristics of depth maps. We also discuss a few important design choices that boost the performance of general MRF-based methods. Experimental results show that our proposed approach achieves high resolution depth maps at more desirable quality, both qualitatively and quantitatively. It can also be applied to enhance the depth maps derived with state-of-the-art stereo methods, resulting in the raised ranking based on the Middlebury benchmark.
Jiangbo Lu, Dongbo Min, Ramanpreet Singh Pahwa, Minh N. Do
ICASSP1
2011 A revisit to cost aggregation in stereo matching: How far can we reduce its computational redundancy?
abstract
This paper presents a novel method for performing an efficient cost aggregation in stereo matching. The cost aggregation problem is re-formulated with a perspective of a histogram, and it gives us a potential to reduce the complexity of the cost aggregation significantly. Different from the previous methods which have tried to reduce the complexity in terms of the size of an image and a matching window, our approach focuses on reducing the computational redundancy which exists among the search range, caused by a repeated filtering for all disparity hypotheses. Moreover, we also reduce the complexity of the window-based filtering through an efficient sampling scheme inside the matching window. The trade-off between accuracy and complexity is extensively investigated into parameters used in the proposed method. Experimental results show that the proposed method provides high-quality disparity maps with low complexity. This work provides new insights into complexity-constrained stereo matching algorithm design.
Dongbo Min, Jiangbo Lu, Minh N. Do
ICCV2
2011 CuteChat: a lightweight tele-immersive video chat system
abstract
This paper presents a lightweight tele-immersive video chat system named CuteChat. Based on our recently developed video object cutout technology, the CuteChat system is designed and optimized to provide a radically new video chat experience by merging each participant in the same shared space, allowing them to interact more naturally in an integrated manner. With the goal to make the system easily accessible by massive consumers, we address the challenges in the whole pipeline of video processing, coding, communication, composition, and playback. Extensive experiments have shown that the proposed CuteChat system runs reliably and comfortably in real time on one's laptop or desktop PC, and it needs only a commodity webcam for video acquisition and just public Internet for tele-immersive video conferencing. With such a really minimal deployment requirement, we present a variety of interesting applications and user experiences created by the CuteChat system.
Jiangbo Lu, Zeping Niu, Bhavdeep Singh, Zhiping Luo, Minh N. Do
ACM Multimedia1
2011 Real-Time and Accurate Stereo: A Scalable Approach With Bitwise Fast Voting on CUDA
abstract
This paper proposes a real-time design for accurate stereo matching on compute unified device architecture (CUDA). We present a leading local algorithm and then accelerate it by parallel computing. High matching accuracy is achieved by cost aggregation over shape-adaptive support regions and disparity refinement using reliable initial estimates. A novel sample-and-restore scheme is proposed to make the algorithm scalable, capable of attaining several times speedup at the expense of minor accuracy degradation. The refinement and the restoration are jointly realized by a local voting method. To accelerate the voting on CUDA, a graphics processing unit (GPU)-oriented bitwise fast voting method is proposed, faster than the traditional histogram-based approach with two orders of magnitude. The whole algorithm is parallelized on CUDA at a fine granularity, efficiently exploiting the computing resources of GPUs. Our design is among the fastest stereo matching methods on GPUs. Evaluated in the Middlebury stereo benchmark, the proposed design produces the most accurate results among the real-time methods. The advantages of speed, accuracy, and desirable scalability advocate our design for practical applications such as robotics systems and multiview teleconferencing.
Ke Zhang 0012, Jiangbo Lu, Qiong Yang, Gauthier Lafruit, Rudy Lauwereins, Luc Van Gool
IEEE Trans. Circuits Syst. Video Technol.2
2010 Scheduling transmissions of real-time video coded frames in video conferencing applications over the Internet
abstract
This paper presents a novel mechanism for scheduling packet transmissions of real-time video conferencing over the public Internet. By scheduling transmissions of video packets at different priorities in the presence of packet losses and delays, we aim to deliver videos with perceptually pleasing and consistent quality at receivers. Fundamental to this new mechanism is a realization of uneven packet transmission rates (UPTR) in the Internet, which allow an instantaneous packet transmission rate to fluctuate around an average rate, without significantly affecting the average loss rate and delay in transmissions. By exploiting this generic UPTR idea, we propose two application scenarios that provide positive design alternatives for robust realtime video delivery over the lossy Internet. Experimental results clearly show the advantages of our proposed methods over those previous solutions without awareness of UPTR.
Jiangbo Lu, Benjamin W. Wah
ICME1
2009 Interpolation error as a quality metric for stereo: Robust, or not?
abstract
To properly benchmark and stimulate current stereo algorithms specifically in the application context of view interpolation, a robust quantitative evaluation approach is important. As a prevailing quality assessment method, interpolation error has been widely used. It measures the distortions between an interpolated image and a real camera image for a desired virtual viewpoint. However, is it a robust quality metric, especially when state-of-the-art stereo technology is developing so fast? This paper hence focuses on revealing several rarely attended weaknesses that make the interpolation error evaluation paradigm vulnerable. In addition, we propose an alternative evaluation method as an early attempt at addressing these challenges, from a perspective of communication system. Evaluation of representative stereo methods from the Middlebury Web site shows that the new approach yields consistent quality assessment outcomes.
Jiangbo Lu, Qiong Yang, Gauthier Lafruit
ICASSP1
2009 Real-time stereo matching: A cross-based local approach
abstract
We propose an area-based local stereo matching algorithm that yields accurate disparity estimates, while achieving the real-time speed completely on the graphics processing unit (GPU). For a local stereo method, the key challenge is to decide an appropriate support window for the pixel under consideration. Our stereo method starts with computing an upright local cross adaptively for each anchor pixel, which defines a per-pixel support skeleton. Next, based on this compact local cross representation, we aggregate the matching costs in a shape adaptive full support region using two orthogonal integration steps. Approximating scene structures accurately, the proposed method is among the best-performing real-time stereo methods according to the benchmark Middlebury stereo evaluation. Additionally, our method is very easy to implement, memory efficient, and hence it is promising for many practical applications.
Jiangbo Lu, Ke Zhang 0012, Gauthier Lafruit, Francky Catthoor
ICASSP1
2009 Robust stereo matching with fast Normalized Cross-Correlation over shape-adaptive regions
abstract
Normalized cross-correlation (NCC) is a common matching technique to tolerate radiometric differences between stereo images. However, traditional rectangle-based NCC tends to blur the depth discontinuities. This paper proposes an efficient stereo algorithm with NCC over shape-adaptive matching regions, producing depth-discontinuity preserving disparity maps while remaining the advantage of robustness to radiometric differences. To alleviate the computational intensity, we propose an acceleration algorithm using an orthogonal integral image technique, achieving a speedup factor of 10~27. In addition, a voting scheme on reliable estimates is applied to refine the initial estimates. Experiments show that, besides the robustness, the proposed method obtains accurate disparity maps at fast speed. Our method highly ranks among the local approaches in the Middlebury stereo benchmark.
Ke Zhang 0012, Jiangbo Lu, Gauthier Lafruit, Rudy Lauwereins, Luc Van Gool
ICIP2
2009 Accurate and efficient stereo matching with robust piecewise voting
abstract
In this paper, we propose an efficient local stereo algorithm for accurate disparity estimation. First, we attain initial disparity estimates by iterating a cross-based cost aggregation process. Then, we propose a robust voting scheme to refine the initial estimates based on a piecewise smoothness prior, improving the quality in occluded regions and low-textured regions effectively. The refinement is guided by the segmentation result of input images. Unreliable initial estimates, which are detected using an efficient left-right consistency check, are rejected to increase the reliability of the voting results. Evaluated with the Middlebury stereo benchmark, our method is among the top performing local methods in accuracy. Compared to other local methods with similar accuracy, our method is faster by a factor of about two orders.
Ke Zhang 0012, Jiangbo Lu, Gauthier Lafruit, Rudy Lauwereins, Luc Van Gool
ICME2
2009 Real-time stereo-based view synthesis algorithms: A unified framework and evaluation on commodity GPUs
Sammy Rogmans, Jiangbo Lu, Philippe Bekaert, Gauthier Lafruit
Signal Process. Image Commun.2
2009 Stream-Centric Stereo Matching and View Synthesis: A High-Speed Approach on GPUs
abstract
In this paper, we propose a real-time image-based rendering (IBR) system. It is specifically designed for photorealistic view synthesis at high-speed on the graphics processing unit (GPU). We steer the proposed IBR system design with two high-level ideas. First, for cost-effective IBR, as long as the synthesized views look visually plausible, the estimated disparity and occlusion need not be correct. Hence, we jointly optimize stereo matching and view synthesis for a favorable end-to-end performance. Second, for great real-time acceleration on GPUs, all functional modules need be shaped at an early design stage, fitting the massively parallel streaming architecture of GPUs. Based on these two guidelines, we first propose a stream-centric local stereo matching algorithm. The key idea is to construct a versatile set of variable support patterns in a highly efficient manner, and then an optimal local support pattern is selected to approximate varying image structures adaptively. Next, a low-complexity adaptive view synthesis technique is proposed. It efficiently tackles visual artifacts in synthesized images, using a novel photometric outlier detection and handling scheme. We evaluated both the disparity estimation accuracy and novel view synthesis quality of the proposed approach, based on the benchmark Middlebury stereo datasets. The experiments show that our local stereo method produces consistently reliable disparity estimates for both homogeneous regions and depth discontinuities, outperforming several previous GPU-based local methods. More importantly, visually plausible intermediate views are generated by our IBR approach at high-speed on the GPU. With stereo matching and view synthesis completely running on an NVIDIA GeForce 8800 GT graphics card, the proposed IBR system reaches about 100 f/s for 450times375 stereo images with 60 disparity levels.
Jiangbo Lu, Sammy Rogmans, Gauthier Lafruit, Francky Catthoor
IEEE Trans. Circuits Syst. Video Technol.1
2009 Cross-Based Local Stereo Matching Using Orthogonal Integral Images
abstract
We propose an area-based local stereo matching algorithm for accurate disparity estimation across all image regions. A well-known challenge to local stereo methods is to decide an appropriate support window for the pixel under consideration, adapting the window shape or the pixelwise support weight to the underlying scene structures. Our stereo method tackles this problem with two key contributions. First, for each anchor pixel an upright cross local support skeleton is adaptively constructed, with four varying arm lengths decided on color similarity and connectivity constraints. Second, given the local cross-decision results, we dynamically construct a shape-adaptive full support region on the fly, merging horizontal segments of the crosses in the vertical neighborhood. Approximating image structures accurately, the proposed method is among the best performing local stereo methods according to the benchmark Middlebury stereo evaluation. Additionally, it reduces memory consumption significantly thanks to our compact local cross representation. To accelerate matching cost aggregation performed in an arbitrarily shaped 2-D region, we also propose an orthogonal integral image technique, yielding a speedup factor of 5-15 over the straightforward integration.
Ke Zhang 0012, Jiangbo Lu, Gauthier Lafruit
IEEE Trans. Circuits Syst. Video Technol.2
2008 Scalable stereo matching with Locally Adaptive Polygon Approximation
abstract
We present a scalable stereo matching algorithm based on a Locally Adaptive Polygon Approximation (LAPA) technique. For accurate local stereo matching, pixel-wise adaptive polygon-based support windows are constructed to approximate spatially varying image structures. Central to building these pixel-wise polygons is a fast algorithm that adaptively decides a set of directional scales, utilizing intensity and spatial information. Thanks to the locally adaptive support window, the proposed method achieves high stereo reconstruction quality both in depth-discontinuity regions and homogenous regions. Moreover, our LAPA-based method offers flexible scalability in terms of quality-complexity trade-off. As a specific instantiation favoring high-quality stereo estimation, our 8-direction stereo method outperforms most of the other local stereo methods and even some global optimization techniques. Another low-complexity alternative is also presented, achieving a significant speedup of up to a factor 20 with graceful accuracy degradation. Within a unified LAPA framework, our stereo method hence facilitates more flexibility in conciliating different algorithm design needs with processing performance issues.
Ke Zhang 0012, Jiangbo Lu, Gauthier Lafruit
ICIP2
2007 Fast Reliable Multi-Scale Motion Region Detection in Video Processing
abstract
Motion region detection is an important vision topic usually tackled by a background subtraction principle, which has some practical restrictions. We hence propose a multi-scale motion region detection technique that can fast and reliably segment foreground motion regions from two successive video frames. The key idea is to leverage multi-scale structural aggregation to effectively accentuate real motion changes while suppressing trivial noisy changes. Consequently, this technique can be effectively applied to motion region-of-interest (ROI) based video coding. Our experiments show that the proposed algorithm can reliably extract motion regions and is less sensitive to thresholds than single-scale methods. Compared with a H.264/AVC encoder, the proposed semantic video encoder achieves a bitrate saving ratio of up to 34% at the similar video quality, besides an overall speedup factor of 2.6 to 3.6. The motion-ROI detection can process a 352 × 288 size video at 20 fps on an Intel Pentium 4 processor.
Jiangbo Lu, Gauthier Lafruit, Francky Catthoor
ICASSP (1)1
2007 Fast Variable Center-Biased Windowing for High-Speed Stereo on Programmable Graphics Hardware
abstract
We present a high-speed dense stereo algorithm that achieves both good quality results and very high disparity estimation throughput on the graphics processing unit (GPU). The key idea is a variable center-biased windowing approach, enabling an adaptive selection of the most suitable support patterns with varying sizes and shapes. As the fundamental construct for variable windows, a truncated separable Laplacian kernel approximation is proposed for the efficient pixel-wise weighted cost aggregation. We also present a number of critical optimization schemes to boost the real-time speed on GPUs. Our method outperforms previous GPU-based local stereo methods and even some methods using global optimization on the Middlebury stereo database. Our optimized implementation completely running on an Nvidia GeForce 7900 graphics card achieves over 605 million disparity estimations per second (Mde/s) including all the overhead, about 2.1 to 12.1 times faster than the existing GPU-based solutions.
Jiangbo Lu, Gauthier Lafruit, Francky Catthoor
ICIP (6)1
2007 Real-Time Stereo Correspondence using a Truncated Separable Laplacian Kernel Approximation on Graphics Hardware
abstract
We present a novel real-time stereo algorithm that achieves both good quality results and very high disparity estimation throughput on the graphics processing unit (GPU). As the key idea of this paper, a truncated separable approximation to an isotropic Laplacian kernel is proposed. This truncated 2D Laplacian kernel variant combines the advantages of large support windows and shiftable windows, while support-weights on geometric proximity can still be appropriately applied to each pixel in truncated support windows. Our method outperforms previous GPU-based local stereo methods and even some methods using global optimization on the benchmark Middlebury stereo database. Because of its separable and regular property, the proposed kernel can be very efficiently implemented on CPUs. Our optimized implementation completely running on an Nvidia GeForce 7900 graphics card achieves over 668 million disparity estimations per second (Mde/s) including all the overhead, about 2.3 to 13.4 times faster than the existing GPU-based solutions.
Jiangbo Lu, Sammy Rogmans, Gauthier Lafruit, Francky Catthoor
ICME1
2007 High-Speed Stream-Centric Dense Stereo and View Synthesis on Graphics Hardware
abstract
This paper presents an efficient image-based rendering system capable of performing online stereo matching and view synthesis at high speed, completely on the graphics processing unit (GPU). Given two rectified stereo images, our algorithm first extracts the disparity map with a stream-centric dense depth estimation approach. For high-quality view synthesis, multi-label masks are then automatically generated to postprocess occlusions and ambiguously estimated regions adaptively. To allow even faster interactive view generation, an alternative forward warping method is also integrated. The experiments show that photorealistic intermediate views of high image quality are yielded by our algorithm. The optimized implementation also provides the state-of-the-art stereo analysis and view synthesis speed, achieving over 47 fps with 450x375 stereo images and 60 disparity levels on an Nvidia GeForce 7900 graphics card.
Jiangbo Lu, Sammy Rogmans, Gauthier Lafruit, Francky Catthoor
MMSP1
2007 An Epipolar Geometry-Based Fast Disparity Estimation Algorithm for Multiview Image and Video Coding
abstract
Effectively coding multiview visual content is an indispensable research topic because multiview image and video that provide greatly enhanced viewing experiences often contain huge amounts of data. Generally, conventional hybrid predictive-coding methodologies are adopted to address the compression by exploiting the temporal and interviewpoint redundancy existing in a multiview image or video sequences. However, their key yet time-consuming component, motion estimation (ME), is usually not efficient in interviewpoint prediction or disparity estimation (DE), because interviewpoint disparity is completely different from temporal motion existing in the conventional video. Targeting a generic fast DE framework for interviewpoint prediction, we propose a novel DE technique in this paper to accelerate the disparity search by employing epipolar geometry. Theoretical analysis, optimal disparity vector distribution histograms, and experimental results show that the proposed epipolar geometry-based DE can greatly reduce search region and effectively track large and irregular disparity, which is typical in convergent multiview camera setups. Compared with the existing state-of-the-art fast ME approaches, our proposed DE can obtain a similar coding efficiency while achieving a significant speedup for interviewpoint prediction and coding. Moreover, a robustness study shows that the proposed DE algorithm is insensitive to the epipolar geometry estimation noise. Hence, its wide application for multiview image and video coding is promising
Jiangbo Lu, Hua Cai, Jianguang Lou, Jiang Li 0008
IEEE Trans. Circuits Syst. Video Technol.1
2006 An Effective Epipolar Geometry Assisted Motion Estimation Technique for Multi-View Image and Video Coding
abstract
To efficiently encode data-intensive multi-view imaging content, conventional hybrid predictive coding methodologies choose to address the compression by exploiting temporal and inter-viewpoint redundancy. However, their key yet time-consuming component, motion estimation (ME), is usually not efficient in inter-viewpoint prediction because inter-viewpoint motion is quite different from temporal motion. In essence, inter-viewpoint correlation is subject to epipolar geometry, which provides constraints for multi-view image sequences. A fast inter-viewpoint ME technique is hence proposed in this paper to accelerate the encoding by employing epipolar geometry. Theoretical analysis and experimental results prove that the proposed ME algorithm can greatly reduce search region and effectively track large and irregular motion that is typical for convergent multi-view camera setups. As a result, compared with fast full search at large search size adopted in H.264, our proposed ME algorithm can obtain a similar coding efficiency while achieving a speedup ratio of 2.9.
Jiangbo Lu, Hua Cai, Jian-Guang Lou, Jiang Li 0008
ICIP1
2003 Practical real-time video codec for mobile devices
abstract
Real-time software-based video codec is widely used on PCs with relatively strong computing capability. However, mobile devices, such as pocket PCs and handheld PCs, still suffer from weak computational power, short battery lifetime and limited display capability. We developed a practical low-complexity real-time video codec for mobile devices. Several methods that can significantly reduce the computational cost are adopted in this codec and described in this paper, including a predictive algorithm for motion estimation, the integer discrete cosine transform (IntDCT), and a DCT/quantizer bypass technique. A real-time video communication implementation of the proposed coded is also introduced. Experiments show that substantial computation reduction is achieved while the loss in video quality is negligible. The proposed codec is very suitable for scenarios where low-complexity computing is required.
Keman Yu, Jiangbo Lu, Jiang Li 0008, Shipeng Li 0001
ICME2