VLDB 2026 Research / reviewers in the wild / expert
Yuchao Dai
dblp:65/7804
· DBLP profile ↗
162ranked-venue papers
9as first author
101since 2021 · last 2026
0000-0002-4432-7406ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 115 · 7 first-author · 67 since 2021Artificial intelligence and machine learning · 111 · 6 first-author · 69 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Spatial Decay for Vision TransformersabstractVision Transformers (ViTs) have revolutionized computer vision, yet their self-attention mechanism lacks explicit spatial inductive biases, leading to suboptimal performance on spatially-structured tasks. Existing approaches introduce data-independent spatial decay based on fixed distance metrics, applying uniform attention weighting regardless of image content and limiting adaptability to diverse visual scenarios. Inspired by recent advances in large language models where content-aware gating mechanisms (e.g., GLA, HGRN2, FOX) significantly outperform static alternatives, we present the first successful adaptation of data-dependent spatial decay to 2D vision transformers. We introduce Spatial Decay Transformer (SDT), featuring a novel Context-Aware Gating (CAG) mechanism that generates dynamic, data-dependent decay for patch interactions. Our approach learns to modulate spatial attention based on both content relevance and spatial proximity. We address the fundamental challenge of 1D-to-2D adaptation through a unified spatial-content fusion framework that integrates manhattan distance-based spatial priors with learned content representations. Extensive experiments on ImageNet-1K classification and generation tasks demonstrate consistent improvements over strong baselines. Our work establishes data-dependent spatial decay as a new paradigm for enhancing spatial attention in vision transformers. Yuxin Mao, Zhen Qin 0003, Jinxing Zhou, Bin Fan 0002, Jing Zhang 0052, Yiran Zhong, Yuchao Dai |
AAAI | 7 |
| 2026 | EC-MVSNet: Enhanced Cascaded Multi-View Stereo with Cross-Scale Relevance IntegrationabstractCascade-based multi-scale architectures are currently the mainstream in Multi-view Stereo (MVS), achieving a balance between computational efficiency and reconstruction accuracy. However, existing cascade MVS methods suffer from significant limitations in cross-scale information utilization, where depth estimation processes operate independently across scales without fully exploiting the rich relevance between adjacent scales. To address this fundamental limitation, we propose an Enhanced Cascade Multi-View Stereo framework (EC-MVSNet), which introduces a novel cross-scale relevance integration strategy. Specifically, we introduce a Cross-Scale Feature-based Joint Construction (CFC) module to synergistically combine features from adjacent scales to build more reliable cost volumes. Additionally, a Cross-Scale Probability-guided Enhancement (CPE) module is proposed to propagate depth probability distributions across scales to guide cost volume enhancement. Furthermore, we propose a Monocular Feature-based Refinement (MFR) module to further enhance depth prediction accuracy by leveraging monocular priors. Extensive experiments demonstrate that EC-MVSNet achieves state-of-the-art performance on multiple benchmarks, validating the effectiveness of the cross-scale integration in improving MVS reconstruction quality. Shaoqian Wang, Jiadai Sun, Bin Fan 0002, Qiang Wang 0023, Yuchao Dai |
AAAI | 6 |
| 2026 | A Generative Victim Model for Segmentation
Aixuan Li, Jing Zhang 0052, Zhexiong Wan, Yiran Zhong, Yuchao Dai |
Int. J. Comput. Vis. | 6 |
| 2026 | Multi-event representation and multi-level fusion for robust RGB-event object tracking
Bin Fan 0002, Zhexiong Wan, Qi Liu 0054, Yuchao Dai |
Knowl. Based Syst. | 5 |
| 2026 | Exploiting Continuity for Unsupervised Single Depth Map Super-ResolutionabstractDepth map super-resolution (DSR) aims at reconstructing high-resolution depth maps from low-resolution input. Existing DSR methods rely on using the high-resolution RGB images as guidance and are typically trained in a supervised manner. However, obtaining well-aligned RGB-Depth pairs is challenging, and these approaches also suffer from texture over-transfer issues. To address these limitations, we propose the Implicit Depth Fitting Network (IDFN), a zero-shot, unsupervised framework that relies solely on a single low-resolution depth map. We formulate DSR as learning a resolution-independent continuous depth field. To accurately capture complex scene geometries, our framework combines Fourier feature encoding with periodic activation functions, effectively balancing smooth surface reconstruction with sharp edge preservation. Furthermore, we introduce a Dithering Strategy to model the sensor degradation process. This strategy enables the network to learn the underlying continuous signal from discrete area-integrated observations, effectively suppressing aliasing artifacts. Extensive experiments show that IDFN not only achieves performance comparable to state-of-the-art unsupervised approaches, but also yields results competitive with leading supervised approaches, demonstrating the effectiveness of our method. Ruobing Jian, Jing Zhang 0052, Yuchao Dai |
IEEE Signal Process. Lett. | 3 |
| 2026 | Boosting Few-Shot Hyperspectral Image Classification Through Dynamic Fusion and Hierarchical EnhancementabstractFew-shot learning has garnered increasing attention in hyperspectral image classification (HSIC) due to its potential to reduce dependency on labor-intensive and costly labeled data. However, most existing methods are constrained to feature extraction using a single image patch of fixed size, and typically neglect the pivotal role of the central pixel in feature fusion, leading to inefficient information utilization. In addition, the correlations among sample features have not been fully explored, thereby weakening feature expressiveness and hindering cross-domain knowledge transfer. To address these issues, we propose a novel few-shot HSIC framework incorporating dynamic fusion and hierarchical enhancement. Specifically, we first introduce a robust feature extraction module, which effectively combines the content concentration of small patches with the noise robustness of large patches, and further captures local spatial correlations through a central-pixel-guided dynamic pooling strategy. Such patch-to-pixel dynamic fusion enables a more comprehensive and robust extraction of ground object information. Then, we develop a support-query hierarchical enhancement module that integrates intraclass self-attention and interclass cross-attention mechanisms. This process not only enhances support-level and query-level feature representation but also facilitates the learning of more informative prior knowledge from the abundantly labeled source domain. Moreover, to further increase feature discriminability, we design an intraclass consistency loss and an interclass orthogonality loss, which collaboratively encourage intraclass samples to be closer together and interclass samples to be more separable in the metric space. Experimental results on four benchmark datasets demonstrate that our method substantially improves classification accuracy and consistently outperforms competing approaches. Code is available at https://github.com/guoying918/DFHE2025. Ying Guo 0014, Bin Fan 0002, Yuchao Dai, Yan Feng 0005, Mingyi He |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Deep Non-Rigid Structure-from-Motion Revisited: Canonicalization and Sequence ModelingabstractNon-Rigid Structure-from-Motion (NRSfM) is a classic 3D vision problem, where a 2D sequence is taken as input to estimate the corresponding 3D sequence. Recently, the deep neural networks have greatly advanced the task of NRSfM. However, existing deep NRSfM methods still have limitations in handling the inherent sequence property and motion ambiguity associated with the NRSfM problem. In this paper, we revisit deep NRSfM from two perspectives to address the limitations of current deep NRSfM methods : (1) canonicalization and (2) sequence modeling. We propose an easy-to-implement per-sequence canonicalization method as opposed to the previous per-dataset canonicalization approaches. With this in mind, we propose a sequence modeling method that combines temporal information and subspace constraint. As a result, we have achieved a more optimal NRSfM reconstruction pipeline compared to previous efforts. The effectiveness of our method is verified by testing the sequence-to-sequence deep NRSfM pipeline with corresponding regularization modules on several commonly used datasets. Zhen Qin 0003, Yiran Zhong, Yuchao Dai |
AAAI | 5 |
| 2025 | Geometry-Aware 3D Salient Object Detection NetworkabstractPoint cloud salient object detection has attracted the attention of researchers in recent years. Since existing works do not fully utilize the geometry context of 3D objects, blurry boundaries are generated when segmenting objects with complex backgrounds. In this paper, we propose a geometry-aware 3D salient object detection network that explicitly clusters points into superpoints to enhance the geometric boundaries of objects, thereby segmenting complete objects with clear boundaries. Specifically, we first propose a simple yet effective superpoint partition module to cluster points into superpoints. In order to improve the quality of superpoints, we present a point cloud class-agnostic loss to learn discriminative point features for clustering superpoints from the object. After obtaining superpoints, we then propose a geometry enhancement module that utilizes superpoint-point attention to aggregate geometric information into point features for predicting the salient map of the object with clear boundaries. Extensive experiments show that our method achieves new state-of-the-art performance on the PCSOD dataset. Chen Wang 0049, Le Hui, Qi Liu 0054, Yuchao Dai |
AAAI | 5 |
| 2025 | MixRI: Mixing Features of Reference Images for Novel Object Pose Estimation
Xinhang Liu, Zheng Dang, Yuchao Dai |
ICCV | 4 |
| 2025 | PanoSplatt3R: Leveraging Perspective Pretraining for Generalized Unposed Wide-Baseline Panorama Reconstruction
Jiahui Ren, Mochu Xiang, Yuchao Dai |
ICCV | 4 |
| 2025 | Event-Aided Dense and Continuous Point Tracking: Everywhere and Anytime
Zhexiong Wan, Jianqin Luo, Yuchao Dai, Gim Hee Lee |
ICCV | 3 |
| 2025 | MRM-RETrack: Hybrid Multi-scale Residual and Mamba for RGB-Event Tracking
Bin Fan 0002, Zhexiong Wan, Zhiyuan Zhang 0002, Yuchao Dai |
PRCV (18) | 5 |
| 2025 | Instance-Level Moving Object Segmentation from a Single Image with Events
Zhexiong Wan, Bin Fan 0002, Le Hui, Yuchao Dai, Gim Hee Lee |
Int. J. Comput. Vis. | 4 |
| 2025 | Constraining multimodal distribution for domain adaptation in stereo matching
Zhelun Shen, Chenming Wu, Zhibo Rao, Lina Liu 0010, Yuchao Dai, Liangjun Zhang |
Pattern Recognit. | 6 |
| 2025 | Self-Supervised Learning for Rolling Shutter Temporal Super-ResolutionabstractMost cameras on portable devices adopt a rolling shutter (RS) mechanism, encoding sufficient temporal dynamic information through sequential readouts. This advantage can be exploited to recover a temporal sequence of latent global shutter (GS) images. Existing methods rely on fully supervised learning, necessitating specialized optical devices to collect paired RS-GS images as ground-truth, which is too costly to scale. In this paper, we propose a self-supervised learning framework for the first time to produce a high frame rate GS video from two consecutive RS images, unleashing the potential of RS cameras. Specifically, we first develop the unified warping model of RS2GS and GS2RS, enabling the complement conversions of RS2GS and GS2RS to be incorporated into a uniform network model. Then, based on the cycle consistency constraint, given a triplet of consecutive RS frames, we minimize the discrepancy between the input middle RS frame and its cycle reconstruction, generated by interpolating back from the predicted two intermediate GS frames. Experiments on various benchmarks show that our approach achieves comparable or better performance than state-of-the-art supervised methods while enjoying stronger generalization capabilities. Moreover, our approach makes it possible to recover smooth and distortion-free videos from two adjacent RS frames in the real-world BS-RSC dataset, surpassing prior limitations. Bin Fan 0002, Ying Guo 0014, Yuchao Dai, Chao Xu 0006, Boxin Shi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Generative Transformer for Accurate and Reliable Salient Object DetectionabstractWe explore the impact of transformers on accurate and reliable salient object detection. For accuracy, we integrate the transformer with a deterministic model and delineate its advantages in structural modeling. Regarding reliability, we address the transformer’s tendency to produce overly confident, incorrect predictions. To gauge reliability implicitly, we introduce a latent variable model within the transformer framework, termed the inferential generative adversarial network (iGAN). The stochastic nature of the latent variable facilitates the estimation of predictive uncertainty, which serves as an auxiliary measure of the model’s prediction reliability. Different from the conventional GAN, which defines the distribution of the latent variable as fixed standard normal distribution$\mathcal {N}(0,\mathbf {I})$. The proposed iGAN infers the latent variable by gradient-based Markov Chain Monte Carlo (MCMC), namely Langevin dynamics, leading to an input-dependent latent variable model. We apply our proposed iGAN to fully supervised salient object detection, explaining that iGAN within the transformer framework leads to both accurate and reliable salient object detection. The source code and experimental results are publicly available via our project page:https://npucvr.github.io/TransformerSOD. Yuxin Mao, Jing Zhang 0052, Zhexiong Wan, Aixuan Li, Yunqiu Lv, Yuchao Dai |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Contrastive Conditional Latent Diffusion for Audio-Visual SegmentationabstractAudio-visual Segmentation (AVS) is conceptualized as a conditional generation task, where audio is considered as the conditional variable for segmenting the sound producer(s). In this case, audio should be extensively explored to maximize its contribution for the final segmentation task. We propose a contrastive conditional latent diffusion model for audio-visual segmentation (AVS) to thoroughly investigate the impact of audio, where the correlation between audio and the final segmentation map is modeled to guarantee the strong correlation between them. To achieve semantic-correlated representation learning, our framework incorporates a latent diffusion model. The diffusion model learns the conditional generation process of the ground-truth segmentation map, resulting in ground-truth aware inference during the denoising process at the test stage. As our model is conditional, it is vital to ensure that the conditional variable contributes to the model output. We thus extensively model the contribution of the audio signal by minimizing the density ratio between the conditional probability of the multimodal data, e.g. conditioned on the audio-visual data, and that of the unimodal data, e.g. conditioned on the audio data only. In this way, our latent diffusion model via density ratio optimization explicitly maximizes the contribution of audio for AVS, which can then be achieved with contrastive learning as a constraint, where the diffusion part serves as the main objective to achieve maximum likelihood estimation, and the density ratio optimization part imposes the constraint. By adopting this latent diffusion model via contrastive learning, we effectively enhance the contribution of audio for AVS. The effectiveness of our solution is validated through experimental results on the benchmark dataset. Code and results are online via our project page: https://github.com/OpenNLPLab/DiffusionAVS. Yuxin Mao, Jing Zhang 0052, Mochu Xiang, Yunqiu Lv, Dong Li 0033, Yiran Zhong, Yuchao Dai |
IEEE Trans. Image Process. | 7 |
| 2024 | Improving Audio-Visual Segmentation with Bidirectional GenerationabstractThe aim of audio-visual segmentation (AVS) is to precisely differentiate audible objects within videos down to the pixel level. Traditional approaches often tackle this challenge by combining information from various modalities, where the contribution of each modality is implicitly or explicitly modeled. Nevertheless, the interconnections between different modalities tend to be overlooked in audio-visual modeling. In this paper, inspired by the human ability to mentally simulate the sound of an object and its visual appearance, we introduce a bidirectional generation framework. This framework establishes robust correlations between an object's visual characteristics and its associated sound, thereby enhancing the performance of AVS. To achieve this, we employ a visual-to-audio projection component that reconstructs audio features from object segmentation masks and minimizes reconstruction errors. Moreover, recognizing that many sounds are linked to object movements, we introduce an implicit volumetric motion estimation module to handle temporal dynamics that may be challenging to capture using conventional optical flow methods. To showcase the effectiveness of our approach, we conduct comprehensive experiments and analyses on the widely recognized AVSBench benchmark. As a result, we establish a new state-of-the-art performance level in the AVS benchmark, particularly excelling in the challenging MS3 subset which involves segmenting multiple sound sources. Code is released in: https://github.com/OpenNLPLab/AVS-bidirectional. Dawei Hao, Yuxin Mao, Xiaodong Han, Yuchao Dai, Yiran Zhong |
AAAI | 5 |
| 2024 | Video Frame Prediction from a Single Image and EventsabstractRecently, the task of Video Frame Prediction (VFP), which predicts future video frames from previous ones through extrapolation, has made remarkable progress. However, the performance of existing VFP methods is still far from satisfactory due to the fixed framerate video used: 1) they have difficulties in handling complex dynamic scenes; 2) they cannot predict future frames with flexible prediction time intervals. The event cameras can record the intensity changes asynchronously with a very high temporal resolution, which provides rich dynamic information about the observed scenes. In this paper, we propose to predict video frames from a single image and the following events, which can not only handle complex dynamic scenes but also predict future frames with flexible prediction time intervals. First, we introduce a symmetrical cross-modal attention augmentation module to enhance the complementary information between images and events. Second, we propose to jointly achieve optical flow estimation and frame generation by combining the motion information of events and the semantic information of the image, then inpainting the holes produced by forward warping to obtain an ideal prediction frame. Based on these, we propose a lightweight pyramidal coarse-to-fine model that can predict a 720P frame within 25 ms. Extensive experiments show that our proposed model significantly outperforms the state-of-the-art frame-based and event-based VFP methods and has the fastest runtime. Code is available at https://npucvr.github.io/VFPSIE/. Juanjuan Zhu, Zhexiong Wan, Yuchao Dai |
AAAI | 3 |
| 2024 | 3D Geometry-aware Deformable Gaussian Splatting for Dynamic View SynthesisabstractIn this paper, we propose a 3D geometry-aware deformable Gaussian Splatting method for dynamic view synthesis. Existing neural radiance fields (NeRF) based solutions learn the deformation in an implicit manner, which cannot incorporate 3D scene geometry. Therefore, the learned deformation is not necessarily geometrically coherent, which results in unsatisfactory dynamic view synthesis and 3D dynamic reconstruction. Recently, 3D Gaussian Splatting provides a new representation of the 3D scene, building upon which the 3D geometry could be exploited in learning the complex 3D deformation. Specifically, the scenes are represented as a collection of 3D Gaussian, where each 3D Gaussian is optimized to move and rotate over time to model the deformation. To enforce the 3D scene geometry constraint during deformation, we explicitly extract 3D geometry features and integrate them in learning the 3D deformation. In this way, our solution achieves 3D geometry-aware deformation modeling, which enables improved dynamic view synthesis and 3D dynamic reconstruction. Extensive experimental results on both synthetic and real datasets prove the superiority of our solution, which achieves new state-of-the-art performance. The project is available at https://npucvr.github.io/GaGS/. Zhicheng Lu, Le Hui, Yuchao Dai |
CVPR | 8 |
| 2024 | Non-rigid Structure-from-Motion: Temporally-smooth Procrustean Alignment and Spatially-variant Deformation ModelingabstractEven though Non-rigid Structure-from-Motion (NRSfM) has been extensively studied and great progress has been made, there are still key challenges that hinder their broad real-world applications: 1) the inherent motion/rotation ambiguity requires either explicit camera motion recovery with extra constraint or complex Procrustean Alignment; 2) existing low-rank modeling of the global shape can over-penalize drastic deformations in the 3D shape sequence. This paper proposes to resolve the above issues from a spatial-temporal modeling perspective. First, we propose a novel Temporally-smooth Procrustean Alignment module that estimates 3D deforming shapes and adjusts the camera motion by aligning the 3D shape sequence consecutively. Our new alignment module remedies the requirement of complex reference 3D shape during alignment, which is more conductive to non-isotropic deformation modeling. Second, we propose a spatial-weighted approach to enforce the low-rank constraint adaptively at different locations to accommodate drastic spatially-variant deformation reconstruction better. Our modeling outperform existing low-rank based methods, and extensive experiments across different datasets validate the effectiveness of our method1.1Project page: https://npucvr.github.io/TSM-NRSfM. Yuchao Dai |
CVPR | 3 |
| 2024 | PaReNeRF: Toward Fast Large-Scale Dynamic NeRF with Patch-Based ReferenceabstractWith photo-realistic image generation, Neural Radiance Field (NeRF) is widely used for large-scale dynamic scene reconstruction as autonomous driving simulator. However, large-scale scene reconstruction still suffers from extremely long training time and rendering time. Low-resolution (L-R) rendering combined with upsampling can alleviate this problem but it degrades image quality. In this paper, we design a lightweight reference decoder which exploits prior information from known views to improve image reconstruction quality of new views. In addition, to speed up prior information search, we propose an optical flow and structural similarity based prior information search method. Results on KITTI and VKITTI2 datasets show that our method significantly outperforms the baseline method in terms of training speed, rendering speed and rendering quality. Penghui Sun, Yuchao Dai, Hojae Lee |
CVPR | 5 |
| 2024 | Improving Depth Completion via Depth Feature UpsamplingabstractThe encoder-decoder network (ED-Net) is a commonly employed choice for existing depth completion methods, but its working mechanism is ambiguous. In this paper, we vi-sualize the internal feature maps to analyze how the net-work densifies the input sparse depth. We find that the en-coder feature of ED-Net focus on the areas with input depth points around. To obtain a dense feature and thus esti-mate complete depth, the decoder feature tends to comple-ment and enhance the encoder feature by skip-connection to make the fused encoder-decoder feature dense, resulting in the decoder feature also exhibits sparse. However, ED-Net obtains the sparse decoder feature from the dense fused feature at the previous stage, where the “dense-i-sparse‘’ process destroys the completeness of features and loses in-formation. To address this issue, we present a depth feature upsampling network (DFU) that explicitly utilizes these dense features to guide the upsampling of a low-resolution (LR) depth feature to a high-resolution (HR) one. The completeness of features is maintained throughout the up-sampling process, thus avoiding information loss. Fur-thermore, we propose a confidence-aware guidance module (CGM), which is confidence-aware and performs guidance with adaptive receptive fields (GARF), to fully exploit the potential of these dense features as guidance. Experimental results show that our DFU, a plug-and-play module, can significantly improve the performance of existing ED-Net based methods with limited computational overheads, and new SOTA results are achieved. Besides, the generalization capability on sparser depth is also enhanced. Project page: https://npucvr.github.iolDFU. Ge Zhang 0006, Shaoqian Wang, Bo Li 0090, Qi Liu 0054, Le Hui, Yuchao Dai |
CVPR | 7 |
| 2024 | TAVGBench: Benchmarking Text to Audible-Video GenerationabstractThe Text to Audible-Video Generation (TAVG) task involves generating videos with accompanying audio based on text descriptions. Achieving this requires skillful alignment of both audio and video elements. To support research in this field, we have developed a comprehensive Text to Audible-Video Generation Benchmark (TAVGBench), which contains over 1.7 million clips with a total duration of 11.8 thousand hours. We propose an automatic annotation pipeline to ensure each audible video has detailed descriptions for both its audio and video contents. We also introduce the Audio-Visual Harmoni score (AVHScore) to provide a quantitative measure of the alignment between the generated audio and video modalities. Additionally, we present a baseline model for TAVG called TAVDiffusion, which uses a two-stream latent diffusion model to provide a fundamental starting point for further research in this area. We achieve the alignment of audio and video by employing cross-attention and contrastive learning. Through extensive experiments and evaluations on TAVGBench, we demonstrate the effectiveness of our proposed model under both conventional metrics and our proposed metrics. The dataset and code can be found on this page https://npucvr.github.io/TAVGBench/ and on github https://github.com/OpenNLPLab/TAVGBench. Yuxin Mao, Xuyang Shen, Jing Zhang 0052, Zhen Qin 0003, Jinxing Zhou, Mochu Xiang, Yiran Zhong, Yuchao Dai |
ACM Multimedia | 8 |
| 2024 | Spatio-Temporal Interactive Learning for Efficient Image Reconstruction of Spiking CamerasabstractThe spiking camera is an emerging neuromorphic vision sensor that records high-speed motion scenes by asynchronously firing continuous binary spike streams. Prevailing image reconstruction methods, generating intermediate frames from these spike streams, often rely on complex step-by-step network architectures that overlook the intrinsic collaboration of spatio-temporal complementary information. In this paper, we propose an efficient spatio-temporal interactive reconstruction network to jointly perform inter-frame feature alignment and intra-frame feature filtering in a coarse-to-fine manner. Specifically, it starts by extracting hierarchical features from a concise hybrid spike representation, then refines the motion fields and target frames scale-by-scale, ultimately obtaining a full-resolution output. Meanwhile, we introduce a symmetric interactive attention block and a multi-motion field estimation block to further enhance the interaction capability of the overall network. Experiments on synthetic and real-captured data show that our approach exhibits excellent performance while maintaining low model complexity. Bin Fan 0002, Jiaoyang Yin, Yuchao Dai, Chao Xu 0006, Tiejun Huang 0001, Boxin Shi |
NeurIPS | 3 |
| 2024 | 3D Focusing-and-Matching Network for Multi-Instance Point Cloud RegistrationabstractMulti-instance point cloud registration aims to estimate the pose of all instances of a model point cloud in the whole scene. Existing methods all adopt the strategy of first obtaining the global correspondence and then clustering to obtain the pose of each instance. However, due to the cluttered and occluded objects in the scene, it is difficult to obtain an accurate correspondence between the model point cloud and all instances in the scene. To this end, we propose a simple yet powerful 3D focusing-and-matching network for multi-instance point cloud registration by learning the multiple pair-wise point cloud registration. Specifically, we first present a 3D multi-object focusing module to locate the center of each object and generate object proposals. By using self-attention and cross-attention to associate the model point cloud with structurally similar objects, we can locate potential matching instances by regressing object centers. Then, we propose a 3D dual-masking instance matching module to estimate the pose between the model point cloud and each object proposal. It performs instance mask and overlap mask masks to accurately predict the pair-wise correspondence. Extensive experiments on two public benchmarks, Scan2CAD and ROBI, show that our method achieves a new state-of-the-art performance on the multi-instance point cloud registration task. Le Hui, Qi Liu 0054, Bo Li 0090, Yuchao Dai |
NeurIPS | 5 |
| 2024 | Unsupervised 3D Pose Estimation with Non-Rigid Structure-from-Motion ModelingabstractMost existing 3D human pose estimation work rely heavily on the powerful memory capability of networks to obtain suitable 2D-3D mappings from the training data. Few works have studied the modeling of human posture deformation in motion. In this paper, we propose a new modeling method for human pose deformations and design an accompanying diffusion-based motion prior. Inspired by the field of non-rigid structure-from-motion, we divide the task of reconstructing 3D human skeletons in motion into the estimation of a 3D reference skeleton, and a frame-by-frame skeleton deformation. A mixed spatial-temporal NRSfMformer is used to simultaneously estimate the 3D reference skeleton and the skeleton deformation of each frame from 2D observations sequence, and then sum them up to obtain the pose of each frame. Subsequently, a loss term based on the diffusion model is used to ensure that the pipeline learns the correct prior motion knowledge. Finally, we have evaluated our proposed method on mainstream datasets and obtained superior results outperforming the state-of-the-art. Haorui Ji, Yuchao Dai, Hongdong Li |
WACV | 3 |
| 2024 | Towards a Unified Network for Robust Monocular Depth Estimation: Network Architecture, Training Strategy and Dataset
Mochu Xiang, Yuchao Dai, Zhensong Zhang |
Int. J. Comput. Vis. | 2 |
| 2024 | Deep Non-Rigid Structure-From-Motion: A Sequence-to-Sequence Translation PerspectiveabstractDirectly regressing the non-rigid shape and camera pose from the individual 2D frame is ill-suited to the Non-Rigid Structure-from-Motion (NRSfM) problem. This frame-by-frame 3D reconstruction pipeline overlooks the inherent spatial-temporal nature of NRSfM, i.e., reconstructing the 3D sequence from the input 2D sequence. In this paper, we propose to solve deep sparse NRSfM from a sequence-to-sequence translation perspective, where the input 2D keypoints sequence is taken as a whole to reconstruct the corresponding 3D keypoints sequence in a self-supervised manner. First, we apply a shape-motion predictor on the input sequence to obtain an initial sequence of shapes and corresponding motions. Then, we propose the Context Layer, which enables the deep learning framework to effectively impose overall constraints on sequences based on the structural characteristics of non-rigid sequences. The Context Layer constructs modules for imposing the self-expressiveness regularity on non-rigid sequences with multi-head attention (MHA) as the core, together with the use of temporal encoding, both of which act simultaneously to constitute constraints on non-rigid sequences in the deep framework. Experimental results across different datasets such as Human3.6M, CMU Mocap, and InterHand prove the superiority of our framework. The code will be made publicly available. Tong Zhang 0023, Yuchao Dai, Yiran Zhong, Hongdong Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Learning Bilateral Cost Volume for Rolling Shutter Temporal Super-ResolutionabstractRolling shutter temporal super-resolution (RSSR), which aims to synthesize intermediate global shutter (GS) video frames between two consecutive rolling shutter (RS) frames, has made remarkable progress with the development of deep convolutional neural networks over the past years. Existing methods cascade multiple separated networks to sequentially estimate intermediate motion fields and synthesize target GS frames. Nevertheless, they are typically complex, do not facilitate the interaction of complementary motion and appearance information, and suffer from problems such as pixel aliasing or poor interpretation. In this paper, we derive the uniform bilateral motion fields for RS-aware backward warping, which endows our network a more explicit geometric meaning by injecting spatio-temporal consistency information through time-offset embedding. More importantly, we develop a unified, single-stage RSSR pipeline to recover the latent GS video in a coarse-to-fine manner. It first extracts pyramid features from given inputs, and then refines the bilateral motion fields together with the anchor frame until generating the desired output. With the help of our proposed bilateral cost volume, which uses the anchor frame as a common reference to model the correlation with two RS frames, the gradually refined anchor frames not only facilitate intermediate motion estimation, but also compensate for contextual details, making additional frame synthesis or refinement networks unnecessary. Meanwhile, an asymmetric bilateral motion model built on top of the symmetric bilateral motion model further improves the generality and adaptability, yielding better GS video reconstruction performance. Extensive quantitative and qualitative experiments on synthetic and real data demonstrate that our method achieves new state-of-the-art results. Bin Fan 0002, Yuchao Dai, Hongdong Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | AGDF-Net: Learning Domain Generalizable Depth Features With Adaptive Guidance FusionabstractCross-domain generalizable depth estimation aims to estimate the depth of target domains (i.e., real-world) using models trained on the source domains (i.e., synthetic). Previous methods mainly use additional real-world domain datasets to extract depth specific information for cross-domain generalizable depth estimation. Unfortunately, due to the large domain gap, adequate depth specific information is hard to obtain and interference is difficult to remove, which limits the performance. To relieve these problems, we propose a domain generalizable feature extraction network with adaptive guidance fusion (AGDF-Net) to fully acquire essential features for depth estimation at multi-scale feature levels. Specifically, our AGDF-Net first separates the image into initial depth and weak-related depth components with reconstruction and contrary losses. Subsequently, an adaptive guidance fusion module is designed to sufficiently intensify the initial depth features for domain generalizable intensified depth features acquisition. Finally, taking intensified depth features as input, an arbitrary depth estimation network can be used for real-world depth estimation. Using only synthetic datasets, our AGDF-Net can be applied to various real-world datasets (i.e., KITTI, NYUDv2, NuScenes, DrivingStereo and CityScapes) with state-of-the-art performances. Furthermore, experiments with a small amount of real-world data in a semi-supervised setting also demonstrate the superiority of AGDF-Net over state-of-the-art approaches. Lina Liu 0010, Xibin Song, Mengmeng Wang 0005, Yuchao Dai, Yong Liu 0007, Liangjun Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Mutual Information Regularization for Weakly-Supervised RGB-D Salient Object DetectionabstractIn this paper, we present a weakly-supervised RGB-D salient object detection model via scribble supervision. Specifically, as a multimodal learning task, we focus on effective multimodal representation learning via inter-modal mutual information regularization. In particular, following the principle of disentangled representation learning, we introduce a mutual information upper bound with a mutual information minimization regularizer to encourage the disentangled representation of each modality for salient object detection. Based on our multimodal representation learning framework, we introduce an asymmetric feature extractor for our multimodal data, which is proven more effective than the conventional symmetric backbone setting. We also introduce multimodal variational auto-encoder as stochastic prediction refinement techniques, which takes pseudo labels from the first training stage as supervision and generates refined prediction. Experimental results on benchmark RGB-D salient object detection datasets verify both effectiveness of our explicit multimodal disentangled representation learning method and the stochastic prediction refinement strategy, achieving comparable performance with the state-of-the-art fully supervised models. Our code and data are available at:https://npucvr.github.io/MIRV/. Aixuan Li, Yuxin Mao, Jing Zhang 0052, Yuchao Dai |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Efficient Multi-View Stereo by Dynamic Cost Volume and Cross-Scale PropagationabstractCurrently, learning-based multi-view stereo (MVS) has been dominated by the pipeline of 3D cost volume and regularization network over thestatic cost volumefor depth regression. However, this methodology is plagued by heavy time and memory consumption, which greatly hinders the applications of these methods for real-world high-resolution images. To address these challenges, we present Effi-MVS+, an efficient multi-scaledynamic cost volumebased MVS method. Firstly, instead of constructing a static cost volume and predicting a probability distribution map for depth regression, we update the depth map by iteratively predicting depth residuals. In each iteration, we construct a lightweight dynamic cost volume by encoding local matching and regularization information. The dynamic cost volume is subsequently processed using a 2D convolution-based GRU, which owns significant advantages in computational complexity and efficiency. Secondly, we propose a cross-scale propagation mechanism to enhance the multi-scale dynamic cost volume. This mechanism facilitates the progressive aggregation of multi-scale information, thereby providing enhanced matching and regularization information. Thirdly, to further improve the efficiency, we provide a reliable initial depth map to launch the framework and guarantee fast convergence. Extensive experiments on the DTU and Tanks & Temples benchmarks demonstrate the superiority of our method, which outperforms other state-of-the-art methods by a large margin in terms ofreconstruction quality, speed, and memory usage. Code will be released at https://github.com/npucvr/Effi-MVS-plus. Shaoqian Wang, Bo Li 0090, Yuchao Dai |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Decomposed Guided Dynamic Filters for Efficient RGB-Guided Depth CompletionabstractRGB-guided depth completion aims at predicting dense depth maps from sparse depth measurements and corresponding RGB images, where how to effectively and efficiently exploit the multi-modal information is a key issue. Guided dynamic filters, which generate spatially-variant depth-wise separable convolutional filters from RGB features to guide depth features, have been proven to be effective in this task. However, the dynamically generated filters require massive model parameters, computational costs and memory footprints when the number of feature channels is large. In this paper, we propose to decompose the guided dynamic filters into a spatially-shared component multiplied by content-adaptive adaptors at each spatial location. Based on the proposed idea, we introduce two decomposition schemes$\mathcal {A}$and$\mathcal {B}$, which decompose the filters by splitting the filter structure and using spatial-wise attention, respectively. The decomposed filters not only maintain the favorable properties of guided dynamic filters as being content-dependent and spatially-variant, but also reduce model parameters and hardware costs, as the learned adaptors are decoupled with the number of feature channels. Extensive experimental results demonstrate that the methods using our schemes outperform state-of-the-art methods on the KITTI dataset, and rank 1st and 2nd on the KITTI benchmark at the time of submission. Meanwhile, they also achieve comparable performance on the NYUv2 dataset. In addition, our proposed methods are general and could be employed as plug-and-play feature fusion blocks in other multi-modal fusion tasks such as RGB-D salient object detection. Yuxin Mao, Qi Liu 0054, Yuchao Dai |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Measuring and Modeling Uncertainty Degree for Monocular Depth EstimationabstractEffectively measuring and modeling the reliability of a trained model is essential to the real-world deployment of monocular depth estimation (MDE) models. However, the intrinsic ill-posedness and ordinal-sensitive nature of MDE pose major challenges to the estimation of uncertainty degree of the trained models. On the one hand, utilizing current uncertainty modeling methods may increase memory consumption and usually take more time. On the other hand, measuring the uncertainty based on model accuracy can also be problematic, where uncertainty reliability and prediction accuracy are not well decoupled. In this paper, we propose to model the uncertainty of MDE models from the perspective of the inherent probability distributions originating from the depth probability volume and its extensions, and to assess it more fairly with more comprehensive metrics. By simply introducing additional training regularization terms, our model, with surprisingly simple formations and without requiring extra modules or multiple inferences, can provide uncertainty estimations with state-of-the-art reliability, and can be further improved when combined with ensemble or sampling methods. A series of experiments demonstrate the effectiveness of our methods. Code and results are available at https://github.com/npucvr/MDEUncertainty. Mochu Xiang, Jing Zhang 0052, Nick Barnes, Yuchao Dai |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Unified Video Reconstruction for Rolling Shutter and Global Shutter CamerasabstractCurrently, the general domain of video reconstruction (VR) is fragmented into different shutters spanning global shutter and rolling shutter cameras. Despite rapid progress in the state-of-the-art, existing methods overwhelmingly follow shutter-specific paradigms and cannot conceptually generalize to other shutter types, hindering the uniformity of VR models. In this paper, we propose UniVR, a versatile framework to handle various shutters through unified modeling and shared parameters. Specifically, UniVR encodes diverse shutter types into a unified space via a tractable shutter adapter, which is parameter-free and thus can be seamlessly delivered to current well-established VR architectures for cross-shutter transfer. To demonstrate its effectiveness, we conceptualize UniVR as three shutter-generic VR methods, namely Uni-SoftSplat, Uni-SuperSloMo, and Uni-RIFE. Extensive experimental results demonstrate that the pre-trained model without any fine-tuning can achieve reasonable performance even on novel shutters. After fine-tuning, new state-of-the-art performances are established that go beyond shutter-specific methods and enjoy strong generalization. The code is available at https://github.com/GitCVfb/UniVR. Bin Fan 0002, Zhexiong Wan, Boxin Shi, Chao Xu 0006, Yuchao Dai |
IEEE Trans. Image Process. | 5 |
| 2024 | Weakly-Supervised Contrastive Learning for Unsupervised Object DiscoveryabstractUnsupervised object discovery (UOD) refers to the task of discriminating the whole region of objects from the background within a scene without relying on labeled datasets, which benefits the task of bounding-box-level localization and pixel-level segmentation. This task is promising due to its ability to discover objects in a generic manner. We roughly categorize existing techniques into two main directions, namely the generative solutions based on image resynthesis, and the clustering methods based on self-supervised models. We have observed that the former heavily relies on the quality of image reconstruction, while the latter shows limitations in effectively modeling semantic correlations. To directly target at object discovery, we focus on the latter approach and propose a novel solution by incorporating weakly-supervised contrastive learning (WCL) to enhance semantic information exploration. We design a semantic-guided self-supervised learning model to extract high-level semantic features from images, which is achieved by fine-tuning the feature encoder of a self-supervised model, namely DINO, via WCL. Subsequently, we introduce Principal Component Analysis (PCA) to localize object regions. The principal projection direction, corresponding to the maximal eigenvalue, serves as an indicator of the object region(s). Extensive experiments on benchmark unsupervised object discovery datasets demonstrate the effectiveness of our proposed solution. The source code and experimental results are publicly available via our project page at https://github.com/npucvr/WSCUOD.git. Yunqiu Lv, Jing Zhang 0052, Nick Barnes, Yuchao Dai |
IEEE Trans. Image Process. | 4 |
| 2023 | Joint Appearance and Motion Learning for Efficient Rolling Shutter CorrectionabstractRolling shutter correction (RSC) is becoming increasingly popular for RS cameras that are widely used in commercial and industrial applications. Despite the promising performance, existing RSC methods typically employ a two-stage network structure that ignores intrinsic infor-mation interactions and hinders fast inference. In this pa-per, we propose a single-stage encoder-decoder-based network, named JAMNet, for efficient RSC. It first extracts pyramid features from consecutive RS inputs, and then simultaneously refines the two complementary information (i.e., global shutter appearance and undistortion motion field) to achieve mutual promotion in a joint learning de-coder. To inject sufficient motion cues for guiding joint learning, we introduce a transformer-based motion embed-ding module and propose to pass hidden states across pyra-mid levels. Moreover, we present a new data augmentation strategy “vertical flip + inverse order” to release the potential of the RSC datasets. Experiments on various benchmarks show that our approach surpasses the state-of-the-art methods by a large margin, especially with a 4.7 dB PSNR leap on real-world RSC. Code is available at https://github.com/GitCVfb/JAMNet. Bin Fan 0002, Yuxin Mao, Yuchao Dai, Zhexiong Wan, Qi Liu 0054 |
CVPR | 3 |
| 2023 | Masked Representation Learning for Domain Generalized Stereo MatchingabstractRecently, many deep stereo matching methods have begun to focus on cross-domain performance, achieving impressive achievements. However, these methods did not deal with the significant volatility of generalization performance among different training epochs. Inspired by masked representation learning and multi-task learning, this paper designs a simple and effective masked representation for domain generalized stereo matching. First, we feed the masked left and complete right images as input into the models. Then, we add a lightweight and simple decoder following the feature extraction module to recover the original left image. Finally, we train the models with two tasks (stereo matching and image reconstruction) as a pseudo-multi-task learning framework, promoting models to learn structure information and to improve generalization performance. We implement our method on two well-known architectures (CFNet and LacGwcNet) to demonstrate its effectiveness. Experimental results on multi-datasets show that: (1) our method can be easily plugged into the current various stereo matching models to improve generalization performance; (2) our method can reduce the significant volatility of generalization performance among different training epochs; (3) we find that the current methods prefer to choose the best results among different training epochs as generalization performance, but it is impossible to select the best performance by ground truth in practice. Zhibo Rao, Mingyi He, Yuchao Dai, Zhelun Shen, Xing Li 0040 |
CVPR | 4 |
| 2023 | Fine-grained Audible Video DescriptionabstractWe explore a new task for audio-visual-language modeling called fine-grained audible video description (FAVD). It aims to provide detailed textual descriptions for the given audible videos, including the appearance and spatial locations of each object, the actions of moving objects, and the sounds in videos. Existing visual-language modeling tasks often concentrate on visual cues in videos while undervaluing the language and audio modalities. On the other hand, FAVD requires not only audio-visual-language modeling skills but also paragraph-level language generation abilities. We construct the first fine-grained audible video description benchmark (FAVDBench) to facilitate this research. For each video clip, we first provide a one-sentence summary of the video, i.e., the caption, followed by 4–6 sentences describing the visual details and 1–2 audio-related descriptions at the end. The descriptions are provided in both English and Chinese. We create two new metrics for this task: an EntityScore to gauge the completeness of entities in the visual descriptions, and an AudioScore to assess the audio descriptions. As a preliminary approach to this task, we propose an audio-visual-language transformer that extends existing video captioning model with an additional audio branch. We combine the masked language modeling and auto-regressive language modeling losses to optimize our model so that it can produce paragraph-level descriptions. We illustrate the efficiency of our model in audio-visual-language modeling by evaluating it against the proposed benchmark using both conventional captioning metrics and our proposed metrics. We further put our benchmark to the test in video generation models, demonstrating that employing fine-grained video descriptions can create more intricate videos than using captions. Code and dataset are available at https://github.com/OpenNLPLab/FAVDBench. Our online benchmark is available at www.avlbench.opennlplab.cn. Xuyang Shen, Dong Li 0033, Jinxing Zhou, Zhen Qin 0003, Xiaodong Han, Aixuan Li, Yuchao Dai, Lingpeng Kong, Meng Wang 0001, Yu Qiao 0001, Yiran Zhong |
CVPR | 8 |
| 2023 | Modeling the Distributional Uncertainty for Salient Object Detection ModelsabstractMost of the existing salient object detection (SOD) models focus on improving the overall model performance, without explicitly explaining the discrepancy between the training and testing distributions. In this paper, we investigate a particular type of epistemic uncertainty, namely distributional uncertainty, for salient object detection. Specifically, for the first time, we explore the existing class-aware distribution gap exploration techniques, i.e. long-tail learning, single-model uncertainty modeling and test-time strategies, and adapt them to model the distributional uncertainty for our class-agnostic task. We define test sample that is dissimilar to the training dataset as being “out-of-distribution” (OOD) samples. Different from the conventional OOD definition, where OOD samples are those not belonging to the closed-world training categories, OOD samples for SOD are those break the basic priors of saliency, i.e. center prior, color contrast prior, compactness prior and etc., indicating OOD as being “continuous” instead of being discrete for our task. We've carried out extensive experimental results to verify effectiveness of existing distribution gap modeling techniques for SOD, and conclude that both train-time single-model uncertainty estimation techniques and weight-regularization solutions that preventing model activation from drifting too much are promising directions for modeling distributional uncertainty for SOD. Jing Zhang 0052, Mochu Xiang, Yuchao Dai |
CVPR | 4 |
| 2023 | Forward Flow for Novel View Synthesis of Dynamic ScenesabstractThis paper proposes a neural radiance field (NeRF) approach for novel view synthesis of dynamic scenes using forward warping. Existing methods often adopt a static NeRF to represent the canonical space, and render dynamic images at other time steps by mapping the sampled 3D points back to the canonical space with the learned backward flow field. However, this backward flow field is non-smooth and discontinuous, which is difficult to be fitted by commonly used smooth motion models. To address this problem, we propose to estimate the forward flow field and directly warp the canonical radiance field to other time steps. Such forward flow field is smooth and continuous within the object region, which benefits the motion model learning. To achieve this goal, we represent the canonical radiance field with voxel grids to enable efficient forward warping, and propose a differentiable warping process, including an average splatting operation and an inpaint network, to resolve the many-to-one and one-to-many mapping issues. Thorough experiments show that our method outperforms existing methods in both novel view rendering and motion modeling, demonstrating the effectiveness of our forward flow motion modeling. Project page: https://npucvr.github.io/ForwardFlowDNeRF. Jiadai Sun, Yuchao Dai, Guanying Chen, Xiaoqing Ye, Xiao Tan 0001, Errui Ding, Jingdong Wang 0001 |
ICCV | 3 |
| 2023 | Efficient LiDAR Point Cloud Oversegmentation NetworkabstractPoint cloud oversegmentation is a challenging task since it needs to produce perceptually meaningful partitions (i.e., superpoints) of a point cloud. Most existing oversegmentation methods cannot efficiently generate superpoints from large-scale LiDAR point clouds due to complex and inefficient procedures. In this paper, we propose a simple yet efficient end-to-end LiDAR oversegmentation network, which segments superpoints from the LiDAR point cloud by grouping points based on low-level point embeddings. Specifically, we first learn the similarity of points from the constructed local neighborhoods to obtain low-level point embeddings through the local discriminative loss. Then, to generate homogeneous superpoints from the sparse LiDAR point cloud, we propose a LiDAR point grouping algorithm that simultaneously considers the similarity of point embeddings and the Euclidean distance of points in 3D space. Finally, we design a superpoint refinement module for accurately assigning the hard boundary points to the corresponding superpoints. Extensive results on two large-scale outdoor datasets, SemanticKITTI and nuScenes, show that our method achieves a new state-of-the-art in LiDAR oversegmentation. Notably, the inference time of our method is 100× faster than that of other methods. Furthermore, we apply the learned superpoints to the LiDAR semantic segmentation task and the results show that using superpoints can significantly improve the LiDAR semantic segmentation of the baseline network. Code is available at https://github.com/fpthink/SuperLiDAR. Le Hui, Linghua Tang, Yuchao Dai, Jin Xie 0001, Jian Yang 0003 |
ICCV | 3 |
| 2023 | Multimodal Variational Auto-encoder based Audio-Visual SegmentationabstractWe propose an Explicit Conditional Multimodal Variational Auto-Encoder (ECMVAE) for audio-visual segmentation (AVS), aiming to segment sound sources in the video sequence. Existing AVS methods focus on implicit feature fusion strategies, where models are trained to fit the discrete samples in the dataset. With a limited and less diverse dataset, the resulting performance is usually unsatisfactory. In contrast, we address this problem from an effective representation learning perspective, aiming to model the contribution of each modality explicitly. Specifically, we find that audio contains critical category information of the sound producers, and visual data provides candidate sound producer(s). Their shared information corresponds to the target sound producer(s) shown in the visual data. In this case, cross-modal shared representation learning is especially important for AVS. To achieve this, our ECMVAE factorizes the representations of each modality with a modality-shared representation and a modality-specific representation. An orthogonality constraint is applied between the shared and specific representations to maintain the exclusive attribute of the factorized latent code. Further, a mutual information maximization regularizer is introduced to achieve extensive exploration of each modality. Quantitative and qualitative evaluations on the AVSBench demonstrate the effectiveness of our approach, leading to a new state-of-the-art for AVS, with a 3.84 mIOU performance leap on the challenging MS3 subset for multiple sound source segmentation. Yuxin Mao, Jing Zhang 0052, Mochu Xiang, Yiran Zhong, Yuchao Dai |
ICCV | 5 |
| 2023 | RPEFlow: Multimodal Fusion of RGB-PointCloud-Event for Joint Optical Flow and Scene Flow EstimationabstractRecently, the RGB images and point clouds fusion methods have been proposed to jointly estimate 2D optical flow and 3D scene flow. However, as both conventional RGB cameras and LiDAR sensors adopt a frame-based data acquisition mechanism, their performance is limited by the fixed low sampling rates, especially in highly-dynamic scenes. By contrast, the event camera can asynchronously capture the intensity changes with a very high temporal resolution, providing complementary dynamic information of the observed scenes. In this paper, we incorporate RGB images, Point clouds and Events for joint optical flow and scene flow estimation with our proposed multi-stage multimodal fusion model, RPEFlow. First, we present an attention fusion module with a cross-attention mechanism to implicitly explore the internal cross-modal correlation for 2D and 3D branches, respectively. Second, we introduce a mutual information regularization term to explicitly model the complementary information of three modalities for effective multimodal feature learning. We also contribute a new synthetic dataset to advocate further research. Experiments on both synthetic and real datasets show that our model outperforms the existing state-of-theart by a wide margin. Code and dataset is available at https://npucvr.github.io/RPEFlow. Zhexiong Wan, Yuxin Mao, Jing Zhang 0052, Yuchao Dai |
ICCV | 4 |
| 2023 | LRRU: Long-short Range Recurrent Updating Networks for Depth CompletionabstractExisting deep learning-based depth completion methods generally employ massive stacked layers to predict the dense depth map from sparse input data. Although such approaches greatly advance this task, their accompanied huge computational complexity hinders their practical applications. To accomplish depth completion more efficiently, we propose a novel lightweight deep network framework, the Long-short Range Recurrent Updating (LRRU) network. Without learning complex feature representations, LRRU first roughly fills the sparse input to obtain an initial dense depth map, and then iteratively updates it through learned spatially-variant kernels. Our iterative update process is content-adaptive and highly flexible, where the kernel weights are learned by jointly considering the guidance RGB images and the depth map to be updated, and large-to-small kernel scopes are dynamically adjusted to capture long-to-short range dependencies. Our initial depth map has coarse but complete scene depth information, which helps relieve the burden of directly regressing the dense depth from sparse ones, while our proposed method can effectively refine it to an accurate depth map with less learnable parameters and inference time. Experimental results demonstrate that our proposed LRRU variants achieve state-of-the-art performance across different parameter regimes. In particular, the LRRU-Base model outperforms competing approaches on the NYUv2 dataset, and ranks 1st on the KITTI depth completion benchmark at the time of submission. Project page: https://npucvr.github.io/LRRU/. Bo Li 0090, Ge Zhang 0006, Qi Liu 0054, Tao Gao 0001, Yuchao Dai |
ICCV | 6 |
| 2023 | Toeplitz Neural Network for Sequence Modeling
Zhen Qin 0003, Xiaodong Han, Weixuan Sun, Dong Li 0033, Dongxu Li 0003, Yuchao Dai, Lingpeng Kong, Yiran Zhong |
ICLR | 7 |
| 2023 | Digging into Depth Priors for Outdoor Neural Radiance FieldsabstractNeural Radiance Fields (NeRFs) have demonstrated impressive performance in vision and graphics tasks, such as novel view synthesis and immersive reality. However, the shape-radiance ambiguity of radiance fields remains a challenge, especially in the sparse viewpoints setting. Recent work resorts to integrating depth priors into outdoor NeRF training to alleviate the issue. However, the criteria for selecting depth priors and the relative merits of different priors have not been thoroughly investigated. Moreover, the relative merits of selecting different approaches to use the depth priors is also an unexplored problem. In this paper, we provide a comprehensive study and evaluation of employing depth priors to outdoor neural radiance fields, covering common depth sensing technologies and most application ways. Specifically, we conduct extensive experiments with two representative NeRF methods equipped with four commonly-used depth priors and different depth usages on two widely used outdoor datasets. Our experimental results reveal several interesting findings that can potentially benefit practitioners and researchers in training their NeRF models with depth priors. Project page: https://cwchenwang.github.io/outdoor-nerf-depth Chen Wang 0049, Jiadai Sun, Lina Liu 0010, Chenming Wu, Zhelun Shen, Dayan Wu, Yuchao Dai, Liangjun Zhang |
ACM Multimedia | 7 |
| 2023 | Continuous Parametric Optical FlowabstractIn this paper, we present continuous parametric optical flow, a parametric representation of dense and continuous motion over arbitrary time interval. In contrast to existing discrete-time representations (i.e., flow in between consecutive frames), this new representation transforms the frame-to-frame pixel correspondences to dense continuous flow. In particular, we present a temporal-parametric model that employs B-splines to fit point trajectories using a limited number of frames. To further improve the stability and robustness of the trajectories, we also add an encoder with a neural ordinary differential equation (NODE) to represent features associated with specific times. We also contribute a synthetic dataset and introduce two evaluation perspectives to measure the accuracy and robustness of continuous flow estimation. Benefiting from the combination of explicit parametric modeling and implicit feature optimization, our model focuses on motion continuity and outperforms the flow-based and point-tracking approaches for fitting long-term and variable sequences. Jianqin Luo, Zhexiong Wan, Yuxin Mao, Bo Li 0090, Yuchao Dai |
NeurIPS | 5 |
| 2023 | NetPanel: Traffic Measurement of Exchange Online Service
Liqun Li, Yu Kang 0006, Boyang Zheng, Yehan Wang, More Zhou, Yuchao Dai, Zhenguo Yang, Brad Rutkowski, Jeff Mealiffe, Qingwei Lin |
NSDI | 7 |
| 2023 | Event-guided Multi-patch Network with Self-supervision for Non-uniform Motion Deblurring
Limeng Zhang, Yuchao Dai, Hongdong Li, Piotr Koniusz |
Int. J. Comput. Vis. | 3 |
| 2023 | Rolling Shutter Inversion: Bring Rolling Shutter Images to High Framerate Global Shutter VideoabstractA single rolling-shutter (RS) image may be viewed as a row-wise combination of a sequence of global-shutter (GS) images captured by a (virtual) moving GS camera within the exposure duration. Although rolling-shutter cameras are widely used, the RS effect causes obvious image distortion especially in the presence of fast camera motion, hindering downstream computer vision tasks. In this paper, we propose to invert the rolling-shutter image capture mechanism, i.e., recovering a continuous high framerate global-shutter video from two time-consecutive RS frames. We call this task the RS temporal super-resolution (RSSR) problem. The RSSR is a very challenging task, and to our knowledge, no practical solution exists to date. This paper presents a novel deep-learning based solution. By leveraging the multi-view geometry relationship of the RS imaging process, our learning based framework successfully achieves high framerate GS generation. Specifically, three novel contributions can be identified: (i) novel formulations for bidirectional RS undistortion flows under constant velocity as well as constant acceleration motion model. (ii) a simple linear scaling operation, which bridges the RS undistortion flow and regular optical flow. (iii) a new mutual conversion scheme between varying RS undistortion flows that correspond to different scanlines. Our method also exploits the underlying spatial-temporal geometric relationships within a deep learning framework, where no additional supervision is required beyond the necessary middle-scanline GS image. Building upon these contributions, this paper represents the very first rolling-shutter temporal super-resolution deep-network that is able to recover high framerate global-shutter videos from just two RS frames. Extensive experimental results on both synthetic and real data show that our proposed method can produce high-quality GS image sequences with rich details, outperforming the state-of-the-art methods. Bin Fan 0002, Yuchao Dai, Hongdong Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Digging Into Uncertainty-Based Pseudo-Label for Robust Stereo MatchingabstractDue to the domain differences and unbalanced disparity distribution across multiple datasets, current stereo matching approaches are commonly limited to a specific dataset and generalize poorly to others. Such domain shift issue is usually addressed by substantial adaptation on costly target-domain ground-truth data, which cannot be easily obtained in practical settings. In this paper, we propose to dig into uncertainty estimation for robust stereo matching. Specifically, to balance the disparity distribution, we employ a pixel-level uncertainty estimation to adaptively adjust the next stage disparity searching space, in this way driving the network progressively prune out the space of unlikely correspondences. Then, to solve the limited ground truth data, an uncertainty-based pseudo-label is proposed to adapt the pre-trained model to the new domain, where pixel-level and area-level uncertainty estimation are proposed to filter out the high-uncertainty pixels of predicted disparity maps and generate sparse while reliable pseudo-labels to align the domain gap. Experimentally, our method shows strong cross-domain, adapt, and joint generalization and obtains 1st place on the stereo task of Robust Vision Challenge 2020. Additionally, our uncertainty-based pseudo-labels can be extended to train monocular depth estimation networks in an unsupervised way and even achieves comparable performance with the supervised methods. Zhelun Shen, Xibin Song, Yuchao Dai, Dingfu Zhou, Zhibo Rao, Liangjun Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | MUNet: Motion uncertainty-aware semi-supervised video object segmentation
Jiadai Sun, Yuxin Mao, Yuchao Dai, Yiran Zhong |
Pattern Recognit. | 3 |
| 2023 | Toward Deeper Understanding of Camouflaged Object DetectionabstractPreys in the wild evolve to be camouflaged to avoid being recognized by predators. In this way, camouflage acts as a key defence mechanism across species that is critical to survival. To detect and segment the whole scope of a camouflaged object, camouflaged object detection (COD) is introduced as a binary segmentation task, with the binary ground truth camouflage map indicating the exact regions of the camouflaged objects. In this paper, we revisit this task and argue that the binary segmentation setting fails to fully understand the concept of camouflage. We find that explicitly modeling the conspicuousness of camouflaged objects against their particular backgrounds can not only lead to a better understanding about camouflage, but also provide guidance to designing more sophisticated camouflage techniques. Furthermore, we observe that it is some specific parts of camouflaged objects that make them detectable by predators. With the above understanding about camouflaged objects, we present the first triple-task learning framework to simultaneouslylocalize, segment, and rankcamouflaged objects, indicating the conspicuousness level of camouflage. As no corresponding datasets exist for either the localization model or the ranking model, we generate localization maps with an eye tracker, which are then processed according to the instance level labels to generate our ranking-based training and testing dataset. We also contribute the largest COD testing set to comprehensively analyse performance of the COD models. Experimental results show that our triple-task learning framework achieves new state-of-the-art, leading to a more explainable COD network. Our code, data, and results are available at:https://github.com/JingZhang617/COD-Rank-Localize-and-Segment. Yunqiu Lv, Jing Zhang 0052, Yuchao Dai, Aixuan Li, Nick Barnes, Deng-Ping Fan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Deep Idempotent Network for Efficient Single Image Blind DeblurringabstractSingle image blind deblurring is highly ill-posed as neither the latent sharp image nor the blur kernel is known. Even though considerable progress has been made, several major difficulties remain for blind deblurring, including the trade-off between high-performance deblurring and real-time processing. Besides, we observe that current single image blind deblurring networks cannot further improve or stabilize the performance but significantly degrades the performance when re-deblurring is repeatedly applied. This implies the limitation of these networks in modeling an ideal deblurring process. In this work, we make two contributions to tackle the above difficulties: (1) We introduce the idempotent constraint into the deblurring framework and present a deep idempotent network to achieve improved blind non-uniform deblurring performance with stable re-deblurring. (2) We propose a simple yet efficient deblurring network with lightweight encoder-decoder units and a recurrent structure that can deblur images in a progressive residual fashion. Extensive experiments on synthetic and realistic datasets prove the superiority of our proposed framework. Remarkably, our proposed network is nearly$6.5\times $smaller and$6.4\times $faster than the state-of-the-art while achieving comparable high performance. Yuxin Mao, Zhexiong Wan, Yuchao Dai, Xin Yu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | WSAMF-Net: Wavelet Spatial Attention-Based MultiStream Feedback Network for Single Image DehazingabstractSingle image-based dehazing has achieved remarkable progress with the development of deep learning technologies. End-to-end neural networks have been proposed to learn a direct hazy-to-clear image translation to recover the clear structures and edges cues from the hazy inputs. However, the frequency domain information is explored insufficiently and lots of intermediate structure and texture related cues of current dehazing networks are ignored, which limits the performances of current approaches. To handle these limitations mentioned above, a wavelet spatial attention based multi-stream feedback network (WSAMF-Net) is proposed for effective single image dehazing. Specifically, the proposed wavelet spatial attention utilizes both frequency-domain and spatial-domain information to enhance the extracted features for better structures and edges. Meanwhile, an enhanced multi-stream based cross feature fusion strategy, including vertical and horizontal attentions, is proposed to reweight and fuse the intermediate features of each stream to acquire more meaningful aggregated features, while the weight sharing strategy is used to achieve a good trade-off between performance and parameters. Besides, feedback mechanism is also designed to provide strong reconstruction ability. Furthermore, we propose a critical real-world industrial dataset (IDS) with images captured in real-world industrial quarry scenarios for research uses. Extensive experiments on various benchmarking datasets, including both synthetic and real-world datasets, demonstrate the superiority of our WSAMF-Net over state-of-the-art single image dehazing methods. The IDS dataset will be available athttps://github.com/XBSong/IDS-Datasethttps://github.com/XBSong/IDS-Dataset. Xibin Song, Dingfu Zhou, Wei Li 0143, Haodong Ding, Yuchao Dai, Liangjun Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | TUSR-Net: Triple Unfolding Single Image Dehazing With Self-Regularization and Dual Feature to Pixel AttentionabstractSingle image dehazing is a challenging and ill-posed problem due to severe information degeneration of images captured in hazy conditions. Remarkable progresses have been achieved by deep-learning based image dehazing methods, where residual learning is commonly used to separate the hazy image into clear and haze components. However, the nature of low similarity between haze and clear components is commonly neglected, while the lack of constraint of contrastive peculiarity between the two components always restricts the performance of these approaches. To deal with these problems, we propose an end-to-end self-regularized network (TUSR-Net) which exploits the contrastive peculiarity of different components of the hazy image, i.e, self-regularization (SR). In specific, the hazy image is separated into clear and hazy components and constraint between different image components, i.e., self-regularization, is leveraged to pull the recovered clear image closer to groundtruth, which largely promotes the performance of image dehazing. Meanwhile, an effective triple unfolding framework combined with dual feature to pixel attention is proposed to intensify and fuse the intermediate information in feature, channel and pixel levels, respectively, thus features with better representational ability can be obtained. Our TUSR-Net achieves better trade-off between performance and parameter size with weight-sharing strategy and is much more flexible. Experiments on various benchmarking datasets demonstrate the superiority of our TUSR-Net over state-of-the-art single image dehazing methods. Xibin Song, Dingfu Zhou, Wei Li 0143, Yuchao Dai, Zhelun Shen, Liangjun Zhang, Hongdong Li |
IEEE Trans. Image Process. | 4 |
| 2023 | Rethinking Training Strategy in Stereo MatchingabstractIn stereo matching, various learning-based approaches have shown impressive performance in solving traditional difficulties on multiple datasets. While most progress is obtained on a specific dataset with a dataset-specific network design, the performance on the single dataset and cross dataset affected by training strategy is often ignored. In this article, we analyze the relationship between different training strategies and performance by retraining some representative state-of-the-art methods (e.g., geometry and context network (GC-Net), pyramid stereo matching network (PSM-Net), and guided aggregation network (GA-Net), etc.). According to our research, it is surprising that the performance of networks on single or cross datasets is significantly improved by pre-training and data augmentation without any particular structure acquirement. Based on this discovery, we improve our previous non-local context attention network (NLCA-Net) to NLCA-Net v2 and train it with the novel strategy and rethink the training strategy of stereo matching concurrently. The quantitative experiments demonstrate that: 1) our model is capable of reaching top performance on both the single dataset and the multiple datasets with the same parameters in this study, which also won the 2nd place in the stereo task of the ECCV Robust vision Challenge 2020 (RVC 2020); and 2) on small datasets (e.g., KITTI, ETH3D, and Middlebury), the model's generalization and robustness are significantly affected by pre-training and data augmentation, even exceeding the network structure's influence in some cases. These observations present a challenge to the conventional wisdom of network architectures in this stage. We expect these discoveries to encourage researchers to rethink the current paradigm of "excessive attention on the performance of a single small dataset" in stereo matching. Zhibo Rao, Yuchao Dai, Zhelun Shen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | End-to-End Learning the Partial Permutation Matrix for Robust 3D Point Cloud RegistrationabstractEven though considerable progress has been made in deep learning-based 3D point cloud processing, how to obtain accurate correspondences for robust registration remains a major challenge because existing hard assignment methods cannot deal with outliers naturally. Alternatively, the soft matching-based methods have been proposed to learn the matching probability rather than hard assignment. However, in this paper, we prove that these methods have an inherent ambiguity causing many deceptive correspondences. To address the above challenges, we propose to learn a partial permutation matching matrix, which does not assign corresponding points to outliers, and implements hard assignment to prevent ambiguity. However, this proposal poses two new problems, i.e. existing hard assignment algorithms can only solve a full rank permutation matrix rather than a partial permutation matrix, and this desired matrix is defined in the discrete space, which is non-differentiable. In response, we design a dedicated soft-to-hard (S2H) matching procedure within the registration pipeline consisting of two steps: solving the soft matching matrix (S-step) and projecting this soft matrix to the partial permutation matrix (H-step). Specifically, we augment the profit matrix before the hard assignment to solve an augmented permutation matrix, which is cropped to achieve the final partial permutation matrix. Moreover, to guarantee end-to-end learning, we supervise the learned partial permutation matrix but propagate the gradient to the soft matrix instead. Our S2H matching procedure can be easily integrated with existing registration frameworks, which has been verified in representative frameworks including DCP, RPMNet, and DGR. Extensive experiments have validated our method, which creates a new state-of-the-art performance. Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Dingfu Zhou, Xibin Song, Mingyi He |
AAAI | 3 |
| 2022 | Neural Deformable Voxel Grid for Fast Optimization of Dynamic View Synthesis
Guanying Chen, Yuchao Dai, Xiaoqing Ye, Jiadai Sun, Xiao Tan 0001, Errui Ding |
ACCV (1) | 3 |
| 2022 | A General Divergence Modeling Strategy for Salient Object Detection
Jing Zhang 0052, Yuchao Dai |
ACCV (7) | 3 |
| 2022 | Context-Aware Video Reconstruction for Rolling Shutter CamerasabstractWith the ubiquity of rolling shutter (RS) cameras, it is becoming increasingly attractive to recover the latent global shutter (GS) video from two consecutive RS frames, which also places a higher demand on realism. Existing solutions, using deep neural networks or optimization, achieve promising performance. However, these methods generate intermediate GS frames through image warping based on the RS model, which inevitably result in black holes and noticeable motion artifacts. In this paper, we alleviate these issues by proposing a context-aware GS video reconstruction architecture. It facilitates the advantages such as occlusion reasoning, motion compensation, and temporal abstraction. Specifically, we first estimate the bilateral motion field so that the pixels of the two RS frames are warped to a common GS frame accordingly. Then, a refinement scheme is proposed to guide the GS frame synthesis along with bilateral occlusion masks to produce high-fidelity GS video frames at arbitrary times. Furthermore, we derive an approximated bilateral motion field model, which can serve as an alternative to provide a simple but effective GS frame initialization for related tasks. Experiments on synthetic and real data show that our approach achieves superior performance over state-of-the-art methods in terms of objective metrics and subjective visual quality. Code is available at https://github.com/GitCVfb/CVR. Bin Fan 0002, Yuchao Dai, Zhiyuan Zhang 0002, Qi Liu 0054, Mingyi He |
CVPR | 2 |
| 2022 | Efficient Multi-view Stereo by Iterative Dynamic Cost VolumeabstractIn this paper, we propose a novel iterative dynamic cost volume for multi-view stereo. Compared with other works, our cost volume is much lighter, thus could be processed with 2D convolution based GRU. Notably, the every-step output of the GRU could be further used to generate new cost volume. In this way, an iterative GRU-based optimizer is constructed. Furthermore, we present a cascade and hierarchical refinement architecture to utilize the multiscale information and speed up the convergence. Specifically, a lightweight 3D CNN is utilized to generate the coarsest initial depth map which is essential to launch the GRU and guarantee a fast convergence. Then the depth map is refined by multi-stage GRUs which work on the pyramid feature maps. Extensive experiments on the DTU and Tanks & Temples benchmarks demonstrate that our method could achieve state-of-the-art results in terms of accuracy, speed and memory usage. Code will be released at https://github.com/bdwsq1996/Effi-MVS. Shaoqian Wang, Bo Li 0090, Yuchao Dai |
CVPR | 3 |
| 2022 | PCW-Net: Pyramid Combination and Warping Cost Volume for Stereo Matching
Zhelun Shen, Yuchao Dai, Xibin Song, Zhibo Rao, Dingfu Zhou, Liangjun Zhang |
ECCV (32) | 2 |
| 2022 | Efficient Spatial-Temporal Information Fusion for LiDAR-Based 3D Moving Object SegmentationabstractAccurate moving object segmentation is an es-sential task for autonomous driving. It can provide effective information for many downstream tasks, such as collision avoidance, path planning, and static map construction. How to effectively exploit the spatial-temporal information is a critical question for 3D LiDAR moving object segmentation (LiDAR-MOS). In this work, we propose a novel deep neural network exploiting both spatial-temporal information and different representation modalities of LiDAR scans to improve LiDAR-MOS performance. Specifically, we first use a range image-based dual-branch structure to separately deal with spatial and temporal information that can be obtained from sequential LiDAR scans, and later combine them using motion-guided attention modules. We also use a point refinement module via 3D sparse convolution to fuse the information from both LiDAR range image and point cloud representations and reduce the artifacts on the borders of the objects. We verify the effectiveness of our proposed approach on the LiDAR-MOS benchmark of SemanticKITTI. Our method outperforms the state-of-the-art methods significantly in terms of LiDAR-MOS IoU. Benefiting from the devised coarse-to-fine architecture, our method operates online at sensor frame rate. Code is available at: https://github.com/haomo-ai/MotionSeg3D. Jiadai Sun, Yuchao Dai, Xianjing Zhang, Jintao Xu 0001, Rui Ai 0001, Weihao Gu, Xieyuanli Chen |
IROS | 2 |
| 2022 | Displacement-Invariant Cost Computation for Stereo MatchingabstractAbstract Although deep learning-based methods have dominated stereo matching leaderboards by yielding unprecedented disparity accuracy, their inference time is typically slow, i.e., less than 4 FPS for a pair of 540p images. The main reason is that the leading methods employ time-consuming 3D convolutions applied to a 4D feature volume. A common way to speed up the computation is to downsample the feature volume, but this loses high-frequency details. To overcome these challenges, we propose a displacement-invariant cost computation module to compute the matching costs without needing a 4D feature volume. Rather, costs are computed by applying the same 2D convolution network on each disparity-shifted feature map pair independently. Unlike previous 2D convolution-based methods that simply perform context mapping between inputs and disparity maps, our proposed approach learns to match features between the two images. We also propose an entropy-based refinement strategy to refine the computed disparity map, which further improves the speed by avoiding the need to compute a second disparity map on the right image. Extensive experiments on standard datasets (SceneFlow, KITTI, ETH3D, and Middlebury) demonstrate that our method achieves competitive accuracy with much less inference time. On typical image sizes (e.g., $$540\times 960$$ 540 × 960 ), our method processes over 100 FPS on a desktop GPU, making our method suitable for time-critical applications such as autonomous driving. We also show that our approach generalizes well to unseen datasets, outperforming 4D-volumetric methods. We will release the source code to ensure the reproducibility. Yiran Zhong, Charles T. Loop, Wonmin Byeon, Stanley T. Birchfield, Yuchao Dai, Kaihao Zhang, Alexey Kamenev, Thomas M. Breuel, Hongdong Li, Jan Kautz |
Int. J. Comput. Vis. | 5 |
| 2022 | Differential SfM and image correction for a rolling shutter stereo rig
Bin Fan 0002, Yuchao Dai, Zhiyuan Zhang 0002 |
Image Vis. Comput. | 2 |
| 2022 | A Representation Separation Perspective to Correspondence-Free Unsupervised 3-D Point Cloud Registrationabstract3-D point cloud registration in remote sensing field has been greatly advanced by deep learning-based methods, where the rigid transformation is either directly regressed from the two point clouds (correspondences-free approaches) or computed from the learned correspondences (correspondences-based approaches). Existing correspondence-free methods generally learn the holistic representation of the entire point cloud, which is fragile for partial and noisy point clouds. In this letter, we propose a correspondence-free unsupervised point cloud registration (UPCR) method from the representation separation perspective. First, we model the input point cloud as a combination of pose-invariant representation and pose-related representation. Second, the pose-related representation is used to learn the relative pose w.r.t. a “latent canonical shape” for thesourceandtargetpoint clouds, respectively. Third, the rigid transformation is obtained from the above two learned relative poses. Our method not only filters out the disturbance in pose-invariant representation but also is robust to partial-to-partial point clouds or noise. Experiments on benchmark datasets demonstrate that our unsupervised method achieves comparable if not better performance than state-of-the-art supervised registration methods.The source code will be made public. Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Dingfu Zhou, Xibin Song, Mingyi He |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Sliding space-disparity transformer for stereo matching
Zhibo Rao, Mingyi He, Yuchao Dai, Zhelun Shen |
Neural Comput. Appl. | 3 |
| 2022 | High Frame Rate Video Reconstruction Based on an Event CameraabstractEvent-based cameras measure intensity changes (called 'events') with microsecond accuracy under high-speed motion and challenging lighting conditions. With the 'active pixel sensor' (APS), the 'Dynamic and Active-pixel Vision Sensor' (DAVIS) allows the simultaneous output of intensity frames and events. However, the output images are captured at a relatively low frame rate and often suffer from motion blur. A blurred image can be regarded as the integral of a sequence of latent images, while events indicate changes between the latent images. Thus, we are able to model the blur-generation process by associating event data to a latent sharp image. Based on the abundant event data alongside a low frame rate, easily blurred images, we propose a simple yet effective approach to reconstruct high-quality and high frame rate sharp videos. Starting with a single blurred frame and its event data from DAVIS, we propose the Event-based Double Integral (EDI) model and solve it by adding regularization terms. Then, we extend it to multiple Event-based Double Integral (mEDI) model to get more smooth results based on multiple images and their events. Furthermore, we provide a new and more efficient solver to minimize the proposed energy model. By optimizing the energy function, we achieve significant improvements in removing blur and the reconstruction of a high temporal resolution video. The video generation is based on solving a simple non-convex optimization problem in a single scalar variable. Experimental results on both synthetic and real datasets demonstrate the superiority of our mEDI model and optimization method compared to the state-of-the-art. Liyuan Pan, Richard I. Hartley, Cedric Scheerlinck, Miaomiao Liu 0001, Xin Yu 0002, Yuchao Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Uncertainty Inspired RGB-D Saliency DetectionabstractWe propose the first stochastic framework to employ uncertainty for RGB-D saliency detection by learning from the data labeling process. Existing RGB-D saliency detection models treat this task as a point estimation problem by predicting a single saliency map following a deterministic learning pipeline. We argue that, however, the deterministic solution is relatively ill-posed. Inspired by the saliency data labeling process, we propose a generative architecture to achieve probabilistic RGB-D saliency detection which utilizes a latent variable to model the labeling variations. Our framework includes two main models: 1) a generator model, which maps the input image and latent variable to stochastic saliency prediction, and 2) an inference model, which gradually updates the latent variable by sampling it from the true or approximate posterior distribution. The generator model is an encoder-decoder saliency network. To infer the latent variable, we introduce two different solutions: i) a Conditional Variational Auto-encoder with an extra encoder to approximate the posterior distribution of the latent variable; and ii) an Alternating Back-Propagation technique, which directly samples the latent variable from the true posterior distribution. Qualitative and quantitative results on six challenging RGB-D benchmark datasets show our approach's superior performance in learning the distribution of saliency maps. The source code is publicly available via our project page: https://github.com/JingZhang617/UCNet. Jing Zhang 0052, Deng-Ping Fan, Yuchao Dai, Saeed Anwar, Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Nick Barnes |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Semi-supervised Active Salient Object Detection
Yunqiu Lv, Bowen Liu 0012, Jing Zhang 0052, Yuchao Dai, Aixuan Li, Tong Zhang 0023 |
Pattern Recognit. | 4 |
| 2022 | Self-supervised rigid transformation equivariance for accurate 3D point cloud registration
Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Dingfu Zhou, Xibin Song, Mingyi He |
Pattern Recognit. | 3 |
| 2022 | Fast and Robust Differential Relative Pose Estimation With Radial DistortionabstractIn this letter, we address the differential two-view geometry problem of estimating the relative pose between two consecutive frames in the presence of radial distortion. This problem is of both theoretical and practical interests and has not been solved. We derive its parameterization and present an effective and robust generalized eigenvalue solver based on the hidden variable technique. Furthermore, we propose a nonlinear refinement scheme within the maximum likelihood criterion to produce more accurate estimates of the relative pose and radial distortion. Compared with the standard differential solutions without modeling the radial distortion, our approach can recover more geometrically correct point correspondences for a pair of radially distorted images. Moreover, our differential solution runs an order of magnitude faster than the discrete solution in terms of recovering the full camera motion. Experiment results on both synthetic and real data demonstrate the effectiveness of our model and method in dealing with the radial distortion. Bin Fan 0002, Yuchao Dai, Zhiyuan Zhang 0002, Mingyi He |
IEEE Signal Process. Lett. | 2 |
| 2022 | Searching Dense Point Correspondences via Permutation Matrix LearningabstractAlthough 3D point cloud data has received widespread attentions as a general form of 3D signal expression, applying point clouds to the task of dense correspondence estimation between 3D shapes has not been investigated widely. Furthermore, even in the few existing 3D point cloud-based methods, an important and widely acknowledged principle,i.e. one-to-one matching, is usually ignored. In response, this paper presents a novel end-to-end learning-based method to estimate the dense correspondence of 3D point clouds, in which the problem of point matching is formulated as a zero-one assignment problem to achieve a permutation matching matrix to implement the one-to-one principle fundamentally. Note that the classical solutions of this assignment problem are always non-differentiable, which is fatal for deep learning frameworks. Thus we design a special matching module, which solves a doubly stochastic matrix at first and then projects this obtained approximate solution to the desired permutation matrix. Moreover, to guarantee end-to-end learning and the accuracy of the calculated loss, we calculate the loss from the learned permutation matrix but propagate the gradient to the doubly stochastic matrix directly which bypasses the permutation matrix during the backward propagation. Our method can be applied to both non-rigid and rigid 3D point cloud data and extensive experiments show that our method achieves state-of-the-art performance for dense correspondence learning.The code will be released. Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Bin Fan 0002, Qi Liu 0054 |
IEEE Signal Process. Lett. | 3 |
| 2022 | Learning a Task-Specific Descriptor for Robust Matching of 3D Point CloudsabstractExisting learning-based point feature descriptors are usually task-agnostic, which pursue describing the individual 3D point clouds as accurate as possible. However, the matching task aims at describing the corresponding points consistently across different 3D point clouds. Therefore these too accurate features may play a counterproductive role due to the inconsistent point feature representations of correspondences caused by the unpredictable noise, partiality, deformation, etc., in the local geometry. In this paper, we propose to learn a robust task-specific feature descriptor to consistently describe the correct point correspondence under interference. Born with anEncoder and aDynamicFusion module, our method EDFNet develops from two aspects. First, we augment the matchability of correspondences by utilizing their repetitive local structure. To this end, a special encoder is designed to exploit two input point clouds jointly for each point descriptor. It not only captures the local geometry of each point in the current point cloud by convolution, but also exploits the repetitive structure from paired point cloud by Transformer. Second, we propose a dynamical fusion module to jointly use different scale features. There is an inevitable struggle between robustness and discriminativeness of the single scale feature. Specifically, the small scale feature is robust since little interference exists in this small receptive field. But it is not sufficiently discriminative as there are many repetitive local structures within a point cloud. Thus the resultant descriptors will lead to many incorrect matches. In contrast, the large scale feature is more discriminative by integrating more neighborhood information. But it is easier to be disturbed since there is much more interference in the large receptive field. Compared with the conventional fusion strategy that handles multiple scale features equally, we analyze the consistency of them to judge the clean ones and perform larger aggregation weights on them during fusion. Then, a robust and discriminative feature descriptor is achieved by focusing on multiple clean scale features. Extensive evaluations validate that EDFNet learns a task-specific descriptor, which achieves state-of-the-art or comparable performance for robust matching of 3D point clouds. Zhiyuan Zhang 0002, Yuchao Dai, Bin Fan 0002, Jiadai Sun, Mingyi He |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | VRNet: Learning the Rectified Virtual Corresponding Points for 3D Point Cloud Registrationabstract3D point cloud registration is fragile to outliers, which are labeled as the points without corresponding points. To handle this problem, a widely adopted strategy is to estimate the relative pose based only on some accurate correspondences, which is achieved by building correspondences on the identified inliers or by selecting reliable ones. However, these approaches are usually complicated and time-consuming. By contrast, the virtual point-based methods learn the virtual corresponding points (VCPs) for allsourcepoints uniformly without distinguishing the outliers and the inliers. Although this strategy is time-efficient, the learned VCPs usually exhibit serious collapse degeneration due to insufficient supervision and the inherent distribution limitation. In this paper, we propose to exploit the best of both worlds and present a novel robust 3D point cloud registration framework. We follow the idea of the virtual point-based methods but learn a new type of virtual points called rectified virtual corresponding points (RCPs), which are defined as the point set with the same shape as thesourceand with the same pose as thetarget. Hence, a pair of consistent point clouds,i.e.sourceand RCPs, is formed by rectifying VCPs to RCPs (VRNet), through which reliable correspondences betweensourceand RCPs can be accurately obtained. Since the relative pose betweensourceand RCPs is the same as the relative pose betweensourceandtarget, the input point clouds can be registered naturally. Specifically, we first construct the initial VCPs by using an estimated soft matching matrix to perform a weighted average on thetargetpoints. Then, we design a correction-walk module to learn an offset to rectify VCPs to RCPs, which effectively breaks the distribution limitation of VCPs. Finally, we develop a hybrid loss function to enforce the shape and geometry structure consistency of the learned RCPs and thesourceto provide sufficient supervision. Extensive experiments on several benchmark datasets demonstrate that our method achieves advanced registration performance and time-efficiency simultaneously.The code will be made public. Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Bin Fan 0002, Mingyi He |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Learning Dense and Continuous Optical Flow From an Event CameraabstractEvent cameras such as DAVIS can simultaneously output high temporal resolution events and low frame-rate intensity images, which own great potential in capturing scene motion, such as optical flow estimation. Most of the existing optical flow estimation methods are based on two consecutive image frames and can only estimate discrete flow at a fixed time interval. Previous work has shown that continuous flow estimation can be achieved by changing the quantities or time intervals of events. However, they are difficult to estimate reliable dense flow, especially in the regions without any triggered events. In this paper, we propose a novel deep learning-based dense and continuous optical flow estimation framework from a single image with event streams, which facilitates the accurate perception of high-speed motion. Specifically, we first propose an event-image fusion and correlation module to effectively exploit the internal motion from two different modalities of data. Then we propose an iterative update network structure with bidirectional training for optical flow prediction. Therefore, our model can estimate reliable dense flow as two-frame-based methods, as well as estimate temporal continuous flow as event-based methods. Extensive experimental results on both synthetic and real captured datasets demonstrate that our model outperforms existing event-based state-of-the-art methods and our designed baselines for accurate dense and continuous optical flow estimation. Zhexiong Wan, Yuchao Dai, Yuxin Mao |
IEEE Trans. Image Process. | 2 |
| 2022 | Relative Pose Estimation for Light Field Cameras Based on LF-Point-LF-Point Correspondence ModelabstractIn this paper, we propose a relative pose estimation algorithm for micro-lens array (MLA)-based conventional light field (LF) cameras. First, by employing the matched LF-point pairs, we establish the LF-point-LF-point correspondence model to represent the correlation between LF features of the same 3D scene point in a pair of LFs. Then, we employ the proposed correspondence model to estimate the relative camera pose, which includes a linear solution and a non-linear optimization on manifold. Unlike prior related algorithms, which estimated relative poses based on the recovered depths of scene points, we adopt the estimated disparities to avoid the inaccuracy in recovering depths due to the ultra-small baseline between sub-aperture images of LF cameras. Experimental results on both simulated and real scene data have demonstrated the effectiveness of the proposed algorithm compared with classical as well as state-of-art relative pose estimation algorithms. Saiping Zhang, Dongyang Jin, Yuchao Dai, Fuzheng Yang 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | Context-Aware 3D Object Detection From a Single Image in Autonomous DrivingabstractCamera sensors have been widely used in Driver-Assistance and Autonomous Driving Systems due to their rich texture information. Recently, with the development of deep learning techniques, many approaches have been proposed to detect objects in 3D from a single frame, however, there is still much room for improvement. In this paper, we generally review the recently proposed state-of-the-art monocular-based 3D object detection approaches first. Based on the analysis of the disadvantage of previous center-based frameworks, a novel feature aggregation strategy has been proposed to boost the 3D object detection by exploring the context information. Specifically, an Instance-Guided Spatial Attention (IGSA) module is proposed to collect the local instance information and the Channel-Wise Feature Attention (CWFA) module is employed for aggregating the global context information. In addition, an instance-guided object regression strategy is also proposed to alleviate the influence of center location prediction uncertainty in the inference process. Finally, the proposed approach has been verified on the public 3D object detection benchmark. The experimental results show that the proposed approach can significantly boost the performance of the baseline method on both 3D detection and 2D Bird’s-Eye View among all three categories. Furthermore, our method outperforms all the monocular-based methods (even these trained with depth as auxiliary inputs) and achieves state-of-the-art performance on the KITTI benchmark. Dingfu Zhou, Xibin Song, Yuchao Dai, Hongdong Li, Liangjun Zhang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | WAFP-Net: Weighted Attention Fusion Based Progressive Residual Learning for Depth Map Super-ResolutionabstractDespite the remarkable progresses achieved in depth map super-resolution (DSR), it remains a major challenge to tackle with real-world degradation of low-resolution (LR) depth maps. Synthetic datasets are mainly used in existing DSR approaches, which is quite different from what would get from a real depth sensor. Besides, the enhancements of features in existing DSR approaches are not sufficiently enough, which also limit the performance. To alleviate these problems, we first propose two types of degradation models to describe the generation of LR depth maps, including bi-cubic down-sampling with noise and interval down-sampling, and different DSR models are learned correspondingly. Then, we propose a weighted attention fusion strategy that is embedded into a progressive residual learning framework, which guarantees that the high-resolution (HR) depth maps can be well recovered in a coarse-to-fine manner. The weighted attention fusion strategy can enhance the features with abundant high-frequency components in both global and local manners, thus better HR depth maps can be expected. Besides, to re-use the effective information in the progressive process sufficiently, a multi-stage fusion module is combined into the proposed framework, and the Total Generalized Variation (TGV) regularization and input loss are exploited to further improve the performance of our method. Extensive experiments of different benchmarks demonstrate the superiority of our approach over the state-of-the-art (SOTA) approaches. Xibin Song, Dingfu Zhou, Wei Li 0111, Yuchao Dai, Liu Liu 0009, Hongdong Li, Ruigang Yang, Liangjun Zhang |
IEEE Trans. Multim. | 4 |
| 2022 | Patch attention network with generative adversarial model for semi-supervised binocular disparity prediction
Zhibo Rao, Mingyi He, Yuchao Dai, Zhelun Shen |
Vis. Comput. | 3 |
| 2021 | Uncertainty-Aware Joint Salient Object and Camouflaged Object DetectionabstractVisual salient object detection (SOD) aims at finding the salient object(s) that attract human attention, while camouflaged object detection (COD) on the contrary intends to discover the camouflaged object(s) that hidden in the surrounding. In this paper, we propose a paradigm of lever-aging the contradictory information to enhance the detection ability of both salient object detection and camouflaged object detection. We start by exploiting the easy positive samples in the COD dataset to serve as hard positive samples in the SOD task to improve the robustness of the SOD model. Then, we introduce a "similarity measure" module to explicitly model the contradicting attributes of these two tasks. Furthermore, considering the uncertainty of labeling in both tasks’ datasets, we propose an adversarial learning network to achieve both higher order similarity measure and network confidence estimation. Experimental results on benchmark datasets demonstrate that our solution leads to state-of-the-art (SOTA) performance for both tasks1. Aixuan Li, Jing Zhang 0052, Yunqiu Lv, Bowen Liu 0012, Tong Zhang 0023, Yuchao Dai |
CVPR | 6 |
| 2021 | Simultaneously Localize, Segment and Rank the Camouflaged ObjectsabstractCamouflage is a key defence mechanism across species that is critical to survival. Common strategies for camouflage include background matching, imitating the color and pattern of the environment, and disruptive coloration, disguising body outlines [37]. Camouflaged object detection (COD) aims to segment camouflaged objects hiding in their surroundings. Existing COD models are built upon binary ground truth to segment the camouflaged objects without illustrating the level of camouflage. In this paper, we revisit this task and argue that explicitly modeling the conspicuousness of camouflaged objects against their particular backgrounds can not only lead to a better understanding about camouflage and evolution of animals, but also provide guidance to design more sophisticated camouflage techniques. Furthermore, we observe that it is some specific parts of the camouflaged objects that make them detectable by predators. With the above understanding about camouflaged objects, we present the first ranking based COD network (Rank-Net) to simultaneously localize, segment and rank camouflaged objects. The localization model is proposed to find the discriminative regions that make the camouflaged object obvious. The segmentation model segments the full scope of the camouflaged objects. Further, the ranking model infers the detectability of different camouflaged objects. Moreover, we contribute a large COD testing set to evaluate the generalization ability of COD models. Experimental results show that our model achieves new state-of-the-art, leading to a more interpretable COD network1. Yunqiu Lv, Jing Zhang 0052, Yuchao Dai, Aixuan Li, Bowen Liu 0012, Nick Barnes, Deng-Ping Fan |
CVPR | 3 |
| 2021 | CFNet: Cascade and Fused Cost Volume for Robust Stereo MatchingabstractRecently, the ever-increasing capacity of large-scale annotated datasets has led to profound progress in stereo matching. However, most of these successes are limited to a specific dataset and cannot generalize well to other datasets. The main difficulties lie in the large domain differences and unbalanced disparity distribution across a variety of datasets, which greatly limit the real-world applicability of current deep stereo matching models. In this paper, we propose CFNet, a Cascade and Fused cost volume based network to improve the robustness of the stereo matching network. First, we propose a fused cost volume representation to deal with the large domain difference. By fusing multiple low-resolution dense cost volumes to enlarge the receptive field, we can extract robust structural representations for initial disparity estimation. Second, we propose a cascade cost volume representation to alleviate the unbalanced disparity distribution. Specifically, we employ a variance-based uncertainty estimation to adaptively adjust the next stage disparity search space, in this way driving the network progressively prune out the space of unlikely correspondences. By iteratively narrowing down the disparity search space and improving the cost volume resolution, the disparity estimation is gradually refined in a coarse-to-fine manner. When trained on the same training images and evaluated on KITTI, ETH3D, and Middlebury datasets with the fixed model parameters and hyperparameters, our proposed method achieves the state-of-the-art overall performance and obtains the 1st place on the stereo task of Robust Vision Challenge 2020. The code will be available at https://github.com/gallenszl/CFNet. Zhelun Shen, Yuchao Dai, Zhibo Rao |
CVPR | 2 |
| 2021 | Deep Two-View Structure-From-Motion RevisitedabstractTwo-view structure-from-motion (SfM) is the cornerstone of 3D reconstruction and visual SLAM. Existing deep learning-based approaches formulate the problem by either recovering absolute pose scales from two consecutive frames or predicting a depth map from a single image, both of which are ill-posed problems. In contrast, we propose to revisit the problem of deep two-view SfM by leveraging the well-posedness of the classic pipeline. Our method consists of 1) an optical flow estimation network that predicts dense correspondences between two frames; 2) a normalized pose estimation module that computes relative camera poses from the 2D optical flow correspondences, and 3) a scale-invariant depth estimation network that leverages epipolar geometry to reduce the search space, refine the dense correspondences, and estimate relative depth maps. Extensive experiments show that our method outperforms all state-of-the-art two-view SfM methods by a clear margin on KITTI depth, KITTI VO, MVS, Scenes11, and SUN3D datasets in both relative pose and depth estimation. Yiran Zhong, Yuchao Dai, Stanley T. Birchfield, Kaihao Zhang, Nikolai Smolyanskiy, Hongdong Li |
CVPR | 3 |
| 2021 | RGB-D Saliency Detection via Cascaded Mutual Information MinimizationabstractExisting RGB-D saliency detection models do not explicitly encourage RGB and depth to achieve effective multi-modal learning. In this paper, we introduce a novel multistage cascaded learning framework via mutual information minimization to explicitly model the multi-modal information between RGB image and depth data. Specifically, we first map the feature of each mode to a lower dimensional feature vector, and adopt mutual information minimization as a regularizer to reduce the redundancy between appearance features from RGB and geometric features from depth. We then perform multi-stage cascaded learning to impose the mutual information minimization constraint at every stage of the network. Extensive experiments on benchmark RGB-D saliency datasets illustrate the effectiveness of our framework. Further, to prosper the development of this field, we contribute the largest (7× larger than NJU2K) COME15K dataset, which contains 15,625 image pairs with high quality polygon-/scribble-/object-/instance-/rank-level annotations. Based on these rich labels, we additionally construct four new benchmarks with strong baselines and observe some interesting phenomena, which can motivate future model design. Source code and dataset are available at https://github.com/JingZhang617/cascaded_rgbd_sod. Jing Zhang 0052, Deng-Ping Fan, Yuchao Dai, Xin Yu 0002, Yiran Zhong, Nick Barnes, Ling Shao 0001 |
ICCV | 3 |
| 2021 | Inverting a Rolling Shutter Camera: Bring Rolling Shutter Images to High Framerate Global Shutter VideoabstractRolling shutter (RS) images can be viewed as the result of the row-wise combination of global shutter (GS) images captured by a virtual moving GS camera over the period of camera readout time. The RS effect brings tremendous difficulties for the downstream applications. In this paper, we propose to invert the above RS imaging mechanism, i.e., recovering a high framerate GS video from consecutive RS images to achieve RS temporal super-resolution (RSSR). This extremely challenging problem, e.g., recovering 1440 GS images from two 720-height RS images, is far from being solved end-to-end. To address this challenge, we exploit the geometric constraint in the RS camera model, thus achieving geometry-aware inversion. Specifically, we make three contributions in resolving the above difficulties: (i) formulating the bidirectional RS undistortion flows under the constant velocity motion model, (ii) building the connection between the RS undistortion flow and optical flow via a scaling operation, and (iii) developing a mutual conversion scheme between varying RS undistortion flows that correspond to different scanlines. Building upon these formulations, we propose the first RS temporal super-resolution network in a cascaded structure to extract high framerate global shutter video. Our method explores the underlying spatio-temporal geometric relationships within a deep learning framework, where no extra supervision besides the middle-scanline ground truth GS image is needed. Essentially, our method can be very efficient for explicit propagation to generate GS images under any scanline. Experimental results on both synthetic and real data show that our method can produce high-quality GS image sequences with rich details, outperforming state-of-the-art methods. Bin Fan 0002, Yuchao Dai |
ICCV | 2 |
| 2021 | SUNet: Symmetric Undistortion Network for Rolling Shutter CorrectionabstractThe vast majority of modern consumer-grade cameras employ a rolling shutter mechanism, leading to image distortions if the camera moves during image acquisition. In this paper, we present a novel deep network to solve the generic rolling shutter correction problem with two consecutive frames. Our pipeline is symmetrically designed to predict the global shutter image corresponding to the intermediate time of these two frames, which is difficult for existing methods because it corresponds to a camera pose that differs most from the two frames. First, two time-symmetric dense undistortion flows are estimated by using well-established principles: pyramidal construction, warping, and cost volume processing. Then, both rolling shutter images are warped into a common global shutter one in the feature space, respectively. Finally, a symmetric consistency constraint is constructed in the image decoder to effectively aggregate the contextual cues of two rolling shutter images, thereby recovering the high-quality global shutter image. Extensive experiments with both synthetic and real data from public benchmarks demonstrate the superiority of our proposed approach over the state-of-the-art methods. Bin Fan 0002, Yuchao Dai, Mingyi He |
ICCV | 2 |
| 2021 | Neural Image Compression via Attentional Multi-scale Back Projection and Frequency DecompositionabstractIn recent years, neural image compression emerges as a rapidly developing topic in computer vision, where the state-of-the-art approaches now exhibit superior compression performance than their conventional counterparts. Despite the great progress, current methods still have limitations in preserving fine spatial details for optimal reconstruction, especially at low compression rates. We make three contributions in tackling this issue. First, we develop a novel back projection method with attentional and multi-scale feature fusion for augmented representation power. Our back projection method recalibrates the current estimation by establishing feedback connections between high-level and low-level attributes in an attentional and discriminative manner. Second, we propose to decompose the input image and separately process the distinct frequency components, whose derived latents are recombined using a novel dual attention module, so that details inside regions of interest could be explicitly manipulated. Third, we propose a novel training scheme for reducing the latent rounding residual. Experimental results show that, when measured in PSNR, our model reduces BD-rate by 9.88% and 10.32% over the state-of-the-art method, and 4.12% and 4.32% over the latest coding standard Versatile Video Coding (VVC) on the Kodak and CLIC2020 Professional Validation dataset, respectively. Our approach also produces more visually pleasant images when optimized for MS-SSIM. The significant improvement upon existing methods shows the effectiveness of our method in preserving and remedying spatial information for enhanced compression quality. Pei You, Shunyuan Han, Yuchao Dai, Hojae Lee |
ICCV | 6 |
| 2021 | UASNet: Uncertainty Adaptive Sampling Network for Deep Stereo MatchingabstractRecent studies have shown that cascade cost volume can play a vital role in deep stereo matching to achieve high resolution depth map with efficient hardware usage. However, how to construct good cascade volume as well as effective sampling for them are still under in-depth study. Previous cascade-based methods usually perform uniform sampling in a predicted disparity range based on variance, which easily misses the ground truth disparity and decreases disparity map accuracy. In this paper, we propose an uncertainty adaptive sampling network (UASNet) featuring two modules: an uncertainty distribution-guided range prediction (URP) model and an uncertainty-based disparity sampler (UDS) module. The URP explores the more discriminative uncertainty distribution to handle the complex matching ambiguities and to improve disparity range prediction. The UDS adaptively adjusts sampling interval to localize disparity with improved accuracy. With the proposed modules, our UASNet learns to construct cascade cost volume and predict full-resolution disparity map directly. Extensive experiments show that the proposed method achieves the highest ground truth covering ratio compared with other cascade cost volume based stereo matching methods. Our method also achieves top performance on both SceneFlow dataset and KITTI benchmark. Yamin Mao, Yuchao Dai, Qiang Wang 0023, Yun-Tae Kim, Hong-Seok Lee |
ICCV | 4 |
| 2021 | PR-RRN: Pairwise-Regularized Residual-Recursive Networks for Non-rigid Structure-from-MotionabstractWe propose PR-RRN, a novel neural-network based method for Non-rigid Structure-from-Motion (NRSfM). PR-RRN consists of Residual-Recursive Networks (RRN) and two extra regularization losses. RRN is designed to effectively recover 3D shape and camera from 2D keypoints with novel residual-recursive structure. As NRSfM is a highly under-constrained problem, we propose two new pairwise regularization to further regularize the reconstruction. The Rigidity-based Pairwise Contrastive Loss regularizes the shape representation by encouraging higher similarity between the representations of high-rigidity pairs of frames than low-rigidity pairs. We propose minimum singular-value ratio to measure the pairwise rigidity. The Pairwise Consistency Loss enforces the reconstruction to be consistent when the estimated shapes and cameras are exchanged between pairs. Our approach achieves state-of-the-art performance on CMU MOCAP and PASCAL3D+ dataset. Haitian Zeng, Yuchao Dai, Xin Yu 0002, Yi Yang 0001 |
ICCV | 2 |
| 2021 | Complementary Patch for Weakly Supervised Semantic SegmentationabstractWeakly Supervised Semantic Segmentation (WSSS) based on image-level labels has been greatly advanced by exploiting the outputs of Class Activation Map (CAM) to generate the pseudo labels for semantic segmentation. However, CAM merely discovers seeds from a small number of regions, which may be insufficient to serve as pseudo masks for semantic segmentation. In this paper, we formulate the expansion of object regions in CAM as an increase in information. From the perspective of information theory, we propose a novel Complementary Patch (CP) Representation and prove that the information of the sum of the CAMs by a pair of input images with complementary hidden (patched) parts, namely CP Pair, is greater than or equal to the information of the baseline CAM. Therefore, a CAM with more information related to object seeds can be obtained by narrowing down the gap between the sum of CAMs generated by the CP Pair and the original CAM. We propose a CP Network (CPN) implemented by a triplet network and three regularization functions. To further improve the quality of the CAMs, we propose a Pixel-Region Correlation Module (PRCM) to augment the contextual in-formation by using object-region relations between the feature maps and the CAMs. Experimental results on the PAS-CAL VOC 2012 datasets show that our proposed method achieves a new state-of-the-art in WSSS, validating the effectiveness of our CP Representation and CPN. Fei Zhang 0016, Chaochen Gu, Chenyue Zhang, Yuchao Dai |
ICCV | 4 |
| 2021 | Rolling-Shutter-stereo-aware motion estimation and image correction
Bin Fan 0002, Yuchao Dai |
Comput. Vis. Image Underst. | 2 |
| 2021 | Self-supervised multi-body scene flow estimation
Jihuang Dai, Yuchao Dai, Bin Fan 0002 |
Neurocomputing | 2 |
| 2021 | Superpixel Soup: Monocular Dense 3D Reconstruction of a Complex Dynamic SceneabstractThis work addresses the task of dense 3D reconstruction of a complex dynamic scene from images. The prevailing idea to solve this task is composed of a sequence of steps and is dependent on the success of several pipelines in its execution. To overcome such limitations with the existing algorithm, we propose a unified approach to solve this problem. We assume that a dynamic scene can be approximated by numerous piecewise planar surfaces, where each planar surface enjoys its own rigid motion, and the global change in the scene between two frames is as-rigid-as-possible (ARAP). Consequently, our model of a dynamic scene reduces to a soup of planar structures and rigid motion of these local planar structures. Using planar over-segmentation of the scene, we reduce this task to solving a "3D jigsaw puzzle" problem. Hence, the task boils down to correctly assemble each rigid piece to construct a 3D shape that complies with the geometry of the scene under the ARAP assumption. Further, we show that our approach provides an effective solution to the inherent scale-ambiguity in structure-from-motion under perspective projection. We provide extensive experimental results and evaluation on several benchmark datasets. Quantitative comparison with competing approaches shows state-of-the-art performance. Suryansh Kumar 0001, Yuchao Dai, Hongdong Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Learning Saliency From Single Noisy Labelling: A Robust Model Fitting PerspectiveabstractThe advances made in predicting visual saliency using deep neural networks come at the expense of collecting large-scale annotated data. However, pixel-wise annotation is labor-intensive and overwhelming. In this paper, we propose to learn saliency prediction from a single noisy labelling, which is easy to obtain (e.g., from imperfect human annotation or from unsupervised saliency prediction methods). With this goal, we address a natural question: Can we learn saliency prediction while identifying clean labels in a unified framework? To answer this question, we call on the theory of robust model fitting and formulate deep saliency prediction from a single noisy labelling as robust network learning and exploit model consistency across iterations to identify inliers and outliers (i.e., noisy labels). Extensive experiments on different benchmark datasets demonstrate the superiority of our proposed framework, which can learn comparable saliency prediction with state-of-the-art fully supervised saliency methods. Furthermore, we show that simply by treating ground truth annotations as noisy labelling, our framework achieves tangible improvements over state-of-the-art methods. Jing Zhang 0052, Yuchao Dai, Tong Zhang 0023, Mehrtash Harandi, Nick Barnes, Richard I. Hartley |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | RS-DPSNet: Deep Plane Sweep Network for Rolling Shutter Stereo ImagesabstractSince the rolling shutter (RS) camera successively exposes each scanline, accurately reconstructing scene depth from an RS stereo image pair remains a great challenge. Directly applying the deep-learning-based depth estimation methods tailored for the global shutter (GS) stereo images leads to undesirable RS depth results due to inherent flaws in the network structure. In this letter, we fill this gap by developing an end-to-end RS-stereo-aware plane sweep network to improve the accuracy of the classic GS-based algorithm (i.e.DPSNet) in estimating the RS depth map. Specifically, we derive the RS-stereo-aware plane sweep model and further produce a more accurate and efficient cost volume through the effective incorporation of this model within DPSNet. Furthermore, to enable learning-based approaches to address the depth estimation problem in the context of RS stereo images, we contribute the first RS stereo dataset, CARLA-RSS. Experimental results demonstrate that our proposed pipeline achieves state-of-the-art performance. Bin Fan 0002, Yuchao Dai, Mingyi He |
IEEE Signal Process. Lett. | 3 |
| 2021 | Bidirectional Guided Attention Network for 3-D Semantic Detection of Remote Sensing ImagesabstractSemantic segmentation and disparity estimation are in the research frontier of the computer vision and remote sensing (RS) fields. However, existing methods mostly deal with these two problems separately or use a combination of multiple models to solve these two tasks. Due to a lack of sufficient information sharing and fusion, they still have difficulties in coping with seasonal appearance differences in 3-D RS problems. In this article, we propose a novel multitask learning architecture that considers the bottom–up and up–bottom visual attention mechanism for 3-D semantic detection, named bidirectional guided attention network (BGA-Net). BGA-Net consists of five modules: unified backbone module (UBM), bidirectional guided attention module (BGAM), semantic segmentation module (SSM), feature matching module (FMM), and bidirectional fusion module (BFM). First, in UBM, we use a shared backbone to extract unified features and share them with three branches/modules (BGAM, SSM, and FMM). Then, SSM and FMM branches are applied to estimate segmentation and disparity maps, whereas the third branch/module (BGAM) shares the global features to guide the task-specific learning via attention mechanism. Finally, we fuse the results of the two tasks by BFM to improve the final performance. Extensive experiments demonstrate that: 1) our BGA-Net can handle the two tasks simultaneously and can be trained in an end-to-end way; 2) these modules fully take advantage of the two tasks’ information to share features and enhance the scene understanding ability, effectively against seasons change of RS images; and 3) BGA-Net has notable superiority and greater flexibility and also sets a new state of the art on the urban semantic 3-D (US3D) benchmark. Moreover, BGA-Net also provides insights into the intelligent interpretation of RS data images. Zhibo Rao, Mingyi He, Zhidong Zhu, Yuchao Dai |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | MLDA-Net: Multi-Level Dual Attention-Based Network for Self-Supervised Monocular Depth EstimationabstractThe success of supervised learning-based single image depth estimation methods critically depends on the availability of large-scale dense per-pixel depth annotations, which requires both laborious and expensive annotation process. Therefore, the self-supervised methods are much desirable, which attract significant attention recently. However, depth maps predicted by existing self-supervised methods tend to be blurry with many depth details lost. To overcome these limitations, we propose a novel framework, named MLDA-Net, to obtain per-pixel depth maps with shaper boundaries and richer depth details. Our first innovation is a multi-level feature extraction (MLFE) strategy which can learn rich hierarchical representation. Then, a dual-attention strategy, combining global attention and structure attention, is proposed to intensify the obtained features both globally and locally, resulting in improved depth maps with sharper boundaries. Finally, a reweighted loss strategy based on multi-level outputs is proposed to conduct effective supervision for self-supervised depth estimation. Experimental results demonstrate that our MLDA-Net framework achieves state-of-the-art depth prediction results on the KITTI benchmark for self-supervised monocular depth estimation with different input modes and training modes. Extensive experiments on other benchmark datasets further confirm the superiority of our proposed approach. Xibin Song, Wei Li 0143, Dingfu Zhou, Yuchao Dai, Hongdong Li, Liangjun Zhang |
IEEE Trans. Image Process. | 4 |
| 2020 | IAFA: Instance-Aware Feature Aggregation for 3D Object Detection from a Single Image
Dingfu Zhou, Xibin Song, Yuchao Dai, Junbo Yin, Feixiang Lu, Miao Liao, Liangjun Zhang |
ACCV (1) | 3 |
| 2020 | Channel Attention Based Iterative Residual Learning for Depth Map Super-ResolutionabstractDespite the remarkable progresses made in deep learning based depth map super-resolution (DSR), how to tackle real-world degradation in low-resolution (LR) depth maps remains a major challenge. Existing DSR model is generally trained and tested on synthetic dataset, which is very different from what would get from a real depth sensor. In this paper, we argue that DSR models trained under this setting are restrictive and not effective in dealing with realworld DSR tasks. We make two contributions in tackling real-world degradation of different depth sensors. First, we propose to classify the generation of LR depth maps into two types: non-linear downsampling with noise and interval downsampling, for which DSR models are learned correspondingly. Second, we propose a new framework for real-world DSR, which consists of four modules : 1) An iterative residual learning module with deep supervision to learn effective high-frequency components of depth maps in a coarse-to-fine manner; 2) A channel attention strategy to enhance channels with abundant high-frequency components; 3) A multi-stage fusion module to effectively reexploit the results in the coarse-to-fine process; and 4) A depth refinement module to improve the depth map by TGV regularization and input loss. Extensive experiments on benchmarking datasets demonstrate the superiority of our method over current state-of-the-art DSR methods. Xibin Song, Yuchao Dai, Dingfu Zhou, Liu Liu 0009, Wei Li 0111, Hongdong Li, Ruigang Yang |
CVPR | 2 |
| 2020 | UC-Net: Uncertainty Inspired RGB-D Saliency Detection via Conditional Variational AutoencodersabstractIn this paper, we propose the first framework (UCNet) to employ uncertainty for RGB-D saliency detection by learning from the data labeling process. Existing RGB-D saliency detection methods treat the saliency detection task as a point estimation problem, and produce a single saliency map following a deterministic learning pipeline. Inspired by the saliency data labeling process, we propose probabilistic RGB-D saliency detection network via conditional variational autoencoders to model human annotation uncertainty and generate multiple saliency maps for each input image by sampling in the latent space. With the proposed saliency consensus process, we are able to generate an accurate saliency map based on these multiple predictions. Quantitative and qualitative evaluations on six challenging benchmark datasets against 18 competing algorithms demonstrate the effectiveness of our approach in learning the distribution of saliency maps, leading to a new state-of-the-art in RGB-D saliency detection. Jing Zhang 0052, Deng-Ping Fan, Yuchao Dai, Saeed Anwar, Fatemehsadat Saleh, Tong Zhang 0023, Nick Barnes |
CVPR | 3 |
| 2020 | Weakly-Supervised Salient Object Detection via Scribble AnnotationsabstractCompared with laborious pixel-wise dense labeling, it is much easier to label data by scribbles, which only costs 1~2 seconds to label one image. However, using scribble labels to learn salient object detection has not been explored. In this paper, we propose a weakly-supervised salient object detection model to learn saliency from such annotations. In doing so, we first relabel an existing large-scale salient object detection dataset with scribbles, namely S-DUTS dataset. Since object structure and detail information is not identified by scribbles, directly training with scribble labels will lead to saliency maps of poor boundary localization. To mitigate this problem, we propose an auxiliary edge detection task to localize object edges explicitly, and a gated structure-aware loss to place constraints on the scope of structure to be recovered. Moreover, we design a scribble boosting scheme to iteratively consolidate our scribble annotations, which are then employed as supervision to learn high-quality saliency maps. As existing saliency evaluation metrics neglect to measure structure alignment of the predictions, the saliency map ranking may not comply with human perception. We present a new metric, termed saliency structure measure, as a complementary metric to evaluate sharpness of the prediction. Extensive experiments on six benchmark datasets demonstrate that our method not only outperforms existing weakly-supervised/unsupervised methods, but also is on par with several fully-supervised state-of-the-art models (Our code and data is publicly available at: https://github.com/JingZhang617/Scribble_Saliency). Jing Zhang 0052, Xin Yu 0002, Aixuan Li, Peipei Song, Bowen Liu 0012, Yuchao Dai |
CVPR | 6 |
| 2020 | Joint 3D Instance Segmentation and Object Detection for Autonomous DrivingabstractCurrently, in Autonomous Driving (AD), most of the 3D object detection frameworks (either anchor- or anchor-free-based) consider the detection as a Bounding Box (BBox) regression problem. However, this compact representation is not sufficient to explore all the information of the objects. To tackle this problem, we propose a simple but practical detection framework to jointly predict the 3D BBox and instance segmentation. For instance segmentation, we propose a Spatial Embeddings (SEs) strategy to assemble all foreground points into their corresponding object centers. Base on the SE results, the object proposals can be generated based on a simple clustering strategy. For each cluster, only one proposal is generated. Therefore, the Non-Maximum Suppression (NMS) process is no longer needed here. Finally, with our proposed instance-aware ROI pooling, the BBox is refined by a second-stage network. Experimental results on the public KITTI dataset show that the proposed SEs can significantly improve the instance segmentation results compared with other feature embedding-based method. Meanwhile, it also outperforms most of the 3D object detectors on the KITTI testing benchmark. Dingfu Zhou, Xibin Song, Liu Liu 0009, Junbo Yin, Yuchao Dai, Hongdong Li, Ruigang Yang |
CVPR | 6 |
| 2020 | Relative Pose Estimation For Stereo Rolling Shutter CamerasabstractIn this paper, we present a novel linear algorithm to estimate the 6 DoF relative pose from consecutive frames of stereo rolling shutter (RS) cameras. Our method is derived based on the assumption that stereo cameras undergo motion with constant velocity around the center of the baseline, which needs 9 pairs of correspondences on both left and right consecutive frames. The stereo RS images enable the recovery of depth maps from the semi-global matching (SGM) algorithm. With the estimated camera motion and depth map, we can correct the RS images to get the undistorted images without any scene structure assumption. Experiments on both simulated points and synthetic RS images demonstrate the effectiveness of our algorithm in relative pose estimation. Bin Fan 0002, Yuchao Dai |
ICIP | 3 |
| 2020 | Novel View Synthesis from only a 6-DoF Camera Pose by Two-stage Networks
Bo Li 0090, Yuchao Dai, Tongxin Zhang |
ICPR | 3 |
| 2020 | Hierarchical Neural Architecture Search for Deep Stereo MatchingabstractTo reduce the human efforts in neural network design, Neural Architecture Search (NAS) has been applied with remarkable success to various high-level vision tasks such as classification and semantic segmentation. The underlying idea for the NAS algorithm is straightforward, namely, to allow the network the ability to choose among a set of operations (\eg convolution with different filter sizes), one is able to find an optimal architecture that is better adapted to the problem at hand. However, so far the success of NAS has not been enjoyed by low-level geometric vision tasks such as stereo matching. This is partly due to the fact that state-of-the-art deep stereo matching networks, designed by humans, are already sheer in size. Directly applying the NAS to such massive structures is computationally prohibitive based on the currently available mainstream computing resources. In this paper, we propose the first \emph{end-to-end} hierarchical NAS framework for deep stereo matching by incorporating task-specific human knowledge into the neural architecture search framework. Specifically, following the gold standard pipeline for deep stereo matching (\ie, feature extraction -- feature volume construction and dense matching), we optimize the architectures of the entire pipeline jointly. Extensive experiments show that our searched network outperforms all state-of-the-art deep stereo matching architectures and is ranked at the top 1 accuracy on KITTI stereo 2012, 2015, and Middlebury benchmarks, as well as the top 1 on SceneFlow dataset with a substantial improvement on the size of the network and the speed of inference. Code available at https://github.com/XuelianCheng/LEAStereo. Xuelian Cheng, Yiran Zhong, Mehrtash Harandi, Yuchao Dai, Xiaojun Chang, Hongdong Li, Tom Drummond, ZongYuan Ge |
NeurIPS | 4 |
| 2020 | Displacement-Invariant Matching Cost Learning for Accurate Optical Flow EstimationabstractLearning matching costs has been shown to be critical to the success of the state-of-the-art deep stereo matching methods, in which 3D convolutions are applied on a 4D feature volume to learn a 3D cost volume. However, this mechanism has never been employed for the optical flow task. This is mainly due to the significantly increased search dimension in the case of optical flow computation, \ie, a straightforward extension would require dense 4D convolutions in order to process a 5D feature volume, which is computationally prohibitive. This paper proposes a novel solution that is able to bypass the requirement of building a 5D feature volume while still allowing the network to learn suitable matching costs from data. Our key innovation is to decouple the connection between 2D displacements and learn the matching costs at each 2D displacement hypothesis independently, \ie, displacement-invariant cost learning. Specifically, we apply the same 2D convolution-based matching net independently on each 2D displacement hypothesis to learn a 4D cost volume. Moreover, we propose a displacement-aware projection layer to scale the learned cost volume, which reconsiders the correlation between different displacement candidates and mitigates the multi-modal problem in the learned cost volume. The cost volume is then projected to optical flow estimation through a 2D soft-argmin layer. Extensive experiments show that our approach achieves state-of-the-art accuracy on various datasets, and outperforms all published optical flow methods on the Sintel benchmark. The code is available at https://github.com/jytime/DICL-Flow. Yiran Zhong, Yuchao Dai, Kaihao Zhang, Pan Ji, Hongdong Li |
NeurIPS | 3 |
| 2020 | Hyperspectral Image Super-Resolution by Band Attention Through Adversarial LearningabstractHyperspectral image (HSI) super-resolution (SR) is a challenging task due to the problems of texture blur and spectral distortion when the upscaling factor is large. To meet these two challenges, band attention through the adversarial learning method is proposed in this article. First, we put the SR process in a generative adversarial network (GAN) framework, so that the resulted high-resolution HSI can keep more texture details. Second, different from the other band-by-band SR method, the input of our method is of full bands. In order to explore the correlation of spectral bands and avoid the spectral distortion, a band attention mechanism is proposed in our generative network. A series of spatial-spectral constraints or loss functions is imposed to guide the training of our generative network so as to further alleviate spectral distortion and texture blur. The experiments on the Pavia and Cave data sets demonstrate that the proposed GAN-based SR method can yield very high-quality results, even under large upscaling factor (e.g., $8\times $ ). More importantly, it can outperform the other state-of-the-art methods by a margin which demonstrates its superiority and effectiveness. Jiaojiao Li 0001, Ruxing Cui, Bo Li 0090, Rui Song 0003, Yunsong Li 0001, Yuchao Dai, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2020 | Joint Stereo Video Deblurring, Scene Flow Estimation and Moving Object SegmentationabstractStereo videos for the dynamic scenes often show unpleasant blurred effects due to the camera motion and the multiple moving objects with large depth variations. Given consecutive blurred stereo video frames, we aim to recover the latent clean images, estimate the 3D scene flow and segment the multiple moving objects. These three tasks have been previously addressed separately, which fail to exploit the internal connections among these tasks and cannot achieve optimality. In this paper, we propose to jointly solve these three tasks in a unified framework by exploiting their intrinsic connections. To this end, we represent the dynamic scenes with the piece-wise planar model, which exploits the local structure of the scene and expresses various dynamic scenes. Under our model, these three tasks are naturally connected and expressed as the parameter estimation of 3D scene structure and camera motion (structure and motion for the dynamic scenes). By exploiting the blur model constraint, the moving objects and the 3D scene structure, we reach an energy minimization formulation for joint deblurring, scene flow and segmentation. We evaluate our approach extensively on both synthetic datasets and publicly available real datasets with fast-moving objects, camera motion, uncontrolled lighting conditions and shadows. Experimental results demonstrate that our method can achieve significant improvement in stereo video deblurring, scene flow estimation and moving object segmentation, over state-of-the-art methods. Liyuan Pan, Yuchao Dai, Miaomiao Liu 0001, Fatih Porikli, Quan Pan 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Ground-Plane-Based Absolute Scale Estimation for Monocular Visual OdometryabstractRecovering an absolute metric scale from a monocular camera is a challenging but highly desirable problem for monocular camera-based systems. By using different kinds of cues, various approaches have been proposed for scale estimation, such as camera height and object size. In this paper, first, we summarize different kinds of scale estimation approaches. Then, we propose a robust divide-and-conquer absolute scale estimation method based on the ground plane and camera height by analyzing the advantages and disadvantages of different approaches. By using the estimated scale, an effective scale correction strategy has been proposed to reduce the scale drift during the monocular visual odometry estimation process. Finally, the effectiveness and robustness of the proposed method have been verified on both public and self-collected image sequences. Dingfu Zhou, Yuchao Dai, Hongdong Li |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2020 | Deep learning based point cloud registration: an overviewabstractPoint cloud registration aims at finding a rigid transformation to align one point cloud to another one. It is a fundamental problem in computer vision and robotics, which has been widely used in various applications, such as 3D reconstruction, SLAM (simultaneous localization and mapping), and autonomous driving. Over the last decades, many researchers have devoted themselves to tackle this challenging problem. Recently, the success of deep learning in high-level vision tasks has been extended to different geometric vision tasks. Various kinds of deep learning based point cloud registration methods have been proposed to exploit different aspects of the problem. However, a comprehensive overview of these approaches is still missing. To this end, in this paper, we summarize recent progress and present a comprehensive overview for deep learning based point cloud registration. We classify the popular approaches into different categories such as, correspondences-based or correspondences-free, effective modules: feature extractor, matching, outlier rejection, and motion estimation. Furthermore, we discuss the merits and demerits in detail. We provide a systematic and compact framework towards currently proposed methods and discuss future research directions. Zhiyuan Zhang 0002, Yuchao Dai, Jiadai Sun |
Virtual Real. Intell. Hardw. | 2 |
| 2019 | MVS2: Deep Unsupervised Multi-View Stereo with Multi-View SymmetryabstractThe success of existing deep-learning based multi-view stereo (MVS) approaches greatly depends on the availability of large-scale supervision in the form of dense depth maps. Such supervision, while not always possible, tends to hinder the generalization ability of the learned models in never-seen-before scenarios. In this paper, we propose the first unsupervised learning based MVS network, which learns the multi-view depth maps from the input multi-view images and does not need ground-truth 3D training data. Our network is symmetric in predicting depth maps for all views simultaneously, where we enforce cross-view consistency of multi-view depth maps during both training and testing stages. Thus, the learned multi-view depth maps naturally comply with the underlying 3D scene geometry. Besides, our network also learns the multi-view occlusion maps, which further improves the robustness of our network in handling real-world occlusions. Experimental results on multiple benchmarking datasets demonstrate the effectiveness of our network and the excellent generalization ability. Yuchao Dai, Zhidong Zhu, Zhibo Rao, Bo Li 0090 |
3DV | 1 |
| 2019 | IoU Loss for 2D/3D Object DetectionabstractIn the 2D/3D object detection task, Intersection-over-Union (IoU) has been widely employed as an evaluation metric to evaluate the performance of different detectors in the testing stage. However, during the training stage, the common distance loss (e.g, L_1 or L_2) is often adopted as the loss function to minimize the discrepancy between the predicted and ground truth Bounding Box (Bbox). To eliminate the performance gap between training and testing, the IoU loss has been introduced for 2D object detection in [1] and [2]. Unfortunately, all these approaches only work for axis-aligned 2D Boxes, which cannot be applied for more general object detection task with rotated Boxes. To resolve this issue, we investigate the IoU computation for two rotated Boxes first and then implement a unified framework, IoU loss layer for both 2D and 3D object detection tasks. By integrating the implemented IoU loss into several state-of-the-art 3D object detectors, consistent improvements have been achieved for both bird-eye-view 2D detection and point cloud 3D detection on the public KITTI [3] benchmark. Dingfu Zhou, Xibin Song, Chenye Guan, Junbo Yin, Yuchao Dai, Ruigang Yang |
3DV | 6 |
| 2019 | Noise-Aware Unsupervised Deep Lidar-Stereo FusionabstractIn this paper, we present LidarStereoNet, the first unsupervised Lidar-stereo fusion network, which can be trained in an end-to-end manner without the need of ground truth depth maps. By introducing a novel ``Feedback Loop'' to connect the network input with output, LidarStereoNet could tackle both noisy Lidar points and misalignment between sensors that have been ignored in existing Lidar-stereo fusion work. Besides, we propose to incorporate the piecewise planar model into the network learning to further constrain depths to conform to the underlying 3D geometry. Extensive quantitative and qualitative evaluations on both real and synthetic datasets demonstrate the superiority of our method, which outperforms state-of-the-art stereo matching, depth completion and Lidar-Stereo fusion approaches significantly. Xuelian Cheng, Yiran Zhong, Yuchao Dai, Pan Ji, Hongdong Li |
CVPR | 3 |
| 2019 | Phase-Only Image Based Kernel Estimation for Single Image Blind DeblurringabstractThe image motion blurring process is generally modelled as the convolution of a blur kernel with a latent image. Therefore, the estimation of the blur kernel is essentially important for blind image deblurring. Unlike existing approaches which focus on approaching the problem by enforcing various priors on the blur kernel and the latent image, we are aiming at obtaining a high quality blur kernel directly by studying the problem in the frequency domain. We show that the auto-correlation of the absolute phase-only image 1 can provide faithful information about the motion (e.g., the motion direction and magnitude, we call it the motion pattern in this paper.) that caused the blur, leading to a new and efficient blur kernel estimation approach. The blur kernel is then refined and the sharp image is estimated by solving an optimization problem by enforcing a regularization on the blur kernel and the latent image. We further extend our approach to handle non-uniform blur, which involves spatially varying blur kernels. Our approach is evaluated extensively on synthetic and real data and shows good results compared to the state-of-the-art deblurring approaches. Liyuan Pan, Richard I. Hartley, Miaomiao Liu 0001, Yuchao Dai |
CVPR | 4 |
| 2019 | Bringing a Blurry Frame Alive at High Frame-Rate With an Event CameraabstractEvent-based cameras can measure intensity changes (called ‘events’) with microsecond accuracy under high-speed motion and challenging lighting conditions. With the active pixel sensor (APS), the event camera allows simultaneous output of the intensity frames. However, the output images are captured at a relatively low frame-rate and often suffer from motion blur. A blurry image can be regarded as the integral of a sequence of latent images, while the events indicate the changes between the latent images. Therefore, we are able to model the blur-generation process by associating event data to a latent image. In this paper, we propose a simple and effective approach, the Event-based Double Integral (EDI) model, to reconstruct a high frame-rate, sharp video from a single blurry frame and its event data. The video generation is based on solving a simple non-convex optimization problem in a single scalar variable. Experimental results on both synthetic and real images demonstrate the superiority of our EDI model and optimization method in comparison to the state-of-the-art. Liyuan Pan, Cedric Scheerlinck, Xin Yu 0002, Richard I. Hartley, Miaomiao Liu 0001, Yuchao Dai |
CVPR | 6 |
| 2019 | ApolloCar3D: A Large 3D Car Instance Understanding Benchmark for Autonomous DrivingabstractAutonomous driving has attracted remarkable attention from both industry and academia. An important task is to estimate 3D properties (e.g. translation, rotation and shape) of a moving or parked vehicle on the road. This task, while critical, is still under-researched in the computer vision community – partially owing to the lack of large scale and fully-annotated 3D car database suitable for autonomous driving research. In this paper, we contribute the first large scale database suitable for 3D car instance understanding – ApolloCar3D. The dataset contains 5,277 driving images and over 60K car instances, where each car is fitted with an industry-grade 3D CAD model with absolute model size and semantically labelled keypoints. This dataset is above 20× larger than PASCAL3D+ and KITTI, the current state-of-the-art. To enable efficient labelling in 3D, we build a pipeline by considering 2D-3D keypoint correspondences for a single instance and 3D relationship among multiple instances. Equipped with such dataset, we build various baseline algorithms with the state-of-the-art deep convolutional neural networks. Specifically, we first segment each car with a pre-trained Mask R-CNN, and then regress towards its 3D pose and shape based on a deformable 3D car model with or without using semantic keypoints. We show that using keypoints significantly improves fitting performance. Finally, we develop a new 3D metric jointly considering 3D pose and 3D shape, allowing for comprehensive evaluation and ablation study. Xibin Song, Peng Wang 0001, Dingfu Zhou, Chenye Guan, Yuchao Dai, Hongdong Li, Ruigang Yang |
CVPR | 6 |
| 2019 | Deep Stacked Hierarchical Multi-Patch Network for Image DeblurringabstractDespite deep end-to-end learning methods have shown their superiority in removing non-uniform motion blur, there still exist major challenges with the current multi-scale and scale-recurrent models: 1) Deconvolution/upsampling operations in the coarse-to-fine scheme result in expensive runtime; 2) Simply increasing the model depth with finer-scale levels cannot improve the quality of deblurring. To tackle the above problems, we present a deep {hierarchical multi-patch network} inspired by Spatial Pyramid Matching to deal with blurry images via a fine-to-coarse hierarchical representation. To deal with the performance saturation w.r.t. depth, we propose a stacked version of our multi-patch model. Our proposed basic multi-patch model achieves the state-of-the-art performance on the GoPro dataset while enjoying a 40$\times$ faster runtime compared to current multi-scale methods. With 30ms to process an image at 1280$\times$720 resolution, it is the first real-time deep motion deblurring model for 720p images at 30fps. For stacked networks, significant improvements (over 1.2dB) are achieved on the GoPro dataset by increasing the network depth. Moreover, by varying the depth of the stacked model, one can adapt the performance and runtime of the same network for different application scenarios. Yuchao Dai, Hongdong Li, Piotr Koniusz |
CVPR | 2 |
| 2019 | Unsupervised Deep Epipolar Flow for Stationary or Dynamic ScenesabstractUnsupervised deep learning for optical flow computation has achieved promising results. Most existing deep-net based methods rely on image brightness consistency and local smoothness constraint to train the networks. Their performance degrades at regions where repetitive textures or occlusions occur. In this paper, we propose Deep Epipolar Flow, an unsupervised optical flow method which incorporates global geometric constraints into network learning. In particular, we investigate multiple ways of enforcing the epipolar constraint in flow estimation. To alleviate a ``chicken-and-egg'' type of problem encountered in dynamic scenes where multiple motions may be present, we propose a low-rank constraint as well as a union-of-subspaces constraint for training. Experimental results on various benchmarking datasets show that our method achieves competitive performance compared with supervised methods and outperforms state-of-the-art unsupervised deep-learning methods. Yiran Zhong, Pan Ji, Yuchao Dai, Hongdong Li |
CVPR | 4 |
| 2019 | Stochastic Attraction-Repulsion Embedding for Large Scale Image LocalizationabstractThis paper tackles the problem of large-scale image-based localization (IBL) where the spatial location of a query image is determined by finding out the most similar reference images in a large database. For solving this problem, a critical task is to learn discriminative image representation that captures informative information relevant for localization. We propose a novel representation learning method having higher location-discriminating power. It provides the following contributions: 1) we represent a place (location) as a set of exemplar images depicting the same landmarks and aim to maximize similarities among intra-place images while minimizing similarities among inter-place images; 2) we model a similarity measure as a probability distribution on L2-metric distances between intra-place and inter-place image representations; 3) we propose a new Stochastic Attraction and Repulsion Embedding (SARE) loss function minimizing the KL divergence between the learned and the actual probability distributions; 4) we give theoretical comparisons between SARE, triplet ranking and contrastive losses. It provides insights into why SARE is better by analyzing gradients. Our SARE loss is easy to implement and pluggable to any CNN. Experiments show that our proposed method improves the localization performance on standard benchmarks by a large margin. Demonstrating the broad applicability of our method, we obtained the third place out of 209 teams in the 2018 Google Landmark Retrieval Challenge. Our code and model are available at https://github.com/Liumouliu/deepIBL. Liu Liu 0009, Hongdong Li, Yuchao Dai |
ICCV | 3 |
| 2019 | Single Image Deblurring and Camera Motion Estimation With Depth MapabstractCamera shake during exposure is a major problem in hand-held photography, as it causes image blur that destroys details in the captured images. In the real world, such blur is mainly caused by both the camera motion and the complex scene structure. While considerable existing approaches have been proposed based on various assumptions regarding the scene structure or the camera motion, few existing methods could handle the real 6 DoF camera motion. In this paper, we propose to jointly estimate the 6 DoF camera motion and remove the non-uniform blur caused by camera motion by exploiting their underlying geometric relationships, with a single blurry image and its depth map (either direct depth measurements, or a learned depth map) as input. We formulate our joint deblurring and 6 DoF camera motion estimation as an energy minimization problem which is solved in an alternative manner. Our model enables the recovery of the 6 DoF camera motion and the latent clean image, which could also achieve the goal of generating a sharp sequence from a single blurry image. Experiments on challenging real-world and synthetic datasets demonstrate that image blur from camera shake can be well addressed within our proposed framework. Liyuan Pan, Yuchao Dai, Miaomiao Liu 0001 |
WACV | 2 |
| 2019 | Deeply Supervised Depth Map Super-Resolution as Novel View SynthesisabstractDeep convolutional neural network (DCNN) has been successfully applied to depth map super-resolution and outperforms existing methods by a wide margin. However, there still exist two major issues with these DCNN-based depth map super-resolution methods that hinder the performance: 1) the low-resolution depth maps either need to be up-sampled before feeding into the network or substantial deconvolution has to be used and 2) the supervision (high-resolution depth maps) is only applied at the end of the network, thus it is difficult to handle large up-sampling factors, such as x8 and x16. In this paper, we propose a new framework to tackle the above problems. First, we propose to represent the task of depth map superresolution as a series of novel view synthesis sub-tasks. The novel view synthesis sub-task aims at generating (synthesizing) a depth map from a different camera pose, which could be learned in parallel. Second, to handle large up-sampling factors, we present a deeply supervised network structure to enforce strong supervision in each stage of the network. Third, a multiscale fusion strategy is proposed to effectively exploit the feature maps at different scales and handle the blocking effect. In this way, our proposed framework could deal with challenging depth map super-resolution efficiently under large up-sampling factors (e.g., x8 and x16). Our method only uses the low-resolution depth map as input, and the support of color image is not needed, which greatly reduces the restriction of our method. Extensive experiments on various benchmarking data sets demonstrate the superiority of our method over current state-of-the-art depth map super-resolution methods. Xibin Song, Yuchao Dai, Xueying Qin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Scalable Dense Non-Rigid Structure-From-Motion: A Grassmannian PerspectiveabstractThis paper addresses the task of dense non-rigid structure-front-motion (NRSfM) using multiple images. State-of-the-art methods to this problem are often hurdled by scalability, expensive computations, and noisy measurements. Further, recent methods to NRSfM usually either assume a small number of sparse feature points or ignore local non-linearities of shape deformations, and thus cannot reliably model complex non-rigid deformations. To address these issues, in this paper, we propose a new approach for dense NRSfM by modeling the problem on a Grassmann manifold. Specifically, we assume the complex non-rigid deformations lie on a union of local linear subspaces both spatially and temporally. This naturally allows for a compact representation of the complex non-rigid deformation over frames. We provide experimental results on several synthetic and real benchmark datasets. The procured results clearly demonstrate that our method, apart from being scalable and more accurate than state-of-the-art methods, is also more robust to noise and generalizes to highly nonlinear deformations. Suryansh Kumar 0001, Anoop Cherian, Yuchao Dai, Hongdong Li |
CVPR | 3 |
| 2018 | Deep Unsupervised Saliency Detection: A Multiple Noisy Labeling PerspectiveabstractThe success of current deep saliency detection methods heavily depends on the availability of large-scale supervision in the form of per-pixel labeling. Such supervision, while labor-intensive and not always possible, tends to hinder the generalization ability of the learned models. By contrast, traditional handcrafted features based unsupervised saliency detection methods, even though have been surpassed by the deep supervised methods, are generally dataset-independent and could be applied in the wild. This raises a natural question that "Is it possible to learn saliency maps without using labeled data while improving the generalization ability?". To this end, we present a novel perspective to unsupervised saliency detection through learning from multiple noisy labeling generated by "weak" and "noisy" unsupervised handcrafted saliency methods. Our end-to-end deep learning framework for unsupervised saliency detection consists of a latent saliency prediction module and a noise modeling module that work collaboratively and are optimized jointly. Explicit noise modeling enables us to deal with noisy saliency maps in a probabilistic way. Extensive experimental results on various benchmarking datasets show that our model not only outperforms all the unsupervised saliency methods with a large margin but also achieves comparable performance with the recent state-of-the-art supervised deep saliency methods. Jing Zhang 0052, Tong Zhang 0023, Yuchao Dai, Mehrtash Harandi, Richard I. Hartley |
CVPR | 3 |
| 2018 | Stereo Computation for a Single Mixture Image
Yiran Zhong, Yuchao Dai, Hongdong Li |
ECCV (9) | 2 |
| 2018 | Open-World Stereo Video Matching with Deep RNN
Yiran Zhong, Hongdong Li, Yuchao Dai |
ECCV (2) | 3 |
| 2018 | Occluded Joints Recovery in 3D Human Pose Estimation based on Distance MatrixabstractAlbeit the recent progress in single image 3D human pose estimation due to the convolutional neural network, it is still challenging to handle real scenarios such as highly occluded scenes. In this paper, we propose to address the problem of single image 3D human pose estimation with occluded measurements by exploiting the Euclidean distance matrix (EDM). Specifically, we present two approaches based on EDM, which could effectively handle occluded joints in 2D images. The first approach is based on 2D-to-2D distance matrix regression achieved by a simple CNN architecture. The second approach is based on sparse coding along with a learned over-complete dictionary. Experiments on the Human3.6M dataset show the excellent performance of these two approaches in recovering occluded observations and demonstrate the improvements in accuracy for 3D human pose estimation with occluded joints.1 Yuchao Dai |
ICPR | 2 |
| 2018 | 3D Geometry-Aware Semantic Labeling of Outdoor Street ScenesabstractThis paper is concerned with the problem of how to better exploit 3D geometric information for dense semantic image labeling. Existing methods often treat the available 3D geometry information (e.g., 3D depth-map) simply as an additional image channel besides the R-G-B color channels, and apply the same technique for RGB image labeling. In this paper, we demonstrate that directly performing 3D convolution in the framework of a residual connected 3D voxel top-down modulation network can lead to superior results. Specifically, we propose a 3D semantic labeling method to label outdoor street scenes whenever a dense depth map is available. Experiments on the “Synthia” and “Cityscape” datasets show our method outperforms the state-of-the-art methods, suggesting such a simple 3D representation is effective in incorporating 3D geometric information. Yiran Zhong, Yuchao Dai, Hongdong Li |
ICPR | 2 |
| 2018 | Depth Map Completion by Jointly Exploiting Blurry Color Images and Sparse Depth MapsabstractWe aim at predicting a complete and high-resolution depth map from incomplete, sparse and noisy depth measurements. Existing methods handle this problem either by exploiting various regularizations on the depth maps directly or resorting to learning based methods. When the corresponding color images are available, the correlation between the depth maps and the color images are used to improve the completion performance, assuming the color images are clean and sharp. However, in real world dynamic scenes, color images are often blurry due to the camera motion and the moving objects in the scene. In this paper, we propose to tackle the problem of depth map completion by jointly exploiting the blurry color image sequences and the sparse depth map measurements, and present an energy minimization based formulation to simultaneously complete the depth maps, estimate the scene flow and deblur the color images. Our experimental evaluations on both outdoor and indoor scenarios demonstrate the state-of-the-art performance of our approach. Liyuan Pan, Yuchao Dai, Miaomiao Liu 0001, Fatih Porikli |
WACV | 2 |
| 2018 | 3D skeleton based action recognition by video-domain translation-scale invariant mapping and multi-scale dilated CNN
Bo Li 0090, Mingyi He, Yuchao Dai, Xuelian Cheng |
Multim. Tools Appl. | 3 |
| 2018 | Monocular depth estimation with hierarchical fusion of dilated CNNs and soft-weighted-sum inference
Bo Li 0090, Yuchao Dai, Mingyi He |
Pattern Recognit. | 2 |
| 2018 | Robust and Efficient Relative Pose With a Multi-Camera System for Autonomous Driving in Highly Dynamic EnvironmentsabstractThis paper studies the relative pose problem for autonomous vehicles driving in highly dynamic and possibly cluttered environments. This is a challenging scenario due to the existence of multiple, large, and independently moving objects in the environment, which often leads to an excessive portion of outliers and results in erroneous motion estimation. Existing algorithms cannot cope with such situations well. This paper proposes a new algorithm for relative pose estimation using a multi-camera system with multiple non-overlapping cameras. The method works robustly even when the number of outliers is overwhelming. By exploiting specific prior knowledge of the autonomous driving scene, we have developed an efficient 4-point algorithm for multi-camera relative pose estimation, which admits analytic solutions by solving a polynomial root finding equation, and runs extremely fast (at about 0.5 μs per root). When the solver is used in combination with a new random sample consensus sampling scheme by exploiting the conjugate motion constraint, we are able to quickly prune unpromising hypotheses and significantly improve the chance of finding inliers. Experiments on synthetic data have validated the performance of the proposed algorithm. Tests on real data further confirm the method's practical relevance. Liu Liu 0009, Hongdong Li, Yuchao Dai, Quan Pan 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2017 | Simultaneous Stereo Video Deblurring and Scene Flow EstimationabstractVideos for outdoor scene often show unpleasant blur effects due to the large relative motion between the camera and the dynamic objects and large depth variations. Existing works typically focus monocular video deblurring. In this paper, we propose a novel approach to deblurring from stereo videos. In particular, we exploit the piece-wise planar assumption about the scene and leverage the scene flow information to deblur the image. Unlike the existing approach [31] which used a pre-computed scene flow, we propose a single framework to jointly estimate the scene flow and deblur the image, where the motion cues from scene flow estimation and blur information could reinforce each other, and produce superior results than the conventional scene flow estimation or stereo deblurring methods. We evaluate our method extensively on two available datasets and achieve significant improvement in flow estimation and removing the blur effect over the state-of-the-art methods. Liyuan Pan, Yuchao Dai, Miaomiao Liu 0001, Fatih Porikli |
CVPR | 2 |
| 2017 | "Maximizing Rigidity" Revisited: A Convex Programming Approach for Generic 3D Shape Reconstruction from Multiple Perspective ViewsabstractRigid structure-from-motion (RSfM) and non-rigid structure-from-motion (NRSfM) have long been treated in the literature as separate (different) problems. Inspired by a previous work which solved directly for 3D scene structure by factoring the relative camera poses out, we revisit the principle of “maximizing rigidity” in structure-from-motion literature, and develop a unified theory which is applicable to both rigid and non-rigid structure reconstruction in a rigidity-agnostic way. We formulate these problems as a convex semi-definite program, imposing constraints that seek to apply the principle of minimizing non-rigidity. Our results demonstrate the efficacy of the approach, with stateof- the-art accuracy on various 3D reconstruction problems. Pan Ji, Hongdong Li, Yuchao Dai, Ian D. Reid 0001 |
ICCV | 3 |
| 2017 | Monocular Dense 3D Reconstruction of a Complex Dynamic Scene from Two Perspective FramesabstractThis paper proposes a new approach for monocular dense 3D reconstruction of a complex dynamic scene from two perspective frames. By applying superpixel over-segmentation to the image, we model a generically dynamic (hence non-rigid) scene with a piecewise planar and rigid approximation. In this way, we reduce the dynamic reconstruction problem to a “3D jigsaw puzzle ” problem which takes pieces from an unorganized “soup of superpixels". We show that our method provides an effective solution to the inherent relative scale ambiguity in structure-from-motion. Since our method does not assume a template prior, or per-object segmentation, or knowledge about the rigidity of the dynamic scene, it is applicable to a wide range of scenarios. Extensive experiments on both synthetic and real monocular sequences demonstrate the superiority of our method compared with the state-of-the-art methods. Suryansh Kumar 0001, Yuchao Dai, Hongdong Li |
ICCV | 2 |
| 2017 | Efficient Global 2D-3D Matching for Camera Localization in a Large-Scale 3D MapabstractGiven an image of a street scene in a city, this paper develops a new method that can quickly and precisely pinpoint at which location (as well as viewing direction) the image was taken, against a pre-stored large-scale 3D point-cloud map of the city. We adopt the recently developed 2D-3D direct feature matching framework for this task [23,31,32,42-44]. This is a challenging task especially for large-scale problems. As the map size grows bigger, many 3D points in the wider geographical area can be visually very similar-or even identical-causing severe ambiguities in 2D-3D feature matching. The key is to quickly and unambiguously find the correct matches between a query image and the large 3D map. Existing methods solve this problem mainly via comparing individual features' visual similarities in a local and per feature manner, thus only local solutions can be found, inadequate for large-scale applications. In this paper, we introduce a global method which harnesses global contextual information exhibited both within the query image and among all the 3D points in the map. This is achieved by a novel global ranking algorithm, applied to a Markov network built upon the 3D map, which takes account of not only visual similarities between individual 2D-3D matches, but also their global compatibilities (as measured by co-visibility) among all matching pairs found in the scene. Tests on standard benchmark datasets show that our method achieved both higher precision and comparable recall, compared with the state-of-the-art. Liu Liu 0009, Hongdong Li, Yuchao Dai |
ICCV | 3 |
| 2017 | Dense non-rigid structure-from-motion made easy - A spatial-temporal smoothness based solutionabstractThis paper proposes a simple spatial-temporal smoothness based method for solving dense non-rigid structure-frommotion (NRSfM). First, we revisit the temporal smoothness and demonstrate that it can be extended to dense case directly. Second, we propose to exploit the spatial smoothness by resorting to the Laplacian of the 3D non-rigid shape. Third, to handle real world noise and outliers in measurements, we robustify the data term by using the L1norm. In this way, our method could robustly exploit both spatial and temporal smoothness effectively and make dense non-rigid reconstruction easy. Our method is very easy to implement, which involves solving a series of least squares problems. Experimental results on both synthetic and real image dense NRSfM tasks show that the proposed method outperforms state-of-the-art dense non-rigid reconstruction methods. Yuchao Dai, Huizhong Deng, Mingyi He |
ICIP | 1 |
| 2017 | Integrated deep and shallow networks for salient object detectionabstractDeep convolutional neural network (CNN) based salient object detection methods have achieved state-of-the-art performance and outperform those unsupervised methods with a wide margin. In this paper, we propose to integrate deep and unsupervised saliency for salient object detection under a unified framework. Specifically, our method takes results of unsupervised saliency (Robust Background Detection, RBD) and normalized color images as inputs, and directly learns an end-to-end mapping between inputs and the corresponding saliency maps. The color images are fed into a Fully Convolutional Neural Networks (FCNN) adapted from semantic segmentation to exploit high-level semantic cues for salient object detection. Then the results from deep FCNN and RBD are concatenated to feed into a shallow network to map the concatenated feature maps to saliency maps. Finally, to obtain a spatially consistent saliency map with sharp object boundaries, we fuse superpixel level saliency map at multi-scale. Extensive experimental results on 8 benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art approaches with a margin. Jing Zhang 0052, Bo Li 0090, Yuchao Dai, Fatih Porikli, Mingyi He |
ICIP | 3 |
| 2017 | Accurate extrinsic calibration between monocular camera and sparse 3D Lidar points without markersabstractIt is of practical interest to automatically calibrate the multiple sensors in autonomous vehicles. In this paper, we deal with an interesting case when used low-resolution Lidar and present a practical approach to extrinsic calibration between monocular camera and Lidar with sparse 3D measurements. We formulate the problem as directly minimizing the feature error evaluated between frames following the way of image warping. To overcome the difficulties in the optimization problem, we propose to use the distance transform and further projection error model to obtain the key approximated edge points that are sensitive to the loss function. Finally, the loss minimization is solved by an efficient random selection algorithm. Experimental results on KITTI dataset show that our proposed method can achieve competitive results and an improvement in translation estimation particularly. Zhipeng Xiao, Hongdong Li, Dingfu Zhou, Yuchao Dai, Bin Dai 0001 |
Intelligent Vehicles Symposium | 4 |
| 2017 | Deep Salient Object Detection by Integrating Multi-level CuesabstractA key problem in salient object detection is how to effectively exploit the multi-level saliency cues in a unified and data-driven manner. In this paper, building upon the recent success of deep neural networks, we propose a fully convolutional neural network based approach empowered with multi-level fusion to salient object detection. By integrating saliency cues at different levels through fully convolutional neural networks and multi-level fusion, our approach could effectively exploit both learned semantic cues and higher-order region statistics for edge-accurate salient object detection. First, we fine-tune a fully convolutional neural network for semantic segmentation to adapt it to salient object detection to learn a suitable yet coarse perpixel saliency prediction map. This map is often smeared across salient object boundaries since the local receptive fields in the convolutional network apply naturally on both sides of such boundaries. Second, to enhance the resolution of the learned saliency prediction and to incorporate higher-order cues that are omitted by the neural network, we propose a multi-level fusion approach where super-pixel level coherency in saliency is exploited. Our extensive experimental results on various benchmark datasets demonstrate that the proposed method outperforms the state-of the-art approaches. Jing Zhang 0052, Yuchao Dai, Fatih Porikli |
WACV | 2 |
| 2017 | Moving object detection and segmentation in urban environments from a moving platform
Dingfu Zhou, Vincent Frémont, Benjamin Quost, Yuchao Dai, Hongdong Li |
Image Vis. Comput. | 4 |
| 2017 | Spatio-temporal union of subspaces for multi-body non-rigid structure-from-motion
Suryansh Kumar 0001, Yuchao Dai, Hongdong Li |
Pattern Recognit. | 2 |
| 2016 | Multi-Body Non-Rigid Structure-from-MotionabstractIn this paper, we present the first multi-body non-rigid structure-from-motion (SFM) method, which simultaneously reconstructs and segments multiple objects that are undergoing non-rigid deformation over time. Under our formulation, 3D trajectories for each non-rigid object can be well approximated with a sparse affine combination of other 3D trajectories from the same object. The resultant optimization is solved by the alternating direction method of multipliers (ADMM). We demonstrate the efficacy of the proposed method through extensive experiments on both synthetic and real data sequences. Our method outperforms other alternative methods, such as first clustering the 2D feature tracks to groups and then doing non-rigid reconstruction in each group or first conducting 3D reconstruction by using single subspace assumption and then clustering the 3D trajectories into groups. Suryansh Kumar 0001, Yuchao Dai, Hongdong Li |
3DV | 2 |
| 2016 | Deep Depth Super-Resolution: Learning Depth Super-Resolution Using Deep Convolutional Neural Network
Xibin Song, Yuchao Dai, Xueying Qin |
ACCV (4) | 2 |
| 2016 | Rolling Shutter Camera Relative Pose: Generalized Epipolar GeometryabstractThe vast majority of modern consumer-grade cameras employ a rolling shutter mechanism. In dynamic geometric computer vision applications such as visual SLAM, the so-called rolling shutter effect therefore needs to be properly taken into account. A dedicated relative pose solver appears to be the first problem to solve, as it is of eminent importance to bootstrap any derivation of multi-view geometry. However, despite its significance, it has received inadequate attention to date. This paper presents a detailed investigation of the geometry of the rolling shutter relative pose problem. We introduce the rolling shutter essential matrix, and establish its link to existing models such as the push-broom cameras, summarized in a clean hierarchy of multi-perspective cameras. The generalization of well-established concepts from epipolar geometry is completed by a definition of the Sampson distance in the rolling shutter case. The work is concluded with a careful investigation of the introduced epipolar geometry for rolling shutter cameras on several dedicated benchmarks. Yuchao Dai, Hongdong Li, Laurent Kneip |
CVPR | 1 |
| 2016 | Robust Optical Flow Estimation of Double-Layer Images under Transparency or ReflectionabstractThis paper deals with a challenging, frequently encountered, yet not properly investigated problem in two-frame optical flow estimation. That is, the input frames are compounds of two imaging layers - one desired background layer of the scene, and one distracting, possibly moving layer due to transparency or reflection. In this situation, the conventional brightness constancy constraint - the cornerstone of most existing optical flow methods - will no longer be valid. In this paper, we propose a robust solution to this problem. The proposed method performs both optical flow estimation, and image layer separation. It exploits a generalized double-layer brightness consistency constraint connecting these two tasks, and utilizes the priors for both of them. Experiments on both synthetic data and real images have confirmed the efficacy of the proposed method. To the best of our knowledge, this is the first attempt towards handling generic optical flow fields of two-frame images containing transparency or reflection. Jiaolong Yang, Hongdong Li, Yuchao Dai, Robby T. Tan |
CVPR | 3 |
| 2016 | Pushing the limit of non-rigid structure-from-motion by shape clusteringabstractRecovering both camera motions and non-rigid 3D shapes from 2D feature tracks is a challenging problem in computer vision. Long-term, complex non-rigid shape variations in real world videos further increase the difficulty for Non-rigid structure-from-motion (NRSfM). Furthermore, there does not exist a criterion to characterize the possibility in recovering the non-rigid shapes and camera motions (i.e., how easy or how difficult the problem could be). In this paper, we first present an analysis to the "reconstructability" measure for NRSfM, where we show that 3D shape complexity and camera motion complexity can be used to index the re-constructability. We propose an iterative shape clustering based method to NRSfM, which alternates between 3D shape clustering and 3D shape reconstruction. Thus, the global reconstructability has been improved and better reconstruction can be achieved. Experimental results on long-term, complex non-rigid motion sequences show that our method outperforms the current state-of-the-art methods by a margin. Huizhong Deng, Yuchao Dai |
ICASSP | 2 |
| 2016 | Reliable scale estimation and correction for monocular Visual OdometryabstractRecovering absolute scale (i.e. metric information) from monocular vision system is a very challenging problem yet is highly desirable for vision-based autonomous driving. This paper proposes a new method for scale recovery, based on the idea of knowing camera height (relative to ground-plane). While this idea of using known camera height is not new in this context, existing implementations of this idea suffer significantly from severe numerical instability arisen in the ground plane homography decomposition stage. Our novel contribution of this work is to alleviate this issue by a divide and conquer approach, i.e. decomposing the motion parameters in the homography from the structure parameters of the ground plane. We also describe a robust procedure to correct scale drift in the monocular visual odometry system. Experimental results on KITTI standard benchmark dataset [1] and our self-collected driving dataset both show significant improvements. Dingfu Zhou, Yuchao Dai, Hongdong Li |
Intelligent Vehicles Symposium | 2 |
| 2015 | Depth and surface normal estimation from monocular images using regression on deep features and hierarchical CRFsabstractPredicting the depth (or surface normal) of a scene from single monocular color images is a challenging task. This paper tackles this challenging and essentially underdetermined problem by regression on deep convolutional neural network (DCNN) features, combined with a post-processing refining step using conditional random fields (CRF). Our framework works at two levels, super-pixel level and pixel level. First, we design a DCNN model to learn the mapping from multi-scale image patches to depth or surface normal values at the super-pixel level. Second, the estimated super-pixel depth or surface normal is refined to the pixel level by exploiting various potentials on the depth or surface normal map, which includes a data term, a smoothness term among super-pixels and an auto-regression term characterizing the local structure of the estimation map. The inference problem can be efficiently solved because it admits a closed-form solution. Experiments on the Make3D and NYU Depth V2 datasets show competitive results compared with recent state-of-the-art methods. Bo Li 0090, Chunhua Shen, Yuchao Dai, Anton van den Hengel, Mingyi He |
CVPR | 3 |
| 2014 | Robust Motion Segmentation with Unknown Correspondences
Pan Ji, Hongdong Li, Mathieu Salzmann, Yuchao Dai |
ECCV (6) | 4 |
| 2014 | A Simple Prior-Free Method for Non-rigid Structure-from-Motion Factorization
Yuchao Dai, Hongdong Li, Mingyi He |
Int. J. Comput. Vis. | 1 |
| 2013 | Multi-view 3D Reconstruction from Uncalibrated Radially-Symmetric CamerasabstractWe present a new multi-view 3D Euclidean reconstruction method for arbitrary uncalibrated radially-symmetric cameras, which needs no calibration or any camera model parameters other than radial symmetry. It is built on the radial 1D camera model [25], a unified mathematical abstraction to different types of radially-symmetric cameras. We formulate the problem of multi-view reconstruction for radial 1D cameras as a matrix rank minimization problem. Efficient implementation based on alternating direction continuation is proposed to handle scalability issue for real-world applications. Our method applies to a wide range of omni directional cameras including both dioptric and catadioptric (central and non-central) cameras. Additionally, our method deals with complete and incomplete measurements under a unified framework elegantly. Experiments on both synthetic and real images from various types of cameras validate the superior performance of our new method, in terms of numerical accuracy and robustness. Jae-Hak Kim, Yuchao Dai, Hongdong Li, Jonghyuk Kim |
ICCV | 2 |
| 2013 | Single-shot extrinsic calibration of a generically configured RGB-D camera rig from scene constraintsabstractWith the increasing use of commodity RGB-D cameras for computer vision, robotics, mixed and augmented reality and other areas, it is of significant practical interest to calibrate the relative pose between a depth (D) camera and an RGB camera in these types of setups. In this paper, we propose a new single-shot, correspondence-free method to extrinsically calibrate a generically configured RGB-D camera rig. We formulate the extrinsic calibration problem as one of geometric 2D-3D registration which exploits scene constraints to achieve single-shot extrinsic calibration. Our method first reconstructs sparse point clouds from a single-view 2D image. These sparse point clouds are then registered with dense point clouds from the depth camera. Finally, we directly optimize the warping quality by evaluating scene constraints in 3D point clouds. Our single-shot extrinsic calibration method does not require correspondences across multiple color images or across different modalities and it is more flexible than existing methods. The scene constraints can be very simple and we demonstrate that a scene containing three sheets of paper is sufficient to obtain reliable calibration and with a lower geometric error than existing methods. Jiaolong Yang, Yuchao Dai, Hongdong Li, Henry J. Gardner, Yunde Jia |
ISMAR | 2 |
| 2013 | Rotation Averaging
Richard I. Hartley, Jochen Trumpf, Yuchao Dai, Hongdong Li |
Int. J. Comput. Vis. | 3 |
| 2013 | Projective Multiview Structure and Motion from Element-Wise FactorizationabstractThe Sturm-Triggs type iteration is a classic approach for solving the projective structure-from-motion (SfM) factorization problem, which iteratively solves the projective depths, scene structure, and camera motions in an alternated fashion. Like many other iterative algorithms, the Sturm-Triggs iteration suffers from common drawbacks, such as requiring a good initialization, the iteration may not converge or may only converge to a local minimum, and so on. In this paper, we formulate the projective SfM problem as a novel and original element-wise factorization (i.e., Hadamard factorization) problem, as opposed to the conventional matrix factorization. Thanks to this formulation, we are able to solve the projective depths, structure, and camera motions simultaneously by convex optimization. To address the scalability issue, we adopt a continuation-based algorithm. Our method is a global method, in the sense that it is guaranteed to obtain a globally optimal solution up to relaxation gap. Another advantage is that our method can handle challenging real-world situations such as missing data and outliers quite easily, and all in a natural and unified manner. Extensive experiments on both synthetic and real images show comparable results compared with the state-of-the-art methods. Yuchao Dai, Hongdong Li, Mingyi He |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | A simple prior-free method for non-rigid structure-from-motion factorizationabstractThis paper proposes a simple “prior-free” method for solving non-rigid structure-from-motion factorization problems. Other than using the basic low-rank condition, our method does not assume any extra prior knowledge about the nonrigid scene or about the camera motions. Yet, it runs reliably, produces optimal result, and does not suffer from the inherent basis-ambiguity issue which plagued many conventional nonrigid factorization techniques. Our method is easy to implement, which involves solving no more than an SDP (semi-definite programming) of small and fixed size, a linear Least-Squares or trace-norm minimization. Extensive experiments have demonstrated that it outperforms most of the existing linear methods of nonrigid factorization. This paper offers not only new theoretical insight, but also a practical, everyday solution, to non-rigid structure-from-motion. Yuchao Dai, Hongdong Li, Mingyi He |
CVPR | 1 |
| 2010 | Element-Wise Factorization for N-View Projective Reconstruction
Yuchao Dai, Hongdong Li, Mingyi He |
ECCV (4) | 1 |
| 2009 | Rotation Averaging with Application to Camera-Rig Calibration
Yuchao Dai, Jochen Trumpf, Hongdong Li, Nick Barnes, Richard I. Hartley |
ACCV (2) | 1 |
| 2009 | Two Efficient Algorithms for Outlier Removal in Multi-view Geometry Using L-Infinity NormabstractL∞ norm has been recently introduced to multi-view geometry computation to achieve globally optimal computation. It however suffers from a serious sensitivity to outliers. A few remedies have been proposed but with high computational complexity. This paper presents two efficient algorithms to overcome these problems. Our first algorithm is based on a cheap and effective local descent method (as opposed to the conventional but expensive SOCP(Second Order Cone Programming)). The second algorithm further improves the first one by using a Depth-first search heuristics. Both algorithms retain the nice property of global optimality of the L∞ scheme, while at cost only a small fraction of the original computation. Experiments on both synthetic data and real images have validated the proposed algorithms. Yuchao Dai, Mingyi He, Hongdong Li |
ICIG | 1 |