Bin Fan 0002

dblp:60/105-2 · DBLP profile ↗
← Back
34ranked-venue papers
13as first author
33since 2021 · last 2026
0000-0002-8028-0166ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 9 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 8 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Learning Spatial Decay for Vision Transformers
abstract
Vision Transformers (ViTs) have revolutionized computer vision, yet their self-attention mechanism lacks explicit spatial inductive biases, leading to suboptimal performance on spatially-structured tasks. Existing approaches introduce data-independent spatial decay based on fixed distance metrics, applying uniform attention weighting regardless of image content and limiting adaptability to diverse visual scenarios. Inspired by recent advances in large language models where content-aware gating mechanisms (e.g., GLA, HGRN2, FOX) significantly outperform static alternatives, we present the first successful adaptation of data-dependent spatial decay to 2D vision transformers. We introduce Spatial Decay Transformer (SDT), featuring a novel Context-Aware Gating (CAG) mechanism that generates dynamic, data-dependent decay for patch interactions. Our approach learns to modulate spatial attention based on both content relevance and spatial proximity. We address the fundamental challenge of 1D-to-2D adaptation through a unified spatial-content fusion framework that integrates manhattan distance-based spatial priors with learned content representations. Extensive experiments on ImageNet-1K classification and generation tasks demonstrate consistent improvements over strong baselines. Our work establishes data-dependent spatial decay as a new paradigm for enhancing spatial attention in vision transformers.
Yuxin Mao, Zhen Qin 0003, Jinxing Zhou, Bin Fan 0002, Jing Zhang 0052, Yiran Zhong, Yuchao Dai
AAAI4
2026 EC-MVSNet: Enhanced Cascaded Multi-View Stereo with Cross-Scale Relevance Integration
abstract
Cascade-based multi-scale architectures are currently the mainstream in Multi-view Stereo (MVS), achieving a balance between computational efficiency and reconstruction accuracy. However, existing cascade MVS methods suffer from significant limitations in cross-scale information utilization, where depth estimation processes operate independently across scales without fully exploiting the rich relevance between adjacent scales. To address this fundamental limitation, we propose an Enhanced Cascade Multi-View Stereo framework (EC-MVSNet), which introduces a novel cross-scale relevance integration strategy. Specifically, we introduce a Cross-Scale Feature-based Joint Construction (CFC) module to synergistically combine features from adjacent scales to build more reliable cost volumes. Additionally, a Cross-Scale Probability-guided Enhancement (CPE) module is proposed to propagate depth probability distributions across scales to guide cost volume enhancement. Furthermore, we propose a Monocular Feature-based Refinement (MFR) module to further enhance depth prediction accuracy by leveraging monocular priors. Extensive experiments demonstrate that EC-MVSNet achieves state-of-the-art performance on multiple benchmarks, validating the effectiveness of the cross-scale integration in improving MVS reconstruction quality.
Shaoqian Wang, Jiadai Sun, Bin Fan 0002, Qiang Wang 0023, Yuchao Dai
AAAI3
2026 Multi-event representation and multi-level fusion for robust RGB-event object tracking
Bin Fan 0002, Zhexiong Wan, Qi Liu 0054, Yuchao Dai
Knowl. Based Syst.2
2026 Coded Event Focal Stack for Continuous Refocusing in Dynamic Scene
abstract
Traditional cameras face limitations in maintaining focus across dynamic scenes, especially during rapid motion, due to the constraints of their lenses. Post-capture refocusing techniques, including deep learning-based methods and light field cameras, have been explored to mitigate these challenges. However, these approaches frequently struggle with temporal consistency or experience a trade-off in spatial resolution. In this paper, we introduce the coded event focal stack, a novel approach that captures both motion and depth information through event streams recorded during a modulated focal sweep. Our coded event focal stack enables the generation of full-time intermediate frames refocused at arbitrary focal distances. Extensive experiments on both synthetic and real-world datasets demonstrate the superior refocusing capability of our method over state-of-the-art techniques, particularly in dynamic scenes with complex motion and depth variations.
Minggui Teng, Suhang Xuan, Zhiang Yan, Hanyue Lou, Bin Fan 0002, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Boosting Few-Shot Hyperspectral Image Classification Through Dynamic Fusion and Hierarchical Enhancement
abstract
Few-shot learning has garnered increasing attention in hyperspectral image classification (HSIC) due to its potential to reduce dependency on labor-intensive and costly labeled data. However, most existing methods are constrained to feature extraction using a single image patch of fixed size, and typically neglect the pivotal role of the central pixel in feature fusion, leading to inefficient information utilization. In addition, the correlations among sample features have not been fully explored, thereby weakening feature expressiveness and hindering cross-domain knowledge transfer. To address these issues, we propose a novel few-shot HSIC framework incorporating dynamic fusion and hierarchical enhancement. Specifically, we first introduce a robust feature extraction module, which effectively combines the content concentration of small patches with the noise robustness of large patches, and further captures local spatial correlations through a central-pixel-guided dynamic pooling strategy. Such patch-to-pixel dynamic fusion enables a more comprehensive and robust extraction of ground object information. Then, we develop a support-query hierarchical enhancement module that integrates intraclass self-attention and interclass cross-attention mechanisms. This process not only enhances support-level and query-level feature representation but also facilitates the learning of more informative prior knowledge from the abundantly labeled source domain. Moreover, to further increase feature discriminability, we design an intraclass consistency loss and an interclass orthogonality loss, which collaboratively encourage intraclass samples to be closer together and interclass samples to be more separable in the metric space. Experimental results on four benchmark datasets demonstrate that our method substantially improves classification accuracy and consistently outperforms competing approaches. Code is available at https://github.com/guoying918/DFHE2025.
Ying Guo 0014, Bin Fan 0002, Yuchao Dai, Yan Feng 0005, Mingyi He
IEEE Trans. Neural Networks Learn. Syst.2
2025 Active Hyperspectral Imaging Using an Event Camera
abstract
Hyperspectral imaging plays a critical role in numerous scientific and industrial fields. Conventional hyperspectral imaging systems often struggle with the trade-off between capture speed, spectral resolution, and bandwidth, particularly in dynamic environments. In this work, we present a novel event-based active hyperspectral imaging system designed for real-time capture with low bandwidth in dynamic scenes. By combining an event camera with a dynamic illumination strategy, our system achieves unprecedented temporal resolution while maintaining high spectral fidelity, all at a fraction of the bandwidth requirements of traditional systems. Unlike basis-based methods that sacrifice spectral resolution for efficiency, our approach enables continuous spectral sampling through an innovative "sweeping rainbow" illumination pattern synchronized with a rotating mirror array. The key insight is leveraging the sparse, asynchronous nature of event cameras to encode spectral variations as temporal contrasts, effectively transforming the spectral reconstruction problem into a series of geometric constraints. Extensive evaluations of both synthetic and real data demonstrate that our system outperforms state-of-the-art methods in temporal resolution while maintaining competitive spectral reconstruction quality.
Bohan Yu, Jinxiu Liang, Zhuofeng Wang, Bin Fan 0002, Art Subpa-Asa, Boxin Shi, Imari Sato
CVPR4
2025 MRM-RETrack: Hybrid Multi-scale Residual and Mamba for RGB-Event Tracking
Bin Fan 0002, Zhexiong Wan, Zhiyuan Zhang 0002, Yuchao Dai
PRCV (18)2
2025 Spk2ImgMamba: Spiking Camera Image Reconstruction with Multi-Scale State Space Models
abstract
As a bio-inspired vision sensor, the spiking camera has showcased remarkable capability in high-speed imaging with a sampling rate of 40,000 Hz. Reconstructing clear images from continuous spike streams, which is obtained by each photosensor continuously detecting photons and firing them asynchronously, has garnered significant attention. Despite promising results, existing spike-to-image reconstruction methods face challenges in balancing global receptive fields and efficient computation due to the inherent limitations of their backbones. Recently, due to powerful long-range modeling and linear complexity, the state space model (SSM) has emerged as a competitive alternative to CNNs and Transformers. In this paper, we propose a lightweight spike-to-image reconstruction network that harnesses Mamba as the backbone. Our approach sequentially executes three core modules: temporal information integration, spatial feature enhancement, and progressive image reconstruction. The former accumulates cues across diverse temporal windows to explore both long-term and short-term contexts. Subsequently, to model global dependencies while heightening local detail perception, we develop a multi-scale SSM block characterized by multi-scale multi-direction scanning, which effectively boosts spatial feature representations. Finally, intensity images are decoded progressively from the enhanced light-intensity features. Extensive experiments on both synthetic and real-captured data demonstrate that our approach achieves state-of-the-art performance, with only 10% of the network parameters and nearly two orders of magnitude less computational effort. The code will be available at https://github.com/interstellarH/Spk2ImgMamba.
Jiaoyang Yin, Bin Fan 0002, Chao Xu 0006, Tiejun Huang 0001, Boxin Shi
WACV2
2025 Instance-Level Moving Object Segmentation from a Single Image with Events
Zhexiong Wan, Bin Fan 0002, Le Hui, Yuchao Dai, Gim Hee Lee
Int. J. Comput. Vis.2
2025 Self-Supervised Learning for Rolling Shutter Temporal Super-Resolution
abstract
Most cameras on portable devices adopt a rolling shutter (RS) mechanism, encoding sufficient temporal dynamic information through sequential readouts. This advantage can be exploited to recover a temporal sequence of latent global shutter (GS) images. Existing methods rely on fully supervised learning, necessitating specialized optical devices to collect paired RS-GS images as ground-truth, which is too costly to scale. In this paper, we propose a self-supervised learning framework for the first time to produce a high frame rate GS video from two consecutive RS images, unleashing the potential of RS cameras. Specifically, we first develop the unified warping model of RS2GS and GS2RS, enabling the complement conversions of RS2GS and GS2RS to be incorporated into a uniform network model. Then, based on the cycle consistency constraint, given a triplet of consecutive RS frames, we minimize the discrepancy between the input middle RS frame and its cycle reconstruction, generated by interpolating back from the predicted two intermediate GS frames. Experiments on various benchmarks show that our approach achieves comparable or better performance than state-of-the-art supervised methods while enjoying stronger generalization capabilities. Moreover, our approach makes it possible to recover smooth and distortion-free videos from two adjacent RS frames in the real-world BS-RSC dataset, surpassing prior limitations.
Bin Fan 0002, Ying Guo 0014, Yuchao Dai, Chao Xu 0006, Boxin Shi
IEEE Trans. Circuits Syst. Video Technol.1
2024 Towards HDR and HFR Video from Rolling-Mixed-Bit Spikings
abstract
The spiking cameras offer the benefits of high dynamic range (HDR), high temporal resolution, and low data redundancy. However, reconstructing HDR videos in high-speed conditions using single-bit spikings presents challenges due to the limited bit depth. Increasing the bit depth of the spikings is advantageous for boosting HDR performance, but the readout efficiency will be decreased, which is unfavorable for achieving a high frame rate (HFR) video. To address these challenges, we propose a readout mechanism to obtain rolling-mixed-bit (RMB) spikings, which involves inter-leaving multi-bit spikings within the single-bit spikings in a rolling manner, thereby combining the characteristics of high bit depth and efficient readout. Furthermore, we introduce RMB-Net for reconstructing HDR and HFR videos. RMB-Net comprises a cross-bit attention block for fusing mixed-bit spikings and a cross-time attention block for achieving temporal fusion. Extensive experiments conducted on synthetic and real-synthetic data demonstrate the superiority of our method. For instance, pure 3 -bit spikings result in 3 times of data volume, whereas our method achieves comparable performance with less than 2% increase in data volume.
Yakun Chang, Yeliduosi Xiaokaiti, Yujia Liu 0005, Bin Fan 0002, Zhaojun Huang, Tiejun Huang 0001, Boxin Shi
CVPR4
2024 Boosting Gaze Object Prediction via Pixel-Level Supervision from Vision Foundation Model
Lei Zhang 0166, Shi Yan 0005, Bin Fan 0002, Binglu Wang
ECCV (69)4
2024 Real-Data-Driven 2000 FPS Color Video from Mosaicked Chromatic Spikes
Zhaojun Huang, Yakun Chang, Bin Fan 0002, Zhaofei Yu, Boxin Shi
ECCV (12)4
2024 Spatio-Temporal Interactive Learning for Efficient Image Reconstruction of Spiking Cameras
abstract
The spiking camera is an emerging neuromorphic vision sensor that records high-speed motion scenes by asynchronously firing continuous binary spike streams. Prevailing image reconstruction methods, generating intermediate frames from these spike streams, often rely on complex step-by-step network architectures that overlook the intrinsic collaboration of spatio-temporal complementary information. In this paper, we propose an efficient spatio-temporal interactive reconstruction network to jointly perform inter-frame feature alignment and intra-frame feature filtering in a coarse-to-fine manner. Specifically, it starts by extracting hierarchical features from a concise hybrid spike representation, then refines the motion fields and target frames scale-by-scale, ultimately obtaining a full-resolution output. Meanwhile, we introduce a symmetric interactive attention block and a multi-motion field estimation block to further enhance the interaction capability of the overall network. Experiments on synthetic and real-captured data show that our approach exhibits excellent performance while maintaining low model complexity.
Bin Fan 0002, Jiaoyang Yin, Yuchao Dai, Chao Xu 0006, Tiejun Huang 0001, Boxin Shi
NeurIPS1
2024 Zero-Shot Event-Intensity Asymmetric Stereo via Visual Prompting from Image Domain
abstract
Event-intensity asymmetric stereo systems have emerged as a promising approach for robust 3D perception in dynamic and challenging environments by integrating event cameras with frame-based sensors in different views. However, existing methods often suffer from overfitting and poor generalization due to limited dataset sizes and lack of scene diversity in the event domain. To address these issues, we propose a zero-shot framework that utilizes monocular depth estimation and stereo matching models pretrained on diverse image datasets. Our approach introduces a visual prompting technique to align the representations of frames and events, allowing the use of off-the-shelf stereo models without additional training. Furthermore, we introduce a monocular cue-guided disparity refinement module to improve robustness across static and dynamic regions by incorporating monocular depth information from foundation models. Extensive experiments on real-world datasets demonstrate the superior zero-shot evaluation performance and enhanced generalization ability of our method compared to existing approaches.
Hanyue Lou, Jinxiu Liang, Minggui Teng, Bin Fan 0002, Yong Xu 0007, Boxin Shi
NeurIPS4
2024 Learning Bilateral Cost Volume for Rolling Shutter Temporal Super-Resolution
abstract
Rolling shutter temporal super-resolution (RSSR), which aims to synthesize intermediate global shutter (GS) video frames between two consecutive rolling shutter (RS) frames, has made remarkable progress with the development of deep convolutional neural networks over the past years. Existing methods cascade multiple separated networks to sequentially estimate intermediate motion fields and synthesize target GS frames. Nevertheless, they are typically complex, do not facilitate the interaction of complementary motion and appearance information, and suffer from problems such as pixel aliasing or poor interpretation. In this paper, we derive the uniform bilateral motion fields for RS-aware backward warping, which endows our network a more explicit geometric meaning by injecting spatio-temporal consistency information through time-offset embedding. More importantly, we develop a unified, single-stage RSSR pipeline to recover the latent GS video in a coarse-to-fine manner. It first extracts pyramid features from given inputs, and then refines the bilateral motion fields together with the anchor frame until generating the desired output. With the help of our proposed bilateral cost volume, which uses the anchor frame as a common reference to model the correlation with two RS frames, the gradually refined anchor frames not only facilitate intermediate motion estimation, but also compensate for contextual details, making additional frame synthesis or refinement networks unnecessary. Meanwhile, an asymmetric bilateral motion model built on top of the symmetric bilateral motion model further improves the generality and adaptability, yielding better GS video reconstruction performance. Extensive quantitative and qualitative experiments on synthetic and real data demonstrate that our method achieves new state-of-the-art results.
Bin Fan 0002, Yuchao Dai, Hongdong Li
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Distribution-Aware and Class-Adaptive Aggregation for Few-Shot Hyperspectral Image Classification
abstract
Recently, few-shot learning based on meta-learning has shown great potential in hyperspectral image classification (HSIC) due to its excellent adaptability to limited training samples. Despite achieving promising results, the existing methods ignore the interaction between the source domain (with abundant-labeled base-class samples) and the target domain (with few-labeled novel-class samples), as well as between the support set and the query set. This issue makes the resulting model usually biased toward the source domain and not robust to the sample variance of novel classes, posing a bottleneck to the improvement of HSIC performance. To overcome these limitations, we propose a flexible and effective distribution-aware and class-adaptive aggregation (DA-CAA) method for few-shot HSIC by transferring the class-level distribution information learned from the base classes to the novel classes. Specifically, we first employ a variational autoencoder (VAE), which is pretrained on abundant-labeled base-class samples, to encode the support set samples as class distributions. Subsequently, we sample class-level features from the learned distribution and adaptively aggregate them with sample-specific query features. This operation not only enhances cross-domain information interaction in a distribution-learning manner, but also ensures that the aggregated features across classes inherit both class-level and sample-specific information. Our proposed class-adaptive aggregation (CAA) encourages complementary fusion of features from all classes, which is beneficial for reducing class confusion. Experiments on four benchmark datasets demonstrate the effectiveness and flexibility of our approach.
Ying Guo 0014, Bin Fan 0002, Yan Feng 0005, Xiuping Jia, Mingyi He
IEEE Trans. Geosci. Remote. Sens.2
2024 Unified Video Reconstruction for Rolling Shutter and Global Shutter Cameras
abstract
Currently, the general domain of video reconstruction (VR) is fragmented into different shutters spanning global shutter and rolling shutter cameras. Despite rapid progress in the state-of-the-art, existing methods overwhelmingly follow shutter-specific paradigms and cannot conceptually generalize to other shutter types, hindering the uniformity of VR models. In this paper, we propose UniVR, a versatile framework to handle various shutters through unified modeling and shared parameters. Specifically, UniVR encodes diverse shutter types into a unified space via a tractable shutter adapter, which is parameter-free and thus can be seamlessly delivered to current well-established VR architectures for cross-shutter transfer. To demonstrate its effectiveness, we conceptualize UniVR as three shutter-generic VR methods, namely Uni-SoftSplat, Uni-SuperSloMo, and Uni-RIFE. Extensive experimental results demonstrate that the pre-trained model without any fine-tuning can achieve reasonable performance even on novel shutters. After fine-tuning, new state-of-the-art performances are established that go beyond shutter-specific methods and enjoy strong generalization. The code is available at https://github.com/GitCVfb/UniVR.
Bin Fan 0002, Zhexiong Wan, Boxin Shi, Chao Xu 0006, Yuchao Dai
IEEE Trans. Image Process.1
2023 Joint Appearance and Motion Learning for Efficient Rolling Shutter Correction
abstract
Rolling shutter correction (RSC) is becoming increasingly popular for RS cameras that are widely used in commercial and industrial applications. Despite the promising performance, existing RSC methods typically employ a two-stage network structure that ignores intrinsic infor-mation interactions and hinders fast inference. In this pa-per, we propose a single-stage encoder-decoder-based network, named JAMNet, for efficient RSC. It first extracts pyramid features from consecutive RS inputs, and then simultaneously refines the two complementary information (i.e., global shutter appearance and undistortion motion field) to achieve mutual promotion in a joint learning de-coder. To inject sufficient motion cues for guiding joint learning, we introduce a transformer-based motion embed-ding module and propose to pass hidden states across pyra-mid levels. Moreover, we present a new data augmentation strategy “vertical flip + inverse order” to release the potential of the RSC datasets. Experiments on various benchmarks show that our approach surpasses the state-of-the-art methods by a large margin, especially with a 4.7 dB PSNR leap on real-world RSC. Code is available at https://github.com/GitCVfb/JAMNet.
Bin Fan 0002, Yuxin Mao, Yuchao Dai, Zhexiong Wan, Qi Liu 0054
CVPR1
2023 Grid-Transformer for Few-Shot Hyperspectral Image Classification
abstract
The application of few-shot learning to hyperspectral image (HSI) classification tasks has gradually become a research hotspot due to the difficulties in acquiring and labeling HSI data. Existing methods tend to cascade a large number of convolutional neural networks. However, such operations can only focus on local information and cannot accurately capture the strong correlation between spectra. To address this problem, we propose Grid-transformer, an efficient spatial-spectral feature extraction model. Specifically, we first introduce a more powerful transformer to compute non-local self-similarity along the spectral dimension, which is beneficial to mine more discriminative spectral features. Then, they are embedded into a grid-like network architecture to fully aggregate multi-scale contextual information, resulting in a more complete spatial-spectral feature representation. Experiments on two benchmark datasets demonstrate that our approach achieves state-of-the-art classification performance.
Ying Guo 0014, Mingyi He, Bin Fan 0002
ICIP3
2023 Rolling Shutter Inversion: Bring Rolling Shutter Images to High Framerate Global Shutter Video
abstract
A single rolling-shutter (RS) image may be viewed as a row-wise combination of a sequence of global-shutter (GS) images captured by a (virtual) moving GS camera within the exposure duration. Although rolling-shutter cameras are widely used, the RS effect causes obvious image distortion especially in the presence of fast camera motion, hindering downstream computer vision tasks. In this paper, we propose to invert the rolling-shutter image capture mechanism, i.e., recovering a continuous high framerate global-shutter video from two time-consecutive RS frames. We call this task the RS temporal super-resolution (RSSR) problem. The RSSR is a very challenging task, and to our knowledge, no practical solution exists to date. This paper presents a novel deep-learning based solution. By leveraging the multi-view geometry relationship of the RS imaging process, our learning based framework successfully achieves high framerate GS generation. Specifically, three novel contributions can be identified: (i) novel formulations for bidirectional RS undistortion flows under constant velocity as well as constant acceleration motion model. (ii) a simple linear scaling operation, which bridges the RS undistortion flow and regular optical flow. (iii) a new mutual conversion scheme between varying RS undistortion flows that correspond to different scanlines. Our method also exploits the underlying spatial-temporal geometric relationships within a deep learning framework, where no additional supervision is required beyond the necessary middle-scanline GS image. Building upon these contributions, this paper represents the very first rolling-shutter temporal super-resolution deep-network that is able to recover high framerate global-shutter videos from just two RS frames. Extensive experimental results on both synthetic and real data show that our proposed method can produce high-quality GS image sequences with rich details, outperforming the state-of-the-art methods.
Bin Fan 0002, Yuchao Dai, Hongdong Li
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Context-Aware Video Reconstruction for Rolling Shutter Cameras
abstract
With the ubiquity of rolling shutter (RS) cameras, it is becoming increasingly attractive to recover the latent global shutter (GS) video from two consecutive RS frames, which also places a higher demand on realism. Existing solutions, using deep neural networks or optimization, achieve promising performance. However, these methods generate intermediate GS frames through image warping based on the RS model, which inevitably result in black holes and noticeable motion artifacts. In this paper, we alleviate these issues by proposing a context-aware GS video reconstruction architecture. It facilitates the advantages such as occlusion reasoning, motion compensation, and temporal abstraction. Specifically, we first estimate the bilateral motion field so that the pixels of the two RS frames are warped to a common GS frame accordingly. Then, a refinement scheme is proposed to guide the GS frame synthesis along with bilateral occlusion masks to produce high-fidelity GS video frames at arbitrary times. Furthermore, we derive an approximated bilateral motion field model, which can serve as an alternative to provide a simple but effective GS frame initialization for related tasks. Experiments on synthetic and real data show that our approach achieves superior performance over state-of-the-art methods in terms of objective metrics and subjective visual quality. Code is available at https://github.com/GitCVfb/CVR.
Bin Fan 0002, Yuchao Dai, Zhiyuan Zhang 0002, Qi Liu 0054, Mingyi He
CVPR1
2022 Differential SfM and image correction for a rolling shutter stereo rig
Bin Fan 0002, Yuchao Dai, Zhiyuan Zhang 0002
Image Vis. Comput.1
2022 Fast and Robust Differential Relative Pose Estimation With Radial Distortion
abstract
In this letter, we address the differential two-view geometry problem of estimating the relative pose between two consecutive frames in the presence of radial distortion. This problem is of both theoretical and practical interests and has not been solved. We derive its parameterization and present an effective and robust generalized eigenvalue solver based on the hidden variable technique. Furthermore, we propose a nonlinear refinement scheme within the maximum likelihood criterion to produce more accurate estimates of the relative pose and radial distortion. Compared with the standard differential solutions without modeling the radial distortion, our approach can recover more geometrically correct point correspondences for a pair of radially distorted images. Moreover, our differential solution runs an order of magnitude faster than the discrete solution in terms of recovering the full camera motion. Experiment results on both synthetic and real data demonstrate the effectiveness of our model and method in dealing with the radial distortion.
Bin Fan 0002, Yuchao Dai, Zhiyuan Zhang 0002, Mingyi He
IEEE Signal Process. Lett.1
2022 Searching Dense Point Correspondences via Permutation Matrix Learning
abstract
Although 3D point cloud data has received widespread attentions as a general form of 3D signal expression, applying point clouds to the task of dense correspondence estimation between 3D shapes has not been investigated widely. Furthermore, even in the few existing 3D point cloud-based methods, an important and widely acknowledged principle,i.e. one-to-one matching, is usually ignored. In response, this paper presents a novel end-to-end learning-based method to estimate the dense correspondence of 3D point clouds, in which the problem of point matching is formulated as a zero-one assignment problem to achieve a permutation matching matrix to implement the one-to-one principle fundamentally. Note that the classical solutions of this assignment problem are always non-differentiable, which is fatal for deep learning frameworks. Thus we design a special matching module, which solves a doubly stochastic matrix at first and then projects this obtained approximate solution to the desired permutation matrix. Moreover, to guarantee end-to-end learning and the accuracy of the calculated loss, we calculate the loss from the learned permutation matrix but propagate the gradient to the doubly stochastic matrix directly which bypasses the permutation matrix during the backward propagation. Our method can be applied to both non-rigid and rigid 3D point cloud data and extensive experiments show that our method achieves state-of-the-art performance for dense correspondence learning.The code will be released.
Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Bin Fan 0002, Qi Liu 0054
IEEE Signal Process. Lett.4
2022 Learning a Task-Specific Descriptor for Robust Matching of 3D Point Clouds
abstract
Existing learning-based point feature descriptors are usually task-agnostic, which pursue describing the individual 3D point clouds as accurate as possible. However, the matching task aims at describing the corresponding points consistently across different 3D point clouds. Therefore these too accurate features may play a counterproductive role due to the inconsistent point feature representations of correspondences caused by the unpredictable noise, partiality, deformation, etc., in the local geometry. In this paper, we propose to learn a robust task-specific feature descriptor to consistently describe the correct point correspondence under interference. Born with anEncoder and aDynamicFusion module, our method EDFNet develops from two aspects. First, we augment the matchability of correspondences by utilizing their repetitive local structure. To this end, a special encoder is designed to exploit two input point clouds jointly for each point descriptor. It not only captures the local geometry of each point in the current point cloud by convolution, but also exploits the repetitive structure from paired point cloud by Transformer. Second, we propose a dynamical fusion module to jointly use different scale features. There is an inevitable struggle between robustness and discriminativeness of the single scale feature. Specifically, the small scale feature is robust since little interference exists in this small receptive field. But it is not sufficiently discriminative as there are many repetitive local structures within a point cloud. Thus the resultant descriptors will lead to many incorrect matches. In contrast, the large scale feature is more discriminative by integrating more neighborhood information. But it is easier to be disturbed since there is much more interference in the large receptive field. Compared with the conventional fusion strategy that handles multiple scale features equally, we analyze the consistency of them to judge the clean ones and perform larger aggregation weights on them during fusion. Then, a robust and discriminative feature descriptor is achieved by focusing on multiple clean scale features. Extensive evaluations validate that EDFNet learns a task-specific descriptor, which achieves state-of-the-art or comparable performance for robust matching of 3D point clouds.
Zhiyuan Zhang 0002, Yuchao Dai, Bin Fan 0002, Jiadai Sun, Mingyi He
IEEE Trans. Circuits Syst. Video Technol.3
2022 VRNet: Learning the Rectified Virtual Corresponding Points for 3D Point Cloud Registration
abstract
3D point cloud registration is fragile to outliers, which are labeled as the points without corresponding points. To handle this problem, a widely adopted strategy is to estimate the relative pose based only on some accurate correspondences, which is achieved by building correspondences on the identified inliers or by selecting reliable ones. However, these approaches are usually complicated and time-consuming. By contrast, the virtual point-based methods learn the virtual corresponding points (VCPs) for allsourcepoints uniformly without distinguishing the outliers and the inliers. Although this strategy is time-efficient, the learned VCPs usually exhibit serious collapse degeneration due to insufficient supervision and the inherent distribution limitation. In this paper, we propose to exploit the best of both worlds and present a novel robust 3D point cloud registration framework. We follow the idea of the virtual point-based methods but learn a new type of virtual points called rectified virtual corresponding points (RCPs), which are defined as the point set with the same shape as thesourceand with the same pose as thetarget. Hence, a pair of consistent point clouds,i.e.sourceand RCPs, is formed by rectifying VCPs to RCPs (VRNet), through which reliable correspondences betweensourceand RCPs can be accurately obtained. Since the relative pose betweensourceand RCPs is the same as the relative pose betweensourceandtarget, the input point clouds can be registered naturally. Specifically, we first construct the initial VCPs by using an estimated soft matching matrix to perform a weighted average on thetargetpoints. Then, we design a correction-walk module to learn an offset to rectify VCPs to RCPs, which effectively breaks the distribution limitation of VCPs. Finally, we develop a hybrid loss function to enforce the shape and geometry structure consistency of the learned RCPs and thesourceto provide sufficient supervision. Extensive experiments on several benchmark datasets demonstrate that our method achieves advanced registration performance and time-efficiency simultaneously.The code will be made public.
Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Bin Fan 0002, Mingyi He
IEEE Trans. Circuits Syst. Video Technol.4
2021 Inverting a Rolling Shutter Camera: Bring Rolling Shutter Images to High Framerate Global Shutter Video
abstract
Rolling shutter (RS) images can be viewed as the result of the row-wise combination of global shutter (GS) images captured by a virtual moving GS camera over the period of camera readout time. The RS effect brings tremendous difficulties for the downstream applications. In this paper, we propose to invert the above RS imaging mechanism, i.e., recovering a high framerate GS video from consecutive RS images to achieve RS temporal super-resolution (RSSR). This extremely challenging problem, e.g., recovering 1440 GS images from two 720-height RS images, is far from being solved end-to-end. To address this challenge, we exploit the geometric constraint in the RS camera model, thus achieving geometry-aware inversion. Specifically, we make three contributions in resolving the above difficulties: (i) formulating the bidirectional RS undistortion flows under the constant velocity motion model, (ii) building the connection between the RS undistortion flow and optical flow via a scaling operation, and (iii) developing a mutual conversion scheme between varying RS undistortion flows that correspond to different scanlines. Building upon these formulations, we propose the first RS temporal super-resolution network in a cascaded structure to extract high framerate global shutter video. Our method explores the underlying spatio-temporal geometric relationships within a deep learning framework, where no extra supervision besides the middle-scanline ground truth GS image is needed. Essentially, our method can be very efficient for explicit propagation to generate GS images under any scanline. Experimental results on both synthetic and real data show that our method can produce high-quality GS image sequences with rich details, outperforming state-of-the-art methods.
Bin Fan 0002, Yuchao Dai
ICCV1
2021 SUNet: Symmetric Undistortion Network for Rolling Shutter Correction
abstract
The vast majority of modern consumer-grade cameras employ a rolling shutter mechanism, leading to image distortions if the camera moves during image acquisition. In this paper, we present a novel deep network to solve the generic rolling shutter correction problem with two consecutive frames. Our pipeline is symmetrically designed to predict the global shutter image corresponding to the intermediate time of these two frames, which is difficult for existing methods because it corresponds to a camera pose that differs most from the two frames. First, two time-symmetric dense undistortion flows are estimated by using well-established principles: pyramidal construction, warping, and cost volume processing. Then, both rolling shutter images are warped into a common global shutter one in the feature space, respectively. Finally, a symmetric consistency constraint is constructed in the image decoder to effectively aggregate the contextual cues of two rolling shutter images, thereby recovering the high-quality global shutter image. Extensive experiments with both synthetic and real data from public benchmarks demonstrate the superiority of our proposed approach over the state-of-the-art methods.
Bin Fan 0002, Yuchao Dai, Mingyi He
ICCV1
2021 Homography-based camera pose estimation with known gravity direction for UAV navigation
Chunhui Zhao 0002, Bin Fan 0002, Jinwen Hu, Quan Pan 0001
Sci. China Inf. Sci.2
2021 Rolling-Shutter-stereo-aware motion estimation and image correction
Bin Fan 0002, Yuchao Dai
Comput. Vis. Image Underst.1
2021 Self-supervised multi-body scene flow estimation
Jihuang Dai, Yuchao Dai, Bin Fan 0002
Neurocomputing3
2021 RS-DPSNet: Deep Plane Sweep Network for Rolling Shutter Stereo Images
abstract
Since the rolling shutter (RS) camera successively exposes each scanline, accurately reconstructing scene depth from an RS stereo image pair remains a great challenge. Directly applying the deep-learning-based depth estimation methods tailored for the global shutter (GS) stereo images leads to undesirable RS depth results due to inherent flaws in the network structure. In this letter, we fill this gap by developing an end-to-end RS-stereo-aware plane sweep network to improve the accuracy of the classic GS-based algorithm (i.e.DPSNet) in estimating the RS depth map. Specifically, we derive the RS-stereo-aware plane sweep model and further produce a more accurate and efficient cost volume through the effective incorporation of this model within DPSNet. Furthermore, to enable learning-based approaches to address the depth estimation problem in the context of RS stereo images, we contribute the first RS stereo dataset, CARLA-RSS. Experimental results demonstrate that our proposed pipeline achieves state-of-the-art performance.
Bin Fan 0002, Yuchao Dai, Mingyi He
IEEE Signal Process. Lett.1
2020 Relative Pose Estimation For Stereo Rolling Shutter Cameras
abstract
In this paper, we present a novel linear algorithm to estimate the 6 DoF relative pose from consecutive frames of stereo rolling shutter (RS) cameras. Our method is derived based on the assumption that stereo cameras undergo motion with constant velocity around the center of the baseline, which needs 9 pairs of correspondences on both left and right consecutive frames. The stereo RS images enable the recovery of depth maps from the semi-global matching (SGM) algorithm. With the estimated camera motion and depth map, we can correct the RS images to get the undistorted images without any scene structure assumption. Experiments on both simulated points and synthetic RS images demonstrate the effectiveness of our algorithm in relative pose estimation.
Bin Fan 0002, Yuchao Dai
ICIP2