Zhen Chen 0004

dblp:11/1266-4 · DBLP profile ↗
← Back
30ranked-venue papers
0as first author
25since 2021 · last 2026
0000-0003-1020-0615ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 23 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 MonoSISTR: Monocular 3D Object Detection via Staged Iterative Structure and Target Refinement
abstract
Monocular 3D object detection aims to predict object category, position, size, and orientation from a single RGB image. Existing DETR-based monocular 3D detectors suffer from maintaining consistent high-confidence responses due to weak or incomplete target features, resulting in information loss for distant and occluded objects during encoding-decoding. Firstly, to address the core challenge of insufficient target feature perception in complex scenes, we propose a staged iterative monocular 3D detector that progressively refines targets from coarse to fine through multiple paired encoding-decoding stages, significantly improving both feature utilization and network convergence. Furthermore, each stage integrates a dynamic target iteration module that continuously enhances query representation by focusing on high-confidence regional features, thereby enhancing the model's perception of potential targets. Finally, we design a dual-branch depth estimator with parallel global and local processing for a comprehensive representation of the depth feature. Experimental results show that our method achieves superior performance over prior approaches on challenging scenarios (e.g., distant and occluded objects) in the KITTI dataset without auxiliary data, while maintaining competitive accuracy on the nuScenes benchmark under frontal-view settings.
Genlin Zhou, Zige Wang, Zhen Chen 0004, Congxuan Zhang
3DV5
2026 FlowAnyTime: Efficient Fine-tuning with Intra-Inter Frame Distillation for All-Weather Optical Flow Estimation
abstract
Motion estimation in degraded scenes has long been a significant challenge, primarily attributed to substantial scene variations and insufficient training data. Existing approaches typically address this limitation by incorporating additional training strategies or modifying network architectures within conventional frameworks. However, these solutions not only require cumbersome training procedures or additional modal inputs, but also lack generalization capabilities. To address this problem, we propose a unified optical flow estimation framework specifically designed for degraded scenes. In this work, we employ large-scale pre-trained optical flow foundation models as both teacher and student networks. Our objective is to compensate for feature incompleteness during image degradation through pre-trained large models. Subsequently, we leverage supervised signals for fine-tuning and introduce an intra-inter frame distillation method to enable the student network to adapt to diverse cross-domain scenarios. Our proposed methodology provides deeper insights into learning style-invariant features from these learnable fine-tuning layers. Extensive experiments demonstrate that our approach achieves superior generalization performance and state-of-the-art results in degraded scenes (including low-light, rain, fog and other conditions) while requiring minimal training resources.
Hongye Chen, Xiaochun Zou, Congxuan Zhang, Zhen Chen 0004
AAAI5
2026 SCG-SSC: Semantic Scene Completion via Self-and-Cross Gated Fusion of Depth Maps and Semantic Priors
Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ke Lu 0002, Liyue Ge
ICPR (9)3
2026 SAP-DQR: Joining Spatial-Adaptive Pyramid and Adaptive Query Reorganization for Speed-Accuracy Instance Segmentation
Jiahao Zou, Congxuan Zhang, Liyue Ge, Jiawen Yang, Zhen Chen 0004, Ke Lu 0002
MMM (1)6
2026 Iter3DDet: Depth-Guided Iterative Fusion and Refinement for Monocular 3D Object Detection
abstract
Monocular 3D object detection offers significant potential for autonomous systems due to its inherent cost-effectiveness and scalability. While DETR-based architectures excel in 2D vision tasks, critical limitations persist in extending them effectively to monocular 3D detection, as evidenced in existing frameworks like MonoDETR and MonoDGP. These methods typically suffer from inefficient serial fusion of multimodal features and lack iterative refinement mechanisms, limiting their performance, especially for mid-to-long range targets. To overcome these shortcomings, we propose Iter3DDet, a novel depth-guided iterative refinement framework that integrates fine-grained feature fusion to significantly enhance detection performance. The core novelty of our approach lies in two key innovations: (1) A hybrid feature encoder combining MonoDGP’s region segmentation head with MonoDETR’s visual backbone, augmented by a multi-scale context attention module that dynamically aggregates structural and semantic cues across pyramid levels, eliminating heuristic fusion rules; (2) A depth-guided adaptive cross-modal decoder that iteratively fuses depth and context features through prioritized attention mechanisms, coupled with a novel iterative refinement training strategy that progressively refines 3D detection hypotheses, substantially improving accuracy across targets of varying difficulty levels. Extensive experiments on the KITTI, nuScenes, and Waymo benchmarks demonstrate Iter3DDet’s state-of-the-art performance, validating the effectiveness of our iterative refinement paradigm. The code will be open-sourced at https://github.com/PCwenyue.
Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ke Lu 0002, Liyue Ge
IEEE Trans. Circuits Syst. Video Technol.3
2026 DRDFNet: A Degradation-Aware Restoration and Detail-Preserving Fusion Network for Infrared and Visible Image
abstract
Multi-source image fusion combines infrared and visible information to improve scene perception in applications such as drone reconnaissance and autonomous driving. However, most existing infrared-visible image fusion methods are developed under ideal imaging assumptions. In adverse environments, visible images often lose structural and textural details, whereas infrared images are affected by noise, stripe artifacts, and low contrast, leading to degraded fusion quality and weakened downstream perception performance. To address these limitations, we propose a unified Degradation-aware Restoration and Detail-preserving Fusion Network (DRDFNet), which consists of a Degradation-Aware Restoration Transformer and a Detail-Preserving Fusion Mamba. The restoration branch uses a Compound Degradation Restoration Module (CDRM) to remove complex degradations, while the fusion branch employs a Dynamic Feature Fusion Module (DFFM) to integrate local complementary cues and global correlations across modalities. A two-stage training strategy is further introduced to reduce the optimization conflict between restoration and fusion. In addition, we construct DIVIF, a large-scale degraded IVIF benchmark generated by a physics-based imaging simulator. Experiments on the DIVIF and AWMM-100k benchmarks demonstrate that DRDFNet achieves robust and competitive performance compared with SOTA methods. Both the dataset and source code will be made publicly available at https://github.com/Liupeng97/DRDFNet.
Peng Liu 0024, An Wei, Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ke Lu 0002
IEEE Trans. Image Process.4
2025 MotionFlow: Joint Motion Priors and Appearance Enhancement for High-Accuracy Optical Flow Estimation
abstract
Although optical flow estimation has improved significantly in recent years, large displacements and occlusions remain challenging for current methods due to motion discontinuities that may hinder accurate feature correspondences in these regions, leading to degraded performance. To address this challenge, we propose a novel method named MotionFlow for high-accuracy optical flow estimation. In the encoding stage, we integrate multi-scale features enhance motion and context appearance information via cross- and inter-enhancement module. Subsequently, cross-frame features are utilized to establish motion priors, thereby providing essential prior knowledge for motion estimation. During the decoding stage, we align appearance features of the target frame with the reference frame through warping to retrieve missing context crucial for motion decoding. Experimental results demonstrate the efficacy of our approach, achieving state-of-the-art performance, particularly outperforming online benchmarks on Sintel Final pass and KITTI-2015 datasets.
Congxuan Zhang, Zhen Chen 0004, Hongye Chen, Liyue Ge, Ke Lu 0002
ICASSP3
2025 ICIMG-Net: Inject Context Information to Motion Generation for Optical Flow Estimation
abstract
Although the overall performance of existing optical flow estimation methods has improved rapidly, motion discontinuities caused by large displacements and occlusions remain significant challenges for accurate optical flow estimation. To address this issue, we propose a novel Inject Context Information for Motion Generation Network (ICIMG-Net). Firstly, we design a Context Injection Module (CIM), which enhances the generation of correlation volumes by injecting context information. This process supplements the semantic details needed for accurate pixel matching, improving matching precision. Then, we construct a Dual-Injection-GRU (DIG), which facilitates dual interactions between semantic and motion information to address the motion discontinuities problem. Finally, we conduct a comprehensively evaluat of our ICIMG-Net against state-of-the-art methods on the MPI-Sintel and KITTI benchmarks, demonstrating that our method achieves competitive results, particularly in complex synthetic and real-world traffic scenarios.
Wenbo Yin, Congxuan Zhang, Zhen Chen 0004, Liyue Ge, Zige Wang
ICASSP3
2025 WCG-Net: Warping Consistency Compensation Guided Multi-Feature Fusion For Stereo Matching
abstract
Despite the significant progress achieved by iterative optimization-based stereo matching methods, a critical challenge persists: these state-of-the-art models continue to face difficulties when handling ill-posed regions. This stems from lighting variations and viewpoint differences, which may cause the feature distributions extracted from the left and right images to differ, resulting in unreliable cost volume construction in ill-posed regions. To remedy this issue, we propose a novel network for stereo matching, named WCG-Net. In WCG-Net, we develop a warping consistency compensation module (WCCM) that employs consistency attention to identify feature differences and generate cross-view features, which are then used to construct the warping correlation volume. By introducing warping correlation volume into warping-guided fusion recurrent unit (WFRU), our method refines the disparity map iteratively, focusing on correcting errors in ill-posed regions. Extensive experimental evaluation demonstrates that WCG-Net achieves competitive results on the KITTI and ETH3D datasets, with particularly superior performance compared to other state-of-the-art algorithms on the KITTI dataset.
Zhibo Rao, Zhen Chen 0004, Congxuan Zhang
ICME4
2025 SRSA-Depth: shape and region similarity awareness for outdoor monocular depth estimation
Liyue Ge, Congxuan Zhang, Zhen Chen 0004, Ke Lu 0002
Multim. Syst.3
2025 DABF-Net: A Dual-Branch Attention-Guided and Bi-Directional Feature Enhancement Network for Infrared Small-Target Detection With Air-to-Ground Benchmark
abstract
Infrared small-target detection (IRSTD) is a critical, yet challenging task with significant applications in both military and civilian domains. Despite advancements in existing methods, two major limitations remain: the difficulty of achieving an optimal balance between detection probability and false alarm rate, and the lack of specialized datasets for air-to-ground scenarios. To address these limitations, this article presents a dual-pronged solution. At the algorithmic level, we propose a novel dual-branch attention-guided and bi-directional feature enhancement network (DABF-Net). First, we design a dual-branch high-low frequency attention (DHLA), which enhances the discriminability between the target and the background by preserving high-frequency edge features and modeling low-frequency contextual information in a complementary manner. Subsequently, we construct a bi-directional fusion module (BFM) to optimize multiscale feature compatibility while suppressing redundant information propagation. Furthermore, we introduce a small-target feature enhancement branch (STEB) employing space-to-depth (SPD) convolution and a feature integration module (FIM) to amplify latent target signatures through exponentially expanding receptive fields. At the data level, we contribute the NCHU-A2G-SIRST benchmark, the first comprehensive dataset specifically designed for the air-to-ground IRSTD task. The dataset contains four different scenes with two types of annotations, enabling evaluation and training of detection models in real-world conditions. Extensive experiments on several challenging datasets, including the self-built NCHU-A2G-SIRST dataset and three public datasets (NCHU-SIRST, NUAA-SIRST, and IRSTD-1K), demonstrate that the DABF-Net can outperform many state-of-the-art competing methods. The code and dataset are publicly available athttps://github.com/PCwenyue/DABF-Net
Fagan Wang, Congxuan Zhang, Peng Liu 0024, Zhen Chen 0004, Weiming Hu 0004
IEEE Trans. Geosci. Remote. Sens.5
2025 Self-Supervised Monocular Depth Estimation With Dual-Path Encoders and Offset Field Interpolation
abstract
Although self-supervised learning approaches have demonstrated tremendous potential in multi-frame depth estimation scenarios, existing methods struggle to perform well in cases involving dynamic targets and static ego-camera conditions. To address this issue, we propose a self-supervised monocular depth estimation method featuring dual-path encoders and learnable offset interpolation (LOI). First, we construct a dual-path encoding scheme that utilizes residual and transformer blocks to extract both single- and multi-frame features from the input frames. We design a contrastive learning strategy to effectively decouple single- and multi-frame features, enabling weighted fusion guided by a confidence map. Next, we explore two distinct decoding heads for simultaneously generating low-resolution predictions and offset fields. We then design an LOI module to directly upsample a low-resolution depth map to a full-resolution map. This one-step decoding framework enables accurate and efficient depth prediction. Finally, we evaluate our proposed method on the KITTI and Cityscapes benchmarks, conducting a comprehensive comparison with state-of-the-art approaches. The experimental results demonstrate that our DualDepth method achieves competitive performance in terms of both estimation accuracy and efficiency.
Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ke Lu 0002, Liyue Ge
IEEE Trans. Image Process.3
2025 GPDF-Net: geometric prior-guided stereo matching with disparity fusion refinement
Congxuan Zhang, Zhibo Rao, Zhen Chen 0004, Zige Wang, Ke Lu 0002
Vis. Comput.4
2024 Real-Time Monocular Depth Estimation on Embedded Systems
abstract
Depth sensing is of paramount importance for unmanned aerial and autonomous vehicles. Nonetheless, contemporary monocular depth estimation methods employing complex deep neural networks within Convolutional Neural Networks are inadequately expedient for real-time inference on embedded platforms. This paper endeavors to surmount this challenge by proposing two efficient and lightweight architectures, RT-MonoDepth and RT-MonoDepth-S, thereby mitigating computational complexity and latency. Our methodologies not only attain accuracy comparable to prior depth estimation methods but also yield faster inference speeds. Specifically, RT-MonoDepth and RT-MonoDepth-S achieve frame rates of 18.4&30.5 FPS on NVIDIA Jetson Nano and 253.0&364.1 FPS on Jetson AGX Orin, utilizing a single RGB image of resolution $640 \times 192$. The experimental results underscore the superior accuracy and faster inference speed of our methods in comparison to existing fast monocular depth estimation methodologies on the KITTI dataset.
Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Liyue Ge
ICIP3
2024 A Robust and Real-Time RGB-D SLAM Method with Dynamic Point Recognition and Depth Segmentation Optimization
Shuaixin Chen, Baolin Gan, Congxuan Zhang, Zhen Chen 0004, Ke Lu 0002
PRCV (9)4
2024 IterDepth: Iterative Residual Refinement for Outdoor Self-Supervised Multi-Frame Monocular Depth Estimation
abstract
Self-supervised monocular depth estimation has been a challenging task in computer vision for a long time, and it relies on only monocular or stereo video for its supervision. To address the challenge, we propose a novel multi-frame monocular depth estimation method called IterDepth, which is based on an iterative residual refinement network. IterDepth extracts depth features from consecutive frames and computes a 3D cost volume measuring the difference between current and previous features transformed by PoseCNN (pose estimation convolutional neural network). We reformulate depth prediction as a residual learning problem, revamping the dominating depth regression to enable high-accuracy multi-frame monocular depth estimation. Specifically, we design a gated recurrent depth fusion unit that seamlessly blends depth features from the cost volume, image features, and the depth prediction. The unit updates the hidden states and refines the depth map through iterative refinement, achieving more accurate predictions than existing methods. Our experiments on the KITTI dataset demonstrate that IterDepth is$7\times $faster in terms of FPS (frames per second) than the recent state-of-the-art DepthFormer model with competitive performance. We also test IterDepth on the Cityscapes dataset to showcase its generalization capability in other real-world environments. Moreover, IterDepth can balance accuracy and computational efficiency by adjusting the number of refinement iterations and performs competitively with other CNN-based monocular depth estimation approaches. Source code is available athttps://github.com/PCwenyue/IterDepth-TCSVT.
Zhen Chen 0004, Congxuan Zhang, Weiming Hu 0004, Bing Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 ACR-Net: Learning High-Accuracy Optical Flow via Adaptive-Aware Correlation Recurrent Network
abstract
Although recurrent network-based optical flow estimation methods have shown great success in recent years, most of these methods have difficulty handling large displacements and occlusions because the existing recurrent networks are usually restricted to coarse-resolution single-scale models while ignoring the multiscale features brought by hierarchical concepts in previous coarse-to-fine approaches. In this paper, we propose an adaptive-aware correlation recurrent network for optical flow estimation, named ACR-Net, which preserves fine motion features with a single-scale resolution recurrent framework and adaptively incorporates multiscale features at different stages to achieve high-accuracy optical flow estimation. First, our proposed self-adaptation scale-aware correlation module can incorporate the adaptive correlation of multiscale inter- and intra-motion features, which makes the features more discriminative for capturing long-range dependencies between pixels. Second, our presented adaptive-aware motion module can effectively extract the required features of different kinds of motion from multilevel correspondence. Third, our introduced cross-guide motion and fusion modules can accurately guide the propagation of reliable pixels towards unreliable pixels and dynamically determine the most suitable expression to address the occlusion challenges. Comprehensive experiments demonstrate that ACR-Net outperforms existing two-view models, striking a good balance between speed and accuracy and achieving the best performance on the MPI-Sintel final pass and KITTI-2015 test datasets. The code will be made publicly available.
Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ke Lu 0002, Liyue Ge, Zige Wang
IEEE Trans. Circuits Syst. Video Technol.3
2023 LCIF-Net: Local criss-cross attention based optical flow method using multi-scale image features and feature pyramid
Zige Wang, Zhen Chen 0004, Congxuan Zhang, Zhongkai Zhou, Hao Chen 0060
Signal Process. Image Commun.2
2023 Jointing Recurrent Across-Channel and Spatial Attention for Multi-Object Tracking With Block-Erasing Data Augmentation
abstract
Although deep-learning-based multi-object tracking (MOT) approaches have achieved remarkable performances in terms of accuracy and efficiency, the issue of object occlusions remains an open challenge for most one-shot MOT methods. To address the problem of object occlusions, in this paper we present a recurrent across-channel and spatial attention-based one-shot multi-object tracking method with block-erasing data augmentation. First, we construct a multiattention feature learning module, named RASFL, that combines recurrent across -channel attention with spatial attention. The RASFL extracts both the correlations of the feature channels and the differences of the spatial locations to improve the accuracy of the re-identification (Re-ID) task. Second, we adopt a block-erasing data augmentation strategy to handle object occlusions by using random pixel blocks to simulate occlusion cases during the network training process. This block-erasing data augmentation assists the network to be more robust under object occlusions. By integrating the proposed RASFL module and the block-erasing data augmentation strategy into a one-shot online MOT system, we build an accurate and robust MOT model called DcMOT. Finally, we run our method on the MOT16, MOT17 and MOT20 datasets to conduct a comprehensive comparison with some of the state-of-the-art MOT methods. The experimental results demonstrate that the proposed DcMOT model achieves a competitive performance in terms of both accuracy and efficiency; with especially good performances in the occlusion cases.
Keyu Deng, Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Bing Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 A Robust Infrared Small Target Detection Method Jointing Multiple Information and Noise Prediction: Algorithm and Benchmark
abstract
Infrared small target detection plays an important role in many military and civilian applications. Despite the great advances made by infrared small target detection studies in recent years, most of the existing methods have difficulty in balancing detection probabilities and false alarms. Moreover, there are only a few public datasets for infrared small targets, which limits the development of infrared small target detection research. To address the abovementioned issues, in this paper, we propose a robust infrared small target detection method that joins multiple pieces of information and noise predictions, named MINP-Net. Specifically, we first design a gradient and contextual information extraction module to extract multiscale features from an input infrared image. Second, we construct a noise prediction network to model the background noise. Third, we plan a regional positioning branch to provide a coarse target location to decrease the false alarm ratio. In addition, we build a new infrared small target detection benchmark to advance the research in this field, named the NCHU-Seg dataset. To the best of our knowledge, the NCHU-Seg dataset is the largest real-world scene dataset for evaluating infrared small target segmentation methods. For a comprehensive evaluation, we compare our method with some of the state-of-the-art methods on both the well-known NUAA-SIRST dataset and our NCHU-Seg dataset. The experimental results demonstrate that the proposed MINP-Net method performs better in terms of detection effectiveness and segmentation accuracy and effectively balances the detection probabilities and false alarms with complex backgrounds. (The code and dataset are available at https://github.com/PCwenyue.).
Siqiang Meng, Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004
IEEE Trans. Geosci. Remote. Sens.4
2023 Whole brain segmentation method from 2.5D brain MRI slice image based on Triple U-Net
Xingyan Chen, Shaofeng Jiang, Lanting Guo, Zhen Chen 0004, Congxuan Zhang
Vis. Comput.4
2022 Using optical flow algorithm based on dynamic illumination mode to examine defects on highly reflective turbine blade surface
abstract
Abstract Notion of optical flow literally refers to the displacements of intensity patterns. In that sense, extracting interested information from 2D scene is analogy to modulation/demodulation in random signal processing. To address the limitations presented in computer vision based on static image, we propose a novel metal component defect detection method, specified as the instance of turbine blade surface detection, using optical flow estimation.To start the specified pattern recognition in 2D presentation, we modulate the brightness constancy assumption equation as illumination varying model, by sampling the second image with function whose frequency was chosen according to the Nyquist sampling theorem, and a sinusoidal factor was introduced as an additive factor. This tunable channel based on 2D image transfers intensity features into optical modes. Then, we implement optical flow estimation on two sequential images. Experimental results reveal grayscale space shows completness in representing the optical modes of turbine blade with various kinds of surface characteristics. By modifying the index of information content, we propose quantitative index to evaluate the performance of our method. Evaluation reveals optical flow algorithm is qualified to examine defects on highly reflective turbine blade, and our method extends the application of optical flow.
Xuan Shang, Wenyan Song, Zhen Chen 0004, Congxuan Zhang
IET Image Process.3
2022 Parallel multiscale context-based edge-preserving optical flow estimation with occlusion detection
Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ming Li 0056
Signal Process. Image Commun.3
2022 Self-Attention-Based Multiscale Feature Learning Optical Flow With Occlusion Feature Map Prediction
abstract
Even though optical flow approaches based on convolutional neural networks have achieved remarkable performance with respect to both accuracy and efficiency, large displacements and motion occlusions remain challenges for most existing learning-based models. To address the abovementioned issues, we propose in this paper a self-attention-based multiscale feature learning optical flow computation method with occlusion feature map prediction. First, we exploit a self-attention mechanism-based multiscale feature learning module to compensate for large displacement optical flows, and the presented module is able to capture long-range dependencies from the input frames. Second, we design a simple but effective self-learning module to acquire an occlusion feature map, in which the predicted occlusion map is utilized to correct the optical flow estimation in occluded areas. Third, we explore a hybrid loss function that integrates the photometric and smoothness losses into the classical endpoint error (EPE)-based loss to ensure the accuracy and robustness of the presented network. Finally, we compare the proposed method with some state-of-the-art approaches using the MPI-Sintel and KITTI test databases. The experimental results demonstrate that the proposed method achieved competitive performance with respect to both accuracy and robustness, and it produced the better results compared to other methods under large displacements and motion occlusions.
Congxuan Zhang, Zhongkai Zhou, Zhen Chen 0004, Weiming Hu 0004, Ming Li 0056, Shaofeng Jiang
IEEE Trans. Multim.3
2021 Dense-CNN: Dense convolutional neural network for stereo matching using multiscale feature connection
Congxuan Zhang, Junjie Wu 0005, Zhen Chen 0004, Wen Liu 0004, Ming Li 0056, Shaofeng Jiang
Signal Process. Image Commun.3
2020 STDC-Flow: large displacement flow field estimation using similarity transformation-based dense correspondence
abstract
In order to improve the accuracy and robustness of optical flow computation under large displacements and motion occlusions, the authors present in this study a large displacement flow field estimation approach using similarity transformation‐based dense correspondence, named STDC‐Flow approach. First, the authors compute an initial nearest‐neighbour field by using the STDC‐Flow of the consecutive two frames, and then extract the consistent regions as the robust nearest‐neighbour field and label the inconsistent regions as the occlusion areas. Second, they improve a non‐local total variation with the L 1 norm optical flow model by using the occlusion information to modify the weighted median filtering optimisation. Third, they fuse the robust nearest‐neighbour field and the computed flow field of the improved variational optical flow model to construct the final flow field by using the quadratic pseudo‐boolean optimisation fusion algorithm. Finally, the authors compare the proposed STDC‐Flow method with several state‐of‐the‐art approaches including the variational and deep learning‐based optical flow models by using the MPI‐Sintel and KITTI evaluation databases. The comparison results demonstrate that the proposed STDC‐Flow method has a high accuracy for flow field computation, especially the capacity of dealing with large displacements and motion occlusions.
Congxuan Zhang, Zhen Chen 0004, Fan Xiong, Wen Liu 0004, Ming Li 0056, Liyue Ge
IET Comput. Vis.2
2020 Refined TV-L1 Optical Flow Estimation Using Joint Filtering
abstract
Though the accuracy and robustness of optical flow has been dramatically enhanced over the past few years, the issue of edge-blurring near the image and motion boundaries has remained a challenge in flow field estimation. In this paper, we propose a refined total variation withL1norm (TV-L1) optical flow estimation approach using joint filtering, named JOF. First, we divide the image into three categorized regions: mutual-structure regions, inconsistent structure regions, and smooth regions. The mutual-structure guided filter for optical flow estimation is constructed by extracting the mutual-structure regions of the flow field. Second, the refined TV-L1optical flow model is proposed by incorporating the non-local term and mutual-structure guided filter objective function into the classical TV-L1energy function. Furthermore, the novel TV-L1optical flow objective function is minimized using a joint filtering program composed of a weighted median filter and a mutual-structure guided filter to optimize the estimated flow field during the coarse-to-fine optical flow computation scheme. Finally, we compare the proposed JOF method with several state-of-the-art approaches including variational and deep learning based optical flow models using the Middlebury, MPI-Sintel, and UCF101 test databases. The evaluation results indicate that the proposed method has high accuracy and good robustness for flow field computation and, especially, the significant benefit of edge-preserving.
Congxuan Zhang, Liyue Ge, Zhen Chen 0004, Ming Li 0056, Wen Liu 0004, Hao Chen 0060
IEEE Trans. Multim.3
2017 Robust Non-Local TV-L1 Optical Flow Estimation With Occlusion Detection
abstract
In this paper, we propose a robust non-local TV-$L^{1}$optical flow method with occlusion detection to address the problem of weak robustness of optical flow estimation with motion occlusion. First, a TV-$L^{1}$form for flow estimation is defined using a combination of the brightness constancy and gradient constancy assumptions in the data term and by varying the weight under the Charbonnier function in the smoothing term. Second, to handle the potential risk of the outlier in the flow field, a general non-local term is added in the TV-$L^{1}$optical flow model to engender the typical non-local TV-$L^{1}$form. Third, an occlusion detection method based on triangulation is presented to detect the occlusion regions of the sequence. The proposed non-local TV-$L^{1}$optical flow model is performed in a linearizing iterative scheme using improved median filtering and a coarse-to-fine computing strategy. The results of the complex experiment indicate that the proposed method can overcome the significant influence of non-rigid motion, motion occlusion, and large displacement motion. Results of experiments comparing the proposed method and existing state-of-the-art methods by, respectively, using Middlebury and MPI Sintel database test sequences show that the proposed method has higher accuracy and better robustness.
Congxuan Zhang, Zhen Chen 0004, Mingrun Wang, Ming Li 0056, Shaofeng Jiang
IEEE Trans. Image Process.2
2014 Real-time brain extraction method from cerebral MRI volume based on graphic processing units
Shaofeng Jiang, Yu Wang 0030, Zhen Chen 0004, Kaiqiong Sun
Neural Comput. Appl.3
2011 Design and Implement of the Intelligent Network Caring System
Zhen Chen 0004, Yu Wang 0030
WASA3