VLDB 2026 Research / reviewers in the wild / expert
Mingliang Zhai
dblp:221/3748
· DBLP profile ↗
35ranked-venue papers
20as first author
27since 2021 · last 2025
0000-0003-2609-9961ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 15 first-author · 18 since 2021Artificial intelligence and machine learning · 12 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | World Knowledge-Enhanced Reasoning Using Instruction-Guided Interactor in Autonomous DrivingabstractThe Multi-modal Large Language Models (MLLMs) with extensive world knowledge have revitalized autonomous driving, particularly in reasoning tasks within perceivable regions. However, when faced with perception-limited areas (dynamic or static occlusion regions), MLLMs struggle to effectively integrate perception ability with world knowledge for reasoning. These perception-limited regions can conceal crucial safety information, especially for vulnerable road users. In this paper, we propose a framework, which aims to improve autonomous driving performance under perception-limited conditions by enhancing the integration of perception capabilities and world knowledge. Specifically, we propose a plug-and-play instruction-guided interaction module that bridges modality gaps and significantly reduces the input sequence length, allowing it to adapt effectively to multi-view video inputs. Furthermore, to better integrate world knowledge with driving-related tasks, we have collected and refined a large-scale multi-modal dataset that includes 2 million natural language QA pairs, 1.7 million grounding task data. To evaluate the model’s utilization of world knowledge, we introduce an object-level risk assessment dataset comprising 200K QA pairs, where the questions necessitate multi-step reasoning leveraging world knowledge for resolution. Extensive experiments validate the effectiveness of our proposed method. Mingliang Zhai, Zengyuan Guo, Ningrui Yang, Xiameng Qin, Sanyuan Zhao, Junyu Han, Ji Tao, Yuwei Wu 0001, Yunde Jia |
AAAI | 1 |
| 2025 | Mamba-SF: Monocular Scene Flow Learning with State Space ModelsabstractMonocular scene flow estimation has been a long-standing problem in computer vision. Methods based on the RAFT architecture are currently the mainstream approaches, while often overlooking the long-range dependencies in motion and texture features and fail to fully utilize the spatial information in texture features. In this paper, we consider that using Transformers introduces high computational complexity. Therefore, we propose the Mamba Motion Module based on State Space Models design, which first models long-range dependencies in motion and texture features and then fully leverages the spatial information in texture features to enhance the motion features, generating global motion features while maintaining low computational complexity. Additionally, texture features play a significant role in constraining motion boundaries. Therefore, we propose the Enhanced Texture Module, which predicts a set of channel weights from the global motion features to enrich the channel properties of the texture features and concatenates texture features with the global motion features along the channels, thereby improving scene flow accuracy. Experimental results show that our method achieves highly competitive results on the KITTI 2015 and Eigen Split datasets, increasing by 18.82% and 2.15% compared to the baseline, respectively. Xuezhi Xiang, Xianye Ben, Insha Hassan, Mingliang Zhai, Lei Zhang 0093, Xiantong Zhen |
ICIP | 5 |
| 2025 | FGRFlow: Learning Fine-Grained Rigidity Scene Flow from 4D Radar Point CloudabstractScene flow estimation using 4D millimeter-wave radar has emerged as a prominent research focus for 3D dynamic perception. However, compared to LiDAR point clouds, the drastic sparsity of radar point clouds poses challenges in enforcing local rigidity constraints, which are crucial for accurate 3D motion estimation. To address this issue, we propose a novel Gaussian-based pseudo-point generation method that fully leverages two distinct yet complementary data modalities, 3D coordinates and Doppler velocity, to support multi-body rigidity assumptions, effectively capturing fine-grained and structured motion patterns from highly sparse radar point clouds. Furthermore, a velocity calibration mechanism is designed to improve the reliability of fine-grained rigid motion velocity estimation. In addition, a progressive fusion strategy is introduced to systematically integrate fine-grained rigid motion priors at multiple levels, enhancing the robustness of matching costs and motion features while effectively compensating for coarse flows. Experimental results on real-world radar scans from the View-of-Delft (VoD) dataset demonstrate the promising performance of our FGRFlow compared to other leading 4D radar-based approaches, validating the advantages of our design choices. Mingliang Zhai, Haidong Hu, Chi-Man Pun, Hao Gao 0005 |
ACM Multimedia | 1 |
| 2025 | Learning optical flow from spiking camera with direction disassemblyabstractAbstract Conventional optical flow estimation methods typically recover two‐dimensional motion from RGB image sequences. Recently, due to the rise and widespread use of spike cameras, learning optical flow from spiking cameras has become a hot topic in the field of two‐dimensional motion estimation. Although existing methods have been designed to learn optical flow by designing feature processing methods for spike streams, there is still insufficient consideration for flow field post‐processing, resulting in limited accuracy of optical flow estimation. To address this problem, an optical flow estimation method based on directional disassembly is proposed. Specifically, the estimated flow fields along the horizontal and vertical directions are disassembled and the motion vectors along the two directions are denoised separately to reduce the burden of post‐processing for complex two‐dimensional motion information. In addition, contextual information is introduced in the post‐processing so that the scene information can effectively contribute to the results of the flow post‐processing. Experimental results show that this proposed method is capable of achieving comparable performance on spike‐based public datasets. Mingliang Zhai, Xuezhi Xiang, Kang Ni, Hao Gao 0005 |
IET Image Process. | 1 |
| 2025 | Multi-object tracking with scale-aware transformer and enhanced association strategy
Xuezhi Xiang, Xiankun Zhou, Mingliang Zhai, Abdulmotaleb El Saddik |
Multim. Syst. | 4 |
| 2025 | Frequency Decoupled Masked Auto-Encoder for Self-Supervised Skeleton-Based Action RecognitionabstractIn 3D skeleton-based action recognition, the limited availability of supervised data has driven interest in self-supervised learning methods. The reconstruction paradigm using masked auto-encoder (MAE) is an effective and mainstream self-supervised learning approach. However, recent studies indicate that MAE models tend to focus on features within a certain frequency range, which may result in the loss of important information. To address this issue, we propose a frequency decoupled MAE. Specifically, by incorporating a scale-specific frequency feature reconstruction module, we delve into leveraging frequency information as a direct and explicit target for reconstruction, which augments the MAE's capability to discern and accurately reproduce diverse frequency attributes within the data. Moreover, in order to address the issue of unstable gradient updates caused by more complex optimization objectives with frequency reconstruction, we introduce a dual-path network combined with an exponential moving average (EMA) parameter updating strategy to guide the model in stabilizing the training process. We have conducted extensive experiments which have demonstrated the effectiveness of the proposed method. Ye Liu 0005, Tianhao Shi, Mingliang Zhai, Jun Liu 0036 |
IEEE Signal Process. Lett. | 3 |
| 2025 | DMRFlow: 4D Radar Scene Flow Estimation With Decoupled Matching and RefinementabstractScene flow estimation from 4D radar sensors has become increasingly popular in recent years. In this paper, we propose a matching and refinement decoupling method to estimate scene flow from 4D radar point clouds. Since 4D radar point clouds are much sparser and noisier than LiDAR point clouds, it is challenging to effectively establish correspondences between two frames and properly refine flow fields in the 3D space. To address this issue, we present decoupled correlation fields and decoupled flow fields for scene flow estimation, named DMRFlow. On the one hand, we propose a position-velocity decoupled matching approach that decouples the positional features from the velocity features of two adjacent point clouds and matches them separately. On the other hand, we design a dynamic-static decoupled refinement approach that splits initial flow fields into two groups according to motion segmentation maps and refines them separately. By integrating the matching and refinement decoupling method, our DMRFlow is able to effectively reduce mutual interference between different features during the matching and refinement process. We evaluate the proposed approach on the View-of-Delft (VoD) dataset. Experimental results show that DMRFlow yields competitive performance in autonomous driving scenarios compared to recent 4D radar scene flow estimation methods. Mingliang Zhai, Bing-Kun Bao, Xuezhi Xiang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Scene Flow Estimation for Autonomous Driving via Correlation Compensation and Initial Motion CheckabstractScene flow estimation from LiDAR sensors is a crucial task for dynamic environmental perception in autonomous driving scenarios. Recently, self-supervised approaches have gained attention for their ability to reduce the burden of point-wise annotation. Although existing methods have been able to generate initial flow fields by constructing point-to-point correspondences between adjacent frames of point clouds, the reliability of correlation extraction and initial motion measurement has not been adequately considered. To address this problem, we propose a novel deep neural network to estimate scene flow from LiDAR sensors. Unlike previous works, our approach incorporates a Statistical-based Correlation Compensation Module (SCCM) that leverages statistical features to capture more reasonable correspondences. Furthermore, we design a Holistic Correlation Compensation Module (HCCM) to capture the overall correspondence that reflects most of the rigid motion in dynamic environments. In addition, an Initial Motion Check Mechanism (IMCM) is introduced to calibrate the initial flow and provide more reliable motion priors for subsequent flow refinement. Extensive experimental results on public scene flow benchmarks show that the proposed approach achieves competitive performance in autonomous driving scenarios. Mingliang Zhai, Bing-Kun Bao, Xuezhi Xiang |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | An Automatic Assessment of Parkinson's Disease in Arising from Chair Task via Refined Diffusion-based Pose EstimatorabstractParkinson’s disease (PD) is a progressively common neurodegenerative disorder characterized by a decline in motor function. The diagnosis of PD typically relies on the Movement Disorder Society-Unified Parkinson’s Disease Rating Scale (MDS-UPDRS), which involves subjective scoring through observation of targeted movements. However, this objective method heavily depends on professional experience and has relatively high misdiagnosis rates. In this paper, we introduce a novel vision-based architecture for automated assessment of the ‘arising from chair’ task, which is one of the key MDS-UPDRS components. First, a diffusion-based 2D pose estimator is proposed to enhance keypoint accuracy by iteratively learning the distribution of ground-truth data and then denoising noisy poses. Second, a keypoint trajectory refinement network is introduced to eliminate the jitter error by considering motion information such as position, velocity, acceleration, and jerk. Finally, based on the predicted skeleton keypoint trajectories, we propose several objective indicators to assess the movement characteristics and perform the final rating using the classifier. The experiment substantiates the proposed algorithm, achieving a precision of 98.7% and an accuracy of 95.8% in classifying the ‘arising from chair’ task. Furthermore, the classification results and the proposed objective indicators have been validated as effective aids for neurologists to provide more precise diagnoses. Chi-Man Pun, Haolun Li 0001, Mingliang Zhai, Feng Xu 0005, Hao Gao 0005 |
BIBM | 4 |
| 2024 | Compositional Substitutivity of Visual Reasoning for Visual Question Answering
Chuanhao Li 0001, Zhen Li 0026, Chenchen Jing, Yuwei Wu 0001, Mingliang Zhai, Yunde Jia |
ECCV (48) | 5 |
| 2024 | In-Context Compositional Generalization for Large Vision-Language ModelsabstractRecent work has revealed that in-context learning for large language models exhibits compositional generalization capacity, which can be enhanced by selecting in-context demonstrations similar to test cases to provide contextual information.However, how to exhibit in-context compositional generalization (ICCG) of large vision-language models (LVLMs) is non-trival.Due to the inherent asymmetry between visual and linguistic modalities, ICCG in LVLMs faces an inevitable challenge-redundant information on the visual modality.The redundant information affects in-context learning from two aspects: (1) Similarity calculation may be dominated by redundant information, resulting in sub-optimal demonstration selection.(2) Redundant information in in-context demonstrations brings misleading contextual information to in-context learning.To alleviate these problems, we propose a demonstration selection method to achieve ICCG for LVLMs, by considering two key factors of demonstrations: content and structure, from a multimodal perspective.Specifically, we design a diversity-coverage-based matching score to select demonstrations with maximum coverage, and avoid selecting demonstrations with redundant information via their content redundancy and structural complexity.We build a GQA-ICCG dataset to simulate the ICCG setting, and conduct experiments on GQA-ICCG and the VQA v2 dataset.Experimental results demonstrate the effectiveness of our method. Chuanhao Li 0001, Chenchen Jing, Zhen Li 0026, Mingliang Zhai, Yuwei Wu 0001, Yunde Jia |
EMNLP | 4 |
| 2024 | Self-Supervised Multi-Scale Hierarchical Refinement Method for Joint Learning of Optical Flow and DepthabstractRecurrently refining the optical flow based on a single high-resolution feature demonstrates high performance. We exploit the strength of this strategy to build a novel architecture for the joint learning of optical flow and depth. Our pro-posed architecture is improved to work in the case of training on unlabeled data, which is extremely challenging. The loss is computed for the iterations carried out over a single high-resolution feature, where the reconstruction loss fails to optimize the accuracy particularity in occluded regions. Therefore, we propose to hierarchically refine the optical flow across multiple scales while feeding the rigid flow calculated from depth and camera pose to provide more refinement. We further propose a self-supervised patch-based similarity loss to be optimized with the reconstruction loss to improve accuracy in the occluded regions. Our proposed method demonstrates efficient performance on the KITTI 2015 dataset, with more improvement in the occluded regions. Rokia Abdein, Xuezhi Xiang, Mingliang Zhai, Abdulmotaleb El Saddik |
ICASSP | 4 |
| 2024 | Visual-Guided Reasoning Path Generation for Visual Question Answering
Chenchen Jing, Mingliang Zhai, Yuwei Wu 0001, Yunde Jia |
PRCV (1) | 3 |
| 2024 | Scene flow estimation from 3D point clouds based on dual-branch implicit neural representationsabstractAbstract Recently, online optimisation‐based scene flow estimation has attracted significant attention due to its strong domain adaptivity. Although online optimisation‐based methods have made significant advances, the performance is far from satisfactory as only flow priors are considered, neglecting scene priors that are crucial for the representations of dynamic scenes. To address this problem, the authors introduce a dual‐branch MLP‐based architecture to encode implicit scene representations from a source 3D point cloud, which can additionally synthesise a target 3D point cloud. Thus, the mapping function between the source and synthesised target 3D point clouds is established as an extra implicit regulariser to capture scene priors. Moreover, their model infers both flow and scene priors in a stronger bidirectional manner. It can effectively establish spatiotemporal constraints among the synthesised, source, and target 3D point clouds. Experiments on four challenging datasets, including KITTI scene flow, FlyingThings3D, Argoverse, and nuScenes, show that our method can achieve potential and comparable results, proving its effectiveness and generality. Mingliang Zhai, Kang Ni, Jiucheng Xie, Hao Gao 0005 |
IET Comput. Vis. | 1 |
| 2024 | GloFP-MSF: monocular scene flow estimation with global feature perception
Xuezhi Xiang, Mingliang Zhai, Abdulmotaleb El Saddik |
Multim. Syst. | 4 |
| 2024 | Learning graph-based representations for scene flow estimation
Mingliang Zhai, Hao Gao 0005, Ye Liu 0005, Jianhui Nie, Kang Ni |
Multim. Tools Appl. | 1 |
| 2023 | Spike-Based Optical Flow Estimation Via Contrastive LearningabstractSpiking cameras have shown promising advantages for optical flow estimation in high-speed scenarios. The recent work SCFlow [1] attempts to train an optical flow model using spike frames based on a multi-scale flow reconstruction loss. However, only using the flow reconstruction loss is unable to effectively deal with the details of motion, which may lead to noise and blur in the estimated flow fields. To address this issue, we introduce a contrastive loss into spike-based optical flow estimation, which exploits both the information of positive samples and negative samples. Moreover, we propose a refinement step with flexible reception fields to effectively refine the initial flow fields. Experiments on the spiking optical flow dataset PHM demonstrate that the proposed network is effective for spike-based optical flow estimation. In addition, our method achieves competitive performance compared to recent spike-based, frame-based, and event-based methods. Mingliang Zhai, Kang Ni, Jiucheng Xie, Hao Gao 0005 |
ICASSP | 1 |
| 2023 | Cross-Modal Optical Flow Estimation via Modality Compensation and AlignmentabstractCross-modal optical flow estimation aims to predict motion fields between two frames collected from different modalities, recently attracting intensive attention. However, a substantial yet challenging problem is how to match images across a large modal discrepancy. In this paper, we propose a modality compensation module (MCM) to extract complementary features from different modalities adaptively. Moreover, a cross-modal feature alignment loss is introduced into our network, pulling the compensative features of two cross-modal frames closer and effectively reducing the modal discrepancy. The experimental results demonstrate that our method can achieve competitive performance on the cross-modal optical flow dataset CrossKITTI. Moreover, we experimentally verify that the proposed MCM and cross-modal feature alignment loss are effective for cross-modal optical flow estimation. Mingliang Zhai, Kang Ni, Jiucheng Xie, Hao Gao 0005 |
ICASSP | 1 |
| 2023 | Learning Scene Flow from 3d Point Clouds with Cross-Transformer and Global Motion CuesabstractScene flow estimation is critical for real-world vision problems such as autonomous driving and augmented reality. Due to the popularity of 3D LiDAR sensors, scene flow estimation from 3D point clouds arouses increasing attention. Existing methods usually use a flow embedding-based layer to find correspondences between point pairs. However, only using a flow embedding-based layer is not enough to model the global mutual relationship between two features due to local matching. In this paper, we introduce a cross-transformer to capture more reliable dependencies for point pairs. Moreover, a global motion-aware module is adopted to learn large displacements with a non-local approach. The experimental results demonstrate that the proposed method achieves comparable performance on public datasets and confirm the effectiveness of exploiting the cross-transformer and global motion cues for scene flow estimation. Mingliang Zhai, Kang Ni, Jiucheng Xie, Hao Gao 0005 |
ICASSP | 1 |
| 2023 | Scene Flow Estimation from Point Clouds with Contrastive Loss and Dual Pseudo LabelsabstractScene flow estimation aims to extract the 3D motion vector between each surface point in two consecutive point clouds. Pseudo-label-based approaches usually exploit point-to-point relations and 3D geometry information to generate the pseudo label for self-supervised learning. However, unreasonable results are still obtained due to the unexploited information of negative samples. Moreover, previous approaches are limited by the fact that pseudo labels are only generated along the forward direction, ignoring the backward direction that has strong spatiotemporal correlations with the forward direction. In this paper, we address these issues in a simple yet effective manner. Specifically, we introduce a contrastive loss to exploit both the information of positive samples and negative samples. Furthermore, we design a dual pseudo labels generation strategy to provide a bidirectional self-supervision for scene flow estimation. Experiments on FlyingThings3D and KITTI datasets show that our method can achieve competitive performance compared to recent self-supervised methods. Mingliang Zhai, Kang Ni, Jiucheng Xie, Xuezhi Xiang, Hao Gao 0005 |
ICIP | 1 |
| 2023 | Fast-StrucTexT: An Efficient Hourglass Transformer with Modality-guided Dynamic Token Merge for Document UnderstandingabstractTransformers achieve promising performance in document understanding because of their high effectiveness and still suffer from quadratic computational complexity dependency on the sequence length. General efficient transformers are challenging to be directly adapted to model document. They are unable to handle the layout representation in documents, e.g. word, line and paragraph, on different granularity levels and seem hard to achieve a good trade-off between efficiency and performance. To tackle the concerns, we propose Fast-StrucTexT, an efficient multi-modal framework based on the StrucTexT algorithm with an hourglass transformer architecture, for visual document understanding. Specifically, we design a modality-guided dynamic token merging block to make the model learn multi-granularity representation and prunes redundant tokens. Additionally, we present a multi-modal interaction module called Symmetry Cross-Attention (SCA) to consider multi-modal fusion and efficiently guide the token mergence. The SCA allows one modality input as query to calculate cross attention with another modality in a dual phase. Extensive experiments on FUNSD, SROIE, and CORD datasets demonstrate that our model achieves the state-of-the-art performance and almost 1.9x faster inference time than the state-of-the-art methods. Mingliang Zhai, Yulin Li 0004, Xiameng Qin, Qunyi Xie, Chengquan Zhang, Yuwei Wu 0001, Yunde Jia |
IJCAI | 1 |
| 2023 | DJSPNet: Deep Joint Statistical-Spatial Pooling Network for High-Resolution SAR Image ClassificationabstractThe previous approaches based on statistical features or spatial features have achieved promising performance on pixel-wise high-resolution (HR) synthetic aperture radar (SAR) image classification, but these methods always cannot capture local spatial features and global statistical properties efficiently because of the complex spatial structural patterns and statistical nature in SAR patches. Inspired by this, we propose a deep joint statistical–spatial pooling network (DJSPNet), for HR SAR image classification, which combines a group second-order statistical feature learning (GSFL) block and an efficient feature-fusion style (EFS) into an end-to-end feature learning block. GSFL block is designed with a group second-order feature learning method in two steps, where the first step divides convolutional channels into several semantic groups. The second step collects second-order feature statistics by calculating pairwise feature interactions within each group. EFS models second-order attentional statistics between statistical characteristics and spatial features by polynomial kernel approximation and guides the discriminative feature activations in SAR patches. More specifically, both GSFL and EFS are stacked and plugged into the encoder stage of conventional U-Net for distinguishable feature learning. Experimental results suggest that the proposed DJSPNet gives better classification performance compared with related deep feature learning networks on a real TerraSAR-X dataset. Kang Ni, Mingliang Zhai, Minrui Zou, Qianqian Wu 0009, Peng Wang 0030 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Flow Learning Based Dual Networks for Low-Light Image Enhancement
Changhui Hu 0001, Weilin Yi, Ziyun Cai, Mingliang Zhai, Wankou Yang |
Neural Process. Lett. | 5 |
| 2022 | Synthesizing Counterfactual Samples for Overcoming Moment Biases in Temporal Video Grounding
Mingliang Zhai, Chuanhao Li 0001, Chenchen Jing, Yuwei Wu 0001 |
PRCV (1) | 1 |
| 2022 | 3D Point Convolutional Network for Dense Scene Flow Estimation
Xuezhi Xiang, Rokia Abdein, Mingliang Zhai, Ning Lv 0001 |
Neural Process. Lett. | 3 |
| 2021 | Geometry understanding from autonomous driving scenarios based on feature refinement
Mingliang Zhai, Xuezhi Xiang |
Neural Comput. Appl. | 1 |
| 2021 | Optical flow and scene flow estimation: A survey
Mingliang Zhai, Xuezhi Xiang, Ning Lv 0001 |
Pattern Recognit. | 1 |
| 2020 | Multi-Task Learning in Autonomous Driving Scenarios Via Adaptive Feature Refinement NetworksabstractMany deep learning applications benefit from multi-task learning with several related objectives. In autonomous driving scenarios, being able to accurately infer motion and spatial information is essential for scene understanding. In this paper, we combine an adaptive feature refinement module and a unified framework for joint learning of optical flow, depth and camera pose estimation in an unsupervised manner. The feature refinement module is embedded into motion estimation and depth prediction sub-networks, which can exploit more channel-wise relationships and contextual information for feature learning. Given a monocular video, our network firstly estimates depth and camera motion, and calculates rigid optical flow. Then, we design an auxiliary flow network for inferring non-rigid flow fields. In addition, a forward-backward consistency check is adopted for occlusion reasoning. Extensive experiments on KITTI dataset demonstrate that the proposed method achieves potential results comparing to recent deep learning networks. Mingliang Zhai, Xuezhi Xiang, Ning Lv 0001, Abdulmotaleb El Saddik |
ICASSP | 1 |
| 2020 | A CNNs-based method for optical flow estimation with prior constraints and stacked U-Nets
Xuezhi Xiang, Mingliang Zhai, Rongfang Zhang, Yulong Qiao, Abdulmotaleb El Saddik |
Neural Comput. Appl. | 2 |
| 2020 | Dual-Path Part-Level Method for Visible-Infrared Person Re-identification
Xuezhi Xiang, Ning Lv 0001, Mingliang Zhai, Rokia Abdein, Abdulmotaleb El Saddik |
Neural Process. Lett. | 3 |
| 2020 | Optical Flow Estimation Using Dual Self-Attention Pyramid NetworksabstractRecently, optical flow estimation benefits greatly from deep learning based techniques. Most approaches use encoder-decoder architecture (U-Net) or spatial pyramid network (SPN) to learn optical flow. Both U-Net and SPN can extract multi-scale features and can predict optical flow directly. However, existing networks ignore to exploit the global information among channel features and inter-spatial relationship of features. In this paper, we propose a dual self-attention pyramid network, which adaptively integrates local features with their global dependencies and focuses on important features and suppresses unimportant features. Specifically, we introduce two types of attention modules into SPN, which emphasizes meaningful features along channel and spatial axes. The channel attention can adaptively re-weight channel-wise features by considering interdependencies among channels. Moreover, the spatial attention can utilize global contextual information to emphasize or suppress features in different spatial locations. In addition, two attention modules are embedded into each pyramidal level, which can refine features at different scale. We evaluate our method on MPI-Sintel and KITTI. The experimental results show that using the dual self-attention module can improve the representation power of network and further increase the accuracy of optical flow estimation. Mingliang Zhai, Xuezhi Xiang, Rongfang Zhang, Ning Lv 0001, Abdulmotaleb El Saddik |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | An Object Context Integrated Network for Joint Learning of Depth and Optical FlowabstractSupervised depth prediction and optical flow estimation have achieved promising performance due to the advanced deep network architectures. Since the ground truths are difficult to be collected, many recent works try to learn the depth and flow in an unsupervised manner. However, existing methods only use features from convolutional layers or a simple aggregation of multi-level features to predict the depth and flow maps, which is insufficient to exploit context information. In this paper, we attempt to exploit object contextual information and investigate the effect of the object context for joint learning of depth and optical flow. Specifically, we present a novel combination of object context and the framework of joint learning depth and optical flow. Our proposed network can exploit and integrate the object context for both tasks by aggregating the context according to pair-wise similarities. Furthermore, we adopt the existing spatial pyramid network (SPN) to estimate the depth and flow in a coarse-to-fine strategy effectively. Given temporally adjacent stereo pairs, our network can be trained end-to-end in an unsupervised manner and can predict the depth and flow maps simultaneously. We conduct experiments on two publicly available datasets, KITTI2012 and KITTI2015. Our proposed approach yields comparable performance on both depth and flow tasks, compared to the recent deep learning-based approaches. Experimental results demonstrate that exploiting object contextual information is useful and beneficial for depth and optical flow estimation. Mingliang Zhai, Xuezhi Xiang, Ning Lv 0001, Abdulmotaleb El Saddik |
IEEE Trans. Image Process. | 1 |
| 2019 | Ad-net: Attention Guided Network for Optical Flow Estimation Using Dilated ConvolutionabstractVariational models for optical flow estimation usually define an energy function that contains prior assumptions to explore rudimentary statistics of images. However, such methods cannot learn motion knowledge from the pre-prepared data and have many parameters that need to be set manually. Nowadays, convolutional neural networks (CNNs) have been used in optical flow estimation successfully, which can learn weights from the training dataset and can predict optical flow end-to-end. In this paper, we propose an attention guided network for learning optical flow, named AD-Net, which contains several attention units for modelling the relativities between the channels. Further, we introduce dilated convolution into supervised network for reducing the loss of motion details. In addition, some prior auxiliary constraints are embedded in the supervised network as auxiliary loss terms. Our proposed approach is tested on MPI-Sintel and KITTI2012 datasets and can preserve motion edges and details effectively. Mingliang Zhai, Xuezhi Xiang, Rongfang Zhang, Ning Lv 0001, Abdulmotaleb El Saddik |
ICASSP | 1 |
| 2019 | Optical Flow Estimation Using Spatial-Channel Combinational Attention-Based Pyramid NetworksabstractRecently, learning to estimate optical flow via deep convolutional networks is attracting significant attention. In this paper, we introduce a spatial-channel attention module into optical flow estimation, which infers attention maps along two separated dimensions, channel and spatial, and then integrates these separated attention maps into a fusion attention map for feature refinement. We embed this module into spatial pyramid network, which can adaptively learn the channel and spatial attention maps at each level for modifying the different scaled features and can further improve the accuracy of optical flow estimation. Our network is trained on FlyingChairs and FlyingThings3D datasets with a supervised manner, and is further tested on MPI-Sintel benchmark. The experimental results show that using the spatial-channel attention unit is beneficial for dense flow estimation and our approach is comparable with the state-of-the-art methods. Xuezhi Xiang, Mingliang Zhai, Rongfang Zhang, Ning Lv 0001, Abdulmotaleb El Saddik |
ICIP | 2 |
| 2019 | Optical flow estimation using channel attention mechanism and dilated convolutional neural networks
Mingliang Zhai, Xuezhi Xiang, Rongfang Zhang, Ning Lv 0001, Abdulmotaleb El Saddik |
Neurocomputing | 1 |