VLDB 2026 Research / reviewers in the wild / expert
Xuezhi Xiang
dblp:192/2034
· DBLP profile ↗
61ranked-venue papers
31as first author
44since 2021 · last 2026
0000-0002-6185-833XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 21 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 9 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Person search with deep learning
Ning Lv 0001, Xuezhi Xiang, Yulong Qiao, Abdulmotaleb El Saddik |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | UniAD: Unified cross-modal prompt regularization for zero-shot anomaly detection across domains
Xuezhi Xiang, Songran Luo, Yingjun Du, Lei Zhang 0093, Xiantong Zhen |
Neurocomputing | 1 |
| 2025 | LKA-ReID: Vehicle Re-Identification with Large Kernel AttentionabstractWith the rapid development of intelligent transportation systems and the popularity of smart city infrastructure, Vehicle Re-ID technology has become an important research field. The vehicle Re-ID task faces an important challenge, which is the high similarity between different vehicles. Existing methods use additional detection or segmentation models to extract differentiated local features. However, these methods either rely on additional annotations or greatly increase the computational cost. Using attention mechanism to capture global and local features is crucial to solve the challenge of high similarity between classes in vehicle Re-ID tasks. In this paper, we propose LKA-ReID with large kernel attention. Specifically, the large kernel attention (LKA) utilizes the advantages of self-attention and also benefits from the advantages of convolution, which can extract the global and local features of the vehicle more comprehensively. We also introduce hybrid channel attention (HCA), which combines channel attention with spatial information, so that the model can better focus on channels and feature regions, and ignore background and other disturbing information. Experiments on VeRi-776 dataset demonstrated the effectiveness of LKA-ReID, with mAP reaches 86.65% and Rank-1 reaches 98.03%. Xuezhi Xiang, Zhushan Ma, Lei Zhang 0093, Denis Ombati, Himaloy Himu, Xiantong Zhen |
ICASSP | 1 |
| 2025 | Mamba-SF: Monocular Scene Flow Learning with State Space ModelsabstractMonocular scene flow estimation has been a long-standing problem in computer vision. Methods based on the RAFT architecture are currently the mainstream approaches, while often overlooking the long-range dependencies in motion and texture features and fail to fully utilize the spatial information in texture features. In this paper, we consider that using Transformers introduces high computational complexity. Therefore, we propose the Mamba Motion Module based on State Space Models design, which first models long-range dependencies in motion and texture features and then fully leverages the spatial information in texture features to enhance the motion features, generating global motion features while maintaining low computational complexity. Additionally, texture features play a significant role in constraining motion boundaries. Therefore, we propose the Enhanced Texture Module, which predicts a set of channel weights from the global motion features to enrich the channel properties of the texture features and concatenates texture features with the global motion features along the channels, thereby improving scene flow accuracy. Experimental results show that our method achieves highly competitive results on the KITTI 2015 and Eigen Split datasets, increasing by 18.82% and 2.15% compared to the baseline, respectively. Xuezhi Xiang, Xianye Ben, Insha Hassan, Mingliang Zhai, Lei Zhang 0093, Xiantong Zhen |
ICIP | 2 |
| 2025 | Learning optical flow from spiking camera with direction disassemblyabstractAbstract Conventional optical flow estimation methods typically recover two‐dimensional motion from RGB image sequences. Recently, due to the rise and widespread use of spike cameras, learning optical flow from spiking cameras has become a hot topic in the field of two‐dimensional motion estimation. Although existing methods have been designed to learn optical flow by designing feature processing methods for spike streams, there is still insufficient consideration for flow field post‐processing, resulting in limited accuracy of optical flow estimation. To address this problem, an optical flow estimation method based on directional disassembly is proposed. Specifically, the estimated flow fields along the horizontal and vertical directions are disassembled and the motion vectors along the two directions are denoised separately to reduce the burden of post‐processing for complex two‐dimensional motion information. In addition, contextual information is introduced in the post‐processing so that the scene information can effectively contribute to the results of the flow post‐processing. Experimental results show that this proposed method is capable of achieving comparable performance on spike‐based public datasets. Mingliang Zhai, Xuezhi Xiang, Kang Ni, Hao Gao 0005 |
IET Image Process. | 3 |
| 2025 | Deep scene flow learning from point cloud with Transformer
Xuezhi Xiang, Rokia Abdein, Lei Zhang 0093, Xiantong Zhen |
Neurocomputing | 1 |
| 2025 | PVFT-Net: A point-voxel fusion method for self-supervised scene flow estimation with transformer
Xuezhi Xiang, Xiaoheng Li, Xiankun Zhou, Lei Zhang 0093, Xiantong Zhen |
Neurocomputing | 1 |
| 2025 | Vehicle re-identification with large separable kernel attention and hybrid channel attention
Xuezhi Xiang, Zhushan Ma, Xiaoheng Li, Lei Zhang 0093, Xiantong Zhen |
Image Vis. Comput. | 1 |
| 2025 | Scene flow estimation from point cloud based on grouped relative self-attention
Xuezhi Xiang, Xiankun Zhou, Yingxin Wei, Yulong Qiao |
Image Vis. Comput. | 1 |
| 2025 | Self-supervised monocular depth estimation with large kernel attention and dynamic scene perception
Xuezhi Xiang, Xiaoheng Li, Lei Zhang 0093, Xiantong Zhen |
J. Vis. Commun. Image Represent. | 1 |
| 2025 | Multi-object tracking with scale-aware transformer and enhanced association strategy
Xuezhi Xiang, Xiankun Zhou, Mingliang Zhai, Abdulmotaleb El Saddik |
Multim. Syst. | 1 |
| 2025 | DMRFlow: 4D Radar Scene Flow Estimation With Decoupled Matching and RefinementabstractScene flow estimation from 4D radar sensors has become increasingly popular in recent years. In this paper, we propose a matching and refinement decoupling method to estimate scene flow from 4D radar point clouds. Since 4D radar point clouds are much sparser and noisier than LiDAR point clouds, it is challenging to effectively establish correspondences between two frames and properly refine flow fields in the 3D space. To address this issue, we present decoupled correlation fields and decoupled flow fields for scene flow estimation, named DMRFlow. On the one hand, we propose a position-velocity decoupled matching approach that decouples the positional features from the velocity features of two adjacent point clouds and matches them separately. On the other hand, we design a dynamic-static decoupled refinement approach that splits initial flow fields into two groups according to motion segmentation maps and refines them separately. By integrating the matching and refinement decoupling method, our DMRFlow is able to effectively reduce mutual interference between different features during the matching and refinement process. We evaluate the proposed approach on the View-of-Delft (VoD) dataset. Experimental results show that DMRFlow yields competitive performance in autonomous driving scenarios compared to recent 4D radar scene flow estimation methods. Mingliang Zhai, Bing-Kun Bao, Xuezhi Xiang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Scene Flow Estimation for Autonomous Driving via Correlation Compensation and Initial Motion CheckabstractScene flow estimation from LiDAR sensors is a crucial task for dynamic environmental perception in autonomous driving scenarios. Recently, self-supervised approaches have gained attention for their ability to reduce the burden of point-wise annotation. Although existing methods have been able to generate initial flow fields by constructing point-to-point correspondences between adjacent frames of point clouds, the reliability of correlation extraction and initial motion measurement has not been adequately considered. To address this problem, we propose a novel deep neural network to estimate scene flow from LiDAR sensors. Unlike previous works, our approach incorporates a Statistical-based Correlation Compensation Module (SCCM) that leverages statistical features to capture more reasonable correspondences. Furthermore, we design a Holistic Correlation Compensation Module (HCCM) to capture the overall correspondence that reflects most of the rigid motion in dynamic environments. In addition, an Initial Motion Check Mechanism (IMCM) is introduced to calibrate the initial flow and provide more reliable motion priors for subsequent flow refinement. Extensive experimental results on public scene flow benchmarks show that the proposed approach achieves competitive performance in autonomous driving scenarios. Mingliang Zhai, Bing-Kun Bao, Xuezhi Xiang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Self-Supervised Multi-Scale Hierarchical Refinement Method for Joint Learning of Optical Flow and DepthabstractRecurrently refining the optical flow based on a single high-resolution feature demonstrates high performance. We exploit the strength of this strategy to build a novel architecture for the joint learning of optical flow and depth. Our pro-posed architecture is improved to work in the case of training on unlabeled data, which is extremely challenging. The loss is computed for the iterations carried out over a single high-resolution feature, where the reconstruction loss fails to optimize the accuracy particularity in occluded regions. Therefore, we propose to hierarchically refine the optical flow across multiple scales while feeding the rigid flow calculated from depth and camera pose to provide more refinement. We further propose a self-supervised patch-based similarity loss to be optimized with the reconstruction loss to improve accuracy in the occluded regions. Our proposed method demonstrates efficient performance on the KITTI 2015 dataset, with more improvement in the occluded regions. Rokia Abdein, Xuezhi Xiang, Mingliang Zhai, Abdulmotaleb El Saddik |
ICASSP | 2 |
| 2024 | Deep Optical Flow Learning With Deformable Large-Kernel Cross-AttentionabstractOptical flow estimation from image sequences is a fundamental problem in computer vision. In recent years, some methods have utilized Transformer to model global dependencies and improve optical flow, achieving impressive performance. However, in these methods, Transformers typically treat two-dimensional image features as one-dimensional sequences. While position encoding partially mitigates the loss of position information between different feature patches, Transformer still lacks inherent biases for modeling local visual patterns and tend to overlook channel characteristics in image features. Therefore, this paper introduces a deformable large kernel attention module, combining the strengths of convolution and attention mechanisms, which can preserve feature channel adaptability while modeling global dependencies without compromising the two-dimensional structure of features, significantly enhancing optical flow estimation. Additionally, the introduced deformable mechanism allows the model to adapt appropriately to different data patterns. Experimental results demonstrate that our optical flow estimation method achieves competitive results on publicly available benchmarks such as Sintel and KITTI. Xuezhi Xiang, Denis Ombati, Lei Zhang 0093, Xiantong Zhen |
ICIP | 1 |
| 2024 | Learning feature contexts by transformer and CNN hybrid deep network for weakly supervised person search
Ning Lv 0001, Xuezhi Xiang, Yulong Qiao, Abdulmotaleb El Saddik |
Comput. Vis. Image Underst. | 2 |
| 2024 | DBMHT: A double-branch multi-hypothesis transformer for 3D human pose estimation in video
Xuezhi Xiang, Xiaoheng Li, Weijie Bao, Yulong Qiao, Abdulmotaleb El Saddik |
Comput. Vis. Image Underst. | 1 |
| 2024 | A GCN and Transformer complementary network for skeleton-based action recognition
Xuezhi Xiang, Xiaoheng Li, Xuzhao Liu, Yulong Qiao, Abdulmotaleb El Saddik |
Comput. Vis. Image Underst. | 1 |
| 2024 | Self-supervised monocular depth estimation with self-distillation and dense skip connection
Xuezhi Xiang, Wei Li 0109, Abdulmotaleb El Saddik |
Comput. Vis. Image Underst. | 1 |
| 2024 | VIDF-Net: A Voxel-Image Dynamic Fusion method for 3D object detection
Xuezhi Xiang, Dianang Li, Xiankun Zhou, Yulong Qiao |
Comput. Vis. Image Underst. | 1 |
| 2024 | Temporal adaptive feature pyramid network for action detection
Xuezhi Xiang, Yulong Qiao, Abdulmotaleb El Saddik |
Comput. Vis. Image Underst. | 1 |
| 2024 | 3D hand pose estimation and reconstruction based on multi-feature fusion
Jiye Wang, Xuezhi Xiang, Abdulmotaleb El Saddik |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | GloFP-MSF: monocular scene flow estimation with global feature perception
Xuezhi Xiang, Mingliang Zhai, Abdulmotaleb El Saddik |
Multim. Syst. | 1 |
| 2024 | Deep Scene Flow Learning: From 2D Images to 3D Point CloudsabstractScene flow describes the 3D motion in a scene. It can be modeled as a single task or as a composite of the auxiliary tasks of depth, camera motion, and optical flow estimation. Deep learning's emergence in recent years has broadened the horizons for new methodologies in estimating these tasks, either as separate tasks or as joint tasks to reconstruct the scene flow. The sequence of images that are either synthesized or captured by a camera is used as input for these methods, which face the challenge of dealing with various situations in images to provide the most accurate motion, such as image quality. Nowadays, images have been superseded by point clouds, which provide 3D information, thereby expediting and enhancing the estimated motion. In this paper, we dig deeply into scene flow estimation in the deep learning era. We provide a comprehensive overview of the important topics regarding both image-based and point-cloud-based methods. In addition, we cover the methodologies for each category, highlighting the network architecture. Furthermore, we provide a comparison between these methods in terms of performance and efficiency. Finally, we conclude this survey with insights and discussions on the open issues and future research directions. Xuezhi Xiang, Rokia Abdein, Wei Li 0109, Abdulmotaleb El Saddik |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | InvFlow: Involution and multi-scale interaction for unsupervised learning of optical flow
Xuezhi Xiang, Rokia Abdein, Ning Lv 0001, Abdulmotaleb El Saddik |
Pattern Recognit. | 1 |
| 2023 | Deformable Cross Attention for Learning Optical FlowabstractOptical flow is the process of estimating motion in scenes. Each object in the scene has a homogeneous motion, i.e., moves in the same direction with the same velocity. Therefore, connecting the parts of an image globally provides an essential cue for learning accurate motion. Convolution-based methods estimate the motion features from the local regions, which miss this important cue. Recently, some methods used Transformer to model global dependencies to improve optical flow. However, Transformer suffers from excessive attention computations and still brings irrelevant parts into the region of interest. Therefore, we propose a deformable cross-attention for optical flow estimation, which provides two important advantages: connecting the parts of the image globally while deforming the attention to the objects’ shapes in the image and reducing the memory consumption. Our proposed method achieved competitive performance on Sintel and KITTI 2015 datasets in terms of accuracy and efficiency. Rokia Abdein, Xuezhi Xiang, Ning Lv 0001, Abdulmotaleb El Saddik |
ICASSP | 2 |
| 2023 | Transpointflow: Learning Scene Flow from Point Clouds with TransformerabstractScene flow estimation is the task of obtaining 3D motion from a dynamic scene. Due to the sparseness of point clouds, extracting features for a local group of points separately may result in different features that may all belong to the same object. This difference makes global correlation prone to producing an unacceptable flow. Local correlation restricts the algorithm to capturing limited movements and fails when fast movement or large deformation of an object occurs. Therefore, we propose a transformer-based scene flow method that can perform global feature modeling through a self-attention layer. Moreover, we propose a cross-attention-based flow embedding layer for global feature matching. We further propose a learnable attention-based up-sampling layer to up-sample the estimated flow to higher resolution based on a single feature scale, eliminating the need to model global dependencies at all scales. Experimental results show that our model produces competitive results on Flyingthings3D and KITTI datasets with efficient performance. Rokia Abdein, Xuezhi Xiang, Abdulmotaleb El Saddik |
ICIP | 2 |
| 2023 | Scene Flow Estimation from Point Clouds with Contrastive Loss and Dual Pseudo LabelsabstractScene flow estimation aims to extract the 3D motion vector between each surface point in two consecutive point clouds. Pseudo-label-based approaches usually exploit point-to-point relations and 3D geometry information to generate the pseudo label for self-supervised learning. However, unreasonable results are still obtained due to the unexploited information of negative samples. Moreover, previous approaches are limited by the fact that pseudo labels are only generated along the forward direction, ignoring the backward direction that has strong spatiotemporal correlations with the forward direction. In this paper, we address these issues in a simple yet effective manner. Specifically, we introduce a contrastive loss to exploit both the information of positive samples and negative samples. Furthermore, we design a dual pseudo labels generation strategy to provide a bidirectional self-supervision for scene flow estimation. Experiments on FlyingThings3D and KITTI datasets show that our method can achieve competitive performance compared to recent self-supervised methods. Mingliang Zhai, Kang Ni, Jiucheng Xie, Xuezhi Xiang, Hao Gao 0005 |
ICIP | 4 |
| 2023 | Global-aware and local-aware enhancement network for person search
Ning Lv 0001, Xuezhi Xiang, Yulong Qiao, Abdulmotaleb El Saddik |
Comput. Vis. Image Underst. | 2 |
| 2023 | EMHIFormer: An Enhanced Multi-Hypothesis Interaction Transformer for 3D human pose estimation in video
Xuezhi Xiang, Kaixu Zhang, Yulong Qiao, Abdulmotaleb El Saddik |
J. Vis. Commun. Image Represent. | 1 |
| 2023 | Engineering Vehicles Detection for Warehouse Surveillance System Based on Modified YOLOv4-Tiny
Xuezhi Xiang, Fanda Meng, Ning Lv 0001 |
Neural Process. Lett. | 1 |
| 2023 | Self-supervised learning of scene flow with occlusion handling through feature masking
Xuezhi Xiang, Rokia Abdein, Ning Lv 0001 |
Pattern Recognit. | 1 |
| 2022 | Transformer-Based Person Search Model with Symmetric Online Instance MatchingabstractPerson search is a challenging retrieval problem which aims at matching pedestrians with the same identity over non-overlapping camera views. In this paper, we adopt Swin Transformer as the backbone network to extract discriminative features. We propose a symmetric online instance matching loss which transfers the symmetric idea from KL divergence to the online instance matching loss. The purpose is to strengthen the robustness of the person search model under the condition of limited training identities. We compared with the state-of-the-arts on two mainstream benchmarks: CUHK-SYSU and PRW datasets. Experimental results demonstrate the effectiveness of our method. Especially, we achieve better performance on the PRW dataset with an improvement of 6.5% and 3.5% at the mAP and top-1 accuracy, respectively. Xuezhi Xiang, Ning Lv 0001, Yulong Qiao |
ICASSP | 1 |
| 2022 | Self-Supervised Learning of Optical Flow, Depth, Camera Pose and Rigidity Segmentation with Occlusion HandlingabstractIn this work, we propose a self-supervised scene flow framework for joint learning of optical flow, stereo depth, camera pose, and rigidity map and handle the occlusion during training. Specifically, we propose a feature masking method to alleviate the occlusion impact on the correlation result and reduce the outliers in both optical flow and depth map. We use the improved optical flow and depth to estimate the camera motion directly using the Perspective-n-Point method, which improves it accordingly. Furthermore, we recursively update the optical flow in both occluded and non-occluded regions with self-supervised cues learned from the rigid and optical flows. Reducing the error in the occluded regions enhances the rigidity map and improves the final optical flow accordingly. Our model achieved the state-of-the-art performance on KITTI 2015 benchmark for optical flow and produced competitive results for the depth, pose, and segmentation tasks. Rokia Abdein, Xuezhi Xiang, Ning Lv 0001 |
ICIP | 2 |
| 2022 | Efficient person search via learning-to-normalize deep representation
Ning Lv 0001, Xuezhi Xiang, Rokia Abdein |
Neurocomputing | 2 |
| 2022 | FCDNet: A Change Detection Network Based on Full-Scale Skip Connections and Coordinate AttentionabstractChange detection (CD) is an important means to monitor environmental changes on Earth. Recently, many methods based on deep learning have been proposed for CD tasks. However, existing CD methods still have difficulties in the detection of changing target edges and small changing targets. In this letter, we propose a novel CD method named FCDNet based on full-scale skip connections (FSC) and coordinate attention (CA). FCDNet makes full use of shallow information and high-level semantics through FSC between different levels of encoder features and decoder features, which can effectively alleviate the missed detection of small targets in the CD task. In addition, we design a multi-receptive field position enhancement module (MRPEM) based on CA. MRPEM enhances the local relationship of features through convolution operations of different kernel sizes, and establishes long-distance dependency of features with the application of CA, thus facilitate accurate detection of the edges of changing targets. We also introduce Depthwise Over-parameterized Convolutional Layer (DOConv) in our network architecture, which can improve model performance without increasing computational complexity during inference. The experimental results show that our method is comparable to state-of-the-art (SOTA) methods on the Season-Varying Change Detection (SVCD) dataset. Xuezhi Xiang, Dashuai Tian, Ning Lv 0001, Qiannan Yan |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Unsupervised optical flow estimation method based on transformer and occlusion compensation
Xuezhi Xiang, Rokia Abdein, Ning Lv 0001 |
Neural Comput. Appl. | 1 |
| 2022 | 3D Point Convolutional Network for Dense Scene Flow Estimation
Xuezhi Xiang, Rokia Abdein, Mingliang Zhai, Ning Lv 0001 |
Neural Process. Lett. | 1 |
| 2021 | Stable and Effective One-Step Method for Person SearchabstractPerson search, which requires both pedestrian detection and person re-identification, is a challenging computer vision task applied to real-world scenarios. The challenges faced by detection and re-identification, such as occlusion, poor illumination, confusing background, are still urgent for person search. In addition, one-step methods for person search need to deal with the divergence between two tasks. In this work, we propose an end-to-end model containing the feature extractor, the region proposal network, and the multi-task learning module. In order to process divergence between detection and re-identification, we introduce switchable normalization and gradient centralization to improve the stability of the model. To solve the imbalance problem of hard examples, we introduce focal loss as a classification loss in the multi-task learning module. The experimental results on two bench-marks, i.e., CUHK-SYSU and PRW, well demonstrate that our method outperforms the state-of-the-art one-step methods. Ning Lv 0001, Xuezhi Xiang, Rokia Abdein, Abdulmotaleb El Saddik |
ICASSP | 2 |
| 2021 | Graph wavelet transform for image texture classificationabstractAbstract Graph is a data structure that can represent complex relationships among data. Graph signal processing, unlike traditional signal processing, explicitly considers the structure and relationship among the signal samples. Graph wavelet transform can provide a multiscale analysis for the graph signal. It is well known that texture is a region property in an image, which is characterized with the intensity and relationship among pixels. In this context of the graph signal processing framework, an image texture can be considered as the signal on the graph. Therefore, a texture classification method based on graph wavelet transform is proposed. Specifically, image textures are decomposed into multiscale components by using two‐channel graph wavelet filter banks. Then the local singular value decomposition is applied to each subband. In order to improve the noise‐resistant ability, the maximum, mean and median values of the local singular values of graph‐wavelet transformation coefficients are extracted. Finally, the Weibull distributions are used to model those extracted values to describe the image textures. The experiments on the benchmark texture datasets are conducted to demonstrate the effectiveness of the proposed method. Yu-Long Qiao, Xuezhi Xiang |
IET Image Process. | 5 |
| 2021 | Geometry understanding from autonomous driving scenarios based on feature refinement
Mingliang Zhai, Xuezhi Xiang |
Neural Comput. Appl. | 2 |
| 2021 | Self-supervised Monocular Trained Depth Estimation Using Triplet Attention and Funnel Activation
Xuezhi Xiang, Yujian Qiu, Kaixu Zhang, Ning Lv 0001 |
Neural Process. Lett. | 1 |
| 2021 | Multi-object Tracking Method Based on Efficient Channel Attention and Switchable Atrous Convolution
Xuezhi Xiang, Wenkai Ren, Yujian Qiu, Kaixu Zhang, Ning Lv 0001 |
Neural Process. Lett. | 1 |
| 2021 | Optical flow and scene flow estimation: A survey
Mingliang Zhai, Xuezhi Xiang, Ning Lv 0001 |
Pattern Recognit. | 2 |
| 2020 | Multi-Task Learning in Autonomous Driving Scenarios Via Adaptive Feature Refinement NetworksabstractMany deep learning applications benefit from multi-task learning with several related objectives. In autonomous driving scenarios, being able to accurately infer motion and spatial information is essential for scene understanding. In this paper, we combine an adaptive feature refinement module and a unified framework for joint learning of optical flow, depth and camera pose estimation in an unsupervised manner. The feature refinement module is embedded into motion estimation and depth prediction sub-networks, which can exploit more channel-wise relationships and contextual information for feature learning. Given a monocular video, our network firstly estimates depth and camera motion, and calculates rigid optical flow. Then, we design an auxiliary flow network for inferring non-rigid flow fields. In addition, a forward-backward consistency check is adopted for occlusion reasoning. Extensive experiments on KITTI dataset demonstrate that the proposed method achieves potential results comparing to recent deep learning networks. Mingliang Zhai, Xuezhi Xiang, Ning Lv 0001, Abdulmotaleb El Saddik |
ICASSP | 2 |
| 2020 | Pavement crack detection network based on pyramid structure and attention mechanismabstractAutomatic detection of pavement crack is an important task for conducting road maintenance. However, as an important part of the intelligent transportation system, automatic pavement crack detection is challenging due to the poor continuity of cracks, the different width of cracks, and the low contrast between cracks and the surrounding pavement. This study proposes a novel pavement crack detection method based on an end‐to‐end trainable deep convolution neural network. The authors build the network using the encoder–decoder architecture and adopt a pyramid module to exploit global context information for the complex topology structures of cracks. Moreover, they introduce a spatial‐channel combinational attention module into the encoder–decoder network for refining crack features. Further, the dilated convolution is used to reduce the loss of crack details due to the pooling operation in the encoder network. In addition, they introduce a lovász hinge loss function, which is suitable for small objects. They train the authors' network on the CRACK500 dataset and evaluate it on three pavement crack datasets. Among the methods they compare, their method can achieve the best experimental results. Xuezhi Xiang, Abdulmotaleb El Saddik |
IET Image Process. | 1 |
| 2020 | Video event classification based on two-stage neural network
Lei Zhang 0093, Xuezhi Xiang |
Multim. Tools Appl. | 2 |
| 2020 | A CNNs-based method for optical flow estimation with prior constraints and stacked U-Nets
Xuezhi Xiang, Mingliang Zhai, Rongfang Zhang, Yulong Qiao, Abdulmotaleb El Saddik |
Neural Comput. Appl. | 1 |
| 2020 | Dual-Path Part-Level Method for Visible-Infrared Person Re-identification
Xuezhi Xiang, Ning Lv 0001, Mingliang Zhai, Rokia Abdein, Abdulmotaleb El Saddik |
Neural Process. Lett. | 1 |
| 2020 | Attention-Based Generative Adversarial Network for Semi-supervised Image Classification
Xuezhi Xiang, Zeting Yu, Ning Lv 0001, Abdulmotaleb El Saddik |
Neural Process. Lett. | 1 |
| 2020 | Unsupervised Optical Flow Estimation Based on Improved Feature Pyramid
Nuohan Li, Anchang Liu, Kuan Ye, Xuezhi Xiang |
Neural Process. Lett. | 9 |
| 2020 | Optical Flow Estimation Using Dual Self-Attention Pyramid NetworksabstractRecently, optical flow estimation benefits greatly from deep learning based techniques. Most approaches use encoder-decoder architecture (U-Net) or spatial pyramid network (SPN) to learn optical flow. Both U-Net and SPN can extract multi-scale features and can predict optical flow directly. However, existing networks ignore to exploit the global information among channel features and inter-spatial relationship of features. In this paper, we propose a dual self-attention pyramid network, which adaptively integrates local features with their global dependencies and focuses on important features and suppresses unimportant features. Specifically, we introduce two types of attention modules into SPN, which emphasizes meaningful features along channel and spatial axes. The channel attention can adaptively re-weight channel-wise features by considering interdependencies among channels. Moreover, the spatial attention can utilize global contextual information to emphasize or suppress features in different spatial locations. In addition, two attention modules are embedded into each pyramidal level, which can refine features at different scale. We evaluate our method on MPI-Sintel and KITTI. The experimental results show that using the dual self-attention module can improve the representation power of network and further increase the accuracy of optical flow estimation. Mingliang Zhai, Xuezhi Xiang, Rongfang Zhang, Ning Lv 0001, Abdulmotaleb El Saddik |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | An Object Context Integrated Network for Joint Learning of Depth and Optical FlowabstractSupervised depth prediction and optical flow estimation have achieved promising performance due to the advanced deep network architectures. Since the ground truths are difficult to be collected, many recent works try to learn the depth and flow in an unsupervised manner. However, existing methods only use features from convolutional layers or a simple aggregation of multi-level features to predict the depth and flow maps, which is insufficient to exploit context information. In this paper, we attempt to exploit object contextual information and investigate the effect of the object context for joint learning of depth and optical flow. Specifically, we present a novel combination of object context and the framework of joint learning depth and optical flow. Our proposed network can exploit and integrate the object context for both tasks by aggregating the context according to pair-wise similarities. Furthermore, we adopt the existing spatial pyramid network (SPN) to estimate the depth and flow in a coarse-to-fine strategy effectively. Given temporally adjacent stereo pairs, our network can be trained end-to-end in an unsupervised manner and can predict the depth and flow maps simultaneously. We conduct experiments on two publicly available datasets, KITTI2012 and KITTI2015. Our proposed approach yields comparable performance on both depth and flow tasks, compared to the recent deep learning-based approaches. Experimental results demonstrate that exploiting object contextual information is useful and beneficial for depth and optical flow estimation. Mingliang Zhai, Xuezhi Xiang, Ning Lv 0001, Abdulmotaleb El Saddik |
IEEE Trans. Image Process. | 2 |
| 2019 | Ad-net: Attention Guided Network for Optical Flow Estimation Using Dilated ConvolutionabstractVariational models for optical flow estimation usually define an energy function that contains prior assumptions to explore rudimentary statistics of images. However, such methods cannot learn motion knowledge from the pre-prepared data and have many parameters that need to be set manually. Nowadays, convolutional neural networks (CNNs) have been used in optical flow estimation successfully, which can learn weights from the training dataset and can predict optical flow end-to-end. In this paper, we propose an attention guided network for learning optical flow, named AD-Net, which contains several attention units for modelling the relativities between the channels. Further, we introduce dilated convolution into supervised network for reducing the loss of motion details. In addition, some prior auxiliary constraints are embedded in the supervised network as auxiliary loss terms. Our proposed approach is tested on MPI-Sintel and KITTI2012 datasets and can preserve motion edges and details effectively. Mingliang Zhai, Xuezhi Xiang, Rongfang Zhang, Ning Lv 0001, Abdulmotaleb El Saddik |
ICASSP | 2 |
| 2019 | Optical Flow Estimation Using Spatial-Channel Combinational Attention-Based Pyramid NetworksabstractRecently, learning to estimate optical flow via deep convolutional networks is attracting significant attention. In this paper, we introduce a spatial-channel attention module into optical flow estimation, which infers attention maps along two separated dimensions, channel and spatial, and then integrates these separated attention maps into a fusion attention map for feature refinement. We embed this module into spatial pyramid network, which can adaptively learn the channel and spatial attention maps at each level for modifying the different scaled features and can further improve the accuracy of optical flow estimation. Our network is trained on FlyingChairs and FlyingThings3D datasets with a supervised manner, and is further tested on MPI-Sintel benchmark. The experimental results show that using the spatial-channel attention unit is beneficial for dense flow estimation and our approach is comparable with the state-of-the-art methods. Xuezhi Xiang, Mingliang Zhai, Rongfang Zhang, Ning Lv 0001, Abdulmotaleb El Saddik |
ICIP | 1 |
| 2019 | Optical flow estimation using channel attention mechanism and dilated convolutional neural networks
Mingliang Zhai, Xuezhi Xiang, Rongfang Zhang, Ning Lv 0001, Abdulmotaleb El Saddik |
Neurocomputing | 2 |
| 2017 | Realistic human action recognition: When CNNS meet LDSabstractIn this paper, we proposed new framework for human action representation, which leverages the strengths of convolutional neural networks (CNNs) and the linear dynamical system (LDS) to represent both spatial and temporal structures of actions in videos. We make two principal contributions: first, we incorporate image-trained CNNs to detect action clip concepts, which takes advantage of different levels of information by combining the two layers in CNNs trained from images; Second, we further propose adopting a linear dynamical system (LDS) to model the relationships between these clip concepts, which captures temporal structures of actions. We have applied the proposed method on two challenging realistic benchmark datasets, and our method achieves high performance up to 86.16% on the YouTube and 82.76% UCF50 datasets, which largely outperforms most of the state-of-the-art algorithms with more sophisticated techniques. Lei Zhang 0093, Yangyang Feng, Xuezhi Xiang, Xiantong Zhen |
ICASSP | 3 |
| 2017 | Learning discriminant grassmann kernels for image-set classificationabstractImage-set classification has recently generated great popularity due to widespread application to challenging tasks in computer vision. The great challenges arise from measuring the similarity between image sets which usually exhibit huge inter-class ambiguity and intra-class variation. In this paper, based on the assumption that each image set as a linear subspace can be treated as a point on a Grassmann manifold, we propose discriminant Grassmann kernels (DGK) of principal angles between subspaces. To tackle the ambiguity and variation, we propose learning the DGK via kernel target alignment, which achieves kernels of great discrimination by maximizing correlations with class labels. The proposed DGK has been evaluated on two challenging datasets including the ETH-80 and UCSD datasets for object recognition and video-based traffic congestion recognition, respectively. Extensive experiments have shown that the proposed DGKs achieves state-of-the-art performance and surpasses most of previous methods, which demonstrates the great effectiveness of the DGKs for image-set classification. Lei Zhang 0093, Xuezhi Xiang, Xiantong Zhen |
ICIP | 3 |
| 2016 | Towards optimal VLAD for human action recognition from still images
Lei Zhang 0093, Changxi Li, Peipei Peng, Xuezhi Xiang, Jingkuan Song |
Image Vis. Comput. | 4 |
| 2015 | Dimensionality reduction by supervised locality analysisabstractHigh-dimensional feature representations have recently been widely used for image classification, which not only induce large storage requirement and high computational complexity, but also tend to be lack of discrimination due to redundant and noisy features. In this paper, we propose a novel algorithm named supervised locality analysis (SLA) for dimensionality reduction. In contrast to conventional dimensionality reduction methods, the proposed SLA incorporates supervision into locality analysis by fully exploring multi-class distributions, which can handle the non-linear data structure while preserving intrinsic discriminative information. The obtained compact and highly discriminative features by the SLA is enables more accurate and efficient classification. Moreover, the SLA can be used for supervised dimensionality reduction of both handcrafted and deep learning based features. We have conduced experiments to evaluate the proposed SLA on three datasets for image classification. The SLA has produced state-of-the-art performance and largely outperformed widely-used dimensionality reduction methods. Lei Zhang 0093, Peipei Peng, Xuezhi Xiang, Xiantong Zhen |
ICIP | 3 |
| 2014 | Learning semantic kernels for scene classificationabstractIn this paper we propose to learn semantic kernels for scene classification. We first decompose the Object Bank representation into subspaces associated with each object, Anchor Objects are then created by clustering for each scene class separately. The Anchor Distances are computed to measure the distance between objects to scene classes. In order to take the advantage of the discriminative information from different scene classes, we propose semantic kernels based on the anchor distances to different classes for scene classification. Through extensive experiments on two benchmark datasets: UIUC-Sports dataset and 15-Scene dataset, we prove that the proposed Semantic Kernels can significantly improve the original Object Bank and achieve state-of-the-art performance. Lei Zhang 0093, Xiantong Zhen, Jiqing Han 0001, Xuezhi Xiang |
ICASSP | 4 |