Qiulei Dong

dblp:05/3703 · DBLP profile ↗
← Back
66ranked-venue papers
15as first author
49since 2021 · last 2026
0000-0003-4015-1615ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 8 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 5 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 4 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CAG-GS: Consistent Anchor Guided Gaussian Splatting for Large-scale Scene Rendering
abstract
Recently, 3D Gaussian Splatting for scene rendering has attracted much attention in computer vision and graphics, but generally suffers from large burdens of both computation and storage when handling large-scale scenes. Some existing works in literature employ a divide-and-conquer strategy for alleviating this issue, where an input large scene is divided into lots of local blocks, and each block is handled separately. However, such a strategy generally leads to limited performance due to the inevitable inconsistency among the 3D Gaussians from different blocks. To address this problem, we propose a Consistent Anchor Guided Gaussian Splatting for large-scale scene rendering under the divide-and-conquer strategy, called CAG-GS. In CAG-GS, a set of learnable anchors for each local block is injected with the corresponding semantic features from a pre-trained semantic segmentation model SAM2 through an explored semantic mapping module, and then these anchors are used to predict the attributes of 3D Gaussians. Moreover, we explore a coarse-to-fine training strategy for CAG-GS, where each local block is optimized independently while being guided by globally consistent semantics. Extensive experimental results on five large-scale scenes demonstrate the superiority of the proposed method over five state-of-the-art methods in most cases.
Qiulei Dong
AAAI2
2026 Multi-view consistent feature learning for open-set semantic image segmentation
Haixia Wang 0003, Yuqin Chen, Mengyu Gao, Xiao Lu 0003, Zhiguo Zhang 0005, Qiulei Dong
Expert Syst. Appl.7
2026 Active learning-based structure parsing of ancient Chinese architectures: a benchmark dataset and baseline
Wei Wang 0347, Yixing Wang, Qiulei Dong, Zhanyi Hu
Expert Syst. Appl.5
2026 DeepTA: High-Speed Deep Camera Translation Averaging with Reverse Direction Invariance
Qiulei Dong
Int. J. Comput. Vis.2
2026 SNOP-GS: Self-refining novel object pose via 3D Gaussian splatting
Qiulei Dong
Pattern Recognit.3
2026 Distortion-Aware Depth Self-Updating for Self-Supervised Fisheye Monocular Depth Estimation
abstract
Self-supervised monocular depth estimation for fisheye cameras has attracted much attention in recent years due to their large view range. However, the performances of existing methods in this field are generally limited due to the inevitable severe distortions in fisheye images. To address this problem, we propose a distortion-aware depth self-updating network for self-supervised fisheye monocular depth estimation called DDS-Net. The proposed DDS-Net method employs a coarse-to-fine learning strategy, in which an explored fine depth predictor for predicting final depth is optimized with the predicted scene depths by a pretrained coarse depth predictor. The fine depth predictor contains a distortion-aware fisheye cost volume construction module and a depth self-updating module. The distortion-aware fisheye cost volume construction module is designed to construct a fisheye cost volume by learning the corresponding feature matching cost between continuous fisheye frames, which enables more accurate pixel-level depth cues to be captured under severe distortions. Based on the constructed cost volume and the initial depth estimated by the pretrained coarse depth predictor, the depth self-updating module is designed to self-update the depth map in an iterative manner. Extensive experimental results on 3 fisheye datasets demonstrate that the proposed method significantly outperforms 14 state-of-the-art methods for fisheye monocular depth estimation.
Qiulei Dong
IEEE Trans. Image Process.2
2026 Bimodal Commuting Alignment Network for Zero-Shot Recognition
abstract
Zero-shot recognition has attracted increasing attention for recognizing unseen classes by aligning image visual features and class semantic features. Although various visual and semantic feature types have been explored in literature,e.g., semantic-unaware/aware visual features and attribute/name-based semantic features, most existing methods are tailored to specific bimodal feature pairs, leading to degraded performance when generalized to other feature combinations. To address this issue, we analyze the commonly-used visual and semantic features and reveal that single-modal features generally contain either redundant or deficient information, and bimodal features generally suffer from cross-modal inconsistency. Accordingly, we propose a bimodal commuting alignment network, called BCA-Net, which could accommodate various bimodal feature combinations. Specifically, an entropy-based perception module is first designed to learn bimodal prompts by strengthening discriminative visual-channel features and semantic tokens respectively. Then, a bimodal complementing module is proposed to refine the learned bimodal prompts for complementing inconsistent information across feature modalities. Extensive experimental results on 13 public datasets demonstrate that BCA-Net consistently outperforms 20 state-of-the-art methods by utilizing 4 bimodal feature combinations in most cases, highlighting its strong generalizability beyond designs tailored to specific bimodal pairs.
Mengyu Gao, Qiulei Dong
IEEE Trans. Multim.2
2025 EquiPose: Exploiting Permutation Equivariance for Relative Camera Pose Estimation
abstract
Relative camera pose estimation between two images is a fundamental task in 3D computer vision. Recently, many relative pose estimation networks have been explored for learning a mapping from two input images to their corresponding relative pose, however, the estimated relative poses by these methods do not have the intrinsic Pose Permutation Equivariance (PPE) property: the estimated relative pose from Image A to Image B should be the inverse of that from Image B to Image A. It means that permuting the input order of two images would cause these methods to obtain inconsistent relative poses. To address this problem, we firstly introduce the concept of PPE mapping, which indicates such a mapping that captures the intrinsic PPE property of relative poses. Then by enforcing the aforementioned PPE property, we propose a general framework for relative pose estimation, called EquiPose, which could easily accommodate various relative pose estimation networks in literature as its baseline models. We further theoretically prove that the proposed EquiPose framework could guarantee that its obtained mapping is a PPE mapping. Given a pre-trained baseline model, the proposed EquiPose framework could improve its performance even without fine-tuning, and could further boost its performance with fine-tuning. Experimental results on four public datasets demonstrate that EquiPose could significantly improve the performances of various state-of-the-art models.
Qiulei Dong
CVPR2
2025 ONDA-Pose: Occlusion-Aware Neural Domain Adaptation for Self-Supervised 6D Object Pose Estimation
abstract
Self-supervised 6D object pose estimation has received increasing attention in computer vision recently. Some typical works in literature attempt to translate the synthetic images with object pose labels generated by object CAD models into the real domain, and then use the translated data for training. However, their performance is generally limited, since (i) there still exists a domain gap between the translated images and the real images and (ii) the translated images can not sufficiently reflect occlusions that exist in many real images. To address these problems, we propose an Occlusion-Aware Neural Domain Adaptation method for self-supervised 6D object Pose estimation, called ONDA-Pose. The proposed method comprises three main steps. Firstly, by utilizing both the training real images without pose labels and a CAD model, we explore a CAD-like radiance field for rendering corresponding synthetic images that have similar textures to those generated by the CAD model. Then, a backbone pose estimator trained on the synthetic data is employed to provide initial pose estimations for the synthetic images rendered from the CAD-like radiance field, and the initial object poses are refined by a global object pose refiner to generate pseudo object pose labels. Finally, the backbone pose estimator is further self-supervised as the final pose estimator by jointly utilizing the real images with pseudo object pose labels and the synthetic images rendered from the CAD-like radiance field. Experimental results on three public datasets demonstrate that ONDA-Pose significantly outperforms the comparative state-of-the-art methods in most cases.
Tao Tan 0002, Qiulei Dong
CVPR2
2025 Causality-Guided Prompt Learning for Vision-Language Models via Visual Granulation
Mengyu Gao, Qiulei Dong
ICCV2
2025 Rectified self-supervised monocular depth estimation loss for nighttime and dynamic scenes
Xiaofei Qin, Yongchao Zhu, Fan Mao, Xuedian Zhang, Changxiang He, Qiulei Dong
Eng. Appl. Artif. Intell.7
2025 Self-supervised monocular depth learning from unknown cameras: Leveraging the power of raw data
Xiaofei Qin, Yongchao Zhu, Xuedian Zhang, Changxiang He, Qiulei Dong
Image Vis. Comput.6
2025 Recursive Counterfactual Deconfounding for image recognition
abstract
Image recognition is a classic and common task that has been widely applied in computer vision in the past decade. Most existing methods in literature aim to learn discriminative features for classification from labeled images; however, they generally neglect confounders that infiltrate the learned features, resulting in low discrimination performances on test images. To address this problem, we propose a Recursive Counterfactual Deconfounding model for image recognition in both closed- and open-set scenarios based on counterfactual analysis, called RCD. The proposed model consists of a factual graph and a counterfactual graph, where the relationships among image features , model predictions, and confounders are built and updated recursively to learn more discriminative features. It performs recursively so that subtler counterfactual features can be learned and progressively eliminated. Both the discriminability and generalizability of the proposed model can be improved accordingly. In addition, a negative correlation constraint is designed to alleviate the negative effects of counterfactual features during model training. Extensive experimental results in both the closed-set and open-set image recognition tasks demonstrate the effectiveness of the proposed method. Particularly under the hard testing mode of the three public datasets (Aircraft, CUB and ImageNet), the proposed method obtains a higher AUROC (Area Under the Receiver Operating Characteristic curves) by 0.7%–3.5% and a higher OSCR (Open-Set Classification Rate) by 0.2%–4.0%.
Jiayin Sun, Mengyu Gao, Qiulei Dong
Knowl. Based Syst.4
2025 Semantic-guided compositional scene representation framework
Qiulei Dong, Yangyong Zhang, Xiao Lu 0003, Zhiguo Zhang 0005, Yuqin Chen, Huanzhou Shu, Haixia Wang 0003
Neural Networks2
2025 Self-Assembled Generative Framework for Generalized Zero-Shot Learning
abstract
Generative models have attracted much attention for handling the generalized zero-shot learning (GZSL) task recently. Most of the existing generative GZSL models are trained for visual feature synthesis by utilizing the unique semantic feature of each object class as input but its kaleidoscopic real visual features as supervisions. However, since the real visual features are inevitably infiltrated by some class-irrelevant information, the trained generative models could not guarantee the discriminability of their synthesized visual features. In this paper, we firstly provide an empirical analysis on this problem, finding that among the elements of the real visual features, some elements contain more class-irrelevant information than the others, resulting in ambiguous visual feature synthesis. Then according to this finding, we propose a self-assembled generative GZSL framework, where both the real and synthesized visual features are re-assembled by identifying and updating the class-irrelevant elements in a self-learning manner, called SaG. Moreover, an element-affinity regularizer is explored for constraining the affinity among different elements, so that the synthesized visual features under the SaG framework approach the updated feature elements. In principle, different generative GZSL models could be seamlessly embedded into the SaG framework, resulting in different GZSL methods. Extensive experimental results demonstrate that the derived methods, by embedding three baseline generative GZSL models into SaG respectively, could boost the performances of their baselines significantly, and one of the derived methods outperforms 20 state-of-the-art GZSL methods in most cases.
Mengyu Gao, Qiulei Dong
IEEE Trans. Image Process.2
2025 Noise-Aware Epileptic Seizure Prediction Network via Self-Attention Feature Alignment
abstract
Recently, deep neural networks have been extensively used to extract features from EEG data for epileptic seizure prediction in the epilepsy diagnosis community. Many existing works in literature either use the ultimate-layer feature or aggregate multi-layer features via straightforward concatenation or element-wise addition, but they do not pay a special attention to the contextual consistency between these features as well as the involved noise in these features. To address the above problem, we propose a Noise-aware epileptic seizure prediction network via Self-attention Feature Alignment, called NSFA-Net. The NSFA-Net consists of two modules: a self-attention backbone module to extract multi-layer features from the input EEG data, and a time-frequency feature alignment module to align these features for maintaining the contextual consistency. In addition, during the training process, a noise-aware regularizer is introduced to alleviate the negative influence of noise that is generally inevitable in EEG data. The average sensitivities of the proposed method on the CHB-MIT and Kaggle datasets are 98.68% and 93.57% respectively, and the average false prediction rates are 0.038/h and 0.060/h respectively. These experimental results show the superiority of the proposed method to some state-of-the-art methods.
Qiulei Dong, Zhixi Wang, Mengyu Gao
IEEE J. Biomed. Health Informatics1
2025 Multi-Scale Spatio-Temporal Attention Network for Epileptic Seizure Prediction
abstract
Epilepticseizure prediction from electroencephalogram (EEG) data has attracted much attention in the clinical diagnosis and treatment of epilepsy. Most of the existing methods in literature extract either spatial or temporal features at a single scale from EEG data, however, their learned features are generally less discriminative since the EEG data is complex and severely noisy in general, leading to low-accuracy predictions. To address this problem, we propose a Multi-scale Spatio-temporal Attention Network to learn discriminative features for seizure prediction, called MSAN, which contains a backbone module, a spatial pyramid module, and a multi-scale sequential aggregation module. The backbone module is to extract initial spatial features from the input EEG spectrograms, and the pyramid module is introduced to learn multi-scale features from the initial features. Then by taking these multi-scale features as input temporal features, the sequential aggregation module employs multiple Long Short-Term Memory(LSTM) blocks to aggregate these features. In addition, a dual-loss function is introduced to alleviate the class imbalance problem. The proposed method achieves an average sensitivity of 96.27% with a mean false prediction rate of 0.00/h on the CHB-MIT dataset and an average sensitivity of 93.57% with a mean false prediction rate of 0.044/h on the Kaggle dataset. The comparative results demonstrate that the proposed method outperforms 10 state-of-the-art epileptic seizure prediction models.
Qiulei Dong, Han Zhang 0059, Jun Xiao 0005, Jiayin Sun
IEEE J. Biomed. Health Informatics1
2025 Towards Semi-Supervised Dual-Modal Semantic Segmentation
abstract
With the development of 3D and 2D data acquisition techniques, it has become easy to obtain point clouds and images of scenes simultaneously, which further facilitates dual-modal semantic segmentation. Most existing methods for simultaneously segmenting point clouds and images rely heavily on the quantity and quality of the labeled training data. However, massive point-wise and pixel-wise labeling procedures are time-consuming and labor-intensive. To address this issue, we propose a parallel dual-stream network to handle the semi-supervised dual-modal semantic segmentation task, called PD-Net, by jointly utilizing a small number of labeled point clouds, a large number of unlabeled point clouds, and unlabeled images. The proposed PD-Net consists of two parallel streams (called original stream and pseudo-label prediction stream). The pseudo-label prediction stream predicts the pseudo labels of unlabeled point clouds and their corresponding images. Then, the unlabeled data is sent to the original stream for self-training. Each stream contains two encoder-decoder branches for 3D and 2D data respectively. In each stream, multiple dual-modal fusion modules are explored for fusing the dual-modal features. In addition, a pseudo-label optimization module is explored to optimize the pseudo labels output by the pseudo-label prediction stream. Experimental results on two public datasets demonstrate that the proposed PD-Net not only outperforms the comparative semi-supervised methods but also achieves competitive performances with some fully-supervised methods in most cases.
Qiulei Dong, Shuang Deng
IEEE Trans. Multim.1
2025 A Survey on Self-Supervised Monocular Depth Estimation Based on Deep Neural Networks
abstract
Monocular depth estimation aims to predict the corresponding scene depth map to an input image, which has wide application prospects in various fields, such as robot navigation, autonomous driving, and augmented reality. Due to the advantage that only images rather than ground truth depth maps are required for model training, self-supervised monocular depth estimation methods have received more and more attention in recent years. Although numerous self-supervised monocular depth estimation methods were proposed, there has been no a comprehensive survey on them yet. Addressing this issue, we review recent developments in the community of self-supervised monocular depth estimation in this article. First, 89 existing works in the literature are categorized and reviewed. Then, we introduce the public datasets and evaluation metrics used in monocular depth estimation. Next, the performances of some state-of-the-art methods are compared and analyzed. Finally, we summarize several open problems and possible future developments in this community.
Qiulei Dong, Zhengming Zhou, Xiaolan Qiu
IEEE Trans. Neural Networks Learn. Syst.1
2024 Density-Guided Semi-Supervised 3D Semantic Segmentation with Dual-Space Hardness Sampling
abstract
Densely annotating the large-scale point clouds is laborious. To alleviate the annotation burden, contrastive learning has attracted increasing attention for tackling semi-supervised 3D semantic segmentation. However, existing point-to-point contrastive learning techniques in literature are generally sensitive to outliers, resulting in insufficient modeling of the point-wise representations. To address this problem, we propose a method named DDSemi for semi-supervised 3D semantic segmentation, where a density-guided contrastive learning technique is explored. This technique calculates the contrastive loss in a point-to-anchor manner by estimating an anchor for each class from the memory bank based on the finding that the cluster centers tend to be located in dense regions. In this technique, an inter-contrast loss is derived from the perturbed unlabeled point cloud pairs, while an intra-contrast loss is derived from a single unlabeled point cloud. The derived losses could enhance the discriminability of the features and implicitly constrain the semantic consistency between the perturbed unlabeled point cloud pairs. In addition, we propose a dual-space hardness sampling strategy to pay more attention to the hard samples located in sparse regions of both the geometric space and feature space by reweighting the point-wise intra-contrast loss. Experimental results on both indoor-scene and outdoor-scene datasets demonstrate that the proposed method outperforms the comparative state-of-the-art semi-supervised methods.
Qiulei Dong
CVPR2
2024 LASS3D: Language-Assisted Semi-Supervised 3D Semantic Segmentation with Progressive Unreliable Data Exploitation
Qiulei Dong
ECCV (3)2
2024 Descriptor Distillation: A Teacher-Student-Regularized Framework for Learning Local Descriptors
Qiulei Dong
Int. J. Comput. Vis.2
2024 Conditional feature generation for transductive open-set recognition via dual-space consistent sampling
Jiayin Sun, Qiulei Dong
Pattern Recognit.2
2024 Adaptive Conditional Denoising Diffusion Model With Hybrid Affinity Regularizer for Generalized Zero-Shot Learning
abstract
Generalized zero-shot learning (GZSL) is a challenging topic in both computer vision and machine learning. Recently, generative models (e.g., GAN and VAE) have attracted much attention for handling the GZSL task, however, they are sometimes prone to either model collapse or ambiguous distribution modeling. Inspired by the feature generation ability of denoising diffusion models in other visual tasks, we propose an Adaptive Conditional Denoising Diffusion Model to synthesize unseen-class visual features for GZSL on condition of a set of semantic features in this paper, called AC-DDM. Unlike traditional denoising diffusion models whose reverse process has both a fixed time interval and a fixed number of total denoising time steps, the proposed AC-DDM has a learnable distribution-constrained predictor which could adaptively learn the time interval and the number of total denoising time steps for each unseen class, so that it could synthesize more discriminative features for sample classification. In order to improve the discrimination ability of the synthesized visual features further, we also explore a hybrid affinity regularizer under the proposed AC-DDM, which forces the differences among the affinity matrices of the real and synthesized visual features to be small. Extensive experimental results on four public benchmark datasets demonstrate the superiority of the proposed model over 20 state-of-the-art models in both the ZSL and GZSL tasks.
Mengyu Gao, Qiulei Dong
IEEE Trans. Circuits Syst. Video Technol.2
2024 Hierarchical Attention Network for Open-Set Fine-Grained Image Recognition
abstract
Triggered by the success of transformers in various visual tasks, the spatial self-attention mechanism has recently attracted more and more attention in the computer vision community. However, we empirically found that a typical vision transformer with the spatial self-attention mechanism could not learn accurate attention maps for distinguishing different categories of fine-grained images. To address this problem, motivated by the temporal attention mechanism in brains, we propose a hierarchical attention network for learning fine-grained feature representations, called HAN, where the features learnt by implementing a sequence of spatial self-attention operations corresponding to multiple moments are aggregated progressively. The proposed HAN consists of four modules: a self-attention backbone module for learning a sequence of features with self-attention operations, a spatial feature self-organizing module for facilitating the model training, a hierarchical aggregation module for aggregating the re-organized features via a Long Short-Term Memory network, and a context-aware module that is implemented as the forget block of the hierarchical aggregation module for preserving/forgetting the long-term memory by utilizing contextual information. Then, we propose a HAN-based method for open-set fine-grained recognition by integrating the proposed HAN network with a linear classifier, called HAN-OSFGR. Extensive experimental results on 3 fine-grained datasets and 2 coarse-grained datasets demonstrate that the proposed HAN-OSFGR outperforms 9 state-of-the-art open-set recognition methods significantly in most cases.
Jiayin Sun, Qiulei Dong
IEEE Trans. Circuits Syst. Video Technol.3
2024 Efficient Region-Based 3-D Urban Building Reconstruction From TomoSAR Images
abstract
The tomographic synthetic aperture radar (TomoSAR) technique has been gaining attention because it can retrieve the 3-D structures of urban buildings by synthesizing apertures along the elevation direction. However, most existing TomoSAR methods in literature calculate elevations pixel by pixel and overlook the correlation between elevations, leading to low accuracy and efficiency. To solve these problems, this study introduces an efficient region-based 3-D urban building reconstruction method that incorporates different geometric primitives (i.e., points, planes, and models). Specifically, the proposed method, under the constraints constructed by different geometric primitives, follows three steps to reconstruct the box-like models of urban buildings: 1) it detects double-bounce regions and reconstructs box-like models based on plane sweeping and region growing; 2) it reconstructs box-like models based on multiplane fitting and optimization for nondouble-bounce regions; and 3) it regularizes box-like models based on building layout priors (e.g., collinearity and proximity). The experimental results on two datasets show that the proposed method can efficiently produce reliable results and outperforms several existing methods both qualitatively and quantitatively.
Wei Wang 0347, Liankun Yu, Qiulei Dong, Zhanyi Hu
IEEE Trans. Geosci. Remote. Sens.3
2024 Cayley Rotation Averaging: Multiple Camera Averaging Under the Cayley Framework
abstract
Rotation averaging, which aims to calculate the absolute rotations of a set of cameras from a redundant set of their relative rotations, is an important and challenging topic arising in the study of structure from motion. A central problem in rotation averaging is how to alleviate the influence of noise and outliers. Addressing this problem, we investigate rotation averaging under the Cayley framework in this paper, inspired by the extra-constraint-free nature of the Cayley rotation representation. Firstly, for the relative rotation of an arbitrary pair of cameras regardless of whether it is corrupted by noise/outliers or not, a general Cayley rotation constraint equation is derived for reflecting the relationship between this relative rotation and the absolute rotations of the two cameras, according to the Cayley rotation representation. Then based on such a set of Cayley rotation constraint equations, a Cayley-based approach for Rotation Averaging is proposed, called CRA, where an adaptive regularizer is designed for further alleviating the influence of outliers. Finally, a unified iterative algorithm for minimizing some commonly-used loss functions is proposed under this approach. Experimental results on 16 real-world datasets and multiple synthetic datasets demonstrate that the proposed CRA approach achieves a better accuracy in comparison to several typical rotation averaging approaches in most cases.
Qiulei Dong, Shuang Deng
IEEE Trans. Image Process.1
2023 Open-set Semantic Segmentation for Point Clouds via Adversarial Prototype Framework
abstract
Recently, point cloud semantic segmentation has attracted much attention in computer vision. Most of the existing works in literature assume that the training and testing point clouds have the same object classes, but they are generally invalid in many real-world scenarios for identifying the 3D objects whose classes are not seen in the training set. To address this problem, we propose an Adversarial Prototype Framework (APF) for handling the open-set 3D semantic segmentation task, which aims to identify 3D unseen-class points while maintaining the segmentation performance on seen-class points. The proposed APF consists of a feature extraction module for extracting point features, a prototypical constraint module, and a feature adversarial module. The prototypical constraint module is designed to learn prototypes for each seen class from point features. The feature adversarial module utilizes generative adversarial networks to estimate the distribution of unseenclass features implicitly, and the synthetic unseen-class features are utilized to prompt the model to learn more effective point features and prototypes for discriminating unseen-class samples from the seen-class ones. Experimental results on two public datasets demonstrate that the proposed APF outperforms the comparative methods by a large margin in most cases.
Qiulei Dong
CVPR2
2023 SMOC-Net: Leveraging Camera Pose for Self-Supervised Monocular Object Pose Estimation
abstract
Recently, self-supervised 6D object pose estimation, where synthetic images with object poses (sometimes jointly with un-annotated real images) are used for training, has attracted much attention in computer vision. Some typical works in literature employ a time-consuming differentiable renderer for object pose prediction at the training stage, so that (i) their performances on real images are generally limited due to the gap between their rendered images and real images and (ii) their training process is computationally expensive. To address the two problems, we propose a novel Network for Self-supervised Monocular Object pose estimation by utilizing the predicted Camera poses from unannotated real images, called SMOC-Net. The proposed network is explored under a knowledge distillation framework, consisting of a teacher model and a student model. The teacher model contains a backbone estimation module for initial object pose estimation, and an object pose refiner for refining the initial object poses using a geometric constraint (called relative-pose constraint) derived from relative camera poses. The student model gains knowledge for object pose estimation from the teacher model by imposing the relative-pose constraint. Thanks to the relative-pose constraint, SMOC-Net could not only narrow the domain gap between synthetic and real data but also reduce the training cost. Experimental results on two public datasets demonstrate that SMOC-Net outperforms several state-of-the-art methods by a large margin while requiring much less training time than the differentiable-renderer-based methods.
Tao Tan 0002, Qiulei Dong
CVPR2
2023 Two-in-One Depth: Bridging the Gap Between Monocular and Binocular Self-supervised Depth Estimation
abstract
Monocular and binocular self-supervised depth estimations are two important and related tasks in computer vision, which aim to predict scene depths from single images and stereo image pairs respectively. In literature, the two tasks are usually tackled separately by two different kinds of models, and binocular models generally fail to predict depth from single images, while the prediction accuracy of monocular models is generally inferior to binocular models. In this paper, we propose a Two-in-One self-supervised depth estimation network, called TiO-Depth, which could not only compatibly handle the two tasks, but also improve the prediction accuracy. TiO-Depth employs a Siamese architecture and each sub-network of it could be used as a monocular depth estimation model. For binocular depth estimation, a Monocular Feature Matching module is proposed for incorporating the stereo knowledge between the two images, and the full TiO-Depth is used to predict depths. We also design a multi-stage joint-training strategy for improving the performances of TiO-Depth in both two tasks by combining the relative advantages of them. Experimental results on the KITTI, Cityscapes, and DDAD datasets demonstrate that TiO-Depth outperforms both the monocular and binocular state-of-the-art methods in most cases, and further verify the feasibility of a two-in-one network for monocular and binocular depth estimation. The code is available at https://github.com/ZM-Zhou/TiO-Depth_pytorch.
Zhengming Zhou, Qiulei Dong
ICCV2
2023 ASPPR: active single-image piecewise planar 3D reconstruction based on geometric priors
Wei Wang 0347, Qiulei Dong, Zhanyi Hu
Sci. China Inf. Sci.2
2023 Interactive piecewise planar building reconstruction from a single image based on geometric priors
Wei Wang 0347, Qiulei Dong, Zhanyi Hu
Expert Syst. Appl.2
2023 Distilled Heterogeneous Feature Alignment Network for SAR Image Semantic Segmentation
abstract
SAR (Synthetic Aperture Radar) image semantic segmentation has attracted increasing attention in the remote sensing community recently, due to SAR’s all-time and all-weather imaging capability. However, SAR images are generally more difficult to be segmented than their EO (Electro-Optical) counterparts, since speckle noises and layovers are inevitably involved in SAR images. On the other hand, EO images could only be obtained under cloud-free conditions, which limits their applications. To this end, this letter investigates how to introduce EO features to assist the training of a SAR-segmentation model so that the model could segment SAR images without their EO counterparts in application, and proposes a distilled heterogeneous feature alignment network (DHFA-Net), where a SAR-segmentation student model learns and aligns the features from a pre-trained EO-segmentation teacher model. In the proposed DHFA-Net, both the student and teacher models employ an identical architecture but different parameter configurations, and a heterogeneous feature distillation module is explored for transferring latent EO features from the teacher model to the student model through heterogeneous feature distillation and then supervising the training of the SAR-segmentation model. Moreover, a heterogeneous feature alignment module is designed to aggregate multi-scale features for segmentation by feature alignment approach in each of the student and teacher models. By enabling the multi-scale heterogeneous feature aggregation, the SAR segmentation performance could be boosted. Experimental results on two public datasets demonstrate the superiority of the proposed DHFA-Net.
Mengyu Gao, Jiping Xu, Qiulei Dong
IEEE Geosci. Remote. Sens. Lett.4
2023 MoEP-AE: Autoencoding Mixtures of Exponential Power Distributions for Open-Set Recognition
abstract
Open-set recognition aims to identify unknown classes while maintaining classification performance on known classes and has attracted increasing attention in the pattern recognition field. However, how to learn effective feature representations whose distributions are usually complex for classifying both known-class and unknown-class samples when only the known-class samples are available for training is an ongoing issue in open-set recognition. In contrast to methods implementing a single Gaussian, a mixture of Gaussians (MoG), or multiple MoGs, we propose a novel autoencoder that learns feature representations by modeling them as mixtures of exponential power distributions (MoEPs) in latent spaces called MoEP-AE. The proposed autoencoder considers that many real-world distributions are sub-Gaussian or super-Gaussian and can thus be represented by MoEPs rather than a single Gaussian or an MoG or multiple MoGs. We design a differentiable sampler that can sample from an MoEP to guarantee that the proposed autoencoder is trained effectively. Furthermore, we propose an MoEP-AE-based method for open-set recognition by introducing a discrimination strategy, where the MoEP-AE is used to model the distributions of the features extracted from the input known-class samples by minimizing a designed loss function at the training stage, called MoEP-AE-OSR. Extensive experimental results in both standard-dataset and cross-dataset settings demonstrate that the MoEP-AE-OSR method outperforms 14 existing open-set recognition methods in most cases in both open-set recognition and closed-set recognition tasks.
Jiayin Sun, Qiulei Dong
IEEE Trans. Circuits Syst. Video Technol.3
2022 Self-distilled Feature Aggregation for Self-supervised Monocular Depth Estimation
Zhengming Zhou, Qiulei Dong
ECCV (1)2
2022 Superpoint-guided Semi-supervised Semantic Segmentation of 3D Point Clouds
abstract
3D point cloud semantic segmentation is a challenging topic in the computer vision field. Most of the existing methods in literature require a large amount of fully labeled training data, but it is extremely time-consuming to obtain these training data by manually labeling massive point clouds. Addressing this problem, we propose a superpoint-guided semi-supervised segmentation network for 3D point clouds, which jointly utilizes a small portion of labeled scene point clouds and a large number of unlabeled point clouds for network training. The proposed network is iteratively updated with its predicted pseudo labels, where a superpoint generation module is introduced for extracting superpoints from 3D point clouds, and a pseudo-label optimization module is explored for automatically assigning pseudo labels to the unlabeled points under the constraint of the extracted superpoints. Additionally, there are some 3D points without pseudo-label supervision. We propose an edge prediction module to constrain features of edge points. A superpoint feature aggregation module and a superpoint feature consistency loss function are introduced to smooth superpoint features. Extensive experimental results on two 3D public datasets demonstrate that our method can achieve better performance than several state-of-the-art point cloud segmentation networks and several popular semi-supervised segmentation methods with few labeled scenes.
Shuang Deng, Qiulei Dong, Bo Liu 0035, Zhanyi Hu
ICRA2
2022 Learning Occlusion-aware Coarse-to-Fine Depth Map for Self-supervised Monocular Depth Estimation
abstract
Self-supervised monocular depth estimation, aiming to learn scene depths from single images in a self-supervised manner, has received much attention recently. In spite of recent efforts in this field, how to learn accurate scene depths and alleviate the negative influence of occlusions for self-supervised depth estimation, still remains an open problem. Addressing this problem, we firstly empirically analyze the effects of both the continuous and discrete depth constraints which are widely used in the training process of many existing works. Then inspired by the above empirical analysis, we propose a novel network to learn an Occlusion-aware Coarse-to-Fine Depth map for self-supervised monocular depth estimation, called OCFD-Net. Given an arbitrary training set of stereo image pairs, the proposed OCFD-Net does not only employ a discrete depth constraint for learning a coarse-level depth map, but also employ a continuous depth constraint for learning a scene depth residual, resulting in a fine-level depth map. In addition, an occlusion-aware module is designed under the proposed OCFD-Net, which is able to improve the capability of the learnt fine-level depth map for handling occlusions. Experimental results on KITTI demonstrate that the proposed method outperforms the comparative state-of-the-art methods under seven commonly used metrics in most cases. In addition, experimental results on Make3D demonstrate the effectiveness of the proposed method in terms of the cross-dataset generalization ability under four commonly used metrics. The code is available at https://github.com/ZM-Zhou/OCFD-Net_pytorch.
Zhengming Zhou, Qiulei Dong
ACM Multimedia2
2022 MCFINet: Multidepth Convolution Network With Shallow-Deep Feature Integration for Semantic Labeling in Remote Sensing Images
abstract
Semantic labeling in remote sensing images is an important and challenging technique, which has attracted increasing attention recently in earth detection, environmental protection, land utilization, and so on. However, it remains a challenge on how to effectively label objects with varied scales and similar textures in literature. Addressing this challenge, we propose a multidepth convolution network with shallow-deep feature integration, called MCFINet, which could effectively integrate multiscale contexts and shallow-layer/deep-layer features for labeling various objects. In the proposed network, we design two new modules—a multidepth convolutional module (MDCM) and an adaptive feature integration module (AFIM). The MDCM employs multilayer convolutions with varied layer numbers but fixed small-sized kernels in parallel to capture multiscale contexts, while the AFIM adaptively integrates the shallow-layer and deep-layer features of the proposed network to capture more discriminant features for segmenting objects with similar textures. Extensive experimental results on two benchmark data sets demonstrate that MCFINet could achieve better performances than seven existing methods in most cases.
Dongji Wang, Qiulei Dong
IEEE Geosci. Remote. Sens. Lett.2
2022 Gated Feature Aggregation for Height Estimation From Single Aerial Images
abstract
Height estimation from single images, strictly speaking, is an ill-posed problem. However, recently, it is shown that it is both possible and feasible to learn a mapping from image statistics to height information. In spite of recent efforts in this field, how to learn fine-shape preserving features, such as object boundaries and contours, is still an open issue. In this work, we propose a progressive learning network to estimate height information from single aerial images in a coarse-to-fine manner. In particular, a gated feature aggregation module is introduced to effectively combine low-level and high-level features. The proposed method is validated on three public datasets, including the Vaihingen dataset, the Potsdam dataset, and the DFC2019 dataset. Both quantitative and qualitative experimental results demonstrate that the proposed method can achieve more accurate height estimation from single aerial images, especially with better object boundary and contour preserving capability, than four related height estimation methods.
Qiulei Dong, Zhanyi Hu
IEEE Geosci. Remote. Sens. Lett.2
2022 SG-SRNs: Superpixel-Guided Scene Representation Networks
abstract
Recently, Scene Representation Networks (SRNs) have attracted increasing attention in computer vision, due to their continuous and light-weight scene representation ability. However, SRNs generally perform poorly on low-texture image regions. Addressing this problem, we propose superpixel-guided scene representation networks in this paper, called SG-SRNs, consisting of a backbone module (SRNs), a superpixel segmentation module, and a superpixel regularization module. In the proposed method, except for the novel view synthesis task, the task of representation-aware superpixel segmentation mask generation is realized by the proposed superpixel segmentation module. Then, the superpixel regularization module utilizes the superpixel segmentation mask to guide the backbone to be learned in a locally smooth way, and optimizes the scene representations of the local regions to indirectly alleviate the structure distortion of low-texture regions in a self-supervised manner. Extensive experimental results on both our constructed datasets and the public Synthetic-NeRF dataset demonstrated that the proposed SG-SRNs achieved a significantly better 3D structure representing performance.
Xiao Lu 0003, Qiulei Dong, Yangyong Zhang, Haixia Wang 0003
IEEE Signal Process. Lett.3
2022 Robust Camera Translation Estimation via Rank Enforcement
abstract
Camera translation averaging, aiming to recover the global camera locations from a given set of camera translation directions, is a challenging problem for Structure from Motion (SfM) in the field of computer vision, largely due to the fact that the given relative translation directions from a set of noisy essential matrices are generally of low accuracy. To tackle this problem, we first reveal a novel but a simple property of the camera translation matrix consisting of all the pairwise camera translations among an arbitrary set of cameras that the rank of this translation matrix is always smaller or equal to 4. Then, by explicitly enforcing this rank property, a novel translation estimation method for computing global camera locations is proposed, called TERE. Moreover, to further improve the performances of the explored TERE in the two aspects of accuracy and speed, an iterative batch-based translation estimation method is proposed, called B-TERE, where a small-scale batch of cameras is selected without replacement from the given set of cameras according to a simple camera selection strategy at each iterative step, and the locations of the selected cameras are estimated by the proposed TERE accordingly. Extensive experimental results on various datasets demonstrate that our proposed methods could achieve better performances in comparison to several state-of-the-art methods.
Qiulei Dong, Xiang Gao 0009, Hainan Cui, Zhanyi Hu
IEEE Trans. Cybern.1
2022 Pursuing 3-D Scene Structures With Optical Satellite Images From Affine Reconstruction to Euclidean Reconstruction
abstract
How to use multiple optical satellite images to recover the 3D scene structure is a challenging and important problem in the remote sensing field. Most existing methods in literature have been explored based on the classical RPC (Rational Polynomial Coefficients) camera model which requires at least 39 GCPs (ground control points), however, it is non-trivial to obtain such a large number of GCPs in many real scenes. Addressing this problem, we propose a hierarchical reconstruction framework based on multiple optical satellite images, which needs only 4 GCPs to fully-automated reconstruct the 3D scene structure. The proposed framework is independent from the RPC model and composed of a dense affine reconstruction stage and a followed affine-to-Euclidean upgrading stage: At the dense affine reconstruction stage, a dense affine reconstruction approach is explored for pursuing the 3D affine scene structure without any GCP from input satellite images. Then at the affine-to-Euclidean upgrading stage, the obtained 3D affine structure is upgraded to a Euclidean one with 4 GCPs. Experimental results on two public datasets demonstrate that the proposed method significantly outperforms several state-of-the-art methods in most cases.
Pinhe Wang, Limin Shi, Bao Chen, Zhanyi Hu, Jianzhong Qiao, Qiulei Dong
IEEE Trans. Geosci. Remote. Sens.6
2022 SAR-to-Optical Image Translation With Hierarchical Latent Features
abstract
Due to the all-weather and all-time imaging capability of Synthetic Aperture Radar (SAR), SAR remote sensing analysis has attracted much attention recently. However, compared with optical images, SAR images are more difficult to be interpreted. If a SAR image could be translated into its corresponding optical image, then the generated optical image would be helpful for assisting the interpretation. Addressing this issue, we investigate how to translate SAR images to optical ones in this work, and propose a parallel generative adversarial model for SAR-to-optical image translation, called Parallel-GAN, consisting of a backbone image translation sub-network and an adjoint optical image reconstruction sub-network. Under the proposed model, the backbone image translation sub-network is designed to translate SAR images to optical ones, and simultaneously some of its intermediate layers are required to output similar latent features to those from the corresponding layers of the adjoint image reconstruction sub-network. Thanks to the imposed hierarchical latent optical features, the proposed Parallel-GAN could achieve the SAR-to-optical image translation effectively. Extensive experimental results on three public datasets demonstrate that the proposed method outperforms ten state-of-the-art methods for SAR-to-optical image translation.
Haixia Wang 0002, Zhanyi Hu, Qiulei Dong
IEEE Trans. Geosci. Remote. Sens.4
2021 Hardness Sampling for Self-Training Based Transductive Zero-Shot Learning
abstract
Transductive zero-shot learning (T-ZSL) which could alleviate the domain shift problem in existing ZSL works, has received much attention recently. However, an open problem in T-ZSL: how to effectively make use of unseen-class samples for training, still remains. Addressing this problem, we first empirically analyze the roles of unseen-class samples with different degrees of hardness in the training process based on the uneven prediction phenomenon found in many ZSL methods, resulting in three observations. Then, we propose two hardness sampling approaches for selecting a subset of diverse and hard samples from a given unseen-class dataset according to these observations. The first one identifies the samples based on the class-level frequency of the model predictions while the second enhances the former by normalizing the class frequency via an approximate class prior estimated by an explored prior estimation algorithm. Finally, we design a new Self-Training framework with Hardness Sampling for T-ZSL, called STHS, where an arbitrary inductive ZSL method could be seamlessly embedded and it is iteratively trained with unseen-class samples selected by the hardness sampling approach. We introduce two typical ZSL methods into the STHS framework and extensive experiments demonstrate that the derived T-ZSL methods outperform many state-of-the-art methods on three public benchmarks. Besides, we note that the unseen-class dataset is separately used for training in some existing transductive generalized ZSL (T-GZSL) methods, which is not strict for a GZSL task. Hence, we suggest a more strict T-GZSL data setting and establish a competitive baseline on this setting by introducing the proposed STHS framework to T-GZSL.
Bo Liu 0035, Qiulei Dong, Zhanyi Hu
CVPR2
2021 SCF-Net: Learning Spatial Contextual Features for Large-Scale Point Cloud Segmentation
abstract
How to learn effective features from large-scale point clouds for semantic segmentation has attracted increasing attention in recent years. Addressing this problem, we propose a learnable module that learns Spatial Contextual Features from large-scale point clouds, called SCF in this paper. The proposed module mainly consists of three blocks, including the local polar representation block, the dual-distance attentive pooling block, and the global contextual feature block. For each 3D point, the local polar representation block is firstly explored to construct a spatial representation that is invariant to the z-axis rotation, then the dual-distance attentive pooling block is designed to utilize the representations of its neighbors for learning more discriminative local features according to both the geometric and feature distances among them, and finally, the global contextual feature block is designed to learn a global context for each 3D point by utilizing its spatial location and the volume ratio of the neighborhood to the global point cloud. The proposed module could be easily embedded into various network architectures for point cloud segmentation, naturally resulting in a new 3D semantic segmentation network with an encoder-decoder architecture, called SCF-Net in this work. Extensive experimental results on two public datasets demonstrate that the proposed SCF-Net performs better than several state-of-the-art methods in most cases.
Siqi Fan 0002, Qiulei Dong, Fenghua Zhu, Peijun Ye 0001, Fei-Yue Wang 0001
CVPR2
2021 Rotation Transformation Network: Learning View-Invariant Point Cloud For Classification And Segmentation
abstract
Many recent works show that a spatial manipulation module could boost the performances of deep neural networks (DNNs) for 3D point cloud analysis. In this paper, we aim to provide an insight into spatial manipulation modules. Firstly, we find that the smaller the rotational degree of freedom (RDF) of objects is, the more easily these objects are handled by these DNNs. Then, we investigate the effect of the popular T-Net module and find that it could not reduce the RDF of objects. Motivated by the above two issues, we propose a rotation transformation network for point cloud analysis, called RTN, which could reduce the RDF of input 3D objects to 0. The RTN could be seamlessly inserted into many existing DNNs for point cloud analysis. Extensive experimental results on 3D point cloud classification and segmentation tasks demonstrate that the proposed RTN could improve the performances of several state-of-the-art methods significantly.
Shuang Deng, Bo Liu 0035, Qiulei Dong, Zhanyi Hu
ICME3
2021 Semantic-diversity transfer network for generalized zero-shot learning via inner disagreement based OOD detector
Bo Liu 0035, Qiulei Dong, Zhanyi Hu
Knowl. Based Syst.2
2021 GA-NET: Global Attention Network for Point Cloud Semantic Segmentation
abstract
How to learn long-range dependencies from 3D point clouds is a challenging problem in 3D point cloud analysis. Addressing this problem, we propose a global attention network for point cloud semantic segmentation, named as GA-Net, consisting of a point-independent global attention module and a point-dependent global attention module for obtaining contextual information of 3D point clouds in this paper. The point-independent global attention module simply shares a global attention map for all 3D points. In the point-dependent global attention module, for each point, a novel random cross attention block using only two randomly sampled subsets is exploited to learn the contextual information of all the points. Additionally, we design a novel point-adaptive aggregation block to replace linear skip connection for aggregating more discriminate features. Extensive experimental results on three 3D public datasets demonstrate that our method outperforms state-of-the-art methods in most cases.
Shuang Deng, Qiulei Dong
IEEE Signal Process. Lett.2
2021 An Iterative Co-Training Transductive Framework for Zero Shot Learning
abstract
In zero-shot learning (ZSL) community, it is generally recognized that transductive learning performs better than inductive one as the unseen-class samples are also used in its training stage. How to generate pseudo labels for unseen-class samples and how to use such usually noisy pseudo labels are two critical issues in transductive learning. In this work, we introduce an iterative co-training framework which contains two different base ZSL models and an exchanging module. At each iteration, the two different ZSL models are co-trained to separately predict pseudo labels for the unseen-class samples, and the exchanging module exchanges the predicted pseudo labels, then the exchanged pseudo-labeled samples are added into the training sets for the next iteration. By such, our framework can gradually boost the ZSL performance by fully exploiting the potential complementarity of the two models' classification capabilities. In addition, our co-training framework is also applied to the generalized ZSL (GZSL), in which a semantic-guided OOD detector is proposed to pick out the most likely unseen-class samples before class-level classification to alleviate the bias problem in GZSL. Extensive experiments on three benchmarks show that our proposed methods could significantly outperform about 31 state-of-the-art ones.
Bo Liu 0035, Lihua Hu, Qiulei Dong, Zhanyi Hu
IEEE Trans. Image Process.3
2020 Zero-Shot Learning from Adversarial Feature Residual to Compact Visual Feature
abstract
Recently, many zero-shot learning (ZSL) methods focused on learning discriminative object features in an embedding feature space, however, the distributions of the unseen-class features learned by these methods are prone to be partly overlapped, resulting in inaccurate object recognition. Addressing this problem, we propose a novel adversarial network to synthesize compact semantic visual features for ZSL, consisting of a residual generator, a prototype predictor, and a discriminator. The residual generator is to generate the visual feature residual, which is integrated with a visual prototype predicted via the prototype predictor for synthesizing the visual feature. The discriminator is to distinguish the synthetic visual features from the real ones extracted from an existing categorization CNN. Since the generated residuals are generally numerically much smaller than the distances among all the prototypes, the distributions of the unseen-class features synthesized by the proposed network are less overlapped. In addition, considering that the visual features from categorization CNNs are generally inconsistent with their semantic features, a simple feature selection strategy is introduced for extracting more compact semantic visual features. Extensive experimental results on six benchmark datasets demonstrate that our method could achieve a significantly better performance than existing state-of-the-art methods by ∼1.2-13.2% in most cases.
Bo Liu 0035, Qiulei Dong, Zhanyi Hu
AAAI2
2020 Face-sketch learning with human sketch-drawing order enforcement
Liang Chang 0001, Lihua Jin, Lifen Weng, Wentao Chao, Xiaoming Deng 0001, Qiulei Dong
Sci. China Inf. Sci.7
2019 Latent-Smoothness Nonrigid Structure From Motion by Revisiting Multilinear Factorization
abstract
How to implement an effective factorization for nonrigid structure from motion (NRSFM) has attracted much attention in recent years. A straightforward factorization scheme is to multilinearly solve NRSFM in an alternating manner, where each of the unknown variables in NRSFM is updated by fixing the others at each iteration. However, recent works show that most existing multilinear factorization (MLF) methods achieve poorer performances than some state-of-the-art sequential factorization methods. In this paper, we reinvestigate the MLF scheme for improving factorization accuracy, and first propose an MLF method with the only low-rank prior for NRSFM in the presence of missing data. Then, for further improving the performances of such MLF methods, a latent "smoothness" characteristic on unknown 3-D deformable shapes is investigated, which is independent of temporal relations among deformable shapes. Accordingly, a latent-smoothness prior for solving NRSFM is derived from the latent smoothness characteristic, and it is able to effectively recover 3-D deformable shapes from unordered data, which is hard for the traditional temporal-smoothness prior to handle. Finally, a regularized factorization method is proposed by integrating MLF with the explored latent-smoothness prior for further pursuing better performances. Extensive experimental results show the effectiveness of our methods in comparison to eight existing multilinear/sequential methods.
Qiulei Dong
IEEE Trans. Cybern.1
2018 Learning stratified 3D reconstruction
Qiulei Dong, Mao Shu, Hainan Cui, Huarong Xu, Zhanyi Hu
Sci. China Inf. Sci.1
2018 Statistics of Visual Responses to Image Object Stimuli from Primate AIT Neurons to DNN Neurons
abstract
Under the goal-driven paradigm, Yamins et al. ( 2014 ; Yamins & DiCarlo, 2016 ) have shown that by optimizing only the final eight-way categorization performance of a four-layer hierarchical network, not only can its top output layer quantitatively predict IT neuron responses but its penultimate layer can also automatically predict V4 neuron responses. Currently, deep neural networks (DNNs) in the field of computer vision have reached image object categorization performance comparable to that of human beings on ImageNet, a data set that contains 1.3 million training images of 1000 categories. We explore whether the DNN neurons (units in DNNs) possess image object representational statistics similar to monkey IT neurons, particularly when the network becomes deeper and the number of image categories becomes larger, using VGG19, a typical and widely used deep network of 19 layers in the computer vision field. Following Lehky, Kiani, Esteky, and Tanaka ( 2011 , 2014 ), where the response statistics of 674 IT neurons to 806 image stimuli are analyzed using three measures (kurtosis, Pareto tail index, and intrinsic dimensionality), we investigate the three issues in this letter using the same three measures: (1) the similarities and differences of the neural response statistics between VGG19 and primate IT cortex, (2) the variation trends of the response statistics of VGG19 neurons at different layers from low to high, and (3) the variation trends of the response statistics of VGG19 neurons when the numbers of stimuli and neurons increase. We find that the response statistics on both single-neuron selectivity and population sparseness of VGG19 neurons are fundamentally different from those of IT neurons in most cases; by increasing the number of neurons in different layers and the number of stimuli, the response statistics of neurons at different layers from low to high do not substantially change; and the estimated intrinsic dimensionality values at the low convolutional layers of VGG19 are considerably larger than the value of approximately 100 reported for IT neurons in Lehky et al. ( 2014 ), whereas those at the high fully connected layers are close to or lower than 100. To the best of our knowledge, this work is the first attempt to analyze the response statistics of DNN neurons with respect to primate IT neurons in image object representation.
Qiulei Dong, Zhanyi Hu
Neural Comput.1
2018 Deep Disentangling Siamese Network for Frontal Face Synthesis Under Neutral Illumination
abstract
Recently, it has been observed that face-recognition performance can be noticeably enhanced through frontal face synthesis. In this letter, a deep disentangling Siamese network (DSN) is proposed for frontal face synthesis. More specifically, we cast frontal face synthesis as the encoder-disentangling-decoder process of profile faces, and a salient feature of our DSN is that a pair of images with arbitrary identities, poses, and illuminations is used as the input rather than a single image used in many existing works. The process consists of the following three stages:1)First, the representations of a pair of face images are learned by the Siamese encoder module.2)Then, the interpretable representations are disentangled into the identity, pose, and illumination representations separately.3)Finally, the identity representation is transferred into the frontal face image using the decoder module.The proposed network is evaluated on the face recognition and synthesis problem. Quantitative and qualitative evaluations on the benchmarks demonstrate that the proposed face synthesis network performs better than the state-of-the-art methods.
Ting Zhang 0006, Qiulei Dong
IEEE Signal Process. Lett.3
2017 Two-Stream Deep Correlation Network for Frontal Face Recovery
abstract
Pose and textural variations are two dominant factors to affect the performance of face recognition. It is widely believed that generating the corresponding frontal face from a face image of an arbitrary pose is an effective step toward improving the recognition performance. In the literature, however, the frontal face is generally recovered by only exploring textural characteristic. In this letter, we propose a two-stream deep correlation network, which incorporates both geometric and textural features for frontal face recovery. Given a face image under an arbitrary pose as input, geometric and textural characteristics are first extracted from two separate streams. The extracted characteristics are then fused through the proposed multiplicative patch correlation layer. These two steps are integrated into one network for end-to-end training and prediction, which is demonstrated effective compared with state-of-the-art methods on the benchmark datasets.
Ting Zhang 0006, Qiulei Dong, Ming Tang 0001, Zhanyi Hu
IEEE Signal Process. Lett.2
2016 Pursuing face identity from view-specific representation to view-invariant representation
abstract
How to learn view-invariant facial representations is an important task for view-invariant face recognition. The recent work [1] discovered that the brain of the macaque monkey has a face-processing network, where some neurons are view-specific. Motivated by this discovery, this paper proposes a deep convolutional learning model for face recognition, which explicitly enforces this view-specific mechanism for learning view-invariant facial representations. The proposed model consists of two concatenated modules: the first one is a convolutional neural network (CNN) for learning the corresponding viewing pose to the input face image; the second one consists of multiple CNNs, each of which learns the corresponding frontal image of an image under a specific viewing pose. This method is of low computational cost, and it can be well trained with a relatively small number of samples. The experimental results on the MultiPIE dataset demonstrate the effectiveness of our proposed convolutional model in contrast to three state-of-the-art works.
Ting Zhang 0006, Qiulei Dong, Zhanyi Hu
ICIP2
2016 Sequential factorization for nonrigid structure from motion via LBFGS
abstract
How to implement an effective factorization for nonrigid structure from motion(NRSFM) has attracted much attention in recent years. Addressing this problem, we propose a novel sequential factorization method without extra priors other than the basis low-rank prior, consisting of a motion estimation module and a 3D shape recovery module. In the motion estimation module, for improving the estimation accuracy, a novel objective function is designed for jointly pursuing the Euclidean corrective matrix and the shape coefficient matrix. And an iterative minimization algorithm is explored to solve the designed objective function based on the Limited-memory Broyden-Fletcher-Goldfarb-Shanno approach(LBFGS), naturally leading to the rotation matrix. In the 3D shape recovery module, a simple iterative algorithm is introduced for effectively computing the 3D deformable shapes with the estimated rotation matrix. The proposed extra-prior-free method is easy to implement and it can achieve an effective tradeoff between estimation accuracy and computational speed, since only several classic techniques are involved. Extensive experimental results demonstrate the effectiveness of the proposed method in comparison to five state-of-the-art methods.
Qiulei Dong
ICPR1
2015 Biologically inspired deep stereo model
abstract
The human stereo vision process begins in primary visual cortex where complex cells are deemed as disparity detectors. Besides correct matches, complex cells could also respond to false matches which cause ambiguous depth perception. This reveals that there exist some inherent mechanisms in the higher visual processing areas for eliminating false matches and thus recovering true disparity. Due to the least understanding of this hierarchical process, this paper aims to investigate this problem by proposing a deep stereo vision model. There are three characteristics of the proposed model: 1) the obtained disparity maps become more accurate from lower layers to higher layers; 2) it tends to promote neural responses of correct matches while inhibiting those of incorrect matches; 3) it generalizes well to unseen data.
Qingqun Kong, Yi Zeng 0001, Qiulei Dong
ICIP3
2014 Euclidean upgrading from segment lengths: DLT-like algorithm and its variants
Kunfeng Shi, Qiulei Dong, Fuchao Wu
Image Vis. Comput.2
2013 Two-dimensional relaxed representation
Qiulei Dong
Neurocomputing1
2013 Smooth incomplete matrix factorization and its applications in image/video denoising
Qiulei Dong
Neurocomputing1
2012 Automatic real-time SLAM relocalization based on a hierarchical bipartite graph model
Qiulei Dong, Zhaopeng Gu, Zhanyi Hu
Sci. China Inf. Sci.1
2012 Weighted Similarity-Invariant Linear Algorithm for Camera Calibration With Rotating 1-D Objects
abstract
In this paper, a weighted similarity-invariant linear algorithm for camera calibration with rotating 1D objects is proposed. First, we propose a new estimation method for computing the relative depth of the free endpoint on the 1D object and prove its robustness against noise compared with those used in previous literature. The introduced estimator is invariant to image similarity transforms, resulting in a similarity-invariant linear calibration algorithm which is slightly more accurate than the well-known normalized linear algorithm. Then, we use the reciprocals of the standard deviations of the estimated relative depths from different images as the weights on the constraint equations of the similarity-invariant linear calibration algorithm, and propose a weighted similarity-invariant linear calibration algorithm with higher accuracy. Experimental results on synthetic data as well as on real image data show the effectiveness of our proposed algorithm.
Kunfeng Shi, Qiulei Dong, Fuchao Wu
IEEE Trans. Image Process.2
2009 Pointwise Motion Image (PMI): A Novel Motion Representation and Its Applications to Abnormality Detection and Behavior Recognition
abstract
In this paper, we propose a novel motion representation and apply it to abnormality detection and behavior recognition. At first, pointwise correspondences for the foreground in two consecutive video frames are established by performing a salient-region-based pointwise matching algorithm. Then, based on the established pointwise correspondences, a pointwise motion image (PMI) for each frame is built up to represent the motion status of the foreground. The PMI is more suitable for video analysis as it encapsulates a variety of motion information such as pointwise motion speed, pointwise motion orientation, pointwise motion duration, as well as the global shape of the foreground. In addition, it represents all of these pieces of information by a color image in the HSV space, by which many popular techniques in the image processing field can be straightforwardly adopted. By combining the PMI and AdaBoost, a method for abnormality detection and behavior recognition is proposed. The proposed method is shown to possess a high discriminative ability and is capable of dealing with local motion, global motion, and similar motions with different speeds. Experiments including a comparison with two existing methods demonstrate the effectiveness of the proposed representation in abnormality detection and behavior recognition.
Qiulei Dong, Yihong Wu 0002, Zhanyi Hu
IEEE Trans. Circuits Syst. Video Technol.1
2006 Gesture Recognition Using Quadratic Curves
Qiulei Dong, Yihong Wu 0002, Zhanyi Hu
ACCV (1)1