VLDB 2026 Research / reviewers in the wild / expert
Chenqiang Gao
dblp:82/9201
· DBLP profile ↗
85ranked-venue papers
8as first author
47since 2021 · last 2026
0000-0003-4174-4148ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 46 · 3 first-author · 21 since 2021Artificial intelligence and machine learning · 33 · 4 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 1 first-author · 17 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ground-to-Aerial Scene Adaptation: Unsupervised drone video action recognition via domain adaptation
Feng Yang 0015, Zhijia Li, Fulin Luo, Anyong Qin, Tiecheng Song, Yue Zhao 0012, Chenqiang Gao |
Eng. Appl. Artif. Intell. | 8 |
| 2026 | Progressive spectral-frequency-spatial guidance network for hyperspectral image dehazing
Qianru Liu, Tiecheng Song, Kaizhao Zhang, Anyong Qin, Feng Yang 0015, Chenqiang Gao |
Expert Syst. Appl. | 7 |
| 2026 | Dual-Path Hierarchical Attention Fusion Module for Cooperative Object Detection Under Heterogeneous Coupled DisturbancesabstractCollaborative 3D object detection leverages information sharing among connected autonomous vehicles (CAVs) to overcome occlusions and limited sensing range. In real-world deployments confront multiple heterogeneously coupled communication disturbances, such as asynchronicity, localization errors, and transmission noise. Existing fusion methods typically model and mitigate each disturbance in isolation, assuming ideal conditions and overlooking their joint effects. As a result, these interactions cause feature misalignment and degraded detection performance. To address this challenge, we propose the Dual-path Hierarchical Attention Fusion Module (DHAFM), which features a cooperative path to enrich CAV feature representations, a stable path to emphasize reliable ego-vehicle features, and a NoiseGate Fusion Module that dynamically reweights these paths based on estimated disturbance levels, prioritizing stability under severe noise and collaboration when communication is reliable. Extensive evaluations on OPV2V, V2XSet, and V2V4REAL benchmarks demonstrate that DHAFM achieves state-of-the-art 3D detection accuracy and markedly improved robustness under both individual and coupled disturbance scenarios. Ziao Li, Chenqiang Gao, Junyin Zhang, Jingyi Yuan |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2026 | A Fusion-Enhanced Network for Infrared and Visible High-Level Vision TasksabstractInfrared and visible dual-modality vision tasks such as semantic segmentation, object detection, and salient object detection can achieve robust performance even in extreme scenes by leveraging complementary information. However, most existing image fusion-based methods and task-specific frameworks exhibit limited generalization across multiple tasks. Moreover, summing the general representations obtained from foundation models poses challenges, including insufficient semantic information mining and feature fusion. In this paper, we propose a fusion-enhanced network, which effectively enriches semantic information and integrates features based on the complementary characteristics of infrared and visible modalities. The proposed network can extend to high-level vision tasks, showing strong generalization capabilities. Firstly, we adopt the infrared and visible foundation models to extract the general representations. Then, to enrich the semantic information of these general representations for high-level vision tasks, we design the feature enhancement module and the token enhancement module for feature maps and tokens, respectively. Besides, the attention-guided fusion module is proposed for effective fusion by exploring the complementary information of two modalities. Moreover, we adopt the cutout&mix augmentation strategy to conduct the data augmentation, which further improves the ability of the model to mine the regional complementarity between the two modalities. Extensive experiments show that the proposed method outperforms state-of-the-art dual-modality methods in the semantic segmentation, object detection, and salient object detection tasks. Fangcen Liu, Chenqiang Gao, Pengcheng Li 0017, Junjie Guo, Deyu Meng |
IEEE Trans. Multim. | 2 |
| 2025 | Frequency-prompt guided spectral-spatial transformer for hyperspectral image classification
Tiecheng Song, Longlong Zhang, Anyong Qin, Feng Yang 0015, Chenqiang Gao |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | Aerial video classification with Window Semantic Enhanced Video Transformers
Feng Yang 0015, Botong Zhou, Xuehua Guan, Anyong Qin, Tiecheng Song, Yue Zhao 0012, Chenqiang Gao |
Expert Syst. Appl. | 9 |
| 2025 | Global-local prompts guided image-text embedding, alignment and aggregation for multi-label zero-shot learning
Tiecheng Song, Feng Yang 0015, Anyong Qin, Yue Zhao 0012, Chenqiang Gao |
J. Vis. Commun. Image Represent. | 6 |
| 2025 | Occlusion-aware multi-person pose estimation with keypoint grouping and dual-prompt guidance in crowded scenes
Tiecheng Song, Anyong Qin, Yue Zhao 0012, Feng Yang 0015, Chenqiang Gao |
J. Vis. Commun. Image Represent. | 7 |
| 2025 | Dual-Branch Residual Network for Cross-Domain Few-Shot Hyperspectral Image Classification With Refined Prototype
Anyong Qin, Chaoqi Yuan, Feng Yang 0015, Tiecheng Song, Chenqiang Gao |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2025 | Dual-branch vision transformer for low-resolution action recognition
Ruixin Chen, Chenqiang Gao, Zhuolin Tan, Fangxin Liu, Jiayi Yu |
Multim. Tools Appl. | 2 |
| 2025 | A Point-Neighborhood Learning Framework for Nasal Endoscopic Image SegmentationabstractLesion segmentation on nasal endoscopic images is challenging due to its complex lesion features. Fully-supervised learning methods achieve promising performance with pixel-level annotations but impose a significant annotation burden on experts. Although weakly supervised or semi-supervised methods can reduce the labelling burden, their performance is still limited. Some weakly semi-supervised methods employ a novel annotation strategy that labels weak single-point annotations for the entire training set while providing pixel-level annotations for a small subset of the data. However, the relevant weakly semi-supervised methods only mine the limited information of the point itself, while ignoring its label property and surrounding reliable information. This paper proposes a simple yet efficient weakly semi-supervised method called the Point-Neighborhood Learning (PNL) framework. PNL incorporates the surrounding area of the point, referred to as the point-neighborhood, into the learning process. In PNL, we propose a point-neighborhood supervision loss and a pseudo-label scoring mechanism to explicitly guide the model’s training. Meanwhile, we proposed a more reliable data augmentation scheme. The proposed method obviously improves performance without increasing the parameters of the segmentation neural network. Experimental results indicate that our method consistently achieves better performance compared to SOTA methods. Additional validation on colonoscopic polyp segmentation datasets confirms our method’s generalizability. Pengyu Jie, Wanquan Liu, Chenqiang Gao, Yihui Wen, Weiping Wen, Pengcheng Li 0017, Deyu Meng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | ConvFormer-CD: Hybrid CNN-Transformer With Temporal Attention for Detecting Changes in Remote Sensing ImageryabstractRecently, the combination of Transformers and convolutional neural networks (CNNs) has witnessed significant advancements in change detection (CD) tasks. However, it remains unexplored how to interactively integrate long-range dependency and local information to enhance the model’s global-local context awareness for effectively mitigating pseudo-changes. In addition, accurate identification and distinction of building changes from complex backgrounds still pose challenges due to the insufficient semantic context modeling across time between bi-temporal images. To address these issues, we propose a hybrid model ConvFormer-CD with parallel convolution and multihead self-attention (MSA). This combination enables better interaction of global and local information, thereby enhancing the adaptability to complex scenarios. Moreover, we introduce a novel module called Temporal Attention to establish cross-temporal semantic relationships between image pairs, effectively highlighting change regions by learning shared and nonshared semantics. This enables our model to accurately detect changed targets even in scenarios characterized by intricate geo-spatial arrangements and distributions. To further refine the differences in bi-temporal images, we propose a difference integration module (DIM) that connects the encoder and the decoder to fuse high-level semantic features across channels. We conduct extensive experiments on four benchmark datasets, including LEVIR-CD, LEVIR-CD+, WHU-CD, and S2Looking-CD, which demonstrates that the proposed ConvFormer-CD outperforms other state-of-the-art (SOTA) methods. Our codes will be available athttps://github.com/taomi-lab/ConvFormer-CD. Feng Yang 0015, Mengtao Li, Wenqiang Shu, Anyong Qin, Tiecheng Song, Chenqiang Gao, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Towards Student Actions in Classroom Scenes: New Dataset and BaselineabstractAnalyzing student actions is an important and challenging task in educational research. Existing efforts have been hampered by the lack of accessible datasets to capture the nuanced action dynamics in classrooms. In this paper, we present a new multi-labelStudent Action Video(SAV) dataset, specifically designed for action detection in classroom settings. The SAV dataset consists of 4,324 carefully trimmed video clips from 758 different classrooms, annotated with 15 distinct student actions. Compared to existing action detection datasets, the SAV dataset stands out by providing a wide range of real classroom scenarios, high-quality video data, and unique challenges, including subtle movement differences, dense object engagement, significant scale differences, varied shooting angles, and visual occlusion. These complexities introduce new opportunities and challenges to advance action detection methods. To benchmark this, we propose a novel baseline method based on a visual transformer, designed to enhance attention to key local details within small and dense object regions. Our method demonstrates excellent performance with a mean Average Precision (mAP) of 67.9% and 27.4% on the SAV and AVA datasets, respectively. This paper not only provides the dataset but also calls for further research into AI-driven educational tools that may transform teaching methodologies and learning outcomes. The code and dataset are released athttps://github.com/Ritatanz/SAV. Zhuolin Tan, Chenqiang Gao, Anyong Qin, Ruixin Chen, Tiecheng Song, Feng Yang 0015, Deyu Meng |
IEEE Trans. Multim. | 2 |
| 2024 | DAMSDet: Dynamic Adaptive Multispectral Detection Transformer with Competitive Query Selection and Adaptive Feature Fusion
Junjie Guo, Chenqiang Gao, Fangcen Liu, Deyu Meng, Xinbo Gao 0001 |
ECCV (27) | 2 |
| 2024 | InfMAE: A Foundation Model in the Infrared Modality
Fangcen Liu, Chenqiang Gao, Yaming Zhang, Junjie Guo, Deyu Meng |
ECCV (18) | 2 |
| 2024 | Learning Interaction-aware 3D Gaussian Splatting for One-shot Hand AvatarsabstractIn this paper, we propose to create animatable avatars for interacting hands with 3D Gaussian Splatting (GS) and single-image inputs. Existing GS-based methods designed for single subjects often yield unsatisfactory results due to limited input views, various hand poses, and occlusions. To address these challenges, we introduce a novel two-stage interaction-aware GS framework that exploits cross-subject hand priors and refines 3D Gaussians in interacting areas. Particularly, to handle hand variations, we disentangle the 3D presentation of hands into optimization-based identity maps and learning-based latent geometric features and neural texture maps. Learning-based features are captured by trained networks to provide reliable priors for poses, shapes, and textures, while optimization-based identity maps enable efficient one-shot fitting of out-of-distribution hands. Furthermore, we devise an interaction-aware attention module and a self-adaptive Gaussian refinement module. These modules enhance image rendering quality in areas with intra- and inter-hand interactions, overcoming the limitations of existing GS-based methods. Our proposed method is validated via extensive experiments on the large-scale InterHand2.6M dataset, and it significantly improves the state-of-the-art performance in image quality. Code and models will be released upon acceptance. Wanquan Liu, Xiaodan Liang, Yiqiang Yan, Yuhao Cheng, Chenqiang Gao |
NeurIPS | 7 |
| 2024 | TopologyFormer: structure transformer assisted topology reconstruction for point cloud completion
Zhenwei Jiang, Chenqiang Gao, Chuandong Liu, Fangcen Liu, Lijie Zhu |
Multim. Tools Appl. | 2 |
| 2024 | THISNet: Tooth Instance Segmentation on 3D Dental Models via Highlighting Tooth RegionsabstractAutomatic tooth instance segmentation on 3D dental models is crucial for digitizing dental treatments and enabling computer-assisted treatment planning. However, It is challenging since the tight arrangement of dental structures and the consequential impact of dental ailments on their morphological characteristics. To address these challenges, we propose a novel method called THISNet. Unlike existing methods, THISNet focuses on highlighting tooth regions rather than relying on bounding box detection, leading to improved accuracy in tooth segmentation and labeling. By incorporating the highlighted tooth regions with a tooth object affinity module, our method effectively integrates global contextual information, considering the relationships between neighboring teeth and their surrounding structures. THISNet adopts an end-to-end learning approach, reducing complexity and enhancing segmentation efficiency compared to multi-stage training methods. Experimental results demonstrate the superiority of THISNet over existing approaches, highlighting its potential in various dental clinical applications. Pengcheng Li 0017, Chenqiang Gao, Fangcen Liu, Deyu Meng, Yan Yan 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Deep Updated Subspace Networks for Few-Shot Remote Sensing Scene ClassificationabstractDue to the difficulty of manually labeling remote sensing scene images and the demand for the ability to recognize new scene classes, few-shot remote sensing scene classification (FSRSSC) has attracted more and more attention. At present, metric-based FSRSSC methods have made promising progress, especially the prototypical networks-based methods. However, due to the complexity of the background of remote sensing scene images, the prototype classifier, which takes the average features of support samples as the metric benchmark, retains the features of category-irrelevant objects and other background information in the image. This leads to a bad classification result. Therefore, in this work, we propose a FSRSSC method based on the deep updated subspace network (DUSN), which uses class subspace as a metric benchmark to represent the commonality of a category and can effectively mitigate the negative impact of irrelevant objects on the classifier. In addition, for the higher inter-class similarity and larger intra-class variance of remote sensing scene images, we further propose an inter-class constraint and an intra-class constraint to mitigate the classification confusion. We leverage the inter-class constraint to make the images of different classes as far apart as possible, and the intra-class constraint to keep the images of the same class clustered as closely together as possible. Experimental results on three public benchmark datasets demonstrate that our method performs better than the state-of-the-art methods for FSRSSC. Anyong Qin, Fuyang Chen, Lingyun Tang, Feng Yang 0015, Yue Zhao 0012, Chenqiang Gao |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Few-Shot Learning With Prototype Rectification for Cross-Domain Hyperspectral Image ClassificationabstractDeep learning has been extensively applied to hyperspectral image (HSI) classification and has achieved significant success. However, the number of labeled samples available for HSI classification tasks is typically limited in practical applications, which makes the high-accuracy of HSI small-sample classification still a challenging research task. Therefore, metric-based prototypical networks for few-shot learning (FSL) have become increasingly popular. However, the majority of existing FSL methods typically have problems with biased prototypes and domain shifts. To address these issues, this article proposed a prototype rectification network framework for cross-domain few-shot HSI classification. Specifically, to obtain more representative prototypes, we designed a query-guided prototype rectification module, which can rectify the feature distribution of the support set prototype and obtain a more representative prototype for subsequent training tasks. Then, we introduced a prototype-based interclass loss function to alleviate the interclass confusion that may result from prototype rectification. Furthermore, we construct an intermediate domain between the source domain and the target domain to alleviate domain shift, which helps mitigate the difficulties of domain transfer and achieve a more comprehensive domain alignment. The experimental results on four publicly available HSI datasets demonstrate that our proposed method outperforms the existing FSL methods. Anyong Qin, Chaoqi Yuan, Xiaoliu Luo, Feng Yang 0015, Tiecheng Song, Chenqiang Gao |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Exploring Hybrid Contrastive Learning and Scene-to-Label Information for Multilabel Remote Sensing Image ClassificationabstractMultilabel remote sensing (RS) image classification aims to predict multiple semantic labels from an RS image. Previous methods [e.g., graph convolution networks (GCNs)] focus on mining the relationships of multiple labels, neglecting that the scene information is closely related to labels. To remedy this deficiency, in this article we propose a novel end-to-end deep neural network for multilabel RS image classification. In the proposed network, we use the GCN as the base model and introduce several new components to improve the classification performance. First, we explore hybrid contrastive learning (CL), including supervised transformation-based CL and unsupervised mix-based CL, to explicitly learn discriminative scene representations. Then, we apply the GCN-based classifier to the learned scene representations to obtain initial label prediction scores. Meanwhile, we pass the scene representations to a softmax layer to predict the probability that each image belongs to each specific scene class and use the scene-to-label information with the law of total probability to calibrate the initial label prediction scores. Finally, we incorporate CL, scene classification, and multilabel classification into a unified learning framework using uncertainty to weigh different losses. Experimental results on two benchmark RS datasets demonstrate the superiority of our proposed network for multilabel image classification. Tiecheng Song, Shufen Bai, Feng Yang 0015, Chenqiang Gao, Haonan Chen 0001, Jun Li 0009 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Joint Classification of Hyperspectral and LiDAR Data Using Height Information Guided Hierarchical Fusion-and-Separation NetworkabstractHyperspectral image (HSI) and LiDAR data are complementary to each other, which can be combined to improve the classification performance. However, existing deep network models do not sufficiently consider their complementarity to design the network structure and loss functions. Moreover, there lacks a hierarchical mutual-assistance learning mechanism that leverages the modality-shared features to enhance the modality-specific ones and vice versa. In view of these, we propose a novel height information guided hierarchical fusion-and-separation network (HFSNet) for joint classification of HSI and LiDAR data. HFSNet consists of three major components, i.e., dual-structure feature encoders (DSFEs), feature fusion-and-separation blocks (F2SBs), and an edge decoder (ED). Specifically, the transformer and convolutional neural network are introduced in DSFEs to encode the spectral and spatial information of HSI and LiDAR data, respectively. In F2SBs, the deformable convolution-based height information guided fusion module and the modality separation refinement module are proposed to sequentially extract modality-shared and modality-specific features. Additionally, the ED is incorporated into our model to predict the LiDAR edge map from the HSI feature to improve the model’s generalization ability. As such, the learned features from HSI and LiDAR data are deeply fused and mutually enhanced. Experiments on three benchmark datasets show the superiority of HFSNet to the state-of-the-art methods for jointly classifying HSI and LiDAR data with limited training samples. Tiecheng Song, Chenqiang Gao, Haonan Chen 0001, Jun Li 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Change-Aware Cascaded Dual-Decoder Network for Remote Sensing Image Change DetectionabstractChange detection aims to detect changes of objects or scenes in remote sensing images, which is critical for observing the Earth’s surface. However, due to the insufficient correlation and aggregation of bitemporal features, the existing deep learning methods are still impacted by varied imaging conditions and complicated boundaries of ground objects in high-resolution remote sensing images. To tackle these challenges, we propose a change-aware cascaded dual-decoder network (CACD2Net), which integrates bitemporal features at different levels to facilitate learning change maps from coarse to fine, thus empowering the network to effectively identify changes and refine pixelwise boundaries in a progressive manner. Within the cascaded dual-decoder architecture, the change location decoder utilizes high-level features to generate a coarse change map, which approximates changes’ localization, while the mask refinement decoder further leverages low-level features to create a texture-aware map that captures more texture and structural information about the change regions. By using the coarse change map as guidance and directing the texture-aware map to focus on the details of changes, the boundaries can be gradually refined, ultimately resulting in an accurate change detection mask. We test our model on the season-varying change detection (SVCD) dataset and the Sun Yat-sen University change detection (SYSU-CD) dataset, and the experimental results show that our model surpasses other state-of-the-art change detection methods. Our codes will be available athttps://github.com/Moonquakes0/CACD2Net. Feng Yang 0015, Yifeng Yuan, Anyong Qin, Yue Zhao 0012, Tiecheng Song, Chenqiang Gao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Spatial Prior-Guided Bi-Directional Cross-Attention Transformers for Tooth Instance SegmentationabstractTooth instance segmentation of dental panoramic X-ray images is of significant clinical importance. Teeth exhibit symmetry within the upper and lower jawbones and are arranged in a specific order. However, previous studies frequently overlook this crucial spatial prior information, resulting in the misidentifications of tooth categories, especially for adjacent or similarly shaped teeth. In this paper, we propose SPGTNet, a spatial prior-guided transformer method, designed to both the extracted tooth positional features from CNNs and the long-range contextual information from vision transformers, specifically for dental panoramic X-ray image segmentation. Initially, a center-based spatial prior perception module is employed to identify the centroid of each tooth, thereby enhancing the spatial prior information for the CNN sequence features. Subsequently, a bi-directional cross-attention module is designed to facilitate the interaction between the spatial prior information of the CNN sequence features and the long-range contextual features of the vision transformer sequence features. Finally, an instance identification head is employed to generate the tooth segmentation results. Extensive experiments on three public benchmark datasets demonstrate the effectiveness and superiority of our proposed method compared to other state-of-the-art approaches. The proposed method accurately identifies and analyzes tooth structures, thereby providing crucial information for dental diagnosis, treatment planning, and research. Pengcheng Li 0017, Chenqiang Gao, Chunfeng Lian, Deyu Meng |
IEEE Trans. Medical Imaging | 2 |
| 2023 | Hierarchical Supervision and Shuffle Data Augmentation for 3D Semi-Supervised Object DetectionabstractState-of-the-art 3D object detectors are usually trained on large-scale datasets with high-quality 3D annotations. However, such 3D annotations are often expensive and time-consuming, which may not be practical for real applications. A natural remedy is to adopt semi-supervised learning (SSL) by leveraging a limited amount of labeled samples and abundant unlabeled samples. Current pseudo-labeling-based SSL object detection methods mainly adopt a teacher-student framework, with a single fixed threshold strategy to generate supervision signals, which inevitably brings confused supervision when guiding the student network training. Besides, the data augmentation of the point cloud in the typical teacher-student framework is too weak, and only contains basic down sampling and flip-and-shift (i.e., rotate and scaling), which hinders the effective learning of feature information. Hence, we address these issues by introducing a novel approach of Hierarchical Supervision and Shuffle Data Augmentation (HSSDA), which is a simple yet effective teacher-student framework. The teacher network generates more reasonable supervision for the student network by designing a dynamic dual-threshold strategy. Besides, the shuffle data augmentation strategy is designed to strengthen the feature representation ability of the student network. Extensive experiments show that HSSDA consistently outperforms the recent state-of-the-art methods on different datasets. The code will be released at https://github.com/azhuantou/HSSDA. Chuandong Liu, Chenqiang Gao, Fangcen Liu, Pengcheng Li 0017, Deyu Meng, Xinbo Gao 0001 |
CVPR | 2 |
| 2023 | TransMIN: Transformer-Guided Multi-Interaction Network for Remote Sensing Object DetectionabstractRemote sensing (RS) object detectors based on convolutional neural networks (CNNs) are hard to model global context dependencies. Transformer-based detectors can overcome this problem via global pairwise interactions, but more comprehensive information interactions are not systemically investigated to boost the detection performance. In view of this issue, we propose a transformer-guided multi-interaction network (TransMIN), which uses ResNet50-feature pyramid network (FPN) as the backbone for remote sensing object detection (RSOD). Specifically, we implement local–global feature interactions (LGFIs) by combining convolution and transformer in the residual blocks of ResNet50 to learn complementary features. We implement cross-view feature interactions (CVFIs) via transformers in the pyramid layers of FPN to capture the correlation between reference features (spatial edge priors and channel statistics) and pyramid features. This enhances edge information and suppresses background interference. In the detection head, we adopt a task-interactive sample assigner (TISA) by considering the interactions of classification and localization losses to obtain high-quality predictions. Experiments on two benchmark datasets demonstrate the superior detection performance of TransMIN over state-of-the-art methods. Guangming Xu, Tiecheng Song, Chenqiang Gao |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Pareto Refocusing for Drone-View Object DetectionabstractDrone-view Object Detection (DOD) is a meaningful but challenging task. It hits a bottleneck due to two main reasons: (1) The high proportion of difficult objects (e.g., small objects, occluded objects, etc.) makes the detection performance unsatisfactory. (2) The unevenly distributed objects make detection inefficient. These two factors also lead to a phenomenon, obeying the Pareto principle, that some challenging regions occupying a low area proportion of the image have a significant impact on the final detection while the vanilla regions occupying the major area have a negligible impact due to the limited room for performance improvement. Motivated by the human visual system that naturally attempts to invest unequal energies in things of hierarchical difficulty for recognizing objects effectively, this paper presents a novel Pareto Refocusing Detection (PRDet) network that distinguishes the challenging regions from the vanilla regions under reverse-attention guidance and refocuses the challenging regions with the assistance of the region-specific context. Specifically, we first propose a Reverse-attention Exploration Module (REM) that excavates the potential position of difficult objects by suppressing the features which are salient to the commonly used detector. Then, we propose a Region-specific Context Learning Module (RCLM) that learns to generate specific contexts for strengthening the understanding of challenging regions. It is noteworthy that the specific context is not shared globally but unique for each challenging region with the exploration of spatial and appearance cues. Extensive experiments and comprehensive evaluations on the VisDrone2021-DET and UAVDT datasets demonstrate that the proposed PRDet can effectively improve the detection performance, especially for those difficult objects, outperforming state-of-the-art detectors. Furthermore, our method also achieves significant performance improvements on the DTU-Drone dataset for power inspection. Jiaxu Leng, Mengjingcheng Mo, Yinghua Zhou, Chenqiang Gao, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Distance Constraint-Based Generative Adversarial Networks for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification suffers from two serious problems, one is the limited labeled pixels, and the other is the class imbalance problem. As a result, the number of labeled pixels in many categories is not sufficient to characterize the spectral-spatial information, and train a satisfying deep model. By making full use of the information of unlabeled pixels, semi-supervised methods can provide better classification performance in the case of limited labeled pixels. However, they do not take into account the imbalance in the HSI data. As a method of data enhancement, generative adversarial networks focus on the above two problems and have also been widely used for the task of the HSI classification. In this work, we propose a distance constraints-based generative adversarial networks (DGAN) method for HSI classification to address these two problems. The DGAN employs the convolution autoencoder (AE) to extract the latent features of the HSI samples, and considers the reconstructed samples from the AE as the real samples for the later classifier and discriminator. In addition, the DGAN uses two distance constraints to solve the problems of the few labeled samples and class imbalance, the one latent-data distance constraint enforcing the generator to generate HSI samples for each class (especially the minority class), another discriminator-score distance constraint guiding the generator to synthesize samples that resemble the real HSI samples. Finally, the generated samples are combined classwise with the reconstructed samples and the real HSI samples to learn the parameters of the classifier and discriminator. Experimental results show that our method achieves state-of-the-art performance in terms of overall accuracy (OA) when trained with only 0.5%-4% of data sets from Indian Pines, Pavia University, and Botswana. Specifically, our method demonstrates improvements of 5.48%, 8.79%, and 0.91% on these three datasets, respectively. It reveals the great potential of the DGAN model in generating the HSI samples for each class, which contributes to improving the classification performance of the HSI data. Anyong Qin, Zhuolin Tan, Yongqing Sun, Feng Yang 0015, Yue Zhao 0012, Chenqiang Gao |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | Infrared Small and Dim Target Detection With Transformer Under Complex BackgroundsabstractThe infrared small and dim (S&D) target detection is one of the key techniques in the infrared search and tracking system. Since the local regions similar to infrared S&D targets spread over the whole background, exploring the correlation amongst image features in large-range dependencies to mine the difference between the target and background is crucial for robust detection. However, existing deep learning-based methods are limited by the locality of convolutional neural networks, which impairs the ability to capture large-range dependencies. Additionally, the S&D appearance of the infrared target makes the detection model highly possible to miss detection. To this end, we propose a robust and general infrared S&D target detection method with the transformer. We adopt the self-attention mechanism of the transformer to learn the correlation of image features in a larger range. Moreover, we design a feature enhancement module to learn discriminative features of S&D targets to avoid miss-detections. After that, to avoid the loss of the target information, we adopt a decoder with the U-Net-like skip connection operation to contain more information of S&D targets. Finally, we get the detection result by a segmentation head. Extensive experiments on two public datasets show the obvious superiority of the proposed method over state-of-the-art methods, and the proposed method has a stronger generalization ability and better noise tolerance. Fangcen Liu, Chenqiang Gao, Deyu Meng, Wangmeng Zuo, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | SS3D: Sparsely-Supervised 3D Object Detection from Point CloudabstractConventional deep learning based methods for 3D object detection require a large amount of 3D bounding box annotations for training, which is expensive to obtain in practice. Sparsely annotated object detection, which can largely reduce the annotations, is very challenging since the missing-annotated instances would be regarded as the background during training. In this paper, we propose a sparsely-supervised 3D object detection method, named SS3D. Aiming to eliminate the negative supervision caused by the missing annotations, we design a missing-annotated instance mining module with strict filtering strategies to mine positive instances. In the meantime, we design a reliable background mining module and a point cloud filling data augmentation strategy to generate the confident data for iteratively learning with reliable supervision. The proposed SS3D is a general framework that can be used to learn any modern 3D object detector. Extensive experiments on the KITTI dataset reveal that on different 3D detectors, the proposed SS3D framework with only 20% annotations required can achieve onpar performance comparing to fully-supervised methods. Comparing with the state-of-the-art semi-supervised 3D objection detection on KITTI, our SS3D improves the benchmarks by significant margins under the same annotation workload. Moreover, our SS3D also out-performs the state-of-the-art weakly-supervised method by remarkable margins, highlighting its effectiveness. Chuandong Liu, Chenqiang Gao, Fangcen Liu, Jiang Liu 0011, Deyu Meng, Xinbo Gao 0001 |
CVPR | 2 |
| 2022 | Active Learning for Hyperspectral Image Classification via Hypergraph Neural NetworkabstractGraph convolution network (GCN) has been extensively applied to the area of hyperspectral image (HSI) classification. However, the graph can not effectively describe the complex relationships between HSI pixels and the GCN still faces the challenge of insufficient labeled pixels. In order to alleviate the above two issues faced by the GCN in HSI classification, we propose a novel framework that integrates the active learning and the hypergraph neural network. First, we construct a hypergraph that can reveal the complex non-pairwise relationships embedded in the hyperspectral images. Next, we train a semi-supervised hypergraph neural network (GNN) with the fewer labeled training set. Then, exploiting the local structural properties of the hypergraph, the most useful HSI pixels are actively selected for labeling. Finally, we fine-tune the GNN with original training set along with the newly labeled pixels. And the last three steps are iteratively carried on. Compared with the other traditional and active learning approaches of HSI classification, the proposed active hypergraph neural network (ACGNN) can achieve better performance on the three HSI datasets. Yongqing Sun, Anyong Qin, Yukihiro Bandoh, Chenqiang Gao, Yusuke Hiwasaki |
ICIP | 4 |
| 2022 | Multiscale Spatio-Temporal Network for Aerial Video Event RecognitionabstractUnmanned aerial vehicles (UAVs) are widely used in the field of remote sensing because of their advantages of providing real-time and high-resolution videos at a low cost. Compared with generic video understanding, aerial video event recognition is faced with emerging challenges: 1) aerial videos contain richer scene information; 2) the scale variations between different videos are large. To address these issues, we propose a Multiscale Spatio-Temporal Network (MSTN) in this paper. More precisely, the MSTN consists of a Pyramid Spatio-Temporal (PST) module and a Multi-Time Scale Decision (MTSD) module, which learn multi-scale spatio-temporal features together. The two modules can better learn spatio-temporal characteristics and boost the performance by 3.7% compared with the baseline method. In ERA, an aerial event recognition dataset, our method achieves the state-of-the-art results. Feng Yang 0015, Yue Zhao 0012, Anyong Qin, Chenqiang Gao |
IGARSS | 5 |
| 2022 | ICNet: Joint Alignment and Reconstruction via Iterative Collaboration for Video Super-ResolutionabstractMost previous frameworks either cost too much time or adopt some fixed modules resulting in alignment error in video super-resolution (VSR). In this paper, we propose a novel many-to-many VSR framework with Iterative Collaboration (ICNet), which employs the concurrent operation by iterative collaboration between alignment and reconstruction proving to be more efficient and effective than existing recurrent and sliding-window frameworks. With the proposed iterative collaboration, alignment can be conducted on super-resolved features from reconstruction while accurate alignment boosts reconstruction in return. In each iteration, the features of low-resolution video frames are first fed into the alignment and reconstruction subnetworks, which can generate temporal aligned features and spatial super-resolved features. Then, both outputs are fed into the proposed Tidy Two-stream Fusion (TTF) subnetwork that shares inter-frame temporal information and intra-frame spatial information without redundancy. Moreover, we design the Frequency Separation Reconstruction (FSR) subnetwork to not only model high-frequency and low-frequency information separately but also take benefit of each other for better reconstruction. Extensive experiments on benchmark datasets demonstrate that the proposed ICNet outperforms state-of-the-art VSR methods in terms of PSNR/SSIM values and visual quality, respectively. Jiaxu Leng, Jia Wang 0036, Xinbo Gao 0001, Bo Hu 0008, Ji Gan, Chenqiang Gao |
ACM Multimedia | 6 |
| 2022 | Spectral-Spatial Residual Graph Attention Network for Hyperspectral Image ClassificationabstractHyperspectral images (HSIs) not only possess abundant spectral features but also present a detailed spatial distribution of land cover, and they have significant advantages in the fine classification of ground materials. Recently, using convolutional neural networks (CNNs) to extract spectral–spatial features has become an effective way for HSI classification. However, conventional convolution kernels learn features from fixed regular square regions, and rich spatial information has not been effectively explored. In this letter, an end-to-end model named spectral–spatial residual graph attention network (S2RGANet) is developed for HSI classification, and it has two crucial elements, including spectral residual and graph attention convolution modules. At first, two spectral residual modules are employed to capture discriminant spectral features. Then, graphs are constructed to reveal the relationship between points in local neighborhoods. By graph attention mechanism, local spatial information is adaptively aggregated from neighboring nodes. Experiments on two public HSI datasets demonstrate that the S2RGANet is significantly superior to some state-of-the-art (SOTA) methods with limited training samples. Kejie Xu, Yue Zhao 0012, Chenqiang Gao, Hong Huang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | M2FN: A Multilayer and Multiattention Fusion Network for Remote Sensing Image Scene ClassificationabstractDeep convolutional neural networks (CNNs) have made great progress in remote sensing (RS) image scene classification. However, by visualizing the learned feature maps, we find that the popular CNN of ResNet can capture incomplete and inaccurate semantic information for classifying scene images with complex spatial distributions and varying object scales. In this letter, we propose a multilayer and multiattention fusion network (M2FN) to alleviate this issue. Specifically, we first introduce a multilayer adaptive feature fusion (MLAFF) module to model the information interaction between different layers and enhance the network’s multiscale representation ability. Then, we design a multidimensional attention (MA) module to weight the multilayer fused features by comprehensively considering their interdependencies between all possible dimensions. The proposed MA module extends the traditional spatial and channel attentions to a more comprehensive one. Experiments on two benchmark data sets demonstrate the superiority of M2FN for RS scene classification over many state-of-the-art methods. Tiecheng Song, Chenqiang Gao, Tan Guo |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | MSLAN: A Two-Branch Multidirectional Spectral-Spatial LSTM Attention Network for Hyperspectral Image ClassificationabstractRecurrent neural networks (RNNs) have been widely used for hyperspectral image (HSI) classification via sequence modeling. However, most of the RNN methods focus on modeling long-range dependencies along the spectral direction, without fully exploring multi-directional dependencies in the joint spectral-spatial domain. To tackle this issue, we propose MSLAN, a two-branch multi-directional spectral-spatial long short-term memory (LSTM) attention network, for HSI classification. In particular, we employ LSTMs to extract six-directional spatial-spectral features which simultaneously capture the spectral-spatial dependencies along different directions. We then design an attention-based feature fuse module to integrate these directional features, followed by a fully connected layer with cross-entropy loss for classification. Additionally, we incorporate an auxiliary branch into our model to enhance the generalization capability. In this branch, random spatial shuffle and a cosine loss are explored for feature consistency learning by taking into account the varying spatial distributions. The resulting two branch networks, sharing the same network structure and weights, are incorporated into a unified deep learning architecture for training. Experiments show the superiority of MSLAN to the state-of-the-art methods for HSI classification with limited training samples. Tiecheng Song, Yuanlin Wang, Chenqiang Gao, Haonan Chen 0001, Jun Li 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Semantic Graph Attention With Explicit Anatomical Association Modeling for Tooth Segmentation From CBCT ImagesabstractAccurate tooth identification and delineation in dental CBCT images are essential in clinical oral diagnosis and treatment. Teeth are positioned in the alveolar bone in a particular order, featuring similar appearances across adjacent and bilaterally symmetric teeth. However, existing tooth segmentation methods ignored such specific anatomical topology, which hampers the segmentation accuracy. Here we propose a semantic graph-based method to explicitly model the spatial associations between different anatomical targets (i.e., teeth) for their precise delineation in a coarse-to-fine fashion. First, to efficiently control the bilaterally symmetric confusion in segmentation, we employ a lightweight network to roughly separate teeth as four quadrants. Then, designing a semantic graph attention mechanism to explicitly model the anatomical topology of the teeth in each quadrant, based on which voxel-wise discriminative feature embeddings are learned for the accurate delineation of teeth boundaries. Extensive experiments on a clinical dental CBCT dataset demonstrate the superior performance of the proposed method compared with other state-of-the-art approaches. Pengcheng Li 0017, Yang Liu 0157, Zhiming Cui 0001, Feng Yang 0015, Yue Zhao 0012, Chunfeng Lian, Chenqiang Gao |
IEEE Trans. Medical Imaging | 7 |
| 2022 | Two-Stream Graph Convolutional Network for Intra-Oral Scanner Image SegmentationabstractPrecise segmentation of teeth from intra-oral scanner images is an essential task in computer-aided orthodontic surgical planning. The state-of-the-art deep learning-based methods often simply concatenate the raw geometric attributes (i.e., coordinates and normal vectors) of mesh cells to train a single-stream network for automatic intra-oral scanner image segmentation. However, since different raw attributes reveal completely different geometric information, the naive concatenation of different raw attributes at the (low-level) input stage may bring unnecessary confusion in describing and differentiating between mesh cells, thus hampering the learning of high-level geometric representations for the segmentation task. To address this issue, we design a two-stream graph convolutional network (i.e., TSGCN), which can effectively handle inter-view confusion between different raw attributes to more effectively fuse their complementary information and learn discriminative multi-view geometric representations. Specifically, our TSGCN adopts two input-specific graph-learning streams to extract complementary high-level geometric representations from coordinates and normal vectors, respectively. Then, these single-view representations are further fused by a self-attention module to adaptively balance the contributions of different views in learning more discriminative multi-view representations for accurate and fully automatic tooth segmentation. We have evaluated our TSGCN on a real-patient dataset of dental (mesh) models acquired by 3D intraoral scanners. Experimental results show that our TSGCN significantly outperforms state-of-the-art methods in 3D tooth (surface) segmentation. Yue Zhao 0012, Yang Liu 0157, Deyu Meng, Zhiming Cui 0001, Chenqiang Gao, Xinbo Gao 0001, Chunfeng Lian, Dinggang Shen |
IEEE Trans. Medical Imaging | 6 |
| 2022 | Infrared Action Detection in the Dark via Cross-Stream Attention MechanismabstractAction detection plays an important role in the field of video understanding and attracts considerable attention in the last decade. However, current action detection methods are mainly based on visible videos, and few of them consider scenes with low-light, where actions are difficult to be detected by existing methods, or even by human eyes. Compared with visible videos, infrared videos are more suitable for the dark environment and resistant to background clutter. In this paper, we investigate the temporal action detection problem in the dark by using infrared videos, which is, to the best of our knowledge, the first attempt in the action detection community. Our model takes the whole video as input, a Flow Estimation Network (FEN) is employed to generate the optical flow for infrared data, and it is optimized with the whole network to obtain action-related motion representations. After feature extraction, the infrared stream and flow stream are fed into a Selective Cross-stream Attention (SCA) module to narrow the performance gap between infrared and visible videos. The SCA emphasizes informative snippets and focuses on the more discriminative stream automatically. Then we adopt a snippet-level classifier to obtain action scores for all snippets and link continuous snippets into final detection results. All these modules are trained in an end-to-end manner. We collect an Infrared action Detection (InfDet) dataset obtained in the dark and conduct extensive experiments to verify the effectiveness of the proposed method. Experimental results show that our proposed method surpasses the state-of-the-art temporal action detection methods designed for visible videos, and it also achieves the best performance compared with other infrared action recognition methods on both InfAR and Infrared-Visible datasets. Xu Chen 0053, Chenqiang Gao, Chaoyu Li, Yi Yang 0001, Deyu Meng |
IEEE Trans. Multim. | 2 |
| 2021 | TSGCNet: Discriminative Geometric Feature Learning With Two-Stream Graph Convolutional Network for 3D Dental Model SegmentationabstractThe ability to segment teeth precisely from digitized 3D dental models is an essential task in computer-aided orthodontic surgical planning. To date, deep learning based methods have been popularly used to handle this task. State-of-the-art methods directly concatenate the raw attributes of 3D inputs, namely coordinates and normal vectors of mesh cells, to train a single-stream network for fully-automated tooth segmentation. This, however, has the drawback of ignoring the different geometric meanings provided by those raw attributes. This issue might possibly confuse the network in learning discriminative geometric features and result in many isolated false predictions on the dental model. Against this issue, we propose a two-stream graph convolutional network (TSGCNet) to learn multi-view geometric information from different geometric attributes. Our TSGCNet adopts two graph-learning streams, designed in an input-aware fashion, to extract more discriminative high-level geometric representations from coordinates and normal vectors, respectively. These feature representations learned from the designed two different streams are further fused to integrate the multi-view complementary information for the cell-wise dense prediction task. We evaluate our proposed TSGCNet on a real-patient dataset of dental models acquired by 3D intraoral scanners, and experimental results demonstrate that our method significantly outperforms state-of-the-art methods for 3D shape segmentation. Yue Zhao 0012, Deyu Meng, Zhiming Cui 0001, Chenqiang Gao, Xinbo Gao 0001, Chunfeng Lian, Dinggang Shen |
CVPR | 5 |
| 2021 | Video-to-Image Casting: A Flatting Method for Video AnalysisabstractPrevious mainstream video analysis methods, especially 3D CNNs-based models, mainly aim to transfer frameworks from the image domain to the video domain, and they follow the regime which has been succeeded in image processing, i.e., large-scale benchmarks and deep networks. However, processing videos is still time-consuming due to the increased computational cost. In this paper, we propose to flat the video and construct a Spatio-temporal Image (STI), i.e., squeezing the temporal dimension into a spatial plane. To pursuit the video-level modeling and efficient architecture, we devise a Collective Convolution (CoConv) operation to replace the 2D convolution. With the holistic sampling strategy, this novel operation can extract the video-level spatio-temporal representation. Moreover, we ensure that each CoConv operation has the same number of parameters as the original 2D filter, thus we can utilize a 2D network equipped with CoConv to analyze videos without additional computations. To verify the effectiveness of our method for the general video analysis, we evaluate it on three typical tasks, i.e., supervised action recognition, self-supervised action recognition, and dynamic texture recognition. Extensive experimental results show that our method can achieve comparable or state-of-the-art performances on these benchmarks while using much fewer computations compared with its 3D counterpart. Xu Chen 0053, Chenqiang Gao, Feng Yang 0015, Yi Yang 0001, Yahong Han |
ACM Multimedia | 2 |
| 2021 | Multi-scale single-stage pose detection with adaptive sample training in the classroom scene
Chenqiang Gao, Hang Tian, Yan Yan 0002 |
Knowl. Based Syst. | 1 |
| 2021 | MSLPNet: multi-scale location perception network for dental panoramic X-ray image segmentation
Qiaoyi Chen, Yue Zhao 0012, Yang Liu 0157, Yongqing Sun, Chongshi Yang, Pengcheng Li 0017, Chenqiang Gao |
Neural Comput. Appl. | 8 |
| 2021 | Quaternionic extended local binary pattern with adaptive structural pyramid pooling for color image representation
Tiecheng Song, Liangliang Xin, Chenqiang Gao |
Pattern Recognit. | 3 |
| 2021 | 3D Dental model segmentation with graph attentional convolution network
Yue Zhao 0012, Chongshi Yang, Yingyun Tan, Yang Liu 0157, Pengcheng Li 0017, Chenqiang Gao |
Pattern Recognit. Lett. | 8 |
| 2021 | Infrared and Visible Cross-Modal Image Retrieval Through Shared FeaturesabstractImage retrieval is one of the key techniques of computer vision, and has been studied for a long time. Nevertheless, little attention is paid to infrared and visible cross-modal retrieval which can be widely used in various applications, e.g., infrared and visible surveillance systems. In this paper, we propose a shared features based infrared-visible cross-modal image retrieval method. The similar visual features are extracted from infrared and visible images as the shared features, and the Euclidean distance is used to measure the similarity between these features. The core of the proposed method comes from three aspects: 1) Feature separation network can separate image features into shared features and exclusive features; 2) Maximum Mean Discrepancy (MMD) loss is employed to constrain the distribution of shared features, which can reduce the retrieval error caused by different imaging angles and similarity of infrared images. 3) The cross-layer fusion encoder compensates for the context loss in the convolution of infrared images. Experimental results on the Infrared-Visible dataset demonstrate the proposed method is effective and outperforms the state-of-the-art approaches. Fangcen Liu, Chenqiang Gao, Yongqing Sun, Yue Zhao 0012, Feng Yang 0015, Anyong Qin, Deyu Meng |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Robust Texture Description Using Local Grouped Order Pattern and Non-Local Binary PatternabstractLocal binary pattern (LBP) and its many variants have shown effectiveness for texture classification. However, most of these LBP methods focus on encoding local intensity differences between a central pixel and its neighboring sampling points and consequently have two major problems: 1) they are unable to describe the intensity order relationships among neighboring sampling points, and 2) they fail to capture long-range pixel interactions that take place outside a compact neighborhood. In view of these problems, in this paper we propose two novel operators, called local grouped order pattern (LGOP) and non-local binary pattern (NLBP), for texture description. For the first problem, LGOP groups the neighboring sampling points by referring to a dominant direction and encodes the groupwise intensity order relationships. For the second problem, NLBP computes several anchors based on global image statistics and progressively encodes non-local intensity differences between the neighboring sampling points and anchors. Finally, we combine LGOP and NLBP via central pixel encoding to construct discriminative histogram features as texture descriptor LGONBP. Experiments on four texture benchmark databases (i.e., Outex, CUReT, UMD and KTH-TIPS) demonstrate the superiority of LGONBP over state-of-the-art LBP variants for texture classification under both noise-free and noisy conditions. The code is available athttps://github.com/stc-cqupt/LGONBP. Tiecheng Song, Jie Feng 0007, Lin Luo 0007, Chenqiang Gao, Hongliang Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Color Texture Description Based on Holistic and Hierarchical Order-Encoding PatternsabstractLocal binary pattern (LBP), as one of the most representative texture operators, has attracted much attention in computer vision and pattern recognition. Many LBP variants were developed in the literature. However, most of them were designed for gray images and their performance remains to be improved for color images. In this paper, we propose a novel color image descriptor named Holistic and Hierarchical Order-Encoding Patterns (H2OEP) for texture classification. In H2OEP, the holistic order-encoding pattern compactly encodes color order variation tendencies for each pixel in color space. The hierarchical order-encoding pattern leverages min ordering, median ordering and max ordering to encode local neighboring relationships across different color channels. Finally, the generated order-encoding patterns are aggregated via central pixel encoding to build 3D joint histograms for image representation. Experiments on four benchmark texture databases demonstrate the effectiveness of the proposed descriptor for color texture classification. Tiecheng Song, Jie Feng 0007, Yuanlin Wang, Chenqiang Gao |
ICPR | 4 |
| 2020 | First- and Second-Order Sorted Local Binary Pattern Features for Grayscale-Inversion and Rotation Invariant Texture ClassificationabstractLocal binary pattern (LBP) is sensitive to inverse grayscale changes. Several methods address this problem by mapping each LBP code and its complement to the minimum one. However, without distinguishing LBP codes and their complements, these methods show limited discriminative power. In this paper, we introduce a histogram sorting method to preserve the distribution information of LBP codes and their complements. Based on this method, we propose first- and second-order sorted LBP (SLBP) features which are robust to inverse grayscale changes and image rotation. The proposed method focuses on encoding difference-sign information and it can be generalized to embed other difference-magnitude features to obtain complementary representations. Experiments demonstrate the effectiveness of our method for texture classification under (linear or nonlinear) grayscale-inversion and rotation changes. Tiecheng Song, Yuanjing Han, Jie Feng 0007, Yuanlin Wang, Chenqiang Gao |
ICPR | 5 |
| 2020 | Adaptive Fusion and Mask Refinement Instance Segmentation Network for High Resolution Remote Sensing ImagesabstractInstance segmentation of remote sensing images (RSIs) is an active yet challenging task because of the huge scale variation and arbitrary complex shapes of objects. To address these issues, we propose an adaptive fusion and mask refinement (AFMR) instance segmentation network for RSIs in this paper. More precisely, AFMR consists of an adaptive fusion module to learn multi-scale complementary spatial features in an unsupervised manner, and a content-aware module for segmentation mask refinement. These two modules enable a better feature learning of convolutional neural network and boost the performance by 1.5% compared with the baseline method. In iSAID, a large-scale dataset for RSIs instance segmentation, our AFMR framework achieves the state-of-the-art accuracy, which verifies the superiority of the proposed method. Jie Ran, Feng Yang 0015, Chenqiang Gao, Yue Zhao 0012, Anyong Qin |
IGARSS | 3 |
| 2020 | TSASNet: Tooth segmentation on dental panoramic X-ray images by Two-Stage Attention Segmentation Network
Yue Zhao 0012, Pengcheng Li 0017, Chenqiang Gao, Yang Liu 0157, Qiaoyi Chen, Feng Yang 0015, Deyu Meng |
Knowl. Based Syst. | 3 |
| 2020 | Spatio-temporal fall event detection in complex scenes using attention guided LSTM
Chenqiang Gao, Yue Zhao 0012, Tiecheng Song |
Pattern Recognit. Lett. | 2 |
| 2020 | Polycrystalline silicon wafer defect segmentation based on deep convolutional neural networks
Chenqiang Gao, Yue Zhao 0012, Shisha Liao, Xindou Li |
Pattern Recognit. Lett. | 2 |
| 2020 | Bayesian query expansion for multi-camera person re-identification
Yutian Lin, Zhedong Zheng, Chenqiang Gao, Yi Yang 0001 |
Pattern Recognit. Lett. | 4 |
| 2019 | Weakly Supervised Instance Segmentation Using Hybrid NetworksabstractWeakly-supervised instance segmentation, which could greatly save labor and time cost of pixel mask annotation, has attracted increasing attention in recent years. The commonly used pipeline firstly utilizes conventional image segmentation methods to automatically generate initial masks and then use them to train an off-the-shelf segmentation network in an iterative way. However, the initial generated masks usually contains a notable proportion of invalid masks which are mainly caused by small object instances. Directly using these initial masks to train segmentation models is harmful for the performance. To address this problem, we propose a kind of hybrid networks in this paper. In our architecture, there is a principle segmentation network which is used to handle the normal samples with valid generated masks. In addition, a complementary branch is added to handle the small and dim objects without valid masks. Experimental results indicate that our method can achieve significantly performance improvement both on the small object instances and large ones, and outperforms all state-of-the-art methods. Shisha Liao, Yongqing Sun, Chenqiang Gao, Pranav Shenoy K. P, Song Mu, Jun Shimamura, Atsushi Sagata |
ICASSP | 3 |
| 2019 | Texture Representation Using Local Binary Encoding Across Scales, Frequency Bands and Image DomainsabstractMost of the local binary pattern (LBP) variants improve LBP by simply concatenating multi-scale LBP histograms or by encoding complementary components in one single image domain, thereby ignoring the correlation information between different scales and image domains. In this paper, we propose a novel LBP-based texture representation by exploring local binary encoding across scales, frequency bands and image domains. Specifically, given a texture image, the multi-scale low- and high-frequency images are obtained by Gaussian filtering and image subtraction. Meanwhile, the multi-scale gradient images are computed based on Gaussian derivative filtering. Then, the LBP code maps are extracted from the low-frequency, high-frequency and gradient images. Finally, the joint LBP encoding across scales, frequency bands and image domains is explored to construct histogram features for texture representation. Experimental results for texture classification demonstrate the superiority of our method over the state-of-the-art LBP variants under both noise-free and noisy conditions. Tiecheng Song, Lin Luo 0007, Chenqiang Gao |
ICIP | 3 |
| 2019 | Pose detection in complex classroom environment based on improved Faster R-CNNabstractPose detection of small targets in poor imaging conditions like heavy occlusion and low resolution is still an open and challenging task in computer vision. For instance, detection of students' poses in classrooms that are even indistinguishable to human eyes remains a rather difficult task. Motivated by the success of convolutional feature merging and locality preserving, the authors propose a pose detection framework combining merged region of interest (ROI) pooling and locality preserving learning. Unlike usual object detection algorithms which use general top‐level convolutional features as inputs, their method uses a merged ROI pooling structure to merge semantic feature and high‐resolution feature from the last two levels of convolutional feature maps, so that this merged feature is made more expressive than the single‐level feature. In addition, the locality feature‐preserving learning is used in the last fully‐connected layer. Through locality preserving learning, features belonging to the same class would be forced to be closer in the feature space, which enables the model with stronger classification ability. Experimental results show that the proposed method outperforms the state‐of‐the‐art methods. Chenqiang Gao, Xu Chen 0053, Yue Zhao 0012 |
IET Image Process. | 2 |
| 2018 | DecideNet: Counting Varying Density Crowds Through Attention Guided Detection and Density EstimationabstractIn real-world crowd counting applications, the crowd densities vary greatly in spatial and temporal domains. A detection based counting method will estimate crowds accurately in low density scenes, while its reliability in congested areas is downgraded. A regression based approach, on the other hand, captures the general density information in crowded regions. Without knowing the location of each person, it tends to overestimate the count in low density areas. Thus, exclusively using either one of them is not sufficient to handle all kinds of scenes with varying densities. To address this issue, a novel end-to-end crowd counting framework, named DecideNet (DEteCtIon and Density Estimation Network) is proposed. It can adaptively decide the appropriate counting mode for different locations on the image based on its real density conditions. DecideNet starts with estimating the crowd density by generating detection and regression based density maps separately. To capture inevitable variation in densities, it incorporates an attention module, meant to adaptively assess the reliability of the two types of estimations. The final crowd counts are obtained with the guidance of the attention module to adopt suitable estimations from the two kinds of density maps. Experimental results show that our method achieves state-of-the-art performance on three challenging crowd counting datasets. Jiang Liu 0011, Chenqiang Gao, Deyu Meng, Alex Hauptmann 0001 |
CVPR | 2 |
| 2018 | PM-GANs: Discriminative Representation Learning for Action Recognition Using Partial-Modalities
Chenqiang Gao, Luyu Yang, Yue Zhao 0012, Wangmeng Zuo, Deyu Meng |
ECCV (6) | 2 |
| 2018 | Infrared and Visible Image Registration Using Transformer Adversarial NetworkabstractIn this paper we address the task of infrared and visible image registration in complex scenes. Due to the difference of infrared and visible images, it is neither easy to reliably find features nor suitable for directly training in deep learning architecture. Thus, we propose a two-stage adversarial network, which first conducts a multi-spectral image transfer to obtain a mapped image. And then the proposed network incorporate a transformer module into the conditional adversarial network architecture to get the refined warped image. Our method can back propagate the multi-spectral registration loss and achieve end-to-end training. Experiments on our multi -spectral dataset demonstrate that this approach is effective and robust, which outperforms other state-of-the-art methods. Chenqiang Gao, Yue Zhao 0012, Tiecheng Song |
ICIP | 2 |
| 2018 | Multi-Scale Cross-Band Encoding of Sectored Local Binary Pattern for Robust Texture ClassificationabstractThe original Local Binary Pattern (LBP) has limited discriminative power and is sensitive to noise. In view of this., this paper proposes a novel image descriptor called Multi-Scale Cross-Band Encoding of Sectored Local Binary Pattern (MCE-SLBP) for robust texture classification. First., the pyramid decomposition is explored to obtain multi-scale low-frequency and high-frequency (difference) images. To encode more discriminative features., these high-frequency images are further decomposed into positive and negative high-frequency images via the polarity splitting. Then., a robust Sectored Local Binary Pattern (SLBP) is proposed to compute texture feature codes on the decomposed images via cross-band joint coding. Finally., a multi-scale histogram representation is obtained by concatenating histograms of texture codes computed at all decomposition levels. Experiments on three benchmark texture databases (i.e.., Outex., Brodatz and CUReT) demonstrate that the proposed method achieves the state-of-the-art classification accuracies both under noise-free conditions and in the presence of different levels of Gaussian noise. Tiecheng Song, Lin Luo 0007, Liangliang Xin, Chenqiang Gao |
ICPR | 4 |
| 2018 | Completed Grayscale-Inversion and Rotation Invariant Local Binary Pattern for Texture ClassificationabstractLocal binary pattern (LBP) and its variants (e.g., LTP and CLBP) are powerful descriptors for texture analysis. However, most of these LBP-based methods are sensitive to inverse grayscale changes. To overcome this problem, we present a novel texture descriptor named Completed Grayscale-Inversion and Rotation Invariant Local Binary Pattern (CGRI-LBP). CGRI-LBP is based on the framework of CLBP which jointly encodes three components (i.e., the signs and magnitudes of local differences as well as central pixels) but with two significant improvements: 1) the sign information of local differences is encoded by a rotation-invariant complementary coding scheme, and 2) the intensity information of central pixels is encoded via a dominant intensity order measure. Extensive experiments on three texture databases (Outex, CUReT and KTH-TIPS) demonstrate that the proposed descriptor achieves the state-of-the-art classification performance in the presence of linear and even nonlinear grayscale-inversion changes. Tiecheng Song, Liangliang Xin, Lin Luo 0007, Chenqiang Gao |
ICPR | 4 |
| 2018 | Semantic feature based multi-spectral saliency detection
Chenqiang Gao, Jie Jian, Jiang Liu 0011 |
Multim. Tools Appl. | 2 |
| 2018 | Action detection based on tracklets with the two-stream CNN
Minwen Zhang, Chenqiang Gao, Jiayao Zhang 0003 |
Multim. Tools Appl. | 2 |
| 2018 | Infrared small-dim target detection based on Markov random field guided noise modeling
Chenqiang Gao, Yongxing Xiao, Qian Zhao 0002, Deyu Meng |
Pattern Recognit. | 1 |
| 2018 | Grayscale-Inversion and Rotation Invariant Texture Description Using Sorted Local Gradient PatternabstractThis letter introduces a novel grayscale-inversion and rotation invariant descriptor, called sorted local gradient pattern, for texture classification. First, we propose two complementary local gradient patterns (LGP), the center-to-ring LGP (LGP_CR) and the ring-to-ring LGP (LGP_RR), to encode rich gradient information present in a local neighborhood. Then, we propose to enhance LGP by encoding pixels’ intensity information. This is achieved by sorting image pixels into two categories via a dominant intensity order measure, followed by extracting LGP features over the categorized pixels. As a result, local gradient information and global intensity order information are both encoded into our descriptor in a way that is robust to grayscale-inversion and rotation changes. Experiments on three texture databases demonstrate that the proposed descriptor achieves state-of-the-art classification results in the presence of linear and even nonlinear grayscale-inversion changes. Tiecheng Song, Liangliang Xin, Chenqiang Gao |
IEEE Signal Process. Lett. | 3 |
| 2017 | Rewind to track: Parallelized apprenticeship learning with backward trackletsabstractData association, which could be categorized into offline approaches and the online counterparts, is a crucial part of a multi-object tracker in the tracking-by-detection framework. On the one hand, classical offline data association methods exploit all the video data and have high computation cost, which makes them unscalable to long-term offline video data. On the other hand, online approaches have much lower computation cost, but they suffer from ID-switches and tracklet drifting problem when directly applied to offline data as they are only aware of “past” observations. In this paper, we propose a mixed style tracker, which is not only as efficient as the online tracker but also aware of “future” observations in offline setting. We start from a Markov Decision Process (MDP) online tracker and design a parallelized apprenticeship learning algorithm to learn both the reward function and transition policy in MDP. By proposing a rewind to track strategy to generate backward tracklets, future detections in offline data are efficiently utilized to obtain a more stable similarity measurement for association. Experiment results show that our approach achieves the state-of-the-art performance on challenging datasets. Jiang Liu 0011, Jia Chen 0001, De Cheng, Chenqiang Gao, Alex Hauptmann 0001 |
ICME | 4 |
| 2017 | Semi-supervised manifold-embedded hashing with joint feature representation and classifier learning
Tiecheng Song, Jianfei Cai 0001, Chenqiang Gao, Fanman Meng, Qingbo Wu 0001 |
Pattern Recognit. | 4 |
| 2017 | A novel learning-based frame pooling method for event detection
Chenqiang Gao, Jiang Liu 0011, Deyu Meng |
Signal Process. | 2 |
| 2016 | Two-Stream Contextualized CNN for Fine-Grained Image ClassificationabstractHuman's cognition system prompts that context information provides potentially powerful clue while recognizing objects. However, for fine-grained image classification, the contribution of context may vary over different images, and sometimes the context even confuses the classification result. To alleviate this problem, in our work, we develop a novel approach, two-stream contextualized Convolutional Neural Network, which provides a simple but efficient context-content joint classification model under deep learning framework. The network merely requires the raw image and a coarse segmentation as input to extract both content and context features without need of human interaction. Moreover, our network adopts a weighted fusion scheme to combine the content and the context classifiers, while a subnetwork is introduced to adaptively determine the weight for each image. According to our experiments on public datasets, our approach achieves considerable high recognition accuracy without any tedious human's involvements, as compared with the state-of-the-art approaches. Jiang Liu 0011, Chenqiang Gao, Deyu Meng, Wangmeng Zuo |
AAAI | 2 |
| 2016 | InfAR dataset: Infrared action recognition at different times
Chenqiang Gao, Yinhe Du, Jiang Liu 0011, Jing Lv, Luyu Yang, Deyu Meng, Alex Hauptmann 0001 |
Neurocomputing | 1 |
| 2016 | People counting based on head detection combining Adaboost and CNN in crowded surveillance environment
Chenqiang Gao, Jiang Liu 0011 |
Neurocomputing | 1 |
| 2016 | Unified discriminating feature analysis for visual category recognition
Wenhe Liu, Chenqiang Gao, Xiaojun Chang |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | People-flow counting in complex environments by combining depth and color information
Chenqiang Gao, Jing Lv |
Multim. Tools Appl. | 1 |
| 2016 | Image Classification by Cross-Media Active Learning With Privileged InformationabstractIn this paper, we propose a novel cross-media active learning algorithm to reduce the effort on labeling images for training. The Internet images are often associated with rich textual descriptions. Even though such textual information is not available in test images, it is still useful for learning robust classifiers. In light of this, we apply the recently proposed supervised learning paradigm, learning using privileged information, to the active learning task. Specifically, we train classifiers on both visual features and privileged information, and measure the uncertainty of unlabeled data by exploiting the learned classifiers and slacking function. Then, we propose to select unlabeled samples by jointly measuring the cross-media uncertainty and the visual diversity. Our method automatically learns the optimal tradeoff parameter between the two measurements, which in turn makes our algorithms particularly suitable for real-world applications. Extensive experiments demonstrate the effectiveness of our approach. Yan Yan 0006, Feiping Nie 0001, Wen Li 0001, Chenqiang Gao, Yi Yang 0001, Dong Xu 0001 |
IEEE Trans. Multim. | 4 |
| 2016 | From constrained to unconstrained datasets: an evaluation of local action descriptors and fusion strategies for interaction recognition
Chenqiang Gao, Luyu Yang, Yinhe Du, Zeming Feng, Jiang Liu 0011 |
World Wide Web | 1 |
| 2015 | Robust low-rank tensor factorization by cyclic weighted median
Deyu Meng, Biao Zhang 0005, Zongben Xu, Lei Zhang 0006, Chenqiang Gao |
Sci. China Inf. Sci. | 5 |
| 2015 | A block coordinate descent approach for sparse principal component analysis
Qian Zhao 0002, Deyu Meng, Zongben Xu, Chenqiang Gao |
Neurocomputing | 4 |
| 2014 | A Novel Group-Sparsity-Optimization-Based Feature Selection Model for Complex Interaction Recognition
Luyu Yang, Chenqiang Gao, Deyu Meng, Lu Jiang 0004 |
ACCV (5) | 2 |
| 2014 | Decomposable Nonlocal Tensor Dictionary Learning for Multispectral Image DenoisingabstractAs compared to the conventional RGB or gray-scale images, multispectral images (MSI) can deliver more faithful representation for real scenes, and enhance the performance of many computer vision tasks. In practice, however, an MSI is always corrupted by various noises. In this paper we propose an effective MSI denoising approach by combinatorially considering two intrinsic characteristics underlying an MSI: the nonlocal similarity over space and the global correlation across spectrum. In specific, by explicitly considering spatial self-similarity of an MSI we construct a nonlocal tensor dictionary learning model with a group-block-sparsity constraint, which makes similar full-band patches (FBP) share the same atoms from the spatial and spectral dictionaries. Furthermore, through exploiting spectral correlation of an MSI and assuming over-redundancy of dictionaries, the constrained nonlocal MSI dictionary learning model can be decomposed into a series of unconstrained low-rank tensor approximation problems, which can be readily solved by off-the-shelf higher order statistics. Experimental results show that our method outperforms all state-of-the-art MSI denoising methods under comprehensive quantitative performance measures. Deyu Meng, Zongben Xu, Chenqiang Gao, Yi Yang 0001, Biao Zhang 0005 |
CVPR | 4 |
| 2014 | Interactive Surveillance Event Detection through Mid-level Discriminative RepresentationabstractEvent detection from real surveillance videos with complicated background environment is always a very hard task. Different from the traditional retrospective and interactive systems designed on this task, which are mainly executed on video fragments located within the event-occurrence time, in this paper we propose a new interactive system constructed on the mid-level discriminative representations (patches/shots) which are closely related to the event (might occur beyond the event-occurrence period) and are easier to be detected than video fragments. By virtue of such easily-distinguished mid-level patterns, our framework realizes an effective labor division between computers and human participants. The task of computers is to train classifiers on a bunch of mid-level discriminative representations, and to sort all the possible mid-level representations in the evaluation sets based on the classifier scores. The task of human participants is then to readily search the events based on the clues offered by these sorted mid-level representations. For computers, such mid-level representations, with more concise and consistent patterns, can be more accurately detected than video fragments utilized in the conventional framework, and on the other hand, a human participant can always much more easily search the events of interest implicated by these location-anchored mid-level representations than conventional video fragments containing entire scenes. Both of these two properties facilitate the availability of our framework in real surveillance event detection applications. Chenqiang Gao, Deyu Meng, Yi Yang 0001, Yang Cai 0002, Haoquan Shen, Gaowen Liu, Alex Hauptmann 0001 |
ICMR | 1 |
| 2014 | The Mystery of Faces: Investigating Face Contribution for Multimedia Event DetectionabstractMultimedia event detection (MED) is a retrieval task with the goal of finding videos of a particular event in a large scale internet video archive, given example videos and text descriptions. Nowadays, different multimodal fusion schemes of low-level and high-level features are extensively investigated and evaluated for MED. For most of events in MED, people are usually the central subjects in videos. The face of a person can be considered as the most important factor which brings a lot of information describing the video events. However, face information has not been systematically investigated in the previous research for MED. In this paper, we investigate the possibility of using the high-level face information to assist multimedia event detection. Moreover, since the labeled data in TRECVID MED dataset are limited, we propose a semi-supervised kernel ridge regression which works well in practice to explore the useful information from unlabeled data to assist the event detection. Extensive experimental results on TRECVID MED dataset show that our proposed method outperforms the state-of-the-art methods by up to 4%. Gaowen Liu, Yan Yan 0002, Chenqiang Gao, Alex Hauptmann 0001, Nicu Sebe |
ICMR | 3 |
| 2014 | GLocal tells you more: Coupling GLocal structural for feature selection with sparsity for image and video classification
Yan Yan 0002, Haoquan Shen, Gaowen Liu, Zhigang Ma, Chenqiang Gao, Nicu Sebe |
Comput. Vis. Image Underst. | 5 |
| 2013 | Infrared Patch-Image Model for Small Target Detection in a Single ImageabstractThe robust detection of small targets is one of the key techniques in infrared search and tracking applications. A novel small target detection method in a single infrared image is proposed in this paper. Initially, the traditional infrared image model is generalized to a new infrared patch-image model using local patch construction. Then, because of the non-local self-correlation property of the infrared background image, based on the new model small target detection is formulated as an optimization problem of recovering low-rank and sparse matrices, which is effectively solved using stable principle component pursuit. Finally, a simple adaptive segmentation method is used to segment the target image and the segmentation result can be refined by post-processing. Extensive synthetic and real data experiments show that under different clutter backgrounds the proposed method not only works more stably for different target sizes and signal-to-clutter ratio values, but also has better detection performance compared with conventional baseline methods. Chenqiang Gao, Deyu Meng, Yi Yang 0001, Yongtao Wang, Xiaofang Zhou 0001, Alex Hauptmann 0001 |
IEEE Trans. Image Process. | 1 |
| 2011 | Large Disparity Motion Layer Extraction via Topological ClusteringabstractIn this paper, we present a robust and efficient approach to extract motion layers from a pair of images with large disparity motion. First, motion models are established as: 1) initial SIFT matches are obtained and grouped into a set of clusters using our developed topological clustering algorithm; 2) for each cluster with no less than three matches, an affine transformation is estimated with least-square solution as tentative motion model; and 3) the tentative motion models are refined and the invalid models are pruned. Then, with the obtained motion models, a graph cuts based layer assignment algorithm is employed to segment the scene into several motion layers. Experimental results demonstrate that our method can successfully segment scenes containing objects with large interframe motion or even with significant interframe scale and pose changes. Furthermore, compared with the previous method invented by Wills and its modified version, our method is much faster and more robust. Yongtao Wang, Junbin Gong, Dazhi Zhang, Chenqiang Gao, Jinwen Tian, Huanqiang Zeng |
IEEE Trans. Image Process. | 4 |