EDBT 2026 Demo / reviewers in the wild / expert
Yue Zhao 0012
dblp:48/76-12
· DBLP profile ↗
29ranked-venue papers
4as first author
22since 2021 · last 2026
0000-0003-0342-2797ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ground-to-Aerial Scene Adaptation: Unsupervised drone video action recognition via domain adaptation
Feng Yang 0015, Zhijia Li, Fulin Luo, Anyong Qin, Tiecheng Song, Yue Zhao 0012, Chenqiang Gao |
Eng. Appl. Artif. Intell. | 7 |
| 2026 | TumorAL: Evidence-aware active learning for 3D tumor segmentation
Hongyi Wang 0006, Jiaxu Leng, Yue Zhao 0012, Weikai Li 0003, Weisheng Li 0001, Xinbo Gao 0001 |
Neurocomputing | 4 |
| 2026 | ToothAxis: Generalizable Tooth Axis Estimation Network From CBCT or IOS ModelsabstractTooth axes, indicating the orientation of teeth, are crucial in orthodontics and dental implants. The precise and automated estimation of tooth axes in 3D dental models is of significant importance. In clinical settings, Cone-beam computed tomography (CBCT) images and intraoral scanning (IOS) models are the two primary forms of digital data, providing 3D volumetric and surface information of the oral cavity, respectively. However, the detection of tooth axes remains largely manual annotation due to the complexities associated with geometric definitions and the variations among different tooth types and individuals. In this paper, we propose a novel two-stage network, named ToothAxis, for tooth axis estimation using either CBCT or IOS models. Given that IOS models only capture the tooth crown surface and lack information about the tooth roots, we initially employ an implicit-function tooth completion module for 3D tooth completion in the first stage. Subsequently, with the 3D tooth models segmented from CBCT images or completed from IOS models, a point-wise offset-based module is proposed in the second stage to accurately estimate the tooth axes. This design aims to encode tooth orientation into a dense representation, which is better suited for sparse information regression tasks, such as tooth axis estimation. Additionally, we incorporate a class-specific feature attention module to integrate global context representation, thereby enhancing robustness in managing diverse tooth shapes. We evaluated ToothAxis on a dataset obtained from real-world dental clinics, comprising 529 tooth models with corresponding CBCT images and paired IOS models. Finally, the ToothAxis achieves angle errors of LA ($2.921^{\circ }$), PSA ($4.801^{\circ }$), and LSA ($5.074^{\circ }$) on tooth models extracted from CBCT images, and LA ($5.326^{\circ }$), PSA ($6.360^{\circ }$), and LSA ($6.520^{\circ }$) on partial crowns extracted from IOS models. Extensive evaluations, ablation studies, and comparative analyses demonstrate that our method achieves accurate tooth axis estimations and surpasses state-of-the-art approaches. Qingyao Luo, Zhiming Cui 0001, Yue Zhao 0012 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Aerial video classification with Window Semantic Enhanced Video Transformers
Feng Yang 0015, Botong Zhou, Xuehua Guan, Anyong Qin, Tiecheng Song, Yue Zhao 0012, Chenqiang Gao |
Expert Syst. Appl. | 7 |
| 2025 | Global-local prompts guided image-text embedding, alignment and aggregation for multi-label zero-shot learning
Tiecheng Song, Feng Yang 0015, Anyong Qin, Yue Zhao 0012, Chenqiang Gao |
J. Vis. Commun. Image Represent. | 5 |
| 2025 | Occlusion-aware multi-person pose estimation with keypoint grouping and dual-prompt guidance in crowded scenes
Tiecheng Song, Anyong Qin, Yue Zhao 0012, Feng Yang 0015, Chenqiang Gao |
J. Vis. Commun. Image Represent. | 5 |
| 2025 | Asymmetric Adaptive Heterogeneous Network for Multi-Modality Medical Image SegmentationabstractExisting studies of multi-modality medical image segmentation tend to aggregate all modalities without discrimination and employ multiple symmetric encoders or decoders for feature extraction and fusion. They often overlook the different contributions to visual representation and intelligent decisions among multi-modality images. Motivated by this discovery, this paper proposes an asymmetric adaptive heterogeneous network for multi-modality image feature extraction with modality discrimination and adaptive fusion. For feature extraction, it uses a heterogeneous two-stream asymmetric feature-bridging network to extract complementary features from auxiliary multi-modality and leading single-modality images, respectively. For feature adaptive fusion, the proposed Transformer-CNN Feature Alignment and Fusion (T-CFAF) module enhances the leading single-modality information, and the Cross-Modality Heterogeneous Graph Fusion (CMHGF) module further fuses multi-modality features at a high-level semantic layer adaptively. Comparative evaluation with ten segmentation models on six datasets demonstrates significant efficiency gains as well as highly competitive segmentation accuracy. (Our code is publicly available at https://github.com/joker-527/AAHN). Shenhai Zheng, Chaohui Yang, Weisheng Li 0001, Xinbo Gao 0001, Yue Zhao 0012 |
IEEE Trans. Medical Imaging | 7 |
| 2025 | A Novel Hierarchical Cross-Stream Aggregation Neural Network for Semantic Segmentation of 3-D Dental Surface ModelsabstractAccurate teeth delineation on 3-D dental models is essential for individualized orthodontic treatment planning. Pioneering works like PointNet suggest a promising direction to conduct efficient and accurate 3-D dental model analyses in end-to-end learnable fashions. Recent studies further imply that multistream architectures to concurrently learn geometric representations from different inputs/views (e.g., coordinates and normals) are beneficial for segmenting teeth with varying conditions. However, such multistream networks typically adopt simple late-fusion strategies to combine features captured from raw inputs that encode complementary but fundamentally different geometric information, potentially hampering their accuracy in end-to-end semantic segmentation. This article presents a hierarchical cross-stream aggregation (HiCA) network to learn more discriminative point/cell-wise representations from multiview inputs for fine-grained 3-D semantic segmentation. Specifically, based upon our multistream backbone with input-tailored feature extractors, we first design a contextual cross-steam aggregation (CA) module conditioned on interstream consistency to boost each view's contextual representation learning jointly. Then, before the late fusion of different streams' outputs for segmentation, we further deploy a discriminative cross-stream aggregation (DA) module to concurrently update all views' discriminative representation learning by leveraging a specific graph attention strategy induced by multiview prototype learning. On both public and in-house datasets of real-patient dental models, our method significantly outperformed state-of-the-art (SOTA) deep learning methods for teeth semantic segmentation. In addition, extended experimental results suggest the applicability of HiCA to other general 3-D shape segmentation tasks. The code is available at https://github.com/ladderlab-xjtu/HiCA. Kehan Li 0009, Jihua Zhu, Zhiming Cui 0001, Xinning Chen, Yang Liu 0157, Fan Wang 0038, Yue Zhao 0012 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | SMTLNet: Domain Prior-Inspired Tooth Segmentation Based on Self-Supervised Manifold Transfer LearningabstractAccurate identification and delineation of teeth in cone-beam computed tomography (CBCT) images are crucial in the advancement of digital dentistry technology. Teeth exhibit high interclass similarity and often have fuzzy boundaries. In addition, it is difficult to obtain teeth samples due to the time-consuming annotation process. However, existing methods typically fail to incorporate this domain-specific prior information under limited labeled samples, which limits the improvement of segmentation performance. Based on the intrinsic characteristics of the tooth CBCT images, a self-supervised manifold transfer learning network (SMTLNet) is proposed to improve segmentation accuracy. Initially, an object-oriented self-supervised pretraining approach is designed to fully explore valuable image representations from unannotated images, and this helps reduce dependence on labeled samples. Furthermore, a manifold optimization strategy is employed to regularize the segmentation model to separate interclass samples while compacting intraclass neighbors. Finally, to address the issue of blurred tooth boundaries, a multiscale boundary constraint module is developed to extract multiscale boundary-aware features, and more discriminative tooth descriptions can be acquired in this way. The proposed SMTLNet method is evaluated on clinical datasets containing diverse challenging cases (e.g., impacted wisdom teeth, crowded dentition), and it achieves state-of-the-art performance with dice similarity coefficients (DSCs) of 91.8%/89.08% and Jaccard similarities (JSs) of 86.71%/82.87% under full (100%) and limited (20%) training data regimes, respectively. The method maintains anatomical precision with Hausdorff distances (HDs) of 1.41 mm (high-resource) and 2.35 mm (low-resource), demonstrating strong clinical applicability in digital dentistry workflows. Yue Zhao 0012, Pengyu Dai, Hong Huang 0002, Yang Liu 0157 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Deep Updated Subspace Networks for Few-Shot Remote Sensing Scene ClassificationabstractDue to the difficulty of manually labeling remote sensing scene images and the demand for the ability to recognize new scene classes, few-shot remote sensing scene classification (FSRSSC) has attracted more and more attention. At present, metric-based FSRSSC methods have made promising progress, especially the prototypical networks-based methods. However, due to the complexity of the background of remote sensing scene images, the prototype classifier, which takes the average features of support samples as the metric benchmark, retains the features of category-irrelevant objects and other background information in the image. This leads to a bad classification result. Therefore, in this work, we propose a FSRSSC method based on the deep updated subspace network (DUSN), which uses class subspace as a metric benchmark to represent the commonality of a category and can effectively mitigate the negative impact of irrelevant objects on the classifier. In addition, for the higher inter-class similarity and larger intra-class variance of remote sensing scene images, we further propose an inter-class constraint and an intra-class constraint to mitigate the classification confusion. We leverage the inter-class constraint to make the images of different classes as far apart as possible, and the intra-class constraint to keep the images of the same class clustered as closely together as possible. Experimental results on three public benchmark datasets demonstrate that our method performs better than the state-of-the-art methods for FSRSSC. Anyong Qin, Fuyang Chen, Lingyun Tang, Feng Yang 0015, Yue Zhao 0012, Chenqiang Gao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Change-Aware Cascaded Dual-Decoder Network for Remote Sensing Image Change DetectionabstractChange detection aims to detect changes of objects or scenes in remote sensing images, which is critical for observing the Earth’s surface. However, due to the insufficient correlation and aggregation of bitemporal features, the existing deep learning methods are still impacted by varied imaging conditions and complicated boundaries of ground objects in high-resolution remote sensing images. To tackle these challenges, we propose a change-aware cascaded dual-decoder network (CACD2Net), which integrates bitemporal features at different levels to facilitate learning change maps from coarse to fine, thus empowering the network to effectively identify changes and refine pixelwise boundaries in a progressive manner. Within the cascaded dual-decoder architecture, the change location decoder utilizes high-level features to generate a coarse change map, which approximates changes’ localization, while the mask refinement decoder further leverages low-level features to create a texture-aware map that captures more texture and structural information about the change regions. By using the coarse change map as guidance and directing the texture-aware map to focus on the details of changes, the boundaries can be gradually refined, ultimately resulting in an accurate change detection mask. We test our model on the season-varying change detection (SVCD) dataset and the Sun Yat-sen University change detection (SYSU-CD) dataset, and the experimental results show that our model surpasses other state-of-the-art change detection methods. Our codes will be available athttps://github.com/Moonquakes0/CACD2Net. Feng Yang 0015, Yifeng Yuan, Anyong Qin, Yue Zhao 0012, Tiecheng Song, Chenqiang Gao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | DTR-Net: Dual-Space 3D Tooth Model Reconstruction From Panoramic X-Ray ImagesabstractIn digital dentistry, cone-beam computed tomography (CBCT) can provide complete 3D tooth models, yet suffers from a long concern of requiring excessive radiation dose and higher expense. Therefore, 3D tooth model reconstruction from 2D panoramic X-ray image is more cost-effective, and has attracted great interest in clinical applications. In this paper, we propose a novel dual-space framework, namely DTR-Net, to reconstruct 3D tooth model from 2D panoramic X-ray images in both image and geometric spaces. Specifically, in the image space, we apply a 2D-to-3D generative model to recover intensities of CBCT image, guided by a task-oriented tooth segmentation network in a collaborative training manner. Meanwhile, in the geometric space, we benefit from an implicit function network in the continuous space, learning using points to capture complicated tooth shapes with geometric properties. Experimental results demonstrate that our proposed DTR-Net achieves state-of-the-art performance both quantitatively and qualitatively in 3D tooth model reconstruction, indicating its potential application in dental practice. Lanzhuju Mei, Yu Fang 0008, Yue Zhao 0012, Xiang Sean Zhou, Zhiming Cui 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 3 |
| 2023 | Distance Constraint-Based Generative Adversarial Networks for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification suffers from two serious problems, one is the limited labeled pixels, and the other is the class imbalance problem. As a result, the number of labeled pixels in many categories is not sufficient to characterize the spectral-spatial information, and train a satisfying deep model. By making full use of the information of unlabeled pixels, semi-supervised methods can provide better classification performance in the case of limited labeled pixels. However, they do not take into account the imbalance in the HSI data. As a method of data enhancement, generative adversarial networks focus on the above two problems and have also been widely used for the task of the HSI classification. In this work, we propose a distance constraints-based generative adversarial networks (DGAN) method for HSI classification to address these two problems. The DGAN employs the convolution autoencoder (AE) to extract the latent features of the HSI samples, and considers the reconstructed samples from the AE as the real samples for the later classifier and discriminator. In addition, the DGAN uses two distance constraints to solve the problems of the few labeled samples and class imbalance, the one latent-data distance constraint enforcing the generator to generate HSI samples for each class (especially the minority class), another discriminator-score distance constraint guiding the generator to synthesize samples that resemble the real HSI samples. Finally, the generated samples are combined classwise with the reconstructed samples and the real HSI samples to learn the parameters of the classifier and discriminator. Experimental results show that our method achieves state-of-the-art performance in terms of overall accuracy (OA) when trained with only 0.5%-4% of data sets from Indian Pines, Pavia University, and Botswana. Specifically, our method demonstrates improvements of 5.48%, 8.79%, and 0.91% on these three datasets, respectively. It reveals the great potential of the DGAN model in generating the HSI samples for each class, which contributes to improving the classification performance of the HSI data. Anyong Qin, Zhuolin Tan, Yongqing Sun, Feng Yang 0015, Yue Zhao 0012, Chenqiang Gao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Multiscale Spatio-Temporal Network for Aerial Video Event RecognitionabstractUnmanned aerial vehicles (UAVs) are widely used in the field of remote sensing because of their advantages of providing real-time and high-resolution videos at a low cost. Compared with generic video understanding, aerial video event recognition is faced with emerging challenges: 1) aerial videos contain richer scene information; 2) the scale variations between different videos are large. To address these issues, we propose a Multiscale Spatio-Temporal Network (MSTN) in this paper. More precisely, the MSTN consists of a Pyramid Spatio-Temporal (PST) module and a Multi-Time Scale Decision (MTSD) module, which learn multi-scale spatio-temporal features together. The two modules can better learn spatio-temporal characteristics and boost the performance by 3.7% compared with the baseline method. In ERA, an aerial event recognition dataset, our method achieves the state-of-the-art results. Feng Yang 0015, Yue Zhao 0012, Anyong Qin, Chenqiang Gao |
IGARSS | 3 |
| 2022 | Curvature-Enhanced Implicit Function Network for High-quality Tooth Model Generation from CBCT Images
Yu Fang 0008, Zhiming Cui 0001, Lei Ma 0006, Lanzhuju Mei, Yue Zhao 0012, Zhihao Jiang 0001, Yiqiang Zhan, Yongsheng Pan, Dinggang Shen |
MICCAI (5) | 6 |
| 2022 | Spectral-Spatial Residual Graph Attention Network for Hyperspectral Image ClassificationabstractHyperspectral images (HSIs) not only possess abundant spectral features but also present a detailed spatial distribution of land cover, and they have significant advantages in the fine classification of ground materials. Recently, using convolutional neural networks (CNNs) to extract spectral–spatial features has become an effective way for HSI classification. However, conventional convolution kernels learn features from fixed regular square regions, and rich spatial information has not been effectively explored. In this letter, an end-to-end model named spectral–spatial residual graph attention network (S2RGANet) is developed for HSI classification, and it has two crucial elements, including spectral residual and graph attention convolution modules. At first, two spectral residual modules are employed to capture discriminant spectral features. Then, graphs are constructed to reveal the relationship between points in local neighborhoods. By graph attention mechanism, local spatial information is adaptively aggregated from neighboring nodes. Experiments on two public HSI datasets demonstrate that the S2RGANet is significantly superior to some state-of-the-art (SOTA) methods with limited training samples. Kejie Xu, Yue Zhao 0012, Chenqiang Gao, Hong Huang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Semantic Graph Attention With Explicit Anatomical Association Modeling for Tooth Segmentation From CBCT ImagesabstractAccurate tooth identification and delineation in dental CBCT images are essential in clinical oral diagnosis and treatment. Teeth are positioned in the alveolar bone in a particular order, featuring similar appearances across adjacent and bilaterally symmetric teeth. However, existing tooth segmentation methods ignored such specific anatomical topology, which hampers the segmentation accuracy. Here we propose a semantic graph-based method to explicitly model the spatial associations between different anatomical targets (i.e., teeth) for their precise delineation in a coarse-to-fine fashion. First, to efficiently control the bilaterally symmetric confusion in segmentation, we employ a lightweight network to roughly separate teeth as four quadrants. Then, designing a semantic graph attention mechanism to explicitly model the anatomical topology of the teeth in each quadrant, based on which voxel-wise discriminative feature embeddings are learned for the accurate delineation of teeth boundaries. Extensive experiments on a clinical dental CBCT dataset demonstrate the superior performance of the proposed method compared with other state-of-the-art approaches. Pengcheng Li 0017, Yang Liu 0157, Zhiming Cui 0001, Feng Yang 0015, Yue Zhao 0012, Chunfeng Lian, Chenqiang Gao |
IEEE Trans. Medical Imaging | 5 |
| 2022 | Two-Stream Graph Convolutional Network for Intra-Oral Scanner Image SegmentationabstractPrecise segmentation of teeth from intra-oral scanner images is an essential task in computer-aided orthodontic surgical planning. The state-of-the-art deep learning-based methods often simply concatenate the raw geometric attributes (i.e., coordinates and normal vectors) of mesh cells to train a single-stream network for automatic intra-oral scanner image segmentation. However, since different raw attributes reveal completely different geometric information, the naive concatenation of different raw attributes at the (low-level) input stage may bring unnecessary confusion in describing and differentiating between mesh cells, thus hampering the learning of high-level geometric representations for the segmentation task. To address this issue, we design a two-stream graph convolutional network (i.e., TSGCN), which can effectively handle inter-view confusion between different raw attributes to more effectively fuse their complementary information and learn discriminative multi-view geometric representations. Specifically, our TSGCN adopts two input-specific graph-learning streams to extract complementary high-level geometric representations from coordinates and normal vectors, respectively. Then, these single-view representations are further fused by a self-attention module to adaptively balance the contributions of different views in learning more discriminative multi-view representations for accurate and fully automatic tooth segmentation. We have evaluated our TSGCN on a real-patient dataset of dental (mesh) models acquired by 3D intraoral scanners. Experimental results show that our TSGCN significantly outperforms state-of-the-art methods in 3D tooth (surface) segmentation. Yue Zhao 0012, Yang Liu 0157, Deyu Meng, Zhiming Cui 0001, Chenqiang Gao, Xinbo Gao 0001, Chunfeng Lian, Dinggang Shen |
IEEE Trans. Medical Imaging | 1 |
| 2021 | TSGCNet: Discriminative Geometric Feature Learning With Two-Stream Graph Convolutional Network for 3D Dental Model SegmentationabstractThe ability to segment teeth precisely from digitized 3D dental models is an essential task in computer-aided orthodontic surgical planning. To date, deep learning based methods have been popularly used to handle this task. State-of-the-art methods directly concatenate the raw attributes of 3D inputs, namely coordinates and normal vectors of mesh cells, to train a single-stream network for fully-automated tooth segmentation. This, however, has the drawback of ignoring the different geometric meanings provided by those raw attributes. This issue might possibly confuse the network in learning discriminative geometric features and result in many isolated false predictions on the dental model. Against this issue, we propose a two-stream graph convolutional network (TSGCNet) to learn multi-view geometric information from different geometric attributes. Our TSGCNet adopts two graph-learning streams, designed in an input-aware fashion, to extract more discriminative high-level geometric representations from coordinates and normal vectors, respectively. These feature representations learned from the designed two different streams are further fused to integrate the multi-view complementary information for the cell-wise dense prediction task. We evaluate our proposed TSGCNet on a real-patient dataset of dental models acquired by 3D intraoral scanners, and experimental results demonstrate that our method significantly outperforms state-of-the-art methods for 3D shape segmentation. Yue Zhao 0012, Deyu Meng, Zhiming Cui 0001, Chenqiang Gao, Xinbo Gao 0001, Chunfeng Lian, Dinggang Shen |
CVPR | 2 |
| 2021 | MSLPNet: multi-scale location perception network for dental panoramic X-ray image segmentation
Qiaoyi Chen, Yue Zhao 0012, Yang Liu 0157, Yongqing Sun, Chongshi Yang, Pengcheng Li 0017, Chenqiang Gao |
Neural Comput. Appl. | 2 |
| 2021 | 3D Dental model segmentation with graph attentional convolution network
Yue Zhao 0012, Chongshi Yang, Yingyun Tan, Yang Liu 0157, Pengcheng Li 0017, Chenqiang Gao |
Pattern Recognit. Lett. | 1 |
| 2021 | Infrared and Visible Cross-Modal Image Retrieval Through Shared FeaturesabstractImage retrieval is one of the key techniques of computer vision, and has been studied for a long time. Nevertheless, little attention is paid to infrared and visible cross-modal retrieval which can be widely used in various applications, e.g., infrared and visible surveillance systems. In this paper, we propose a shared features based infrared-visible cross-modal image retrieval method. The similar visual features are extracted from infrared and visible images as the shared features, and the Euclidean distance is used to measure the similarity between these features. The core of the proposed method comes from three aspects: 1) Feature separation network can separate image features into shared features and exclusive features; 2) Maximum Mean Discrepancy (MMD) loss is employed to constrain the distribution of shared features, which can reduce the retrieval error caused by different imaging angles and similarity of infrared images. 3) The cross-layer fusion encoder compensates for the context loss in the convolution of infrared images. Experimental results on the Infrared-Visible dataset demonstrate the proposed method is effective and outperforms the state-of-the-art approaches. Fangcen Liu, Chenqiang Gao, Yongqing Sun, Yue Zhao 0012, Feng Yang 0015, Anyong Qin, Deyu Meng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Adaptive Fusion and Mask Refinement Instance Segmentation Network for High Resolution Remote Sensing ImagesabstractInstance segmentation of remote sensing images (RSIs) is an active yet challenging task because of the huge scale variation and arbitrary complex shapes of objects. To address these issues, we propose an adaptive fusion and mask refinement (AFMR) instance segmentation network for RSIs in this paper. More precisely, AFMR consists of an adaptive fusion module to learn multi-scale complementary spatial features in an unsupervised manner, and a content-aware module for segmentation mask refinement. These two modules enable a better feature learning of convolutional neural network and boost the performance by 1.5% compared with the baseline method. In iSAID, a large-scale dataset for RSIs instance segmentation, our AFMR framework achieves the state-of-the-art accuracy, which verifies the superiority of the proposed method. Jie Ran, Feng Yang 0015, Chenqiang Gao, Yue Zhao 0012, Anyong Qin |
IGARSS | 4 |
| 2020 | TSASNet: Tooth segmentation on dental panoramic X-ray images by Two-Stage Attention Segmentation Network
Yue Zhao 0012, Pengcheng Li 0017, Chenqiang Gao, Yang Liu 0157, Qiaoyi Chen, Feng Yang 0015, Deyu Meng |
Knowl. Based Syst. | 1 |
| 2020 | Spatio-temporal fall event detection in complex scenes using attention guided LSTM
Chenqiang Gao, Yue Zhao 0012, Tiecheng Song |
Pattern Recognit. Lett. | 4 |
| 2020 | Polycrystalline silicon wafer defect segmentation based on deep convolutional neural networks
Chenqiang Gao, Yue Zhao 0012, Shisha Liao, Xindou Li |
Pattern Recognit. Lett. | 3 |
| 2019 | Pose detection in complex classroom environment based on improved Faster R-CNNabstractPose detection of small targets in poor imaging conditions like heavy occlusion and low resolution is still an open and challenging task in computer vision. For instance, detection of students' poses in classrooms that are even indistinguishable to human eyes remains a rather difficult task. Motivated by the success of convolutional feature merging and locality preserving, the authors propose a pose detection framework combining merged region of interest (ROI) pooling and locality preserving learning. Unlike usual object detection algorithms which use general top‐level convolutional features as inputs, their method uses a merged ROI pooling structure to merge semantic feature and high‐resolution feature from the last two levels of convolutional feature maps, so that this merged feature is made more expressive than the single‐level feature. In addition, the locality feature‐preserving learning is used in the last fully‐connected layer. Through locality preserving learning, features belonging to the same class would be forced to be closer in the feature space, which enables the model with stronger classification ability. Experimental results show that the proposed method outperforms the state‐of‐the‐art methods. Chenqiang Gao, Xu Chen 0053, Yue Zhao 0012 |
IET Image Process. | 4 |
| 2018 | PM-GANs: Discriminative Representation Learning for Action Recognition Using Partial-Modalities
Chenqiang Gao, Luyu Yang, Yue Zhao 0012, Wangmeng Zuo, Deyu Meng |
ECCV (6) | 4 |
| 2018 | Infrared and Visible Image Registration Using Transformer Adversarial NetworkabstractIn this paper we address the task of infrared and visible image registration in complex scenes. Due to the difference of infrared and visible images, it is neither easy to reliably find features nor suitable for directly training in deep learning architecture. Thus, we propose a two-stage adversarial network, which first conducts a multi-spectral image transfer to obtain a mapped image. And then the proposed network incorporate a transformer module into the conditional adversarial network architecture to get the refined warped image. Our method can back propagate the multi-spectral registration loss and achieve end-to-end training. Experiments on our multi -spectral dataset demonstrate that this approach is effective and robust, which outperforms other state-of-the-art methods. Chenqiang Gao, Yue Zhao 0012, Tiecheng Song |
ICIP | 3 |