EDBT 2026 Demo / reviewers in the wild / expert
Feng Yang 0015
dblp:22/4613-15
· DBLP profile ↗
25ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0003-0413-8640ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ground-to-Aerial Scene Adaptation: Unsupervised drone video action recognition via domain adaptation
Feng Yang 0015, Zhijia Li, Fulin Luo, Anyong Qin, Tiecheng Song, Yue Zhao 0012, Chenqiang Gao |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | Progressive spectral-frequency-spatial guidance network for hyperspectral image dehazing
Qianru Liu, Tiecheng Song, Kaizhao Zhang, Anyong Qin, Feng Yang 0015, Chenqiang Gao |
Expert Syst. Appl. | 6 |
| 2025 | Motion Artifact Removal in Pixel-Frequency Domain via Alternate Masks and Diffusion ModelabstractMotion artifacts present in magnetic resonance imaging (MRI) can seriously interfere with clinical diagnosis. Removing motion artifacts is a straightforward solution and has been extensively studied. However, paired data are still heavily relied on in recent works and the perturbations in k-space (frequency domain) are not well considered, which limits their applications in the clinical field. To address these issues, we propose a novel unsupervised purification method which leverages pixel-frequency information of noisy MRI images to guide a pre-trained diffusion model to recover clean MRI images. Specifically, considering that motion artifacts are mainly concentrated in high-frequency components in k-space, we utilize the low-frequency components as the guide to ensure correct tissue textures. Additionally, given that high-frequency and pixel information are helpful for recovering shape and detail textures, we design alternate complementary masks to simultaneously destroy the artifact structure and exploit useful information. Quantitative experiments are performed on datasets from different tissues and show that our method achieves superior performance on several metrics. Qualitative evaluations with radiologists also show that our method provides better clinical feedback. Dawei Zhou 0004, Lei Hu 0002, Feng Yang 0015, Zaiyi Liu, Nannan Wang 0001, Xinbo Gao 0001 |
AAAI | 5 |
| 2025 | Frequency-prompt guided spectral-spatial transformer for hyperspectral image classification
Tiecheng Song, Longlong Zhang, Anyong Qin, Feng Yang 0015, Chenqiang Gao |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Aerial video classification with Window Semantic Enhanced Video Transformers
Feng Yang 0015, Botong Zhou, Xuehua Guan, Anyong Qin, Tiecheng Song, Yue Zhao 0012, Chenqiang Gao |
Expert Syst. Appl. | 1 |
| 2025 | Global-local prompts guided image-text embedding, alignment and aggregation for multi-label zero-shot learning
Tiecheng Song, Feng Yang 0015, Anyong Qin, Yue Zhao 0012, Chenqiang Gao |
J. Vis. Commun. Image Represent. | 3 |
| 2025 | Occlusion-aware multi-person pose estimation with keypoint grouping and dual-prompt guidance in crowded scenes
Tiecheng Song, Anyong Qin, Yue Zhao 0012, Feng Yang 0015, Chenqiang Gao |
J. Vis. Commun. Image Represent. | 6 |
| 2025 | Dual-Branch Residual Network for Cross-Domain Few-Shot Hyperspectral Image Classification With Refined Prototype
Anyong Qin, Chaoqi Yuan, Feng Yang 0015, Tiecheng Song, Chenqiang Gao |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | ConvFormer-CD: Hybrid CNN-Transformer With Temporal Attention for Detecting Changes in Remote Sensing ImageryabstractRecently, the combination of Transformers and convolutional neural networks (CNNs) has witnessed significant advancements in change detection (CD) tasks. However, it remains unexplored how to interactively integrate long-range dependency and local information to enhance the model’s global-local context awareness for effectively mitigating pseudo-changes. In addition, accurate identification and distinction of building changes from complex backgrounds still pose challenges due to the insufficient semantic context modeling across time between bi-temporal images. To address these issues, we propose a hybrid model ConvFormer-CD with parallel convolution and multihead self-attention (MSA). This combination enables better interaction of global and local information, thereby enhancing the adaptability to complex scenarios. Moreover, we introduce a novel module called Temporal Attention to establish cross-temporal semantic relationships between image pairs, effectively highlighting change regions by learning shared and nonshared semantics. This enables our model to accurately detect changed targets even in scenarios characterized by intricate geo-spatial arrangements and distributions. To further refine the differences in bi-temporal images, we propose a difference integration module (DIM) that connects the encoder and the decoder to fuse high-level semantic features across channels. We conduct extensive experiments on four benchmark datasets, including LEVIR-CD, LEVIR-CD+, WHU-CD, and S2Looking-CD, which demonstrates that the proposed ConvFormer-CD outperforms other state-of-the-art (SOTA) methods. Our codes will be available athttps://github.com/taomi-lab/ConvFormer-CD. Feng Yang 0015, Mengtao Li, Wenqiang Shu, Anyong Qin, Tiecheng Song, Chenqiang Gao, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Towards Student Actions in Classroom Scenes: New Dataset and BaselineabstractAnalyzing student actions is an important and challenging task in educational research. Existing efforts have been hampered by the lack of accessible datasets to capture the nuanced action dynamics in classrooms. In this paper, we present a new multi-labelStudent Action Video(SAV) dataset, specifically designed for action detection in classroom settings. The SAV dataset consists of 4,324 carefully trimmed video clips from 758 different classrooms, annotated with 15 distinct student actions. Compared to existing action detection datasets, the SAV dataset stands out by providing a wide range of real classroom scenarios, high-quality video data, and unique challenges, including subtle movement differences, dense object engagement, significant scale differences, varied shooting angles, and visual occlusion. These complexities introduce new opportunities and challenges to advance action detection methods. To benchmark this, we propose a novel baseline method based on a visual transformer, designed to enhance attention to key local details within small and dense object regions. Our method demonstrates excellent performance with a mean Average Precision (mAP) of 67.9% and 27.4% on the SAV and AVA datasets, respectively. This paper not only provides the dataset but also calls for further research into AI-driven educational tools that may transform teaching methodologies and learning outcomes. The code and dataset are released athttps://github.com/Ritatanz/SAV. Zhuolin Tan, Chenqiang Gao, Anyong Qin, Ruixin Chen, Tiecheng Song, Feng Yang 0015, Deyu Meng |
IEEE Trans. Multim. | 6 |
| 2024 | ESMS-Net: Enhancing Semantic-Mask Segmentation Network With Pyramid Atrousformer for Remote Sensing ImageabstractTransformers has gained widespread adoption in remote sensing image (RSI) segmentation. However, RSI has densely overlapping terrain and significant shadow, making it challenging to segment the blended boundaries of terrains that are the hard classes. Currently, most transformer-based methods construct the self-attention with a sliding window, which influences the feature receptive fields to conquer the intersecting and overlapping objects. Additionally, they often rarely focus specifically on the representation of these hard segmentation objects. To overcome these challenges, we propose a novel Enhancing Semantic Mask Segmentation Network (ESMS-Net) framework including a local-global joint encoder, an auxiliary enhanced encoder, and a multiscale dense decoder. In the local-global joint encoder, we construct a Pyramid Pooling AtrousFormer (PPAFormer) that performs the self-attention with a pyramid-structured atrous sliding window, which enhances the range of receptive fields and the global representation performance. Meanwhile, we construct the dual-feature fusion module (DFFM) and multilevel feature weighted fusion (MFWF) in the multiscale dense decoder to reduce information loss and facilitate the interaction of deep semantic information. For the auxiliary enhanced encoder, we develop a semantic mask based on the predicted results to maintain the hard segmentation classes, and then use the same structure as the first two stages of the local-global joint encoder to learn the hard regions again. Extensive experiments demonstrate the proposed ESMS-Net can achieve significant improvements for segmentation performance compared with the state-of-the-art methods on the ISPRS-Vaihingen and Potsdam datasets. The code will be available athttps://github.com/Wzysaber/ESMS-Net. Fulin Luo, Tan Guo, Feng Yang 0015, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Deep Updated Subspace Networks for Few-Shot Remote Sensing Scene ClassificationabstractDue to the difficulty of manually labeling remote sensing scene images and the demand for the ability to recognize new scene classes, few-shot remote sensing scene classification (FSRSSC) has attracted more and more attention. At present, metric-based FSRSSC methods have made promising progress, especially the prototypical networks-based methods. However, due to the complexity of the background of remote sensing scene images, the prototype classifier, which takes the average features of support samples as the metric benchmark, retains the features of category-irrelevant objects and other background information in the image. This leads to a bad classification result. Therefore, in this work, we propose a FSRSSC method based on the deep updated subspace network (DUSN), which uses class subspace as a metric benchmark to represent the commonality of a category and can effectively mitigate the negative impact of irrelevant objects on the classifier. In addition, for the higher inter-class similarity and larger intra-class variance of remote sensing scene images, we further propose an inter-class constraint and an intra-class constraint to mitigate the classification confusion. We leverage the inter-class constraint to make the images of different classes as far apart as possible, and the intra-class constraint to keep the images of the same class clustered as closely together as possible. Experimental results on three public benchmark datasets demonstrate that our method performs better than the state-of-the-art methods for FSRSSC. Anyong Qin, Fuyang Chen, Lingyun Tang, Feng Yang 0015, Yue Zhao 0012, Chenqiang Gao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Few-Shot Learning With Prototype Rectification for Cross-Domain Hyperspectral Image ClassificationabstractDeep learning has been extensively applied to hyperspectral image (HSI) classification and has achieved significant success. However, the number of labeled samples available for HSI classification tasks is typically limited in practical applications, which makes the high-accuracy of HSI small-sample classification still a challenging research task. Therefore, metric-based prototypical networks for few-shot learning (FSL) have become increasingly popular. However, the majority of existing FSL methods typically have problems with biased prototypes and domain shifts. To address these issues, this article proposed a prototype rectification network framework for cross-domain few-shot HSI classification. Specifically, to obtain more representative prototypes, we designed a query-guided prototype rectification module, which can rectify the feature distribution of the support set prototype and obtain a more representative prototype for subsequent training tasks. Then, we introduced a prototype-based interclass loss function to alleviate the interclass confusion that may result from prototype rectification. Furthermore, we construct an intermediate domain between the source domain and the target domain to alleviate domain shift, which helps mitigate the difficulties of domain transfer and achieve a more comprehensive domain alignment. The experimental results on four publicly available HSI datasets demonstrate that our proposed method outperforms the existing FSL methods. Anyong Qin, Chaoqi Yuan, Xiaoliu Luo, Feng Yang 0015, Tiecheng Song, Chenqiang Gao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Exploring Hybrid Contrastive Learning and Scene-to-Label Information for Multilabel Remote Sensing Image ClassificationabstractMultilabel remote sensing (RS) image classification aims to predict multiple semantic labels from an RS image. Previous methods [e.g., graph convolution networks (GCNs)] focus on mining the relationships of multiple labels, neglecting that the scene information is closely related to labels. To remedy this deficiency, in this article we propose a novel end-to-end deep neural network for multilabel RS image classification. In the proposed network, we use the GCN as the base model and introduce several new components to improve the classification performance. First, we explore hybrid contrastive learning (CL), including supervised transformation-based CL and unsupervised mix-based CL, to explicitly learn discriminative scene representations. Then, we apply the GCN-based classifier to the learned scene representations to obtain initial label prediction scores. Meanwhile, we pass the scene representations to a softmax layer to predict the probability that each image belongs to each specific scene class and use the scene-to-label information with the law of total probability to calibrate the initial label prediction scores. Finally, we incorporate CL, scene classification, and multilabel classification into a unified learning framework using uncertainty to weigh different losses. Experimental results on two benchmark RS datasets demonstrate the superiority of our proposed network for multilabel image classification. Tiecheng Song, Shufen Bai, Feng Yang 0015, Chenqiang Gao, Haonan Chen 0001, Jun Li 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Change-Aware Cascaded Dual-Decoder Network for Remote Sensing Image Change DetectionabstractChange detection aims to detect changes of objects or scenes in remote sensing images, which is critical for observing the Earth’s surface. However, due to the insufficient correlation and aggregation of bitemporal features, the existing deep learning methods are still impacted by varied imaging conditions and complicated boundaries of ground objects in high-resolution remote sensing images. To tackle these challenges, we propose a change-aware cascaded dual-decoder network (CACD2Net), which integrates bitemporal features at different levels to facilitate learning change maps from coarse to fine, thus empowering the network to effectively identify changes and refine pixelwise boundaries in a progressive manner. Within the cascaded dual-decoder architecture, the change location decoder utilizes high-level features to generate a coarse change map, which approximates changes’ localization, while the mask refinement decoder further leverages low-level features to create a texture-aware map that captures more texture and structural information about the change regions. By using the coarse change map as guidance and directing the texture-aware map to focus on the details of changes, the boundaries can be gradually refined, ultimately resulting in an accurate change detection mask. We test our model on the season-varying change detection (SVCD) dataset and the Sun Yat-sen University change detection (SYSU-CD) dataset, and the experimental results show that our model surpasses other state-of-the-art change detection methods. Our codes will be available athttps://github.com/Moonquakes0/CACD2Net. Feng Yang 0015, Yifeng Yuan, Anyong Qin, Yue Zhao 0012, Tiecheng Song, Chenqiang Gao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Distance Constraint-Based Generative Adversarial Networks for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification suffers from two serious problems, one is the limited labeled pixels, and the other is the class imbalance problem. As a result, the number of labeled pixels in many categories is not sufficient to characterize the spectral-spatial information, and train a satisfying deep model. By making full use of the information of unlabeled pixels, semi-supervised methods can provide better classification performance in the case of limited labeled pixels. However, they do not take into account the imbalance in the HSI data. As a method of data enhancement, generative adversarial networks focus on the above two problems and have also been widely used for the task of the HSI classification. In this work, we propose a distance constraints-based generative adversarial networks (DGAN) method for HSI classification to address these two problems. The DGAN employs the convolution autoencoder (AE) to extract the latent features of the HSI samples, and considers the reconstructed samples from the AE as the real samples for the later classifier and discriminator. In addition, the DGAN uses two distance constraints to solve the problems of the few labeled samples and class imbalance, the one latent-data distance constraint enforcing the generator to generate HSI samples for each class (especially the minority class), another discriminator-score distance constraint guiding the generator to synthesize samples that resemble the real HSI samples. Finally, the generated samples are combined classwise with the reconstructed samples and the real HSI samples to learn the parameters of the classifier and discriminator. Experimental results show that our method achieves state-of-the-art performance in terms of overall accuracy (OA) when trained with only 0.5%-4% of data sets from Indian Pines, Pavia University, and Botswana. Specifically, our method demonstrates improvements of 5.48%, 8.79%, and 0.91% on these three datasets, respectively. It reveals the great potential of the DGAN model in generating the HSI samples for each class, which contributes to improving the classification performance of the HSI data. Anyong Qin, Zhuolin Tan, Yongqing Sun, Feng Yang 0015, Yue Zhao 0012, Chenqiang Gao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Multiscale Spatio-Temporal Network for Aerial Video Event RecognitionabstractUnmanned aerial vehicles (UAVs) are widely used in the field of remote sensing because of their advantages of providing real-time and high-resolution videos at a low cost. Compared with generic video understanding, aerial video event recognition is faced with emerging challenges: 1) aerial videos contain richer scene information; 2) the scale variations between different videos are large. To address these issues, we propose a Multiscale Spatio-Temporal Network (MSTN) in this paper. More precisely, the MSTN consists of a Pyramid Spatio-Temporal (PST) module and a Multi-Time Scale Decision (MTSD) module, which learn multi-scale spatio-temporal features together. The two modules can better learn spatio-temporal characteristics and boost the performance by 3.7% compared with the baseline method. In ERA, an aerial event recognition dataset, our method achieves the state-of-the-art results. Feng Yang 0015, Yue Zhao 0012, Anyong Qin, Chenqiang Gao |
IGARSS | 1 |
| 2022 | Semantic Graph Attention With Explicit Anatomical Association Modeling for Tooth Segmentation From CBCT ImagesabstractAccurate tooth identification and delineation in dental CBCT images are essential in clinical oral diagnosis and treatment. Teeth are positioned in the alveolar bone in a particular order, featuring similar appearances across adjacent and bilaterally symmetric teeth. However, existing tooth segmentation methods ignored such specific anatomical topology, which hampers the segmentation accuracy. Here we propose a semantic graph-based method to explicitly model the spatial associations between different anatomical targets (i.e., teeth) for their precise delineation in a coarse-to-fine fashion. First, to efficiently control the bilaterally symmetric confusion in segmentation, we employ a lightweight network to roughly separate teeth as four quadrants. Then, designing a semantic graph attention mechanism to explicitly model the anatomical topology of the teeth in each quadrant, based on which voxel-wise discriminative feature embeddings are learned for the accurate delineation of teeth boundaries. Extensive experiments on a clinical dental CBCT dataset demonstrate the superior performance of the proposed method compared with other state-of-the-art approaches. Pengcheng Li 0017, Yang Liu 0157, Zhiming Cui 0001, Feng Yang 0015, Yue Zhao 0012, Chunfeng Lian, Chenqiang Gao |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Video-to-Image Casting: A Flatting Method for Video AnalysisabstractPrevious mainstream video analysis methods, especially 3D CNNs-based models, mainly aim to transfer frameworks from the image domain to the video domain, and they follow the regime which has been succeeded in image processing, i.e., large-scale benchmarks and deep networks. However, processing videos is still time-consuming due to the increased computational cost. In this paper, we propose to flat the video and construct a Spatio-temporal Image (STI), i.e., squeezing the temporal dimension into a spatial plane. To pursuit the video-level modeling and efficient architecture, we devise a Collective Convolution (CoConv) operation to replace the 2D convolution. With the holistic sampling strategy, this novel operation can extract the video-level spatio-temporal representation. Moreover, we ensure that each CoConv operation has the same number of parameters as the original 2D filter, thus we can utilize a 2D network equipped with CoConv to analyze videos without additional computations. To verify the effectiveness of our method for the general video analysis, we evaluate it on three typical tasks, i.e., supervised action recognition, self-supervised action recognition, and dynamic texture recognition. Extensive experimental results show that our method can achieve comparable or state-of-the-art performances on these benchmarks while using much fewer computations compared with its 3D counterpart. Xu Chen 0053, Chenqiang Gao, Feng Yang 0015, Yi Yang 0001, Yahong Han |
ACM Multimedia | 3 |
| 2021 | Infrared and Visible Cross-Modal Image Retrieval Through Shared FeaturesabstractImage retrieval is one of the key techniques of computer vision, and has been studied for a long time. Nevertheless, little attention is paid to infrared and visible cross-modal retrieval which can be widely used in various applications, e.g., infrared and visible surveillance systems. In this paper, we propose a shared features based infrared-visible cross-modal image retrieval method. The similar visual features are extracted from infrared and visible images as the shared features, and the Euclidean distance is used to measure the similarity between these features. The core of the proposed method comes from three aspects: 1) Feature separation network can separate image features into shared features and exclusive features; 2) Maximum Mean Discrepancy (MMD) loss is employed to constrain the distribution of shared features, which can reduce the retrieval error caused by different imaging angles and similarity of infrared images. 3) The cross-layer fusion encoder compensates for the context loss in the convolution of infrared images. Experimental results on the Infrared-Visible dataset demonstrate the proposed method is effective and outperforms the state-of-the-art approaches. Fangcen Liu, Chenqiang Gao, Yongqing Sun, Yue Zhao 0012, Feng Yang 0015, Anyong Qin, Deyu Meng |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Adaptive Fusion and Mask Refinement Instance Segmentation Network for High Resolution Remote Sensing ImagesabstractInstance segmentation of remote sensing images (RSIs) is an active yet challenging task because of the huge scale variation and arbitrary complex shapes of objects. To address these issues, we propose an adaptive fusion and mask refinement (AFMR) instance segmentation network for RSIs in this paper. More precisely, AFMR consists of an adaptive fusion module to learn multi-scale complementary spatial features in an unsupervised manner, and a content-aware module for segmentation mask refinement. These two modules enable a better feature learning of convolutional neural network and boost the performance by 1.5% compared with the baseline method. In iSAID, a large-scale dataset for RSIs instance segmentation, our AFMR framework achieves the state-of-the-art accuracy, which verifies the superiority of the proposed method. Jie Ran, Feng Yang 0015, Chenqiang Gao, Yue Zhao 0012, Anyong Qin |
IGARSS | 2 |
| 2020 | TSASNet: Tooth segmentation on dental panoramic X-ray images by Two-Stage Attention Segmentation Network
Yue Zhao 0012, Pengcheng Li 0017, Chenqiang Gao, Yang Liu 0157, Qiaoyi Chen, Feng Yang 0015, Deyu Meng |
Knowl. Based Syst. | 6 |
| 2019 | Learning the Synthesizability of Dynamic Texture SamplesabstractExemplar-based dynamic texture synthesis (EDTS) is targeted to generate new samples of high quality that are perceptually similar to a given input dynamic texture exemplar. This paper addresses the issue of learning the synthesizability of dynamic texture samples. Given a dynamic texture sample, how is its possibility of being synthesized by EDTS methods estimated, and what is the most suitable EDTS algorithm to complete the task? To this end, we propose associating dynamic texture samples with synthesizability scores by learning regression models on a compiled dynamic texture dataset annotated in terms of synthesizability. More precisely, we first define the synthesizability of DT samples and characterize them by a set of spatiotemporal features. We then train regression models on the annotated dataset with feature representation to predict the synthesizability scores of the DT samples and learn classifiers to select the most suitable EDTS algorithm. We further complete the selection, partition and synthesizability prediction of the DT samples in a hierarchical scheme. The learned synthesizability is finally applied to detecting synthesizable regions in videos. Both quantitative and qualitative experiments demonstrate that our method can efficiently learn and predict the synthesizability of DT samples. Feng Yang 0015, Gui-Song Xia, Dengxin Dai, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Delving into the Synthesizability of Dynamic Texture SamplesabstractThe example-based dynamic texture synthesis (EDTS) methods have emerged in multitude, dedicated to generating new dynamic textures (DTs) of high quality from an input exemplar. The problem of EDTS has been studied for several decades, but none of the existing synthesis methods are able to tackle all kinds of dynamic textures equally well. Rather than focus on new synthesis methods, we turn to another way to help EDTS by investigating dynamic texture synthesizability - how synthesizable a specific dynamic texture sample is by EDTS. We propose to predict synthesizability score of a given dynamic texture sample, and suggest which EDTS method is best suited to synthesize it. To this end, we compiled a dynamic texture dataset and annotated each DT in terms of synthesizability. We address the problem of learning dynamic texture synthesizability by using regression model to train a predictor on the data collection. More precisely, we first characterize DT samples by a set of spatiotemporal features. Then, based on dynamic texture descriptors, we train regression models to estimate synthesizability scores and use an additional classifier to choose the optimal EDTS methods. The experiments demonstrate that our method can predict the synthesizability of DT samples effectively. Feng Yang 0015, Gui-Song Xia, Dengxin Dai, Liangpei Zhang 0001 |
ICPR | 1 |
| 2016 | Dynamic texture recognition by aggregating spatial and temporal features via ensemble SVMs
Feng Yang 0015, Gui-Song Xia, Gang Liu 0013, Liangpei Zhang 0001, Xin Huang 0002 |
Neurocomputing | 1 |