VLDB 2026 Research / reviewers in the wild / expert
Junjie Zhang 0011
dblp:99/6243-11
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-0895-3127ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Frequency-Aware Vision-Language Multimodality Generalization Network for Remote Sensing Image ClassificationabstractThe booming remote sensing (RS) technology is giving rise to a novel multimodality generalization task, which requires the model to overcome data heterogeneity while possessing powerful cross-scene generalization ability. Moreover, most vision-language models usually describe surface materials using universal texts, lacking proprietary linguistic prior knowledge specific to different RS modalities. In this work, we formalize RS multimodality generalization (RSMG) as a learning paradigm, and propose a frequency-aware vision-language multimodality generalization network (FVMGN) for RS image classification. Specifically, a diffusion-based training-test-time augmentation (DTAug) strategy is designed to reconstruct multimodal land-cover distributions, enriching input information for FVMGN. Following that, to overcome multimodal heterogeneity, a multimodal wavelet disentanglement (MWDis) module is developed to learn cross-domain invariant features by resampling low and high frequency components in the frequency domain. Considering the characteristics of RS vision modalities, shared and proprietary class texts is designed as linguistic inputs for the transformer-based text encoder to extract diverse text features. For multimodal vision inputs, a spatial-frequency-aware image encoder (SFIE) is constructed to realize local-global feature reconstruction and representation. Finally, a multiscale spatial-frequency feature alignment (MSFFA) module is suggested to construct a unified semantic space, ensuring refined multiscale alignment of different text and vision features in spatial and frequency domains. Extensive experiments show that FVMGN has the excellent multimodality generalization ability compared with state-of-the-art methods. Junjie Zhang 0011, Feng Zhao 0005, Hanqiang Liu 0001, Jun Yu 0001 |
AAAI | 1 |
| 2026 | Generative Information-Guided Heterogeneous Cross-Fusion Network With Contrastive Learning for Multimodal Remote Sensing Image ClassificationabstractMultimodal remote sensing (RS) images exhibit distinct structure and distribution characteristics, making it challenging to design an effective multimodal RS image classification algorithm. Moreover, although existing deep learning-based methods have become the darling in the multimodal RS image classification, they usually lack effective exploration and explicit integration for generative information from different modalities. Aiming at the above challenges, a generative information-guided heterogeneous cross-fusion network with contrastive learning (GIHCN) is proposed for multimodal RS image classification. Firstly, to simulate the land-cover distributions from different modal data, a multimodal generative information learning architecture (MGILA) is constructed to capture the unsupervised heterogeneous distribution features. Secondly, to achieve bidirectional modeling between heterogeneous data and the reconstructed land-cover distributions, a heterogeneous data & generative information cross-attention module (HGCM) is designed to explore the complementarity between multimodal data and the reconstructed land-cover distributions. HGCM can provide the heterogeneous generative information for current modal data or provide the heterogeneous data support for current modal generative information, thereby obtaining cross-fusion sources with different attributes. Furthermore, we achieve the effective feature extraction for different cross-fusion sources by a designed multimodal contrastive learning framework (MCLF). Notably, to capture local information and long-range dependencies, a hybrid classification network with convolutional neural network and Mamba (CMNet) is proposed as the feature extraction backbone of each cross-fusion source to further improve the classification performance. Finally, we construct a joint multimodality loss function for MCLF, which can reduce the distribution difference between modalities while focusing on the information flow within and across the modality. Experimental results on four multimodal RS datasets confirm the effectiveness of GIHCN compared with other state-of-the-art methods. The source code will be released at https://github.com/ZJier/GIHCN. Junjie Zhang 0011, Feng Zhao 0005, Hanqiang Liu 0001, Jun Yu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Lightweight anchor-free one-level feature indoor personnel detection method based on transformer
Feng Zhao 0005, Yongheng Li, Hanqiang Liu 0001, Junjie Zhang 0011, Zhenglin Zhu |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Data and knowledge-driven deep multiview fusion network based on diffusion model for hyperspectral image classification
Junjie Zhang 0011, Feng Zhao 0005, Hanqiang Liu 0001, Jun Yu 0001 |
Expert Syst. Appl. | 1 |
| 2024 | Semi-Supervised Co-Training Model Using Convolution and Transformer for Hyperspectral Image ClassificationabstractDeep learning algorithms have shown significant advantages in hyperspectral image (HSI) classification. However, these algorithms usually require a large number of labeled samples and the annotation of these samples consumes massive time and resource costs. To achieve effective classification results in situations with small samples, a semi-supervised co-training model using convolution and transformer (SCM-CT) is proposed in this letter. Firstly, two different networks, namely multi-scale parallel CNN (MPCNN) and global and local transformer fusion network (GLTFN), are designed as co-training learners to extract multi-scale spectral-spatial features and global-local combined features in HSIs, respectively. Secondly, to ensure two learners generate reliable predictions and utilize more unlabeled samples with low confidence pseudo-labels, a self-adaptive threshold and conflict pseudo-labeling (SATCP) strategy is proposed to facilitate the model to learn more valuable spectral-spatial information from conflict predictions and improve the convergence speed and model performance. Finally, to prevent the learners from stepping into the collapse, the discrepancy loss is computed to reduce the similarity between the features extracted by the two learners, forcing them to learn different information from the same input. Experimental results on University of Pavia, Salinas Valley, and Houston 2013 datasets show that SCM-CT achieves overall accuracies of 97.42%, 95.60%, and 95.20%, respectively, outperforming the state-of-the-art methods. Feng Zhao 0005, Xiqun Song, Junjie Zhang 0011, Hanqiang Liu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Multiscale Alignment and Progressive Feature Fusion Network for High Resolution Remote Sensing Images Change DetectionabstractThe advancement of deep learning technology has significantly improved the performance of high-resolution remote sensing (HRRS) image change detection (CD) task. HRRS images on CD are usually captured under varying conditions, which may generate pseudo-changes caused by seasonal variations, changes in lighting angles, and object motion. These factors can interfere with the effective modeling of difference information. However, conventional methods of modeling difference information may not adequately address these challenges, leaving them susceptible to such factors. Therefore, this article proposes a multiscale alignment and progressive feature fusion network (MAPNet) based on convolutional neural network (CNN) and transformer, which can effectively model difference information. First, a flow- and attention-guided bitemporal alignment module (FA-BAFM) is designed to align bitemporal features and capture the differences between them. Second, a progressive difference feature fusion module (PDFFM) is developed to comprehensively fuse the difference information across various scales and levels. Finally, a local feature enhancement module (LFEM) is constructed to improve the ability of the backbone to extract both global and local features. Experimental results on the learning, vision, and remote sensing CD (LEVIR-CD), Wuhan University (WHU), and Sun Yat-Sen University CD (SYSU-CD) datasets show that MAPNet achieves F1-scores of 91.90%, 94.19%, and 83.69%, respectively, outperforming the state-of-the-art (SOTA) methods. The demo code will be released athttps://github.com/zlin9/mapnet. Feng Zhao 0005, Zhenglin Zhu, Hanqiang Liu 0001, Junjie Zhang 0011 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Multiple vision architectures-based hybrid network for hyperspectral image classification
Feng Zhao 0005, Junjie Zhang 0011, Zhe Meng, Hanqiang Liu 0001, Zhenhui Chang, JiuLun Fan 0001 |
Expert Syst. Appl. | 2 |
| 2023 | Convolution Transformer Fusion Splicing Network for Hyperspectral Image ClassificationabstractConvolutional neural networks (CNNs) have attained remarkable performance in hyperspectral image (HSI) classification owing to excellent locally modeling ability. However, the existing CNNs cannot capture global context information from HSI. Recently, vision transformer (ViT) has been proven to be effective in the image field. However, its retrieval of local space information in HSI classification is not satisfactory, and the input mode always leads to the loss of spatial location information and local information. In this letter, we propose a novel convolution transformer fusion splicing network (CTFSN) for HSI classification. From the perspective of local information and global information, this method adopts two feature fusion ways of addition and channel stacking to capture hyperspectral features. First, to effectively utilize shallow features and preserve spatial location information, we propose a residual splicing convolution block to serialize HSI. In addition, the convolutional transformer fusion block (CTFB) is designed to achieve additional local modeling on the basis of capturing global features. Finally, the dual branch fusion splicing module is adopted to fuse and splice the local features from the depthwise residual block and the global features from CTFB. Experimental results on three widely used datasets show that our method is superior to several other state-of-the-art classification methods. Feng Zhao 0005, Junjie Zhang 0011, Hanqiang Liu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Residual Dense Asymmetric Convolutional Neural Network for Hyperspectral Image ClassificationabstractRecently, convolutional neural networks (CNNs) show excellent performance on the hyperspectral image (HSI) classification tasks. However, traditional CNNs usually have insufficient feature discrimination and a large number of network parameters. In response to the above problems, a residual dense asymmetric convolutional network (RDACN) for HSI classification is proposed in this paper. Firstly, we de-sign a novel residual dense asymmetric convolutional block to effectively leverage the information of the previous layers. Moreover, the block adopts two feature fusion methods of addition and channel stacking to capture discriminative hyperspectral feature. Secondly, the ordinary square convolutional kernel is replaced with the asymmetric convolutional kernels, which can reduce CNN parameters. Finally, experimental results on three well-known hyperspectral datasets show that RDACN achieves competitive classification performance compared with the state-of-the-art CNNs. Zhe Meng, Junjie Zhang 0011, Feng Zhao 0005, Hanqiang Liu 0001, Zhenhui Chang |
IGARSS | 2 |
| 2022 | Convolution Transformer Mixer for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) can provide rich spectral information which can be helpful for accurate classification in many applications. Yet, incorporating spatial information in the classification process can improve the classification accuracy even further. Existing convolutional neural network (CNN) usually only focuses on local features in hyperspectral cubes, whereas the burgeoning vision transformer (ViT) is interested in global features in HSIs. In this letter, we propose a deep aggregated framework for HSI classification called convolution transformer mixer (CTMixer) to combine the advantages of the above two paradigms effectively. A group parallel residual block is firstly applied to capture local spectral-spatial features in the HSI patches. Secondly, a double-branch structure, consisting of the CNN and transformer branches, is developed to capture local-global hyperspectral features. Finally, to achieve an elegant combination of CNN and ViT, a novel local-global multi-head self-attention mechanism is proposed by introducing convolution operations in the multi-head self-attention mechanism to further improve the classification accuracy. Extensive experiments demonstrate that the CTMixer achieves competitive classification results on several common HSI datasets compared with other state-of-the-art networks. The source code for this work will be available at https://github.com/ZJier/CTMixer. Junjie Zhang 0011, Zhe Meng, Feng Zhao 0005, Hanqiang Liu 0001, Zhenhui Chang |
IEEE Geosci. Remote. Sens. Lett. | 1 |