VLDB 2026 Research / reviewers in the wild / expert
Tongzhen Zhang
dblp:87/11110
· DBLP profile ↗
14ranked-venue papers
2as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cycle-Refined Multidecision Joint Alignment Network for Unsupervised Domain Adaptive Hyperspectral Change DetectionabstractHyperspectral change detection, which provides abundant information on land cover changes in the Earth's surface, has become one of the most crucial tasks in remote sensing. Recently, deep-learning-based change detection methods have shown remarkable performance, but the acquirement of labeled data is extremely expensive and time-consuming. It is intuitive to learn changes from the scene with sufficient labeled data and adapting them into an unlabeled new scene. However, the nonnegligible domain shift between different scenes leads to inevitable performance degradation. In this article, a cycle-refined multidecision joint alignment network (CMJAN) is proposed for unsupervised domain adaptive hyperspectral change detection, which realizes progressive alignment of the data distributions between the source and target domains with cycle-refined high-confidence labeled samples. There are two key characteristics: 1) progressively mitigate the distribution discrepancy to learn domain-invariant difference feature representation and 2) update the high-confidence training samples of the target domain in a cycle manner. The benefit is that the domain shift between the source and target domains is progressively alleviated to promote change detection performance on the target domain in an unsupervised manner. Experimental results on different datasets demonstrate that the proposed method can achieve better performance than the state-of-the-art change detection methods. Jiahui Qu, Wenqian Dong, Tongzhen Zhang, Yunsong Li 0001, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Matrix Factorization Informed Interpretable Deep Network for Unregistered Hyperspectral and Multispectral Images FusionabstractConsidering the existing issues in unregistered hyper-spectral images (HSI) and multispectral images (MSI) fusion methods: i) the designed registration modules introduce a significant computational burden, and registration errors accumulate in fusion errors; ii) the methods lack model guidance, resulting in poor interpretability of the network. In this paper, we propose a matrix factorization informed interpretable deep network to address the challenges of unregistered HSI and MSI fusion (IUFNet). In particular, we derive an extended matrix factorization model for unregistered fusion (EUMF), which substitutes the abundance matrix of HSI containing low-resolution and distorted spatial information by the high-resolution abundance matrix of MSI. This substitution ingeniously eliminates the dependence of fusion performance on registration accuracy. Subsequently, IUFNet is designed to unfold the iterative results obtained by proximal gradient descent into the deep learning network, where each operation has a clear physical meaning. Overall, this network achieves the fusion of unregistered HSI and MSI and exhibits inter-pretability. Experimental results on the widely used Paiva Center dataset demonstrate the effectiveness and superiority of the proposed method. Tongzhen Zhang, Jiahui Qu, Yunsong Li 0001, Qian Du 0001, Wenqian Dong |
IGARSS | 1 |
| 2024 | Graph Representation Learning-Guided Diffusion Model for Hyperspectral Change DetectionabstractDue to its capability to monitor subtle changes occurring on the Earth’s surface, hyperspectral images change detection (HSI-CD) has emerged as a focal research area in the field of remote sensing. Recently, diffusion models have demonstrated remarkable performance in the field of HSI-CD. However, vanilla diffusion models are mostly constructed by CNN, which struggles to model global context relationships in complex scenes to result in limited change detection accuracy. In order to overcome the shortcomings about vanilla diffusion models, we innovatively design graph representation learning-guided diffusion model (GDM) and propose the GDM-based HSI-CD network (GDMCD). Specially, we utilize graph convolutional to construct the GDM as the feature extractor, which can adequately extract global difference features of HSIs. Then, we design the difference perception amplification module (DPAM) to increase the distinction between difference features extracted by GDM. Finally, we obtain the change map by classifying difference features which are processed by DPAM. Experiments conducted on three publicly available datasets with 1% sample size demonstrate that the proposed method outperforms the other state-of-the-art methods in terms of Overall Accuracy (OA), Kappa Coefficient (KC) achieving improvements of approximately 0.006%, 1.61%, and 0.34%, respectively. Xinyu Ding, Jiahui Qu, Wenqian Dong, Tongzhen Zhang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Progressive Multi-Iteration Registration-Fusion Co-Optimization Network for Unregistered Hyperspectral Image Super-ResolutionabstractExisting fusion-based hyperspectral image super-resolution (fusion-based HSI-SR) methods usually reconstruct high-resolution hyperspectral image (HR-HSI) by integrating the complementary information of low-resolution hyperspectral image (LR-HSI) and high-resolution multispectral image (HR-MSI). However, most of such methods rely on accurately registered images or consider registration and fusion as a two-stage task, which means that fusion must tolerate the accumulation of errors due to misregistration. In this paper, we propose a progressive multi-iteration registration-fusion co-optimization network (PMI-RFCoNet) for unregistered hyperspectral image super-resolution, which progressively refines the registration and fusion result over multiple levels to reconstruct registered HR-HSI. To achieve registration-fusion co-optimization, the registration-fusion cooptimization block (Co-RFB) is designed to iterate continuously over multiple levels. We embed the interactive registration module (IRM) and the spectral recalibration and fusion module (SRFU) in Co-RFB, which can facilitate the network utilizing spatial and spectral features at different levels to generate more accurate HR-HSI. Specifically, IRM generates deformation field based on spatial correlations captured at long distances to repair non-rigid pixel offsets, and SRFU further performs adaptive high-fidelity spectral correction and spatial information fusion on the registration results. We conduct experimental verification on four widely used datasets, and the results show that PMI-RFCoNet can flexibly cope with different types and degrees of non-rigid deformation and achieve superior performance. Code is available at https://github.com/Jiahuiqu/PMI-RFCoNet. Jiahui Qu, Xuyao Liu, Wenqian Dong, Yang Liu 0084, Tongzhen Zhang, Yang Xu 0070, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | TMCFN: Text-Supervised Multidimensional Contrastive Fusion Network for Hyperspectral and LiDAR ClassificationabstractThe joint classification of hyperspectral images (HSIs) and LiDAR data plays a crucial role in earth observation missions. Most advanced methods are based on discrete label supervision. However, since discrete labels only convey limited information that a sample belongs to a single definite class and lack of prior information, it is difficult to supervise the model to capture rich inherent semantic information in complex data distributions, hindering the classification performance. To this end, we propose a text-supervised multidimensional contrastive fusion network, termed as TMCFN, which leverages class text information to guide the learning of visual representations while establishing a semantic association of text and visual features for classification by using multidimensionally incorporated contrastive learning (CL) paradigms. Specifically, TMCFN is composed of text information encoding (TIE), visual features representation (VFR) and text-visual features alignment and classification (TVFAC). TIE is employed to extract semantic information from class text extended from class names, intrinsic attributes and inter-class relationships. VFR mainly comprises a new fusion-based contrastive feature learning module (FCFLM) to extract discriminative visual features and a text-guided attention feature fusion module (TAF2M) to fuse visual features under the guidance of text information. TVFAC optimizes the learning of visual features under the supervision of text information while using a CL paradigm to align text and visual features for establishing the semantic association, and achieves the classification by directly computing the similarity between the visual features and each text feature without an additional classifier. Experiments with three standard datasets verify the effectiveness of TMCFN. Yueguang Yang, Jiahui Qu, Wenqian Dong, Tongzhen Zhang, Song Xiao 0001, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Noise Prior Knowledge Informed Bayesian Inference Network for Hyperspectral Super-ResolutionabstractWell-known deep learning (DL) is widely used in fusion based hyperspectral image super-resolution (HS-SR). However, DL-based HS-SR models have been designed mostly using off-the-shelf components from current deep learning toolkits, which lead to two inherent challenges: i) they have largely ignored the prior information contained in the observed images, which may cause the output of the network to deviate from the general prior configuration; ii) they are not specifically designed for HS-SR, making it hard to intuitively understand its implementation mechanism and therefore uninterpretable. In this paper, we propose a noise prior knowledge informed Bayesian inference network for HS-SR. Instead of designing a "black-box" deep model, our proposed network, termed as BayeSR, reasonably embeds the Bayesian inference with the Gaussian noise prior assumption to the deep neural network. In particular, we first construct a Bayesian inference model with the Gaussian noise prior assumption that can be solved iteratively by the proximal gradient algorithm, and then convert each operator involved in the iterative algorithm into a specific form of network connection to construct an unfolding network. In the process of network unfolding, based on the characteristics of the noise matrix, we ingeniously convert the diagonal noise matrix operation which represents the noise variance of each band into the channel attention. As a result, the proposed BayeSR explicitly encodes the prior knowledge possessed by the observed images and considers the intrinsic generation mechanism of HS-SR through the whole network flow. Qualitative and quantitative experimental results demonstrate the superiority of the proposed BayeSR against some state-of-the-art methods. Wenqian Dong, Jiahui Qu, Song Xiao 0001, Tongzhen Zhang, Yunsong Li 0001, Xiuping Jia |
IEEE Trans. Image Process. | 4 |
| 2022 | A Mutual Guidance Attention-Based Multi-Level Fusion Network for Hyperspectral and LiDAR ClassificationabstractHyperspectral image (HSI) and light detection and ranging (LiDAR) data classification has attracted more and more attention in remote sensing. Convolution neural network (CNN) has been proven to be effective for HSI and LiDAR data classification. In this letter, a novel three-branch CNN is designed to learn spectral, spatial, and elevation features, each of which adopts the multi-level feature fusion (MLF) module to fuse the shallow and deep features. Furthermore, in order to fully fuse the spatial and elevation information, we propose a mutual guidance attention (MGA) module. The MGA module increases the information flow between spatial and elevation branches, highlights the features of interest, and weakens useless features. The proposed method is evaluated on public datasets Houston and Trento. Experimental results demonstrate that our proposed method can provide higher classification accuracy than some existing methods. Tongzhen Zhang, Song Xiao 0001, Wenqian Dong, Jiahui Qu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Context-Aware Guided Attention Based Cross-Feedback Dense Network for Hyperspectral Image Super-ResolutionabstractConvolutional neural networks (CNNs) have shown impressive performance in computer vision due to their non-linearity. Particularly, DenseNet that facilitates feature re-use in a feedforward manner has achieved state-of-the-art reconstruction accuracy for super-resolution (SR). However, most DenseNet based SR models transfer the features generated from each layer to all the subsequent layers, inevitably introducing redundancy, especially for high-dimensional hyperspectral (HS) images. To tackle this problem, we propose a two-branch cross-feedback dense network with context-aware guided attention (CFDcagaNet) for HS super-resolution (HSSR), which allows the network to learn the attention maps of high-level features and refine the low-level features in a feedback manner across two branches. Context-aware guided attention uses high-level posterior information to provide more faithful spatial-spectral guidance for low-level features, which enables CFDcagaNet to learn more effective spatial-spectral features at low levels and yield more effective spatial-spectral transfer in the network. Extensive experiments on widely-used datasets demonstrate that the proposed method outperforms state-of-the-art methods in terms of both quantitative values and visual qualities. Wenqian Dong, Jiahui Qu, Tongzhen Zhang, Yunsong Li 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | A Spatial-Spectral Dual-Optimization Model-Driven Deep Network for Hyperspectral and Multispectral Image FusionabstractDeep learning, especially convolutional neural networks (CNNs), has shown very promising results for multispectral (MS) and hyperspectral (HS) image fusion (MS/HS fusion) task. Most of the existing CNN methods are based on “black-box” models that are not specifically designed for MS/HS fusion, which largely ignore the priors evidently possessed by the observed HS and MS images, and lack clear interpretability, leaving room for further improvement. In this paper, we propose an interpretable network, named as spatial-spectral dual-optimization model driven deep network (S2DMDN), which embeds the intrinsic generation mechanism of the MS/HS fusion to the network. There are two key characteristics: (i) Explicitly encode the spatial prior and spectral prior evidently possessed by the input MS and HS images in the network architecture; (ii) Unfold an iterative spatial-spectral dual-optimization algorithm into a model driven deep network. The benefit is that the network has good interpretability and generalization capability, and the fused image is richer in semantics and more precise in spatial. Extensive experiments are conducted to prove the superiority of our proposed method over other state-of-the-art methods in terms of quantitative evaluation metrics and qualitative visual effects. Wenqian Dong, Tongzhen Zhang, Jiahui Qu, Yunsong Li 0001, Haoming Xia |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Laplacian Pyramid Dense Network for Hyperspectral PansharpeningabstractHyperspectral (HS) pansharpening aims to create a pansharpened image that integrates the spatial details of the panchromatic (PAN) image and the spectral content of the HS image. In this article, we present a deep convolutional network within the mature Gaussian–Laplacian pyramid for pansharpening (LPPNet). The overall structure of LPPNet is a cascade of the Laplacian pyramid dense network with a similar structure at each pyramid level. Following the general idea of multiresolution analysis (MRA), the subband residuals of the desired HS images are extracted from the PAN image and injected into the upsampled HS image to reconstruct the high-resolution HS images level by level. Applying the mature Laplace pyramid decomposition technique to the convolution neural network (CNN) can simplify the pansharpening problem into several pyramid-level learning problems so that the pansharpening problem can be solved with a shallow CNN with fewer parameters. Specifically, the Laplacian pyramid technology is used to decompose the image into different levels that can differentiate large- and small-scale details, and each level is handled by a spatial subnetwork in a divide-and-conquer way to make the network more efficient. Experimental results show that the proposed LPPNet method performs favorably against some state-of-the-art pansharpening methods in terms of objective indexes and subjective visual appearance. Wenqian Dong, Tongzhen Zhang, Jiahui Qu, Song Xiao 0001, Jie Liang 0001, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Multibranch Feature Fusion Network With Self- and Cross-Guided Attention for Hyperspectral and LiDAR ClassificationabstractThe effective fusion of multi-source data helps to improve performance of land cover classification. Most existing convolutional neural network (CNN) based methods adopt an early/late fusion strategy to fuse the low-level/high-level features for classification, which still has two inherent challenges: i) the conventional convolution operation performs a weighted average operation on each pixel in the receptive field, which will reduce the discriminability of the center pixel due to the influence of the interference pixels, and ii) the spatial-spectral features of the hyperspectral image (HSI), the elevation features of light detection and ranging (LiDAR), and the complementary features between the multimodal data are not fully exploited, which results in the reduction of classification accuracy. In this paper, an effective multi-branch feature fusion network with self- and cross-guided attention (MB2FscgaNet) is proposed for joint classification of LiDAR and HSI. The main concern of this paper is how to accurately estimate more effective spectral-spatial-elevation features and yield more effective transfer in network. Specifically, MB2FscgaNet adopts a multi-branch feature fusion architecture to fully exploit the hierarchical features from LiDAR and HSI level by level. At each level of the network, a self- and cross-guided attention (SCGA) is developed to assign higher weight to interesting areas and channels of LiDAR and HSI feature maps to obtain refined spectral-spatial-elevation features and provide complementary information cross guidance between LiDAR and HS. We further designed a spectral supplement module (SeSuM) to improve the discriminative ability of the center pixel. Comparative classification results and ablation studies demonstrate that the proposed MB2FscgaNet achieves competitive performance against state-of-the-art methods. Wenqian Dong, Tian Zhang 0017, Jiahui Qu, Song Xiao 0001, Tongzhen Zhang, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Cross-Guided Pyramid Attention-Based Residual Hyperdense Network for Hyperspectral Image PansharpeningabstractHyperspectral image pansharpening is of great importance in improving spatial resolution for many commercial platforms and remote sensing tasks. Convolutional neural network (CNN) has recently been applied in pansharpening. However, most existing CNN-based pansharpening models followed an early-fusion/late-fusion strategy, which integrates the low-level/high-level features of PAN and HS streams at the input/output of the network. It is difficult to learn more complex combinations between panchromatic (PAN) and hyperspectral (HS) streams. This paper proposes a novel end-to-end residual hyper-dense pansharpening network with a cross-guided pyramid attention (called RHDcgpaNet). The overall architecture of the proposed method is a residual hyper-dense network, which extends the definition of dense connections to two-stream pansharpening problem. The proposed RHDcgpaNet allows guidance from the state of the preceding layers to all the layers in-between PAN and HS streams in a feed-forward manner, significantly increasing the learning representation. A cross-guided pyramid attention is designed and embedded to the proposed residual hyper-dense network to yield more useful spatial-spectral feature transfer in network. Extensive experiments on widely-used datasets demonstrate that the proposed RHDcgpaNet achieves favorable performance in comparison with state-of-the-art methods. Jiahui Qu, Tongzhen Zhang, Wenqian Dong, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | End to end multi-scale convolutional neural network for crowd countingabstractCrowd counting is a challenging task in computer vison field and haven’t been well addressed until now. In this paper, we intend to develop an end to end multi-scale deep convolutional neural network(CNN) model that can accurately estimate the crowd count from an individual image with arbitrary crowd density and perspective. The proposed model extract multi-scale deep CNN features from the input image and regress the crwod count directly, without any post-processing . Hence our model could handle muti-scale targets well in various crowd scene. We evaluate our model on several benchmark datasets and the performance outperforms some state-of-the-art methods. What’s more, due to the end-to-end characteristics, our model demonstrates good practical application performance. Deyi Ji, Hongtao Lu 0001, Tongzhen Zhang |
ICMV | 3 |
| 2017 | Bayesian Supervised HashingabstractAmong learning based hashing methods, supervised hashing seeks compact binary representation of the training data to preserve semantic similarities. Recent years have witnessed various problem formulations and optimization methods for supervised hashing. Most of them optimize a form of loss function with a regulization term, which can be viewed as a maximum a posterior (MAP) estimation of the hashing codes. However, these approaches are prone to overfitting unless hyperparameters are tuned carefully. To address this problem, we present a novel fully Bayesian treatment for supervised hashing problem, named Bayesian Supervised Hashing (BSH), in which hyperparameters are automatically tuned during optimization. Additionally, by utilizing automatic relevance determination (ARD), we can figure out relative discriminating ability of different hashing bits and select most informative bits among them. Experimental results on three real-world image datasets with semantic information show that BSH can achieve superior performance over state-of-the-art methods with comparable training time. Zihao Hu, Junxuan Chen, Tongzhen Zhang |
CVPR | 4 |