Tianjing Zhang

dblp:309/2963 · also Tian-Jing Zhang · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
17since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 10 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021
YearPublicationVenuePosition
2025 Zero-Shot Blind-Spot Image Denoising via Cross-Scale Non-Local Pixel Refilling
abstract
Blind-spot denoising (BSD) method is a powerful paradigm for zero-shot image denoising by training models to predict masked target pixels from their neighbors. However, they struggle with real-world noise exhibiting strong local correlations, where efforts to suppress noise correlation often weaken pixel-value dependencies, adversely affecting denoising performance. This paper presents a theoretical analysis quantifying the impact of replacing masked pixels with observations exhibiting weaker noise correlation but potentially reduced similarity, revealing a trade-off that impacts the statistical risk of the estimation. Guided by this insight, we propose a computational scheme that replaces masked pixels with distant ones of similar appearance and lower noise correlation. This strategy improves the prediction by balancing noise suppression and structural consistency. Experiments confirm the effectiveness of our method, outperforming existing zero-shot BSD methods.
Qilong Guo, Tianjing Zhang, Hui Ji 0002
NeurIPS2
2024 Test-Time Model Adaptation for Image Reconstruction Using Self-supervised Adaptive Layers
Yutian Zhao, Tianjing Zhang, Hui Ji 0002
ECCV (37)2
2024 Cross-Scale Self-Supervised Blind Image Deblurring via Implicit Neural Representation
abstract
Blind image deblurring (BID) is an important yet challenging image recovery problem. Most existing deep learning methods require supervised training with ground truth (GT) images. This paper introduces a self-supervised method for BID that does not require GT images. The key challenge is to regularize the training to prevent over-fitting due to the absence of GT images. By leveraging an exact relationship among the blurred image, latent image, and blur kernel across consecutive scales, we propose an effective cross-scale consistency loss. This is implemented by representing the image and kernel with implicit neural representations (INRs), whose resolution-free property enables consistent yet efficient computation for network training across multiple scales. Combined with a progressively coarse-to-fine training scheme, the proposed method significantly outperforms existing self-supervised methods in extensive experiments.
Tianjing Zhang, Yuhui Quan, Hui Ji 0002
NeurIPS1
2024 KNLConv: Kernel-Space Non-Local Convolution for Hyperspectral Image Super-Resolution
abstract
Pixel-level adaptive convolution, which overcomes the deficiency of the spatial-invariance of standard convolution, is always limited to performing feature extraction from local patches and ignores the latent long-range dependencies imperceptible in the feature space, which are more significant in pixel-level tasks such as hyperspectral image super-resolution (HSISR). To handle such limitations, we propose kernel-space non-local convolution (KNLConv), which explores non-local dependencies in the generated kernel space, to leverage these global information to guide the network to extract image features more flexibly. Technically, the proposed KNLConv first decomposes the convolutional kernel space into spatial and channel dimensions, and designs a depth-wise non-local expansion convolution (NLEC) in the spatial dimension of the kernel-space to explore underlying global correlations. Then introduce an adaptive point-wise convolution (APC), generalizing the NLEC to the pixel-level while integrating features in the channel dimension. In addition, applying KNLConv, we design an effective network architecture for hyperspectral image super-resolution. Extensive experiments demonstrate that our approach performs favorably against current state-of-the-art HSISR methods, both on quantitative indicators and visual quality.
Ran Ran 0001, Liang-Jian Deng, Tianjing Zhang, Jianlong Chang, Qi Tian 0001
IEEE Trans. Multim.3
2023 LGPConv: Learnable Gaussian Perturbation Convolution for Lightweight Pansharpening
abstract
Pansharpening is a crucial and challenging task that aims to obtain a high spatial resolution image by merging a multispectral (MS) image and a panchromatic (PAN) image. Current methods use CNNs with standard convolution, but we've observed strong correlation among channel dimensions in the kernel, leading to computational burden and redundancy. To address this, we propose Learnable Gaussian Perturbation Convolution (LGPConv), surpassing standard convolution. LGPConv leverages two properties of standard convolution kernels: 1) correlations within channels, learning a premier kernel as a base to reduce parameters and training difficulties caused by redundancy; 2) introducing Gaussian noise perturbations to simulate randomness and enhance nonlinear representation within channels. We incorporate LGPConv into a well-designed pansharpening network and demonstrate its superiority through extensive experiments, achieving state-of-the-art performance with minimal parameters (27K). Code is available on the GitHub page of the authors.
Chen-Yu Zhao, Tianjing Zhang, Ran Ran 0001, Zhi-Xuan Chen, Liang-Jian Deng
IJCAI2
2023 Cloud Game Video Coding Based On Human Eye Fixation Point
abstract
Cloud Gaming enables users to run high-quality games on thin clients with limited graphics processing and data computing capabilities. Under the running mode of cloud games, all games are run on the server side, and the rendered game picture is compressed and sent to the user through the network. On the client side, the user's gaming device doesn't need any high-end processors or graphics cards, only basic video extraction capabilities. However, cloud gaming requires a high bandwidth connection to present a good-quality game picture to the user. At present, the main problem is limited bandwidth during transmission, which leads to poor quality of the game image received by users and poor user experience, which has become an important problem hindering the popularity of cloud games. To solve this problem, we observed that when playing a game, the user's eyes are not always focused on the entire picture, and due to the characteristics of visual perception, the user pays more attention to the area around the eye fixation point. In this paper, for the first time, we add information about user interactions with devices to the network to more accurately predict user fixation points. Using visual perception features, more bit rates are assigned to ROI regions near the user's fixation point. Our results show that our network is able to predict more accurate user fixation points, and that our approach can significantly improve the subjective quality of the area near the user's focus of the game frame at the same bitrate, with the VMAF scores is 2 to 3 points higher on average compared to H.264, the most commonly used standard encoder in cloud games today.
Geng Wei, Tianjing Zhang, Ming Lu 0003, Hao Chen 0036
MMSP2
2023 A Triple-Double Convolutional Neural Network for Panchromatic Sharpening
abstract
Pansharpening refers to the fusion of a panchromatic (PAN) image with a high spatial resolution and a multispectral (MS) image with a low spatial resolution, aiming to obtain a high spatial resolution MS (HRMS) image. In this article, we propose a novel deep neural network architecture with level-domain-based loss function for pansharpening by taking into account the following double-type structures, i.e., double-level, double-branch, and double-direction, called as triple-double network (TDNet). By using the structure of TDNet, the spatial details of the PAN image can be fully exploited and utilized to progressively inject into the low spatial resolution MS (LRMS) image, thus yielding the high spatial resolution output. The specific network design is motivated by the physical formula of the traditional multi-resolution analysis (MRA) methods. Hence, an effective MRA fusion module is also integrated into the TDNet. Besides, we adopt a few ResNet blocks and some multi-scale convolution kernels to deepen and widen the network to effectively enhance the feature extraction and the robustness of the proposed TDNet. Extensive experiments on reduced- and full-resolution datasets acquired by WorldView-3, QuickBird, and GaoFen-2 sensors demonstrate the superiority of the proposed TDNet compared with some recent state-of-the-art pansharpening approaches. An ablation study has also corroborated the effectiveness of the proposed approach. The code is available at https://github.com/liangjiandeng/TDNet.
Tianjing Zhang, Liang-Jian Deng, Ting-Zhu Huang, Jocelyn Chanussot, Gemine Vivone
IEEE Trans. Neural Networks Learn. Syst.1
2022 LAGConv: Local-Context Adaptive Convolution Kernels with Global Harmonic Bias for Pansharpening
abstract
Pansharpening is a critical yet challenging low-level vision task that aims to obtain a higher-resolution image by fusing a multispectral (MS) image and a panchromatic (PAN) image. While most pansharpening methods are based on convolutional neural network (CNN) architectures with standard convolution operations, few attempts have been made with context-adaptive/dynamic convolution, which delivers impressive results on high-level vision tasks. In this paper, we propose a novel strategy to generate local-context adaptive (LCA) convolution kernels and introduce a new global harmonic (GH) bias mechanism, exploiting image local specificity as well as integrating global information, dubbed LAGConv. The proposed LAGConv can replace the standard convolution that is context-agnostic to fully perceive the particularity of each pixel for the task of remote sensing pansharpening. Furthermore, by applying the LAGConv, we provide an image fusion network architecture, which is more effective than conventional CNN-based pansharpening approaches. The superiority of the proposed method is demonstrated by extensive experiments implemented on a wide range of datasets compared with state-of-the-art pansharpening methods. Besides, more discussions testify that the proposed LAGConv outperforms recent adaptive convolution techniques for pansharpening.
Zi-Rong Jin, Tianjing Zhang, Tai-Xiang Jiang, Gemine Vivone, Liang-Jian Deng
AAAI2
2022 Cross-Frequency Detail Compensation Network for Pansharpening
abstract
Pansharpening is a fusion technique aiming at improving the spatial resolution of multispectral images while preserving spectral information. Previous attempts to adopt CNNs have led to significant progress in pansharpening, but always with a cumbersome network structure, and there exists redundancy in both spatial and channel of feature maps learned by CNNs. Considering the distinct properties of components with different frequencies in the feature map, we propose a cross-frequency detail compensation network (CFDCNet) by processing low, medium, and high frequency separately. Specifically, a cross-frequency convolution block is designed to produce a representation that captures the different frequency classes while achieving the more efficient detail extraction. Overall pipeline is progressive, and the learned features are fused in an interactive compensation manner to obtain the final output. Experimental results demonstrate the superiority of CFDCNet over state-of-the-art pansharpening methods in terms of visual quality and quantitative metrics.
Xiao-Nan Zhao, Chen-Yu Zhao, Tianjing Zhang, Liang-Jian Deng
IGARSS3
2022 SpanConv: A New Convolution via Spanning Kernel Space for Lightweight Pansharpening
abstract
Standard convolution operations can effectively perform feature extraction and representation but result in high computational cost, largely due to the generation of the original convolution kernel corresponding to the channel dimension of the feature map, which will cause unnecessary redundancy. In this paper, we focus on kernel generation and present an interpretable span strategy, named SpanConv, for the effective construction of kernel space. Specifically, we first learn two navigated kernels with single channel as bases, then extend the two kernels by learnable coefficients, and finally span the two sets of kernels by their linear combination to construct the so-called SpanKernel. The proposed SpanConv is realized by replacing plain convolution kernel by SpanKernel. To verify the effectiveness of SpanConv, we design a simple network with SpanConv. Experiments demonstrate the proposed network significantly reduces parameters comparing with benchmark networks for remote sensing pansharpening, while achieving competitive performance and excellent generalization. Code is available at https://github.com/zhi-xuan-chen/IJCAI-2022 SpanConv.
Zhi-Xuan Chen, Cheng Jin 0003, Tianjing Zhang, Liang-Jian Deng
IJCAI3
2022 A Decoder-free Transformer-like Architecture for High-efficiency Single Image Deraining
abstract
Despite the success of vision Transformers for the image deraining task, they are limited by computation-heavy and slow runtime. In this work, we investigate Transformer decoder is not necessary and has huge computational costs. Therefore, we revisit the standard vision Transformer as well as its successful variants and propose a novel Decoder-Free Transformer-Like (DFTL) architecture for fast and accurate single image deraining. Specifically, we adopt a cheap linear projection to represent visual information with lower computational costs than previous linear projections. Then we replace standard Transformer decoder block with designed Progressive Patch Merging (PPM), which attains comparable performance and efficiency. DFTL could significantly alleviate the computation and GPU memory requirements through proposed modules. Extensive experiments demonstrate the superiority of DFTL compared with competitive Transformer architectures, e.g., ViT, DETR, IPT, Uformer, and Restormer. The code is available at https://github.com/XiaoXiao-Woo/derain.
Ting-Zhu Huang, Liang-Jian Deng, Tianjing Zhang
IJCAI4
2022 A Cascaded Multi-Task Generative Framework for Detecting Aortic Dissection on 3-D Non-Contrast-Enhanced Computed Tomography
abstract
Contrast-enhanced computed tomography (CE-CT) is the gold standard for diagnosing aortic dissection (AD). However, contrast agents can cause allergic reactions or renal failure in some patients. Moreover, AD diagnosis by radiologists using non-contrast-enhanced CT (NCE-CT) images has poor sensitivity. To address this issue, we propose a novel cascaded multi-task generative framework for AD detection using NCE-CT volumes. The framework includes a 3D nnU-Net and a 3D multi-task generative architecture (3D MTGA). Specifically, the 3D nnU-Net was employed to segment aortas from NCE-CT volumes. The 3D MTGA was then employed to simultaneously synthesize CE-CT volumes, segment true & false lumen, and classify the patient as AD or non-AD. A theoretical formulation demonstrated that the 3D MTGA could increase the Jensen-Shannon Divergence (JSD) between AD and non-AD for each NCE-CT volume, thus indirectly improving the AD detection performance. Experiments also showed that the proposed framework could achieve an average accuracy of 0.831, a sensitivity of 0.938, and an F1-score of 0.847 in comparison with seven state-of-the-art classification models used by three radiologists with junior, intermediate, and senior experiences, respectively. The experimental results indicate that the proposed framework obtains superior performance to state-of-the-art models in AD detection. Thus, it has great potential to reduce the misdiagnosis of AD using NCE-CT in clinical practice. The source codes and supplementary materials for our framework are available at https://github.com/yXiangXiong/CMTGF.
Xiangyu Xiong, Chuanqi Sun, Zhuoneng Zhang, Xiuhong Guan, Tianjing Zhang, Hao Chen 0037, Zhangbo Cheng, Xiaohai Ma, Guoxi Xie
IEEE J. Biomed. Health Informatics6
2021 Dynamic Cross Feature Fusion for Remote Sensing Pansharpening
abstract
Deep Convolution Neural Networks have been adopted for pansharpening and achieved state-of-the-art performance. However, most of the existing works mainly focus on single-scale feature fusion, which leads to failure in fully considering relationships of information between high-level semantics and low-level features, despite the network is deep enough. In this paper, we propose a dynamic cross feature fusion network (DCFNet) for pansharpening. Specifically, DCFNet contains multiple parallel branches, including a high-resolution branch served as the backbone, and the low-resolution branches progressively supplemented into the backbone. Thus our DCFNet can represent the overall information well. In order to enhance the relationships of inter-branches, dynamic cross feature transfers are embedded into multiple branches to obtain high-resolution representations. Then contextualized features will be learned to improve the fusion of information. Experimental results indicate that DCFNet significantly outperforms the prior arts in both quantitative indicators and visual qualities.
Ting-Zhu Huang, Liang-Jian Deng, Tianjing Zhang
ICCV4
2021 Weighted Shallow-Deep Feature Fusion Network for Pansharpening
abstract
In this paper, we propose a novel weighted shallow-deep feature fusion convolutional neural network (WSDFNet) for the task of multispectral image pansharpening. This network could effectively overcome the drawback of the common identity skip connection (ISC), and propagate shallow features scaled by a novel adaptive skip weighter (ASW) to deeper layers. By the technique, it could favor the feature fusion in different network depths adequately, as well as yield a promising outcome. Experimental results on reduced- and full-resolution WorldView-3 dataset demonstrate the superiority of the WSDFNet compared with recent state-of-the-art (SOTA) pansharpening approaches. Moreover, WSDFNet is also verified as a lightweight network.
Zi-Rong Jin, Tianjing Zhang, Cheng Jin 0003, Liang-Jian Deng
IGARSS2
2021 Progressive Band-Separated Convolutional Neural Network for Multispectral Pansharpening
abstract
Recently, convolutional neural networks (CNNs) have been introduced to pansharpening for enhancing fusion accuracy and overcoming the drawbacks of the conventional methods. However, most of methods based on CNN fail to distinguish the difference of multispectral bands, and only use a uniform set of convolutional kernels to extract features. In this paper, we design a progressive, band-separated convolutional network architecture for discriminatively learning the features and relation among spectral bands, aiming to address the problem mentioned before. More specifically, the proposed architecture mainly consists of three aspects. First, to accurately preserve the spectral peculiarities, we divide the multispectral input image in terms of its bands into several groups. Second, our original panchromatic and multispectral inputs are filtered by a high-pass operation to further yield more spatial details. Third, we use a spectral fusion module (SFM) for each group and associate them to progressively assemble the whole architecture. It is worth mentioning that the architecture could be integrated into any other competitive CNNs to improve the performance. Both visual and quantitative experiments have demonstrated that our proposed method outperforms recent state-of-the-art pansharpening techniques.
Shishi Xiao, Cheng Jin 0003, Tianjing Zhang, Ran Ran 0001, Liang-Jian Deng
IGARSS3
2021 BAM: Bilateral Activation Mechanism for Image Fusion
abstract
As the conventional activation functions such as ReLU, LeakyReLU, and PReLU, the negative parts in feature maps are simply truncated or linearized, which may result in unflexible structure and undesired information distortion. In this paper, we propose a simple but effective Bilateral Activation Mechanism (BAM) which could be applied to the activation function to offer an efficient feature extraction model. Based on BAM, the Bilateral ReLU Residual Block (BRRB) that still sufficiently keeps the nonlinear characteristic of ReLU is constructed to separate the feature maps into two parts, i.e., the positive and negative components, then adaptively represent and extract the features by two independent convolution layers. Besides, our mechanism will not increase any extra parameters or computational burden in the network. We finally embed the BRRB into a basic ResNet architecture, called BRResNet, it is easy to obtain state-of-the-art performance in two image fusion tasks, i.e., pansharpening and hyperspectral image super-resolution (HISR). Additionally, deeper analysis and ablation study demonstrate the effectiveness of BAM, the lightweight property of the network, etc. Please find the code from the project page1 https://liangjiandeng.github.io/Projects_Res/bam_mm2021.html
Zi-Rong Jin, Liang-Jian Deng, Tianjing Zhang, Xiao-Xu Jin
ACM Multimedia3
2021 SSconv: Explicit Spectral-to-Spatial Convolution for Pansharpening
abstract
Pansharpening aims to fuse a high spatial resolution panchromatic (PAN) image and a low resolution multispectral (LR-MS) image to obtain a multispectral image with the same spatial resolution as the PAN image. Thanks to the flexible structure of convolution neural networks (CNNs), they have been successfully applied to the problem of pansharpening. However, most of the existing methods only simply feed the up-sampled LR-MS into the CNNs and ignore the spatial distortion caused by direct up-sampling. In this paper, we propose an explicit spectral-to-spatial convolution (SSconv) that aggregates spectral features into the spatial domain to perform the up-sampling operation, which can get better performance than the direct up-sampling. Furthermore, SSconv is embedded into a multiscale U-shaped convolution neural network (MUCNN) for fully utilizing the multispectral information of involved images. In particular, multiscale injection branch and mixed loss on cross-scale levels are employed to fuse pixel-wise image information. Benefiting from the distortion-free property of SSconv, the proposed MUCNN can generate state-of-the-art performance with a simple structure, both on reduced-resolution and full-resolution datasets acquired from WorldView-3 and GaoFen-2. Please find the code from the project page.
Liang-Jian Deng, Tianjing Zhang
ACM Multimedia3