EDBT 2026 Demo / reviewers in the wild / expert
Honghui Xu 0002
dblp:28/825-2
· DBLP profile ↗
35ranked-venue papers
9as first author
35since 2021 · last 2026
0000-0002-6213-2979ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 6 first-author · 25 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TRT: Harnessing Tensor Ring Transformer for Hyperspectral Image Super-ResolutionabstractDeep unfolding networks (DUNs) have recently emerged as a promising approach for hyperspectral image super-resolution (HSISR) by combining the benefits of nonlinear deep learning architectures with interpretable optimization techniques. Despite their advantages, current DUNs face significant challenges, particularly in approximating degradation matrices across both spatial and spectral dimensions, which results in complex and cumbersome model construction. By analyzing the difference between the upsampled low-resolution hyperspectral images (LRHS) and the true target image, we observed that the residual image exhibits strong sparsity, akin to noise. Leveraging this insight, we reformulate the HSISR problem as a robust principal component analysis (RPCA)-based denoising task, effectively eliminating the need for the complex approximation of spatial degradation matrix and its transpose. In addition, we introduce a Tensor Ring Transformer based on multilinear products as the prior term, wherein tokens are mapped to a tensor ring factor domain and the traditional dot product is replaced with a multilinear tensor ring product. This significantly reduces the computational complexity of the Transformer model, from \( \mathcal{O}(N^2d) \) to \( \mathcal{O}(Nr^2) \), with \( r Honghui Xu 0002, Yubin Gu, Yueqian Quan, Chuangjie Fang, Hong Qiu, Jianwei Zheng 0001 |
AAAI | 1 |
| 2026 | Subspace-frequency regularization for hyperspectral image super-resolution
Chuangjie Fang, Yan Li 0083, Hong Qiu, Honghui Xu 0002, Jianwei Zheng 0001 |
Knowl. Based Syst. | 5 |
| 2026 | Information-coupled MRI acceleration via multi-modal mapping and progressive masking
Jiawei Jiang 0002, Honghui Xu 0002, Jianwei Zheng 0001 |
Pattern Recognit. | 3 |
| 2026 | Arbitrary-Scale Fusion Operator for High-Resolution Hyperspectral ImagingabstractFor high-resolution hyperspectral (HrHs) imaging, spatial-spectral fusion offers a promising alternative to expensive equipment. However, retraining multiple models for varied scaling factors is currently unavoidable, costing extra computational resource and human labor. To address this issue, we propose Arbitrary-scale Fusion Operator (AFO), a lightweight solution for HrHs fusion given arbitrary scalings, turning the retraining strategy into “training-free” ones. Specifically, AFO treats low-resolution hyperspectral (LrHs) images and high resolution multispectral (HrMs) images as light-wise degraded functions within the spectrum, which are initially embedded into a high-dimensional space to simulate the original light signals, tapping the potential of enriched prior learning. Then, a flow of kernel integration (KI) is meticulously crafted, followed by a rival step of dimension reduction for HrHs generation. For a well-behaved KI computation, an Attention-Driven Convolution Integration (ADCI) is engineered to restore the broken discretization invariance derived by convolutions, yet with the locally inductive bias preserved. In addition, we propose an Implicit Neural Functional Integration (INFI) to achieve cross domain interaction of spatial degradation functions, followed by the use of Galerkin-type Integration (GI) as a decoder to handle high-frequency information. Finally, the bonded activation functions are improved for the principle of continuous-discrete equivalence. Extensive experiments validate the superiority of our approach over cutting-edge methods. Notably, our proposal holds significantly better generalization on arbitrary scaling factors, yet requires only 0.07M parameters. Honghui Xu 0002, Wei Li 0034, Jiawei Jiang 0002, Zhi Liu 0009, Jianwei Zheng 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | DreamText: High Fidelity Scene Text SynthesisabstractScene text synthesis involves rendering specified texts onto arbitrary images. Current methods typically formulate this task in an end-to-end manner but lack effective character-level guidance during training. Besides, their text encoders, pre-trained on a single font type, struggle to adapt to the diverse font styles encountered in practical applications. Consequently, these methods suffer from character distortion, repetition, and absence, particularly in polystylistic scenarios. To this end, this paper proposes DreamText for high-fidelity scene text synthesis. Our key idea is to reconstruct the diffusion training process, introducing more refined guidance tailored to this task, to expose and rectify the model’s attention at the character level and strengthen its learning of text regions. This transformation poses a hybrid optimization challenge, involving both discrete and continuous variables. To effectively tackle this challenge, we employ a heuristic alternate optimization strategy. Meanwhile, we jointly train the text encoder and generator to comprehensively learn and utilize the diverse font present in the training dataset. This joint training is seamlessly integrated into the alternate optimization process, fostering a synergistic relationship between learning character embedding and re-estimating character attention. Specifically, in each step, we first encode potential character-generated position information from cross-attention maps into latent character masks. These masks are then utilized to update the representation of specific characters in the current step, which, in turn, enables the generator to correct the character’s attention in the subsequent steps. Both qualitative and quantitative results demonstrate the superiority of our method to the state of the art. Our project page is here. Honghui Xu 0002, Cheng Jin 0001 |
CVPR | 3 |
| 2025 | Gradient Selection Tuning via Information BottleneckabstractPre-trained visual models enjoy strong representations, yet suffer from massive parameters to be shifted in downstream practices. Many parameter-efficient fine-tuning methods have been proposed, mostly requiring only 1% additional parameters to achieve comparable results. However, current solutions either consider all feature channels equally or detect saliencies with individual layer, leading to many redundancies reserved. To address current issues, this paper proposes a new parameter fine-tuning method named “Gradient Selection Tuning” (GST), which leverages gradients that are capable of capturing the cascading effects across successive channels. Instead of saliency detection, we turn to compress the redundancies for channel selection, since the computed gradient values enjoy much lower mutual information. With GST facilitated, we further elaborate an Information-Guided Adapter following information bottleneck theory, effectively performing parameter compression yet with task-specific features preserved. Experimental results demonstrate that our method outperforms the baseline methods by adding only 0.075M parameters to ViT-B backbone. On domain generalization, our proposal also enjoys strong performance in low-parameter scenarios. Xiaoxu Lin, Wei Li 0034, Ni Xu, Honghui Xu 0002, Jianwei Zheng 0001 |
ECAI | 5 |
| 2025 | Controllable Face Inpainting via Pseudo-Style EmbeddingabstractImage inpainting, a critical facet of computer vision, is in full bloom accompanied by the rapid innovation of convolution neural networks and transformers, revolutionizing the practical management of abnormity disposal, image editing, etc. Of these applications, face inpainting is more challenging due to the higher demand for semantic accuracy in key regions such as eyes and nose. Classical face inpainting methods are celebrated for their fast generation speed and refined texture details. However, they often lack the level of controllability required for complex tasks. In contrast, existing multi-modal controllable inpainting techniques offer enhanced guidance through image-text integration but tend to be time-consuming and produce suboptimal texture refinement. To address these limitations, we propose the Multi-modal Pseudo-style Embedded Transformer (MPET), a novel and efficient multi-modal inpainting algorithm that seamlessly integrates the strengths of both approaches, achieving state-of-the-art performance. Specifically, edge completion facilitates a cost-efficient and simple bridging of the contour continuity. Multi-modal pseudo-style generation amalgamates the image-text modalities, successfully embedding text features within the visual vectors, thereby culminating in the formation of pseudo-style diagrams rich in diverse attributes. On that basis, a controllable style-embedded siamese network is elaborated, effectively orchestrating the interaction among style attributes while ensuring high-precision pixel infusion. Extensive experiments on public datasets demonstrate the superiority of our approach through both quantitative and qualitative evaluations, highlighting its potential to advance the field of face inpainting. Jiawei Jiang 0002, Yueqian Quan, Honghui Xu 0002, Jianwei Zheng 0001 |
ECAI | 4 |
| 2025 | Pipeline-Centered Neighboring Network for Deep Unfolding PansharpeningabstractPansharpening technique is dedicated to enriching the spatial details of low-resolution multispectral images (LRMS) under the guidance of a panchromatic (PAN) image. With the guarantee of promising results, Transformer-based methods have enjoyed a high reputation in this field. However, to reduce computational cost, existing solutions typically divide images into smaller, independent windows, which often weakens inter-window and channel-wise interactions as well as leads to unsmooth edges. To address these issues, we first formulate the pansharpening task as a variational optimization problem, and subsequently solve its data and prior subproblems alternately through an unrolling algorithm. In the prior extractor, we propose a Pipeline-Centered Neighboring Attention (PCNA), which holistically allows all pixels to share the same attention span while fully leveraging channel dependencies, thereby significantly improving the capability to process multispectral images. Moreover, a Multi-Scale Channel-Aware (MSCA) module is designed to capture the edges and structural details. Finally, by sequentially integrating the data and prior modules at each iteration stage, we unroll the iterations into a stage-wise unfolding network. Extensive experiments on three satellite datasets demonstrate the effectiveness and efficiency of our proposal compared to cutting-edge methods. Yan Li 0083, Qiuju Chen, Chuangjie Fang, Ni Xu, Honghui Xu 0002, Jianwei Zheng 0001 |
ICASSP | 5 |
| 2025 | Laboring on Less Labors: RPCA Paradigm for Pan-Sharpening
Honghui Xu 0002, Chuangjie Fang, Jianwei Zheng 0001 |
ICCV | 1 |
| 2025 | C3S3: Complementary Competition and Contrastive Selection for Semi-Supervised Medical Image SegmentationabstractFor the immanent challenge of insufficiently annotated samples in the medical field, semi-supervised medical image segmentation (SSMIS) offers a promising solution. Despite achieving impressive results in delineating primary target areas, most current methodologies struggle to precisely capture the subtle details of boundaries. This deficiency often leads to significant diagnostic inaccuracies. To tackle this issue, we introduce C3S3, a novel semi-supervised segmentation model that synergistically integrates complementary competition and contrastive selection. This design significantly sharpens boundary delineation and enhances overall precision. Specifically, we develop an Outcome-Driven Contrastive Learning module dedicated to refining boundary localization. Additionally, we incorporate a Dynamic Complementary Competition module that leverages two high-performing sub-networks to generate pseudo-labels, thereby further improving segmentation quality. The proposed C3S3 undergoes rigorous validation on two publicly accessible datasets, encompassing the practices of both MRI and CT scans. The results demonstrate that our method achieves superior performance compared to previous cutting-edge competitors. Especially, on the 95HD and ASD metrics, our approach achieves a notable improvement of at least 6%, highlighting the significant advancements. The code is available at https://github.com/Y-TARL/C3S3. Jiaying He, Yitong Lin, Honghui Xu 0002, Jianwei Zheng 0001 |
ICME | 4 |
| 2025 | Enhancing Object Coherence in Layout-to-Image SynthesisabstractLayout-to-image synthesis aims to generate complex scenes, where users require fine control over the layout of the objects. However, it remains challenging to control the object coherence, including semantic coherence (e.g., the cat looks at the flowers or not) and physical coherence (e.g., the hand and the racket should not be misaligned). In this paper, we propose a novel diffusion model with effective global semantic fusion (GSF) and self-similarity feature enhancement modules to guide the object coherence for this task. For semantic coherence, we argue that the image caption contains rich information for defining the semantic relationship within the objects in the images. Instead of simply employing cross-attention between captions and latent images, which addresses the highly relevant layout restriction and semantic coherence requirement separately and thus leads to unsatisfying results shown in our experiments, we develop GSF to fuse the supervision from the layout restriction and semantic coherence requirement and exploit it to guide the image synthesis process. Moreover, to improve the physical coherence, we develop a Self-similarity Coherence Attention (SCA) module to explicitly integrate local contextual physical coherence relation into each pixel’s generation process. Specifically, we adopt a self-similarity map to encode the physical coherence restrictions and employ it to extract coherent features from text embedding. Extensive experiments demonstrate the superiority of our proposed method. Code is available at here. Changhai Zhou, Honghui Xu 0002 |
ICME | 3 |
| 2025 | Collaborative Cross-Complementary Unfolding Network for Pan-sharpening Remote Sensing ImageabstractDue to the acquisition limitations of physical devices, pansharpening serves as a computational alternative, enhancing spatial details in low-resolution hyperspectral images with the guidance of corresponding panchromatic images. By leveraging the benefits of nonlinear network architectures and interpretable optimization schemes, deep unfolding networks (DUNs) have shed new light on pansharpening. However, current DUNs lack a dedicated design for both estimating the degradation matrices and extracting intricate information from the proximal operator. To address these challenges, we propose a novel Collaborative Cross-Complementary Unfolding Network (C3U), which is organized into two main steps: customized multi-scale convolution estimation (MSCE) and a data-driven prior extractor. In the MSCE step, the spatial and spectral degradation matrices are individually adapted through multiscale treatment and point convolution operations. Specifically, the overall estimation undergoes an end-to-end iterative block, allowing for adaptive modeling of complex spatial and spectral structures. Within the prior extractor, a cross-complementary attention mechanism is proposed to enable iterative information interaction between global and local Transformers, capturing holistic features and enhancing inductive capacity. Additionally, a collaborative scale-aware-channel mechanism is designed to enlarge the receptive field and capture multiscale channel features in a lightweight manner. More importantly, the principle of collaborative cross-complementary (CCC) permeates all the sub-assemblies, ensuring a desirable information flow. Experimental results on multiple remote sensing datasets demonstrate the superiority of the proposed method over previous state-of-the-art (SOTA) techniques, achieving a 0.8 dB PSNR gain on the GF-2 dataset. Honghui Xu 0002, Yan Li 0083, Yutao Jia, Chuangjie Fang, Jianwei Zheng 0001 |
ICMR | 1 |
| 2025 | SpecSolver: Solving Spatial-Spectral Fusion via Semantic TransformerabstractBy clustering pixels with locally similar values, superpixel-based approaches have shown great potential in processing hyperspectral images (HSI), thereby reducing the computational burden associated with large spatial dimensions. However, specific for spatial-spectral fusion (SSF), superpixel segmentation is inherently non-differentiable and irreversible; hence it is inapplicable. To address the issues, we propose a semantic transformer-based solver, namely SpecSolver, which is basically inspired by the benefits of superpixel-based approaches, yet with the inner mechanism completely improved. The core idea lies in learning the intrinsic semantic states of HSIs hidden behind discretized pixel representations. Specifically, we propose a new Semantic-Attention to adaptively split the image domain into a series of learnable slices of flexible shapes, where image pixels under similar semantic states will be ascribed to the same slice. By calculating attention to the Semantic-Superpixel tokens encoded from slices, SpecSolver can effectively capture intricate semantic correlations from the vast number of pixels, which also empowers the solver with an endogenous capacity for modeling different magnification scales and allows for efficient computation in linear complexity. On that basis, we elaborate a SpatialNet module, which extracts multiscale local spectral information, and a FreqNet module, which supplements global information, capturing subtle details and variations across different spectra. Experiments on two benchmark SSF datasets verify the state-of-the-art (SOTA) performance of the proposed method, both visually and quantitatively. Also, ablation studies validate the mentioned contributions. Wei Li 0034, Honghui Xu 0002, Jiawei Jiang 0002, Jianwei Zheng 0001 |
ACM Multimedia | 3 |
| 2025 | Arbitrary-scale Fusion Neural OperatorabstractSpatial-spectral fusion offers a promising alternative to expensive equipment in high-resolution hyperspectral (HrHs) imaging. However, training separate models for different scaling factors remains costly. To address this, we propose the Arbitrary-scale Fusion Neural Operator (AFNO), a lightweight solution for HrHs fusion across arbitrary scalings. Instead of entities, AFNO treats low-resolution hyperspectral (LrHs) and high-resolution multispectral (HrMs) images as functions and performs meticulously designed integrations as the mapping operator. The key components include Attention-Driven Convolution Integration (ADCI) to restore discretization invariance disrupted by convolutions, Implicit Neural Functional Integration (INFI) for cross-domain interaction of spatial degradations, and Galerkin-type Integration as a decoder for high-frequency details. Additionally, the bonded activation opeartor are improved for the principle of continuous-discrete equivalence. Extensive experiments validate the superiority of our approach over cutting-edge methods. Notably, AFNO holds significantly better generalization on arbitrary scaling factors, yet requiring only 0.07M parameters. Wei Li 0034, Honghui Xu 0002, Jiawei Jiang 0002, Zhi Liu 0009, Jianwei Zheng 0001 |
ACM Multimedia | 3 |
| 2025 | Nonlinear Learnable Triple-Domain Transform Tensor Nuclear Norm for Hyperspectral Image Super-ResolutionabstractTensor Nuclear Norm (TNN) has been widely employed as a regularization term for hyperspectral image super-resolution (HSISR). However, conventional TNN constraints based on Discrete Fourier Transform (DFT) often suffer from rank estimation biases and an inability to effectively capture complex spectral-spatial correlations, limiting their efficacy in HSISR. To address these challenges, we propose a Nonlinear Learnable Triple-domain (NLT) transform framework that integrates nonlinear transform, DFT, and self-learning adaptation. This multi-stage process promotes singular value concentration, improving low-rank approximation and rank estimation accuracy. Building upon this framework, we develop an NL-transform-oriented tensor product, a truncated singular value decomposition (TSVD) operation, and a novel tensor nuclear norm (NLTN) tailored for HSISR. By incorporating spectral subspace estimation and clustering-based patch grouping, our approach effectively leverages spatial-spectral correlations and non-local self-similarities, leading to enhanced reconstruction quality. To further mitigate singular value over-penalization, we introduce a logarithmic-based generalized NLTNN (GNLTN) and formulate an optimization strategy based on the alternating direction method of multipliers (ADMM). Extensive experiments demonstrate that our method significantly outperforms existing approaches in terms of fusion accuracy and visual fidelity, setting new benchmarks for hyperspectral image super-resolution. The code is available at https://github.com/xuhonghui96/GNLTN. Honghui Xu 0002, Yueqian Quan, Chuangjie Fang, Yan Li 0083, Jianwei Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | SyFormer: Structure-Guided Synergism Transformer for Large-Portion Image InpaintingabstractImage inpainting is in full bloom accompanied by the progress of convolutional neural networks (CNNs) and transformers, revolutionizing the practical management of abnormity disposal, image editing, etc. However, due to the ever-mounting image resolutions and missing areas, the challenges of distorted long-range dependencies from cluttered background distributions and reduced reference information in image domain inevitably rise, which further cause severe performance degradation. To address the challenges, we propose a novel large-portion image inpainting approach, namely the Structure-Guided Synergism Transformer (SyFormer), to rectify the discrepancies in feature representation and enrich the structural cues from limited reference. Specifically, we devise a dual-routing filtering module that employs a progressive filtering strategy to eliminate invalid noise interference and establish global-level texture correlations. Simultaneously, the structurally compact perception module maps an affinity matrix within the introduced structural priors from a structure-aware generator, assisting in matching and filling the corresponding patches of large-proportionally damaged images. Moreover, we carefully assemble the aforementioned modules to achieve feature complementarity. Finally, a feature decoding alignment scheme is introduced in the decoding process, which meticulously achieves texture amalgamation across hierarchical features. Extensive experiments are conducted on two publicly available datasets, i.e., CelebA-HQ and Places2, to qualitatively and quantitatively demonstrate the superiority of our model over state-of-the-arts. Yuchao Feng, Honghui Xu 0002, Chuanmeng Zhu, Jianwei Zheng 0001 |
AAAI | 3 |
| 2024 | HMNet: Hierarchical Microscale-Aware Network for Infrared Small Target DetectionabstractCompared to the natural image community, infrared target detection suffers more challenges due to the severely tiny and low-contrast objects, especially in cases with obscuration from clutter and noise. The traditional solutions are susceptible to noise interference, which yields suboptimal performance lacking of contour and texture details. Meanwhile, due to the spatial invariance of convolutional layers, most deep learning-based methods locate small targets loosely during feature extraction, leading to serious omissions. To address these limitations, we propose a hierarchical microscale-aware network (HMNet) following an encoder-decoder structure that is mainly equipped with two novel modules: the holistic attention-aware (HAA) module and the scale-aware adaptive extraction (SAE) module. HAA integrates local and global cues via self-attention, depthwise separable convolutions, and dilated convolutions, which hammers at enhancing target features and ensuring accurate localization. As a complement, SAE employs multi-scale features and spatial-channel attention to acquire richer texture details while reducing background noise. The experiments on public datasets demonstrate that our method achieves state-of-the-art performance. Yueqian Quan, Honghui Xu 0002, Yidong Yan, Jianwei Zheng 0001 |
ICASSP | 2 |
| 2024 | Robust Principal Component Analysis via High-Order Self-Learning Transform Tensor Nuclear NormabstractIn recent studies, tensor singular value decomposition (TSVD) within the high-order (Ho) algebra has shed light on solving the Tensor Robust Principal Component Analysis (TRPCA) problem. However, the utilization of fixed or data-independent transformations in HoTSVD may result in suboptimal outcomes. To overcome this limitation, we propose a self-learning TSVD method that rectifies computational inefficiencies and learns a lossless transformation, inducing a lower average-rank tensor. This involves multiplying learnable semi-orthogonal matrices obtained through Tucker compression with the original tensor along all modes, resulting in a core tensor with enhanced inherent low rankness and new self-learning transform matrices. The semi-orthogonal transforms, acting as a crucial building block, enhance spatial low-rankness, facilitating the resolution of smaller-scale problems and the design of efficient algorithms. Additionally, a reweighting Schatten-p scheme is integrated into the self-learning HoTSVD to understand global low-rank correlations, offering an effective numerical solution. Finally, we develop an alternating direction method of multipliers (ADMM)-based algorithm as a solver. Experimental results on Light Field Images (LFI), showcase the superiority of our proposed method over previous state-of-the-art approaches. Honghui Xu 0002, Yueqian Quan, Chuangjie Fang, Jianwei Zheng 0001 |
ICME | 1 |
| 2024 | Multi-dimensional visual data completion via weighted hybrid graph-Laplacian
Jiawei Jiang 0002, Yile Xu, Honghui Xu 0002, Guojiang Shen, Jianwei Zheng 0001 |
Signal Process. | 3 |
| 2024 | ORSI Salient Object Detection via Progressive Semantic Flow and Uncertainty-Aware RefinementabstractWith the prosperity of deep learning techniques, salient object detection in remote sensing images (RSI-SOD) is concomitantly in full flourishing. However, due to the inherent challenges such as uncertainty in object quantities and scales, cluttered backgrounds, and blurred edges arising from shadows, most current approaches struggle for salient feature learning with the aid of heavy model architecture, yet often result in barely satisfactory performance. Some methods compromise model complexity to improve efficiency, albeit with significantly degraded results. To earn a satisfactory balance of efficacy and efficiency, we propose a new network for RSI-SOD, namely SFANet, based on progressive semantic flow and uncertainty-aware refinement. Specifically, we design a global semantic enhancement block (GSEB) to reduce background interference and accurately localize salient objects of varying quantities and scales, which further consists of three modularized components, i.e., semantic extraction module (SEM), interscale fusion module (IFM), and deep semantic graph-inference module (DSGM). SEM together with IFM contributes to the effective aggregation of multi-scale contexts by extracting fused and progressive semantic cues. DSGM performs semantic inference to better localize salient objects with irregularities in scale and topological structure. Furthermore, we present an uncertainty-aware refinement module (URM) to recognize salient objects in cluttered backgrounds and effectively suppress shadows. Extensive experiments are conducted on three RSI-SOD datasets, from which superior results can be achieved by our SFANet, outperforming the other cutting-edge methods. The code is available at https://github.com/ZhengJianwei2/SFANet. Yueqian Quan, Honghui Xu 0002, Renfang Wang, Qiu Guan, Jianwei Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Cascade-Transform-Based Tensor Nuclear Norm for Hyperspectral Image Super-ResolutionabstractRecent advancements in tensor nuclear norm (TNN) have led to promising solutions for hyperspectral image super-resolution (HSISR), which produces enriched outputs by fusing low-resolution hyperspectral images (LRHSIs) with high-resolution multispectral images (HRMSIs). However, current TNN, mainly reliant on the discrete Fourier transform (DFT), still suffers from mirroring boundary effects and singleton domain limitation. As a relief, we propose cascade-transform-based tensor nuclear norm (CTNN) with two variants for HSISR, featuring new definitions and algebraic structures for tensor product and TNN operations. The first variant processes tubal elements derived from DFT as inputs in the discrete cosine transform (DCT) domain, allowing for more nuanced feature extraction. The second learns adaptive matrices from the data in each iteration update and links them with a fixed DFT matrix to dynamically update the transform domain, preventing rank estimation bias. Furthermore, the nonconvex form of the proposed CTNN is applied to three modes of each spectral subspace similarity cube, termed log-sum-based full-scale CTNN (LFCTNN), capturing the global low-rank structure of LRHSI and the nonlocal similarities present in HRMSI. Experimental evaluations on various remote sensing datasets indicate that our approach exceeds existing state-of-the-art methods. The code is available athttps://github.com/xuhonghui96/LFCTNN. Honghui Xu 0002, Chuangjie Fang, Yilin Ge, Yubin Gu, Jianwei Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Contrastive Attention-guided Multi-level Feature Registration for Reference-based Super-resolutionabstractGiven low-quality input and assisted by referential images, reference-based super-resolution (RefSR) strives to enlarge the spatial size with the guarantee of realistic textures, for which sophisticated feature-matching strategies are naturally demanded. However, the miserable transformation gap between inputs and references, e.g., texture rotation and scaling within patches, often yields distorted textures and terrible ghosting artifacts, which seriously hampers the visual senses and their further investigation. To circumvent this challenge, we propose a contrastive attention-guided multi-level feature registration for RefSR, explicitly tapping the potential of interacting between inputs and references. Specifically, we develop a multi-level feature warping scheme, involving patch-level coarse feature swapping and pixel-level deformable alignment, to model generalized spatial transformation correspondences steered by contrastive attention. Notably, a spatial registration module is embedded for further calibration against the potential misalignment issue and inter-feature distribution difference. In addition, aiming at suppressing the impacts of irrelevant or superfluous information on cross-scale features, we incorporate a multi-residual feature fusion module to strive for visually plausible textures. Experimental results on four publicly available datasets demonstrate that our method outperforms most state-of-the-art approaches in terms of both efficiency and perceptual effectiveness. Jianwei Zheng 0001, Yu Liu 0151, Yuchao Feng, Honghui Xu 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Building Change Detection Using Cross-Temporal Feature Interaction NetworkabstractBuilding change detection of remote sensing images is in full flourishing accompanied by the prosperity of convolutional neural networks. For spatial-temporal context modeling, existing solutions disregard the inter-image interactions, albeit their positive contribution to the acquisition of differences. To fill the gap, we propose a cross-temporal feature interaction network to effectively derive the change representations. Specifically, we propose a linearized cross-attention, which motivates each counterpart to glimpse the representation of another image while preserving its own features. In addition, to circumvent the misalignment caused by step-down sampling in the backbone, we introduce multi-level feature alignment using learnable affine transformation and stepwise aggregation. Based on a naive backbone (ResNet18) without sophisticated structures, our model outperforms other state-of-the-art methods on three datasets in terms of both efficiency and effectiveness. Yuchao Feng, Jiawei Jiang 0002, Honghui Xu 0002, Jianwei Zheng 0001 |
ICASSP | 3 |
| 2023 | Low-Dose CT Reconstruction Via Optimization-Inspired GANabstractMost research on Low-dose Computed Tomography (LDCT) reconstruction is designed as a black box, lacking controllability and interpretability. In this paper, a Proximal Linear ADMM framework-based Generative Adversarial Network (PLA-GAN) is proposed. Specifically, without loss of interpretability, channel attention blocks and NonLocal Sparse Attention (NLSA) modules are embedded into two regularizers respectively and iterated alternately, driving the network to cope with real and complex CT image degradation through a multi-scale and adaptive way. To further promote the visual quality, a discriminator containing NLSA module is also introduced. The comparisons with state-of-the-arts on the Mayo dataset validate the superiority of our proposed algorithm both numerically and visually. The advantages of generalizability and interpretability are also evident. Jiawei Jiang 0002, Yuchao Feng, Honghui Xu 0002, Jianwei Zheng 0001 |
ICASSP | 3 |
| 2023 | Compact Intertemporal Coupling Network for Remote Sensing Change DetectionabstractChange detection of multi-temporal remote sensing images is in full flourishing accompanied by the popularity and prosperity of deep learning. The prominent challenge lies in the crude distribution of the newly constructed and demolished changes, the interferences of massive irrelevant objects, and the spatial-temporal changes from the passage of time. For context modeling, existing solutions waste massive attention on task-irrelevant features, spotlighting insufficiently on the genuinely changed regions. To fill the gap, we propose a compact intertemporal coupling network (CICNet) to derive the change representations. Specifically, to underpin the interaction of spatial-temporal differences in a global perspective, we detach and innovate the solo-head self-attention into a lightweight intertemporal-attention, favorably bridging the intra-level features. In addition, to circumvent the misalignment imposed by spatial sampling, lightweight global channel- and spatial- attentions are globally incorporated for stepwise calibration between localization seduced low-level information and semantics abundant high-level features. Based on a naive backbone (ResNet18/34) without sophisticated structures, our model outperforms other state-of-the-art methods on four datasets in terms of both efficiency and effectiveness. Yuchao Feng, Honghui Xu 0002, Jiawei Jiang 0002, Jianwei Zheng 0001 |
ICME | 2 |
| 2023 | GA-HQS: MRI reconstruction via a generically accelerated unfolding approachabstractDeep unfolding networks (DUNs) are the foremost methods in the realm of compressed sensing MRI, as they can employ learnable networks to facilitate interpretable forward-inference operators. However, several daunting issues still exist, including the heavy dependency on first-order optimization algorithms, the insufficient information fusion mechanisms, and the limitation of capturing long-range relationships. To address the issues, we propose a Generically Accelerated Half-Quadratic Splitting (GA-HQS) algorithm that incorporates second-order gradient information and pyramid attention modules for the delicate fusion of inputs at the pixel level. Moreover, a multi-scale split transformer is also designed to enhance the global feature representation. Comprehensive experiments demonstrate that our method surpasses previous ones on single-coil MRI acceleration tasks. Jiawei Jiang 0002, Honghui Xu 0002, Yuchao Feng, Jianwei Zheng 0001 |
ICME | 3 |
| 2023 | CA-GAN: Object Placement via Coalescing Attention based Generative Adversarial NetworkabstractLearning to posit a foreground object over a background scene is an intriguing yet challenging problem, which frequently emerges in applications such as image editing and scene parsing. To date, most existing studies are fed up with knotty issues, including the deficiency of harnessing the interaction between the object and the scene, the astriction of involving little prior knowledge during training, etc. To break the shackles, we propose a novel end-to-end framework dubbed Coalescing Attention based Generative Adversarial Network (CA-GAN). Specifically, in our synthesizer, a feature polymerizer is designed to distill multi-scale information from both background and foreground. On that basis, a dual-branch coalescing attention module is proposed for a better exploration of the global feature-interaction relationships between object and scene. In addition, we add a supervised trail to learn the prior knowledge from the positive composite image, which further guides the synthesizer to discover a credible placement for the foreground object. With extensive experiments conducted on the OPA dataset, our proposal presents superiority in both rationality and diversity compared with other state-of-the-art methods. Our code is available at https://github.com/ZhengJianwei2/CA-GAN. Yuchao Feng, Honghui Xu 0002, Jianwei Zheng 0001 |
ICME | 4 |
| 2023 | A Lightweight Collective-attention Network for Change DetectionabstractChange detection of multi-temporal remote sensing images is mushrooming with the innovations of neural networks, whose daunting challenge lies in locating sporadically distributed spatial-temporal changes given sophisticated scenes and various imaging conditions. Unfortunately, instead of devoting full attention to changes, most existing solutions often expend unnecessary resources yet derive task-irrelevant features. To relieve this issue, we propose a collective-attention network, which enjoys lightweight model architecture yet guarantees high performance. Specifically, an inter-temporal collective-attention module is developed for efficient interaction of bi-temporal features, in which a shared attention distribution is derived via the multiplication of temporal-concatenated queries and spatial-subtracted keys. Additionally, we present a non-change consistency-constraint, enforcing a change-oriented attention distribution and a noise-suppressed treatment. With the learned interaction features, bi-temporal differences are captured simply using the operations of spatial absolute error and temporal concatenation. Finally, decoding multi-scale differences is accomplished by lightweight temporal self-attention and spatial self-attention. Experiments on four datasets demonstrate that our model achieves state-of-the-art performance, yet requires only 1.71M parameters and 1.98G FLOPs. Yuchao Feng, Yanyan Shao, Honghui Xu 0002, Jinshan Xu, Jianwei Zheng 0001 |
ACM Multimedia | 3 |
| 2023 | Tensor completion via hybrid shallow-and-deep priors
Honghui Xu 0002, Jiawei Jiang 0002, Yuchao Feng, Yiting Jin, Jianwei Zheng 0001 |
Appl. Intell. | 1 |
| 2023 | Nonlocal B-spline representation of tensor decomposition for hyperspectral image inpainting
Honghui Xu 0002, Yidong Yan, Jianwei Zheng 0001 |
Signal Process. | 1 |
| 2023 | Change Detection on Remote Sensing Images Using Dual-Branch Multilevel Intertemporal NetworkabstractChange detection (CD) of remote sensing (RS) images is mushrooming up accompanied by the on-going innovation of convolutional neural networks (CNNs). Yet with the high-speed technology upgrade, the obstacle that identifies unbalanced variations in foreground–background categories still lies on the table, especially in cases with limited samples and massive interference such as seasonal turnover, illumination intensity, and building reformation. Moreover, to date, neither of the off-the-shelf methods probes the feasibility of direct interaction between bitemporal images before accessing difference features. In this article, we propose a dual-branch multilevel intertemporal network (DMINet) to efficiently and effectively derive the change representations. Specifically, by unifying self-attention (SelfAtt) and cross-attention (CrossAtt) in a single module, we present an intertemporal joint-attention (JointAtt) block to steer the global feature distribution of each input, motivating information coupling between intralevel representations and meanwhile suppressing the task-irrelevant interferences. In addition, centering more on the detection of difference features, a reliable architecture is designed by spotlighting two concerns, i.e., the difference acquisition using subtraction and concatenation as well as the multilevel difference aggregation using incremental feature alignment. Based on a naive backbone without sophisticated structures, i.e., ResNet18, our model outperforms other state-of-the-art (SOTA) methods on four CD datasets, especially in cases with rarely samples. Moreover, the achievement is attained with light overheads. Yuchao Feng, Jiawei Jiang 0002, Honghui Xu 0002, Jianwei Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | ORSI Salient Object Detection via Bidimensional Attention and Full-Stage Semantic GuidanceabstractThe application of optical remote sensing images (ORSIs) is prevalent in many fields. Accordingly, ORSI-oriented salient object detection (SOD) has attracted more attention in recent years. However, yet many previously proposed methods present appealing performance in natural scene images (NSIs), they are difficult to be directly extended to remote sensing images due to the more complex scenes, such as blended backgrounds and diversiform topological shapes. Most specifically designed models often fail to achieve satisfactory results due to the weak usage of edge information and the ignorance of attention loss. Besides, computational inefficiency often causes poor applicability. To solve these problems, we propose a new model, namely, Bidimensional Attention and Full-stage Semantic Guidance Network (BAFS-Net), containing an edge guidance branch and a mainstream detection branch. Concretely, edge guidance generates boundary information, in which supervision with border labels is imposed to highlight the salient regions and plays a complementary role on the main branch. The mainstream detection branch involves two important components, i.e., bidimensional attention modules (BAMs) and semantic-guided fusion modules (SGFMs). Between these two, BAM uniformly assembles channel and spatial attention in an efficient and rational manner, addressing the open issue of dimensionwisely attention computation. SGFM hammers at the fusion of high-level features and low-level features. Moreover, the semantic maps are employed to interact with SGFM in full stages. Our approach surpasses most state-of-the-art RSI-SOD methods proposed in recent years, with respect to the accuracy, parameter size, computational cost, and floating point operations per second (FLOPS). The code is available athttps://github.com/ZhengJianwei2/BAFS-Net. Yubin Gu, Honghui Xu 0002, Yueqian Quan, Jianwei Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Fast Tensor Nuclear Norm for Structured Low-Rank Visual InpaintingabstractLow-rank modeling has achieved great success in visual data completion. However, the low-rank assumption of original visual data may be in approximate mode, which leads to suboptimality for the recovery of underlying details, especially when the missing rate is extremely high. In this paper, we go further by providing a detailed analysis about the rank distributions in Hankel structured and clustered cases, and figure out both non-local similarity and patch-based structuralization play a positive role. This motivates us to develop a new Hankel low-rank tensor recovery method that is competent to truthfully capture the underlying details with sacrifice of slightly more computational burden. First, benefiting from the correlation of different spectral bands and the smoothness of local spatial neighborhood, we divide the visual data into overlapping 3D patches and group the similar ones into individual clusters exploring the non-local similarity. Second, the 3D patches are individually mapped to the structured Hankel tensors for better revealing low-rank property of the image. Finally, we solve the tensor completion model via the well-known alternating direction method of multiplier (ADMM) optimization algorithm. Due to the fact that size expansion happens inevitably in Hankelization operation, we further propose a fast randomized skinny tensor singular value decomposition (rst-SVD) to accelerate the per-iteration running efficiency. Extensive experimental results on real world datasets verify the superiority of our method compared to the state-of-the-art visual inpainting approaches. Honghui Xu 0002, Jianwei Zheng 0001, Xiaomin Yao, Yuchao Feng, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | ICIF-Net: Intra-Scale Cross-Interaction and Inter-Scale Feature Fusion Network for Bitemporal Remote Sensing Images Change DetectionabstractChange detection (CD) of remote sensing (RS) images has enjoyed remarkable success by virtue of convolutional neural networks (CNNs) with promising discriminative capabilities. However, CNNs lack the capability of modeling long-range dependencies in bitemporal image pairs, resulting in inferior identifiability against the same semantic targets yet with varying features. The recently thriving Transformer, on the contrary, is warranted, for practice, with global receptive fields. To jointly harvest the local-global features and circumvent the misalignment issues caused by step-by-step downsampling operations in traditional backbone networks, we propose an intra-scale cross-interaction and inter-scale feature fusion network (ICIF-Net), explicitly tapping the potential of integrating CNN and Transformer. In particular, the local features and global features, respectively, extracted by CNN and Transformer, are interactively communicated at the same spatial resolution using a linearized Conv Attention module, which motivates the counterpart to glimpse the representation of another branch while preserving its own features. In addition, with the introduction of two attention-based inter-scale fusion schemes, including mask-based aggregation and spatial alignment (SA), information integration is enforced at different resolutions. Finally, the integrated features are fed into a conventional change prediction head to generate the output. Extensive experiments conducted on four CD datasets of bitemporal (RS) images demonstrate that our ICIF-Net surpasses the other state-of-the-art (SOTA) approaches. Yuchao Feng, Honghui Xu 0002, Jiawei Jiang 0002, Jianwei Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Tensor completion using patch-wise high order Hankelization and randomized tensor ring initialization
Jianwei Zheng 0001, Honghui Xu 0002, Yuchao Feng, Peijun Chen, Shengyong Chen |
Eng. Appl. Artif. Intell. | 3 |