VLDB 2026 Research / reviewers in the wild / expert
Jianwei Zheng 0001
dblp:60/4818-1
· DBLP profile ↗
77ranked-venue papers
22as first author
63since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 7 first-author · 39 since 2021Artificial intelligence and machine learning · 31 · 10 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 5 first-author · 13 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TRT: Harnessing Tensor Ring Transformer for Hyperspectral Image Super-ResolutionabstractDeep unfolding networks (DUNs) have recently emerged as a promising approach for hyperspectral image super-resolution (HSISR) by combining the benefits of nonlinear deep learning architectures with interpretable optimization techniques. Despite their advantages, current DUNs face significant challenges, particularly in approximating degradation matrices across both spatial and spectral dimensions, which results in complex and cumbersome model construction. By analyzing the difference between the upsampled low-resolution hyperspectral images (LRHS) and the true target image, we observed that the residual image exhibits strong sparsity, akin to noise. Leveraging this insight, we reformulate the HSISR problem as a robust principal component analysis (RPCA)-based denoising task, effectively eliminating the need for the complex approximation of spatial degradation matrix and its transpose. In addition, we introduce a Tensor Ring Transformer based on multilinear products as the prior term, wherein tokens are mapped to a tensor ring factor domain and the traditional dot product is replaced with a multilinear tensor ring product. This significantly reduces the computational complexity of the Transformer model, from \( \mathcal{O}(N^2d) \) to \( \mathcal{O}(Nr^2) \), with \( r Honghui Xu 0002, Yubin Gu, Yueqian Quan, Chuangjie Fang, Hong Qiu, Jianwei Zheng 0001 |
AAAI | 7 |
| 2026 | LaViSE: Language-aware Vision Scale Enhancement for Referring Remote Sensing Image SegmentationabstractReferring Remote Sensing Image Segmentation (RRSIS) aims to segment target objects in aerial imagery based on natural language expressions. Although recent multi-scale feature aggregation methods have improved cross-modal alignment, and existing approaches still struggle with accurate localization and segmentation across scales because interactions between visual scales and language are not sufficiently modeled. To address these challenges, we propose a SAM-based framework termed LaViSE, which incorporates two key modules : the Language-Guided Hierarchical Fusion (LGHF) module integrates cross-modal features at multiple scales for precise object localization by injecting spatial coordinates into visual representations and combining visual features with aligned features; the Language-Attentive Scale-Unified Aggregation (LASA) module globally merges multi-level features while maintaining spatial consistency and boundary fidelity, and further preserves fine-grained structural details through and language-conditioned feature recalibration across scales, which is crucial for segmenting densely distributed and scale-varying targets in complex remote sensing imagery. Experiments on the widely used RefSegRS and RRSIS-D benchmarks demonstrate that LaViSE consistently outperforms state-of-the-art methods, particularly in challenging scenarios with small and densely distributed objects. Our code and pre-trained models will be released upon publication. Yan Li 0083, Zhouchao Fu, Shengjie Yang, Jianwei Zheng 0001 |
ICMR | 6 |
| 2026 | Subspace-frequency regularization for hyperspectral image super-resolution
Chuangjie Fang, Yan Li 0083, Hong Qiu, Honghui Xu 0002, Jianwei Zheng 0001 |
Knowl. Based Syst. | 6 |
| 2026 | Information-coupled MRI acceleration via multi-modal mapping and progressive masking
Jiawei Jiang 0002, Honghui Xu 0002, Jianwei Zheng 0001 |
Pattern Recognit. | 5 |
| 2026 | Lightweight interaction-attention network for colorectal polyp segmentation
Yueqian Quan, Jianwei Zheng 0001 |
Pattern Recognit. Lett. | 4 |
| 2026 | Arbitrary-Scale Fusion Operator for High-Resolution Hyperspectral ImagingabstractFor high-resolution hyperspectral (HrHs) imaging, spatial-spectral fusion offers a promising alternative to expensive equipment. However, retraining multiple models for varied scaling factors is currently unavoidable, costing extra computational resource and human labor. To address this issue, we propose Arbitrary-scale Fusion Operator (AFO), a lightweight solution for HrHs fusion given arbitrary scalings, turning the retraining strategy into “training-free” ones. Specifically, AFO treats low-resolution hyperspectral (LrHs) images and high resolution multispectral (HrMs) images as light-wise degraded functions within the spectrum, which are initially embedded into a high-dimensional space to simulate the original light signals, tapping the potential of enriched prior learning. Then, a flow of kernel integration (KI) is meticulously crafted, followed by a rival step of dimension reduction for HrHs generation. For a well-behaved KI computation, an Attention-Driven Convolution Integration (ADCI) is engineered to restore the broken discretization invariance derived by convolutions, yet with the locally inductive bias preserved. In addition, we propose an Implicit Neural Functional Integration (INFI) to achieve cross domain interaction of spatial degradation functions, followed by the use of Galerkin-type Integration (GI) as a decoder to handle high-frequency information. Finally, the bonded activation functions are improved for the principle of continuous-discrete equivalence. Extensive experiments validate the superiority of our approach over cutting-edge methods. Notably, our proposal holds significantly better generalization on arbitrary scaling factors, yet requires only 0.07M parameters. Honghui Xu 0002, Wei Li 0034, Jiawei Jiang 0002, Zhi Liu 0009, Jianwei Zheng 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | Breaking Information Isolation: Accelerating MRI via Inter-sequence Mapping and Progressive MaskingabstractDeep unfolding network (DUN) has shed new light on multi-sequence MRI reconstruction, providing both high interpretability and acceptable performance. However, current approaches still suffer from the plight of information isolation, i.e., learning features of multi-suquences individually and leaving the mask departed from model updating. In this work, we propose a new unfolding solution, namely Information-coupled MRI Acceleration (IMA), to address the isolation issue. Concretely, two specific mechanisms are presented. On the one hand, the latent connections across different sequences are explicitly molded via two auxiliary matrices. While the first matrix is meticulously engineered to assemble the spatial details, the second one hammers at capturing the depth information conditioned on the enriched channels. On the other hand, following a deep analysis on the non-uniform distribution in low- and high-frequency components of the given mask, we elaborate a new unfolding flow using a progressive masking scheme, featuring a dilation-contraction mechanism during forward propagation of successive stages. Massive experiments are conducted under various sampling patterns and acceleration rates, whose results demonstrate that, without any sophisticated architectures, our IMA outperforms the current cutting-edge methods both visually and numerically. Jianwei Zheng 0001, Xiaomin Yao, Guojiang Shen, Wei Li 0034, Jiawei Jiang 0002 |
AAAI | 1 |
| 2025 | LMO: Linear Mamba Operator for MRI ReconstructionabstractInterpretability and consistency have long been crucial factors in MRI reconstruction. While interpretability has been significantly innovated with the emerging deep unfolding networks, current solutions still suffer from inconsistency issues and produce inferior anatomical structures. Especially in out-of-distribution cases, e.g., when the acceleration rate (AR) varies, the generalization performance is often catastrophic. To counteract the dilemma, we propose an innovative Linear Mamba Operator (LMO) to ensure consistency and generalization, while still enjoying desirable interpretability. Theoretically, we argue that mapping between function spaces, rather than between signal instances, provides a solid foundation of high generalization. Technically, LMO achieves a good balance between global integration facilitated by a state space model that scans the whole function domain, and local integration engaged with an appealing property of continuous-discrete equivalence. On that basis, learning holistic features can be guaranteed, tapping the potential of maximizing data consistency. Quantitative and qualitative results demonstrate that LMO significantly outperforms other state-of-the-arts. More importantly, LMO is the unique model that, with AR changed, achieves retraining performance without retraining steps. Codes are available at https://github.com/ZhengJianwei2/LMO. Wei Li 0034, Jiawei Jiang 0002, Kaihao Yu, Jianwei Zheng 0001 |
CVPR | 5 |
| 2025 | Gradient Selection Tuning via Information BottleneckabstractPre-trained visual models enjoy strong representations, yet suffer from massive parameters to be shifted in downstream practices. Many parameter-efficient fine-tuning methods have been proposed, mostly requiring only 1% additional parameters to achieve comparable results. However, current solutions either consider all feature channels equally or detect saliencies with individual layer, leading to many redundancies reserved. To address current issues, this paper proposes a new parameter fine-tuning method named “Gradient Selection Tuning” (GST), which leverages gradients that are capable of capturing the cascading effects across successive channels. Instead of saliency detection, we turn to compress the redundancies for channel selection, since the computed gradient values enjoy much lower mutual information. With GST facilitated, we further elaborate an Information-Guided Adapter following information bottleneck theory, effectively performing parameter compression yet with task-specific features preserved. Experimental results demonstrate that our method outperforms the baseline methods by adding only 0.075M parameters to ViT-B backbone. On domain generalization, our proposal also enjoys strong performance in low-parameter scenarios. Xiaoxu Lin, Wei Li 0034, Ni Xu, Honghui Xu 0002, Jianwei Zheng 0001 |
ECAI | 6 |
| 2025 | Controllable Face Inpainting via Pseudo-Style EmbeddingabstractImage inpainting, a critical facet of computer vision, is in full bloom accompanied by the rapid innovation of convolution neural networks and transformers, revolutionizing the practical management of abnormity disposal, image editing, etc. Of these applications, face inpainting is more challenging due to the higher demand for semantic accuracy in key regions such as eyes and nose. Classical face inpainting methods are celebrated for their fast generation speed and refined texture details. However, they often lack the level of controllability required for complex tasks. In contrast, existing multi-modal controllable inpainting techniques offer enhanced guidance through image-text integration but tend to be time-consuming and produce suboptimal texture refinement. To address these limitations, we propose the Multi-modal Pseudo-style Embedded Transformer (MPET), a novel and efficient multi-modal inpainting algorithm that seamlessly integrates the strengths of both approaches, achieving state-of-the-art performance. Specifically, edge completion facilitates a cost-efficient and simple bridging of the contour continuity. Multi-modal pseudo-style generation amalgamates the image-text modalities, successfully embedding text features within the visual vectors, thereby culminating in the formation of pseudo-style diagrams rich in diverse attributes. On that basis, a controllable style-embedded siamese network is elaborated, effectively orchestrating the interaction among style attributes while ensuring high-precision pixel infusion. Extensive experiments on public datasets demonstrate the superiority of our approach through both quantitative and qualitative evaluations, highlighting its potential to advance the field of face inpainting. Jiawei Jiang 0002, Yueqian Quan, Honghui Xu 0002, Jianwei Zheng 0001 |
ECAI | 6 |
| 2025 | Object-Level Control for Refined Structure and Appearance in Conditional Image SynthesisabstractRecent advances in pretrained diffusion models, particularly the FreeControl, have enabled fine-grained spatial control in text-to-image generation. However, FreeControl still suffers notable limitations in detail generation and appearance synthesis. With deep analysis performed, we reveal that inadequate feature representation in the early generation phases is the main cause of the insufficient structure elaboration. To address this issue, we dig into the temporal evolution of inverse attention features, then extract more expressive structure information as the input to the guidance function, ensuring formation integrity. Moreover, to tackle the surface degradation and ambiguity caused by the dual guidance of structure and appearance, we engineer the Adaptive Instance Normalization (AdaIN) mechanism into a latent space, rather than the typical feature space, during the intermediate generation stage. This improvement not only guarantees the close alignment between the generated image and structural reference but also significantly strengthens the appearance modeling capability and optimizes the texture representation of both foreground and background elements. Extensive experiments demonstrate that our proposal consistently outperforms existing baseline models across multiple metrics, including Self-sim, CLIP, and LPIPS. Both quantitative and qualitative results confirm that our approach achieves superior performance in terms of content consistency, visual quality, and detail preservation. Ni Xu, Wei Li 0034, Zhouchao Fu, Xiaoxu Lin, Jianwei Zheng 0001 |
ECAI | 6 |
| 2025 | Pipeline-Centered Neighboring Network for Deep Unfolding PansharpeningabstractPansharpening technique is dedicated to enriching the spatial details of low-resolution multispectral images (LRMS) under the guidance of a panchromatic (PAN) image. With the guarantee of promising results, Transformer-based methods have enjoyed a high reputation in this field. However, to reduce computational cost, existing solutions typically divide images into smaller, independent windows, which often weakens inter-window and channel-wise interactions as well as leads to unsmooth edges. To address these issues, we first formulate the pansharpening task as a variational optimization problem, and subsequently solve its data and prior subproblems alternately through an unrolling algorithm. In the prior extractor, we propose a Pipeline-Centered Neighboring Attention (PCNA), which holistically allows all pixels to share the same attention span while fully leveraging channel dependencies, thereby significantly improving the capability to process multispectral images. Moreover, a Multi-Scale Channel-Aware (MSCA) module is designed to capture the edges and structural details. Finally, by sequentially integrating the data and prior modules at each iteration stage, we unroll the iterations into a stage-wise unfolding network. Extensive experiments on three satellite datasets demonstrate the effectiveness and efficiency of our proposal compared to cutting-edge methods. Yan Li 0083, Qiuju Chen, Chuangjie Fang, Ni Xu, Honghui Xu 0002, Jianwei Zheng 0001 |
ICASSP | 6 |
| 2025 | Laboring on Less Labors: RPCA Paradigm for Pan-Sharpening
Honghui Xu 0002, Chuangjie Fang, Jianwei Zheng 0001 |
ICCV | 5 |
| 2025 | C3S3: Complementary Competition and Contrastive Selection for Semi-Supervised Medical Image SegmentationabstractFor the immanent challenge of insufficiently annotated samples in the medical field, semi-supervised medical image segmentation (SSMIS) offers a promising solution. Despite achieving impressive results in delineating primary target areas, most current methodologies struggle to precisely capture the subtle details of boundaries. This deficiency often leads to significant diagnostic inaccuracies. To tackle this issue, we introduce C3S3, a novel semi-supervised segmentation model that synergistically integrates complementary competition and contrastive selection. This design significantly sharpens boundary delineation and enhances overall precision. Specifically, we develop an Outcome-Driven Contrastive Learning module dedicated to refining boundary localization. Additionally, we incorporate a Dynamic Complementary Competition module that leverages two high-performing sub-networks to generate pseudo-labels, thereby further improving segmentation quality. The proposed C3S3 undergoes rigorous validation on two publicly accessible datasets, encompassing the practices of both MRI and CT scans. The results demonstrate that our method achieves superior performance compared to previous cutting-edge competitors. Especially, on the 95HD and ASD metrics, our approach achieves a notable improvement of at least 6%, highlighting the significant advancements. The code is available at https://github.com/Y-TARL/C3S3. Jiaying He, Yitong Lin, Honghui Xu 0002, Jianwei Zheng 0001 |
ICME | 5 |
| 2025 | Spatial-Spectral Fusion Neural OperatorabstractWith the rapid development of deep learning, spatial-spectral fusion (SSF) has emerged as an ideal alternative to traditional, costly hyperspectral image (HSI) acquisition methods. However, current solutions necessitate training and storing multiple models for different scaling factors. Besides, a meticulously designed network architecture to meet desirable performance often suffers from a severe computational burden. To counteract the dilemma, we propose SFNO, a lightweight spatial-spectral fusion neural operator for arbitrary-scale SSF. SFNO leverages approximation theory by embedding features from two degraded functions into a high-dimensional latent space, enabling efficient learning of basis functions. Kernel integration mechanisms are then used to approximate certain priors, followed by dimensionality reduction to generate high-resolution HSIs. Moreover, with the aid of discrete invariance property, we propose a new mechanism of progressive resampling (PR), which allows for the shrinkage of necessary spatial domain without any performance degradation. Extensive experiments on CAVE and Harvard datasets show that SFNO and its variant significantly improve performance, especially in out-of-domain fusion, requiring only 0.098M parameters and 0.966G FLOPs. Wei Li 0034, Jiawei Jiang 0002, Ni Xu, Yan Li 0083, Jianwei Zheng 0001 |
ICME | 6 |
| 2025 | Multi-Scale Core-Peripheral Attention Network for Camouflaged Object DetectionabstractIn recent years, camouflage object detection has remained a significant challenge due to the high similarities between objects and backgrounds. Relying solely on convolutions with limited receptive fields or attentions with fixed ranges is in trouble with handling the size variability of cared objects. Moreover, camouflaged targets are frequently covered by their surroundings, with existing methods prone to erroneously identifying the occluded portions. To break the dilemma, we propose a multi-scale core-peripheral attention network (CPANet), mainly including two elaborations: core-peripheral mask attention (CPMA) and multi-scale weighted fusion (MSWF). CPMA boosts camouflaged features by employing core- and peripheral-based attention mechanisms, mitigating the influence of surrounding obstacles and enabling precise localization of concealed targets. Additionally, MSWF captures multi-scale low-level features to refine local details and manifest complete object representations. Extensive evaluations demonstrate that CPANet outperforms state-of-the-art methods across four widely used benchmarks. Yueqian Quan, Tiancheng Pan, Chuangjie Fang, Yan Li 0083, Jianwei Zheng 0001 |
ICME | 5 |
| 2025 | Collaborative Cross-Complementary Unfolding Network for Pan-sharpening Remote Sensing ImageabstractDue to the acquisition limitations of physical devices, pansharpening serves as a computational alternative, enhancing spatial details in low-resolution hyperspectral images with the guidance of corresponding panchromatic images. By leveraging the benefits of nonlinear network architectures and interpretable optimization schemes, deep unfolding networks (DUNs) have shed new light on pansharpening. However, current DUNs lack a dedicated design for both estimating the degradation matrices and extracting intricate information from the proximal operator. To address these challenges, we propose a novel Collaborative Cross-Complementary Unfolding Network (C3U), which is organized into two main steps: customized multi-scale convolution estimation (MSCE) and a data-driven prior extractor. In the MSCE step, the spatial and spectral degradation matrices are individually adapted through multiscale treatment and point convolution operations. Specifically, the overall estimation undergoes an end-to-end iterative block, allowing for adaptive modeling of complex spatial and spectral structures. Within the prior extractor, a cross-complementary attention mechanism is proposed to enable iterative information interaction between global and local Transformers, capturing holistic features and enhancing inductive capacity. Additionally, a collaborative scale-aware-channel mechanism is designed to enlarge the receptive field and capture multiscale channel features in a lightweight manner. More importantly, the principle of collaborative cross-complementary (CCC) permeates all the sub-assemblies, ensuring a desirable information flow. Experimental results on multiple remote sensing datasets demonstrate the superiority of the proposed method over previous state-of-the-art (SOTA) techniques, achieving a 0.8 dB PSNR gain on the GF-2 dataset. Honghui Xu 0002, Yan Li 0083, Yutao Jia, Chuangjie Fang, Jianwei Zheng 0001 |
ICMR | 6 |
| 2025 | SpecSolver: Solving Spatial-Spectral Fusion via Semantic TransformerabstractBy clustering pixels with locally similar values, superpixel-based approaches have shown great potential in processing hyperspectral images (HSI), thereby reducing the computational burden associated with large spatial dimensions. However, specific for spatial-spectral fusion (SSF), superpixel segmentation is inherently non-differentiable and irreversible; hence it is inapplicable. To address the issues, we propose a semantic transformer-based solver, namely SpecSolver, which is basically inspired by the benefits of superpixel-based approaches, yet with the inner mechanism completely improved. The core idea lies in learning the intrinsic semantic states of HSIs hidden behind discretized pixel representations. Specifically, we propose a new Semantic-Attention to adaptively split the image domain into a series of learnable slices of flexible shapes, where image pixels under similar semantic states will be ascribed to the same slice. By calculating attention to the Semantic-Superpixel tokens encoded from slices, SpecSolver can effectively capture intricate semantic correlations from the vast number of pixels, which also empowers the solver with an endogenous capacity for modeling different magnification scales and allows for efficient computation in linear complexity. On that basis, we elaborate a SpatialNet module, which extracts multiscale local spectral information, and a FreqNet module, which supplements global information, capturing subtle details and variations across different spectra. Experiments on two benchmark SSF datasets verify the state-of-the-art (SOTA) performance of the proposed method, both visually and quantitatively. Also, ablation studies validate the mentioned contributions. Wei Li 0034, Honghui Xu 0002, Jiawei Jiang 0002, Jianwei Zheng 0001 |
ACM Multimedia | 5 |
| 2025 | Arbitrary-scale Fusion Neural OperatorabstractSpatial-spectral fusion offers a promising alternative to expensive equipment in high-resolution hyperspectral (HrHs) imaging. However, training separate models for different scaling factors remains costly. To address this, we propose the Arbitrary-scale Fusion Neural Operator (AFNO), a lightweight solution for HrHs fusion across arbitrary scalings. Instead of entities, AFNO treats low-resolution hyperspectral (LrHs) and high-resolution multispectral (HrMs) images as functions and performs meticulously designed integrations as the mapping operator. The key components include Attention-Driven Convolution Integration (ADCI) to restore discretization invariance disrupted by convolutions, Implicit Neural Functional Integration (INFI) for cross-domain interaction of spatial degradations, and Galerkin-type Integration as a decoder for high-frequency details. Additionally, the bonded activation opeartor are improved for the principle of continuous-discrete equivalence. Extensive experiments validate the superiority of our approach over cutting-edge methods. Notably, AFNO holds significantly better generalization on arbitrary scaling factors, yet requiring only 0.07M parameters. Wei Li 0034, Honghui Xu 0002, Jiawei Jiang 0002, Zhi Liu 0009, Jianwei Zheng 0001 |
ACM Multimedia | 6 |
| 2025 | Solving Partial Differential Equations via Radon Neural OperatorabstractNeural operator is considered a popular data-driven alternative to traditional partial differential equation (PDE) solvers. However, most current solutions, whether fulfilling computations in frequency, Laplacian, and wavelet domains, all deviate far from the intrinsic PDE space. While with meticulous network architecture elaborated, the deviation often leads to biased accuracy. To address the issue, we open a new avenue that pioneers leveraging Radon transform to decompose the input space, finalizing a novel Radon neural operator (RNO) to solve PDEs in infinite-dimensional function space. Distinct from previous solutions, we project the input data into the sinogram domain, shrinking the multi-dimensional transformations to a reduced-dimensional counterpart and fitting compactly with the PDE space. Theoretically, we prove that RNO obeys a property of bilipschitz strongly monotonicity under diffeomorphism, providing deeper insights to guarantee the desired accuracy than typical discrete invariance or continuous-discrete equivalence. Within the sinogram domain, we further evidence that different angles contribute unequally to the overall space, thus engineering a reweighting technique to enable more effective PDE solutions. On that basis, a sinogram-domain convolutional layer is crafted, which operates on a fixed $\theta$-grid that is decoupled from the PDE space, further enjoying a natural guarantee of discrete invariance. Extensive experiments demonstrate that RNO sets new state-of-the-art (SOTA) scores across massive standard benchmarks, with superior generalization performance enjoyed. Code is available at <https://github.com/wenbin-lu/Radon-Neural-Operator>. Wenbin Lu, Junnan Xu, Wei Li 0034, Jianwei Zheng 0001 |
NeurIPS | 6 |
| 2025 | BioGraphFusion: graph knowledge embedding for biological completion and reasoningabstractMOTIVATION: Biomedical knowledge graphs (KGs) are crucial for drug discovery and disease understanding, yet their completion and reasoning are challenging. Knowledge embedding (KE) methods capture global semantics but struggle with dynamic structural integration, while graph neural networks (GNNs) excel locally but often lack semantic understanding. Even ensemble approaches, including those leveraging language models, often fail to achieve a deep, adaptive, and synergistic co-evolution between semantic comprehension and structural learning. Addressing this critical gap in fostering continuous, reciprocal refinement between these two aspects in complex biomedical KGs is paramount. RESULTS: We introduce BioGraphFusion, a novel framework for deeply synergistic semantic and structural learning. BioGraphFusion establishes a global semantic foundation via tensor decomposition, guiding an LSTM-driven mechanism to dynamically refine relation embeddings during graph propagation. This fosters adaptive interplay between semantic understanding and structural learning, further enhanced by query-guided subgraph construction and a hybrid scoring mechanism. Experiments across three key biomedical tasks demonstrate BioGraphFusion's superior performance over state-of-the-art KE, GNN, and ensemble models. A case study on cutaneous malignant melanoma 1 highlights its ability to unveil biologically meaningful pathways. AVAILABILITY AND IMPLEMENTATION: Source code and all data underlying this article are freely available in the GitHub repository at https://github.com/Y-TARL/BioGraphFusion. Yitong Lin, Jiaying He, Xinnan Zhu, Jianwei Zheng 0001 |
Bioinform. | 5 |
| 2025 | Nonlinear Learnable Triple-Domain Transform Tensor Nuclear Norm for Hyperspectral Image Super-ResolutionabstractTensor Nuclear Norm (TNN) has been widely employed as a regularization term for hyperspectral image super-resolution (HSISR). However, conventional TNN constraints based on Discrete Fourier Transform (DFT) often suffer from rank estimation biases and an inability to effectively capture complex spectral-spatial correlations, limiting their efficacy in HSISR. To address these challenges, we propose a Nonlinear Learnable Triple-domain (NLT) transform framework that integrates nonlinear transform, DFT, and self-learning adaptation. This multi-stage process promotes singular value concentration, improving low-rank approximation and rank estimation accuracy. Building upon this framework, we develop an NL-transform-oriented tensor product, a truncated singular value decomposition (TSVD) operation, and a novel tensor nuclear norm (NLTN) tailored for HSISR. By incorporating spectral subspace estimation and clustering-based patch grouping, our approach effectively leverages spatial-spectral correlations and non-local self-similarities, leading to enhanced reconstruction quality. To further mitigate singular value over-penalization, we introduce a logarithmic-based generalized NLTNN (GNLTN) and formulate an optimization strategy based on the alternating direction method of multipliers (ADMM). Extensive experiments demonstrate that our method significantly outperforms existing approaches in terms of fusion accuracy and visual fidelity, setting new benchmarks for hyperspectral image super-resolution. The code is available at https://github.com/xuhonghui96/GNLTN. Honghui Xu 0002, Yueqian Quan, Chuangjie Fang, Yan Li 0083, Jianwei Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Semantic-Spatial Attention for Refined Object Placement in Text-to-Image SynthesisabstractSolely based on given prompts, text-guided diffusion models have enjoyed a unique capability in generating diverse and creative images. Nevertheless, the conveyance of image information through text presents a series of challenges, particularly in controlling the positioning of objects in synthesized images. Despite attempts of recent efforts in exploring alternative conditions, such as bounding box/mask-image pairs, the requirement of a substantial amount of paired data and time-consuming fine-tuning emerge as new issues. Given the observations that not only prompt-related cross-attention maps reveal the spatial arrangement and centroid positions of the objects, but also out-of-prompt markers enjoy rich semantic information, we thus engineer a weighted optimization loss. Specifically, three spatial sub-losses, namely inner box reinforcement loss, outer box attenuation loss, and centroid loss, are devised and seamlessly integrated into the sampling step of current vanilla diffusion models. Without any annotations of layout data required, the final approach runs in a training-free fashion. Extensive experiments with new performance scores demonstrate that our proposal not only successfully addresses the issue of object positioning but also boosts the capabilities of most current models, such as Stable Diffusion and GLIGEN, in high-quality synthesis and coverage of various concepts. Moreover, the proposed mechanism plays a plug-and-play role. Jianwei Zheng 0001, Ni Xu, Wei Li 0034, Jiawei Jiang 0002, Xiaoqin Zhang 0002 |
IEEE Trans. Multim. | 1 |
| 2025 | Axial-shunted Spatial-temporal Conversation for Change DetectionabstractBenefitting from the maturing of intelligence techniques and advanced sensors, recent years have witnessed the full flourishing of change detection (CD) on multi-temporal remote sensing images. However, extraneous interference caused by normal temporal evolution and the extreme sparsity of spatial changes still plague the detection accuracy. To counteract this dilemma, a lightweight axial-shunted spatial-temporal conversation network (ASCNet) is proposed, which models the intrinsic representations in dually augmented images with a parallel treatment of convolutions and attentions. Specifically, for the features of weakly augmented bi-temporal image pairs from Siamese CNN, a roundtable attention-based and intra-scale axial-shunted interaction, with linear complexity, is presented. By splitting horizontally or vertically into multiple chunks and then performing axial-squeeze operation, axial-shunted scheme can achieve fine-grained attention while maintaining linear complexity. Moreover, roundtable attention pursues efficient bi-temporal modeling by incorporating both self-attention and cross-attention in a single attentional computation, while imposing change guiding and difference gating for focusing on changes. Simultaneously, a video transformer is introduced for the modeling of strongly augmented sequences, followed by an inter-scale spatial-temporal alignment to recalibrate the feature responses. ASCNet demonstrates state-of-the-art performance on four publicly available CD datasets while maintaining superior computational efficiency. The source code is available at https://github.com/fengyuchao97/ASCNet . Yuchao Feng, Jiawei Jiang 0002, Jintao Lai, Jianwei Zheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Null Space Matters: Range-Null Decomposition for Consistent Multi-Contrast MRI ReconstructionabstractConsistency and interpretability have long been the critical issues in MRI reconstruction. While interpretability has been dramatically improved with the employment of deep unfolding networks (DUNs), current methods still suffer from inconsistencies and generate inferior anatomical structure. Especially in multi-contrast scenes, different imaging protocols often exacerbate the concerned issue. In this paper, we propose a range-null decomposition-assisted DUN architecture to ensure consistency while still providing desirable interpretability. Given the input decomposed, we argue that the inconsistency could be analytically relieved by feeding solely the null-space component into proximal mapping, while leaving the range-space counterpart fixed. More importantly, a correlation decoupling scheme is further proposed to narrow the information gap for multi-contrast fusion, which dynamically borrows isotropic features from the opponent while maintaining the modality-specific ones. Specifically, the two features are attached to different frequencies and learned individually by the newly designed isotropy encoder and anisotropy encoder. The former strives for the contrast-shared information, while the latter serves to capture the contrast-specific features. The quantitative and qualitative results show that our proposal outperforms most cutting-edge methods by a large margin. Codes will be released on https://github.com/chenjiachengzzz/RNU. Jiawei Jiang 0002, Fei Wu 0026, Jianwei Zheng 0001 |
AAAI | 4 |
| 2024 | SyFormer: Structure-Guided Synergism Transformer for Large-Portion Image InpaintingabstractImage inpainting is in full bloom accompanied by the progress of convolutional neural networks (CNNs) and transformers, revolutionizing the practical management of abnormity disposal, image editing, etc. However, due to the ever-mounting image resolutions and missing areas, the challenges of distorted long-range dependencies from cluttered background distributions and reduced reference information in image domain inevitably rise, which further cause severe performance degradation. To address the challenges, we propose a novel large-portion image inpainting approach, namely the Structure-Guided Synergism Transformer (SyFormer), to rectify the discrepancies in feature representation and enrich the structural cues from limited reference. Specifically, we devise a dual-routing filtering module that employs a progressive filtering strategy to eliminate invalid noise interference and establish global-level texture correlations. Simultaneously, the structurally compact perception module maps an affinity matrix within the introduced structural priors from a structure-aware generator, assisting in matching and filling the corresponding patches of large-proportionally damaged images. Moreover, we carefully assemble the aforementioned modules to achieve feature complementarity. Finally, a feature decoding alignment scheme is introduced in the decoding process, which meticulously achieves texture amalgamation across hierarchical features. Extensive experiments are conducted on two publicly available datasets, i.e., CelebA-HQ and Places2, to qualitatively and quantitatively demonstrate the superiority of our model over state-of-the-arts. Yuchao Feng, Honghui Xu 0002, Chuanmeng Zhu, Jianwei Zheng 0001 |
AAAI | 5 |
| 2024 | Learning Object Placement via Convolution Scoring Attention
Yuchao Feng, Jianwei Zheng 0001 |
BMVC | 3 |
| 2024 | High-fidelity Person-centric Subject-to-Image SynthesisabstractCurrent subject-driven image generation methods en-counter significant challenges in person-centric image generation. The reason is that they learn the semantic scene and person generation by fine-tuning a common pre-trained diffusion, which involves an irreconcilable training imbalance. Precisely, to generate realistic persons, they need to sufficiently tune the pre-trained model, which inevitably causes the model to forget the rich semantic scene prior and makes scene generation over-fit to the training data. Moreover, even with sufficient fine-tuning, these methods can still not generate high-fidelity persons since joint learning of the scene and person generation also lead to quality compromise. In this paper, we propose Face-diffuser, an effective collaborative generation pipeline to eliminate the above training imbal-ance and quality compromise. Specifically, we first develop two specialized pre-trained diffusion models, i.e., Text-driven Diffusion Model (TDM) and Subject-augmented Diffusion Model (SDM), for scene and person generation, respectively. The sampling process is divided into three sequential stages, i.e., semantic scene construction, subject-scene fusion, and subject enhancement. The first and last stages are performed by TDM and SDM respectively. The subject-scene fusion stage, that is the collaboration achieved through a novel and highly effective mechanism, Saliency-adaptive Noise Fusion (SNF). Specifically, it is based on our key observation that there exists a robust link between classifier-free guidance responses and the saliency of generated images. In each time step, SNF leverages the unique strengths of each model and allows for the spatial blending of predicted noises from both models automatically in a saliency-aware manner, all of which can be seamlessly integrated into the DDIM sampling process. Extensive experiments confirm the impressive effectiveness and robustness of the Face-diffuser in gener-ating high-fidelity person images depicting multiple unseen persons with varying contexts. Code is available at https://github.com/CodeGoat24/Face-diffuser. Jianwei Zheng 0001, Cheng Jin 0001 |
CVPR | 3 |
| 2024 | Decoupled Competitive Framework for Semi-Supervised Medical Image SegmentationabstractConfronting the critical challenge of insufficiently annotated samples in medical domain, semi-supervised medical image segmentation (SSMIS) emerges as a promising solution. Specifically, most methodologies following the Mean Teacher (MT) or Dual Students (DS) architecture have achieved commendable results. However, to date, these approaches face a performance bottleneck due to two inherent limitations, e.g., the over-coupling problem within MT structure owing to the employment of exponential moving average (EMA) mechanism, as well as the severe cognitive bias between two students of DS structure, both of which potentially lead to reduced efficacy, or even model collapse eventually. To mitigate these issues, a Decoupled Competitive Framework (DCF) is elaborated in this work, which utilizes a straightforward competition mechanism for the update of EMA, effectively decoupling students and teachers in a dynamical manner. In addition, the seamless exchange of invaluable and precise insights is facilitated among students, guaranteeing a better learning paradigm. The DCF introduced undergoes rigorous validation on three publicly accessible datasets, which encompass both 2D and 3D datasets. The results demonstrate the superiority of our method over previous cutting-edge competitors. Code will be available at https://github.com/Fly-away20/DCF. Jiahe Ying, Jianwei Zheng 0001 |
ECAI | 4 |
| 2024 | Memory-Augmented Dual-Domain Unfolding Network for MRI ReconstructionabstractThe compressed sensing MRI aims to recover high-fidelity images from undersampled k-space data, which enables MRI acceleration and meanwhile mitigates problems caused by prolonged acquisition time, such as physiological motion artifacts, patient discomfort, and delayed medical care. In this regard, the deep unfolding network (DUN) has emerged as the predominant solution due to the benefits of better interpretability and model capacity. However, existing algorithms remain inadequate for two principal reasons. First, directly unrolling a typical optimization algorithm is ill-considered for the structure information and domain knowledge. Second, the incorporation of the MRI-oriented imaging mechanism is inadequate. To tackle these two issues, we propose a Memoryaugmented Dual-domain Unfolding Network (MDUNet). Particularly, the per-iteratively learned memory is held to facilitate a better efficacy of feature representation. Besides, with the scheme of memory augmentation alternatively employed in the k-space and image domain, both the regional structure and global information can be complementarily integrated in a spiral manner. Comprehensive experiments conducted on diverse datasets, sampling rates, and sampling patterns demonstrate that our method, while maintaining a relatively small number of parameters, surpasses the latest methods. Codes will be available on the GitHub homepage of the corresponding author. Jiawei Jiang 0002, Yueqian Quan, Jianwei Zheng 0001 |
ICASSP | 5 |
| 2024 | HMNet: Hierarchical Microscale-Aware Network for Infrared Small Target DetectionabstractCompared to the natural image community, infrared target detection suffers more challenges due to the severely tiny and low-contrast objects, especially in cases with obscuration from clutter and noise. The traditional solutions are susceptible to noise interference, which yields suboptimal performance lacking of contour and texture details. Meanwhile, due to the spatial invariance of convolutional layers, most deep learning-based methods locate small targets loosely during feature extraction, leading to serious omissions. To address these limitations, we propose a hierarchical microscale-aware network (HMNet) following an encoder-decoder structure that is mainly equipped with two novel modules: the holistic attention-aware (HAA) module and the scale-aware adaptive extraction (SAE) module. HAA integrates local and global cues via self-attention, depthwise separable convolutions, and dilated convolutions, which hammers at enhancing target features and ensuring accurate localization. As a complement, SAE employs multi-scale features and spatial-channel attention to acquire richer texture details while reducing background noise. The experiments on public datasets demonstrate that our method achieves state-of-the-art performance. Yueqian Quan, Honghui Xu 0002, Yidong Yan, Jianwei Zheng 0001 |
ICASSP | 5 |
| 2024 | Robust Principal Component Analysis via High-Order Self-Learning Transform Tensor Nuclear NormabstractIn recent studies, tensor singular value decomposition (TSVD) within the high-order (Ho) algebra has shed light on solving the Tensor Robust Principal Component Analysis (TRPCA) problem. However, the utilization of fixed or data-independent transformations in HoTSVD may result in suboptimal outcomes. To overcome this limitation, we propose a self-learning TSVD method that rectifies computational inefficiencies and learns a lossless transformation, inducing a lower average-rank tensor. This involves multiplying learnable semi-orthogonal matrices obtained through Tucker compression with the original tensor along all modes, resulting in a core tensor with enhanced inherent low rankness and new self-learning transform matrices. The semi-orthogonal transforms, acting as a crucial building block, enhance spatial low-rankness, facilitating the resolution of smaller-scale problems and the design of efficient algorithms. Additionally, a reweighting Schatten-p scheme is integrated into the self-learning HoTSVD to understand global low-rank correlations, offering an effective numerical solution. Finally, we develop an alternating direction method of multipliers (ADMM)-based algorithm as a solver. Experimental results on Light Field Images (LFI), showcase the superiority of our proposed method over previous state-of-the-art approaches. Honghui Xu 0002, Yueqian Quan, Chuangjie Fang, Jianwei Zheng 0001 |
ICME | 4 |
| 2024 | PrimeComposer: Faster Progressively Combined Diffusion for Image Composition with Attention Steeringabstractan oil painting of an eiffel tower in the distance, Van Gogh Style an oil painting of a shopping mall in the distance, Van Gogh Style a pencil drawing of a car and a willow, black and white painting a pencil drawing of buildings in the distance, black and white painting a cartoon animation of an elephant in the forest a cartoon animation of a cabinet a professional photograph of a wet puppy in a pool, ultra realistic a professional photograph of a castle in the distance, ultra realistic Figure 1: Displayed are the results generated using PrimeComposer, showcasing its prowess across various domains: oil painting, sketching, cartoon animation, and photorealism. Jianwei Zheng 0001, Cheng Jin 0001 |
ACM Multimedia | 3 |
| 2024 | Alias-Free Mamba Neural OperatorabstractBenefiting from the booming deep learning techniques, neural operators (NO) are considered as an ideal alternative to break the traditions of solving Partial Differential Equations (PDE) with expensive cost.
Yet with the remarkable progress, current solutions concern little on the holistic function features--both global and local information-- during the process of solving PDEs.
Besides, a meticulously designed kernel integration to meet desirable performance often suffers from a severe computational burden, such as GNO with $O(N(N-1))$, FNO with $O(NlogN)$, and Transformer-based NO with $O(N^2)$.
To counteract the dilemma, we propose a mamba neural operator with $O(N)$ computational complexity, namely MambaNO.
Functionally, MambaNO achieves a clever balance between global integration, facilitated by state space model of Mamba that scans the entire function, and local integration, engaged with an alias-free architecture. We prove a property of continuous-discrete equivalence to show the capability of
MambaNO in approximating operators arising from universal PDEs to desired accuracy. MambaNOs are evaluated on a diverse set of benchmarks with possibly multi-scale solutions and set new state-of-the-art scores, yet with fewer parameters and better efficiency. Jianwei Zheng 0001, Wei Li 0034, Ni Xu, Xiaoxu Lin, Xiaoqin Zhang 0002 |
NeurIPS | 1 |
| 2024 | Multi-dimensional visual data completion via weighted hybrid graph-Laplacian
Jiawei Jiang 0002, Yile Xu, Honghui Xu 0002, Guojiang Shen, Jianwei Zheng 0001 |
Signal Process. | 5 |
| 2024 | STENet: A Spatial Selection and Temporal Evolution Network for Change Detection in Remote Sensing ImagesabstractAccompanied by the booming development of remote sensing (RS) imaging techniques, change detection (CD) has emerged as a conspicuous focal point in the realm of geoscience. Traditionally, extensive research has predominantly centered on extracting semantic features from individual images yet neglecting the interinput correlations latent in bitemporal imagery. This oversight gives rise to the occurrence of pseudovariant regions and the blurring of detection boundaries. To address these challenges, we propose a spatial selection and temporal evolution network, named STENet, which aims to unravel semantic correlations between bitemporal images from both spatial and temporal perspectives. Specifically, a dual pathway is crafted. The first one explores precisely the spatial localization of changes in bitemporal pairs, while the other complements the local details by generating a pseudovideo input via specific data augmentation. To boost the precision (Pre) of localizing sparsely changed targets, we further present a dynamic selective attention (DSA) mechanism, which strives for more focus on the positive regions while holding a mild computational demand. Moreover, by leveraging data augmentation to derive the temporal evolution of a set of progressively changed images, we then exploit a 3-D convolution-based encoder to mine the potential details therein, endeavoring to a refinement of target boundaries. Ultimately, STENet enforces the fusion of multiscale spatial and temporal features through a dedicated decoder and generates the final change map. Experiments on four popular datasets show that our proposal scores higher than most state-of-the-art approaches. In addition, the appealing performance is achieved with mild quantities of parameters and computations. The code is available athttps://github.com/ZhengJianwei2/STENet. Jintao Lai, Yiting Jin, Jianwei Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | ORSI Salient Object Detection via Progressive Semantic Flow and Uncertainty-Aware RefinementabstractWith the prosperity of deep learning techniques, salient object detection in remote sensing images (RSI-SOD) is concomitantly in full flourishing. However, due to the inherent challenges such as uncertainty in object quantities and scales, cluttered backgrounds, and blurred edges arising from shadows, most current approaches struggle for salient feature learning with the aid of heavy model architecture, yet often result in barely satisfactory performance. Some methods compromise model complexity to improve efficiency, albeit with significantly degraded results. To earn a satisfactory balance of efficacy and efficiency, we propose a new network for RSI-SOD, namely SFANet, based on progressive semantic flow and uncertainty-aware refinement. Specifically, we design a global semantic enhancement block (GSEB) to reduce background interference and accurately localize salient objects of varying quantities and scales, which further consists of three modularized components, i.e., semantic extraction module (SEM), interscale fusion module (IFM), and deep semantic graph-inference module (DSGM). SEM together with IFM contributes to the effective aggregation of multi-scale contexts by extracting fused and progressive semantic cues. DSGM performs semantic inference to better localize salient objects with irregularities in scale and topological structure. Furthermore, we present an uncertainty-aware refinement module (URM) to recognize salient objects in cluttered backgrounds and effectively suppress shadows. Extensive experiments are conducted on three RSI-SOD datasets, from which superior results can be achieved by our SFANet, outperforming the other cutting-edge methods. The code is available at https://github.com/ZhengJianwei2/SFANet. Yueqian Quan, Honghui Xu 0002, Renfang Wang, Qiu Guan, Jianwei Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Cascade-Transform-Based Tensor Nuclear Norm for Hyperspectral Image Super-ResolutionabstractRecent advancements in tensor nuclear norm (TNN) have led to promising solutions for hyperspectral image super-resolution (HSISR), which produces enriched outputs by fusing low-resolution hyperspectral images (LRHSIs) with high-resolution multispectral images (HRMSIs). However, current TNN, mainly reliant on the discrete Fourier transform (DFT), still suffers from mirroring boundary effects and singleton domain limitation. As a relief, we propose cascade-transform-based tensor nuclear norm (CTNN) with two variants for HSISR, featuring new definitions and algebraic structures for tensor product and TNN operations. The first variant processes tubal elements derived from DFT as inputs in the discrete cosine transform (DCT) domain, allowing for more nuanced feature extraction. The second learns adaptive matrices from the data in each iteration update and links them with a fixed DFT matrix to dynamically update the transform domain, preventing rank estimation bias. Furthermore, the nonconvex form of the proposed CTNN is applied to three modes of each spectral subspace similarity cube, termed log-sum-based full-scale CTNN (LFCTNN), capturing the global low-rank structure of LRHSI and the nonlocal similarities present in HRMSI. Experimental evaluations on various remote sensing datasets indicate that our approach exceeds existing state-of-the-art methods. The code is available athttps://github.com/xuhonghui96/LFCTNN. Honghui Xu 0002, Chuangjie Fang, Yilin Ge, Yubin Gu, Jianwei Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | CGMA-Net: Cross-Level Guidance and Multi-Scale Aggregation Network for Polyp SegmentationabstractColonoscopy is considered the best prevention and control method for colorectal cancer, which suffers extremely high rates of mortality and morbidity. Automated polyp segmentation of colonoscopy images is of great importance since manual polyp segmentation requires a considerable time of experienced specialists. However, due to the high similarity between polyps and mucosa, accompanied by the complex morphological features of colonic polyps, the performance of automatic polyp segmentation is still unsatisfactory. Accordingly, we propose a network, namely Cross-level Guidance and Multi-scale Aggregation (CGMA-Net), to earn a performance promotion. Specifically, three modules, including Cross-level Feature Guidance (CFG), Multi-scale Aggregation Decoder (MAD), and Details Refinement (DR), are individually proposed and synergistically assembled. With CFG, we generate spatial attention maps from the higher-level features and then multiply them with the lower-level features, highlighting the region of interest and suppressing the background information. In MAD, we parallelly use multiple dilated convolutions of different sizes to capture long-range dependencies between features. For DR, an asynchronous convolution is used along with the attention mechanism to enhance both the local details and the global information. The proposed CGMA-Net is evaluated on two benchmark datasets, i.e., CVC-ClinicDB and Kvasir-SEG, whose results demonstrate that our method not only presents state-of-the-art performance but also holds relatively fewer parameters. Concretely, we achieve the Dice Similarity Coefficient (DSC) of 91.85% and 95.73% on Kvasir-SEG and CVC-ClinicDB, respectively. The assessment of model generalization is also conducted, resulting in DSC scores of 86.25% and 86.97% on the two datasets respectively. Jianwei Zheng 0001, Yidong Yan |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | Cascading Blend Network for Image InpaintingabstractImage inpainting refers to filling in unknown regions with known knowledge, which is in full flourish accompanied by the popularity and prosperity of deep convolutional networks. Current inpainting methods have excelled in completing small-sized corruption or specifically masked images. However, for large-proportion corrupted images, most attention-based and structure-based approaches, though reported with state-of-the-art performance, fail to reconstruct high-quality results due to the short consideration of semantic relevance. To relieve the above problem, in this paper, we propose a novel image inpainting approach, namely cascading blend network (CBNet), to strengthen the capacity of feature representation. As a whole, we introduce an adjacent transfer attention (ATA) module in the decoder, which preserves contour structure reasonably from the deep layer and blends structure-texture information from the shadow layer. In a coarse to delicate manner, a multi-scale contextual blend (MCB) block is further designed to felicitously assemble the multi-stage feature information. In addition, to ensure a high qualified hybrid of the feature information, extra deep supervision is applied to the intermediate features through a cascaded loss. Qualitative and quantitative experiments on the Paris StreetView, CelebA, and Places2 datasets demonstrate the superior performance of our approach compared with most state-of-the-art algorithms. Yiting Jin, Wanliang Wang, Yidong Yan, Jiawei Jiang 0002, Jianwei Zheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | Contrastive Attention-guided Multi-level Feature Registration for Reference-based Super-resolutionabstractGiven low-quality input and assisted by referential images, reference-based super-resolution (RefSR) strives to enlarge the spatial size with the guarantee of realistic textures, for which sophisticated feature-matching strategies are naturally demanded. However, the miserable transformation gap between inputs and references, e.g., texture rotation and scaling within patches, often yields distorted textures and terrible ghosting artifacts, which seriously hampers the visual senses and their further investigation. To circumvent this challenge, we propose a contrastive attention-guided multi-level feature registration for RefSR, explicitly tapping the potential of interacting between inputs and references. Specifically, we develop a multi-level feature warping scheme, involving patch-level coarse feature swapping and pixel-level deformable alignment, to model generalized spatial transformation correspondences steered by contrastive attention. Notably, a spatial registration module is embedded for further calibration against the potential misalignment issue and inter-feature distribution difference. In addition, aiming at suppressing the impacts of irrelevant or superfluous information on cross-scale features, we incorporate a multi-residual feature fusion module to strive for visually plausible textures. Experimental results on four publicly available datasets demonstrate that our method outperforms most state-of-the-art approaches in terms of both efficiency and perceptual effectiveness. Jianwei Zheng 0001, Yu Liu 0151, Yuchao Feng, Honghui Xu 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | GCN-based Autism Spectrum Disorder Diagnosis via Convolutional Restructuring Attention
Qianwei Zhou, Yuchao Feng, Jianwei Zheng 0001 |
CogSci | 5 |
| 2023 | Building Change Detection Using Cross-Temporal Feature Interaction NetworkabstractBuilding change detection of remote sensing images is in full flourishing accompanied by the prosperity of convolutional neural networks. For spatial-temporal context modeling, existing solutions disregard the inter-image interactions, albeit their positive contribution to the acquisition of differences. To fill the gap, we propose a cross-temporal feature interaction network to effectively derive the change representations. Specifically, we propose a linearized cross-attention, which motivates each counterpart to glimpse the representation of another image while preserving its own features. In addition, to circumvent the misalignment caused by step-down sampling in the backbone, we introduce multi-level feature alignment using learnable affine transformation and stepwise aggregation. Based on a naive backbone (ResNet18) without sophisticated structures, our model outperforms other state-of-the-art methods on three datasets in terms of both efficiency and effectiveness. Yuchao Feng, Jiawei Jiang 0002, Honghui Xu 0002, Jianwei Zheng 0001 |
ICASSP | 4 |
| 2023 | Low-Dose CT Reconstruction Via Optimization-Inspired GANabstractMost research on Low-dose Computed Tomography (LDCT) reconstruction is designed as a black box, lacking controllability and interpretability. In this paper, a Proximal Linear ADMM framework-based Generative Adversarial Network (PLA-GAN) is proposed. Specifically, without loss of interpretability, channel attention blocks and NonLocal Sparse Attention (NLSA) modules are embedded into two regularizers respectively and iterated alternately, driving the network to cope with real and complex CT image degradation through a multi-scale and adaptive way. To further promote the visual quality, a discriminator containing NLSA module is also introduced. The comparisons with state-of-the-arts on the Mayo dataset validate the superiority of our proposed algorithm both numerically and visually. The advantages of generalizability and interpretability are also evident. Jiawei Jiang 0002, Yuchao Feng, Honghui Xu 0002, Jianwei Zheng 0001 |
ICASSP | 4 |
| 2023 | Compact Intertemporal Coupling Network for Remote Sensing Change DetectionabstractChange detection of multi-temporal remote sensing images is in full flourishing accompanied by the popularity and prosperity of deep learning. The prominent challenge lies in the crude distribution of the newly constructed and demolished changes, the interferences of massive irrelevant objects, and the spatial-temporal changes from the passage of time. For context modeling, existing solutions waste massive attention on task-irrelevant features, spotlighting insufficiently on the genuinely changed regions. To fill the gap, we propose a compact intertemporal coupling network (CICNet) to derive the change representations. Specifically, to underpin the interaction of spatial-temporal differences in a global perspective, we detach and innovate the solo-head self-attention into a lightweight intertemporal-attention, favorably bridging the intra-level features. In addition, to circumvent the misalignment imposed by spatial sampling, lightweight global channel- and spatial- attentions are globally incorporated for stepwise calibration between localization seduced low-level information and semantics abundant high-level features. Based on a naive backbone (ResNet18/34) without sophisticated structures, our model outperforms other state-of-the-art methods on four datasets in terms of both efficiency and effectiveness. Yuchao Feng, Honghui Xu 0002, Jiawei Jiang 0002, Jianwei Zheng 0001 |
ICME | 4 |
| 2023 | GA-HQS: MRI reconstruction via a generically accelerated unfolding approachabstractDeep unfolding networks (DUNs) are the foremost methods in the realm of compressed sensing MRI, as they can employ learnable networks to facilitate interpretable forward-inference operators. However, several daunting issues still exist, including the heavy dependency on first-order optimization algorithms, the insufficient information fusion mechanisms, and the limitation of capturing long-range relationships. To address the issues, we propose a Generically Accelerated Half-Quadratic Splitting (GA-HQS) algorithm that incorporates second-order gradient information and pyramid attention modules for the delicate fusion of inputs at the pixel level. Moreover, a multi-scale split transformer is also designed to enhance the global feature representation. Comprehensive experiments demonstrate that our method surpasses previous ones on single-coil MRI acceleration tasks. Jiawei Jiang 0002, Honghui Xu 0002, Yuchao Feng, Jianwei Zheng 0001 |
ICME | 5 |
| 2023 | CA-GAN: Object Placement via Coalescing Attention based Generative Adversarial NetworkabstractLearning to posit a foreground object over a background scene is an intriguing yet challenging problem, which frequently emerges in applications such as image editing and scene parsing. To date, most existing studies are fed up with knotty issues, including the deficiency of harnessing the interaction between the object and the scene, the astriction of involving little prior knowledge during training, etc. To break the shackles, we propose a novel end-to-end framework dubbed Coalescing Attention based Generative Adversarial Network (CA-GAN). Specifically, in our synthesizer, a feature polymerizer is designed to distill multi-scale information from both background and foreground. On that basis, a dual-branch coalescing attention module is proposed for a better exploration of the global feature-interaction relationships between object and scene. In addition, we add a supervised trail to learn the prior knowledge from the positive composite image, which further guides the synthesizer to discover a credible placement for the foreground object. With extensive experiments conducted on the OPA dataset, our proposal presents superiority in both rationality and diversity compared with other state-of-the-art methods. Our code is available at https://github.com/ZhengJianwei2/CA-GAN. Yuchao Feng, Honghui Xu 0002, Jianwei Zheng 0001 |
ICME | 5 |
| 2023 | A Lightweight Collective-attention Network for Change DetectionabstractChange detection of multi-temporal remote sensing images is mushrooming with the innovations of neural networks, whose daunting challenge lies in locating sporadically distributed spatial-temporal changes given sophisticated scenes and various imaging conditions. Unfortunately, instead of devoting full attention to changes, most existing solutions often expend unnecessary resources yet derive task-irrelevant features. To relieve this issue, we propose a collective-attention network, which enjoys lightweight model architecture yet guarantees high performance. Specifically, an inter-temporal collective-attention module is developed for efficient interaction of bi-temporal features, in which a shared attention distribution is derived via the multiplication of temporal-concatenated queries and spatial-subtracted keys. Additionally, we present a non-change consistency-constraint, enforcing a change-oriented attention distribution and a noise-suppressed treatment. With the learned interaction features, bi-temporal differences are captured simply using the operations of spatial absolute error and temporal concatenation. Finally, decoding multi-scale differences is accomplished by lightweight temporal self-attention and spatial self-attention. Experiments on four datasets demonstrate that our model achieves state-of-the-art performance, yet requires only 1.71M parameters and 1.98G FLOPs. Yuchao Feng, Yanyan Shao, Honghui Xu 0002, Jinshan Xu, Jianwei Zheng 0001 |
ACM Multimedia | 5 |
| 2023 | Latent-space Unfolding for MRI ReconstructionabstractTo circumvent the problems caused by prolonged acquisition periods, compressed sensing MRI enjoys a high usage profile to accelerate the recovery of high-quality images from under-sampled k-space data. Most current solutions dedicate to solving this issue with the pursuit of certain prior properties, yet the treatments are all enforced in the original space, resulting in limited feature information. To achieve a performance promotion yet with the guarantee of running efficiency, in this work, we propose a latent-space unfolding network (LsUNet). Specifically, by an elaborately designed reversible network, the inputs are first mapped to a channel-lifted latent space, which taps the potential of capturing spatial-invariant features sufficiently. Within the latent space, we then unfold an accelerated optimization algorithm to iterate an efficient and feasible solution, in which a parallelly dual-domain update is equipped for better feature fusion. Finally, an inverse embedding transformation of the recovered high-dimensional representation is applied to achieve the expected estimation. LsUNet enjoys high interpretability due to the physically induced modules, which not only facilitates an intuitive understanding of the internal operating mechanism but also endows it with high generalization ability. Comprehensive experiments on different datasets and various sampling rates/patterns demonstrate the advantages of our proposal over the latest methods both visually and numerically. Jiawei Jiang 0002, Yuchao Feng, Dongyan Guo, Jianwei Zheng 0001 |
ACM Multimedia | 5 |
| 2023 | Tensor completion via hybrid shallow-and-deep priors
Honghui Xu 0002, Jiawei Jiang 0002, Yuchao Feng, Yiting Jin, Jianwei Zheng 0001 |
Appl. Intell. | 5 |
| 2023 | ORSI Salient Object Detection via Cross-Scale Interaction and Enlarged Receptive FieldabstractDue to the diversity of scales and shapes, the uncertainty of object position, and the complexity of edge details, the recent merging problem of salient object detection in optical remote sensing image (RSI-SOD) is a considerably challenging topic. To cope with the challenges, we propose a new cross-scale interaction network (CIFNet) equipped with the enlarged receptive field, which mainly contains three modules in an encoder–decoder architecture, including a furcate skip connection module (FSCM), a global leading attention module, and an expansion–integration module (EIM). First, the FSCM uses dilated convolutions to enlarge the receptive field and furcate skip connections to capture more multiscale contextual information, both of which facilitate the adaptability of the model to different sizes, shapes, and quantities of the target objects. Second, on the low-resolution branch, a global leading attention module (GLM) locates the potentially significant object positions in the feature map from a global semantic perspective. Finally, through an attention-guided cascade structure, the EIM seeks more delicate characteristics by refining the features in a coarse-to-fine fashion. Extensive experiments are conducted on two RSI-SOD datasets, from which superior results can be achieved by our CIFNet, outperforming the other state-of-the-art methods. Compared with the second-best method, the performance gain of our method reaches 3.45% on mean absolute error (MAE) and 1.38% on$F_{\beta} ^{\mathrm {adp}}$. Notably, the proposed CIF-Net runs with 40.40-M parameters, 14.8-GFLOPs computational complexity, and 58-frames/s inference speed, which guarantees high efficiency. Jianwei Zheng 0001, Yueqian Quan |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2023 | Nonlocal B-spline representation of tensor decomposition for hyperspectral image inpainting
Honghui Xu 0002, Yidong Yan, Jianwei Zheng 0001 |
Signal Process. | 5 |
| 2023 | Change Detection on Remote Sensing Images Using Dual-Branch Multilevel Intertemporal NetworkabstractChange detection (CD) of remote sensing (RS) images is mushrooming up accompanied by the on-going innovation of convolutional neural networks (CNNs). Yet with the high-speed technology upgrade, the obstacle that identifies unbalanced variations in foreground–background categories still lies on the table, especially in cases with limited samples and massive interference such as seasonal turnover, illumination intensity, and building reformation. Moreover, to date, neither of the off-the-shelf methods probes the feasibility of direct interaction between bitemporal images before accessing difference features. In this article, we propose a dual-branch multilevel intertemporal network (DMINet) to efficiently and effectively derive the change representations. Specifically, by unifying self-attention (SelfAtt) and cross-attention (CrossAtt) in a single module, we present an intertemporal joint-attention (JointAtt) block to steer the global feature distribution of each input, motivating information coupling between intralevel representations and meanwhile suppressing the task-irrelevant interferences. In addition, centering more on the detection of difference features, a reliable architecture is designed by spotlighting two concerns, i.e., the difference acquisition using subtraction and concatenation as well as the multilevel difference aggregation using incremental feature alignment. Based on a naive backbone without sophisticated structures, i.e., ResNet18, our model outperforms other state-of-the-art (SOTA) methods on four CD datasets, especially in cases with rarely samples. Moreover, the achievement is attained with light overheads. Yuchao Feng, Jiawei Jiang 0002, Honghui Xu 0002, Jianwei Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | ORSI Salient Object Detection via Bidimensional Attention and Full-Stage Semantic GuidanceabstractThe application of optical remote sensing images (ORSIs) is prevalent in many fields. Accordingly, ORSI-oriented salient object detection (SOD) has attracted more attention in recent years. However, yet many previously proposed methods present appealing performance in natural scene images (NSIs), they are difficult to be directly extended to remote sensing images due to the more complex scenes, such as blended backgrounds and diversiform topological shapes. Most specifically designed models often fail to achieve satisfactory results due to the weak usage of edge information and the ignorance of attention loss. Besides, computational inefficiency often causes poor applicability. To solve these problems, we propose a new model, namely, Bidimensional Attention and Full-stage Semantic Guidance Network (BAFS-Net), containing an edge guidance branch and a mainstream detection branch. Concretely, edge guidance generates boundary information, in which supervision with border labels is imposed to highlight the salient regions and plays a complementary role on the main branch. The mainstream detection branch involves two important components, i.e., bidimensional attention modules (BAMs) and semantic-guided fusion modules (SGFMs). Between these two, BAM uniformly assembles channel and spatial attention in an efficient and rational manner, addressing the open issue of dimensionwisely attention computation. SGFM hammers at the fusion of high-level features and low-level features. Moreover, the semantic maps are employed to interact with SGFM in full stages. Our approach surpasses most state-of-the-art RSI-SOD methods proposed in recent years, with respect to the accuracy, parameter size, computational cost, and floating point operations per second (FLOPS). The code is available athttps://github.com/ZhengJianwei2/BAFS-Net. Yubin Gu, Honghui Xu 0002, Yueqian Quan, Jianwei Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Remote Sensing Semantic Segmentation via Boundary Supervision-Aided Multiscale Channelwise Cross Attention NetworkabstractHigh spatial resolution (HSR) remote sensing images inevitably pose the challenge of multi-scale transformation, as small objects such as cars and helicopters may occupy only a few pixel points. This incurs a significant hurdle for global context modeling, particularly in backbone networks with large downsampling coefficients. Simple summation or concatenation techniques, such as skip connections, fail to address semantic gaps and even impose negative impacts on multi-scale feature fusion. Meanwhile, due to the complexity of foreground objects, the boundary details of HSR remote sensing images are easy to lose in sampling operations. To overcome these challenges, we propose a Multi-scale Channel-wise Cross Attention Network (MCCANet) assisted by boundary supervision. Technically, MCCA captures the channel attention with various scales, which allows dynamic and adaptive feature fusion in a contextual scale-aware manner and focuses on both large and small objects distributed throughout the inputs. Besides, a Channel and Context Strainer (CCS) module is proposed and embedded in MCCA, filtering channels and contexts for the mitigation of intra-class differences. In addition, we apply a Boundary Supervision (BS) module to recover boundary contour, avoiding the blurring effect during the construction of contextual information. The refined boundary allows for the effective recognition of surrounding pixels, ensuring a better segmentation performance. Extensive experiments on iSAlD, ISPRS Potsdam, and LoveDA datasets demonstrate that our proposed MCCANet achieves a good balance of high accuracy and efficiency. Code will be available at: https://github.com/ZhengJianwei2/MCCANet. Jianwei Zheng 0001, Anhao Shao, Yidong Yan |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Boosting Feature-Aware Network for Salient Object Detection
Jianwei Zheng 0001, Yubin Gu, Yuchao Feng, Jinshan Xu |
ICANN (4) | 1 |
| 2022 | Fast Tensor Nuclear Norm for Structured Low-Rank Visual InpaintingabstractLow-rank modeling has achieved great success in visual data completion. However, the low-rank assumption of original visual data may be in approximate mode, which leads to suboptimality for the recovery of underlying details, especially when the missing rate is extremely high. In this paper, we go further by providing a detailed analysis about the rank distributions in Hankel structured and clustered cases, and figure out both non-local similarity and patch-based structuralization play a positive role. This motivates us to develop a new Hankel low-rank tensor recovery method that is competent to truthfully capture the underlying details with sacrifice of slightly more computational burden. First, benefiting from the correlation of different spectral bands and the smoothness of local spatial neighborhood, we divide the visual data into overlapping 3D patches and group the similar ones into individual clusters exploring the non-local similarity. Second, the 3D patches are individually mapped to the structured Hankel tensors for better revealing low-rank property of the image. Finally, we solve the tensor completion model via the well-known alternating direction method of multiplier (ADMM) optimization algorithm. Due to the fact that size expansion happens inevitably in Hankelization operation, we further propose a fast randomized skinny tensor singular value decomposition (rst-SVD) to accelerate the per-iteration running efficiency. Extensive experimental results on real world datasets verify the superiority of our method compared to the state-of-the-art visual inpainting approaches. Honghui Xu 0002, Jianwei Zheng 0001, Xiaomin Yao, Yuchao Feng, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | LSCIDMR: Large-Scale Satellite Cloud Image Database for Meteorological ResearchabstractPeople can infer the weather from clouds. Various weather phenomena are linked inextricably to clouds, which can be observed by meteorological satellites. Thus, cloud images obtained by meteorological satellites can be used to identify different weather phenomena to provide meteorological status and future projections. How to classify and recognize cloud images automatically, especially with deep learning, is an interesting topic. Generally speaking, large-scale training data are essential for deep learning. However, there is no such cloud images database to date. Thus, we propose a large-scale cloud image database for meteorological research (LSCIDMR). To the best of our knowledge, it is the first publicly available satellite cloud image benchmark database for meteorological research, in which weather systems are linked directly with the cloud images. LSCIDMR contains 104 390 high-resolution images, covering 11 classes with two different annotation methods: 1) single-label annotation and 2) multiple-label annotation, called LSCIDMR-S and LSCIDMR-M, respectively. The labels are annotated manually, and we obtain a total of 414 221 multiple labels and 40 625 single labels. Several representative deep learning methods are evaluated on the proposed LSCIDMR, and the results can serve as useful baselines for future research. Furthermore, experimental results demonstrate that it is possible to learn effective deep learning models from a sufficiently large image database for the cloud image classification. Cong Bai, Minjing Zhang, Jinglin Zhang 0001, Jianwei Zheng 0001, Shengyong Chen |
IEEE Trans. Cybern. | 4 |
| 2022 | ICIF-Net: Intra-Scale Cross-Interaction and Inter-Scale Feature Fusion Network for Bitemporal Remote Sensing Images Change DetectionabstractChange detection (CD) of remote sensing (RS) images has enjoyed remarkable success by virtue of convolutional neural networks (CNNs) with promising discriminative capabilities. However, CNNs lack the capability of modeling long-range dependencies in bitemporal image pairs, resulting in inferior identifiability against the same semantic targets yet with varying features. The recently thriving Transformer, on the contrary, is warranted, for practice, with global receptive fields. To jointly harvest the local-global features and circumvent the misalignment issues caused by step-by-step downsampling operations in traditional backbone networks, we propose an intra-scale cross-interaction and inter-scale feature fusion network (ICIF-Net), explicitly tapping the potential of integrating CNN and Transformer. In particular, the local features and global features, respectively, extracted by CNN and Transformer, are interactively communicated at the same spatial resolution using a linearized Conv Attention module, which motivates the counterpart to glimpse the representation of another branch while preserving its own features. In addition, with the introduction of two attention-based inter-scale fusion schemes, including mask-based aggregation and spatial alignment (SA), information integration is enforced at different resolutions. Finally, the integrated features are fed into a conventional change prediction head to generate the output. Extensive experiments conducted on four CD datasets of bitemporal (RS) images demonstrate that our ICIF-Net surpasses the other state-of-the-art (SOTA) approaches. Yuchao Feng, Honghui Xu 0002, Jiawei Jiang 0002, Jianwei Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | Tensor completion using patch-wise high order Hankelization and randomized tensor ring initialization
Jianwei Zheng 0001, Honghui Xu 0002, Yuchao Feng, Peijun Chen, Shengyong Chen |
Eng. Appl. Artif. Intell. | 1 |
| 2021 | Enhanced low-rank constraint for temporal subspace clustering and its acceleration scheme
Jianwei Zheng 0001, Guojiang Shen, Shengyong Chen |
Pattern Recognit. | 1 |
| 2021 | Adversarial robustness via attention transfer
Zhuorong Li, Chao Feng 0010, Minghui Wu 0001, Hongchuan Yu, Jianwei Zheng 0001, Fanwei Zhu |
Pattern Recognit. Lett. | 5 |
| 2021 | Hyperspectral Image Classification Using Mixed Convolutions and Covariance PoolingabstractRecently, convolution neural network (CNN)-based hyperspectral image (HSI) classification has enjoyed high popularity due to its appealing performance. However, using 2-D or 3-D convolution in a standalone mode may be suboptimal in real applications. On the one hand, the 2-D convolution overlooks the spectral information in extracting feature maps. On the other hand, the 3-D convolution suffers from heavy computation in practice and seems to perform poorly in scenarios having analogous textures along with consecutive spectral bands. To solve these problems, we propose a mixed CNN with covariance pooling for HSI classification. Specifically, our network architecture starts with spectral-spatial 3-D convolutions that followed by a spatial 2-D convolution. Through this mixture operation, we fuse the feature maps generated by 3-D convolutions along the spectral bands for providing complementary information and reducing the dimension of channels. In addition, the covariance pooling technique is adopted to fully extract the second-order information from spectral-spatial feature maps. Motivated by the channel-wise attention mechanism, we further propose two principal component analysis (PCA)-involved strategies, channel-wise shift and channel-wise weighting, to highlight the importance of different spectral bands and recalibrate channel-wise feature response, which can effectively improve the classification accuracy and stability, especially in the case of limited sample size. To verify the effectiveness of the proposed model, we conduct classification experiments on three well-known HSI data sets, Indian Pines, University of Pavia, and Salinas Scene. The experimental results show that our proposal, although with less parameters, achieves better accuracy than other state-of-the-art methods. Jianwei Zheng 0001, Yuchao Feng, Cong Bai, Jinglin Zhang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Hyperspectral Image Restoration via Local Low-Rank Matrix Recovery and Moreau-Enhanced Total VariationabstractIn this letter, we present a hyperspectral image (HSI) mixed-noise removal method named Moreau-enhanced total variation (TV) regularized local low-rank matrix recovery (LLRMTV). The rank-fixed matrix recovery is first adopted to separate the low-rank clean HSI patches from the sparse noise. Then, a Moreau-enhanced TV regularized image reconstruction strategy is utilized to ensure the piecewise smoothness of the reconstructed image from the low-rank patches. The proposed Moreau-enhanced TV restoration method involves a nonconvex penalty designed to maintain the convexity of the objective function. Moreover, the proposed model is integrated into an augmented Lagrange multiplier (ALM) algorithm to produce final results, leading to a complete HSI restoration framework. Examples of restoration illustrate the improvement over the typical TV regularization. Yanhong Yang, Jianwei Zheng 0001, Shengyong Chen |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2020 | Local low-rank matrix recovery for hyperspectral image denoising with ℓ0 gradient constraint
Yanhong Yang, Jianwei Zheng 0001, Shengyong Chen |
Pattern Recognit. Lett. | 2 |
| 2020 | Efficient Implementation of Truncated Reweighting Low-Rank Matrix ApproximationabstractThe weighted nuclear norm minimization and truncated nuclear norm minimization are two well-known low-rank constraint for visual applications. In this paper, by integrating their advantages into a unified formulation, we find a better weighting strategy, namely truncated reweighting norm minimization (TRNM), which provides better approximation to the target rank for some specific task. Albeit nonconvex and truncated, we prove that TRNM is equivalent to certain weighted quadratic programming problems, whose global optimum can be accessed by the newly presented reweighting singular value thresholding operator. More importantly, we design a computationally efficient optimization algorithm, namely momentum update and rank propagation (MURP), for the general TRNM regularized problems. The individual advantages of MURP include, first, reducing iterations through nonmonotonic search, and second, mitigating computational cost by reducing the size of target matrix. Furthermore, the descent property and convergence of MURP are proven. Finally, two practical models, i.e., Matrix Completion Problem via TRNM (MCTRNM) and Space Clustering Model via TRNM (SCTRNM), are presented for visual applications. Extensive experimental results show that our methods achieve better performance, both qualitatively and quantitatively, compared with several state-of-the-art algorithms. Jianwei Zheng 0001, Xiaolong Zhou 0001, Jiafa Mao, Hongchuan Yu |
IEEE Trans. Ind. Informatics | 1 |
| 2020 | Truncated Low-Rank and Total p Variation Constrained Color Image Completion and its Moreau Approximation AlgorithmabstractRecently, low-rank (LR) and total variation (TV) constrained tensor completion algorithms have been broadly studied for image restoration. These algorithms, however, ignore the difference of the intrinsic properties along spatial structure, spectral correlation, and unfolded mode. In this paper, we go further by providing a detailed comparison of the LR and TV properties in matrix and tensor cases, and figure out the LRTV constraints for pixel matrices are more evident and accordant than for others. This inspires us to develop a simple yet effective multichannel LRTV model that is capable of genuinely discovering the intrinsic properties with reduced computational cost. Moreover, due to the suboptimality of nuclear norm and$l_{1}$norm in approximating the essential low rank and low gradient properties, we employ two enhanced constraints, i.e., truncated nuclear norm (TNN) and total$p$variation${\text{T}}_{p}\text{V}$, for a better performance. This results in a challenging problem since that both TNN and${\text{T}}_{p}\text{V}$are nonsmooth and nonconvex. Observing that the Moreau approximation of${\text{T}}_{p}\text{V}$constraint is a continuous difference-of-convex function, we then develop a first-order method by repeatedly computing two simple proximal operators. Under mild assumption, we further prove that the sequence generated by our method clusters at a stationary point. Extensive experimental results on color image completion show the efficacy and efficiency of our method over state-of-the-art competitors. Jianwei Zheng 0001, Xi Yang 0006, Shengyong Chen |
IEEE Trans. Image Process. | 1 |
| 2019 | Double Weighted Low-Rank Representation and Its Efficient Implementation
Jianwei Zheng 0001, Kechen Lou, Wanliang Wang |
PAKDD (2) | 1 |
| 2019 | Weighted Mixed-Norm Regularized Regression for Robust Face IdentificationabstractFace identification (FI) via regression-based classification has been extensively studied during the recent years. Most vector-based methods achieve appealing performance in handing the noncontiguous pixelwise noises, while some matrix-based regression methods show great potential in dealing with contiguous imagewise noises. However, there is a lack of consideration of the mixture noises case, where both contiguous and noncontiguous noises are jointly contained. In this paper, we propose a weighted mixed-norm regression (WMNR) method to cope with the mixture image corruption. WMNR reveals certain essential characteristics of FI problems and bridges the vector- and matrix-based methods. Particularly, WMNR provides two advantages for both theoretical analysis and practical implementation. First, it generalizes possible distributions of the residuals into a unified feature weighted loss function. Second, it constrains the residual image as low-rank structure that can be quantified with general nonconvex functions and a weight factor. Moreover, a new reweighted alternating direction method of multipliers algorithm is derived for the proposed WMNR model. The algorithm exhibits great computational efficiency since it divides the original optimization problem into certain subproblems with analytical solution or can be implemented in a parallel manner. Extensive experiments on several public face databases demonstrate the advantages of WMNR over the state-of-the-art regression-based approaches. More specifically, the WMNR achieves an appealing tradeoff between identification accuracy and computational efficiency. Compared with the pure vector-based methods, our approach achieves more than 10% performance improvement and saves more than 70% of runtime, especially in severe corruption scenarios. Compared with the pure matrix-based methods, although it requires slightly more computation time, the performance benefits are even larger; up to 20% improvement can be obtained. Jianwei Zheng 0001, Kechen Lou, Xi Yang 0006, Cong Bai, Jinhui Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | An Efficient Truncated Nuclear Norm Constrained Matrix Completion For Image InpaintingabstractThe matrix completion problem has found many applications in image and graphics fields. Developing fast and exact algorithms still remains challenging. The truncated nuclear norm (TNN), taking advantage of priori target rank information, is known as a better surrogate function to the rank constraint than the traditional nuclear norm. However, the TNN penalized algorithms always converge slowly due to the two-step scheme for avoiding directly updating the noncovex functions. In this paper, we propose a computationally efficient algorithm, Momentum Adaptive and Rank Revealing (MARR), for the TNN regularized matrix completion problem. The distinct advantages are to (1) reduce iterations by introducing non-monotonic constraints, and (2) decrease computational burden by controlling the size of matrix. Moreover, the descent property and convergence of the search are proven. Experiments on image inpainting, including images and range data, show that the proposed algorithm achieves very competitive results visually and numerically compared to several state-of-the-art approaches, while providing substantial reduction of iterations and runtime, thereby more applicable in real-world problems. Jianwei Zheng 0001, Hongchuan Yu, Wanliang Wang |
CGI | 1 |
| 2018 | Optimization of deep convolutional neural network for large scale image retrieval
Cong Bai, Ling Huang 0003, Jianwei Zheng 0001, Shengyong Chen |
Neurocomputing | 4 |
| 2018 | Kernel group sparse representation classifier via structural and non-convex constraints
Jianwei Zheng 0001, Hong Qiu, Weiguo Sheng 0001, Xi Yang 0006, Hongchuan Yu |
Neurocomputing | 1 |
| 2017 | Iterative Re-Constrained Group Sparse Face Recognition With Adaptive Weights LearningabstractIn this paper, we consider the robust face recognition problem via iterative re-constrained group sparse classifier (IRGSC) with adaptive weights learning. Specifically, we propose a group sparse representation classification (GSRC) approach in which weighted features and groups are collaboratively adopted to encode more structure information and discriminative information than other regression based methods. In addition, we derive an efficient algorithm to optimize the proposed objective function, and theoretically prove the convergence. There are several appealing aspects associated with IRGSC. First, adaptively learned weights can be seamlessly incorporated into the GSRC framework. This integrates the locality structure of the data and validity information of the features into l2,p-norm regularization to form a unified formulation. Second, IRGSC is very flexible to different size of training set as well as feature dimension thanks to the l2,p-norm regularization. Third, the derived solution is proved to be a stationary point (globally optimal if p ≥ 1). Comprehensive experiments on representative data sets demonstrate that IRGSC is a robust discriminative classifier which significantly improves the performance and efficiency compared with the state-of-the-art methods in dealing with face occlusion, corruption, and illumination changes, and so on. Jianwei Zheng 0001, Shengyong Chen, Guojiang Shen, Wanliang Wang |
IEEE Trans. Image Process. | 1 |
| 2016 | Kernel-based discriminative elastic embedding algorithm
Jianwei Zheng 0001, Hong Qiu, Wanliang Wang, Chenchen Kong, Hailun Wang |
Appl. Intell. | 1 |
| 2015 | Efficient kernel discriminative common vectors for classification
Jianwei Zheng 0001, Qiongfang Huang, Shengyong Chen, Wanliang Wang |
Vis. Comput. | 1 |
| 2014 | Incremental min-max projection analysis for classification
Jianwei Zheng 0001, Shengyong Chen, Wanliang Wang |
Neurocomputing | 1 |
| 2013 | Kernel-Based Manifold-Oriented Stochastic Neighbor Projection MethodabstractA new method for performing a nonlinear form of manifold-oriented stochastic neighbor projection method is proposed. By the use of kernel functions, one can operate in the feature space without ever computing the coordinates of the data in that space, but rather by simply computing the inner products between the images of all pairs of data in the feature space. The proposed method is termed as kernel-based manifoldoriented stochastic neighbor projection(KMSNP). By two different strategies, KMSNP is divided into two methods: KMSNP1 and KMSNP2. Experimental results on several databases show that, compared with the relevant methods, the proposed methods obtain higher classification performance and recognition rate. INTRODUCTION Kernel-based methods(kernel methods for short) have become a new hot topic in machine learning fields in recent years, their theoretical basis is statistical learning theory. Kernel methods are a class of algorithms for pattern analysis, whose best known element is the support vector machine(SVM) (Dardas and Georganas 2011).The methods skillfully introduce kernel function which not only reduces the curse of dimensionality (Cherchi and Guevara 2012, Xue et al. 2012), but also effectively solves the local minimum and incomplete statistical analysis in traditional pattern recognition methods on the premise of no additional computational capacity. As an availability way to resolve the problem of nonlinear pattern recognition, kernel methods approach the problem by mapping the data into a highdimensional feature space, where each coordinate corresponds to one feature of the data items, transforming the data into a set of points in a Euclidean space (Chen and Li 2011, Zhang et al. 2008). The theory of kernel methods can be traced back to 1909, Mercer proposed Mercer's theorem (Mercer 1909) which indicates that any ‘reasonable’ kernel function corresponds to some feature space. 1964, the use of Mercer's theorem for interpreting kernels as inner products in a feature space was introduced into machine learning by Aizerman et al. (AizermanI et al. 1964), but no sufficient importance has been attached to it. Until 1992, Vapnik et al. (Boser et al. 1992) successfully extended the SVM to the non-linear SVM by using kernel functions, it began to show its potential and advantages. Subsequently, more and more kernelbased methods were presented, such as: kernel principal component analysis(KPCA) (Xiao et al. 2012), kernel fisher discriminator(KFD) (Yang et al. 2005), kernel independent component analysis (KICA) (Zhang et al. 2013), kernel partial least squares(KPLS) (Helander et al. 2012) and so on. In this paper, we propose to use the kernel idea and present a method called kernel-based manifoldoriented stochastic neighbor projection(KMSNP) method through improving the manifold-oriented stochastic neighbor projection(MSNP) (Wu et al. 2011) technique. MSNP is based on stochastic neighbor embedding(SNE) (Hinton and Roweis 2002) and t-SNE (Maaten and Hinton 2008). The basic principle of SNE is to convert pairwise Euclidean distances into probabilities of selecting neighbors to model pairwise similarities while t-SNE uses student t-distribution to model pairwise dissimilarities in low-dimensional space. Different from SNE and t-SNE, MSNP converts pairwise dissimilarities of inputs to probability distribution related to geodesic distance in highdimensional space and uses Cauchy distribution to model stochastic distribution of features. Furthermore, it recovers the manifold structure through a linear projection by requiring the two distributions to be similar. Experiments demonstrate MSNP has unique advantages in terms of visualization and recognition task, but there are still two drawbacks in it: firstly, MSNP is an unsupervised method and lack of the idea of class label, so it is not suitable for pattern identification; secondly, since MSNP is a linear feature dimensionality reduction algorithm, it cannot effectively settle the nonlinear feature extraction problem. To overcome the disadvantages of MSNP, we have done some preliminary work. On the first, we introduced the idea of class label and presented a method called discriminative stochastic neighbor embedding analysis(DSNE) (Zheng et al. 2012, Chen Proceedings 27th European Conference on Modelling and Simulation ©ECMS Webjorn Rekdalsbakken, Robin T. Bye, Houxiang Zhang (Editors) ISBN: 978-0-9564944-6-7 / ISBN: 978-0-9564944-7-4 (CD) and Wang 2012). On the second, we think KMSNP can overcome the disadvantage mentioned above well. The rest of this paper is organized as follows: in Section 2, we provide a brief review of MSNP. Section 3 describes the detailed algorithm derivation of KMSNP. Furthermore, experiments on various databases are presented in Section 4. Finally, we provide some concluding remarks and describe several issues for future works in Section 5. MSNP Considering the problem of representing d-dimensional data vectors x1, x2, . . . , xN, by r-dimensional (r << d) vectors y1, y2, . . ., yN such that yi represents xi. The basic principle of MSNP is to convert pairwise dissimilarity of inputs to probability distribution related to geodesic distance in high-dimensional space, and then using Cauchy distribution to model stochastic distribution of features, finally, MSNP recovers the manifold structure through a linear projection by requiring the two distributions to be similar. Mathematically, the similarity of datapoint xi to datapoint xj is depicted as the following joint probability pij which means xi how possible to pick xj as its neighbor: exp( / 2) exp( / 2) geo ij ij geo ik k i D p D (1) where Dij is the geodesic distance for xi and xj. In practice, MSNP calculates geodesic distance by using a two-phase method (Wu et al. 2011). Firstly, an adjacency graph G is constructed by K-nearest neighbor strategy. Secondly, the desired geodesic distance is approximated by the shortest path of graph G. This procedure is proposed in Isomap to estimate geodesic distance and the detail calculation steps can be found in (Tenenbaum et al. 2000). For low-dimensional representations, MSNP employs Cauchy distribution with degree of freedom to construct joint probability qij. The probability qij indicates how possible point i and point j can be stochastic neighbors is defined as: Jianwei Zheng 0001, Hong Qiu, Qiongfang Huang, Wanliang Wang, Xinli Xu |
ECMS | 1 |