Guixu Zhang

dblp:44/7527 · DBLP profile ↗
← Back
125ranked-venue papers
1as first author
64since 2021 · last 2026
0000-0003-4720-6607ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 62 · 1 first-author · 30 since 2021Artificial intelligence and machine learning · 56 · 29 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 15 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 4
YearPublicationVenuePosition
2026 A Geometric Perspective on Optimizing Vector Quantized Latent Diffusion Model for Image Restoration
abstract
In this paper, we investigate the limitations of the Vector Quantized Latent Diffusion Model (VQ-LDM) in restoration tasks. We identify a performance gap between the Vector Quantization (VQ) and Diffusion Model components, manifested as a significant discrepancy between the reconstruction quality of ground truth images processed via VQ autoregression and degraded images restored by VQ-LDM. Through experiments, we attribute this gap primarily to the lack of robustness in the mapped points of VQ within the original VQ-LDM framework. To address this issue, we propose a geometric based optimization approach. First, we introduce a simple yet effective method, termed interpolation-based latent initial state optimization, which mitigates the performance gap by replacing the original mapped points with interpolated values, supported by theoretical analysis. Here, the latent initial state refers specifically to the input of the diffusion model. Building upon this, we further propose a Chebyshev center-based latent initial state optimization, an elegant theoretical solution from a geometric perspective, that further enhances restoration performance. Our improvements consistently achieve superior results across nine benchmark datasets.
Chen Hang, Haoming Chen, Xuwei Fang, Weisheng Xie, Xiangxiang Gao, Faming Fang, Guixu Zhang
AAAI7
2026 RPE-PAD: Relative Pose Estimation for Pose-agnostic Anomaly Detection
abstract
Pose-agnostic Anomaly Detection (PAD) aims to detect anomalies when the poses of query images are unknown and differ from those in the training set. Therefore, accurately estimating the camera poses for the query images in the test set is critical for this task. Existing query-specific framework methods require re-optimizing a new set of parameters for each query image, limiting their generalization and increasing computational burden. To overcome these limitations, we propose a novel method, Relative Pose Estimation for Pose-agnostic Anomaly Detection (RPE-PAD), which enhances both generalization and efficiency with a query-independent framework. Specifically, we propose a Random View Synthesis Scheme (RVSS) that generates new poses by adding Gaussian perturbations to the original poses, then renders the corresponding views to augment the dataset. To estimate the relative camera pose between two input images, we introduce an Iterative Relative Pose Refinement Network (IRPRN), which incorporates a hierarchical coarse-to-fine refinement strategy. Furthermore, we employ a Multi-Pair Training Strategy (MPTS) to train the proposed IRPRN, leveraging multiple image pairs to expand the relative pose transformation space during training. Extensive experiments demonstrate that our method achieves robust anomaly detection performance while significantly improving inference efficiency.
Mengzan Qi, Rongkang Ma, Yingying Fang, Guixu Zhang, Tieyong Zeng, Zhi Li 0080
AAAI5
2026 Deep Algorithm Unrolling with Alignment Embedding for Guided Image Super-resolution
Faming Fang, Tingting Wang 0007, Junkang Zhang, Aimin Zhou, Riquan Zhang, Guixu Zhang
Int. J. Comput. Vis.7
2026 Eliciting CLIP's intrinsic attribute knowledge through a dual-cache guided mechanism for class-incremental learning
Shengcheng Ye, Yaomin Huang, Faming Fang, Guixu Zhang
Knowl. Based Syst.5
2026 Task-aware all-in-one guided image super-resolution
Tingting Wang 0007, Jun Wang 0024, Qiuhai Yan, Junkang Zhang, Faming Fang, Guixu Zhang
Pattern Recognit.6
2026 Deep Unfolding Segmentation Network for Under-Sampled Magnetic Resonance Images
abstract
Magnetic Resonance (MR) image segmentation is a critical task in assisting disease diagnosis. Most existing methods assume that the images being segmented are fully-sampled. However, they ignore the fact that MR images obtained in clinics are often reconstructed from under-sampled k-space data. There are artifacts or distorted details in the reconstruction, leading to unsatisfactory segmentation performance. In this paper, we propose an end-to-end deep unfolding framework to segment desired lesions or organs from the under-sampled k-space data. Specifically, we build a new model to combine the compressive sensing-based under-sampled image reconstruction and level-set-based segmentation. In this model, we introduce an L0 norm on the reconstruction images to enforce smoothing while preserving important edge and boundary, boosting downstream segmentation performance. We employ the Augmented Lagrangian Method to seek the solution and unfold the iterative algorithm into a deep neural network, called deep unfolding segmentation network (DUSNet). To further enhance segmentation performance, we introduce a boundary loss function, which encourages the model to effectively capture edge details of the regions of interest and imposes geometric constraints on the segmentation results. Through end-to-end training, DUSNet can efficiently segment target regions from under-sampled k-space data. Comprehensive experiments demonstrate that the proposed DUSNet outperforms existing state-of-the-art methods for under-sampled MR image segmentation, achieving superior segmentation accuracy.
Le Hu, Pengcheng Lei, Faming Fang, Guixu Zhang
IEEE J. Biomed. Health Informatics5
2025 Decoupling Scattering: Pseudo-Label Guided NeRF for Scenes with Scattering Media
abstract
Neural Radiance Fields (NeRF) has been widely used in computer vision and graphics, achieving impressive results in novel view synthesis and multi-view 3D reconstruction. However, despite its excellent performance under ideal conditions, NeRF struggles in challenging environments such as hazy, foggy, and underwater scenes, primarily due to the difficulty in decoupling objects from the scattering medium. To mitigate this limitation, we proposed a novel approach for NeRF in scenes with scattering media. Specifically, we leverage pseudo-labels during the early stage of training to guide NeRF in decoupling the densities of objects and the scattering medium, guiding the model toward a more appropriate search space. Furthermore, we introduce a Cyclical Progressive Dimensional Optimization Strategy (CPDOS) that focuses on optimizing a single or a few variables during specific periods. Experimental results demonstrate that our method can effectively simulate hazy and underwater scenes, accurately decouple the scattering medium from objects, estimate atmospheric parameters, and outperform existing methods in novel view synthesis and image restoration tasks.
Junkang Zhang, Faming Fang, Guixu Zhang
AAAI4
2025 First-order State Space Model for Lightweight Image Super-resolution
abstract
State space models (SSMs), particularly Mamba, have shown promise in NLP tasks and are increasingly applied to vision tasks. However, most Mamba-based vision models focus on network architecture and scan paths, with little attention to the SSM module. In order to explore the potential of SSMs, we modified the calculation process of SSM without increasing the number of parameters to improve the performance on lightweight super-resolution tasks. In this paper, we introduce the First-order State Space Model (FSSM) to improve the original Mamba module, enhancing performance by incorporating token correlations. We apply a first-order hold condition in SSMs, derive the new discretized form, and analyzed cumulative error. Extensive experimental results demonstrate that FSSM improves the performance of MambaIR on five benchmark datasets without additionally increasing the number of parameters, and surpasses current lightweight SR methods, achieving state-of-the-art results.
Yekai Lu, Guang Yang 0068, Faming Fang, Guixu Zhang
ICASSP6
2025 Surface-Aware Feed-Forward Quadratic Gaussian for Frame Interpolation with Large Motion
abstract
Motion in the real world takes place in 3D space. Existing Frame Interpolation methods often estimate global receptive fields in 2D frame space. Due to the limitations of 2D space, these global receptive fields are limited, which makes it difficult to match object correspondences between frames, resulting in sub-optimal performance when handling large-motion scenarios. In this paper, we introduce a novel pipeline for exploring object correspondences based on differential surface theory. The differential surface coordinate system provides a better representation of the real world, enabling effective exploration of object correspondences. Specifically, the pipeline first transforms an input pair of video frames from the image coordinate system to the differential surface coordinate system. Subsequently, within this coordinate system, object correspondences are explored based on surface geometric properties and the surface uniqueness theorem. Experimental findings showcase that our method attains state-of-the-art performance across large motion benchmarks. Our method demonstrates the state-of-the-art performance on these VFI subsets with large motion.
Zaoming Yan, Yaomin Huang, Pengcheng Lei, Qizhou Chen, Guixu Zhang, Faming Fang
NeurIPS5
2025 Deep maximum a posterior estimator for accelerated MRI reconstruction
Tingting Wang 0007, Shengcheng Ye, Faming Fang, Guixu Zhang, Yuanyi Zheng
Knowl. Based Syst.4
2025 Digging Deeper in Gradient for Unrolling-Based Accelerated MRI Reconstruction
abstract
There are two main methods that can be used to accelerate MRI reconstruction: parallel imaging and compressed sensing. To further accelerate the sampling process, the combination of these two methods has been extensively studied in recent years. However, existing MRI reconstruction methods often overlook the exploration of high-frequency information of images, leading to sub-optimal recovery of fine details in the reconstructed results. To address this issue, we conduct an in-depth analysis of image gradients and propose a novel MRI reconstruction model based on Maximum a Posteriori (MAP) estimation. We first establish the Cumulative Deviation from Maximum Gradient magnitude (CDMG) prior for fully sampled MR images through theoretical analysis, then incorporate this explicit CDMG prior along with an implicit deep prior to form the prior probability term. This combination of priors strikes a balance between physically informed constraints and data-driven adaptability, aiding in the recovery of meaningful high-frequency information. Additionally, we introduce a multi-order gradient operator to enhance the observation model, thereby improving the accuracy of the likelihood term. Through MAP estimation, we develop a novel accelerated MRI reconstruction model, the optimization of which is achieved by unrolling it into a convolutional neural network structure, referred to as DDGU-Net. Extensive experimental results demonstrate the effectiveness of our approach in reconstructing high-quality MR images and achieving state-of-the-art (SOTA) results, particularly at higher acceleration factors.
Faming Fang, Tingting Wang 0007, Guixu Zhang, Fang Li 0004
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Homography Estimation With Adaptive Query Transformer and Gated Interaction Module
abstract
Homography estimation is essential for aligning images captured from different viewpoints by accurately modeling the geometric relationship between them. In homography estimation, global information plays a critical role. To establish global correspondences, cross-attention has been widely used in recent studies. However, vanilla cross-attention mechanisms treat queries in redundant and low-texture areas the same as those in richly textured areas, leading to the accumulation and propagation of erroneous information. We define this phenomenon, where the model excessively attends to queries in redundant and low-texture areas, as query over-focusing. To alleviate query over-focusing and achieve fine-grained homography estimation, we propose a novel homography estimation network, termed AGNet, which integrates an Adaptive Query Transformer (AQFormer) and a Gated Interaction Module (GIM). The AQFormer is designed to dynamically adjust attention by applying a mask to queries, allowing the model to adaptively emphasize feature-rich regions while suppressing redundant or weakly textured areas. Meanwhile, the GIM selectively captures local information by adjusting convolutional kernels based on input, enhancing the extraction of shared features between image pairs. Extensive experiments on various datasets demonstrate that AGNet significantly improves accuracy in homography estimation, particularly in challenging scenarios with low overlap and large viewpoint variations.
Faming Fang, Tingting Wang 0007, Guixu Zhang
IEEE Trans. Circuits Syst. Video Technol.4
2025 Robust Deep Convolutional Dictionary Model With Alignment Assistance for Multi-Contrast MRI Super-Resolution
abstract
Multi-contrast magnetic resonance imaging (MCMRI) super-resolution (SR) methods aims to leverage the complementary information present in multi-contrast images. However, existing methods encounter several limitations. Firstly, most current networks fail to appropriately model the correlations of multi-contrast images and lack certain interpretability. Secondly, they often overlook the negative impact of spatial misalignment between modalities in clinical practice. Thirdly, existing methods do not effectively constrain the complementary information learned between multi-contrast images, resulting in information redundancy and limiting their model performance. In this paper, we propose a robust alignment-assisted multi-contrast convolutional dictionary (A2-CDic) model to address these challenges. Specifically, we develop an observation model based on convolutional sparse coding to explicitly represent multi-contrast images as common (e.g., consistent textures) and unique (e.g., inconsistent structures and contrasts) components. Considering there are spatial misalignments in real-world multi-contrast images, we incorporate a spatial alignment module to compensate for the misaligned structures. This approach enables the proposed model to fully exploit the valuable information in the reference image while mitigating interference from inconsistent information. We employ the proximal gradient algorithm to optimize the model and unroll the iterative steps into a multi-scale convolutional dictionary network. Furthermore, we utilize mutual information losses to constrain the extracted common and unique components. This constraint reduces the redundancy between the decomposed components, allowing each sub-module to learn more representative features. We evaluate our model on four publicly available datasets comprising internal, external, spatially aligned, and misaligned MCMRI images. The experimental results demonstrate that our model surpasses existing state-of-the-art MCMRI SR methods in terms of both generalization ability and overall performance. Code is available at https://github.com/lpcccc-cv/A2-CDic.
Pengcheng Lei, Miaomiao Zhang 0002, Faming Fang, Guixu Zhang
IEEE Trans. Medical Imaging4
2025 Auxiliary Representation Guided Network for Visible-Infrared Person Re-Identification
abstract
Visible-Infrared Person Re-identification aims to retrieve images of specific identities across modalities. To relieve the large cross-modality discrepancy, researchers introduce the auxiliary modality within the image space to assist modality-invariant representation learning. However, the challenge persists in constraining the inherent quality of generated auxiliary images, further leading to a bottleneck in retrieval performance. In this paper, we propose a novel Auxiliary Representation Guided Network (ARGN) to explore the potential of auxiliary representations, which are directly generated within the modality-shared embedding space. In contrast to the original visible and infrared representations, which contain information solely from their respective modalities, these auxiliary representations integrate cross-modality information by fusing both modalities. In our framework, we utilize these auxiliary representations as modality guidance to reduce the cross-modality discrepancy. First, we propose a High-quality Auxiliary Representation Learning (HARL) framework to generate identity-consistent auxiliary representations. The primary objective of our HARL is to ensure that auxiliary representations capture diverse modality information from both modalities while concurrently preserving identity-related discrimination. Second, guided by these auxiliary representations, we design an Auxiliary Representation Guided Constraint (ARGC) to optimize the modality-shared embedding space. By incorporating this constraint, the modality-shared embedding space is optimized to achieve enhanced intra-identity compactness and inter-identity separability, further improving the retrieval performance. In addition, to improve the robustness of our framework against the modality variation, we introduce a Part-based Adaptive Gaussian Module (PAGM) to adaptively extract discriminative information across modalities. Finally, extensive experiments are conducted to demonstrate the superiority of our method over state-of-the-art approaches on three VI-ReID datasets.
Mengzan Qi, Sixian Chan 0001, Chen Hang, Guixu Zhang, Tieyong Zeng, Zhi Li 0080
IEEE Trans. Multim.4
2025 FrDiff: Framelet-Based Conditional Diffusion Model for Multispectral and Panchromatic Image Fusion
abstract
The process of fusing low-resolution multispectral (LRMS) and high-resolution panchromatic (PAN) imagery, commonly referred to as pansharpening, is intended to generate high-resolution multispectral (HRMS) imagery. Typically, most pre-existing pansharpening frameworks mainly emphasize the straightforward learning of the mapping relationship among PAN and LRMS images to HRMS images. However, a key limitation of these frameworks is their potential overemphasis on spatial information, particularly the enhancement of low-frequency components. As a result, such an oversight potentially hinders the model's ability to simultaneously restore both spectral and spatial details. To address this issue, we propose a novel pansharpening model based on the denoising diffusion probabilistic model (DDPM), dubbed FrDiff. Specifically, we build a framelet-based conditional diffusion model that leverages the generative power of diffusion models to produce more refine results. Different from conventional methods directly inferring HRMS images, our strategy is designed to project their framelet coefficients, utilizing the available PAN and LRMS images as resources. This approach enables the separation of high-frequency and low-frequency components through framelet transformation, which are subsequently recombined to create a novel set of conditional embeddings that feed into the diffusion process. At the same time, the powerful predictive power of the diffusion model is exploited to simultaneously recover the high-frequency and low-frequency components of the HRMS. Moreover, we introduce a framelet-oriented cross-attention module dedicated to honing spectral fidelity. This module is crucial for improving the spectral precision of the HRMS images, ensuring a balanced emphasis on both spatial and spectral enhancements. Quantitative and qualitative experiments on multiple benchmark datasets demonstrate that the proposed method achieves more robustness and high-quality results than other state-of-the-art pansharpening methods.
Junkang Zhang, Faming Fang, Tingting Wang 0007, Guixu Zhang
IEEE Trans. Multim.4
2024 Triple Feature Disentanglement for One-Stage Adaptive Object Detection
abstract
In recent advancements concerning Domain Adaptive Object Detection (DAOD), unsupervised domain adaptation techniques have proven instrumental. These methods enable enhanced detection capabilities within unlabeled target domains by mitigating distribution differences between source and target domains. A subset of DAOD methods employs disentangled learning to segregate Domain-Specific Representations (DSR) and Domain-Invariant Representations (DIR), with ultimate predictions relying on the latter. Current practices in disentanglement, however, often lead to DIR containing residual domain-specific information. To address this, we introduce the Multi-level Disentanglement Module (MDM) that progressively disentangles DIR, enhancing comprehensive disentanglement. Additionally, our proposed Cyclic Disentanglement Module (CDM) facilitates DSR separation. To refine the process further, we employ the Categorical Features Disentanglement Module (CFDM) to isolate DIR and DSR, coupled with category alignment across scales for improved source-target domain alignment. Given its practical suitability, our model is constructed upon the foundational framework of the Single Shot MultiBox Detector (SSD), which is a one-stage object detection approach. Experimental validation highlights the effectiveness of our method, demonstrating its state-of-the-art performance across three benchmark datasets.
Haoan Wang, Shilong Jia, Tieyong Zeng, Guixu Zhang, Zhi Li 0080
AAAI4
2024 Harmonizing Knowledge Transfer in Neural Network with Unified Distillation
Yaomin Huang, Zaomin Yan, Chaomin Shen 0001, Faming Fang, Guixu Zhang
ECCV (33)5
2024 Three-Stage Temporal Deformable Network for Blurry Video Frame Interpolation
abstract
Blurry video frame interpolation (BVFI) aims to generate high-frame-rate clear videos from low-frame-rate blurry videos, is a challenging but important topic in the computer vision community. Blurry videos not only provide spatial and temporal information like clear videos, but also contain additional motion information hidden in each blurry frame. However, existing BVFI methods usually fail to fully leverage all valuable information, which ultimately hinders their performance. In this paper, we propose a simple three-stage temporal deformable network to fully explore useful information from blurry videos. The frame interpolation stage designs a deformable network to directly sample useful information from blurry inputs and synthesize an intermediate frame at an arbitrary time interval. The temporal feature fusion stage explores the long-term temporal information for each target frame through a bi-directional recurrent deformable alignment network. And the deblurring stage applies a transformer-empowered Taylor approximation network to recursively recover the high-frequency details. Quantitative and qualitative results indicate that our model outperforms existing SOTA methods.
Pengcheng Lei, Zaoming Yan, Tingting Wang 0007, Faming Fang, Guixu Zhang
ICME5
2024 HFF-Net: A High-Frequency Fidelity Model for Accelerated Parallel MRI Reconstruction
abstract
Magnetic Resonance Imaging (MRI) plays a crucial role in diagnosing and treating various diseases. However, the long acquisition time of MRI scans often leads to patient discomfort and motion artifacts. Consequently, accelerating MRI speed is essential. Researchers have combined Deep Learning with Compressed Sensing and Parallel Imaging to advance MRI. However, many existing methods fail to effectively recover the fine details and structures in Magnetic Resonance images. To address these challenges, we propose a novel model for accelerated parallel MRI reconstruction. Our model incorporates a high-frequency fidelity method into the reconstruction process, explicitly emphasizing the recovery of high-frequency information. Additionally, we consider the joint priori distribution between the reconstructed images from each coil. Using the variable splitting approach, the proposed model is unrolled as an end-to-end network termed HFF-Net. Experimental results demonstrate that our method outperforms state-of-the-art techniques, yielding high-quality MR images with enhanced detail and fine structure recovery.
Zhenggang Yang, Faming Fang, Qiaosi Yi, Guixu Zhang, Fang Li 0004
ICME4
2024 Purified Distillation: Bridging Domain Shift and Category Gap in Incremental Object Detection
abstract
Incremental Object Detection (IOD) simulates the dynamic data flow in real-world applications, which require detectors to learn new classes or adapt to new domains while retaining knowledge from previous tasks. Most existing IOD methods focus only on class incremental learning, assuming all data comes from the same domain. However, this is hardly achievable in practical applications, as images collected under different conditions often exhibit completely different characteristics, such as lighting, weather, style, etc. Class IOD methods suffer from performance degradation in these scenarios with domain shifts. To bridge domain shifts and category gaps in IOD, we propose Purified Distillation (PD), where we use a set of trainable queries to transfer the teacher's attention on old tasks to the student and adopt the gradient reversal layer to guide the student to learn the teacher's feature space structure from a micro perspective, which has not been extensively studied in previous works. Meanwhile, PD combines classification confidence with localization confidence to purify the most meaningful output nodes, so that the student model inherits a more comprehensive teacher knowledge. Extensive experiments across various IOD settings on six widely used datasets show that PD significantly outperforms state-of-the-art methods. Even after five steps of incremental learning, our method can preserve 60.6% mAP on the first task, while compared methods can only maintain up to 55.9%.
Shilong Jia, Tingting Wu 0001, Yingying Fang, Tieyong Zeng, Guixu Zhang, Zhi Li 0080
ACM Multimedia5
2024 Student-Oriented Teacher Knowledge Refinement for Knowledge Distillation
abstract
Knowledge distillation has become widely recognized for its ability to transfer knowledge from a large teacher network to a compact and more streamlined student network. Traditional knowledge distillation methods primarily follow a teacher-oriented paradigm that imposes the task of learning the teacher's complex knowledge onto the student network. However, significant disparities in model capacity and architectural design hinder the student's comprehension of the complex knowledge imparted by the teacher, resulting in sub-optimal performance. This paper introduces a novel perspective emphasizing student-oriented and refining the teacher's knowledge to better align with the student's needs, thereby improving knowledge transfer effectiveness. Specifically, we present the Student-Oriented Knowledge Distillation (SoKD), which incorporates a learnable feature augmentation strategy during training to refine the teacher's knowledge of the student dynamically. Furthermore, we deploy the Distinctive Area Detection Module (DAM) to identify areas of mutual interest between the teacher and student, concentrating knowledge transfer within these critical areas to avoid transferring irrelevant information. This customized module ensures a more focused and effective knowledge distillation process. Our approach, functioning as a plug-in, could be integrated with various knowledge distillation methods. Extensive experimental results demonstrate the efficacy and generalizability of our method.
Chaomin Shen 0001, Yaomin Huang, Haokun Zhu, Jinsong Fan, Guixu Zhang
ACM Multimedia5
2024 Exploring Fixed Point in Image Editing: Theoretical Support and Convergence Optimization
abstract
In image editing, Denoising Diffusion Implicit Models (DDIM) inversion has become a widely adopted method and is extensively used in various image editing approaches. The core concept of DDIM inversion stems from the deterministic sampling technique of DDIM, which allows the DDIM process to be viewed as an Ordinary Differential Equation (ODE) process that is reversible. This enables the prediction of corresponding noise from a reference image, ensuring that the restored image from this noise remains consistent with the reference image. Image editing exploits this property by modifying the cross-attention between text and images to edit specific objects while preserving the remaining regions. However, in the DDIM inversion, using the $t-1$ time step to approximate the noise prediction at time step $t$ introduces errors between the restored image and the reference image. Recent approaches have modeled each step of the DDIM inversion process as finding a fixed-point problem of an implicit function. This approach significantly mitigates the error in the restored image but lacks theoretical support regarding the existence of such fixed points. Therefore, this paper focuses on the study of fixed points in DDIM inversion and provides theoretical support. Based on the obtained theoretical insights, we further optimize the loss function for the convergence of fixed points in the original DDIM inversion, improving the visual quality of the edited image. Finally, we extend the fixed-point based image editing to the application of unsupervised image dehazing, introducing a novel text-based approach for unsupervised dehazing.
Chen Hang, Haoming Chen, Xuwei Fang, Vincent Xie, Faming Fang, Guixu Zhang
NeurIPS7
2024 Self-supervised medical slice interpolation network using controllable feature flow
Pengcheng Lei, Faming Fang, Tingting Wang 0007, Cong Liu 0011, Guixu Zhang
Expert Syst. Appl.5
2024 Few sampling meshes-based 3D tooth segmentation via region-aware graph convolutional network
Bodong Cheng, Najun Niu, Jun Wang 0024, Tieyong Zeng, Guixu Zhang, Jun Shi 0004, Juncheng Li 0003
Expert Syst. Appl.6
2024 HFGN: High-Frequency residual Feature Guided Network for fast MRI reconstruction
Faming Fang, Le Hu, Qiaosi Yi, Tieyong Zeng, Guixu Zhang
Pattern Recognit.6
2024 UGNet: Uncertainty aware geometry enhanced networks for stereo matching
Zhengkai Qi, Junkang Zhang, Faming Fang, Tingting Wang 0007, Guixu Zhang
Pattern Recognit.5
2024 WeaFU: Weather-Informed Image Blind Restoration via Multi-Weather Distribution Diffusion
abstract
The extraction of distribution from images with diverse weather conditions is crucial for enhancing the robustness of visual algorithms. When addressing image degradation caused by different weather, accurately perceiving the data distribution of weather-informed degradation becomes a fundamental challenge. However, given the highly stochastic nature, modelling weather distribution poses a formidable task. In this paper, we propose a novel multi-Weather distribution difFUsion blind restoration model, named WeaFU. Firstly, the model employs representation learning to map image distribution into a latent space. Subsequently, WeaFU utilizes a diffusion-based approach, with the assistance of Diffusion Distribution Generator (DDG), to perceive and extract corresponding weather distribution. This strategy ingeniously injects data distribution into the recovery process, significantly enhancing the robustness of the model in diverse weather scenarios. Finally, a Conditional Distribution-Aware Transformer (CDAT) is constructed to align the distribution information with pixels, thereby obtaining clear images. Extensive experiments on real and synthetic datasets demonstrate that WeaFU achieves superior performance.
Bodong Cheng, Juncheng Li 0003, Jun Shi 0004, Yingying Fang, Guixu Zhang, Tieyong Zeng, Zhi Li 0080
IEEE Trans. Circuits Syst. Video Technol.5
2024 MSCSCformer: Multiscale Convolutional Sparse Coding-Based Transformer for Pansharpening
abstract
With the increasing significance of high-quality, high-resolution multispectral images (HRMS) in various domains, pansharpening, which fuses low-resolution multispectral images (LRMS) with high-resolution panchromatic images (PAN), has gained considerable attention. However, current deep learning methods have limitations in capturing global long-range dependencies and incorporating spectral characteristics across different spectral bands of multispectral images (MS). Additionally, model-based approaches do not effectively utilize the multi-scale information between LRMS and HRMS data, limiting their further performance enhancement. To address these limitations, we propose a new observation model based on Multi-Scale Convolutional Sparse Coding (MS-CSC) and design a novel Multi-Scale Hybrid Spatial-spectral Transformer (MSHST) for the unfolding networks. The MS-CSC based observation model aims to fuse multi-scale information, while the MSHST incorporates spatial self-attention to capture global long-range dependencies and spectral self-attention to capture the inter-band correlation. Experimental results demonstrate the superiority of our method over other state-of-the-art approaches in both reduced-resolution and full-resolution evaluations. Ablation experiments further validate the effectiveness of the proposed multi-scale model and MSHST. Code is available at https://github.com/Eternityyx/MSCSCformer.
Yongxu Ye, Tingting Wang 0007, Faming Fang, Guixu Zhang
IEEE Trans. Geosci. Remote. Sens.4
2024 LSRFormer: Efficient Transformer Supply Convolutional Neural Networks With Global Information for Aerial Image Segmentation
abstract
Both local context and global context information are essential for the semantic segmentation of aerial images. Convolutional Neural Networks (CNNs) can capture local context information well but cannot model the global dependencies. Vision transformers (ViTs) are good at extracting global information but cannot retain the spatial details well. In order to leverage the advantages of these two paradigms, we integration them in one model in this study. However, global token interaction of ViT brings high computational cost, which makes it difficult to apply to large-sized aerial images. To handle this problem, we propose a novel efficient ViT block named long-short-range transformer (LSRFormer). Instead of mainstream ViTs designed as backbones, LSRFormer is a pre-training-free and plug-and-play module to be appended after CNN stages to supplement the global information. It is composed of long-range self-attention (LR-SA), short-range self-attention (SR-SA), and multi-scale-convolutional feed-forward-network (MSC-FFN). LR-SA establishes long-range dependencies at the junction of the windows and SR-SA diffuses the long-range information from window boundary to internal. MSC-FFN can capture multi-scale information inside the ViT block. We append LSRFormer block after each CNN stage of a pure convolutional network to build a model named ConvLSR-Net. Compared with existing models which combining CNN and ViTs, our model can learn both local and global representation at all stages of the model. In particular, ConvLSR-Net achieves state-of-the-art (SOTA) results on four challenging aerial image segmentation benchmarks, including iSAID, LoveDA, ISPRS Potsdam and Vaihingen. Code has been released at https://github.com/stdcoutzrh/ConvLSR-Net.
Renhe Zhang, Qian Zhang 0003, Guixu Zhang
IEEE Trans. Geosci. Remote. Sens.3
2024 Joint Under-Sampling Pattern and Dual-Domain Reconstruction for Accelerating Multi-Contrast MRI
abstract
Multi-Contrast Magnetic Resonance Imaging (MCMRI) utilizes the short-time reference image to facilitate the reconstruction of the long-time target one, providing a new solution for fast MRI. Although various methods have been proposed, they still have certain limitations. 1) existing methods featuring the preset under-sampling patterns give rise to redundancy between multi-contrast images and limit their model performance; 2) most methods focus on the information in the image domain, prior knowledge in the k-space domain has not been fully explored; and 3) most networks are manually designed and lack certain physical interpretability. To address these issues, we propose a joint optimization of the under-sampling pattern and a deep-unfolding dual-domain network for accelerating MCMRI. Firstly, to reduce the redundant information and sample more contrast-specific information, we propose a new framework to learn the optimal under-sampling pattern for MCMRI. Secondly, a dual-domain model is established to reconstruct the target image in both the image domain and the k-space frequency domain. The model in the image domain introduces a spatial transformation to explicitly model the inconsistent and unaligned structures of MCMRI. The model in the k-space learns prior knowledge from the frequency domain, enabling the model to capture more global information from the input images. Thirdly, we employ the proximal gradient algorithm to optimize the proposed model and then unfold the iterative results into a deep-unfolding network, called MC-DuDoN. We evaluate the proposed MC-DuDoN on MCMRI super-resolution and reconstruction tasks. Experimental results give credence to the superiority of the current model. In particular, since our approach explicitly models the inconsistent structures, it shows robustness on spatially misaligned MCMRI. In the reconstruction task, compared with conventional masks, the learned mask restores more realistic images, even under an ultra-high acceleration ratio ( ×30 ). Code is available at https://github.com/lpcccc-cv/MC-DuDoNet.
Pengcheng Lei, Le Hu, Faming Fang, Guixu Zhang
IEEE Trans. Image Process.4
2024 Flow Guidance Deformable Compensation Network for Video Frame Interpolation
abstract
Flow-based and deformable convolution (DConv)-based methods are two mainstream approaches for solving the video frame interpolation (VFI) problem, which have made remarkable progress with the development of deep convolutional networks over the past years. However, flow-based VFI methods often suffer from the inaccuracy of flow map estimation, especially in dealing with complex and irregular real-world motions. DConv-based VFI methods have advantages in handling complex motions, while the increased degree of freedom makes the training of the DConv model difficult. To address these problems, in this article, we propose a flow guidance deformable compensation network (FGDCN) for the VFI task. FGDCN decomposes the frame sampling process into two steps: a flow step and a deformation step. Specifically, the flow step utilizes a coarse-to-fine flow estimation network to directly estimate the intermediate flows and synthesizes an anchor frame simultaneously. To ensure the accuracy of the estimated flow, a distillation loss and a task-oriented loss are jointly employed in this step. Under the guidance of the flow priors learned in step one, the deformation step designs a new pyramid deformable compensation network to compensate for the missing details of the flow step. In addition, a pyramid loss is proposed to supervise the model in both the image and frequency domains. Experimental results show that the proposed algorithm achieves excellent performance on various datasets with fewer parameters.
Pengcheng Lei, Faming Fang, Tieyong Zeng, Guixu Zhang
IEEE Trans. Multim.4
2023 CP3: Channel Pruning Plug-in for Point-Based Networks
abstract
Channel pruning can effectively reduce both computational cost and memory footprint of the original network while keeping a comparable accuracy performance. Though great success has been achieved in channel pruning for 2D image-based convolutional networks (CNNs), existing works seldom extend the channel pruning methods to 3D point-based neural networks (PNNs). Directly implementing the 2D CNN channel pruning methods to PNNs undermine the performance of PNNs because of the different representations of 2D images and 3D point clouds as well as the network architecture disparity. In this paper, we proposed CP3, which is a Channel Pruning Plugin for Point-based network. CP3is elaborately designed to leverage the characteristics of point clouds and PNNs in order to enable 2D channel pruning methods for PNNs. Specifically, it presents a coordinate-enhanced channel importance metric to reflect the correlation between dimensional information and individual channel features, and it recycles the discarded points in PNN's sampling process and reconsiders their potentially-exclusive information to enhance the robustness of channel pruning. Experiments on various PNN architectures show that CP3constantly improves state-of-the-art 2D CNN pruning approaches on different point cloud tasks. For instance, our compressed PointNeXt-S on ScanObjectNN achieves an accuracy of 88.52% with a pruning rate of 57.8%, outperforming the baseline pruning methods with an accuracy gain of 1.94%.
Yaomin Huang, Ning Liu 0007, Zhengping Che, Chaomin Shen 0001, Yaxin Peng, Guixu Zhang, Xinmei Liu, Feifei Feng, Jian Tang 0008
CVPR7
2023 Indoor Depth Recovery Based on Deep Unfolding with Non-Local Prior
abstract
In recent years, depth recovery based on deep networks has achieved great success. However, the existing state-of-the-art network designs perform like black boxes in depth recovery tasks, lacking a clear mechanism. Utilizing the property that there is a large amount of non-local common characteristics in depth images, we propose a novel model-guided depth recovery method, namely the DC-NLAR model. A non-local auto-regressive regular term is also embedded into our model to capture more non-local depth information. To fully use the excellent performance of neural networks, we develop a deep image prior to better describe the characteristic of depth images. We also introduce an implicit data consistency term to tackle the degenerate operator with high heterogeneity. We then unfold the proposed model into networks by using the half-quadratic splitting algorithm. This proposed method is experimented on the NYU-Depth V2 and SUN RGB-D datasets, and the experimental results achieve comparable performance to that of deep learning methods.
Yuhui Dai, Junkang Zhang, Faming Fang, Guixu Zhang
ICCV4
2023 Decomposition-Based Variational Network for Multi-Contrast MRI Super-Resolution and Reconstruction
abstract
Multi-contrast MRI super-resolution (SR) and reconstruction methods aim to explore complementary information from the reference image to help the reconstruction of the target image. Existing deep learning-based methods usually manually design fusion rules to aggregate the multi-contrast images, fail to model their correlations accurately and lack certain interpretations. Against these issues, we propose a multi-contrast variational network (MC-VarNet) to explicitly model the relationship of multi-contrast images. Our model is constructed based on an intuitive motivation that multi-contrast images have consistent (edges and structures) and inconsistent (contrast) information. We thus build a model to reconstruct the target image and decompose the reference image as a common component and a unique component. In the feature interaction phase, only the common component is transferred to the target image. We solve the variational model and unfold the iterative solutions into a deep network. Hence, the proposed method combines the good interpretability of model-based methods with the powerful representation ability of deep learning-based methods. Experimental results on the multi-contrast MRI reconstruction and SR demonstrate the effectiveness of the proposed model. Especially, since we explicitly model the multi-contrast images, our model is more robust to the reference images with noises and large inconsistent structures. The code is available at https://github.com/lpcccccv/MC-VarNet.
Pengcheng Lei, Faming Fang, Guixu Zhang, Tieyong Zeng
ICCV3
2023 Swin-ASNet: An Adaptive RGB-selection Network with Swin Transformer for Retinal Vessel Segmentation
abstract
The retinal vasculature reflected by fundus images provides ophthalmologists with information to diagnose eye-related diseases. Therefore, the development of an accurate and automatic retinal vessel segmentation system is critical. Most deep learning-based methods directly use a color image or use a grayscale image simply transformed from the color image as input, and few models focus on the relationship between the three RGB channels. In this paper, we analyze the characteristics of the separated RGB channel images and propose an adaptive RGB-selection network with swin transformer (Swin-ASNet). The input of Swin-ASNet includes the original fundus image and three grayscale images separated from the color image. Our method can select the useful information adaptively through a designed adaptive selection aggregation module. In addition, we adopt the latest swin transformer as the backbone to extract strong features. To better fuse high and low-level features, we design a high-low interaction module, which applies a modified non-local operation under the graph convolution domain. The low-level features are injected into deep semantic information to enhance the detail representation. Experimental results show that our method can achieve state-of-the-art results in three public datasets, comparing with the existing methods.
Qunchao Jin, Hongyu Hou, Guixu Zhang, Haoan Wang, Zhi Li 0080
ICME3
2023 Fine-grained Learning for Visible-Infrared Person Re-identification
abstract
Visible-Infrared Person Re-identification aims to retrieve specific identities from different modalities. In order to relieve the modality discrepancy, previous works mainly concentrate on aligning the distribution of high-level features, while disregarding the exploration of fine-grained information. In this paper, we propose a novel Fine-grained Information Exploration Network (FIENet) to implement discriminative representation, further alleviating the modality discrepancy. Firstly, we propose a Progressive Feature Aggregation Module (PFAM) to progressively aggregate mid-level features, and a Multi-Perception Interaction Module (MPIM) to achieve the interaction with diverse perceptions. Additionally, combined with PFAM and MPIM, more fine-grained information can be extracted, which is beneficial for FIENet to focus on discriminative human parts in both modalities effectively. Secondly, in terms of the feature center, we introduce an Identity-Guided Center Loss (IGCL) to supervise identity representation with intra-identity and inter-identity information. Finally, extensive experiments are conducted to demonstrate that our method achieves state-of-the-art performance.
Mengzan Qi, Sixian Chan 0001, Chen Hang, Guixu Zhang, Zhi Li 0080
ICME4
2023 Deep Unfolding Convolutional Dictionary Model for Multi-Contrast MRI Super-resolution and Reconstruction
abstract
Magnetic resonance imaging (MRI) tasks often involve multiple contrasts. Recently, numerous deep learning-based multi-contrast MRI super-resolution (SR) and reconstruction methods have been proposed to explore the complementary information from the multi-contrast images. However, these methods either construct parameter-sharing networks or manually design fusion rules, failing to accurately model the correlations between multi-contrast images and lacking certain interpretations. In this paper, we propose a multi-contrast convolutional dictionary (MC-CDic) model under the guidance of the optimization algorithm with a well-designed data fidelity term. Specifically, we bulid an observation model for the multi-contrast MR images to explicitly model the multi-contrast images as common features and unique features. In this way, only the useful information in the reference image can be transferred to the target image, while the inconsistent information will be ignored. We employ the proximal gradient algorithm to optimize the model and unroll the iterative steps into a deep CDic model. Especially, the proximal operators are replaced by learnable ResNet. In addition, multi-scale dictionaries are introduced to further improve the model performance. We test our MC-CDic model on multi-contrast MRI SR and reconstruction tasks. Experimental results demonstrate the superior performance of the proposed MC-CDic model against existing SOTA methods. Code is available at https://github.com/lpcccc-cv/MC-CDic.
Pengcheng Lei, Faming Fang, Guixu Zhang, Ming Xu 0010
IJCAI3
2023 Deep Algorithm Unrolling with Registration Embedding for Pansharpening
abstract
Pansharpening aims to sharpen low resolution (LR) multispectral (MS) images with the help of corresponding high resolution (HR) panchromatic (PAN) images to obtain HRMS images. Model-based pansharpening methods manually design objective functions via observation model and hand-crafted priors. However, inevitable performance degradation may occur in the case that the prior is invalid. Although many deep learning based end-to-end pansharpening methods have been proposed recently, they still need to be improved due to the insufficient study on HRMS related domain knowledge. Besides, existing pansharpening methods rarely consider the misalignments between MS and PAN images, leading to poor performance. To tackle these issues, this paper proposes to unrolling the observation model with registration embedding for pansharpening. Inspired by the optical flow estimation, we embed the registration operation into the observation model to reconstruct the pansharpening function with the help of a deep prior of HRMS images, and then unroll the iterative solution into a novel deep convolutional network.. Apart from the single HRMS supervision, we also introduce a consistency loss to supervise the two degradation processes. The use of consistency loss enables the degradation sub-networks to learn more realistic degradation. Experimental results at reduced-resolution and full-resolution are reported to demonstrate the superiority of the proposed method to other state-of-the-art pansharpening methods. In GaoFen-2 dataset evaluation, our method achieves 1.2dB higher PSNR than SOTA techniques.
Tingting Wang 0007, Yongxu Ye, Faming Fang, Guixu Zhang, Ming Xu 0010
ACM Multimedia4
2023 DSAT-Net: Dual Spatial Attention Transformer for Building Extraction From Aerial Images
abstract
Both local and global context dependencies are essential for building extraction from remote sensing (RS) images. Convolutional Neural Network (CNN) can extract local spatial details well but lacks the ability to model long-range dependency. In recent years, Vision Transformer (ViT) have shown great potential in modeling global context dependency. However, it usually brings huge computational cost, and spatial details can not be fully retained in the process of feature extraction. To maximize the advantages of CNNs and ViTs, we propose DSAT-Net, which combine them in one model. In DSAT-Net, we design an efficient Dual Spatial Attention Transformer (DSAFormer) to solve the defects of standard ViT. It has a dual attention structure to complement each other. Specifically, the global attention path (GAP) conducts a large scale down sampling of the feature maps before the global self-attention computing, to reduce the computational cost. The local attention path (LAP) uses efficient stripe convolution to generate local attention, which can alleviate the loss of information caused by down-sampling operation in the GAP and supplement the spatial details. In addition, we design a feature refining module called Channel Mixing Feature Refine Module (CM-FRM) to fuse low-level and high-level features. Our model achieved competitive results on three public building extraction datasets. Code will be available at: https://github.com/stdcoutzrh/BuildingExtraction.
Renhe Zhang, Zhechun Wan, Qian Zhang 0003, Guixu Zhang
IEEE Geosci. Remote. Sens. Lett.4
2023 SDSC-UNet: Dual Skip Connection ViT-Based U-Shaped Model for Building Extraction
abstract
Benefiting from effective global information interaction, vision-transformers (ViTs) have been widely used in the building extraction task. However, buildings in remote sensing (RS) images usually differ greatly in size. Mainstream ViT-based segmentation models for RS images are based on Swin Transformer, which lacks multi-scale information inside the ViT block. In addition, they only connect the output of the entire ViT encoder block to the decoder, which ignore the similarity information of the attention maps inside the ViT encoder block, and are unable to provide better global dependencies for the decoder. To solve above problems, we introduce a novel Shunted Transformer, which enables the model to capture multi-scale information internally while fully establishing global dependencies, to build a pure ViT-based U-shaped model for building extraction. Furthermore, unlike the previous single-skip-connection structure of U-shaped methods, we build a novel dual skip connection structure inside the model. It simultaneously transmits the attention maps inside the ViT encoder block and its entire output to the decoder, thereby fully mining the information of the ViT encoder block and providing better global information guidance for the decoder. Thus, our model is named Shunted Dual Skip Connection UNet (SDSC-UNet). We also design a feature fusion module called Dual Skip Upsample Fusion Module (DSUFM) to aggregate the information. Our model has yields state-of-the-art (SOTA) performance (83.02%IoU) on the Inria Aerial Image Labeling Dataset. Code will be available.
Renhe Zhang, Qian Zhang 0003, Guixu Zhang
IEEE Geosci. Remote. Sens. Lett.3
2023 FrMLNet: Framelet-Based Multilevel Network for Pansharpening
abstract
Most modern satellites can provide two types of images: 1) panchromatic (PAN) image and 2) multispectral (MS) image. The former has high spatial resolution and low spectral resolution, while the latter has high spectral resolution and low spatial resolution. To obtain images with both high spectral and spatial resolution, pansharpening has emerged to fuse the spatial information of the PAN image and the spectral information of the MS image. However, most pansharpening methods fail to preserve spatial and spectral information simultaneously. In this article, we propose a framelet-based convolutional neural network (CNN) for pansharpening which makes it possible to pursue both high spectral and high spatial resolution. Our network consists of three subnetworks: 1) feature embedding net; 2) feature fusion net; and 3) framelet prediction net. Different from conventional CNN methods directly inferring high-resolution MS images, our approach learns to predict their framelet coefficients from available PAN and MS images. The introduction of multilevel feature aggregation and hybrid residual connection makes full use of spatial information of PAN image and spectral information of MS image. Quantitative and qualitative experiments at reduced- and full-resolution demonstrate that the proposed method achieves more appealing results than other state-of-the-art pansharpening methods. The source code and trained models are available at https://github.com/TingMAC/FrMLNet.
Tingting Wang 0007, Faming Fang, Guixu Zhang
IEEE Trans. Cybern.4
2023 DMCSC: Deep Multisource Convolutional Sparse Coding Model for Pansharpening
abstract
Pansharpening aims to produce a high-resolution multispectral (HRMS) image by combining a low-resolution multispectral (LRMS) image with a high-resolution panchromatic (PAN) image through a fusion process. Deep learning (DL)-based pansharpening methods have demonstrated impressive results in generating high-quality HRMS images. However, they suffer from a lack of interpretability due to their black-box network architectures. Recently, model-based deep unrolling networks have been proposed to improve the interpretability of networks. Among these approaches, the multi-source convolutional sparse coding (MCSC)-based models stand out by effectively learning common and unique features from both LRMS and PAN images, showing promising results. As the LRMS image provides limited information in MCSC-based models, it can result in weak feature response and even lead to incorrect fusion outcomes. To address this issue, we propose a novel deep MCSC-based method that enhances the robustness and performance. Specifically, we build an optimization model that integrates MCSC with a degradation model and a deep prior, which can sufficiently capture the common information shared by the latent HRMS images and PAN images, thereby enabling the recovery of more accurate spectral information. To optimize the proposed model, we adopt an iterative optimization strategy that unfolds the iterative solution into networks. Moreover, we propose an enhanced version of our method that utilizes multi-scale dictionaries to capture common and unique features at different scales, thereby facilitating the extraction of more abundant spectral and spatial details. We evaluate the effectiveness of our proposed method on multiple benchmark datasets. Experiment results demonstrate its effectiveness in improving the robustness and performance of MCSC-based models.
Junkang Zhang, Yongxu Ye, Faming Fang, Tingting Wang 0007, Guixu Zhang
IEEE Trans. Geosci. Remote. Sens.5
2023 FEGNet: A Feedback Enhancement Gate Network for Automatic Polyp Segmentation
abstract
Regular colonoscopy is an effective way to prevent colorectal cancer by detecting colorectal polyps. Automatic polyp segmentation significantly aids clinicians in precisely locating polyp areas for further diagnosis. However, polyp segmentation is a challenge problem, since polyps appear in a variety of shapes, sizes and textures, and they tend to have ambiguous boundaries. In this paper, we propose a U-shaped model named Feedback Enhancement Gate Network (FEGNet) for accurate polyp segmentation to overcome these difficulties. Specifically, for the high-level features, we design a novel Recurrent Gate Module (RGM) based on the feedback mechanism, which can refine attention maps without any additional parameters. RGM consists of Feature Aggregation Attention Gate (FAAG) and Multi-Scale Module (MSM). FAAG can aggregate context and feedback information, and MSM is applied for capturing multi-scale information, which is critical for the segmentation task. In addition, we propose a straightforward but effective edge extraction module to detect boundaries of polyps for low-level features, which is used to guide the training of early features. In our experiments, quantitative and qualitative evaluations show that the proposed FEGNet has achieved the best results in polyp segmentation compared to other state-of-the-art models on five colonoscopy datasets.
Qunchao Jin, Hongyu Hou, Guixu Zhang, Zhi Li 0080
IEEE J. Biomed. Health Informatics3
2023 Frequency Learning via Multi-Scale Fourier Transformer for MRI Reconstruction
abstract
Since Magnetic Resonance Imaging (MRI) requires a long acquisition time, various methods were proposed to reduce the time, but they ignored the frequency information and non-local similarity, so that they failed to reconstruct images with a clear structure. In this article, we propose Frequency Learning via Multi-scale Fourier Transformer for MRI Reconstruction (FMTNet), which focuses on repairing the low-frequency and high-frequency information. Specifically, FMTNet is composed of a high-frequency learning branch (HFLB) and a low-frequency learning branch (LFLB). Meanwhile, we propose a Multi-scale Fourier Transformer (MFT) as the basic module to learn the non-local information. Unlike normal Transformers, MFT adopts Fourier convolution to replace self-attention to efficiently learn global information. Moreover, we further introduce a multi-scale learning and cross-scale linear fusion strategy in MFT to interact information between features of different scales and strengthen the representation of features. Compared with normal Transformers, the proposed MFT occupies fewer computing resources. Based on MFT, we design a Residual Multi-scale Fourier Transformer module as the main component of HFLB and LFLB. We conduct several experiments under different acceleration rates and different sampling patterns on different datasets, and the experiment results show that our method is superior to the previous state-of-the-art method.
Qiaosi Yi, Faming Fang, Guixu Zhang, Tieyong Zeng
IEEE J. Biomed. Health Informatics3
2022 Label-Guided Auxiliary Training Improves 3D Object Detector
Yaomin Huang, Xinmei Liu, Yichen Zhu 0001, Chaomin Shen 0001, Zhengping Che, Guixu Zhang, Yaxin Peng, Feifei Feng, Jian Tang 0008
ECCV (9)7
2022 CT image quality enhancement via a dual-channel neural network with jointing denoising and super-resolution
Hongyu Hou, Qunchao Jin, Guixu Zhang, Zhi Li 0080
Neurocomputing3
2022 Adjustable super-resolution network via deep supervised learning and progressive self-distillation
Juncheng Li 0003, Faming Fang, Tieyong Zeng, Guixu Zhang, Xizhao Wang
Neurocomputing4
2022 Patch-based weighted SCAD prior for compressive sensing
Yamin Ru, Fang Li 0004, Faming Fang, Guixu Zhang
Inf. Sci.4
2022 Semantic Segmentation Network Using Local Relationship Upsampling for Remote Sensing Images
abstract
Semantic segmentation is a fundamental task in remote sensing image processing. It provides pixel-level classification, which is important for many applications, such as building extraction and land use mapping. The development of convolutional neural network has considerably improved the performance of semantic segmentation. Most semantic segmentation networks are the encoder–decoder structure. Bilinear interpolation is an ordinary upsampling method in the decoder, but bilinear interpolation only considers its own features and inserts three times its own features. This over-simple and data-independent bilinear upsampling may lead to suboptimal results. In this work, we propose an upsampling method based on local relations to replace bilinear interpolation. Upsampling is performed by correlating the local relationship of feature maps of adjacent stages, which can better integrate local and global information. We also design a fusion module based on local similarity. Our proposed method with ResNet101 as the backbone of the segmentation network can improve the average$F_{1}$score and overall accuracy of the Vaihingen data set by 2.69% and 1.31%, respectively. Our proposed method also has fewer parameters and less inference time.
Baokai Lin, Guang Yang 0068, Qian Zhang 0003, Guixu Zhang
IEEE Geosci. Remote. Sens. Lett.4
2022 Low-Level Feature Enhancement Network for Semantic Segmentation of Buildings
abstract
In recent years, convolutional neural networks (CNNs) have been widely used in extracting buildings from remote sensing images. Both semantic representation and spatial location details are crucial for this task. We propose methods to enhance the performance of semantic segmentation by using these low-level features considering that man-made buildings in aerial images have strong textures and edges. Texture Enhancement Attention Module (TEAM) is proposed to strengthen feature in the position with rich texture and improve the semantic representation. Edge Extraction Module (EEM) is applied for directly guiding spatial details learning, which starts with super-resolution maps created by Super-Resolution Module (SRM). Detail Supplement Module (DSM) is designed to further provide details for decoder. On this basis, we propose a low-level feature enhancement network (LFENet) for semantic segmentation of buildings. Experiment results on two aerial datasets show that our works greatly improve the accuracy over the baseline and other models.
Zhechun Wan, Qian Zhang 0003, Guixu Zhang
IEEE Geosci. Remote. Sens. Lett.3
2022 Homogeneous Aggregation Convolution for Building Extraction From Remote Sensing Images
abstract
An increasing number of convolutional neural networks are being applied to various fields, and they have achieved excellent performance. However, standard convolution cannot recognize the connection between surrounding features and cannot characterize them efficiently. For example, buildings in urban remote sensing images exhibit geometric changes, such as rotation, scaling, and local changes. The concept of dynamic convolution is proposed to solve the aforementioned problem. Existing dynamic convolution methods enhance an expression by dynamically changing the sampling points or weights of convolution. However, these end-to-end training methods do not consider which sampling points are important. In this letter, we propose a homogeneous aggregation convolution (HAC) that gives more attention to the sampling points that belong to the same class as the target point. A generate probability map module is designed to generate a probability map between target and sampling points and share this probability map across convolution layers to save computational cost. Experimental results demonstrate that the proposed HAC outperforms standard convolution, and the intersection over union and F1 score are higher than the standard convolution by 2.23% and 1.28%, respectively, on the WHU and Austin datasets. Compared with other convolutions, the proposed HAC convolution is the most efficient in building extraction.
Rouyu Zhang, Baokai Lin, Qian Zhang 0003, Guixu Zhang
IEEE Geosci. Remote. Sens. Lett.4
2022 Phase retrieval from incomplete data via weighted nuclear norm minimization
Zhi Li 0080, Ming Yan 0006, Tieyong Zeng, Guixu Zhang
Pattern Recognit.4
2022 GJTD-LR: A Trainable Grouped Joint Tensor Dictionary With Low-Rank Prior for Single Hyperspectral Image Super-Resolution
abstract
Reconstructing a high-resolution hyperspectral image (HR-HSI) by using a single low-resolution hyperspectral image (LR-HSI) is a significant technique for increasing the spatial resolution of HSIs and overcoming the physical limitation of the HSI sensor. Most single HSI super-resolution methods have achieved great success recently. However, owning to the difficulty of acquiring an HSI, the available training samples are relatively few, which will inevitably lead to relatively low performance. To address this issue, in the paper, we propose a novel single HSI super-resolution method by combining a trainable grouped joint tensor dictionary and a low-rank prior (GJTD-LR). First, we design a trainable grouped joint tensor dictionary, which can build an accurate mapping relationship between training HR-HSIs and their corresponding LR-HSIs with relatively few training samples. To be specific, the training HR-HSI and LR-HSI pairs are decomposed into a joint tensor dictionary and a set of sparse coefficients by using tensor-tensor product to fully preserve the spectral correlation. In addition, we apply a grouped strategy to divide the training images into several groups and learn a compact joint dictionary for each group. Second, a tensor low-rank model is forced into the reconstruction model to further capture the spatial correlation. At last, GJTD-LR is optimized by employing alternating direction method of multipliers (ADMM), soft threshold algorithm, singular value decomposition and fourier domain transform. The experimental results on both remote sensed HSIs and indoor HSIs show the superiority of GJTD-LR to some other traditional and advanced single HSI super-resolution methods.
Cong Liu 0011, Zhihao Fan, Guixu Zhang
IEEE Trans. Geosci. Remote. Sens.3
2022 Collaborative Network for Super-Resolution and Semantic Segmentation of Remote Sensing Images
abstract
In the past few years, multitask learning (MTL) has been widely used in a single model to solve the problems of multiple businesses. MTL enables each task to achieve high performance and greatly reduces computational resource overhead. In this work, we designed a collaborative network that simultaneously solves the super-resolution semantic segmentation and super-resolution image reconstruction. This algorithm can obtain high-resolution semantic segmentation and super-resolution reconstruction results by taking relatively low-resolution images as input when high-resolution data are inconvenient or computing resources are limited. The framework consists of three parts: the semantic segmentation branch (SSB), the super-resolution branch (SRB), and the structural affinity block (SAB). Specifically, the SSB, SRB, and SAB are responsible for completing super-resolution semantic segmentation, image super-resolution reconstruction, and associated features, respectively. Our proposed method is simple and efficient, and it can replace the different branches with most of the state-of-the-art models. The International Society for Photogrammetry and Remote Sensing (ISPRS) segmentation benchmarks were used to evaluate our models. In particular, super-resolution semantic segmentation on the Potsdam dataset reduced Intersection over Union (IoU) by only 1.8% when the resolution of the input image was reduced by a factor of two. The experimental results showed that our framework can obtain more accurate semantic segmentation and super-resolution reconstruction results than the single model.
Qian Zhang 0003, Guang Yang 0068, Guixu Zhang
IEEE Trans. Geosci. Remote. Sens.3
2022 Multi-Scale Grid Network for Image Deblurring With High-Frequency Guidance
abstract
It has been demonstrated that the blurring process reduces the high-frequency information of the original sharp image, so the main challenge for image deblurring is to reconstruct high-frequency information from the blurry image. In this paper, we propose a novel image deblurring framework to focus on the reconstruction of high-frequency information, which consists of two main subnetworks: a high-frequency reconstruction subnetwork (HFRSN) and a multi-scale grid subnetwork (MSGSN). The HFRSN is built to reconstruct latent high-frequency information from multiple scale blurry images. The MSGSN performs deblurring processes with high-frequency guidance at different scales simultaneously. Besides, in order to better use high-frequency information to restore sharpening images, we designed a high-frequency information aggregation (HFAG) module and a high-frequency information attention (HFAT) module in MSGSN. The HFAG module is designed to fuse high-frequency features and image features at the feature extraction stage, and the HFAT module is built to enhance the feature reconstruction stage. Extensive experiments on different datasets show the effectiveness and efficiency of our method.
Yang Liu 0289, Faming Fang, Tingting Wang 0007, Juncheng Li 0003, Yun Sheng, Guixu Zhang
IEEE Trans. Multim.6
2022 Efficient and Accurate Multi-Scale Topological Network for Single Image Dehazing
abstract
Single image dehazing is a challenging ill-posed problem that has drawn significant attention in the last few years. Recently, convolutional neural networks have achieved great success in image dehazing. However, it is still difficult for these increasingly complex models to recover accurate details from the hazy image. In this paper, we pay attention to the feature extraction and utilization of the input image itself. To achieve this, we propose a Multi-scale Topological Network (MSTN) to fully explore the features at different scales. Meanwhile, we design a Multi-scale Feature Fusion Module (MFFM) and an Adaptive Feature Selection Module (AFSM) to achieve the selection and fusion of features at different scales, so as to achieve progressive image dehazing. This topological network provides a large number of search paths that enable the network to extract abundant image features as well as strong fault tolerance and robustness. In addition, ASFM and MFFM can adaptively select important features and ignore interference information when fusing different scale representations. Extensive experiments are conducted to demonstrate the superiority of our method compared with state-of-the-art methods.
Qiaosi Yi, Juncheng Li 0003, Faming Fang, Aiwen Jiang, Guixu Zhang
IEEE Trans. Multim.5
2021 Structure-Preserving Deraining with Residue Channel Prior Guidance
abstract
Single image deraining is important for many high-level computer vision tasks since the rain streaks can severely degrade the visibility of images, thereby affecting the recognition and analysis of the image. Recently, many CNN-based methods have been proposed for rain removal. Although these methods can remove part of the rain streaks, it is difficult for them to adapt to real-world scenarios and restore high-quality rain-free images with clear and accurate structures. To solve this problem, we propose a Structure-Preserving Deraining Network (SPDNet) with RCP guidance. SPDNet directly generates high-quality rain-free images with clear and accurate structures under the guidance of RCP but does not rely on any rain-generating assumptions. Specifically, we found that the RCP of images contains more accurate structural information than rainy images. Therefore, we introduced it to our deraining network to protect structure information of the rain-free image. Meanwhile, a Wavelet-based Multi-Level Module (WMLM) is proposed as the backbone for learning the background information of rainy images and an Interactive Fusion Module (IFM) is designed to make full use of RCP information. In addition, an iterative guidance strategy is proposed to gradually improve the accuracy of RCP, refining the result in a progressive path. Extensive experimental results on both synthetic and real-world datasets demonstrate that the proposed model achieves new state-of-the-art results. Code: https://github.com/Joyies/SPDNet
Qiaosi Yi, Juncheng Li 0003, Qinyan Dai, Faming Fang, Guixu Zhang, Tieyong Zeng
ICCV5
2021 Feedback Network for Mutually Boosted Stereo Image Super-Resolution and Disparity Estimation
abstract
Under stereo settings, the problem of image super-resolution (SR) and disparity estimation are interrelated that the result of each problem could help to solve the other. The effective exploitation of correspondence between different views facilitates the SR performance, while the high-resolution (HR) features with richer details benefit the correspondence estimation. According to this motivation, we propose a Stereo Super-Resolution and Disparity Estimation Feedback Network (SSRDE-FNet), which simultaneously handles the stereo image super-resolution and disparity estimation in a unified framework and interact them with each other to further improve their performance. Specifically, the SSRDE-FNet is composed of two dual recursive sub-networks for left and right views. Besides the cross-view information exploitation in the low-resolution (LR) space, HR representations produced by the SR process are utilized to perform HR disparity estimation with higher accuracy, through which the HR features can be aggregated to generate a finer SR result. Afterward, the proposed HR Disparity Information Feedback (HRDIF) mechanism delivers information carried by HR disparity back to previous layers to further refine the SR image reconstruction. Extensive experiments demonstrate the effectiveness and advancement of SSRDE-FNet.
Qinyan Dai, Juncheng Li 0003, Qiaosi Yi, Faming Fang, Guixu Zhang
ACM Multimedia5
2021 VAN: Voting and Attention Based Network for Unsupervised Medical Image Registration
Zhiang Zu, Guixu Zhang, Yaxin Peng, Chaomin Shen 0001
PRICAI (1)2
2021 Edge-guided Composition Network for Image Stitching
Qinyan Dai, Faming Fang, Juncheng Li 0003, Guixu Zhang, Aimin Zhou
Pattern Recognit.4
2021 MDCN: Multi-Scale Dense Cross Network for Image Super-Resolution
abstract
Convolutional neural networks have been proven to be of great benefit for single-image super-resolution (SISR). However, previous works do not make full use of multi-scale features and ignore the inter-scale correlation between different upsampling factors, resulting in sub-optimal performance. Instead of blindly increasing the depth of the network, we are committed to mining image features and learning the inter-scale correlation between different upsampling factors. To achieve this, we propose a Multi-scale Dense Cross Network (MDCN), which achieves great performance with fewer parameters and less execution time. MDCN consists of multi-scale dense cross blocks (MDCBs), hierarchical feature distillation block (HFDB), and dynamic reconstruction block (DRB). Among them, MDCB aims to detect multi-scale features and maximize the use of image features flow at different scales, HFDB focuses on adaptively recalibrate channel-wise feature responses to achieve feature distillation, and DRB attempts to reconstruct SR images with different upsampling factors in a single model. It is worth noting that all these modules can run independently. It means that these modules can be selectively plugged into any CNN model to improve model performance. Extensive experiments show that MDCN achieves competitive results in SISR, especially in the reconstruction task with multiple upsampling factors. The code is provided athttps://github.com/MIVRC/MDCN-PyTorch.
Juncheng Li 0003, Faming Fang, Kangfu Mei, Guixu Zhang
IEEE Trans. Circuits Syst. Video Technol.5
2021 Superpixel-Based Seamless Image Stitching for UAV Images
abstract
Image stitching aims to generate a natural seamless high-resolution panoramic image free of distortions or artifacts as fast as possible. In this article, we propose a new seam cutting strategy based on superpixels for unmanned aerial vehicle (UAV) image stitching. Explicitly, we decompose the issue into three steps: image registration, seam cutting, and image blending. First, we employ adaptive as-natural-as-possible (AANAP) warps for registration, obtaining two aligned images in the same coordinate system. Then, we propose a novel superpixel-based energy function that integrates color difference, gradient difference, and texture complexity information to search a perceptually optimal seam located in continuous areas with high similarity. We apply the graph cut algorithm to solve the problem and thereby conceal artifacts in the overlapping area. Finally, we utilize a superpixel-based color blending approach to eliminate visible seams and achieve natural color transitions. Experimental results demonstrate that our method can effectively and efficiently realize seamless stitching, and is superior to several state-of-the-art methods in UAV image stitching.
Yiting Yuan, Faming Fang, Guixu Zhang
IEEE Trans. Geosci. Remote. Sens.3
2021 Luminance-Aware Pyramid Network for Low-Light Image Enhancement
abstract
Low-light image enhancement based on deep convolutional neural networks (CNNs) has revealed prominent performance in recent years. However, it is still a challenging task since the underexposed regions and details are always imperceptible. Moreover, deep learning models are always accompanied by complex structures and enormous computational burden, which hinders their deployment on mobile devices. To remedy these issues, in this paper, we present a lightweight and efficient Luminance-aware Pyramid Network (LPNet) to reconstruct normal-light images in a coarse-to-fine strategy. The architecture is comprised of two coarse feature extraction branches and a luminance-aware refinement branch with an auxiliary subnet learning the luminance map of the input and target images. Besides, we propose a multi-scale contrast feature block (MSCFB) that involves channel split, channel shuffle strategies, and contrast attention mechanism. MSCFB is the essential component of our network, which achieves an excellent balance between image quality and model size. In this way, our method can not only brighten up low-light images with rich details and high contrast but also significantly ameliorate the execution speed. Extensive experiments demonstrate that our LPNet outperforms state-of-the-art methods both qualitatively and quantitatively.
Juncheng Li 0003, Faming Fang, Fang Li 0004, Guixu Zhang
IEEE Trans. Multim.5
2021 Multilevel Edge Features Guided Network for Image Denoising
abstract
Image denoising is a challenging inverse problem due to complex scenes and information loss. Recently, various methods have been considered to solve this problem by building a well-designed convolutional neural network (CNN) or introducing some hand-designed image priors. Different from previous works, we investigate a new framework for image denoising, which integrates edge detection, edge guidance, and image denoising into an end-to-end CNN model. To achieve this goal, we propose a multilevel edge features guided network (MLEFGN). First, we build an edge reconstruction network (Edge-Net) to directly predict clear edges from the noisy image. Then, the Edge-Net is embedded as part of the model to provide edge priors, and a dual-path network is applied to extract the image and edge features, respectively. Finally, we introduce a multilevel edge features guidance mechanism for image denoising. To the best of our knowledge, the Edge-Net is the first CNN model specially designed to reconstruct image edges from the noisy image, which shows good accuracy and robustness on natural images. Extensive experiments clearly illustrate that our MLEFGN achieves favorable performance against other methods and plenty of ablation studies demonstrate the effectiveness of our proposed Edge-Net and MLEFGN. The code is available at https://github.com/MIVRC/MLEFGN-PyTorch.
Faming Fang, Juncheng Li 0003, Yiting Yuan, Tieyong Zeng, Guixu Zhang
IEEE Trans. Neural Networks Learn. Syst.5
2020 Stylization-Based Architecture for Fast Deep Exemplar Colorization
abstract
Exemplar-based colorization aims to add colors to a grayscale image guided by a content related reference image. Existing methods are either sensitive to the selection of reference images (content, position) or extremely time and resource consuming, which limits their practical application. To tackle these problems, we propose a deep exemplar colorization architecture inspired by the characteristics of stylization in feature extracting and blending. Our coarse- to-fine architecture consists of two parts: a fast transfer sub-net and a robust colorization sub-net. The transfer sub- net obtains a coarse chrominance map via matching basic feature statistics of the input pairs in a progressive way. The colorization sub-net refines the map to generate the final results. The proposed end-to-end network can jointly learn faithful colorization with a related reference and plausible color prediction with unrelated reference. Extensive experimental validation demonstrates that our approach outperforms the state-of-the-art methods in less time whether in exemplar-based colorization or image stylization tasks.
Zhongyou Xu, Tingting Wang 0007, Faming Fang, Yun Sheng, Guixu Zhang
CVPR5
2020 Enhanced Sparse Model for Blind Deblurring
Faming Fang, Fang Li 0004, Guixu Zhang
ECCV (25)5
2020 OID: Outlier Identifying and Discarding in Blind Image Deblurring
Faming Fang, Guixu Zhang
ECCV (25)5
2020 Label Smoothing Technique for Ordinal Classification in Cloud Assessment
abstract
Satellite image classification is a challenging task if the input labels are not sufficiently accurate. The automatic cloud cover assessment (ACCA), for example, aims to classify the cloud covers of satellite images as alphabetical categories from A to E showing the escalating levels of clouds; however, those labels for training are often obtained by a subjective qualitative assessment, i.e., they may be not accurate. Therefore, this paper studies how to conduct ACCA under this circumstance. We propose a label smoothing approach and improve the accuracy around 3 percentage points (e.g., from 75.9% to 78.4% for ResNet network) without changing other network structures and parameters.
Yuxuan Wei, Qixuan Liu, Guixu Zhang, Yaxin Peng, Chaomin Shen 0001
IGARSS3
2020 A Classification Surrogate Model based Evolutionary Algorithm for Neural Network Structure Learning
abstract
Designing neural networks often requires a large number of artificial intelligence experts. However, such manual processes are time-consuming and require numerous resources. In this paper, we try to search neural network structures automatically for the image classification task. Moreover, considering the huge computational cost of neural architecture search (NAS), we attempt to apply a classification surrogate model based multi-objective evolutionary algorithm to search neural network architectures (CSMEA-Net). The algorithm combines two objectives, i.e., minimizing the validation error and the computational complexity measured by the number of floating-point operations (FLOPs) to achieve Pareto Optimality. In addition, we improve the components of the cell-based search space. The performance of network architectures discovered by our method is evaluated on CIFAR-10 and CIFAR-100 datasets. The experimental results show that the proposed approach can find a higher-performance neural network architecture compared with both hand-crafted as well as automatically-designed networks.
Wenyue Hu, Aimin Zhou, Guixu Zhang
IJCNN3
2020 Removing moiré patterns from single images
Faming Fang, Tingting Wang 0007, Shuyan Wu, Guixu Zhang
Inf. Sci.4
2020 A Novel Retinex-Based Fractional-Order Variational Model for Images With Severely Low Light
abstract
In this paper, we propose a novel Retinex-based fractional-order variational model for severely low-light images. The proposed method is more flexible in controlling the regularization extent than the existing integer-order regularization methods. Specifically, we decompose directly in the image domain and perform the fractional-order gradient total variation regularization on both the reflectance component and the illumination component to get more appropriate estimated results. The merits of the proposed method are as follows: 1) small-magnitude details are maintained in the estimated reflectance. 2) illumination components are effectively removed from the estimated reflectance. 3) the estimated illumination is more likely piecewise smooth. We compare the proposed method with other closely related Retinex-based methods. Experimental results demonstrate the effectiveness of the proposed method.
Fang Li 0004, Faming Fang, Guixu Zhang
IEEE Trans. Image Process.4
2020 Variational Single Image Dehazing for Enhanced Visualization
abstract
In this paper, we investigate the challenging task of removing haze from a single natural image. The analysis on the haze formation model shows that the atmospheric veil has much less relevance to chrominance than luminance, which motivates us to neglect the haze in the chrominance channel and concentrate on the luminance channel in the dehazing process. Besides, the experimental study illustrates that the YUV color space is most suitable for image dehazing. Accordingly, a variational model is proposed in the Y channel of the YUV color space by combining the reformulation of the haze model and the two effective priors. As we mainly focus on the Y channel, most of the chrominance information of the image is preserved after dehazing. The numerical procedure based on the alternating direction method of multipliers (ADMM) scheme is presented to obtain the optimal solution. Extensive experimental results on real-world hazy images and synthetic dataset demonstrate clearly that our method can unveil the details and recover vivid color information, which is competitive among many existing dehazing algorithms. Further experiments show that our model also can be applied for image enhancement.
Faming Fang, Tingting Wang 0007, Yang Wang 0020, Tieyong Zeng, Guixu Zhang
IEEE Trans. Multim.5
2020 A Superpixel-Based Variational Model for Image Colorization
abstract
Image colorization refers to a computer-assisted process that adds colors to grayscale images. It is a challenging task since there is usually no one-to-one correspondence between color and local texture. In this paper, we tackle this issue by exploiting weighted nonlocal self-similarity and local consistency constraints at the resolution of superpixels. Given a grayscale target image, we first select a color source image containing similar segments to target image and extract multi-level features of each superpixel in both images after superpixel segmentation. Then a set of color candidates for each target superpixel is selected by adopting a top-down feature matching scheme with confidence assignment. Finally, we propose a variational approach to determine the most appropriate color for each target superpixel from color candidates. Experiments demonstrate the effectiveness of the proposed method and show its superiority to other state-of-the-art methods. Furthermore, our method can be easily extended to color transfer between two color images.
Faming Fang, Tingting Wang 0007, Tieyong Zeng, Guixu Zhang
IEEE Trans. Vis. Comput. Graph.4
2019 The Adversarial Attack and Detection under the Fisher Information Metric
abstract
Many deep learning models are vulnerable to the adversarial attack, i.e., imperceptible but intentionally-designed perturbations to the input can cause incorrect output of the networks. In this paper, using information geometry, we provide a reasonable explanation for the vulnerability of deep learning models. By considering the data space as a non-linear space with the Fisher information metric induced from a neural network, we first propose an adversarial attack algorithm termed one-step spectral attack (OSSA). The method is described by a constrained quadratic form of the Fisher information matrix, where the optimal adversarial perturbation is given by the first eigenvector, and the vulnerability is reflected by the eigenvalues. The larger an eigenvalue is, the more vulnerable the model is to be attacked by the corresponding eigenvector. Taking advantage of the property, we also propose an adversarial detection method with the eigenvalues serving as characteristics. Both our attack and detection algorithms are numerically optimized to work efficiently on large datasets. Our evaluations show superior performance compared with other methods, implying that the Fisher information is a promising approach to investigate the adversarial attacks and defenses.
Chenxiao Zhao, P. Thomas Fletcher, Mixue Yu, Yaxin Peng, Guixu Zhang, Chaomin Shen 0001
AAAI5
2019 Fuzzy-Classification Assisted Solution Preselection in Evolutionary Optimization
abstract
In evolutionary optimization, the preselection is an efficient operator to improve the search efficiency, which aims to filter unpromising candidate solutions before fitness evaluation. Most existing preselection operators rely on fitness values, surrogate models, or classification models. Basically, the classification based preselection regards the preselection as a classification procedure, i.e., differentiating promising and unpromising candidate solutions. However, the difference between promising and unpromising classes becomes fuzzy as the running process goes on, as all the left solutions are likely to be promising ones. Facing this challenge, this paper proposes a fuzzy classification based preselection (FCPS) scheme, which utilizes the membership function to measure the quality of candidate solutions. The proposed FCPS scheme is applied to two state-of-the-art evolutionary algorithms on a test suite. The experimental results show the potential of FCPS on improving algorithm performance.
Aimin Zhou, Jianyong Sun, Guixu Zhang
AAAI4
2019 MIHS: A Multiobjective Pan-sharpening Method for Remote Sensing Images
abstract
Pan-sharpening aims to integrate the spatial details of high resolution panchromatic image (PAN) with the spectral information of the corresponding low resolution multispectral image (MS) to produce high resolution multispectral image, which is a challenge real-world task. In this paper, we propose a novel pan-sharpening method, called multiobjective Intensity-Hue-Saturation (MIHS) transformation, which combines adaptive Intensity-Hue-Saturation transformation and evolutionary multi-objective optimization techniques for remote sensing images. The basic idea is to convert the pan-sharpening problem into a mutiobjective optimization problem and deal with it by an multiobjective evolutionary algorithm. The proposed method is applied to two Quick-bird images. Experimental results demonstrate that the proposed method does markedly improve the pan-sharpening performance compared to the state-of-the-art methods and an evolutionary method with single objective in terms of both subjective visual effects and objective quality metrics.
Yingxia Chen, Cong Liu 0011, Aimin Zhou, Guixu Zhang
CEC4
2019 Blind Image Deblurring With Local Maximum Gradient Prior
abstract
Blind image deblurring aims to recover sharp image from a blurred one while the blur kernel is unknown. To solve this ill-posed problem, a great amount of image priors have been explored and employed in this area. In this paper, we present a blind deblurring method based on Local Maximum Gradient (LMG) prior. Our work is inspired by the simple and intuitive observation that the maximum value of a local patch gradient will diminish after the blur process, which is proved to be true both mathematically and empirically. This inherent property of blur process helps us to establish a new energy function. By introducing an liner operator to compute the Local Maximum Gradient, together with an effective optimization scheme, our method can handle various specific scenarios. Extensive experimental results illustrate that our method is able to achieve favorable performance against state-of-the-art algorithms on both synthetic and real-world images.
Faming Fang, Tingting Wang 0007, Guixu Zhang
CVPR4
2019 Pareto Optimal Set Approximation by Models: A Linear Case
Aimin Zhou, Haoying Zhao, Hu Zhang 0002, Guixu Zhang
EMO4
2019 A New Female Body Segmentation and Feature Localisation Method for Image-Based Anthropometry
Yun Sheng, Guixu Zhang
MMM (1)3
2019 Cascaded Dilated Dense Network with Two-step Data Consistency for MRI Reconstruction
abstract
Compressed Sensing MRI (CS-MRI) aims at reconstrcuting de-aliased images from sub-Nyquist sampling k-space data to accelerate MR Imaging. Inspired by recent deep learning methods, we propose a Cascaded Dilated Dense Network (CDDN) for MRI reconstruction. Dense blocks with residual connection are used to restore clear images step by step and dilated convolution is introduced for expanding receptive field without taking more network parameters. After each sub-network, we use a novel two-step Data Consistency (DC) operation in k-space. We convert the complex result from first DC operation to real-valued images and applied another sampled \emph{k}-space data replacement. Extensive experiments demonstrate that the proposed CDDN with two-step DC achieves state-of-art result.
Faming Fang, Guixu Zhang
NeurIPS3
2019 A new area-based convexity measure with distance weighted area integration for planar shapes
Xiayan Shi, Yun Sheng, Guixu Zhang
Comput. Aided Geom. Des.4
2019 Fast Color Blending for Seamless Image Stitching
abstract
In this letter, we propose a fast and robust method for stitching overlapped images captured by the unmanned aerial vehicle. First, we apply the shape-preserving half-projective method to precisely and stably align a pair of partially overlapped input images. Then, an optimal stitching line is searched to remove ghosts caused by the moving objects in the overlapped area. We subsequently propose a color blending method to eliminate all the color inconsistencies in the prealigned image. In accordance with the color differences of the pixels on the optimal stitching seam, we utilize weighted value coordinate interpolation algorithms to compute accurate color changes for all the pixels in the target image. The calculated color changes are then added to the target image to remove the color inconsistency. Furthermore, we introduce the superpixel segmentation to divide the target image into a reduced number of superpixels, and we assign each superpixel the same color change value. Such a superpixel level operation can greatly reduce the computational complexity. Experiments show that our method is promising to achieve effective and efficient stitching results.
Faming Fang, Tingting Wang 0007, Yingying Fang, Guixu Zhang
IEEE Geosci. Remote. Sens. Lett.4
2019 High-Quality Bayesian Pansharpening
abstract
Pansharpening is a process of acquiring a multi-spectral image with high spatial resolution by fusing a low resolution multi-spectral image with a corresponding high resolution panchromatic image. In this paper, a new pansharpening method based on the Bayesian theory is proposed. The algorithm is mainly based on three assumptions: 1) the geometric information contained in the pan-sharpened image is coincident with that contained in the panchromatic image; 2) the pan-sharpened image and the original multi-spectral image should share the same spectral information; and 3) in each pan-sharpened image channel, the neighboring pixels not around the edges are similar. We build our posterior probability model according to above-mentioned assumptions and solve it by the alternating direction method of multipliers. The experiments at reduced and full resolution show that the proposed method outperforms the other state-of-the-art pansharpening methods. Besides, we verify that the new algorithm is effective in preserving spectral and spatial information with high reliability. Further experiments also show that the proposed method can be successfully extended to hyper-spectral image fusion.
Tingting Wang 0007, Faming Fang, Fang Li 0004, Guixu Zhang
IEEE Trans. Image Process.4
2018 Parallel Hashing Using Representative Points in Hyperoctants
abstract
The goal of hashing is to learn a low-dimensional binary representation of high-dimensional information, leading to a tremendous reduction of computational cost. Previous studies usually achieved this goal by applying projection or quantization methods. However, the projection method fails to capture the intrinsic data structures, and the quantization method cannot make full use of complete information by its strategy of partitioning original space. To combine their advantages and avoid their drawbacks, we propose a novel algorithm, termed as representative points quantization (RPQ), by using the representative points defined as the barycenters of points in the hyperoctants. To settle the problem of exponential time complexity with the growth of the coding length, for long hashing codes, we further propose a parallel RPQ (PRPQ) algorithm, by separating a long code into several short codes, re-coding the short codes in different low dimensional subspaces, and then concatenating them to a long code. Experiments on image retrieval tasks demonstrate that RPQ and PRPQ can well capture the main topology structure of data, showing that our algorithm achieves better performance than state-of-the-art methods.
Chaomin Shen 0001, Mixue Yu, Chenxiao Zhao, Yaxin Peng, Guixu Zhang
CIKM5
2018 Multi-scale Residual Network for Image Super-Resolution
Juncheng Li 0003, Faming Fang, Kangfu Mei, Guixu Zhang
ECCV (8)4
2018 Hybrid Noise for LIC-Based Pencil Hatching Simulation
abstract
Line Integral Convolution (LIC) has been widely adopted in pencil hatching generation, where an image degraded by random binary white noise (RBWN) is filtered by LIC along a priori vector field. Nonetheless, an RBWN degraded image through LIC produces hatching graduation only in terms of stroke intensity, and can neither produce hatching graduation explicitly in stroke density nor create visually clear hatching strokes while input pixel values are fairly low. In this paper we address these issues from a noise point of view by assessing several noise models and subsequently constructing a new noise model, called hybrid noise. The new noise model has been experimentally demonstrated in simulating hatching graduation in terms of both stroke intensity and stroke density with quantified graduality measurements. To oversee the effectiveness of hybrid noise we implement the whole pipeline of pencil drawing simulation and compare our results with those state of the art algorithms.
Qunye Kong, Yun Sheng, Guixu Zhang
ICME3
2018 Dual-Way Guided Depth Image Inpainting with RGBD Image Pairs
Yun Sheng, Guixu Zhang
MMM (1)4
2018 Preselection via classification: A case study on evolutionary multiobjective optimization
Aimin Zhou, Ke Tang 0001, Guixu Zhang
Inf. Sci.4
2018 A PDE patch-based spectral method for progressive mesh compression and mesh denoising
Qiqi Shen, Yun Sheng, Congkun Chen, Guixu Zhang, Hassan Ugail
Vis. Comput.4
2017 Accelerating MOEA/D by Nelder-Mead method
abstract
The multiobjective evolutionary algorithm based on decomposition (MOEA/D) converts a multiobjective optimization problem into a set of single-objective subproblems, and tackles them simultaneously. In MOEA/D, the offspring generation is a crucial part to increase the convergence of the algorithm and maintain the diversity of the solution set. Currently, the majority of reproduction operators consider the quality of neighborhood exploration, i.e., the capability to distribute along the population structure, while few operators have good capability for subproblem exploitation, i.e., the ability to push solutions forward along the subproblems. To address this issue in this paper, we introduce one of the derivative-free optimization methods, Nelder-Mead simplex (NMS) method, to MOEA/D to accelerate the algorithm convergence. The NMS operator is combined with a differential evolution (DE) operator in the offspring generation. The comparison study demonstrates that calling the NMS operator occasionally can help to accelerate the convergence.
Aimin Zhou, Guixu Zhang, Hemant K. Singh
CEC3
2017 Illumination-Preserving Embroidery Simulation for Non-photorealistic Rendering
Qiqi Shen, Dele Cui, Yun Sheng, Guixu Zhang
MMM (2)4
2017 A PDE-based head visualization method with CT data
abstract
Abstract In this paper, we extend the use of the partial differential equation (PDE) method to head visualization with computed tomography (CT) data and show how the two primary medical visualization means, surface reconstruction, and volume rendering can be integrated into one single framework through PDEs. Our scheme first performs head segmentation from CT slices using a variational approach, the output of which can be readily used for extraction of a small set of PDE boundary conditions. With the extracted boundary conditions, head surface reconstruction is then executed. Because only a few slices are used, our method can perform head surface reconstruction more efficiently in both computational time and storage cost than the widely used marching cubes algorithm. By elaborately introducing a third parameterwto the PDE method, a solid head can be created, based on which the head volume is subsequently rendered with 3D texture mapping. Instead of designing a transfer function, we associate the alpha value of texels of the 3D texture with the PDE parameterwthrough a linear transform. This association enables the production of a visually translucent head volume. The experimental results demonstrate the feasibility of the developed head visualization method. Copyright © 2015 John Wiley & Sons, Ltd.
Congkun Chen, Yun Sheng, Fang Li 0004, Guixu Zhang, Hassan Ugail
Comput. Animat. Virtual Worlds4
2017 Image-based embroidery modeling and rendering
abstract
Abstract Embroidery is a traditional handicraft of sewing stitches into fabric or other materials in different patterns, and this ancient non‐photorealistic art form has not drawn enough attention thus far. In this paper, we present an image‐based method to simulate the traditional embroidery art. The method combines stroke‐based rendering techniques with the Phong lighting model to create picturesque embroidery‐like images. We first build a 3D stitch model and derive some most commonly used stitch patterns from it. Then we preprocess the input image by segmenting it into regions, from which the parameters to specify stitch patterns are obtained. Finally, we apply stitches back onto the desired regions and render them under a virtual light source. Experimental results show that our method, different from the existing schemes, is capable of performing fine embroidery simulations with the effects of lighting and shading based on an input image. Copyright © 2016 John Wiley & Sons, Ltd.
Dele Cui, Yun Sheng, Guixu Zhang
Comput. Animat. Virtual Worlds3
2017 A heuristic convexity measure for 3D meshes
Yun Sheng, Guixu Zhang
Vis. Comput.4
2016 Joint distribution adaptation based TSK Fuzzy logic system for epileptic EEG signal identification
abstract
Transfer learning based method, which utilizes plenty labeled data in the source domain to build an accuracy classifier for the target domain, serves as an effective means in the epileptic detection by using electroencephalogram (EEG) signals. Among existing approaches, Fuzzy logic system (FLS) based on transductive transfer learning is an efficient method due to its superior interpretability and strong learning abilities. However, this kind of method cannot simultaneously reduce the differences in both marginal distributions and conditional distributions between the training and test datasets of EEG signals. To overcome this problem, in this paper, we construct a Takagi-Sugeno-Kang (TSK) FLS based on the joint distribution adaptation (JDA), which refers to TSK-JDA-FLS. It aims to match both marginal and conditional distributions, and we extend the algorithm to perform a multi-class classification for identifying epileptic EEG signals. Extensive experiments verify that TSK-JDA-FLS significantly outperforms competitive non-transfer learning and transfer learning methods in the epileptic EEG datasets.
Yaxin Peng, Guixu Zhang, Chaomin Shen 0001
BIBM3
2016 Sampling in latent space for a mulitiobjective estimation of distribution algorithm
abstract
A regularity model-based multiobjective estimation of distribution algorithm (RM-MEDA) has been proposed for continuous multiobjective optimization problems. Generating promising solutions to approximate the population is significant to RM-MEDA. In the reproduction of RM-MEDA, it adopts a Latin square design strategy to sample points in the latent space that is extended to cover the whole Pareto set. However, the setting of the extension scale is problem-dependent to some extent. To circumvent this issue, we propose a differential evolution based sampling (DES) scheme for RM-MEDA. DES mutates the projections of the parent solutions in the latent space to generate promising candidate offspring solutions. The empirical experiment results have shown the significant advantages of the DES scheme comparing to the Latin square design.
Bing Dong, Aimin Zhou, Guixu Zhang
CEC3
2016 An estimation of distribution algorithm guided by mean shift
abstract
The estimation of distribution algorithm is widely used to solve global optimization problems in recent years. The basic idea is using machine learning methods to extract relevant features of the search space among the selected individuals and to construct a probabilistic model for sampling new solutions. As we know, EDAs mainly focus on the global distribution information of population and are lack of solution location information. In this paper, we extend our previous work to propose a new EDA guided by the mean shift method, which is originally proposed as a density estimation method and is used as a local search method in this paper. In the new approach, at first a set of candidate solutions are generated by EDA. Then the mean shift method is used to refine some good parent solutions. Finally the sampled candidate solutions and the refined solutions are combined to form the offspring solutions. By this way, the global distribution information and the solution location information are used in offspring reproduction. We apply the new approach to a set of test instances and the experiment results indicate that the new algorithm can obtain good performance in most functions with a faster convergence rate.
Aimin Zhou, Guixu Zhang
CEC3
2016 Self-adaptive spectral cluster number detecting with particle swarm optimization algorithm
abstract
Spectral clustering algorithms have been playing an important role in solving many problems in pattern recognition and image processing. As a well-known spectral clustering algorithm, Normalized Cut has been proved powerful in image segmentation and data clustering. Morever spectral clustering has shown to be more effective in finding clusters than many traditional algorithms such as k-means. However, how to decide the number of clusters is always a crucial problem we confront. It's just yet acknownledge that evolutionary algorithms have a powerful ability to solve such optimization problems. In this paper, we apply a Validity Measure for Fuzzy Clustering(VMFC) to determine the cluster number in spectral clustering with the Particle Swarm Optimization selecting the optimal number of clusters from several possible choices.
Chupeng Zeng, Aimin Zhou, Guixu Zhang
CEC3
2016 A Morphological Building Detection Framework for High-Resolution Optical Imagery Over Urban Areas
abstract
This letter proposes an efficient framework for building detection from coarse to fine using morphological technique for high-resolution optical satellite imagery over urban areas. First, the preliminary result of building regions is obtained by the recently developed morphological building index (MBI) method, which is able to detect potential building structures. However, the raw results derived from the MBI can be subject to a number of false alarms, which are caused by bright soil, roads, and open areas. In this letter, we propose to use morphological spatial pattern analysis as a postprocessing to further optimize the MBI result and remove the commission errors. The original MBI result is then separated into seven mutually exclusive categories-core, islet, loop, bridge, perforation, edge, and branch-by applying a series of morphological transformations such as erosions, geodesic dilation, reconstruction by dilation, anchored skeletonization, etc. The objects corresponding to the generic categories are then analyzed, and the categories corresponding to building parts are maintained, while the others are abandoned. After this postprocessing, the small noisy patches and narrow roads, which were wrongly extracted by the MBI, can be removed. In addition, the shape of the buildings can also be regularized by removing the branches, and the holes contained in the building objects can be identified and filled. Extensive experiments performed on GeoEye-1 and WorldView-2 images confirm the effectiveness and robustness of the proposed morphological building detection framework.
Qian Zhang 0003, Xin Huang 0002, Guixu Zhang
IEEE Geosci. Remote. Sens. Lett.3
2016 Framelet-Based Sparse Unmixing of Hyperspectral Images
abstract
Spectral unmixing aims at estimating the proportions (abundances) of pure spectrums (endmembers) in each mixed pixel of hyperspectral data. Recently, a semi-supervised approach, which takes the spectral library as prior knowledge, has been attracting much attention in unmixing. In this paper, we propose a new semi-supervised unmixing model, termed framelet-based sparse unmixing (FSU), which promotes the abundance sparsity in framelet domain and discriminates the approximation and detail components of hyperspectral data after framelet decomposition. Due to the advantages of the framelet representations, e.g., images have good sparse approximations in framelet domain, and most of the additive noises are included in the detail coefficients, the FSU model has a better antinoise capability, and accordingly leads to more desirable unmixing performance. The existence and uniqueness of the minimizer of the FSU model are then discussed, and the split Bregman algorithm and its convergence property are presented to obtain the minimal solution. Experimental results on both simulated data and real data demonstrate that the FSU model generally performs better than the compared methods.
Guixu Zhang, Faming Fang
IEEE Trans. Image Process.1
2015 A classification and Pareto domination based multiobjective evolutionary algorithm
abstract
In multiobjective evolutionary algorithms, most selection operators are based on the objective values or the approximated objective values. It is arguable that the selection in evolutionary algorithms is a classification problem in nature, i.e., selection equals to classifying the selected solutions into one class and the unselected ones into another class. Following this idea, we propose a classification based preselection for multiobjective evolutionary algorithms. This approach maintains two external populations: one is a positive data set which contains a set of `good' solutions, and the other is a negative data set contains a set of `bad' solutions. In each generation, the two external populations are used to train a classifier firstly, then the classifier is applied to filter the newly generated candidate solutions and only the ones labeled as positive are kept as the offspring solutions. The proposed preselection is integrated into the Pareto domination based algorithm framework in this paper. A systematic empirical study on the influence of different classifiers and different reproduction operators has been done. The experimental results indicate that the classification based preselection can improve the performance of Pareto domination based multiobjective evolutionary algorithms.
Aimin Zhou, Guixu Zhang
CEC3
2015 On neighborhood exploration and subproblem exploitation in decomposition based multiobjective evolutionary algorithms
abstract
The decomposition based multiobjective evolutionary algorithm, denoted as MOEA/D, is an open framework for multiobjective optimization. This paper addresses the reproduction operation in MOEA/D. Generally, the solutions from a neighborhood of a subproblem are chosen as the mating pool for offspring reproduction. Since the Pareto set of an MOP shows some kind of structure in the decision space, the newly generated solutions based on the mating pool are arguable more likely to distribute along the population structure, which is called neighborhood exploration, and less likely to push a solution forward along the subproblem, which is called subproblem exploitation. To balance neighborhood exploration and subproblem exploitation, we propose to utilize both history and neighbor solutions for offspring reproduction. This idea is implemented through two operators based on the multivariate Gaussian distribution model, one is based on neighbor solutions and the other is based on previously visited solutions. When generating a new trial solution for a subproblem, one of the two operators is chosen with a probability. The proposed reproduction strategy is embedded in the MOEA/D framework and applied to a test suite. The comparison study has demonstrated that the new reproduction strategy is promising.
Aimin Zhou, Guixu Zhang, Wenyin Gong
CEC3
2015 Graph Cut Based Mesh Segmentation Using Feature Points and Geodesic Distance
abstract
Both prominent feature points and geodesic distance are key factors for mesh segmentation. With these two factors, this paper proposes a graph cut based mesh segmentation method. The mesh is first preprocessed by Laplacian smoothing. According to the Gaussian curvature, candidate feature points are then selected by a predefined threshold. With DBSCAN (Density-Based Spatial Clustering of Application with Noise), the selected candidate points are separated into some clusters, and the points with the maximum curvature in every cluster are regarded as the final feature points. We label these feature points, and regard the faces in the mesh as nodes for graph cut. Our energy function is constructed by utilizing the ratio between the geodesic distance and the Euclidean distance of vertex pairs of the mesh. The final segmentation result is obtained by minimizing the energy function using graph cut. The proposed algorithm is pose-invariant and can robustly segment the mesh into different parts in line with the selected feature points.
Yun Sheng, Guixu Zhang, Hassan Ugail
CW3
2015 Instant Messenger with Personalized 3D Avatar
abstract
Instant Messengers (IM) have become popular online chat tools in cyber worlds. The existing IMs mainly rely on text, audio, and video for user conversation and communication. The pure text chat is monotonous, while the video chat only occurs when users have their webcoms installed and wish to see each other. In this paper, we propose a new IM with personalized 3D Avatars, and present a prototype of this system. Through our IM, the user can synthesize a personalized 3D avatar by just inputting a 2D self face image. Expressions of the avatar are synthesized by a group of predefined phonemes and visemes, and driven by a text-to-visual speech engine. Moreover, our system supports 3D avatar decoration. During online chat, text is input through the text-to-visual speech engine to drive the generation of voice and animation for the 3D avatar. Tests have demonstrated that our IM with personalized 3D avatars can enhance the fun and vividness of online chat experience.
Yuangang Lu, Yun Sheng, Guixu Zhang
CW3
2015 Change Detection Using L 0 Smoothing and Superpixel Techniques
abstract
We propose an unsupervised change detection method for satellite images using $$L_0$$ smoothing, superpixel techniques and k-means. First, we produce the difference image according to image types (synthetic aperture radar or optical images). Second, we use $$L_0$$ smoothing, an image editing method that can simultaneously sharpen major edges and smooth low-amplitude structures, to generate two difference images with distinct smooth levels. Third, k-means algorithm with $$k=2$$ is applied on one smoothed difference image to cluster all pixels into changed or unchanged classes. Fourth, the Voronoi-Cells (VCells) algorithm is applied on the other difference image to obtain roughly uniform superpixels while preserving local image boundaries. Finally, we calculate the change degree for each superpixel, and the change detection map is produced by using k-means again. The novelties of this paper are that we use the $$L_0$$ smoothing to reduce noise and preserve edges, and utilize the spatial information with the help of superpixel. Experimental results on synthetic aperture radar and optical images show the effectiveness of our approach.
Xiaoliang Shi, Guixu Zhang, Chaomin Shen 0001
KSEM3
2015 Image segmentation framework based on multiple feature spaces
abstract
Image segmentation plays a key role in many fields such as image processing and recognition. Although various segmentation methods have been proposed in recent decades, most of these methods are based on only a single feature space. How to combine various features to image segmentation is a challenge problem. To address this problem, the authors propose to combine different features based on evolutionary multiobjective optimisation. Two optimisation objectives, which are based on colour and texture features, respectively, are therefore designed for image segmentation. The experiments show that the author's method is able to combine multiple features for image segmentation successfully.
Cong Liu 0011, Aimin Zhou, Chunxue Wu, Guixu Zhang
IET Image Process.4
2015 Variational approach for multi-source image fusion
abstract
In this study, the authors propose a variational model for image fusion using a gradient field to describe the features of all input images. The authors’ model is based on energy minimisation and the fused image corresponds to the minimiser of the energy functional. The authors first construct the gradient of fused image by using a weighted sum of the input gradients. Next, to increase the contrast in the fused image, the authors subtract the norm of gradient in the fused image from the functional. Finally, for the purpose of visual uniformity, the authors integrate the inputs using a ‘gray world’ assumption. The authors implement the algorithm using the augmented Lagrangian method. Three sets of images are used to verify the proposed method. Comparisons with other state‐of‐the‐art algorithms show that the proposed algorithm obtains remarkable results.
Sizhang Tang, Faming Fang, Guixu Zhang
IET Image Process.3
2015 An Antinoise Method for Hyperspectral Unmixing
abstract
In this letter, we propose an antinoise method for hyperspectral unmixing. In the antinoise method, all noises are addressed. The following techniques are applied: 1) an endmember dictionary is constructed first to initialize the solution; 2) an approximated L0norm constraint is employed to prune the dictionary and fulfill the sparse coding; and 3) the Itakura-Saito divergence, instead of the Square of Euclidean Distance divergence, is utilized to construct a novel optimization function. The experimental results on both synthetic and real hyperspectral data sets demonstrate the efficacy of the proposed method.
Aimin Zhou, Guixu Zhang, Faming Fang
IEEE Geosci. Remote. Sens. Lett.3
2015 Similarity-Guided and ℓp-Regularized Sparse Unmixing of Hyperspectral Data
abstract
In this letter, we propose a novel sparse unmixing model combined with two effective regularization terms: one is a similarity-weighting constraint, and the other is the ℓp(0p-norm, it has numerical advantages over the convex ℓ1-norm and better approximates the ℓ0-norm theoretically. Moreover, the ℓp-norm regularizer can simultaneously promote sparsity and enforce the abundance sum-to-one constraint. Therefore, this term yields more desirable results in practice. Experimental results on both simulated and real data demonstrate the effectiveness of the proposed model.
Faming Fang, Guixu Zhang
IEEE Geosci. Remote. Sens. Lett.3
2014 An MOEA/D with multiple differential evolution mutation operators
abstract
In evolutionary algorithms, the reproduction operators play an important role. It is arguable that different operators may be suitable for different kinds of problems. Therefore, it is natural to combine multiple operators to achieve better performance. To demonstrate this idea, in this paper, we propose an MOEA/D with multiple differential evolution mutation operators called MOEA/D-MO. MOEA/D aims to decompose a multiobjective optimization problem (MOP) into a number of single objective optimization problems (SOPs) and optimize those SOPs simultaneously. In MOEA/D-MO, we combine multiple operators to do reproduction. Three mutation strategies with randomly selected parameters from a parameter pool are used to generate new trial solutions. The proposed algorithm is applied to a set of test instances with different complexities and characteristics. Experimental results show that the proposed combining method is promising.
Aimin Zhou, Guixu Zhang
IEEE Congress on Evolutionary Computation3
2014 A locally weighted metamodel for pre-selection in evolutionary optimization
abstract
The evolutionary algorithms are usually criticized for their slow convergence. To address this weakness, a variety of strategies have been proposed. Among them, the metamodel or surrogate based approaches are promising since they replace the original optimization objective by a metamodel. However, the metamodel building itself is expensive and therefore the metamodel based evolutionary algorithms are commonly applied to expensive optimization. In this paper, we propose an alternative metamodel, named locally weighted metamodel (LWM), for the pre-selection in evolutionary optimization. The basic idea is to estimate the objective values of candidate offspring solutions for an individual, and choose the most promising one as the offspring solution. Instead of building a global model as many other algorithms do, a LWM is built for each candidate offspring solution in our approach. The LWM based pre-selection is implemented in a multi-operator based evolutionary algorithm, and applied to a set of test instances with different characteristics. Experimental results show that the proposed approach is promising.
Qiuxiao Liao, Aimin Zhou, Guixu Zhang
IEEE Congress on Evolutionary Computation3
2014 Adaptive image segmentation by using mean-shift and evolutionary optimisation
abstract
Undersegmentation or oversegmentation is a challenge faced in image segmentation methods, and it is extreme important to determine the optimal number of regions (clusters) of an image in real‐world applications. In this study, we introduce an adaptive strategy to do so. The basic idea is to firstly oversegment an image by using the Mean‐shift (MS) method, and then segment the obtained oversegmented results by using an evolutionary algorithm. In the second stage, a feature is extracted for each region obtained by the MS method, and a new fitness function is designed to determine the optimal number of clusters. The adaptive approach is applied to a variety of images, and the experimental results show that our method is both efficient and effective for image segmentation.
Cong Liu 0011, Aimin Zhou, Qian Zhang 0003, Guixu Zhang
IET Image Process.4
2014 Framelet based pan-sharpening via a variational method
Faming Fang, Guixu Zhang, Fang Li 0004, Chaomin Shen 0001
Neurocomputing2
2014 A Novel Blind Spectral Unmixing Method Based on Error Analysis of Linear Mixture Model
abstract
It is well known that the linear mixture model (LMM) is attracting much attention due to its simplicity. However, some theoretical analysis reveals that the traditional LMM also impedes the improvement of blind spectral unmixing. For this reason, we propose a novel blind spectral unmixing method (NBSUM) in this letter. NBSUM utilizes the conjugate gradient to calculate end-member spectral and abundance, which can not only overcome some shortcomings of the traditional LMM but also provide more accurate results. NBSUM is compared with some state-of-the-art approaches on both synthetic and real hyperspectral data sets, and the experimental results demonstrate the efficacy of the proposed method.
Faming Fang, Aimin Zhou, Guixu Zhang
IEEE Geosci. Remote. Sens. Lett.4
2013 A differential evolution with an orthogonal local search
abstract
Differential evolution (DE) is a kind of evolutionary algorithms (EAs), which are population based heuristic global optimization methods. EAs, including DE, are usually criticized for their slow convergence comparing to traditional optimization methods. How to speed up the EA convergence while keeping its global search ability is still a challenge in the EA community. In this paper, we propose a differential evolution method with an orthogonal local search (OLSDE), which combines orthogonal design (OD) and EA for global optimization. In each generation of OLSDE, a general DE process is used firstly, and then an OD based local search is utilized to improve the quality of some solutions. The proposed OLSDE is applied to a variety of test instances and compared with a basic DE method and an orthogonal based DE method. The experimental results show that OLSDE is promising for dealing with the given continuous test instances.
Zhenzhen Dai, Aimin Zhou, Guixu Zhang, Sanyi Jiang
IEEE Congress on Evolutionary Computation3
2013 Approximation Model Guided Selection for Evolutionary Multiobjective Optimization
Aimin Zhou, Qingfu Zhang 0001, Guixu Zhang
EMO3
2013 Automatic clustering method based on evolutionary optimisation
abstract
How to set the cluster number plays a key role in many clustering applications. To address this issue, this study introduces an automatic clustering method based on evolutionary algorithms (EAs). The basic idea is to convert a clustering problem into a global optimisation problem and tackle it by an EA. A new validity index, which balances the inter‐cluster consistency and the intra‐cluster consistency, is proposed to be the objective function. Three adaptive coding schemes, which can deal with variable‐length optimisation problems by using a fixed‐length chromosome, are designed to detect the cluster number automatically. The validity index and adaptive coding schemes are incorporated in an EA for automatic clustering. The authors approach is compared with some widely used validity indices and an adaptive coding scheme on some artificial data sets and two real‐world problems. The experimental results suggest that their method not only successfully detects the correct cluster numbers but also achieve stable results for most of test problems.
Cong Liu 0011, Aimin Zhou, Guixu Zhang
IET Comput. Vis.3
2013 A Variational Approach for Pan-Sharpening
abstract
Pan-sharpening is a process of acquiring a high resolution multispectral (MS) image by combining a low resolution MS image with a corresponding high resolution panchromatic (PAN) image. In this paper, we propose a new variational pan-sharpening method based on three basic assumptions: 1) the gradient of PAN image could be a linear combination of those of the pan-sharpened image bands; 2) the upsampled low resolution MS image could be a degraded form of the pan-sharpened image; and 3) the gradient in the spectrum direction of pan-sharpened image should be approximated to those of the upsampled low resolution MS image. An energy functional, whose minimizer is related to the best pan-sharpened result, is built based on these assumptions. We discuss the existence of minimizer of our energy and describe the numerical procedure based on the split Bregman algorithm. To verify the effectiveness of our method, we qualitatively and quantitatively compare it with some state-of-the-art schemes using QuickBird and IKONOS data. Particularly, we classify the existing quantitative measures into four categories and choose two representatives in each category for more reasonable quantitative evaluation. The results demonstrate the effectiveness and stability of our method in terms of the related evaluation benchmarks. Besides, the computation efficiency comparison with other variational methods also shows that our method is remarkable.
Faming Fang, Fang Li 0004, Chaomin Shen 0001, Guixu Zhang
IEEE Trans. Image Process.4
2012 A multiobjective evolutionary algorithm based on decomposition and probability model
abstract
Many real world applications require optimizing multiple objectives simultaneously. Multiobjective evolutionary algorithm based on decomposition (MOEA/D) is a new framework for dealing with such kind of multiobjective optimization problems (MOPs). MOEA/D focuses on how to maintain a set of scalarized sub-problems to approximate the optimum of a MOP. This paper addresses the offspring reproduction operator in MOEA/D. It is arguable that, to design efficient offspring generators, the properties of both the algorithm to use and the problem to tackle should be considered. To illustrate this idea, a generator based on multivariate Gaussian models is proposed under the MOEA/D framework in this paper. In the new generator, both the local and global population distribution information is extracted by a set of Gaussian distribution models; new trial solutions are sampled from the probability models. The proposed approach is applied to a set of benchmark problems with complicated Pareto sets. The comparison study shows that the offspring generator is promising for dealing with continuous MOPs.
Aimin Zhou, Qingfu Zhang 0001, Guixu Zhang
IEEE Congress on Evolutionary Computation3
2012 Lagrangian multipliers and split Bregman methods for minimization problems constrained on Sn-1
Fang Li 0004, Tieyong Zeng, Guixu Zhang
J. Vis. Commun. Image Represent.3
2011 An estimation of distribution algorithm based on nonparametric density estimation
abstract
Probabilistic models play a key role in an estimation of distribution algorithm(EDA). Generally, the form of a probabilistic model has to be chosen before executing an EDA. In each generation, the probabilistic model parameters will be estimated by training the model on a set of selected individuals and new individuals are then sampled from the probabilistic model. In this paper, we propose to use probabilistic models in a different way: firstly generate a set of candidate points, then find some as offspring solutions by a filter which is based on a nonparametric density estimation method. Based on this idea, we propose a nonparametric estimation of distribution algorithm (nEDA) for global optimization. The major differences between nEDA and traditional EDAs are (1) nEDA uses a generating filtering strategy to create new solutions while traditional EDAs use a model building-sampling strategy to generate solutions, and (2) nEDA utilizes a nonparametric density model with traditional EDAs usually utilize parametric density models. nEDA is compared with a traditional EDA which is based on Gaussian model on a set of benchmark problems. The preliminary experimental results show that nEDA is promising for dealing with global optimization problems.
Luhan Zhou, Aimin Zhou, Guixu Zhang, Chuan Shi 0001
IEEE Congress on Evolutionary Computation3
2011 A new information fusion approach for image segmentation
abstract
In this paper we propose a new hybrid image segmentation algorithm that integrate the region-based method with the boundary-based method. More specifically we take an information fusion approach based on the Tensor Voting framework that seamlessly fuse the information from the region-based Mean Shift method with the boundary-based Canny Edge Detection algorithm. We have tested our algorithm on several images from the Caltech 101 database [18]. Experiments results show the new algorithm is very efficient and can achieve very good segmentation results.
Ratchadaporn Kanawong, Ye Duan, Guixu Zhang
ICIP4
2011 Fast image inpainting and colorization by Chambolle's dual method
Fang Li 0004, Ruihua Liu, Guixu Zhang
J. Vis. Commun. Image Represent.4
2010 Variational Color Image Segmentation via Chromaticity-Brightness Decomposition
Yaxin Peng, Guixu Zhang
MMM4
2009 Variational denoising of partly textured images
Fang Li 0004, Chaomin Shen 0001, Chunli Shen, Guixu Zhang
J. Vis. Commun. Image Represent.4