VLDB 2026 Research / reviewers in the wild / expert
Jiangshe Zhang 0001
dblp:74/982-1 · also Jiang-She Zhang 0001
· DBLP profile ↗
133ranked-venue papers
9as first author
64since 2021 · last 2026
0000-0002-8395-1180ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 94 · 7 first-author · 42 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 16 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 1 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scale-invariant information bottleneck for domain generalization
Jiangshe Zhang 0001, Chunxia Zhang 0002, Junmin Liu, Lizhen Ji |
Expert Syst. Appl. | 2 |
| 2026 | Spatial-frequency domain aggregation upsampling for pan-sharpening
Kai Sun 0007, Junying Hu, Junmin Liu, Jiangshe Zhang 0001 |
Neural Networks | 6 |
| 2025 | Task-driven Image Fusion with Learnable Fusion LossabstractMulti-modal image fusion aggregates information from multiple sensor sources, achieving superior visual quality and perceptual features compared to single-source images, often improving downstream tasks. However, current fusion methods for downstream tasks still use predefined fusion objectives that potentially mismatch the downstream tasks, limiting adaptive guidance and reducing model flexibility. To address this, we propose Task-driven Image Fusion (TDFusion), a fusion framework incorporating a learnable fusion loss guided by task loss. Specifically, our fusion loss includes learnable parameters modeled by a neural network called the loss generation module. This module is supervised by the downstream task loss in a meta-learning manner. The learning objective is to minimize the task loss of fused images after optimizing the fusion module with the fusion loss. Iterative updates between the fusion module and the loss module ensure that the fusion network evolves toward minimizing task loss, guiding the fusion process toward the task objectives. TDFusion’s training relies entirely on the downstream task loss, making it adaptable to any specific task. It can be applied to any architecture of fusion and task networks. Experiments demonstrate TDFusion’s performance through fusion experiments conducted on four different datasets, in addition to evaluations on semantic segmentation and object detection tasks. The code is available at https://github.com/HaowenBai/TDFusion. Haowen Bai, Jiangshe Zhang 0001, Zixiang Zhao, Lilun Deng, Yukun Cui |
CVPR | 2 |
| 2025 | Retinex-MEF: Retinex-Based Glare Effects Aware Unsupervised Multi-Exposure Image Fusion
Haowen Bai, Jiangshe Zhang 0001, Zixiang Zhao, Lilun Deng, Yukun Cui |
ICCV | 2 |
| 2025 | S2-AMNet: A lightweight Spatial-Spectral Adaptive Modulation Network for surface defect detection
Jiayong Bao, Chunxia Zhang 0002, Li-Li Bao, Jiangshe Zhang 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Self-similar spectral reasoning network for efficient anti-aliasing seismic data reconstruction
Changpeng Wang, Aoqi Song, Chunxia Zhang 0002, Jiangshe Zhang 0001, Zhiliang Zhou |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Seismic data reconstruction via an adaptive feature fusion networkabstractSeismic data reconstruction is a crucial step in seismic data processing. Traditional methods and deep learning approaches have both been widely used in this field. However, they ignored the interactive learning of inter-channel information, especially in the case of high missing rate where feature extraction became more difficult. To address this issue, we propose an adaptive feature fusion network for the reconstruction of both random and consecutive missing seismic data. The information interaction block is designed into this model to improve the efficiency and adaptability of feature selection. It adaptively emphasizes important feature channels and enables inter-channel information exchange learning. To enhance the ability to capture global and local details, a cross-dimensional feature fusion module is designed at the bottleneck, integrating information from both the channel and spatial dimensions. Additionally, the strategy loss is designed to enable the network to learn the correlations among missing parts of the seismic traces, thereby boosting the reconstruction performance of our model. Compared with other state-of-the-art seismic data reconstruction methods, the proposed algorithm achieves improvements in both qualitative and quantitative evaluations: the reconstruction quality has improved by 20% on both synthetic and field datasets with random missing data. The reconstruction quality has improved by 30% on both synthetic and field datasets with consecutive missing data. At the end of the paper, we conducted ablation experiments, hyperparameter analysis and discussion. Yuting Mu, Changpeng Wang, Chunxia Zhang 0002, Jiangshe Zhang 0001, Junxiong Jia |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Contrastive Learning with Similarity Enhancement for Dimensionality Reduction
Changpeng Wang, Linlin Feng, Lizhen Ji, Jiangshe Zhang 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Automatic fault interpretation method embedded with clustering task in 3D-UNet3+
Chunxia Zhang 0002, Jiangshe Zhang 0001, Chunfeng Tao |
Expert Syst. Appl. | 3 |
| 2025 | ReFusion: Learning Image Fusion from Reconstruction with Learnable Loss Via Meta-Learning
Haowen Bai, Zixiang Zhao, Jiangshe Zhang 0001, Lilun Deng, Yukun Cui, Baisong Jiang |
Int. J. Comput. Vis. | 3 |
| 2025 | CoFM: Alternating convolution and frequency module is also strong for visual recognition
Jiayong Bao, Jiangshe Zhang 0001, Chunxia Zhang 0002 |
Neurocomputing | 2 |
| 2025 | Stabilizing Sharpness-Aware Minimization Through A Simple Renormalization StrategyabstractRecently, sharpness-aware minimization (SAM) has attracted much attention because of its surprising effectiveness in improving generalization performance. However, compared to stochastic gradient descent (SGD), it is more prone to getting stuck at the saddle points, which as a result may lead to performance degradation. To address this issue, we propose a simple renormalization strategy, dubbed Stable SAM (SSAM), so that the gradient norm of the descent step maintains the same as that of the ascent step. Our strategy is easy to implement and flexible enough to integrate with SAM and its variants, almost at no computational cost. With elementary tools from convex optimization and learning theory, we also conduct a theoretical analysis of sharpness-aware training, revealing that compared to SGD, the effectiveness of SAM is only assured in a limited regime of learning rate. In contrast, we show how SSAM extends this regime of learning rate and then it can consistently perform better than SAM with the minor modification. Finally, we demonstrate the improved performance of SSAM on several representative data sets and tasks. Chengli Tan, Jiangshe Zhang 0001, Junmin Liu, Yunda Hao |
J. Mach. Learn. Res. | 2 |
| 2025 | FDSANet: Seismic Data Reconstruction Based on a Frequency-Domain Self-Attention NetworkabstractSeismic data reconstruction is a crucial step in seismic data processing. Most existing methods reconstruct seismic data in the spatial domain, often ignoring some important frequency components in the frequency domain, such as high-frequency texture features. Therefore, we propose a frequency-domain self-attention network (FDSANet) to effectively reconstruct seismic data with high-missing-rate. The wavelet transform is employed in this model to better restore weak signals and provide more information at different resolutions. The fast Fourier transform in the frequency-domain self-attention module (FDSAM) enhances global frequency awareness, especially for high-frequency energy. Different frequency components are element-wise multiplied by dynamic weights, effectively suppressing energy leakage and aliasing. Moreover, the nearest-neighbor similarity loss on adjacent shot gathers is incorporated into the loss function to learn information from neighboring shot gathers, further enhancing the reconstruction performance of our model. Experiments on both synthetic and field datasets demonstrate that FDSANet achieves significant improvement over several state-of-the-art methods. Yuting Mu, Changpeng Wang, Chunxia Zhang 0002, Jiangshe Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | DCTCNet: Sequency discrete cosine transform convolution network for visual recognition
Jiayong Bao, Jiangshe Zhang 0001, Chunxia Zhang 0002, Li-Li Bao |
Neural Networks | 2 |
| 2025 | An information-theoretic learning model based on importance sampling with application in face verification
Jiangshe Zhang 0001, Lizhen Ji, Chunxia Zhang 0002, Yukun Cui |
Pattern Recognit. Lett. | 1 |
| 2025 | Deep Unfolding Multi-Modal Image Fusion Network via Attribution AnalysisabstractMulti-modal image fusion synthesizes information from multiple sources into a single image, facilitating downstream tasks such as semantic segmentation. Current approaches primarily focus on acquiring informative fusion images at the visual display stratum through intricate mappings. Although some approaches attempt to jointly optimize image fusion and downstream tasks, these efforts often lack direct guidance or interaction, serving only to assist with a predefined fusion loss. To address this, we propose an “Unfolding Attribution Analysis Fusion network” (UAAFusion), using attribution analysis to tailor fused images more effectively for semantic segmentation, enhancing the interaction between the fusion and segmentation. Specifically, we utilize attribution analysis techniques to explore the contributions of semantic regions in the source images to task discrimination. At the same time, our fusion algorithm incorporates more beneficial features from the source images, thereby allowing the segmentation to guide the fusion process. Our method constructs a model-driven unfolding network that uses optimization objectives derived from attribution analysis, with an attribution fusion loss calculated from the current state of the segmentation network. We also develop a new pathway function for attribution analysis, specifically tailored to the fusion tasks in our unfolding network. An attribution attention mechanism is integrated at each network stage, allowing the fusion network to prioritize areas and pixels crucial for high-level recognition tasks. Additionally, to mitigate the information loss in traditional unfolding networks, a memory augmentation module is incorporated into our network to improve the information flow across various network layers. Extensive experiments demonstrate our method’s superiority in image fusion and applicability to semantic segmentation. The code is available athttps://github.com/HaowenBai/UAAFusion. Haowen Bai, Zixiang Zhao, Jiangshe Zhang 0001, Baisong Jiang, Lilun Deng, Yukun Cui, Chunxia Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | A Label-Free High-Precision Residual Moveout Picking Method for Depth-Domain Tomography Based on Deep LearningabstractResidual moveout (RMO) provides critical information for depth-domain tomography. The current industry-standard method for fitting RMO involves scanning high-order polynomial equations. However, this analytical approach does not accurately capture abrupt variation of the RMO, leading to low iteration efficiency in tomographic inversion. Supervised learning-based image segmentation methods for picking can effectively capture local variations; however, they encounter challenges such as a scarcity of reliable training samples and the high complexity of post-processing. To address these issues, this study proposes a deep learning-based cascade picking method. It distinguishes accurate and robust RMOs using a segmentation network and a post-processing technique based on trend regression. Additionally, a data synthesis method is introduced, enabling the segmentation network to be trained on synthetic datasets for effective picking in field data. Furthermore, a set of metrics is proposed to quantify the quality of automatically picked RMOs. Experimental results based on both model and real data demonstrate that, compared to semblance-based methods, our approach achieves greater picking density and accuracy. Jiandong Liang, Shuaizhe Liang, Jinping Zhu, Chunxia Zhang 0002, Jiangshe Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | CACNN: Capsule Attention Convolutional Neural Networks for 3D Object RecognitionabstractRecently, view-based approaches, which recognize a 3D object through its projected 2-D images, have been extensively studied and have achieved considerable success in 3D object recognition. Nevertheless, most of them use a pooling operation to aggregate viewwise features, which usually leads to the visual information loss. To tackle this problem, we propose a novel layer called capsule attention layer (CAL) by using attention mechanism to fuse the features expressed by capsules. In detail, instead of dynamic routing algorithm, we use an attention module to transmit information from the lower level capsules to higher level capsules, which obviously improves the speed of capsule networks. In particular, the view pooling layer of multiview convolutional neural network (MVCNN) becomes a special case of our CAL when the trainable weights are chosen on some certain values. Furthermore, based on CAL, we propose a capsule attention convolutional neural network (CACNN) for 3D object recognition. Extensive experimental results on three benchmark datasets demonstrate the efficiency of our CACNN and show that it outperforms many state-of-the-art methods. Kai Sun 0007, Jiangshe Zhang 0001, Zixiang Zhao, Chunxia Zhang 0002, Junmin Liu, Junying Hu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Equivariant Multi-Modality Image FusionabstractMulti-modality image fusion is a technique that combines information from different sensors or modalities, en-abling the fused image to retain complementary features from each modality, such as functional highlights and texture details. However, effective training of such fusion models is challenging due to the scarcity of ground truth fusion data. To tackle this issue, we propose the Equivariant Multi-Modality imAge fusion (EMMA) paradigm for end-to-end self-supervised learning. Our approach is rooted in the prior knowledge that natural imaging responses are equiv-ariant to certain transformations. Consequently, we introduce a novel training paradigm that encompasses a fusion module, a pseudo-sensing module, and an equivariant fusion module. These components enable the net training to follow the principles of the natural sensing-imaging process while satisfying the equivariant imaging prior. Extensive experiments confirm that EMMA yields high-quality fusion results for infraredvisible and medical images, concurrently facilitating downstream multi-modal segmentation and detection tasks. The code is available at https://github.com/Zhaozixiang1228/MMIF-EMMA. Zixiang Zhao, Haowen Bai, Jiangshe Zhang 0001, Yulun Zhang 0001, Kai Zhang 0008, Radu Timofte, Luc Van Gool |
CVPR | 3 |
| 2024 | Image Fusion via Vision-Language ModelabstractImage fusion integrates essential information from multiple images into a single composite, enhancing structures, textures, and refining imperfections. Existing methods predominantly focus on pixel-level and semantic visual features for recognition, but often overlook the deeper text-level semantic information beyond vision. Therefore, we introduce a novel fusion paradigm named image Fusion via vIsion-Language Model (FILM), for the first time, utilizing explicit textual information from source images to guide the fusion process. Specifically, FILM generates semantic prompts from images and inputs them into ChatGPT for comprehensive textual descriptions. These descriptions are fused within the textual domain and guide the visual information fusion, enhancing feature extraction and contextual understanding, directed by textual semantic information via cross-attention. FILM has shown promising results in four image fusion tasks: infrared-visible, medical, multi-exposure, and multi-focus image fusion. We also propose a vision-language dataset containing ChatGPT-generated paragraph descriptions for the eight image fusion datasets across four fusion tasks, facilitating future research in vision-language model-based image fusion. Code and dataset are available at https://github.com/Zhaozixiang1228/IF-FILM. Zixiang Zhao, Lilun Deng, Haowen Bai, Yukun Cui, Yulun Zhang 0001, Haotong Qin, Jiangshe Zhang 0001, Luc Van Gool |
ICML | 9 |
| 2024 | A two-stage spatial prediction modeling approach based on graph neural networks and neural processes
Li-Li Bao, Chunxia Zhang 0002, Jiangshe Zhang 0001 |
Expert Syst. Appl. | 3 |
| 2024 | Attention-based deformable convolutional network for Chinese various dynasties character recognition
Sheng Zhuo, Jiangshe Zhang 0001 |
Expert Syst. Appl. | 2 |
| 2024 | NCLDR: Nearest-Neighbor Contrastive Learning with Dual Correlation Loss for Dimensionality Reduction
Linlin Feng, Changpeng Wang, Kangjian Ge, Jiangshe Zhang 0001 |
Neurocomputing | 5 |
| 2024 | A deep learning-based Monte Carlo simulation scheme for stochastic differential equations driven by fractional Brownian motion
Cornelis W. Oosterlee, Jiangshe Zhang 0001 |
Neurocomputing | 3 |
| 2024 | Automatic Source Point Offset via REINFORCE Based on TransformerabstractSource point offset (SPO) plays a crucial role in geophysical prospecting, as it places the source points away from obstacles to facilitate exploration efforts. However, it is time-consuming for manual work to consider various intricate conditions, such as the smoothness of source lines and uniformity of fold distribution. Moreover, existing methods cost much time in optimizing one specific objective, which limits the applicability to diverse construction areas. To address this challenge, this letter leverages the widely adopted transformer architecture as the model and uses REINFORCE to train this model, by formulating the SPO as a combinatorial optimization problem. To enhance communication among candidate nodes in SPO, a graph attention layer extracts distinct information among these nodes. Experimental results on four field datasets demonstrate comparable performance of our method with the conventional method, while providing a reference for deviated geometry design quickly in field test survey. Li Long, Chunxia Zhang 0002, Li-Li Bao, Jiangshe Zhang 0001, Huibing Zhao |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Spatial multi-attention conditional neural processes
Li-Li Bao, Jiangshe Zhang 0001, Chunxia Zhang 0002 |
Neural Networks | 2 |
| 2024 | Enhancing Sound Source Localization via False Negative EliminationabstractSound source localization aims to localize objects emitting the sound in visual scenes. Recent works obtaining impressive results typically rely on contrastive learning. However, the common practice of randomly sampling negatives in prior arts can lead to the false negative issue, where the sounds semantically similar to visual instance are sampled as negatives and incorrectly pushed away from the visual anchor/query. As a result, this misalignment of audio and visual features could yield inferior performance. To address this issue, we propose a novel audio-visual learning framework which is instantiated with two individual learning schemes: self-supervised predictive learning (SSPL) and semantic-aware contrastive learning (SACL). SSPL explores image-audio positive pairs alone to discover semantically coherent similarities between audio and visual features, while a predictive coding module for feature alignment is introduced to facilitate the positive-only learning. In this regard SSPL acts as a negative-free method to eliminate false negatives. By contrast, SACL is designed to compact visual features and remove false negatives, providing reliable visual anchor and audio negatives for contrast. Different from SSPL, SACL releases the potential of audio-visual contrastive learning, offering an effective alternative to achieve the same goal. Comprehensive experiments demonstrate the superiority of our approach over the state-of-the-arts. Furthermore, we highlight the versatility of the learned representation by extending the approach to audio-visual event classification and object detection tasks. Zengjie Song, Jiangshe Zhang 0001, Yuxi Wang 0001, Junsong Fan, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Sharpness-Aware Lookahead for Accelerating Convergence and Improving GeneralizationabstractLookahead is a popular stochastic optimizer that can accelerate the training process of deep neural networks. However, the solutions found by Lookahead often generalize worse than those found by its base optimizers, such as SGD and Adam. To address this issue, we propose Sharpness-Aware Lookahead (SALA), a novel optimizer that aims to identify flat minima that generalize well. SALA divides the training process into two stages. In the first stage, the direction towards flat regions is determined by leveraging a quadratic approximation of the optimization trajectory, without incurring any extra computational overhead. In the second stage, however, it is determined by Sharpness-Aware Minimization (SAM), which is particularly effective in improving generalization at the terminal phase of training. In contrast to Lookahead, SALA retains the benefits of accelerated convergence while also enjoying superior generalization performance compared to the base optimizer. Theoretical analysis of the expected excess risk, as well as empirical results on canonical neural network architectures and datasets, demonstrate the advantages of SALA over Lookahead. It is noteworthy that with approximately 25% more computational overhead than the base optimizer, SALA can achieve the same generalization performance as SAM which requires twice the training budget of the base optimizer. Chengli Tan, Jiangshe Zhang 0001, Junmin Liu, Yihong Gong |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Convolutional neural network incorporating misclassification information for image recognition
Junying Hu, Rongrong Fei, Fang Du, Peiju Chang, Jiangshe Zhang 0001 |
Soft Comput. | 5 |
| 2024 | Information Theoretical Importance Sampling Clustering and Its Relationship With Fuzzy C-MeansabstractA current assumption of most clustering methods is that the training data and future data are taken from the same distribution. However, this assumption may not hold in most real-world scenarios. In this article, we propose an information theoretical importance sampling based approach for clustering problems (ITISC), which minimizes the worst case of expected distortions under the constraint of distribution deviation. The distribution deviation constraint can be converted to the constraint over a set of weight distributions centered on the uniform distribution derived from importance sampling. The objective of the proposed approach is to minimize the loss under maximum degradation hence the resulting problem is a constrained minimax optimization problem, which can be reformulated as an unconstrained problem using the Lagrange method. The optimization problem can be solved by either an alternating optimization algorithm or a general optimization routine by commercially available software. Experiment results on synthetic datasets and a real-world load forecasting problem validate the effectiveness of the proposed model. Furthermore, we demonstrate that fuzzy c-means is a special case of ITISC with the logarithmic distortion, and this observation provides an interesting physical interpretation for fuzzy exponent$m$. Jiangshe Zhang 0001, Lizhen Ji |
IEEE Trans. Fuzzy Syst. | 1 |
| 2024 | Simultaneous Automatic Picking and Manual Picking Refinement for First-BreakabstractFirst-break picking is a pivotal procedure in processing microseismic data for geophysics and resource exploration. Recent advancements in deep learning have catalyzed the evolution of automated methods for identifying first-break. Nevertheless, the complexity of seismic data acquisition and the requirement for detailed, expert-driven labeling often result in outliers and potential mislabeling within manually labeled datasets. These issues can negatively affect the training of neural networks, necessitating algorithms that handle outliers or mislabeled data effectively. We introduce the Simultaneous Picking and Refinement (SPR) algorithm, designed to handle datasets plagued by outlier samples or even noisy labels. Unlike conventional approaches that regard manual picks as ground truth, our method treats the true first-break as a latent variable within a probabilistic model that includes a first-break labeling prior. SPR aims to uncover this variable, enabling dynamic adjustments and improved accuracy across the dataset. This strategy mitigates the impact of outliers or inaccuracies in manual labels. Intra-site picking experiments and cross-site generalization experiments on publicly available data confirm our method’s performance in identifying first-break and its generalization across different sites. Additionally, our investigations into noisy signals and labels underscore SPR’s resilience to both types of noise and its capability to refine misaligned manual annotations. Moreover, the flexibility of SPR, not being limited to any single network architecture, enhances its adaptability across various deep learning-based picking methods. Focusing on learning from data that may contain outliers or partial inaccuracies, SPR provides a robust solution to some of the principal obstacles in automatic first-break picking. Haowen Bai, Zixiang Zhao, Jiangshe Zhang 0001, Yukun Cui, Chunxia Zhang 0002, Zhenbo Guo |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | UPNet: Uncertainty-Based Picking Deep Learning Network for Robust First Break PickingabstractIn seismic exploration, first break (FB) picking is a crucial aspect in determining subsurface velocity models, significantly influencing the placement of wells. Many deep neural networks (DNNs)-based automatic picking methods have been proposed to accelerate this process. Significantly, the segmentation-based DNN methods provide a segmentation map and then estimate FB from the map using a thresholding technique. However, these automatic methods applied in field datasets cannot ensure robustness, especially in the case of a low signal-to-noise ratio (SNR). In this article, we introduce uncertainty quantification (UQ) into FB picking and propose a novel uncertainty-based picking deep learning network (UPNet). UPNet specifically consists of two DNNs. A Bayesian network infers a posterior distribution by sampling the segmentation map of FB. Subsequently, a regression network integrates the segmentation map, the original trace, and the low-frequency (LF) trace to infer the FB trace by trace. Finally, a decision-making method provides the final FB based on uncertainty analysis, offering robust FB. Importantly, UPNet avoids post-processing to obtain FB using the threshold method, as in the segmentation-based picking methods, and instead provides the FB of each trace end-to-end. Moreover, UPNet not only estimates the uncertainty of the network output but can also filter out predictions with low confidence. Many experiments have shown that UPNet demonstrates higher accuracy and robustness than the deterministic DNN-based model, achieving state-of-the-art (SOTA) performance in field surveys. In addition, we verify that the calculated uncertainty is significant, which can serve as a reference for human decision-making. Jiangshe Zhang 0001, Xiao-Li Wei, Li Long, Chunxia Zhang 0002, Zhenbo Guo |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Seismic Data Interpolation via Denoising Diffusion Implicit Models With Coherence-Corrected ResamplingabstractAccurate interpolation of seismic data is crucial for improving the quality of imaging and interpretation. In recent years, deep learning models such as U-Net and generative adversarial networks (GANs) have been widely applied to seismic data interpolation. However, they often underperform when the training and test missing patterns do not match. To alleviate this issue, here we propose a novel framework that is built upon the multimodal adaptable diffusion models. In the training phase, following the common wisdom, we use the denoising diffusion probabilistic model with a cosine noise schedule. This cosine global noise configuration improves the use of seismic data by reducing the involvement of excessive noise stages. In the inference phase, we introduce the denoising diffusion implicit model (DDIM) to reduce the number of sampling steps. Different from the conventional unconditional generation, we incorporate the known trace information into each reverse sampling step for achieving conditional interpolation. To enhance the coherence and continuity between the revealed traces and the missing traces, we further propose two strategies, including successive coherence correction and resampling. Coherence correction penalizes the mismatches in the revealed traces, while resampling conducts cyclic interpolation between adjacent reverse steps. Extensive experiments on synthetic and field seismic data validate our model’s superiority and demonstrate its generalization capability to various missing patterns and different noise levels with just one training session. In addition, uncertainty quantification and ablation studies are also investigated. Xiao-Li Wei, Chunxia Zhang 0002, Chengli Tan, Deng Xiong, Baisong Jiang, Jiangshe Zhang 0001, Sang-Woon Kim |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Understanding Short-Range Memory Effects in Deep Neural NetworksabstractStochastic gradient descent (SGD) is of fundamental importance in deep learning. Despite its simplicity, elucidating its efficacy remains challenging. Conventionally, the success of SGD is ascribed to the stochastic gradient noise (SGN) incurred in the training process. Based on this consensus, SGD is frequently treated and analyzed as the Euler-Maruyama discretization of stochastic differential equations (SDEs) driven by either Brownian or Lévy stable motion. In this study, we argue that SGN is neither Gaussian nor Lévy stable. Instead, inspired by the short-range correlation emerging in the SGN series, we propose that SGD can be viewed as a discretization of an SDE driven by fractional Brownian motion (FBM). Accordingly, the different convergence behavior of SGD dynamics is well-grounded. Moreover, the first passage time of an SDE driven by FBM is approximately derived. The result suggests a lower escaping rate for a larger Hurst parameter, and thus, SGD stays longer in flat minima. This happens to coincide with the well-known phenomenon that SGD favors flat minima that generalize well. Extensive experiments are conducted to validate our conjecture, and it is demonstrated that short-range memory effects persist across various model architectures, datasets, and training strategies. Our study opens up a new perspective and may contribute to a better understanding of SGD. Chengli Tan, Jiangshe Zhang 0001, Junmin Liu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | CDDFuse: Correlation-Driven Dual-Branch Feature Decomposition for Multi-Modality Image FusionabstractMulti-modality (MM) image fusion aims to render fused images that maintain the merits of different modalities, e.g., functional highlight and detailed textures. To tackle the challenge in modeling cross-modality features and decomposing desirable modality-specific and modality-shared features, we propose a novel Correlation-Driven feature Decomposition Fusion (CDDFuse) network. Firstly, CDDFuse uses Restormer blocks to extract cross-modality shallow features. We then introduce a dual-branch Transformer-CNN feature extractor with Lite Transformer (LT) blocks leveraging long-range attention to handle low-frequency global features and Invertible Neural Networks (INN) blocks focusing on extracting high-frequency local information. A correlation-driven loss is further proposed to make the low-frequency features correlated while the high-frequency features uncorrelated based on the embedded information. Then, the LT-based global fusion and INN-based local fusion layers output the fused image. Extensive experiments demonstrate that our CDDFuse achieves promising results in multiple fusion tasks, including infrared-visible image fusion and medical image fusion. We also show that CDDFuse can boost the performance in downstream infrared-visible semantic segmentation and object detection in a unified benchmark. The code is available at https://github.om/haozixiang1228/MMIF-CDDFuse. Zixiang Zhao, Haowen Bai, Jiangshe Zhang 0001, Yulun Zhang 0001, Zudi Lin, Radu Timofte, Luc Van Gool |
CVPR | 3 |
| 2023 | Spherical Space Feature Decomposition for Guided Depth Map Super-ResolutionabstractGuided depth map super-resolution (GDSR), as a hot topic in multi-modal image processing, aims to upsample low-resolution (LR) depth maps with additional information involved in high-resolution (HR) RGB images from the same scene. The critical step of this task is to effectively extract domain-shared and domain-private RGB/depth features. In addition, three detailed issues, namely blurry edges, noisy surfaces, and over-transferred RGB texture, need to be addressed. In this paper, we propose the Spherical Space feature Decomposition Network (SSDNet) to solve the above issues. To better model cross-modality features, Restormer block-based RGB/depth encoders are employed for extracting local-global features. Then, the extracted features are mapped to the spherical space to complete the separation of private features and the alignment of shared features. Shared features of RGB are fused with the depth features to complete the GDSR task. Subsequently, a spherical contrast refinement (SCR) module is proposed to further address the detail issues. Patches that are classified according to imperfect categories are input into the SCR module, where the patch features are pulled closer to the ground truth and pushed away from the corresponding imperfect samples in the spherical feature space via contrastive learning. Extensive experiments demonstrate that our method can achieve state-of-the-art results on four test datasets, as well as successfully generalize to real-world scenes. The code is available at https://github.com/Zhaozixiang1228/GDSR-SSDNet. Zixiang Zhao, Jiangshe Zhang 0001, Chengli Tan, Yulun Zhang 0001, Radu Timofte, Luc Van Gool |
ICCV | 2 |
| 2023 | DDFM: Denoising Diffusion Model for Multi-Modality Image FusionabstractMulti-modality image fusion aims to combine different modalities to produce fused images that retain the complementary features of each modality, such as functional highlights and texture details. To leverage strong generative priors and address challenges such as unstable training and lack of interpretability for GAN-based generative methods, we propose a novel fusion algorithm based on the denoising diffusion probabilistic model (DDPM). The fusion task is formulated as a conditional generation problem under the DDPM sampling framework, which is further divided into an unconditional generation subproblem and a maximum likelihood subproblem. The latter is modeled in a hierarchical Bayesian manner with latent variables and inferred by the expectation-maximization (EM) algorithm. By integrating the inference solution into the diffusion sampling iteration, our method can generate high-quality fused images with natural image generative priors and cross-modality information from source images. Note that all we required is an unconditional pre-trained generative model, and no fine-tuning is needed. Our extensive experiments indicate that our approach yields promising fusion results in infrared-visible image fusion and medical image fusion. The code is available at https://github.com/Zhaozixiang1228/MMIF-DDFM. Zixiang Zhao, Haowen Bai, Yuanzhi Zhu 0001, Jiangshe Zhang 0001, Yulun Zhang 0001, Kai Zhang 0008, Deyu Meng, Radu Timofte, Luc Van Gool |
ICCV | 4 |
| 2023 | Dimensionality reduction by t-Distribution adaptive manifold embedding
Changpeng Wang, Linlin Feng, Tianjun Wu, Jiangshe Zhang 0001 |
Appl. Intell. | 5 |
| 2023 | MBIAN: Multi-level bilateral interactive attention network for multi-modal image processing
Kai Sun 0007, Jiangshe Zhang 0001, Chunxia Zhang 0002, Junying Hu |
Expert Syst. Appl. | 2 |
| 2023 | Consecutively Missing Seismic Data Reconstruction Via Wavelet-Based Swin Residual NetworkabstractMissing traces reconstruction is a key step for seismic data processing. In recent years, researchers have proposed various interpolation methods for seismic trace reconstruction. However, their models are hard to recover the weak signals in the consecutively missing case. Moreover, convolution operation used in these models is not sensitive to long-term dependencies and global information, which affects the reconstruction of the middle part of the missing area. To solve these problems, we propose a wavelet-based swin residual network (WSRN) for seismic data reconstruction. The swin residual block is designed into the U-net framework to improve the local and non-local modeling ability. Furthermore, by replacing the normal sampling layer, the multi-level wavelet transform is introduced to enhance the recovery ability of weak signals, and data augmentation strategy and a hybrid loss function are used to improve the reconstruction performance of WSRN. Experimental results on synthetic and field datasets illustrate that WSRN achieves significant improvement over some representative deep learning methods. Anguo Dong, Changpeng Wang, Chunxia Zhang 0002, Jiangshe Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2023 | Hybrid Shot2Shot and Re-De-Noising Regularization for Random Noise Attenuation of Seismic DataabstractRandom noise attenuation is essential in seismic data processing. In this paper, we propose an unsupervised method called “shot2shot with re-de-noising regularization” to remove random noise. Shot2Shot (S2S) is a new way to train a denoising neural network. S2S takes a shot-gather and its multiple neighboring shot-gathers as input and labels of the neural network, respectively. The principle that S2S can eliminate noise is the correlation of seismic waves and the independence of random noise between neighboring shot-gathers. Because neural networks are more likely to learn correlated information between inputs and labels rather than independent information. Although S2S is effective in denoising, this mode of training may lead to relatively coarse results. Therefore, we propose re-de-noising regularization to make the results of S2S more refined. The re-de-noising regularization consists of two penalty terms that balance each other, the re-de-noising term and the stability term. The stability term is responsible for introducing more fine content from the observations, such as weak waves, but this can introduce new noise. Thus the re-de-noising term is used to avoid the interference of this new noise. Experimentally, our method outperforms other state-of-the-art methods in terms of quantitative results. Visually, our method not only removes the noise but also reconstructs the noisy data more completely. In addition, we explain the role of S2S and re-de-noising regularization more intuitively through ablation experiments. Finally, the robustness of the key hyperparameters is discussed. Aoqi Song, Changpeng Wang, Chunxia Zhang 0002, Jiangshe Zhang 0001, Xiao-Li Wei, Xiong Deng |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Multi-agent deep reinforcement learning algorithm with trend consistency regularization for portfolio management
Cong Ma 0005, Jiangshe Zhang 0001, Zongxin Li |
Neural Comput. Appl. | 2 |
| 2023 | Long Short-Term Memory Networks with Multiple Variables for Stock Market Prediction
Jiangshe Zhang 0001, Chunxia Zhang 0002, Cong Ma 0005 |
Neural Process. Lett. | 2 |
| 2023 | Regeneration-Constrained Self-Supervised Seismic Data InterpolationabstractSeismic data interpolation is an indispensable part of seismic data processing. In recent years, deep-learning-based interpolation algorithms for seismic data have become popular due to their high accuracy. However, a considerable amount of work has focused on the migration of concepts and algorithms in deep-learning-based methods while ignoring the implicit properties of seismic data itself. In this article, we propose the regeneration prior, which is an implicit property of seismic data with respect to the interpolation function, and are used for self-supervised seismic data interpolation tasks. In mathematical form, the regeneration prior can be considered as a regular term describing the structure of the seismic data. Theoretically, the regeneration prior is a necessary condition to obtain an optimal interpolation function. Experimentally, the proposed method achieves significant improvement in accuracy and intuitive visualization in comparison with advanced unsupervised or self-supervised methods. In addition, we provide an intuitive interpretation of the regeneration prior, and our study shows that the regeneration prior plays an anti-overfitting structuring role in the parameter learning process of the interpolation function. Finally, we analyze the robustness of the regeneration prior. The experimental results show that the performance of the regeneration prior is stable despite the fact that the hyperparameters associated with the regeneration prior are perturbed in a considerable range. Aoqi Song, Changpeng Wang, Chunxia Zhang 0002, Jiangshe Zhang 0001, Xiong Deng, Xiao-Li Wei |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Discrete Cosine Transform Network for Guided Depth Map Super-ResolutionabstractGuided depth super-resolution (GDSR) is an essential topic in multi-modal image processing, which reconstructs high-resolution (HR) depth maps from low-resolution ones collected with suboptimal conditions with the help of HR RGB images of the same scene. To solve the challenges in interpreting the working mechanism, extracting cross-modal features and RGB texture over-transferred, we propose a novel Discrete Cosine Transform Network (DCTNet) to alleviate the problems from three aspects. First, the Discrete Cosine Transform (DCT) module reconstructs the multi-channel HR depth features by using DCT to solve the channel-wise optimization problem derived from the image domain. Second, we introduce a semi-coupled feature extraction module that uses shared convolutional kernels to extract common information and private kernels to extract modality-specific information. Third, we employ an edge attention mechanism to highlight the contours informative for guided upsampling. Extensive quantitative and qualitative evaluations demonstrate the effectiveness of our DCTNet, which outperforms previous state-of-the-art methods with a relatively small number of parameters. The code is available at https://github.com/Zhaozixiang1228/GDSR-DCTNet. Zixiang Zhao, Jiangshe Zhang 0001, Zudi Lin, Hanspeter Pfister |
CVPR | 2 |
| 2022 | Semi-supervised nonnegative matrix factorization with positive and negative label propagations
Changpeng Wang, Jiangshe Zhang 0001, Tianjun Wu |
Appl. Intell. | 2 |
| 2022 | Forecasting stock volatility and value-at-risk based on temporal convolutional networks
Chunxia Zhang 0002, Xingfang Huang, Jiangshe Zhang 0001, Hua-Chuan Huang |
Expert Syst. Appl. | 4 |
| 2022 | Seismic Data Reconstruction via Recurrent Residual Multiscale InferenceabstractSeismic data reconstruction is an important technology in seismic data processing. Existing reconstruction methods have achieved promising performance for regularly/randomly missing cases. However, recovering consecutive missing data remains challenging due to the loss of large amounts of information in local regions. In this paper, we devise a novel network called RRMFI-Net, which is mainly constructed by a Recurrent Residual Multiscale Feature Inference (RRMFI) module and a Recurrence Adjustment Attention (RAA) module. The RRMFI module infers and fills the missing regions multiple times, and uses the result as a clue for the next inference, which makes the result more elegant. To ensure that there is no ambiguity between the results of multiple inferences, we devise an RRA module, which is fused into the RRMFI module to obtain padding information from a long distance. Experimentally, we compare RRMFI-Net with supervised state-of-the-art methods, demonstrating that RRMFI-Net is more effective on multiple indicators. Furthermore, we conduct ablation studies discussing the impact of key network hyperparameters. Aoqi Song, Changpeng Wang, Chunxia Zhang 0002, Jiangshe Zhang 0001, Xiong Deng |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | A Novel Data Augmentation Method for Chinese Character Spatial Structure Recognition by Normalized Deformable Convolutional Networks
Sheng Zhuo, Jiangshe Zhang 0001, Chunxia Zhang 0002 |
Neural Process. Lett. | 2 |
| 2022 | Hyperspectral image denoising by low-rank models with hyper-Laplacian total variation prior
Jiangshe Zhang 0001, Chunxia Zhang 0002 |
Signal Process. | 2 |
| 2022 | Semi-supervised graph regularized nonnegative matrix factorization with local coordinate for image representation
Huirong Li, Yuelin Gao, Junmin Liu, Jiangshe Zhang 0001 |
Signal Process. Image Commun. | 4 |
| 2022 | Efficient and Model-Based Infrared and Visible Image Fusion via Algorithm UnrollingabstractInfrared and visible image fusion (IVIF) expects to obtain images that retain thermal radiation information from infrared images and texture details from visible images. In this paper, a model-based convolutional neural network (CNN) model, referred to as Algorithm Unrolling Image Fusion (AUIF), is proposed to overcome the shortcomings of traditional CNN-based IVIF models. The proposed AUIF model starts with the iterative formulas of two traditional optimization models, which are established to accomplish two-scale decomposition, i.e., separating low-frequency base information and high-frequency detail information from source images. Then the algorithm unrolling is implemented where each iteration is mapped to a CNN layer and each optimization model is transformed into a trainable neural network. Compared with the general network architectures, the proposed framework combines the model-based prior information and is designed more reasonably. After the unrolling operation, our model contains two decomposers (encoders) and an additional reconstructor (decoder). In the training phase, this network is trained to reconstruct the input image. While in the test phase, the base (or detail) decomposed feature maps of infrared/visible images are merged respectively by an extra fusion layer, and then the decoder outputs the fusion image. Qualitative and quantitative comparisons demonstrate the superiority of our model, which can robustly generate fusion images containing highlight targets and legible details, exceeding the state-of-the-art methods. Furthermore, our network has fewer weights and faster speed. Zixiang Zhao, Jiangshe Zhang 0001, Chengyang Liang, Chunxia Zhang 0002, Junmin Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Toward a Controllable Disentanglement NetworkabstractThis article addresses two crucial problems of learning disentangled image representations, namely, controlling the degree of disentanglement during image editing, and balancing the disentanglement strength and the reconstruction quality. To encourage disentanglement, we devise distance covariance-based decorrelation regularization. Further, for the reconstruction step, our model leverages a soft target representation combined with the latent image code. By exploring the real-valued space of the soft target representation, we are able to synthesize novel images with the designated properties. To improve the perceptual quality of images generated by autoencoder (AE)-based models, we extend the encoder-decoder architecture with the generative adversarial network (GAN) by collapsing the AE decoder and the GAN generator into one. We also design a classification-based protocol to quantitatively evaluate the disentanglement strength of our model. The experimental results showcase the benefits of the proposed model. Zengjie Song, Oluwasanmi Koyejo, Jiangshe Zhang 0001 |
IEEE Trans. Cybern. | 3 |
| 2022 | Automatic Velocity Picking Using a Multi-Information Fusion Deep Semantic Segmentation NetworkabstractVelocity picking, a critical step in seismic data processing, has been studied for decades. Although manual picking can produce accurate normal moveout (NMO) velocities from the velocity spectra of prestack gathers, it is time-consuming and becomes infeasible with the emergence of a large amount of seismic data. Numerous automatic velocity picking methods have thus been developed. In recent years, deep learning (DL) methods have produced good results on the seismic data with medium and high signal-to-noise ratios (SNR). Unfortunately, it still lacks a picking method to automatically generate accurate velocities in situations of low SNR. In this paper, we propose a multi-information fusion network (MIFN) to estimate stacking velocity from the fusion information of velocity spectra and stack gather segments (SGS). In particular, we transform the velocity picking problem into a semantic segmentation problem based on the velocity spectrum images. Meanwhile, the information provided by SGS is used as a prior in the network to assist segmentation. The experimental results on two field datasets show that the picking results of MIFN are stable and accurate for the scenarios with medium and high SNR, and it also performs well in low SNR scenarios. Code is made publicly available at https://github.com/newbee-ML/MIFN-Velocity-Picking. Jiangshe Zhang 0001, Zixiang Zhao, Chunxia Zhang 0002, Li Long, Weifeng Geng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Hybrid Loss-Guided Coarse-to-Fine Model for Seismic Data Consecutively Missing Trace ReconstructionabstractSeismic data are generally sampled irregularly and sparsely along spatial coordinates because economic costs and obstacles hinder the regular arrangement of geophones in the field. Thus, the sampled seismic data often contain missing traces which result in difficulties for later processing steps. To alleviate this issue, versatile interpolation methods have been developed to interpolate the missing traces. However, the existing models for recovering seismic data with consecutively missing traces in a large amplitude range tend to produce artifacts and blurred signal details. We propose in this paper a hybrid loss guided coarse-to-fine model which consists of a coarse network and a refinement network to allow different regions of seismic data to be recovered in different stages. The coarse network is designed to reconstruct the strong signals and the refinement network is implemented subsequently to recover the weak signals. In addition, the refinement network focuses its attention on the areas which are not well recovered by the coarse network via a weight-masked mechanism. By resorting to the hybrid loss function L1+SSIM+Relativistic Average Least-Square Generative Adversarial Network (RaLSGAN), our model enables more accurate and realistic signal details to be reconstructed. Experiments with synthetic and field data demonstrate that our model is superior to the existing mainstream approaches and the role of the key components is also investigated through ablation studies. Xiao-Li Wei, Chunxia Zhang 0002, Zixiang Zhao, Xiong Deng, Jiangshe Zhang 0001, Sang-Woon Kim |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2021 | Deep Gradient Projection Networks for Pan-sharpeningabstractPan-sharpening is an important technique for remote sensing imaging systems to obtain high resolution multi-spectral images. Recently, deep learning has become the most popular tool for pan-sharpening. This paper develops a model-based deep pan-sharpening approach. Specifically, two optimization problems regularized by the deep prior are formulated, and they are separately responsible for the generative models for panchromatic images and low resolution multispectral images. Then, the two problems are solved by a gradient projection algorithm, and the iterative steps are generalized into two network blocks. By alternatively stacking the two blocks, a novel network, called gradient projection based pan-sharpening neural network, is constructed. The experimental results on different kinds of satellite datasets demonstrate that the new network out-performs state-of-the-art methods both visually and quantitatively. The codes are available at https://github.com/xsxjtu/GPPNN. Jiangshe Zhang 0001, Zixiang Zhao, Kai Sun 0007, Junmin Liu, Chunxia Zhang 0002 |
CVPR | 2 |
| 2021 | FGF-GAN: A Lightweight Generative Adversarial Network for Pansharpening via Fast Guided FilterabstractPansharpening is a widely used image enhancement technique for remote sensing. Its principle is to fuse the input high-resolution single-channel panchromatic (PAN) image and low-resolution multi-spectral image and to obtain a high-resolution multi-spectral (HRMS) image. The existing deep learning pansharpening method has two shortcomings. First, features of two input images need to be concatenated along the channel dimension to reconstruct the HRMS image, which makes the importance of PAN images not prominent, and also leads to high computational cost. Second, the implicit information of features is difficult to extract through the manually designed loss function. To this end, we propose a generative adversarial network via the fast guided filter (FGF) for pansharpening. In generator, traditional channel concatenation is replaced by FGF to better retain the spatial information while reducing the number of parameters. Meanwhile, the fusion objects can be highlighted by the spatial attention module. In addition, the latent information of features can be preserved effectively through adversarial training. Numerous experiments illustrate that our network generates high-quality HRMS images that can surpass existing methods, and with fewer parameters. Zixiang Zhao, Jiangshe Zhang 0001, Kai Sun 0007, Junmin Liu, Chunxia Zhang 0002 |
ICME | 2 |
| 2021 | Deep Convolutional Sparse Coding Network For Pansharpening With Guidance Of Side InformationabstractPansharpening is a fundamental issue in remote sensing field. This paper proposes a side information partially guided convolutional sparse coding (SCSC) model for pansharpening. The key idea is to split the low resolution multispectral image into a panchromatic image related feature map and a panchromatic image irrelated feature map, where the former one is regularized by the side information from panchromatic images. With the principle of algorithm unrolling techniques, the proposed model is generalized as a deep neural network, called as SCSC pansharpening neural network (SCSC-PNN). Compared with 13 classic and state-of-the-art methods on three satellites, the numerical experiments show that SCSC-PNN is superior to others. The codes are available at https://github.com/xsxjtu/SCSC-PNN. Jiangshe Zhang 0001, Kai Sun 0007, Zixiang Zhao, Junmin Liu, Chunxia Zhang 0002 |
ICME | 2 |
| 2021 | A parallel multi-module deep reinforcement learning algorithm for stock trading
Cong Ma 0005, Jiangshe Zhang 0001, Junmin Liu, Lizhen Ji |
Neurocomputing | 2 |
| 2021 | Saliency-based YOLO for single target detection
Junying Hu, Chuanjin Richard Shi, Jiangshe Zhang 0001 |
Knowl. Inf. Syst. | 3 |
| 2021 | CondenseNet with exclusive lasso regularization
Lizhen Ji, Jiangshe Zhang 0001, Chunxia Zhang 0002, Cong Ma 0005, Kai Sun 0007 |
Neural Comput. Appl. | 2 |
| 2021 | MFIF-GAN: A new generative adversarial network for multi-focus image fusion
Junmin Liu, Zixiang Zhao, Chunxia Zhang 0002, Jiangshe Zhang 0001 |
Signal Process. Image Commun. | 6 |
| 2021 | Global Context-Augmented Objection Detection in VHR Optical Remote Sensing ImagesabstractThe deep learning method, especially convolution neural networks (CNNs), has recently made ground-breaking advances on object detection in very-high-resolution (VHR) optical remote sensing images. However, as CNN is originally designed for the classification of natural images, these methods are not very suitable for object detection of remote sensing images. First, current CNN-based approaches have difficulty to deal with objects that have large rotation variation, which is widely existed in optical remote sensing images. Second, the detectors based on CNN only have limited receptive fields and, thus, can hardly utilize the global contextual information that is essential for accurate detection of small targets. To address these two issues, this article proposes a novel deep learning-based object detection framework, including a geometric transform module (GTM) and a global contextual feature fusion module (GCFM). Especially, the GTM combines rotation and flip transformation to deal with the multiangle characteristics of objects. The GCFM uses a spatial attention mechanism to adaptively involve global contextual information in feature maps to improve the recognition and location accuracy of targets. We introduce the two modules into the YOLOv3 framework to achieve end-to-end detection with high performance and efficiency. Comprehensive evaluations on three publicly available object detection data sets demonstrate the excellent performance of the proposed methods. Jiangshe Zhang 0001, Junmin Liu, Chunxia Zhang 0002, Changsheng Zhou, Shuyun Yang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | DRCNN: Dynamic Routing Convolutional Neural Network for Multi-View 3D Object Recognitionabstract3D object recognition is one of the most important tasks in 3D data processing, and has been extensively studied recently. Researchers have proposed various 3D recognition methods based on deep learning, among which a class of view-based approaches is a typical one. However, in the view-based methods, the commonly used view pooling layer to fuse multi-view features causes a loss of visual information. To alleviate this problem, in this paper, we construct a novel layer called Dynamic Routing Layer (DRL) by modifying the dynamic routing algorithm of capsule network, to more effectively fuse the features of each view. Concretely, in DRL, we use rearrangement and affine transformation to convert features, then leverage the modified dynamic routing algorithm to adaptively choose the converted features, instead of ignoring all but the most active feature in view pooling layer. We also illustrate that the view pooling layer is a special case of our DRL. In addition, based on DRL, we further present a Dynamic Routing Convolutional Neural Network (DRCNN) for multi-view 3D object recognition. Our experiments on three 3D benchmark datasets show that our proposed DRCNN outperforms many state-of-the-arts, which demonstrates the efficacy of our method. Kai Sun 0007, Jiangshe Zhang 0001, Junmin Liu, Ruixuan Yu, Zengjie Song |
IEEE Trans. Image Process. | 2 |
| 2020 | DIDFuse: Deep Image Decomposition for Infrared and Visible Image FusionabstractInfrared and visible image fusion, a hot topic in the field of image processing, aims at obtaining fused images keeping the advantages of source images. This paper proposes a novel auto-encoder (AE) based fusion network. The core idea is that the encoder decomposes an image into background and detail feature maps with low- and high-frequency information, respectively, and that the decoder recovers the original image. To this end, the loss function makes the background/detail feature maps of source images similar/dissimilar. In the test phase, background and detail feature maps are respectively merged via a fusion module, and the fused image is recovered by the decoder. Qualitative and quantitative results illustrate that our method can generate fusion images containing highlighted targets and abundant detail texture information with strong reproducibility and meanwhile surpass state-of-the-art (SOTA) approaches. Zixiang Zhao, Chunxia Zhang 0002, Junmin Liu, Jiangshe Zhang 0001 |
IJCAI | 5 |
| 2020 | Parallel Multi-Environment Shaping Algorithm for Complex Multi-step Task
Cong Ma 0005, Zhizhong Li 0002, Dahua Lin, Jiangshe Zhang 0001 |
Neurocomputing | 4 |
| 2020 | Penetrating the influence of regularizations on neural network based on information bottleneck theory
Jiangshe Zhang 0001, Cong Ma 0005, Junmin Liu |
Neurocomputing | 1 |
| 2020 | Variational Bayesian weighted complex network reconstruction
Chunxia Zhang 0002, Pei Wang 0004, Jiangshe Zhang 0001 |
Inf. Sci. | 4 |
| 2020 | Face clustering via learning a sparsity preserving low-rank graph
Changpeng Wang, Jiangshe Zhang 0001, Xueli Song, Tianjun Wu |
Multim. Tools Appl. | 2 |
| 2020 | Bayesian deep matrix factorization network for multiple images denoising
Chunxia Zhang 0002, Jiangshe Zhang 0001 |
Neural Networks | 3 |
| 2020 | ELMAENet: A Simple, Effective and Fast Deep Architecture for Image Classification
Peiju Chang, Jiangshe Zhang 0001, Rongrong Fei |
Neural Process. Lett. | 2 |
| 2020 | Adaptive quantile low-rank matrix factorization
Chunxia Zhang 0002, Jiangshe Zhang 0001 |
Pattern Recognit. | 3 |
| 2020 | Neural network with multiple connection weights
Jiangshe Zhang 0001, Junying Hu, Junmin Liu |
Pattern Recognit. | 1 |
| 2020 | Weighted-capsule routing via a fuzzy gaussian model
Ouafa Amira, Fang Du, Jiangshe Zhang 0001, Chunxia Zhang 0002, Rafik Hamza |
Pattern Recognit. Lett. | 4 |
| 2020 | Bayesian fusion for infrared and visible images
Zixiang Zhao, Chunxia Zhang 0002, Junmin Liu, Jiangshe Zhang 0001 |
Signal Process. | 5 |
| 2020 | A distributed parallel training method of deep belief networks
Jiangshe Zhang 0001, Chunxia Zhang 0002, Junying Hu |
Soft Comput. | 2 |
| 2020 | Robust CP Tensor Factorization With Skew NoiseabstractThe low-rank tensor factorization (LRTF) technique has received increasing popularity in data science, especially in computer vision applications. Many robust LRTF models have been presented recently. However, none of them take the skewness of data into account. This letter proposes a novel LRTF model for skew data analysis by modeling noise as a Mixture of Asymmetric Laplacians (MoAL). The numerical experiments show that the new model MoAL-LRTF outperforms several state-of-the-art counterparts. The codes for all the experiments are available at https://xsxjtu.github.io/Projects/MoAL/main.html. Xingfang Huang, Chunxia Zhang 0002, Jiangshe Zhang 0001 |
IEEE Signal Process. Lett. | 4 |
| 2020 | HAM-MFN: Hyperspectral and Multispectral Image Multiscale Fusion Network With RAP LossabstractThe fusion of hyperspectral image (HSI) and multispectral image (MSI) is one of the most significant topics in remote sensing image processing. Recently, deep learning (DL) has emerged as an important tool for this task. However, existing DL-based methods have two drawbacks, that is, limited ability for feature extraction and suffering from spectral distortion. To address these issues, this article presents a novel neural network, where sophisticated techniques are employed, including network-in-network convolutional unit, batch normalization, and skip connection. To make full use of the MSI, the proposed model fuses HSI and MSI at different scales. Besides, this article presents a new loss function, called RMSE, angle and Laplacian (RAP) loss (the combination of the relative mean squared error, angle loss, and Laplacian loss), to deal with both spatial and spectral distortions. Experiments conducted on four data sets have verified the rationality of network structure and the proposed loss function and demonstrated that the proposed novel model outperforms state-of-the-art counterparts. Ouafa Amira, Junmin Liu, Chunxia Zhang 0002, Jiangshe Zhang 0001, Guanghai Li |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2020 | Bayesian Transfer Learning for Object Detection in Optical Remote Sensing ImagesabstractIn the literature of object detection in optical remote sensing images, a popular pipeline is first modifying an off-the-shelf deep neural network, then initializing the modified network by pretrained weights on a source data set, and finally fine-tuning the network on a target data set. The procedure works well in practice but might not make full use of underlying knowledge implied by pretrained weights. In this article, we propose a novel method, referred to as Fisher regularization, for efficient knowledge transferring. Based on Bayes' theorem, the method stores underlying knowledge into a Fisher information matrix and fine-tunes parameters based on the knowledge. The proposed method would not introduce extra parameters and is less sensitive to hyperparameters than classical weight decay. Experiments on NWPUVHR-10 and DOTA data sets show that the proposed method is effective and works well with different object detectors. Changsheng Zhou, Jiangshe Zhang 0001, Junmin Liu, Chunxia Zhang 0002, Junying Hu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | A neuro-heuristic approach for recognition of lung diseases from X-ray images
Qiao Ke, Jiangshe Zhang 0001, Wei Wei 0006, Dawid Polap, Marcin Wozniak, Leon Kosmider, Robertas Damasevicius |
Expert Syst. Appl. | 2 |
| 2019 | Discriminative multi-modal deep generative models
Fang Du, Jiangshe Zhang 0001, Junying Hu, Rongrong Fei |
Knowl. Based Syst. | 2 |
| 2019 | FPCANet: Fisher discrimination for Principal Component Analysis Network
Kai Sun 0007, Jiangshe Zhang 0001, Hongwei Yong, Junmin Liu |
Knowl. Based Syst. | 2 |
| 2019 | Convolutional Sparse Representation of Injected Details for PansharpeningabstractIn this letter, we address the pansharpening problem, which focuses on constructing a high-resolution (HR) multispectral (MS) image from a low-resolution (LR) MS and an HR panchromatic (Pan) image. The accuracy of pansharpening method based on sparse representation (SR) mainly depends on the construction of dictionary and the learning of sparse coefficients, while the details injection (DI)-based pansharpening method sharpens the MS bands by adding the proper spatial details from Pan. The combination of SR and DI has been put forward as the pansharpening method based on SR of injected details (SR-D). However, limited to the patch-based manner, pansharpening with traditional SR model faces two disadvantages, i.e., limited ability in detail preservation and high sensitivity to misregistration. In this letter, we replace the traditional SR model with convolutional SR (CSR) as a global SR model in the SR-D method and propose a new pansharpening method called CSR of injected details (CSR-D) to overcome the above-mentioned two drawbacks. Experimental results on the IKONOS and WorldView2 data sets show that the proposed method can achieve remarkable spectral and spatial quality on both reduced scale and full scale. Rongrong Fei, Jiangshe Zhang 0001, Junmin Liu, Fang Du, Peiju Chang, Junying Hu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2019 | Discriminative low-rank representation with Schatten-p norm for image recognition
Changpeng Wang, Jiangshe Zhang 0001 |
Multim. Tools Appl. | 2 |
| 2019 | A new variant of restricted Boltzmann machine with horizontal connections
Jiangshe Zhang 0001, Nannan Ji, Changpeng Wang |
Neural Comput. Appl. | 2 |
| 2019 | Discriminative Representation Learning with Supervised Auto-encoder
Fang Du, Jiangshe Zhang 0001, Nannan Ji, Junying Hu, Chunxia Zhang 0002 |
Neural Process. Lett. | 2 |
| 2019 | Enhance the Performance of Deep Neural Networks via L2 Regularization on the Input of Activations
Jiangshe Zhang 0001, Huirong Li, Changpeng Wang |
Neural Process. Lett. | 2 |
| 2019 | Fast Inference Predictive Coding: A Novel Model for Constructing Deep Neural NetworksabstractAs a biomimetic model of visual information processing, predictive coding (PC) has become increasingly popular for explaining a range of neural responses and many aspects of brain organization. While the development of PC model is encouraging in the neurobiology community, its practical applications in machine learning (e.g., image classification) have not been fully explored yet. In this paper, a novel image processing model called fast inference PC (FIPC) is presented for image representation and classification. Compared with the basic PC model, a regression procedure and a classification layer have been added to the proposed FIPC model. The regression procedure is used to learn regression mappings that achieve fast inference at test time, while the classification layer can instruct the model to extract more discriminative features. In addition, effective learning and fine-tuning algorithms are developed for the proposed model. Experimental results obtained on four image benchmark data sets show that our model is able to directly and fast infer representations and, simultaneously, produce lower error rates on image classification tasks. Zengjie Song, Jiangshe Zhang 0001, Junmin Liu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Robust face recognition via discriminative and common hybrid dictionary learning
Changpeng Wang, Wei Wei 0006, Jiangshe Zhang 0001, Houbing Song |
Appl. Intell. | 3 |
| 2018 | Graph Regularized Nonnegative Matrix Factorization with Sample Diversity for Image Representation
Changpeng Wang, Xueli Song, Jiangshe Zhang 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2018 | An effective hierarchical extreme learning machine based multimodal fusion framework
Fang Du, Jiangshe Zhang 0001, Nannan Ji, Chunxia Zhang 0002 |
Neurocomputing | 2 |
| 2018 | Big data analytics enabled by feature extraction based on partial independence
Qiao Ke, Jiangshe Zhang 0001, Houbing Song, Yan Wan 0001 |
Neurocomputing | 2 |
| 2018 | Symmetric low-rank representation with adaptive distance penalty for semi-supervised learning
Changpeng Wang, Jiangshe Zhang 0001, Fang Du |
Neurocomputing | 2 |
| 2018 | Early stopping aggregation in selective variable selection ensembles for high-dimensional linear regression models
Chunxia Zhang 0002, Jiangshe Zhang 0001, Qingyan Yin |
Knowl. Based Syst. | 2 |
| 2018 | Lp-WGAN: Using Lp-norm normalization to stabilize Wasserstein generative adversarial networks
Changsheng Zhou, Jiangshe Zhang 0001, Junmin Liu |
Knowl. Based Syst. | 2 |
| 2018 | A Deep Neural Network Based on ELM for Semi-supervised Learning of Image Classification
Peiju Chang, Jiangshe Zhang 0001, Junying Hu, Zengjie Song |
Neural Process. Lett. | 2 |
| 2017 | A modified version of Helmholtz machine by using a Restricted Boltzmann Machine to model the generative probability of the top layer
Junying Hu, Jiangshe Zhang 0001, Nannan Ji, Chunxia Zhang 0002 |
Neurocomputing | 2 |
| 2017 | Graph-based discriminative nonnegative matrix factorization with label information
Huirong Li, Jiangshe Zhang 0001, Junmin Liu |
Neurocomputing | 2 |
| 2017 | Generalized extreme learning machine autoencoder and a new deep neural network
Kai Sun 0007, Jiangshe Zhang 0001, Chunxia Zhang 0002, Junying Hu |
Neurocomputing | 2 |
| 2017 | Graph-regularized CF with local coordinate for image representation
Huirong Li, Jiangshe Zhang 0001, Junmin Liu |
J. Vis. Commun. Image Represent. | 2 |
| 2017 | A new regularized restricted Boltzmann machine based on class preserving
Junying Hu, Jiangshe Zhang 0001, Nannan Ji, Chunxia Zhang 0002 |
Knowl. Based Syst. | 2 |
| 2017 | Graph-based discriminative concept factorization for data representation
Huirong Li, Jiangshe Zhang 0001, Junying Hu, Chunxia Zhang 0002, Junmin Liu |
Knowl. Based Syst. | 2 |
| 2017 | A ranking-based strategy to prune variable selection ensembles
Chunxia Zhang 0002, Jiangshe Zhang 0001, Qingyan Yin |
Knowl. Based Syst. | 2 |
| 2016 | A new deep neural network based on a stack of single-hidden-layer feedforward neural networks with randomly fixed hidden neurons
Junying Hu, Jiangshe Zhang 0001, Chunxia Zhang 0002 |
Neurocomputing | 2 |
| 2016 | Class-driven concept factorization for image representation
Huirong Li, Jiangshe Zhang 0001, Junmin Liu |
Neurocomputing | 2 |
| 2016 | On size invariance texture image retrieval by fuzzy logic classifier and scattering statistical features
Jiangshe Zhang 0001 |
Pattern Anal. Appl. | 2 |
| 2015 | Enhancing performance of the backpropagation algorithm via sparse response regularization
Jiangshe Zhang 0001, Nannan Ji, Junmin Liu, Jiyuan Pan, Deyu Meng |
Neurocomputing | 1 |
| 2015 | Learning latent features by nonnegative matrix factorization combining similarity judgments
Jiangshe Zhang 0001, Changpeng Wang, Yu-Qian Yang |
Neurocomputing | 1 |
| 2015 | Singular Value Decomposition Projection for solving the small sample size problem in face recognition
Changpeng Wang, Jiangshe Zhang 0001, Guodong Chang, Qiao Ke |
J. Vis. Commun. Image Represent. | 2 |
| 2014 | Parallel tempering with equi-energy moves for training of restricted boltzmann machinesabstractTraining RBMs is laborious due to the difficulty of sampling from model's distribution. Although using Parallel Tempering (PT) alleviates the problem to some extent, it will result in low swap acceptance ratio when the states' energies of neighboring chains are very different. In this paper, we propose a novel PT algorithm based on the principle of swapping between chains with the same level of energy. This new algorithm partitions the state space obtained by a population of Gibbs sampling chains into several energy rings. In each ring, states have similar energies and swapping of each pair of states are conducted with a probability. Experiments on a toy dataset as well as the MNIST dataset shown that the new algorithm keeps high swap acceptance ration and results in better likelihood scores compared to several training methods. Nannan Ji, Jiangshe Zhang 0001 |
IJCNN | 2 |
| 2014 | Enhancing performance of restricted Boltzmann machines via log-sum regularization
Nannan Ji, Jiangshe Zhang 0001, Chunxia Zhang 0002, Qingyan Yin |
Knowl. Based Syst. | 2 |
| 2014 | A sparse-response deep belief network based on rate distortion theory
Nannan Ji, Jiangshe Zhang 0001, Chunxia Zhang 0002 |
Pattern Recognit. | 2 |
| 2014 | Discriminative restricted Boltzmann machine for invariant pattern recognition with linear transformations
Nannan Ji, Jiangshe Zhang 0001, Chunxia Zhang 0002 |
Pattern Recognit. Lett. | 2 |
| 2014 | Learning ensemble classifiers via restricted Boltzmann machines
Chunxia Zhang 0002, Jiangshe Zhang 0001, Nannan Ji, Gao Guo |
Pattern Recognit. Lett. | 2 |
| 2014 | Spectral Unmixing via Compressive SensingabstractThe recently developed theory of compressive sensing (CS) exhibits enormous potentials in signal recovery. In this paper, we investigate its application on spectral unmixing, which appears in hyperspectral data analysis and is usually based on a linear mixture model (LMM) that assumes that a mixed pixel is a linear combination of a set of pure spectral signatures (called endmembers) weighted by their corresponding abundances. Unlike the classical LMM that is a compact representation, we first extend it to a sparse representation (SR) by using a redundant and known endmember set instead of the complete one. Then, the SR model is multiplied by a random Gaussian measurement matrix, so spectral unmixing is casted in the framework of CS. Finally, the ℓ1-minimization algorithms are used to recover the nonnegative abundances by solving the SR model and our proposed model named CS+SR, respectively. Experimental results on both simulated and real hyperspectral data demonstrate that the CS+SR model, formed by multiplying a random Gaussian matrix on the SR model, can improve, at least in the sense of probability, the ability of the ℓ1-minimization algorithms for recovering the nonnegative sparse abundances. Junmin Liu, Jiangshe Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2014 | Enhancing Low-Rank Subspace Clustering by Manifold RegularizationabstractRecently, low-rank representation (LRR) method has achieved great success in subspace clustering (SC), which aims to cluster the data points that lie in a union of low-dimensional subspace. Given a set of data points, LRR seeks the lowest rank representation among the many possible linear combinations of the bases in a given dictionary or in terms of the data itself. However, LRR only considers the global Euclidean structure, while the local manifold structure, which is often important for many real applications, is ignored. In this paper, to exploit the local manifold structure of the data, a manifold regularization characterized by a Laplacian graph has been incorporated into LRR, leading to our proposed Laplacian regularized LRR (LapLRR). An efficient optimization procedure, which is based on alternating direction method of multipliers (ADMM), is developed for LapLRR. Experimental results on synthetic and real data sets are presented to demonstrate that the performance of LRR has been enhanced by using the manifold regularization. Junmin Liu, Jiangshe Zhang 0001, Zongben Xu |
IEEE Trans. Image Process. | 3 |
| 2012 | Label propagation through sparse neighborhood and its applications
Fei Zang, Jiangshe Zhang 0001 |
Neurocomputing | 2 |
| 2012 | Face recognition using Elasticfaces
Fei Zang, Jiangshe Zhang 0001, Jiyuan Pan |
Pattern Recognit. | 2 |
| 2012 | A New Maximum Simplex Volume Method Based on Householder Transformation for Endmember ExtractionabstractEndmember extraction is very important in hyperspectral image analysis. The accurate identification of endmembers enables target detection and classification and efficient spectral unmixing. Although a number of endmember extraction algorithms have been proposed, such as two state-of-the-art algorithms—vertex component analysis (VCA) and simplex growing algorithm (SGA)—it is still a rather challenging task. In this paper, a new maximum simplex volume method based on Householder transformation (HT), referred to as maximum volume by HT (MVHT), is presented for endmember extraction. The proposed algorithm provides consistent results with low computational complexity, which overcomes the disadvantage of the inconsistent result of VCA and the shortcoming of the high computational cost of SGA resulted from calculating the simplex volume. A comparative study and analysis are conducted among the three endmember extraction algorithms, VCA, SGA, and MVHT, on both simulated and real hyperspectral data. The obtained experimental results demonstrate that the proposed MVHT algorithm generally provides a competitive or even better performance over VCA and SGA. Junmin Liu, Jiangshe Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2011 | Discriminative learning by sparse representation for classification
Fei Zang, Jiangshe Zhang 0001 |
Neurocomputing | 2 |
| 2011 | Large margin based nonnegative matrix factorization and partial least squares regression for face recognition
Jiyuan Pan, Jiangshe Zhang 0001 |
Pattern Recognit. Lett. | 2 |
| 2010 | A method to sparsify the solution of support vector regression
Gao Guo, Jiangshe Zhang 0001, Gai-Ying Zhang |
Neural Comput. Appl. | 2 |
| 2010 | A variant of Rotation Forest for constructing ensemble classifiers
Chunxia Zhang 0002, Jiangshe Zhang 0001 |
Pattern Anal. Appl. | 2 |
| 2009 | A method to sparsify the solution of support vector regression
Gao Guo, Jiangshe Zhang 0001, Gai-Ying Zhang |
Neural Comput. Appl. | 2 |
| 2008 | Enhanced piecewise regression based on deterministic annealing
Jiangshe Zhang 0001, Yuqian Yang, Chenghu Zhou |
Sci. China Ser. F Inf. Sci. | 1 |
| 2008 | RotBoost: A technique for combining Rotation Forest and AdaBoost
Chunxia Zhang 0002, Jiangshe Zhang 0001 |
Pattern Recognit. Lett. | 2 |
| 2007 | Reducing examples to accelerate support vector regression
Gao Guo, Jiangshe Zhang 0001 |
Pattern Recognit. Lett. | 2 |
| 2004 | Improved possibilistic C-means clustering algorithmsabstractA possibilistic approach was proposed in a previous paper for C-means clustering, and two algorithms realizing this approach were reported in two previous papers. Although the possibilistic approach is sound, these two algorithms tend to find identical clusters. In this paper, we modify and improve these algorithms to overcome their shortcoming. The numerical results demonstrate that the improved algorithms can determine proper clusters and they can realize the advantages of the possibilistic approach. Jiangshe Zhang 0001, Yiu-Wing Leung |
IEEE Trans. Fuzzy Syst. | 1 |
| 2003 | Robust clustering by pruning outliersabstractIn many applications of C-means clustering, the given data set often contains noisy points. These noisy points will affect the resulting clusters, especially if they are far away from the data points. In this paper, we develop a pruning approach for robust C-means clustering. This approach identifies and prunes the outliers based on the sizes and shapes of the clusters so that the resulting clusters are least affected by the outliers. The pruning approach is general, and it can improve the robustness of many existing C-means clustering methods. In particular, we apply the pruning approach to improve the robustness of hard C-means clustering, fuzzy C-means clustering, and deterministic-annealing C-means clustering. As a result, we obtain three clustering algorithms that are the robust versions of the existing ones. In addition, we integrate the pruning approach with the fuzzy approach and the possibilistic approach to design two new algorithms for robust C-means clustering. The numerical results demonstrate that the pruning approach can achieve good robustness. Jiangshe Zhang 0001, Yiu-Wing Leung |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2001 | A new data processing method based on a biological model of the compound eye: direction quantization representationabstractThis paper presents a new data representation method called direction quantization representation (DQR) which is motivated by a simplified geometric model of biological compound eye and used in describing the shape of convex hulls of objects. Advantages of DQR include high efficiency and stability in numerical computation, convenience for semidynamic maintenance, suitability for parallel implementation, and applicability to various convex set related problems. Several practical applications are presented which show the feasibility and powerfulness of DQR. Hong Qiao, Jiangshe Zhang 0001, Zongben Xu |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2000 | Clustering by Scale-Space FilteringabstractIn pattern recognition and image processing, the major application areas of cluster analysis, human eyes seem to possess a singular aptitude to group objects and find important structures in an efficient and effective way. Thus, a clustering algorithm simulating a visual system may solve some basic problems in these areas of research. From this point of view, we propose a new approach to data clustering by modeling the blurring effect of lateral retinal interconnections based on scale space theory. In this approach, a data set is considered as an image with each light point located at a datum position. As we blur this image, smaller light blobs merge into larger ones until the whole image becomes one light blob at a low enough level of resolution. By identifying each blob with a cluster, the blurring process generates a family of clustering along the hierarchy. The advantages of the proposed approach are: 1) The derived algorithms are computationally stable and insensitive to initialization and they are totally free from solving difficult global optimization problems. 2) It facilitates the construction of new checks on cluster validity and provides the final clustering a significant degree of robustness to noise in data and change in scale. 3) It is more robust in cases where hyperellipsoidal partitions may not be assumed. 4) it is suitable for the task of preserving the structure and integrity of the outliers in the clustering process. 5) The clustering is highly consistent with that perceived by human eyes. 6) The new approach provides a unified framework for scale-related clustering algorithms derived from many different fields such as estimation theory, recurrent signal processing on self-organization feature maps, information theory and statistical mechanics, and radial basis function neural networks. Yee Leung, Jiangshe Zhang 0001, Zongben Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1998 | Optimal neural network algorithm for on-line string matchingabstractWe consider an online string matching problem in which we find all the occurrences of a pattern of m characters in a text of n characters, where all the characters of the pattern are available before processing, while the characters of the text are input one after the other. We propose a space-time optimal parallel algorithm for this problem using a neural network approach, This algorithm uses m McCulloch-Pitts neurons connected as a linear array. It processes every input character of the text in one step and hence it requires at most n iteration steps. Yiu-Wing Leung, Jiangshe Zhang 0001, Zongben Xu |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 1997 | Neural networks for convex hull computationabstractComputing convex hull is one of the central problems in various applications of computational geometry. In this paper, a convex hull computing neural network (CHCNN) is developed to solve the related problems in the N-dimensional spaces. The algorithm is based on a two-layered neural network, topologically similar to ART, with a newly developed adaptive training strategy called excited learning. The CHCNN provides a parallel online and real-time processing of data which, after training, yields two closely related approximations, one from within and one from outside, of the desired convex hull. It is shown that accuracy of the approximate convex hulls obtained is around O[K(-1)(N-1/)], where K is the number of neurons in the output layer of the CHCNN. When K is taken to be sufficiently large, the CHCNN can generate any accurate approximate convex hull. We also show that an upper bound exists such that the CHCNN will yield the precise convex hull when K is larger than or equal to this bound. A series of simulations and applications is provided to demonstrate the feasibility, effectiveness, and high efficiency of the proposed algorithm. Yee Leung, Jiangshe Zhang 0001, Zongben Xu |
IEEE Trans. Neural Networks | 2 |