Zhang Yi 0001

dblp:63/1807-1 · DBLP profile ↗
← Back
306ranked-venue papers
11as first author
121since 2021 · last 2026
0000-0002-5867-9322ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 231 · 9 first-author · 80 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 25 since 2021Applied, interdisciplinary, general and emerging computing · 34 · 22 since 2021Databases, data management, data science and information retrieval · 19 · 1 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 1 since 2021Theory of computation · 4Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RoSE: A Role Correlation Structure-Enhanced Model for Multi-Event Argument Extraction
abstract
Event co-occurrences have been proven effective for event argument extraction (EAE) in previous studies; however, few have considered intra- and inter-event role correlations. Since role varies among different event types, event structure heterogeneity and overlap pose significant challenges to EAE. To address this issue, we propose a Role Correlation Structure-Enhanced model for Multi-Event Argument Extraction (RoSE), capable of capturing both heterogeneity and overlap of event structures through modeling role correlations. The proposed RoSE model employs a joint context-prompts input, role-centric graph-guided encoder (RoGE), and role-specific information fusion (RoIF). The RoGE is designed to enhance the intra- and inter-event role correlation between prompts and their corresponding event contexts. The RoIF module utilizes intra-event role information to improve multi-event arguments extraction. Extensive experiments on four widely-used benchmarks (RAMS, WikiEvents, MLEE, and ACE05) demonstrate that our proposed approach achieves state-of-the-art performance, validating the effectiveness of incorporating both intra- and inter-event role correlations.
Geting Huang, Kai Zhou 0004, Zhang Yi 0001, Xiuyuan Xu
AAAI4
2026 GeoCoBox: Box-supervised 3D Tumor Segmentation via Geometric Co-embedding
abstract
Data economics drives AI by optimizing data usage, reducing costs, and enhancing efficiency. In 3D tumor segmentation, efficiency is crucial due to the high demand for labor-intensive manual annotations. Box-supervised segmentation offers a promising alternative but is constrained by tumor morphology complexity and boundary ambiguity. In this paper, we propose a novel 3D tumor segmentation model that integrates both positional and embedding features to facilitate inter-task collaboration. We introduce an Anatomical-Driven Class Activation Map to predefine the complex tumor morphology prior, which is further refined by our Geometric Pixel Co-embedding Learner. This learner utilizes contrastive learning to encode semantic information between center and edge pixels, enhancing pixel clustering and progressively refining tumor boundary segmentation in a coarse-to-fine manner. Our approach outperforms existing box-supervised methods in segmentation performance, with extensive experiments on four tumor datasets demonstrating significant improvements. This work provides a cost-effective and efficient solution for tumor segmentation, advancing the application of data economics in medical imaging.
Tianzhong Lan, Zhang Yi 0001, Xiuyuan Xu, Min Zhu 0005
AAAI2
2026 CNM-UNet: Continuous Ordinary Differential Equations for Medical Image Segmentation
abstract
Integrating Ordinary Differential Equations (ODEs) with U-shaped neural networks has emerged as a novel direction in medical image segmentation. Current networks predominantly employ discretization methods incorporating ODEs. However, these methods face inherent trade-offs between model compactness, computational accuracy, and efficiency. Continuous ODE solutions were rarely studied because they face three limitations: high computational costs, long training time, and poor generalization ability. To address these limitations, we propose an innovative Continuous Neural Memory ODE UNet (CNM-UNet), which replaces all hierarchical decoder layers in vanilla UNet with a single Continuous Neural Memory ODEs Block (CNM-Block) decoder, significantly reducing computation costs and improving training efficiency. CNM-UNet leverages ODEs' dynamic properties to establish continuous temporal feature extraction. For alleviating the generalization problem, a DUal SElf-updated (DUSE) strategy based on test-time adaptation principles is introduced to enhance cross-domain generalization. Experimental results demonstrate CNM-UNet's comprehensive advantages in computational capacity, convergence speed, and cross-domain adaptability, offering new insights for practical deployment of continuous ODE methodologies for medical image segmentation.
Yashi Zhu, Quansong He, Kaishen Wang, Zhang Yi 0001, Tao He 0016
AAAI6
2026 Mitigating Entity Hallucinations in 3D Radiology Report Generation via Dual-Stream Alignment
abstract
Entity hallucination poses a major challenge in radiology report generation (RRG), particularly for 3D CT scans where complex spatial contexts amplify factual errors. To address this, medical entity phrases serve as key carriers for multi-modal prompting, integrating expert knowledge into the vision-language model. Current methods use unified cross-attention for volume-phrase alignment, failing to account for anatomical specificity during the alignment process. In this work, we introduce the Dual-stream Entity Alignment Reporting network (DEAR) that separately models organ and lesion entities to resolve anatomical bias. Specifically, the dual-stream entity aligner is designed to partition medical entity phrases into organ and lesion streams, feeding them into separate cross-attention blocks in parallel to achieve fine-grained volume–phrase alignment. For structurally regular and spatially stable organ entities, an organ-guided cross-attention (OGCA) block is proposed to enforce structural consistency by retrieving the top-k voxel tokens via volume–phrase similarity and preserving spatial connectivity through morphological dilation. Meanwhile, a lesion-guided cross-attention (LGCA) block is introduced for structurally irregular and spatially variable lesion entities, enhancing anomaly sensitivity through phrase-weighted attention and refining discriminative boundaries via 3D residual Laplacian filtering. Experiments demonstrate that DEAR significantly reduces entity hallucinations and improves clinical factuality in 3D RRG benchmarks.
Lingyu Zhou, Zhang Yi 0001, Xiuyuan Xu
AAAI3
2026 Class Incremental Medical Image Segmentation via Prototype-Guided Calibration and Dual-Aligned Distillation
abstract
Class incremental medical image segmentation (CIMIS) aims to preserve knowledge of previously learned classes while learning new ones without relying on old-class annotations. However, existing methods 1) either adopt one-size-fits-all strategies that treat all spatial regions and feature channels equally, which may hinder the preservation of accurate old knowledge, 2) or focus solely on aligning local prototypes with global ones for old classes while overlooking their local representations in new data, leading to knowledge degradation. To mitigate the above issues, we propose Prototype-Guided Calibration Distillation (PGCD) and Dual-Aligned Prototype Distillation (DAPD) for CIMIS in this paper. Specifically, PGCD exploits prototype-to-feature similarity to calibrate class-specific distillation intensity in different spatial regions, effectively reinforcing reliable old knowledge and suppressing misleading cues from old classes. Complementarily, DAPD aligns the local prototypes of old classes extracted from the current model with both global historical prototypes and local prototypes, further enhancing segmentation performance on old categories. Comprehensive evaluations on two widely used multi-organ segmentation benchmarks demonstrate that our method outperforms current state-of-the-art methods, highlighting its robustness and generalization capabilities.
Shengqian Zhu, Chengrong Yu, Guangjun Li, Jiafei Wu, Xiaogang Xu 0002, Zhang Yi 0001, Junjie Hu 0004
AAAI8
2026 A neuroscience-based EEG emotion recognition framework with local brain region-guided global learning
Linna Wu, Yuanlun Xie, Kaibo Shi, Nan Zhou 0010, Shiping Wen 0001, Zhang Yi 0001
Expert Syst. Appl.8
2026 Enhancing Exploration and Exploitation in Tumor Treatment Through Action-Guided Deep Reinforcement Learning
abstract
Inverse treatment planning is pivotal in tumor treatment planning. It enables the multi-objective optimization of radiation dose delivery, ensuring precise tumor targeting while sparing surrounding healthy tissues. This process often requires frequent parameter adjustments to achieve the desired balance between objectives, making it both labor-intensive and time-consuming. Deep reinforcement learning (DRL) provides an automated, model-based planning solution, aimed at reducing reliance on human expertise and enhancing the efficiency of objective parameter optimization. However, most current approaches apply DRL to inverse planning without fully leveraging the knowledge embedded in the continuous state-action space, defined by the coupling between nonstationary planning states and continuous decision variables. This may result in insufficient exploration and exploitation, leading to inefficient optimization. This work introduces an innovative action-guided DRL (AgDRL) approach for automatic inverse planning. Our goal is to enhance exploration and exploitation by leveraging insightful guidance from reward-guided actions. The implementation of AgDRL incorporates both exploitation and exploration in the action-state space. For exploitation, high-reward actions are employed as guidance to achieve the optimal action adjustment. For exploration, low-reward actions are recommended as training resets to explore a broader range of the latent state space. Quantitative and qualitative experiments are conducted in various settings to evaluate the proposed method. The results are assessed using DRL-related metrics (e.g. reward gains) and clinical-related measurements (e.g. dose-volume histograms, DVHs). Experimental results on a real-world rectal cancer dataset empirically demonstrate that the proposed AgDRL-based approach significantly improves optimization efficiency through a high-reward strategy while enhancing exploration diversity via a low-reward strategy, consistently outperforming the MatRad treatment planning optimization platform.
Chengrong Yu, Zhonglian Wei, Yuncheng Shen, Yingyong Yin, Zhang Yi 0001, Guangjun Li, Junjie Hu 0004
Int. J. Neural Syst.5
2026 Advancing depression detection in audio through innovative semi-supervised learning technology
Xiang Li 0210, Junjie Hu 0004, Zhang Yi 0001, Yuanyuan Chen 0006
Knowl. Based Syst.6
2026 Enhancing feature fusion of U-like networks with dynamic skip connections
Quansong He, Kaishen Wang, Jianlong Xiong, Zhang Yi 0001, Tao He 0016
Medical Image Anal.5
2026 Layer-wise correlation and attention discrepancy distillation for semantic segmentation
Jianping Gou, Kaijie Chen, Weihua Ou, Xin Luo 0001, Zhang Yi 0001
Pattern Recognit.6
2026 Weighted sample correlation knowledge distillation for visual recognition
Daidai Liu, Jianping Gou, Baosheng Yu, Zhang Yi 0001
Pattern Recognit.6
2026 Cascaded neural memory ODEs for predicting fluence maps in rectal cancer IMRT
Xiangjie Tan, Chengrong Yu, Zhang Yi 0001, Junjie Hu 0004
Pattern Recognit.5
2026 Synergistic knowledge distillation via reciprocal and self learning
Renjie Huang, Jianping Gou, Yibing Zhan, Zhang Yi 0001
Pattern Recognit.7
2026 Dual-Policy Fusion for Multitask Multiagent Reinforcement Learning
abstract
Multiagent reinforcement learning (MARL) has shown strong performance in cooperative tasks. However, most existing approaches are designed for single-task scenarios and struggle to adapt to complex and dynamic environments. Multitask MARL methods aim to improve adaptability by sharing policies across tasks, but they often suffer from negative transfer due to conflicting task-specific knowledge. To address this, we propose dual-policy fusion for multitask MARL (DPF-MTMARL), which explicitly integrates a shared policy for leveraging common knowledge and task-specific policies for capturing task-specific information. Specifically, in DPF-MTMARL, we propose a learning method to efficiently train the task-specific policies and provide corresponding theoretical analysis. Additionally, we derive the theoretical conditions for decentralizing the joint policy and enforce these conditions through a regularization term during training. Extensive experiments demonstrate that DPF-MTMARL significantly outperforms state-of-the-art baselines in both homogeneous and heterogeneous task sets, effectively mitigating negative transfer and enabling robust multitask learning.
Naizhuo Zeng, Mingsheng Fu, Liwei Huang, Hong Qu 0002, Zhang Yi 0001
IEEE Trans. Cybern.7
2026 Rethinking Propagation Methods for Interactive Medical Image Segmentation
abstract
Propagation-based methods have drawn increasing research attention in interactive medical image segmentation. However, existing propagation-based methods face two significant challenges: 1) Due tothe continuous nature of anatomical structures within the organs and tumors throughout the volume, over-propagation is likely to occur as the propagation process reaches the end of structures, leadingto a degradation in segmentation performance. 2) During the multi-round refinement process, selecting the worst-segmented slice for refinement tends to hinder the optimization of segmentation results. To overcome these challenges, we propose the Discrepancy Aware Network (DANet), which includes a Discrepancy Learning Module (DLM) and employs a confidence loss to achieve accurate segmentation. Specifically, DLM captures the temporal-contextual discrepancy between previous and current slices, enabling the model to perceive the variations of the target. Furthermore, the confidence loss is responsible for regularizing the over-confident segmentation at the image level by estimating the target foreground. Additionally, we design a straightforward slice selection strategy to optimize the refinement process. Extensive experimental results on five public medical datasets demonstrate significant improvements over state-of-the-art methods (e.g., with +1.07% improvement on the MSD-Spleen dataset).
Shengqian Zhu, Yuncheng Shen, Yingyong Yin, Zhang Yi 0001, Guangjun Li, Junjie Hu 0004
IEEE J. Biomed. Health Informatics5
2026 Multi-Scale Collaborative Distillation Graph Neural Networks for Session-Based Recommendation
abstract
Session-based recommendation (SBR) in service computing is pivotal in predicting a user's next action based on their current anonymous session. While Graph Neural Network (GNN)-based methods have shown promise in capturing intricate item transformation relationships within sessions, they often fall short in accurately modeling user preferences. This is primarily due to the common practice of solely considering the last item in the session as the user's current interest, neglecting potentially valuable information embedded in other session items which is essential for capturing user global preferences. Moreover, existing models typically optimize performance solely through cross-entropy loss between predicted items and ground truth labels, while overlooking latent valuable knowledge embedded in intermediate features and item-item relationships that lends support to the model in accurately capturing and modeling user preferences. To address these shortcomings, we propose Multi-Scale Collaborative Distillation (MSCD) for SBR. Our approach introduces a current interest adaptive selection module, which dynamically selects appropriate item embeddings as session-local embeddings by evaluating the importance of each item within the session. This allows for a more accurate capture of the user's current true preferences. Additionally, we propose collaborative knowledge distillation, where multiple models are trained concurrently, enabling the transfer of three types of knowledge including response-based, feature-based, and relationship-based knowledge between models, thereby enriching the model's understanding of user preferences. Experimental evaluations conducted on three popular SBR datasets demonstrate that our MSCD model outperforms recent state-of-the-art methods in terms of recommendation accuracy. Our codes are available at:https://github.com/lonely-ice/MSCD.
Jianping Gou, Youhui Cheng, Benteng Ma, Lan Du 0002, Xin Luo 0001, Zhang Yi 0001
IEEE Trans. Serv. Comput.6
2025 ATS and BEMAF-UNet for Accurate and Robust Renal Histopathology Image Segmentation
Linfeng Du, Xiang Li 0210, Zhang Yi 0001, Dehan Li
IEEE Big Data3
2025 Coronary Artery Calcification Segmentation by Using Cross-Frequency Conditioner and Geometric Priors Learning
Weili Jiang, Gadeng Luosang, Yijun Yao, Zhang Yi 0001, Jianyong Wang 0002, Mao Chen 0008
MICCAI (4)7
2025 Domain Generalization for Pulmonary Nodule Detection via Distributionally-Regularized Mamba
Tianzhong Lan, Zhang Yi 0001, Xiuyuan Xu, Min Zhu 0005
MICCAI (6)3
2025 LooBox: Loose-box-supervised 3D Tumor Segmentation with Self-correcting Bidirectional Learning
abstract
Deep learning-based tumor segmentation methods typically require precise pixel-level annotations, which are costly in clinical practice. While bounding box supervision offers a more efficient alternative, existing approaches assume unrealistically tight box annotations, leading to performance degradation when applied to the loose boxes commonly produced by medical annotators. To address this challenge, we propose LooBox, a novel 3D segmentation framework that utilizes loose box annotations through a self-correction and bidirectional rectification paradigm. For the self-correction part, we propose a noise cleaner that comprehensively utilizes deterministic outer box information by integrating three complementary perspectives for predictive self-rectification: entropy mapping, gradient monitoring, and foreground-background affinity measurement. For the bidirectional rectification part, we introduce an augmentation-driven comprehensive consistency constraint strategy. Specifically, the framework incorporates: an asymmetric co-teaching architecture comprising a basic UNet and an enhanced UNet variant with a noise adapter, and an augmentation-driven consistency mechanism that computes pairwise loss between self-corrected predictions after each training iteration to ensure robust tumor feature extraction. Comprehensive evaluations on LIDC-IDRI, MSD-Lung, and MSD-Pancreas datasets demonstrate that LooBox achieves superior segmentation accuracy compared to state-of-the-art box-supervised methods.
Tianzhong Lan, Zhang Yi 0001, Xiuyuan Xu, Min Zhu 0005
ACM Multimedia2
2025 PRIME: Prototype-Driven Class Incremental Learning for Medical Image Segmentation
abstract
Class incremental medical segmentation (CIMS) aims to sequentially learn new classes while preserving knowledge of previously learned categories in the absence of old-class labels. Current methods suffer from performance degradation under class imbalance and require additional segmentation heads to accommodate new categories. Inspired by recent prototype learning that leverages prototypes to achieve robust recognition of new categories under limited-data regimes, we introduce a Prototype-dRIven class increMEntal (PRIME) method. PRIME replaces the incremental segmentation heads with prototypes to mitigate class imbalance, allowing new class learning with the simple addition of new prototypes. Based on prototype learning, PRIME further involves three tailored techniques. First, prototype structure alignment imposes structural constraints on inter-prototype relations to maintain consistent relative distances in the feature space, improving the model's ability to distinguish distinct classes. Second, pixel-wise contrastive loss term groups embeddings of similar samples while separating those of different classes, enhancing segmentation accuracy across all categories. Finally, the consensus-based prototype update mechanism refines the old prototypes during the learning of new classes, preventing performance degradation on the old classes. Extensive experiments on two public multi-organ segmentation datasets demonstrate that our approach significantly outperforms state-of-the-art methods, validating the effectiveness of the proposed PRIME.
Shengqian Zhu, Chengrong Yu, Wenbo Qi, Jiafei Wu, Guangjun Li, Zhang Yi 0001, Xiaogang Xu 0002, Junjie Hu 0004
ACM Multimedia7
2025 MobileODE: An Extra Lightweight Network
abstract
Depthwise-separable convolution has emerged as a significant milestone in the lightweight development of Convolutional Neural Networks (CNNs) over the past decade. This technique consists of two key components: depthwise convolution, which captures spatial information, and pointwise convolution, which enhances channel interactions. In this paper, we propose a novel method to lightweight CNNs through the discretization of Ordinary Differential Equations (ODEs). Specifically, we optimize depthwise-separable convolution by replacing the pointwise convolution with a discrete ODE module, termed the \emph{\textbf{C}hannelwise \textbf{O}DE \textbf{S}olver (COS)}. The COS module is constructed by a simple yet efficient direct differentiation Euler algorithm, using learnable increment parameters. This replacement reduces parameters by over $98.36$\% compared to conventional pointwise convolution. By integrating COS into MobileNet, we develop a new extra lightweight network called MobileODE. With carefully designed basic and inverse residual blocks, the resulting MobileODEV1 and MobileODEV2 reduce channel interaction parameters by $71.0$\% and $69.2$\%, respectively, compared to MobileNetV1, while achieving higher accuracy across various tasks, including image classification, object detection, and semantic segmentation. The code is available at {\url{https://github.com/cashily/MobileODE}}.
Bo Gou, Xiangde Min, Lei Zhang 0005, Zhang Yi 0001, Tao He 0016
NeurIPS6
2025 Simplified Transformer
Lei Xu 0044, Haiying Luo, Zhang Yi 0001
Neurocomputing3
2025 Deformable symmetry attention for nuclear medicine image segmentation
Zeao Zhang, Ruomeng Liu, Huawei Cai, Quan Guo, Zhang Yi 0001
Neurocomputing7
2025 Efficient fine-tuning of vision transformer via path-augmented parameter adaptation
Yao Zhou 0002, Zhang Yi 0001, Gary G. Yen
Inf. Sci.2
2025 Knowledge-embedded large language models for emergency triage
Qing-Yang Shen, Xiaozhi Zhang, Haomin Ren, Quan Guo, Zhang Yi 0001
Knowl. Based Syst.5
2025 Visual prompt-driven universal model for medical image segmentation in radiotherapy
Shengqian Zhu, Chengrong Yu, Zhang Yi 0001, Junjie Hu 0004
Knowl. Based Syst.3
2025 Intra-class progressive and adaptive self-distillation
Jianping Gou, Jiaye Lin, Weihua Ou, Baosheng Yu, Zhang Yi 0001
Neural Networks6
2025 Neural Memory Self-Supervised State Space Models With Learnable Gates
abstract
Discrete Ordinary Differential Equations (ODEs) have been employed to develop lightweight deep neural networks in recent years. In this letter, we introduce a novel lightweight UNet variant called the neural memory Self-Supervised State Space Model (nmS4M-UNet) for medical image segmentation, where discrete ODEs serve as the decoder. The proposed nmS4M block has learnable gates and performs multi-head computation to enhance memory updates. Additionally, the nmS4M-UNet incorporates a self-supervised learning branch to improve feature extraction capabilities. The intermediate features are reused as partial input to the decoder, helping to mitigate network overfitting. The nmS4M-UNet reduces the number of parameters by 29.70% compared to the standard UNet. Experimental results on the PH2, ISIC2018, and BU-COCO datasets demonstrate that the proposed nmS4M-UNet achieves performance comparable to state-of-the-art models.
Zhang Yi 0001, Tao He 0016, Jiajun Bu
IEEE Signal Process. Lett.3
2025 Learning Robust Representations by Autoencoders With Dynamical Implicit Mapping
abstract
Autoencoder is an unsupervised neural network that learns effective representations of data and has wide applications in feature learning, data compression, etc. However, Autoencoder is very sensitive to noise, resulting in low generalization and robustness of the model. To solve this problem, we propose a stable and efficient Autoencoder model called nmFunc-Autoencoder. Inspired by the Neural Memory Ordinary Differential Equation, the Neural Memory Activation Function uses its excellent dynamic nonlinear implicit mapping to establish a mapping relationship between external inputs and stable values to ensure the stability of distinguishable feature extraction, thereby performing better robustness when subjected to noise attacks. We conduct robustness experiments to evaluate its performance. The result showed that compared with other Autoencoder models, the data features extracted by the proposed model are more robust. Subsequently, in the execution efficiency experiments and ablation study, the model was shown to be low-cost and effective.
Jianda Zeng, Weili Jiang, Zhang Yi 0001, Yong-Guo Shi, Jianyong Wang 0002
IEEE Signal Process. Lett.3
2025 Graph Convolutional Networks With Collaborative Feature Fusion for Sequential Recommendation
abstract
Sequential recommendation seeks to understand user preferences based on their past actions and predict future interactions with items. Recently, several techniques for sequential recommendation have emerged, primarily leveraging graph convolutional networks (GCNs) for their ability to model relationships effectively. However, real-world scenarios often involve sparse interactions, where early and recent short-term preferences play distinct roles in the recommendation process. Consequently, vanilla GCNs struggle to effectively capture the explicit correlations between these early and recent short-term preferences. To address these challenges, we introduce a novel approach termed Graph Convolutional Networks with Collaborative Feature Fusion (COFF). Specifically, our method addresses the issue by initially dividing each user interaction sequence into two segments. We then construct two separate graphs for these segments, aiming to capture the user's early and recent short-term preferences independently. To obtain robust prediction, we employ multiple GCNs in a collaborative distillation manner, incorporating a feature fusion module to establish connections between the early and recent short-term preferences. This approach enables a more precise representation of user preferences. Experimental evaluations conducted on five popular sequential recommendation datasets demonstrate that our COFF model outperforms recent state-of-the-art methods in terms of recommendation accuracy.
Jianping Gou, Youhui Cheng, Yibing Zhan, Baosheng Yu, Weihua Ou, Zhang Yi 0001
IEEE Trans. Big Data6
2025 Synthetic Gradient Optimization-Based Implicit Amortized Bayesian Meta-Learning for Few-Shot Pumi Spectrographic Image Recognition
abstract
Meta-learning provides a promising solution to the issue of insufficient training samples in Pumi spectrogram recognition. However, capturing model uncertainty remains a critical challenge, particularly for tasks influenced by lexical ambiguities. To overcome this problem, we propose a novel method, Synthetic Gradient Optimization-Based Implicit Amortized Bayesian Meta-Learning (SGO-IABML), which captures model uncertainty by evaluating posterior distributions within a hierarchical Bayesian framework, thereby facilitating few-shot Pumi spectrogram recognition. Specifically, SGO-IABML reformulates meta-learning as a bi-level variational inference problem, leveraging information bottleneck principles. At the lower level, a generative inference module is developed to implicitly model task-specific variational posteriors, thereby enhancing the model’s expressiveness. Given the lack of analytical forms for implicit distributions, we derive the Fenchel-Bayesian Bound Theorem to measure the divergence between arbitrary distributions. For the meta-learning of variational parameters, SGO-IABML constructs a synthetic gradient optimizer, integrating prior gradient information to facilitate rapid adaptation to new tasks. At the upper level, the model is calibrated by estimating the local geometry of the posterior distribution, utilizing the Generalized Gauss-Newton Matrix to capture the directional sensitivity of the loss function. Comprehensive experimental results on Pumi spectrograms demonstrate that SGO-IABML achieves state-of-the-art performance in generalization, calibration, expressiveness, versatility, and cross-domain adaptability. Furthermore, ablation studies confirm the contribution of each component to the overall performance improvement.
Meijun Fu, Jun Wang 0002, Zhang Yi 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Fourier Boundary Features Network With Wider Catchers for Glass Segmentation
abstract
Glass largely blurs the boundary between the real world and the reflection. The special transmittance and reflectance quality have confused the semantic tasks related to machine vision. Therefore, how to clear the boundary built by glass, and avoid over-capturing features as false positive information in deep structure, matters for constraining the segmentation of reflection surface and penetrating glass. We propose the Fourier Boundary Features Network with Wider Catchers (FBWC), which might represent the first attempt to utilize sufficiently wide horizontal shallow branches without vertical deepening for guiding the fine granularity segmentation boundary through primary glass semantic information. Specifically, we design the Wider Coarse-Catchers (WCC) for anchoring large area segmentation and reducing excessive extraction from a structural perspective. We embed fine-grained features by Cross Transpose Attention (CTA), which is introduced to avoid the incomplete area within the boundary caused by reflection noise. For excavating glass features and balancing high-low layers context, a learnable Fourier Convolution Controller (FCC) is proposed to regulate information integration robustly. The proposed method is validated on three different public glass segmentation datasets. Experimental results reveal that the proposed method yields better segmentation performance compared with the state-of-the-art (SOTA) methods in glass image segmentation.
Xiaolin Qin, Jiacen Liu, Qianlei Wang, Fei Zhu 0004, Zhang Yi 0001
IEEE Trans. Image Process.6
2025 Multi-Teacher Temporal Regulation Network for Surgical Workflow Recognition
abstract
Automatic recognition of surgical workflow plays a vital role in modern operating rooms. Given the complex nature and extended duration of surgical videos, accurate recognition of surgical workflow is highly challenging. Despite being widely studied, existing methods still face two major limitations: insufficient visual feature extraction and performance degradation caused by inconsistency between training and testing features. To address these limitations, this paper proposes a Multi-Teacher Temporal Regulation Network (MTTR-Net) for surgical workflow recognition. To extract discriminative visual features, we introduce a "sequence of clips" training strategy. This strategy employs a set of sparsely sampled video clips as input to train the feature encoder and incorporates an auxiliary temporal regularizer to model long-range temporal dependencies across these clips, ensuring the feature encoder captures critical information from each frame. Then, to mitigate the inconsistency between training and testing features, we further develop a cross-mimicking strategy that iteratively trains multiple feature encoders on different data subsets to generate consistent mimicked features. A temporal encoder is trained on these mimicked features to achieve stable performance during testing. Extensive experiments on eight public surgical video datasets demonstrate that our MTTR-Net outperforms state-of-the-art methods across various metrics. Our code has been released at https://github.com/kaideH/MGTR-Net.
Kaide Huang, Xianglei Yuan, Rui-De Liu, Lian-Song Ye, Yao Zhou 0002, Zhang Yi 0001
IEEE Trans. Medical Imaging7
2025 SA-Seg: Annotation-Efficient Segmentation for Airway Tree Using Saliency-Based Annotation
abstract
Segmentation of the airway tree plays a vital role in clinical practice. However, the complex airway tree structure makes it quite challenging to annotate accurately. Although some annotation-efficient methods have shown promising results in medical image segmentation, most are developed for locally focused segmentation objects and are incompatible with the airway. In this work, we propose an annotation-efficient segmentation method to improve annotation efficiency and tree completeness. It includes a new efficient annotation way and an accompanying segmentation method. The saliency-based annotation method only needs to annotate high-saliency regions, thus greatly improving the annotation efficiency. Inspired by positive-unlabeled learning, we model the dependency relationship between key items in the annotation process to learn from biased weak annotation. The probabilistic model of the annotation process is divided into the score function and the bias function. The score function models the uniform foreground feature representation of the airway, while the bias function models the saliency bias between labeled and unlabeled airway regions. Then, the two models are implemented with convolutional neural networks and optimized by applying an EM algorithm during training. Experimental results reveal that our approach saves 89% annotation time and significantly narrows the performance gap between weak and full annotations. This highlights its potential for clinical applications.
Kai Zhou 0004, Zhang Yi 0001, Xiuyuan Xu
IEEE Trans. Medical Imaging3
2025 Prototype Bayesian Meta-Learning for Few-Shot Image Classification
abstract
Meta-learning aims to leverage prior knowledge from related tasks to enable a base learner to quickly adapt to new tasks with limited labeled samples. However, traditional meta-learning methods have limitations as they provide an optimal initialization for all new tasks, disregarding the inherent uncertainty induced by few-shot tasks and impeding task-specific self-adaptation initialization. In response to this challenge, this article proposes a novel probabilistic meta-learning approach called prototype Bayesian meta-learning (PBML). PBML focuses on meta-learning variational posteriors within a Bayesian framework, guided by prototype-conditioned prior information. Specifically, to capture model uncertainty, PBML treats both meta- and task-specific parameters as random variables and integrates their posterior estimates into hierarchical Bayesian modeling through variational inference (VI). During model inference, PBML employs Laplacian estimation to approximate the integral term over the likelihood loss, deriving a rigorous upper-bound for generalization errors. To enhance the model's expressiveness and enable task-specific adaptive initialization, PBML proposes a data-driven approach to model the task-specific variational posteriors. This is achieved by designing a generative model structure that incorporates prototype-conditioned task-dependent priors into the random generation of task-specific variational posteriors. Additionally, by performing latent embedding optimization, PBML decouples the gradient-based meta-learning from the high-dimensional variational parameter space. Experimental results on benchmark datasets for few-shot image classification illustrate that PBML attains state-of-the-art or competitive performance when compared to other related works. Versatility studies demonstrate the adaptability and applicability of PBML in addressing diverse and challenging few-shot tasks. Furthermore, ablation studies validate the performance gains attributed to the inference and model components.
Meijun Fu, Jun Wang 0002, Zhang Yi 0001
IEEE Trans. Neural Networks Learn. Syst.4
2025 Coexistence of Cyclic Sequential Pattern Recognition and Associative Memory in Neural Networks by Attractor Mechanisms
abstract
Neural networks are developed to model the behavior of the brain. One crucial question in this field pertains to when and how a neural network can memorize a given set of patterns. There are two mechanisms to store information: associative memory and sequential pattern recognition. In the case of associative memory, the neural network operates with dynamical attractors that are point attractors, each corresponding to one of the patterns to be stored within the network. In contrast, sequential pattern recognition involves the network memorizing a set of patterns and subsequently retrieving them in a specific order over time. From a dynamical perspective, this corresponds to the presence of a continuous attractor or a cyclic attractor composed of the sequence of patterns stored within the network in a given order. Evidence suggests that the brain is capable of simultaneously performing both associative memory and sequential pattern recognition. Therefore, these types of attractors coexist within the neural network, signifying that some patterns are stored as point attractors, while others are stored as continuous or cyclic attractors. This article investigates the coexistence of cyclic attractors and continuous or point attractors in certain nonlinear neural networks, enabling the simultaneous emergence of various memory mechanisms. By selectively grouping neurons, conditions are established for the existence of cyclic attractors, continuous attractors, and point attractors, respectively. Furthermore, each attractor is explicitly represented, and a competitive dynamic emerges among these coexisting attractors, primarily regulated by adjustments to external inputs.
Jingyang Huo, Zhang Yi 0001, Jinsong Leng
IEEE Trans. Neural Networks Learn. Syst.4
2025 Q-ADER: An Effective Q-Learning for Recommendation With Diminishing Action Space
abstract
Deep reinforcement learning (RL) has been widely applied to personalized recommender systems (PRSs) as they can capture user preferences progressively. Among RL-based techniques, deep Q-network (DQN) stands out as the most popular choice due to its simple update strategy and superior performance. Typically, many recommendation scenarios are accompanied by the diminishing action space setting, where the available action space will gradually decrease to avoid recommending duplicate items. However, existing DQN-based recommender systems inherently grapple with a discrepancy between the fixed full action space inherent in the Q-network and the diminishing available action space during recommendation. This article elucidates how this discrepancy induces an issue termed action diminishing error in the vanilla temporal difference (TD) operator. Due to this discrepancy, standard DQN methods prove impractical for learning accurate value estimates, rendering them ineffective in the context of diminishing action space. To mitigate this issue, we propose the Q-learning-based action diminishing error reduction (Q-ADER) algorithm to modify the value estimate error at each step. In practice, Q-ADER augments the standard TD learning with an error reduction term which is straightforward to implement on top of the existing DQN algorithms. Experiments are conducted on four real-world datasets to verify the effectiveness of our proposed algorithm.
Hong Qu 0002, Mingsheng Fu, Wenyu Chen 0001, Zhang Yi 0001
IEEE Trans. Neural Networks Learn. Syst.6
2025 Retina-Inspired Lightweight Spiking Convolutional Neural Network for Single-Image Dehazing
abstract
Suspended particles in hazy medium absorb and scatter light, severely degrading imaging quality. Numerous single-image dehazing methods have been proposed to reconstruct clear images from hazy ones. However, most of them focus on increasing depth and width to improve dehazing performance, which incurs high computation and energy costs. To address this issue, we propose a lightweight spiking convolutional neural network (CNN) referred to as retina-inspired spiking CNN (RI-SCNN) for the reconstruction of hazy images. Unlike conventional dehazing techniques, first, our proposed network simulates the hierarchical structure and cellular function of the retina and devises five network modules to efficiently encode and extract image features through ON and OFF roads. Furthermore, the linear reconstruction mechanism is introduced to integrate the outputs from different roads, adaptively preserving regions with optimal details and constructing a comprehensive visual representation. Finally, by the transformed atmospheric scattering formula, our network can generate the dehazy image. Incorporating the microscale spiking mechanism of the brain, the entire network leverages discrete binary spike trains for information encoding and transmission, directly trained by spiking surrogate gradient learning on integrate-and-fire (IF) neurons. Experimental results demonstrate the superiority of the proposed RI-SCNN in terms of quantitative dehazing performance, qualitative visual effect, energy efficiency, and run speed. Considering its lightweight architecture with ultralow computation and energy costs, the network is encouraged to be deployed in the visual sensor hardware to improve overall performance.
Xiaoling Luo 0001, Qian Sun 0014, Hong Qu 0002, Zhang Yi 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 A Lightweight U-like Network Utilizing Neural Memory Ordinary Differential Equations for Slimming the Decoder
Quansong He, Zhang Yi 0001, Tao He 0016
IJCAI4
2024 Strengthening Layer Interaction via Dynamic Layer Attention
Kaishen Wang, Xun Xia, Jian Liu 0041, Zhang Yi 0001, Tao He 0016
IJCAI4
2024 IarCAC: Instance-Aware Representation for Coronary Artery Calcification Segmentation in Cardiac CT Angiography
Weili Jiang, Zhang Yi 0001, Jianyong Wang 0002, Mao Chen 0008
MICCAI (1)3
2024 Efficient and Gender-Adaptive Graph Vision Mamba for Pediatric Bone Age Assessment
Lingyu Zhou, Zhang Yi 0001, Kai Zhou 0004, Xiuyuan Xu
MICCAI (5)2
2024 Ori-Net: Orientation-guided Neural Network for Automated Coronary Arteries Segmentation
Weili Jiang, Yuheng Jia, Zhang Yi 0001, Mao Chen 0008, Jianyong Wang 0002
Expert Syst. Appl.5
2024 Automated Quality Assessment of Medical Images in Echocardiography Using Neural Networks with Adaptive Ranking and Structure-Aware Learning
abstract
The quality of medical images is crucial for accurately diagnosing and treating various diseases. However, current automated methods for assessing image quality are based on neural networks, which often focus solely on pixel distortion and overlook the significance of complex structures within the images. This study introduces a novel neural network model designed explicitly for automated image quality assessment that addresses pixel and semantic distortion. The model introduces an adaptive ranking mechanism enhanced with contrast sensitivity weighting to refine the detection of minor variances in similar images for pixel distortion assessment. More significantly, the model integrates a structure-aware learning module employing graph neural networks. This module is adept at deciphering the intricate relationships between an image's semantic structure and quality. When evaluated on two ultrasound imaging datasets, the proposed method outshines existing leading models in performance. Additionally, it boasts seamless integration into clinical workflows, enabling real-time image quality assessment, crucial for precise disease diagnosis and treatment.
Gadeng Luosang, Jian Liu 0041, Fanxin Zeng, Zhang Yi 0001, Jianyong Wang 0002
Int. J. Neural Syst.5
2024 A Bidirectional Feedforward Neural Network Architecture Using the Discretized Neural Memory Ordinary Differential Equation
abstract
Deep Feedforward Neural Networks (FNNs) with skip connections have revolutionized various image recognition tasks. In this paper, we propose a novel architecture called bidirectional FNN (BiFNN), which utilizes skip connections to aggregate features between its forward and backward paths. The BiFNN accepts any FNN as a plugin that can incorporate any general FNN model into its forward path, introducing only a few additional parameters in the cross-path connections. The backward path is implemented as a nonparameter layer, utilizing a discretized form of the neural memory Ordinary Differential Equation (nmODE), which is named [Formula: see text]-net. We provide a proof of convergence for the [Formula: see text]-net and evaluate its initial value problem. Our proposed architecture is evaluated on diverse image recognition datasets, including Fashion-MNIST, SVHN, CIFAR-10, CIFAR-100, and Tiny-ImageNet. The results demonstrate that BiFNNs offer significant improvements compared to embedded models such as ConvMixer, ResNet, ResNeXt, and Vision Transformer. Furthermore, BiFNNs can be fine-tuned to achieve comparable performance with embedded models on Tiny-ImageNet and ImageNet-1K datasets by loading the same pretrained parameters.
Zhang Yi 0001, Tao He 0016
Int. J. Neural Syst.2
2024 A Forward Learning Algorithm for Neural Memory Ordinary Differential Equations
abstract
The deep neural network, based on the backpropagation learning algorithm, has achieved tremendous success. However, the backpropagation algorithm is consistently considered biologically implausible. Many efforts have recently been made to address these biological implausibility issues, nevertheless, these methods are tailored to discrete neural network structures. Continuous neural networks are crucial for investigating novel neural network models with more biologically dynamic characteristics and for interpretability of large language models. The neural memory ordinary differential equation (nmODE) is a recently proposed continuous neural network model that exhibits several intriguing properties. In this study, we present a forward-learning algorithm, called nmForwardLA, for nmODE. This algorithm boasts lower computational dimensions and greater efficiency. Compared with the other learning algorithms, experimental results on MNIST, CIFAR10, and CIFAR100 demonstrate its potency.
Xiuyuan Xu, Haiying Luo, Zhang Yi 0001, Haixian Zhang
Int. J. Neural Syst.3
2024 Precise Localization for Anatomo-Physiological Hallmarks of the Cervical Spine by Using Neural Memory Ordinary Differential Equation
abstract
In the evaluation of cervical spine disorders, precise positioning of anatomo-physiological hallmarks is fundamental for calculating diverse measurement metrics. Despite the fact that deep learning has achieved impressive results in the field of keypoint localization, there are still many limitations when facing medical image. First, these methods often encounter limitations when faced with the inherent variability in cervical spine datasets, arising from imaging factors. Second, predicting keypoints for only 4% of the entire X-ray image surface area poses a significant challenge. To tackle these issues, we propose a deep neural network architecture, NF-DEKR, specifically tailored for predicting keypoints in cervical spine physiological anatomy. Leveraging neural memory ordinary differential equation with its distinctive memory learning separation and convergence to a singular global attractor characteristic, our design effectively mitigates inherent data variability. Simultaneously, we introduce a Multi-Resolution Focus module to preprocess feature maps before entering the disentangled regression branch and the heatmap branch. Employing a differentiated strategy for feature maps of varying scales, this approach yields more accurate predictions of densely localized keypoints. We construct a medical dataset, SCUSpineXray, comprising X-ray images annotated by orthopedic specialists and conduct similar experiments on the publicly available UWSpineCT dataset. Experimental results demonstrate that compared to the baseline DEKR network, our proposed method enhances average precision by 2% to 3%, accompanied by a marginal increase in model parameters and the floating-point operations (FLOPs). The code (https://github.com/Zhxyi/NF-DEKR) is available.
Dehan Li, Yuexiong Xie, Zhang Yi 0001, Litai Ma, Lei Xu 0044
Int. J. Neural Syst.6
2024 Leveraging denoising diffusion probabilistic model to improve the multi-thickness CT segmentation
Chengrong Yu, Shengqian Zhu, Zhang Yi 0001, Junjie Hu 0004
Neurocomputing5
2024 ABSR: Progressive Alternate Refinement for Blind Cardiac MRI Super-Resolution
abstract
Deep learning-based methods for super-resolution (SR) reconstruction of cardiac magnetic resonance imaging (CMRI) have achieved commendable reconstruction performance owing to the potent learning capability of neural networks. Nonetheless, these methods suffer from performance degradation when handling real-world CMRI images, failing to reconstruct high-fidelity CMRI high-resolution images. This degradation stems from the fact that real-world CMRI images are afflicted with blur and noise, whereas mainstream deep learning-based CMRI SR algorithms are typically trained on images degraded using bicubic methods. To address this problem, we propose a progressive alternate refinement for blind CMRI SR, which implements blind SR reconstruction through progressive alternate refinement optimization of the CMRI image feature extraction process and blur kernel feature extraction process. Moreover, we propose a novel progressive blind SR reconstruction subnetwork, which utilizes the alternate residual attention block (ARAB) to perform deep feature extraction. Meanwhile, we propose an ARAB, which uses channel and pixel attention mechanisms to extract high-frequency features from the extracted CMRI and blur kernel features. Extensive experimental results demonstrate that our proposed alternate refinement for blind CMRI super-resolution outperforms the state-of-the-art SR methods, exhibiting superior reconstruction performance and the ability to reconstruct high-fidelity CMRI images.
Defu Qiu, Zhaoyang Song, Wenjun Zhang 0005, Kelvin K. L. Wong, Zhang Yi 0001
IEEE Internet Things J.5
2024 Multi-instance imbalance semantic segmentation by instance-dependent attention and adaptive hard instance mining
Weili Jiang, Zhang Yi 0001, Mao Chen 0008, Jianyong Wang 0002
Knowl. Based Syst.3
2024 Multiobjective Evolutionary Generative Adversarial Network Compression for Image Translation
abstract
Generative Adversarial Networks (GANs) have achieved remarkable success in image translation tasks. However, its prohibitive computational overhead has been a major hurdle for deployment on resource constrained platforms. Due to the training instability and the complicated network architecture, existing model acceleration techniques cannot appropriately handle the GAN compression problem. To cope with these difficulties, we propose a multi-objective evolutionary algorithm to compress GAN models in the context of image translation tasks, which is termed as MEGC. Particularly, the conflict between the computational cost of the GAN model and the quality of generated image is explicitly modeled as a two-objective optimization problem, and the evolved Pareto set is utilized to guide the sampling process during supernet training, which can in turn divert the focus of the supernet training to well-performing compact subnets. Besides, an evaluation-free strategy is introduced to facilitate exploration in the search space while incurring no extra computational cost. Based on the above design, the proposed MEGC eliminates the requirement of subnet searching in the post-processing procedure. Experiments on image translation tasks under paired and unparied settings demonstrate the effectiveness of the proposed MEGC on reducing the computational cost of GANs while improving the quality of generated images compared to those of the full models.
Yao Zhou 0002, Xianglei Yuan, Kaide Huang, Zhang Yi 0001, Gary G. Yen
IEEE Trans. Evol. Comput.5
2024 Biobjective Optimization Method for Large-Scale Group Decision Making Based on Hesitant Fuzzy Linguistic Preference Relations With Granularity Levels
abstract
Large-scale group decision making becomes increasingly common with the rapid development of society and the increasing complexity of practical problems. However, it is difficult to distinguish the semantic differences between the same linguistic term, and original linguistic term may not express flexible semantics, so that this will affect the precise of decision-making results. With the help of granular computing, this article adopts a new format of linguistic term, named as hesitant fuzzy linguistic term set with granularity level, to endow preference information with flexibility and at a specific granularity. Then, in this study, we propose a novel intelligent biobjective optimization method for large-scale group decision making, considering group consensus degree and group risk degree in the decision-making process, where group risk degree is measured from the motivation of portfolio risk. Differential evolution is used to handle biobjective optimization method to determine the optimal results. We also introduce an additive consistency measure and develop a method to derive the corresponding threshold values through Monte Carlo simulation. Finally, the case study and comparison results are covered to demonstrate the practicality and superiority of the proposed method. This work has some original points: 1) Hesitant fuzzy linguistic term set with granularity level brings flexibility to the decision-making process. 2) Group consensus degree and group risk degree are involved in biobjective optimization method, where the group risk degree is measured from the motivation of portfolio risk. 3) A novel additive consistency measure is proposed and different threshold values of preference relations in different dimensions are derived.
Yuanhang Zheng, Zeshui Xu, Witold Pedrycz, Zhang Yi 0001
IEEE Trans. Fuzzy Syst.5
2024 Adaptive Annotation Correlation Based Multi-Annotation Learning for Calibrated Medical Image Segmentation
abstract
Medical image segmentation is a fundamental task in many clinical applications, yet current automated segmentation methods rely heavily on manual annotations, which are inherently subjective and prone to annotation bias. Recently, modeling annotator preference has garnered great interest, and several methods have been proposed in the past two years. However, the existing methods completely ignore the potential correlation between annotations, such as complementary and discriminative information. In this work, the Adaptive annotation CorrelaTion based multI-annOtation LearNing (ACTION) method is proposed for calibrated medical image segmentation. ACTION employs consensus feature learning and dynamic adaptive weighting to leverage complementary information across annotations and emphasize discriminative information within each annotation based on their correlations, respectively. Meanwhile, memory accumulation-replay is proposed to accumulate the prior knowledge and integrate it into the model to enable the model to accommodate the multi-annotation setting. Two medical image benchmarks with different modalities are utilized to evaluate the performance of ACTION, and extensive experimental results demonstrate that it achieves superior performance compared to several state-of-the-art methods.
Lei Zhang 0005, Xin Shu 0005, Zizhou Wang, Zhang Yi 0001
IEEE J. Biomed. Health Informatics5
2024 ICNoduleNet: Enhancing Pulmonary Nodule Detection Performance on Sharp Kernel CT Imaging
abstract
Thoracic computed tomography (CT) currently plays the primary role in pulmonary nodule detection, where the reconstruction kernel significantly impacts performance in computer-aided pulmonary nodule detectors. The issue of kernel selection affecting performance has been overlooked in pulmonary nodule detection. This paper first introduces a novel pulmonary nodule detection dataset named Reconstruction Kernel Imaging for Pulmonary Nodule Detection (RKPN) for quantifying algorithm differences between the two imaging types. The dataset contains pairs of images taken from the same patient on the same date, featuring both smooth (B31f) and sharp kernel (B60f) reconstructions. All other imaging parameters and pulmonary nodule labels remain entirely consistent across these pairs. Extensive quantification reveals mainstream detectors perform better on smooth kernel imaging than on sharp kernel imaging. To address suboptimal detection on the sharp kernel imaging, we further propose an image conversion-based pulmonary nodule detector called ICNoduleNet. A lightweight 3D slice-channel converter (LSCC) module is introduced to convert sharp kernel images into smooth kernel images, which can sufficiently learn inter-slice and inter-channel feature information while avoiding introducing excessive parameters. We conduct thorough experiments that validate the effectiveness of ICNoduleNet, it takes sharp kernel images as input and can achieve comparable or even superior detection performance to the baseline that uses the smooth kernel images. The evaluation shows promising results and proves the effectiveness of ICNoduleNet.
Tianzhong Lan, Fanxin Zeng, Zhang Yi 0001, Xiuyuan Xu, Min Zhu 0005
IEEE J. Biomed. Health Informatics3
2024 Reciprocal Teacher-Student Learning via Forward and Feedback Knowledge Distillation
abstract
Knowledge distillation (KD) is a prevalent model compression technique in deep learning, aiming to leverage knowledge from a large teacher model to enhance the training of a smaller student model. It has found success in deploying compact deep models in intelligent applications like intelligent transportation, smart health, and distributed intelligence. Current knowledge distillation methods primarily fall into two categories: offline and online knowledge distillation. Offline methods involve a one-way distillation process, transferring unvaried knowledge from teacher to student, while online methods enable the simultaneous training of multiple peer students. However, existing knowledge distillation methods often face challenges where the student may not fully comprehend the teacher's knowledge due to model capacity gaps, and there might be knowledge incongruence among outputs of multiple students without teacher guidance. To address these issues, we propose a novel reciprocal teacher-student learning inspired by human teaching and examining through forward and feedback knowledge distillation (FFKD). Forward knowledge distillation operates offline, while feedback knowledge distillation follows an online scheme. The rationale is that feedback knowledge distillation enables the pre-trained teacher model to receive feedback from students, allowing the teacher to refine its teaching strategies accordingly. To achieve this, we introduce a new weighting constraint to gauge the extent of students' understanding of the teacher's knowledge, which is then utilized to enhance teaching strategies. Experimental results on five visual recognition datasets demonstrate that the proposed FFKD outperforms current state-of-the-art knowledge distillation methods.
Jianping Gou, Baosheng Yu, Jinhua Liu 0001, Lan Du 0002, Shaohua Wan 0001, Zhang Yi 0001
IEEE Trans. Multim.7
2024 Hierarchical Locality-Aware Deep Dictionary Learning for Classification
abstract
Deep dictionary learning (DDL) shows good performance in visual classification tasks. However, almost all existing DDL methods ignore the locality relationships between the input data representations and the learned dictionary atoms, and learn sub-optimal representations in the feature coding stage, which are less conducive to classification. To this end, we propose a hierarchical locality-aware deep dictionary learning (HILADLE) framework for classification, which can learn locality-constrained dictionaries at different abstract levels through hierarchical dictionary learning. The locality constraints play an important role in learning informative dictionary atoms while preserving the data structure in the original input feature space. Moreover, instead of using an identity activation function like existing DDL methods, we further boost the generalization performance of our HILADLE method with a ReLU activation function to deal with the overfitting issue caused by over-parameterization, inspired by its effectiveness in deep neural networks. Finally, the concatenation of all feature representations learned at different layers is used as input to the final classifier. We demonstrate, through an extensive set of experiments on several benchmark face recognition, image classification, and age estimation datasets, that our method is able to surpass several dictionary learning, deep dictionary learning and deep learning methods.
Jianping Gou, Xin He 0034, Lan Du 0002, Baosheng Yu, Zhang Yi 0001
IEEE Trans. Multim.6
2024 Reconstructed Graph Constrained Auto-Encoders for Multi-View Representation Learning
abstract
The application of Auto-Encoder (AE) to multi-view representation learning has gained traction due to advancements in deep learning. While some current AE-based multi-view representation learning algorithms incorporate the geometric structure of the input data into their feature representation learning process, their use of a shallow structured graph regularization term can be restrictive when used in conjunction with deep models. Furthermore, current multi-view representation learning algorithms do not fully utilize the diversity and consistency presented in different views, leading to a reduction in the efficacy of feature learning. This paper introduces a novel approach, reconstructed graph constrained auto-encoders (RGCAE), for multi-view representation learning. Unlike existing methods, our approach incorporates deep adaptive graph regularization based on multi-layer perceptron to ensure the preservation of the geometric similarity graph, which is constructed based on the local invariance principle. By decoupling the feature representation learning from the preservation of the geometric structure among different views, our approach can better leverage the diversity presented in multi-view data. We obtain view-specific representations that preserve the geometric structure and then combine them by averaging to obtain a common representation. To ensure the consistency of the multi-view data, we minimize the loss between the view-specific and common representations. Consequently, our RGCAE approach can maintain the geometric structure of multi-view data and is better suited for integration with deep models. Extensive experiments on six datasets demonstrate that RGCAE obtained promising performance, compared with the state-of-the-art methods.
Jianping Gou, Nannan Xie, Yun-Hao Yuan 0001, Lan Du 0002, Weihua Ou, Zhang Yi 0001
IEEE Trans. Multim.6
2024 Difference-Aware Distillation for Semantic Segmentation
abstract
In recent years, various distillation methods for semantic segmentation have been proposed. However, these methods typically train the student model to imitate the intermediate features or logits of the teacher model directly, thereby overlooking the high-discrepancy regions learned by both models, particularly the differences in instance edges. In this paper, we introduce a novel approach, called Difference-aware Distillation, to address this limitation. Our proposed method detects the discrepancies among the teacher model and the student model in the logit space through two masking mechanisms (i.e., masking by logit differences with respect to the ground truth labels and masking by differences in the predictive class probabilities), and guides the student model to restore the teacher's features with the focus on these highly-discrepant regions, resulting in improved segmentation performance. With the features jointly masked by these two mechanisms, the student model learns to preserve the teacher's features via a feature generation module, thus achieving better representation. Our experimental evaluation on three datasets, Cityscapes, Pascal2012, and ADE20 K, demonstrates our proposed approach outperforms several baselines considered. Further visualization analysis confirms that our method effectively directs the student model's attention to the discrepancies, such as the edges of small objects and the interiors of large objects.
Jianping Gou, Xiabin Zhou, Lan Du 0002, Yibing Zhan, Wu Chen 0005, Zhang Yi 0001
IEEE Trans. Multim.6
2024 DBSR: Quadratic Conditional Diffusion Model for Blind Cardiac MRI Super-Resolution
abstract
Cardiac magnetic resonance imaging (CMRI) can help experts quickly diagnose cardiovascular diseases. Due to the patient's breathing and slight movement during the magnetic resonance imaging scan, the obtained CMRI may be severely blurred, affecting the accuracy of clinical diagnosis. To address this issue, we propose the quadratic conditional diffusion model for blind CMRI super-resolution (DBSR). Specifically, we propose a conditional blur kernel noise predictor, which predicts the blur kernel from low-resolution images by the diffusion model, transforming the unknown blur kernel in low-resolution CMRI into a known one. Meanwhile, we design a novel conditional CMRI noise predictor, which uses the predicted blur kernel as prior knowledge to guide the diffusion model in reconstructing high-resolution CMRI. Furthermore, we propose a cascaded residual attention network feature extractor, which extracts feature information from CMRI low-resolution images for blur kernel prediction and SR reconstruction of CMRI images. Extensive experimental results indicate that our proposed DBSR achieves better blind super-resolution reconstruction results than several state-of-the-art baselines.
Defu Qiu, Yuhu Cheng 0001, Kelvin K. L. Wong, Wenjun Zhang 0005, Zhang Yi 0001, Xuesong Wang 0001
IEEE Trans. Multim.5
2024 Weighted Graph-Structured Semantics Constraint Network for Cross-Modal Retrieval
abstract
Cross-modal retrieval aims to retrieve relevant content of different modalities by giving a query of another modality. The biggest difficulty is how to bridge the heterogeneous gap between different modalities. The commonly-used methods tend to focus on exploiting individual image-text pair and mining the relations of cross-modality data thereof, but ignore the role of multi-sample correlation. Moreover, more global, structural inter-pair knowledge contained by the training dataset will be under-used. To fully exploit graph-structured semantics and mine the semantic information in the dataset for learning discriminative representations, we propose Weighted Graph-structured Semantics Constraint Network (WGSCN), a unified, graph-based, semantic-constrained learning framework, in which GCN is used to mine comprehensive relation information from cross modality data. Our main inspiration is to design a novel two-branch GCN-based Cross-modal Semantic Encoding (GCSE) module to produce semantic embeddings with the both modality-specific and modality-shared correlation. Moreover, a GAN-based dual learning approach is used to further improve the discriminability and model the joint distribution across different modalities. Our proposed GDL uses semantic embeddings as supervisory signal to make the common representation semantically discriminative while adversarial learning and dual learning are used to make the common representation modality-invariant. Through comparative experiments on five commonly used cross-modal datasets, we have shown the superior retrieval accuracy of our WGSCN.
Lei Zhang 0005, Leiting Chen, Chuan Zhou 0004, Xin Li 0079, Fan Yang 0054, Zhang Yi 0001
IEEE Trans. Multim.6
2024 Improving Exploration in Actor-Critic With Weakly Pessimistic Value Estimation and Optimistic Policy Optimization
abstract
Deep off-policy actor-critic algorithms have been successfully applied to challenging tasks in continuous control. However, these methods typically suffer from the poor sample efficiency problem, limiting their widespread adoption in real-world domains. To mitigate this issue, we propose a novel actor-critic algorithm with weakly pessimistic value estimation and optimistic policy optimization (WPVOP) for continuous control. WPVOP integrates two key ingredients: 1) a weakly pessimistic value estimation, which compensates the pessimism of lower confidence bound in conventional value function (i.e., clipped double Q -learning) to trigger exploration in low-value state-action regions and 2) an optimistic policy optimization algorithm by sampling actions that could benefit the policy learning most toward optimal Q -values for efficient exploration. We theoretically analyze that the proposed weakly pessimistic value estimation method is lower and upper bounded, and empirically show that it could avoid extremely over-optimistic value estimates. We show that these two ideas are largely complementary, and can be fruitfully integrated to improve performance and promote sample efficiency of exploration. We evaluate WPVOP on the suite of continuous control tasks from MuJoCo, achieving state-of-the-art sample efficiency and performance.
Mingsheng Fu, Wenyu Chen 0001, Fan Zhang 0068, Haixian Zhang, Hong Qu 0002, Zhang Yi 0001
IEEE Trans. Neural Networks Learn. Syst.7
2024 Minicolumn-Based Episodic Memory Model With Spiking Neurons, Dendrites and Delays
abstract
Episodic memory is fundamental to the brain's cognitive function, but how neuronal activity is temporally organized during its encoding and retrieval is still unknown. In this article, combining hippocampus structure with a spiking neural network (SNN), a new bionic spiking temporal memory (BSTM) model is proposed to explore the encoding, formation, and retrieval of episodic memory. For encoding episodic memory, the spike-timing-dependent-plasticity (STDP) learning algorithm and a proposed minicolumn selection algorithm are used to encode each input item into several active minicolumns. For the formation of episodic memory, a sequential memory algorithm is proposed to store the contexts between items. For retrieval of episodic memory, the local retrieval algorithm and the global retrieval algorithm are proposed to retrieve sequence information, achieving multisentence prediction and multitime step prediction. All functions of BSTM are based on bionic spiking neurons, which have biological characteristics including columnar and dendritic structures, firing and receiving spikes, and delaying transmission. To test the performance of the BSTM model, the Children's Book Test (CBT) data set was used to conduct a series of experiments under different settings, including changing the number of minicolumns, neurons and sequences, modifying sequence items, etc. Compared to other sequence memory algorithms, the experimental results show that the proposed BSTM achieves higher accuracy and better robustness.
Yi Chen 0034, Jilun Zhang, Xiaoling Luo 0001, Malu Zhang, Hong Qu 0002, Zhang Yi 0001
IEEE Trans. Neural Networks Learn. Syst.7
2023 MemGCN: memory-augmented graph neural network for predict conduction disturbance after transcatheter aortic valve replacement
Gadeng Luosang, Yuheng Jia, Jianyong Wang 0002, Mao Chen 0008, Zhang Yi 0001
Appl. Intell.7
2023 GFF-Net: Graph-based feature fusion network for diagnosing plus disease in retinopathy of prematurity
Kaide Huang, Yuanyuan Chen 0006, Jie Zhong 0004, Zhang Yi 0001
Appl. Intell.6
2023 Enhancing Robustness of Medical Image Segmentation Model with Neural Memory Ordinary Differential Equation
abstract
Deep neural networks (DNNs) have emerged as a prominent model in medical image segmentation, achieving remarkable advancements in clinical practice. Despite the promising results reported in the literature, the effectiveness of DNNs necessitates substantial quantities of high-quality annotated training data. During experiments, we observe a significant decline in the performance of DNNs on the test set when there exists disruption in the labels of the training dataset, revealing inherent limitations in the robustness of DNNs. In this paper, we find that the neural memory ordinary differential equation (nmODE), a recently proposed model based on ordinary differential equations (ODEs), not only addresses the robustness limitation but also enhances performance when trained by the clean training dataset. However, it is acknowledged that the ODE-based model tends to be less computationally efficient compared to the conventional discrete models due to the multiple function evaluations required by the ODE solver. Recognizing the efficiency limitation of the ODE-based model, we propose a novel approach called the nmODE-based knowledge distillation (nmODE-KD). The proposed method aims to transfer knowledge from the continuous nmODE to a discrete layer, simultaneously enhancing the model's robustness and efficiency. The core concept of nmODE-KD revolves around enforcing the discrete layer to mimic the continuous nmODE by minimizing the KL divergence between them. Experimental results on 18 organs-at-risk segmentation tasks demonstrate that nmODE-KD exhibits improved robustness compared to ODE-based models while also mitigating the efficiency limitation.
Junjie Hu 0004, Chengrong Yu, Zhang Yi 0001, Haixian Zhang
Int. J. Neural Syst.3
2023 Weight matrix as a switch between line attractor and plane attractor of ring neural networks
Wenshuang Chen, Jinsong Leng, Zhang Yi 0001
Neurocomputing5
2023 Cascade-refine model for cephalometric landmark detection in high-resolution orthodontic images
Tao He 0016, Jixiang Guo, Fanxin Zeng, Zhang Yi 0001
Knowl. Based Syst.7
2023 Multilayer perceptron neural network with regression and ranking loss for patient-specific quality assurance
Wenjie Liu 0010, Lei Zhang 0005, Lizhang Xie, Guangjun Li, Sen Bai, Zhang Yi 0001
Knowl. Based Syst.7
2023 Fine-grained recognition: Multi-granularity labels and category similarity matrix
Xin Shu 0005, Lei Zhang 0005, Zizhou Wang, Lituan Wang, Zhang Yi 0001
Knowl. Based Syst.5
2023 Intra-class consistency and inter-class discrimination feature learning for automatic skin lesion classification
Lituan Wang, Lei Zhang 0005, Xin Shu 0005, Zhang Yi 0001
Medical Image Anal.4
2023 Discriminative and Geometry-Preserving Adaptive Graph Embedding for dimensionality reduction
Jianping Gou, Xia Yuan, Ya Xue, Lan Du 0002, Shuyin Xia, Zhang Yi 0001
Neural Networks7
2023 A Feature Space-Restricted Attention Attack on Medical Deep Learning Systems
abstract
Deep neural network has shown a powerful performance in the medical image analysis of a variety of diseases. However, a number of studies over the past few years have demonstrated that these deep learning systems can be vulnerable to well-designed adversarial attacks, with minor disruptions added to the input. Since both the public and academia have focused on deep learning in the health information economy, these adversarial attacks would prove more important and raise security concerns. In this article, adversarial attacks on deep learning systems in medicine are analyzed from two different points of view: 1) white box and 2) black box. A fast adversarial sample generation method, Feature Space-Restricted Attention Attack is proposed to explore more confusing adversarial samples. It is based on a generative adversarial network with bound classification space to generate perturbations to achieve attacks. Meanwhile, it can employ an attention mechanism to focus this perturbation on the lesion region. This enables the perturbation closely associated with the classification information making the attack more efficient and invisible. The performance and specificity of the proposed attack method are demonstrated by conducting extensive experiments on three different types of medical images. Finally, it is expected that this work can assist practitioners become being of current weaknesses in the deployment of deep learning systems in clinical settings. And, it further investigates domain-specific features of medical deep learning systems to enhance model generalization and resistance to attacks.
Zizhou Wang, Xin Shu 0005, Yan Wang 0015, Yangqin Feng, Lei Zhang 0005, Zhang Yi 0001
IEEE Trans. Cybern.6
2023 Multilevel Attention-Based Sample Correlations for Knowledge Distillation
abstract
Recently, model compression has been widely used for the deployment of cumbersome deep models on resource-limited edge devices in the performance-demanding industrial Internet of Things (IoT) scenarios. As a simple yet effective model compression technique, knowledge distillation (KD) aims to transfer the knowledge (e.g., sample relationships as the relational knowledge) from a large teacher model to a small student model. However, existing relational KD methods usually build sample correlations directly from the feature maps at a certain middle layer in deep neural networks, which tends to overfit the feature maps of the teacher model and fails to address the most important sample regions. Inspired by this, we argue that the characteristics of important regions are of great importance, and thus, introduce attention maps to construct sample correlations for knowledge distillation. Specifically, with attention maps from multiple middle layers, attention-based sample correlations are newly built upon the most informative sample regions, and can be used as an effective and novel relational knowledge for knowledge distillation. We refer to the proposed method as multilevel attention-based sample correlations for knowledge distillation (or MASCKD). We perform extensive experiments on popular KD datasets for image classification, image retrieval, and person reidentification, where the experimental results demonstrate the effectiveness of the proposed method for relational KD.
Jianping Gou, Liyuan Sun 0005, Baosheng Yu, Shaohua Wan 0001, Weihua Ou, Zhang Yi 0001
IEEE Trans. Ind. Informatics6
2023 Spark Rough Hypercuboid Approach for Scalable Feature Selection
abstract
Feature selection refers to choose an optimal non-redundant feature subset with minimal degradation of learning performance and maximal avoidance of data overfitting. The appearance of large data explosion leads to the sequential execution of algorithms are extremely time-consuming, which necessitates the scalable parallelization of algorithms by efficiently exploiting the distributed computational capabilities. In this paper, we present parallel feature selection algorithms underpinned by a rough hypercuboid approach in order to scale for the growing data volumes. Metrics in terms of rough hypercuboid are highly suitable to parallel distributed processing, and fits well with the Apache Spark cluster computing paradigm. Two data parallelism strategies, namely, vertical partitioning and horizontal partitioning, are implemented respectively to decompose the data into concurrent iterative computing streams. Experimental results on representative datasets show that our algorithms significantly faster than its original sequential counterpart while guaranteeing the quality of the results. Furthermore, the proposed algorithms are perfectly capable of exploiting the distributed-memory clusters to accomplish the computation task that fails on a single node due to the memory constraints. Parallel scalability and extensibility analysis have confirmed that our parallelization extends well to process massive amount of data and can scales well with the increase of computational nodes.
Chuan Luo 0001, Sizhao Wang, Tianrui Li 0001, Hongmei Chen 0001, Jiancheng Lv 0001, Zhang Yi 0001
IEEE Trans. Knowl. Data Eng.6
2023 Intra- and Inter-Class Induced Discriminative Deep Dictionary Learning for Visual Recognition
abstract
Deep dictionary learning (DDL) aims to learn dictionaries at different levels and the deepest level representations. However, existing DDL algorithms impose a$l_{1}$-norm constraint on the deepest level representations, ignoring the constraints on different level representations. Meanwhile, they fail to discover effectively the essential discrimination information. Therefore, the obtained representations are less discriminative, which degrades model performance. To tackle those issues, we propose an intra- and inter-class induced discriminative deep dictionary learning (DDDL). Specifically, both intra-class compactness and inter-class separability of layer-wise data representations are newly devised as two discriminative constraints on deep dictionary learning. In a hierarchical structure, we obtain a more informative dictionary and the class-specific representations are thus more discriminative at each layer. Due to the$l_{2}$-norm intra- and inter-class constraints of layer-wise data representation, we devise a layer-wise optimization strategy to efficiently learn the closed-form solution of the deepest representation for classification. Comprehensive experiments and analyses on several visual recognition tasks show that our DDDL model surpasses recent shallow and deep representation learning approaches.
Jianping Gou, Xia Yuan, Baosheng Yu, Zhang Yi 0001
IEEE Trans. Multim.5
2023 Supervised Learning in Multilayer Spiking Neural Networks With Spike Temporal Error Backpropagation
abstract
The brain-inspired spiking neural networks (SNNs) hold the advantages of lower power consumption and powerful computing capability. However, the lack of effective learning algorithms has obstructed the theoretical advance and applications of SNNs. The majority of the existing learning algorithms for SNNs are based on the synaptic weight adjustment. However, neuroscience findings confirm that synaptic delays can also be modulated to play an important role in the learning process. Here, we propose a gradient descent-based learning algorithm for synaptic delays to enhance the sequential learning performance of single spiking neuron. Moreover, we extend the proposed method to multilayer SNNs with spike temporal-based error backpropagation. In the proposed multilayer learning algorithm, information is encoded in the relative timing of individual neuronal spikes, and learning is performed based on the exact derivatives of the postsynaptic spike times with respect to presynaptic spike times. Experimental results on both synthetic and realistic datasets show significant improvements in learning efficiency and accuracy over the existing spike temporal-based learning algorithms. We also evaluate the proposed learning method in an SNN-based multimodal computational model for audiovisual pattern recognition, and it achieves better performance compared with its counterparts.
Xiaoling Luo 0001, Hong Qu 0002, Zhang Yi 0001, Jilun Zhang, Malu Zhang
IEEE Trans. Neural Networks Learn. Syst.4
2023 Large-Scale Meta-Heuristic Feature Selection Based on BPSO Assisted Rough Hypercuboid Approach
abstract
The selection of prominent features for building more compact and efficient models is an important data preprocessing task in the field of data mining. The rough hypercuboid approach is an emerging technique that can be applied to eliminate irrelevant and redundant features, especially for the inexactness problem in approximate numerical classification. By integrating the meta-heuristic-based evolutionary search technique, a novel global search method for numerical feature selection is proposed in this article based on the hybridization of the rough hypercuboid approach and binary particle swarm optimization (BPSO) algorithm, namely RH-BPSO. To further alleviate the issue of high computational cost when processing large-scale datasets, parallelization approaches for calculating the hybrid feature evaluation criteria are presented by decomposing and recombining hypercuboid equivalence partition matrix via horizontal data partitioning. A distributed meta-heuristic optimized rough hypercuboid feature selection (DiRH-BPSO) algorithm is thus developed and embedded in the Apache Spark cloud computing model. Extensive experimental results indicate that RH-BPSO is promising and can significantly outperform the other representative feature selection algorithms in terms of classification accuracy, the cardinality of the selected feature subset, and execution efficiency. Moreover, experiments on distributed-memory multicore clusters show that DiRH-BPSO is significantly faster than its sequential counterpart and is perfectly capable of completing large-scale feature selection tasks that fail on a single node due to memory constraints. Parallel scalability and extensibility analysis also demonstrate that DiRH-BPSO could scale out and extend well with the growth of computational nodes and the volume of data.
Chuan Luo 0001, Sizhao Wang, Tianrui Li 0001, Hongmei Chen 0001, Jiancheng Lv 0001, Zhang Yi 0001
IEEE Trans. Neural Networks Learn. Syst.6
2023 RegNet: Self-Regulated Network for Image Classification
abstract
The ResNet and its variants have achieved remarkable successes in various computer vision tasks. Despite its success in making gradient flow through building blocks, the information communication of intermediate layers of blocks is ignored. To address this issue, in this brief, we propose to introduce a regulator module as a memory mechanism to extract complementary features of the intermediate layers, which are further fed to the ResNet. In particular, the regulator module is composed of convolutional recurrent neural networks (RNNs) [e.g., convolutional long short-term memories (LSTMs) or convolutional gated recurrent units (GRUs)], which are shown to be good at extracting spatio-temporal information. We named the new regulated network as regulated residual network (RegNet). The regulator module can be easily implemented and appended to any ResNet architecture. Experimental results on three image classification datasets have demonstrated the promising performance of the proposed architecture compared with the standard ResNet, squeeze-and-excitation ResNet, and other state-of-the-art architectures.
Yu Pan 0005, Xinglin Pan, Steven C. H. Hoi, Zhang Yi 0001, Zenglin Xu
IEEE Trans. Neural Networks Learn. Syst.5
2023 RHDOFS: A Distributed Online Algorithm Towards Scalable Streaming Feature Selection
abstract
Feature selection is an important topic in data mining and machine learning, which aims to select an optimal feature subset for building effective and explainable prediction models. This article introduces Rough Hypercuboid based Distributed Online Feature Selection (RHDOFS) method to tackle two critical challenges of Volume and Velocity associated with Big Data. By exploring the class separability in the boundary region of rough hypercuboid approach, a novel integrated feature evaluation criterion is proposed by examining not only the explicit patterns contained in the positive region but also the useful implicit patterns derived from the boundary region. An efficient online feature selection method for streaming feature scenario is developed to identify relevant and nonredundant features in an incremental iterative fashion. Furthermore, a parallel optimization mechanism by combining both data and computational independence is further employed to accelerate the original sequential implementation. An efficient distributed online feature selection algorithm is presented and implemented on the Apache Spark platform to scale for massive amount of data by exploiting the computational capabilities of multicore clusters. Encouraging results of extensive experiments indicate the superiority and notable advantages of the proposed algorithm over the relevant and representative online feature selection algorithms. Empirical tests on scalability and extensibility also demonstrate our distributed implementation significantly reduces the computational times requirements while maintaining the prediction accuracy, and is capable of scaling well in volume of data and number of computing nodes.
Chuan Luo 0001, Sizhao Wang, Tianrui Li 0001, Hongmei Chen 0001, Jiancheng Lv 0001, Zhang Yi 0001
IEEE Trans. Parallel Distributed Syst.6
2022 Deep Dictionary Learning with an Intra-Class Constraint
abstract
In recent years, deep dictionary learning (DDL)has attracted a great amount of attention due to its effectiveness for represen-tation learning and visual recognition. However, most existing methods focus on unsupervised deep dictionary learning, failing to further explore the category information. To make full use of the category information of different samples, we pro-pose a novel deep dictionary learning model with an intra-class constraint (DDLIC) for visual classification. Specif-ically, we design the intra-class compactness constraint on the intermediate representation at different levels to encour-age the intra-class representations to be closer to each other, and eventually the learned representation becomes more dis-criminative. Unlike the traditional DDL methods, during the classification stage, our DDLIC performs a layer-wise greedy optimization in a similar way to the training stage. Experi-mental results on four image datasets show that our method is superior to the state-of-the-art methods.
Xia Yuan, Jianping Gou, Baosheng Yu, Zhang Yi 0001
ICME5
2022 Automated assessment of BI-RADS categories for ultrasound images using multi-scale neural networks with an order-constrained loss function
Yong Pi, Xiaofeng Qi, Dan Deng, Zhang Yi 0001
Appl. Intell.5
2022 Multi-view fusion segmentation for brain glioma on CT images
Han Wang 0025, Junjie Hu 0004, Lei Zhang 0005, Sen Bai, Zhang Yi 0001
Appl. Intell.6
2022 Metal artifact reduction for oral and maxillofacial computed tomography images by a generative adversarial network
Lei Xu 0044, Shanluo Zhou, Jixiang Guo, Zhang Yi 0001
Appl. Intell.6
2022 A medical question answering system using large language models and knowledge graphs
abstract
Question answering systems have become prominent in all areas, while in the medical domain it has been challenging because of the abundant domain knowledge. Retrieval based approach has become promising as large pretrained language models come forth. This study focuses on building a retrieval-based medical question answering system, tackling the challenge with large language models and knowledge extensions via graphs. We first retrieve an extensive but coarse set of answers via Elasticsearch efficiently. Then, we utilize semantic matching with pretrained language models to achieve a fine-grained ranking enhanced with named entity recognition and knowledge graphs to exploit the relation of the entities in question and answer. A new architecture based on siamese structures for answer selection is proposed. To evaluate the approach, we train and test the model on two Chinese data sets, NLPCC2017 and cMedQA. We also conduct experiments on two English data sets, TREC-QA and WikiQA. Our model achieves consistent improvement as compared to strong baselines on all data sets. Qualification studies with cMedQA and our in-house data set show that our system gains highly competitive performance. The proposed medical question answering system outperforms baseline models and systems in quantification and qualification evaluations.
Quan Guo, Zhang Yi 0001
Int. J. Intell. Syst.3
2022 An intelligent system for craniomaxillofacial defecting reconstruction
abstract
Craniomaxillofacial defects caused by congenital or acquired reasons seriously affect patients' physical and mental health. How to accurately and objectively repair the morphology of craniomaxillofacial tissues and organs through surgery is a difficult problem, and the preoperative virtual design is crucial. Traditional preoperative virtual design methods include mirror technology, statistical shape model, and deformable template. However, these methods are complex, time-consuming, and only applicable to some types of defects. Therefore, a general, intelligent, and personalized craniomaxillofacial defect virtual reconstruction system is desired. To solve this problem, a novel deep learning method, RecGAN, is proposed in this paper. RecGAN can learn the bone morphology of normal people, repair the defect intelligently based on the patient's remaining bone, and fully adapt to the special conditions of different patients. Currently, there are no open-source maxillofacial data sets available. Thus, a new maxillofacial computed tomography image data set with 500 simulated cases and 100 clinical cases is constructed to train and validate the method. The experimental results show that RecGAN can effectively restore the normal bone and tissue morphology of the patient's craniomaxillofacial defect area, solve the problem that there is no objective repair method for craniomaxillofacial defect, and achieve the best effect in the same type of research. The proposed intelligent craniomaxillofacial defect virtual reconstruction system based on RecGAN is expected to be applied in future clinical practice.
Lei Xu 0044, Yutao Xiong, Jixiang Guo, Kelvin K. L. Wong, Zhang Yi 0001
Int. J. Intell. Syst.6
2022 Deep Multimodal Neural Network Based on Data-Feature Fusion for Patient-Specific Quality Assurance
abstract
Patient-specific quality assurance (QA) for Volumetric Modulated Arc Therapy (VMAT) plans is routinely performed in the clinical. However, it is labor-intensive and time-consuming for medical physicists. QA prediction models can address these shortcomings and improve efficiency. Current approaches mainly focus on single cancer and single modality data. They are not applicable to clinical practice. To assess the accuracy of QA results for VMAT plans, this paper presents a new model that learns complementary features from the multi-modal data to predict the gamma passing rate (GPR). According to the characteristics of VMAT plans, a feature-data fusion approach is designed to fuse the features of imaging and non-imaging information in the model. In this study, 690 VMAT plans are collected encompassing more than ten diseases. The model can accurately predict the most VMAT plans at all three gamma criteria: 2%/2 mm, 3%/2 mm and 3%/3 mm. The mean absolute error between the predicted and measured GPR is 2.17%, 1.16% and 0.71%, respectively. The maximum deviation between the predicted and measured GPR is 3.46%, 4.6%, 8.56%, respectively. The proposed model is effective, and the features of the two modalities significantly influence QA results.
Lizhang Xie, Lei Zhang 0005, Guangjun Li, Zhang Yi 0001
Int. J. Neural Syst.5
2022 Computer-aided diagnosis of breast cancer in ultrasonography images by deep learning
Xiaofeng Qi, Fasheng Yi, Lei Zhang 0005, Yong Pi, Yuanyuan Chen 0006, Jixiang Guo, Jianyong Wang 0002, Quan Guo, Jilan Li, Yi Chen 0034, Zhang Yi 0001
Neurocomputing13
2022 VMAT dose prediction in radiotherapy by using progressive refinement UNet
Jianyong Wang 0002, Junjie Hu 0004, Xiaozhi Zhang, Sen Bai, Zhang Yi 0001
Neurocomputing7
2022 Fusing entropy measures for dynamic feature selection in incomplete approximation spaces
Chuan Luo 0001, Tianrui Li 0001, Hongmei Chen 0001, Jiancheng Lv 0001, Zhang Yi 0001
Knowl. Based Syst.5
2022 Fusing deep and handcrafted features for intelligent recognition of uptake patterns on thyroid scintigraphy
Yong Pi, Jianan Wei, Huawei Cai, Zhang Yi 0001
Knowl. Based Syst.6
2022 Cross-granularity multi-task network for ischemia diagnosis and defect detection in the myocardial perfusion imaging
Jianan Wei, Yong Pi, Huawei Cai, Lisha Jiang, Yongzhao Xiang, Zhang Yi 0001
Knowl. Based Syst.8
2022 Segmentation for regions of interest in radiotherapy by self-supervised learning
Chengrong Yu, Junjie Hu 0004, Guiyuan Li, Shengqian Zhu, Sen Bai, Zhang Yi 0001
Knowl. Based Syst.6
2022 Continuous Recurrent Neural Networks Based on Function Satlins
Jinsong Leng, Bisen Liu, Zhang Yi 0001
Neural Process. Lett.5
2022 Gram regularization for sparse and disentangled representation
Zhentao Gao, Yuanyuan Chen 0006, Quan Guo, Zhang Yi 0001
Pattern Anal. Appl.4
2022 Feature-Sensitive Deep Convolutional Neural Network for Multi-Instance Breast Cancer Detection
abstract
To obtain a well-performed computer-aided detection model for detecting breast cancer, it is usually needed to design an effective and efficient algorithm and a well-labeled dataset to train it. In this paper, first, a multi-instance mammography clinic dataset was constructed. Each case in the dataset includes a different number of instances captured from different views, it is labeled according to the pathological report, and all the instances of one case share one label. Nevertheless, the instances captured from different views may have various levels of contributions to conclude the category of the target case. Motivated by this observation, a feature-sensitive deep convolutional neural network with an end-to-end training manner is proposed to detect breast cancer. The proposed method first uses a pre-train model with some custom layers to extract image features. Then, it adopts a feature fusion module to learn to compute the weight of each feature vector. It makes the different instances of each case have different sensibility on the classifier. Lastly, a classifier module is used to classify the fused features. The experimental results on both our constructed clinic dataset and two public datasets have demonstrated the effectiveness of the proposed method.
Yan Wang 0015, Lei Zhang 0005, Xin Shu 0005, Yangqin Feng, Zhang Yi 0001
IEEE ACM Trans. Comput. Biol. Bioinform.5
2022 Automated Segmentation of the Clinical Target Volume in the Planning CT for Breast Cancer Using Deep Neural Networks
abstract
3-D radiotherapy is an effective treatment modality for breast cancer. In 3-D radiotherapy, delineation of the clinical target volume (CTV) is an essential step in the establishment of treatment plans. However, manual delineation is subjective and time consuming. In this study, we propose an automated segmentation model based on deep neural networks for the breast cancer CTV in planning computed tomography (CT). Our model is composed of three stages that work in a cascade manner, making it applicable to real-world scenarios. The first stage determines which slices contain CTVs, as not all CT slices include breast lesions. The second stage detects the region of the human body in an entire CT slice, eliminating boundary areas, which may have side effects for the segmentation of the CTV. The third stage delineates the CTV. To permit the network to focus on the breast mass in the slice, a novel dynamically strided convolution operation, which shows better performance than standard convolution, is proposed. To train and evaluate the model, a large dataset containing 455 cases and 50 425 CT slices is constructed. The proposed model achieves an average dice similarity coefficient (DSC) of 0.802 and 0.801 for right-0 and left-sided breast, respectively. Our method shows superior performance to that of previous state-of-the-art approaches.
Xiaofeng Qi, Junjie Hu 0004, Lei Zhang 0005, Sen Bai, Zhang Yi 0001
IEEE Trans. Cybern.5
2022 Analysis of Equilibria for a Class of Recurrent Neural Networks With Two Subnetworks
abstract
This article is concerned with the problem of the number and dynamical properties of equilibria for a class of connected recurrent networks with two switching subnetworks. In this network model, parameters serve as switches that allow two subnetworks to be turned ON or OFF among different dynamic states. The two subnetworks are described by a nonlinear coupled equation with a complicated relation among network parameters. Thus, the number and dynamical properties of equilibria have been very hard to investigate. By using Sturm's theorem, together with the geometrical properties of the network equation, we give a complete analysis of equilibria, including the existence, number, and dynamical properties. Necessary and sufficient conditions for the existence and exact number of equilibria are established. Moreover, the dynamical property of each equilibrium point is discussed without prior assumption of their locations. Finally, simulation examples are given to illustrate the theoretical results in this article.
Lingling Liu, Jacek M. Zurada, Zhang Yi 0001
IEEE Trans. Cybern.4
2022 Deep Neural Network With Structural Similarity Difference and Orientation-Based Loss for Position Error Classification in the Radiotherapy of Graves' Ophthalmopathy Patients
abstract
Identifying position errors for Graves' ophthalmopathy (GO) patients using electronic portal imaging device (EPID) transmission fluence maps is helpful in monitoring treatment. However, most of the existing models only extract features from dose difference maps computed from EPID images, which do not fully characterize all information of the positional errors. In addition, the position error has a three-dimensional spatial nature, which has never been explored in previous work. To address the above problems, a deep neural network (DNN) model with structural similarity difference and orientation-based loss is proposed in this paper, which consists of a feature extraction network and a feature enhancement network. To capture more information, three types of Structural SIMilarity (SSIM) sub-index maps are computed to enhance the luminance, contrast, and structural features of EPID images, respectively. These maps and the dose difference maps are fed into different networks to extract radiomic features. To acquire spatial features of the position errors, an orientation-based loss function is proposed for optimal training. It makes the data distribution more consistent with the realistic 3D space by integrating the error deviations of the predicted values in the left-right, superior-inferior, anterior-posterior directions. Experimental results on a constructed dataset demonstrate the effectiveness of the proposed model, compared with other related models and existing state-of-the-art methods.
Wenjie Liu 0010, Lei Zhang 0005, Guyu Dai, Xiangbin Zhang, Guangjun Li, Zhang Yi 0001
IEEE J. Biomed. Health Informatics6
2022 Hierarchical Graph Augmented Deep Collaborative Dictionary Learning for Classification
abstract
Recently, deep dictionary learning (DDL) has aroused attention due to its abilities of learning multiple different dictionaries and extracting multi-level abstract feature representations for samples. It has been applied to many intelligent recognition tasks, such as vehicle detection, traffic sign recognition and driver monitoring. Nevertheless, the off-the-shelf DDL-based methods ignore the essential structural information of data in multi-layer dictionary learning. The learned hierarchical data representations are less discriminative. To address this issue, we develop a new DDL framework, called the hierarchical graph augmented deep collaborative dictionary learning (HGDCDL). Firstly, we propose a new deep collaborative dictionary learning (DCDL) that applies collaborative representation to the deepest-level representation learning. Most importantly, equipped with a simple yet effective hierarchal graph construction mechanism, our HGDCDL uses the structure of data to regularize dictionary learning, and generates more informative dictionaries and discriminative representations at different levels. Extensive experiments show that our HGDCDL performs significantly better than the state-of-the-art shallow and deep representation learning methods for classification.
Jianping Gou, Xia Yuan, Lan Du 0002, Shuyin Xia, Zhang Yi 0001
IEEE Trans. Intell. Transp. Syst.5
2022 RECISTSup: Weakly-Supervised Lesion Volume Segmentation Using RECIST Measurement
abstract
Lesion volume segmentation in medical imaging is an effective tool for assessing lesion/tumor sizes and monitoring changes in growth. Since manually segmentation of lesion volume is not only time-consuming but also requires radiological experience, current practices rely on an imprecise surrogate called response evaluation criteria in solid tumors (RECIST). Although RECIST measurement is coarse compared with voxel-level annotation, it can reflect the lesion's location, length, and width, resulting in a possibility of segmenting lesion volume directly via RECIST measurement. In this study, a novel weakly-supervised method called RECISTSup is proposed to automatically segment lesion volume via RECIST measurement. Based on RECIST measurement, a new RECIST measurement propagation algorithm is proposed to generate pseudo masks, which are then used to train the segmentation networks. Due to the spatial prior knowledge provided by RECIST measurement, two new losses are also designed to make full use of it. In addition, the automatically segmented lesion results are used to supervise the model training iteratively for further improving segmentation performance. A series of experiments are carried out on three datasets to evaluate the proposed method, including ablation experiments, comparison of various methods, annotation cost analyses, visualization of results. Experimental results show that the proposed RECISTSup achieves the state-of-the-art result compared with other weakly-supervised methods. The results also demonstrate that RECIST measurement can produce similar performance to voxel-level annotation while significantly saving the annotation cost.
Han Wang 0025, Fasheng Yi, Jingling Wang, Zhang Yi 0001, Haixian Zhang
IEEE Trans. Medical Imaging4
2022 Subtraction Gates: Another Way to Learn Long-Term Dependencies in Recurrent Neural Networks
abstract
Recurrent neural networks (RNNs) can remember temporal contextual information over various time steps. The well-known gradient vanishing/explosion problem restricts the ability of RNNs to learn long-term dependencies. The gate mechanism is a well-developed method for learning long-term dependencies in long short-term memory (LSTM) models and their variants. These models usually take the multiplication terms as gates to control the input and output of RNNs during forwarding computation and to ensure a constant error flow during training. In this article, we propose the use of subtraction terms as another type of gates to learn long-term dependencies. Specifically, the multiplication gates are replaced by subtraction gates, and the activations of RNNs input and output are directly controlled by subtracting the subtrahend terms. The error flows remain constant, as the linear identity connection is retained during training. The proposed subtraction gates have more flexible options of internal activation functions than the multiplication gates of LSTM. The experimental results using the proposed Subtraction RNN (SRNN) indicate comparable performances to LSTM and gated recurrent unit in the Embedded Reber Grammar, Penn Tree Bank, and Pixel-by-Pixel MNIST experiments. To achieve these results, the SRNN requires approximate three-quarters of the parameters used by LSTM. We also show that a hybrid model combining multiplication forget gates and subtraction gates could achieve good performance.
Tao He 0016, Hua Mao 0001, Zhang Yi 0001
IEEE Trans. Neural Networks Learn. Syst.3
2022 Deep Attention-Based Imbalanced Image Classification
abstract
Class imbalance is a common problem in real-world image classification problems, some classes are with abundant data, and the other classes are not. In this case, the representations of classifiers are likely to be biased toward the majority classes and it is challenging to learn proper features, leading to unpromising performance. To eliminate this biased feature representation, many algorithm-level methods learn to pay more attention to the minority classes explicitly according to the prior knowledge of the data distribution. In this article, an attention-based approach called deep attention-based imbalanced image classification (DAIIC) is proposed to automatically pay more attention to the minority classes in a data-driven manner. In the proposed method, an attention network and a novel attention augmented logistic regression function are employed to encapsulate as many features, which belongs to the minority classes, as possible into the discriminative feature learning process by assigning the attention for different classes jointly in both the prediction and feature spaces. With the proposed object function, DAIIC can automatically learn the misclassification costs for different classes. Then, the learned misclassification costs can be used to guide the training process to learn more discriminative features using the designed attention networks. Furthermore, the proposed method is applicable to various types of networks and data sets. Experimental results on both single-label and multilabel imbalanced image classification data sets show that the proposed method has good generalizability and outperforms several state-of-the-art methods for imbalanced image classification.
Lituan Wang, Lei Zhang 0005, Xiaofeng Qi, Zhang Yi 0001
IEEE Trans. Neural Networks Learn. Syst.4
2022 Evolutionary Shallowing Deep Neural Networks at Block Levels
abstract
Neural networks have been demonstrated to be trainable even with hundreds of layers, which exhibit remarkable improvement on expressive power and provide significant performance gains in a variety of tasks. However, the prohibitive computational cost has become a severe challenge for deploying them on resource-constrained platforms. Meanwhile, widely adopted deep neural network architectures, for example, ResNets or DenseNets, are manually crafted on benchmark datasets, which hamper their generalization ability to other domains. To cope with these issues, we propose an evolutionary algorithm-based method for shallowing deep neural networks (DNNs) at block levels, which is termed as ESNB. Different from existing studies, ESNB utilizes the ensemble view of block-wise DNNs and employs the multiobjective optimization paradigm to reduce the number of blocks while avoiding performance degradation. It automatically discovers shallower network architectures by pruning less informative blocks, and employs knowledge distillation to recover the performance. Moreover, a novel prior knowledge incorporation strategy is proposed to improve the exploration ability of the evolutionary search process, and a correctness-aware knowledge distillation strategy is designed for better knowledge transferring. Experimental results show that the proposed method can effectively accelerate the inference of DNNs while achieving superior performance when compared with the state-of-the-art competing methods.
Yao Zhou 0002, Gary G. Yen, Zhang Yi 0001
IEEE Trans. Neural Networks Learn. Syst.3
2022 An Attention-Based Interactive Learning-to-Rank Model for Document Retrieval
abstract
The core issue of learning-to-rank (LTR) for document retrieval lies in finding an optimal ranking policy to meet the search intent of the user. The majority of proposed LTR approaches treat the ranking as a static process, employing a fixed ranking policy to immediately assign scores to documents. By contrast, ranking is not a static but an interactive process where the user continues interacting with the document retrieval system through information exchange such as search intent (e.g., rating or clicking for the retrieved items). We model the interactive ranking process (IRP), and propose an Attention-Based Interactive LTR model (AIRank) to constitute an intent-aware flexible ranking policy to gratify the user’s need. To enhance the ranking quality, the inherent relations among documents are procured by the self-attention method to contribute to an enriched user intent representation. Furthermore, we mend the policy gradient learning method to train the AIRank in the IRP. Experiments demonstrate the effectiveness of AIRank compared to the state-of-the-art methods in terms of normalized discounted cumulative gain and expected reciprocal rank.
Fan Zhang 0068, Wenyu Chen 0001, Mingsheng Fu, Hong Qu 0002, Zhang Yi 0001
IEEE Trans. Syst. Man Cybern. Syst.6
2021 A prior-based method for colorectal lymph node region classification via deep neural network
abstract
Colorectal cancer (CRC) is a common malignant tumor disease appeared in colon or rectum walls. CRC metastasis often appears with lymph nodes, and the CRC lymph nodes region classification is essential for CRC diagnosis, generally classified as lateral lymph nodes (LLN) and non-lateral lymph nodes (NLLN). Previous CRC diagnosis relied heavily on the physician’s clinical experience, which is a manual and time-consuming process. An automated method based on prior is proposed to CRC lymph node region classification using Convolutional neural networks (CNNs). Two novel priors are proposed, including spatial prior and shape prior. The spatial prior is based on medical domain knowledge to relieve the difficulty of extracting useful features from the complex semantic information of CT images. And the shape prior is proposed through carefully analyzing the dataset, which aims to find an optimal size that can preserve features in the origin CT images and be adaptive to neural network input. Experimental results demonstrate that the proposed method achieves impressive classification performance, in terms of an accuracy of 96.66% and an AUC of 0.9941. Additionally, we apply the proposed method in other medical classification works and it also achieves satisfying results.
Yueyao Huang, Han Wang 0025, Mingtian Wei, Jingling Wang, Haixian Zhang, Zhang Yi 0001
BIBM8
2021 DeepUWF-plus: automatic fundus identification and diagnosis system based on ultrawide-field fundus imaging
Yan Dai 0006, Yuanyuan Chen 0006, Jie Zhong 0004, Zhang Yi 0001
Appl. Intell.6
2021 Class mean-weighted discriminative collaborative representation for classification
abstract
Representation-based classification (RBC) has been attracting a great deal of attention in pattern recognition. As a typical extension to RBC, collaborative representation-based classification (CRC) has demonstrated its superior performance in various image classification tasks. Ideally, we expect that the learned class-specific representations for a testing sample are discriminative, and the representation computed for the true class dominates the final representation of the testing sample. Most existing CRC-based methods can learn pattern discrimination, but cannot differentiate the contribution of class-specific representations to the classification of each testing sample. It is challenging for a representation-based classifier to retain both properties. To address this challenge and further improve CRC's classification performance, we propose a novel CRC-based method, class mean-weighted discriminative collaborative representation-based classifier (CMW-DCRC). Its objective function penalises the standard l 2 -norm residuals with two discriminative regularisation terms. A decorrelating term makes the class-specific representations more discriminative, and a newly designed class mean-weighted term that promotes the training samples from individual classes to competitively reconstruct the testing sample while boosting the contribution of the true class. To further enhance the robustness of CRC, we extend CMW-DCRC by replacing the l2-norm coding residual with a l1-norm coding residual, and solve the optimisation problem with an iteratively reweighted least square algorithm. Extensive experimental results on nine image data sets have shown that our methods outperform the state-of-the-art RBC-based methods.
Jianping Gou, Lan Du 0002, Shaoning Zeng, Yongzhao Zhan 0001, Zhang Yi 0001
Int. J. Intell. Syst.6
2021 An intelligent system of pelvic lymph node detection
abstract
Computed tomography (CT) scanning is a fast and painless procedure that can capture clear imaging information beneath the abdomen and is widely used to help diagnose and monitor disease progress. The pelvic lymph node is a key indicator of colorectal cancer metastasis. In the traditional process, an experienced radiologist must read all the CT scanning images slice by slice to track the lymph nodes for future diagnosis. However, this process is time-consuming, exhausting, and subjective due to the complex pelvic structure, numerous blood vessels, and small lymph nodes. Therefore, automated methods are desirable to make this process easier. Currently, the available open-source CTLNDataset only contains large lymph nodes. Consequently, a new data set called PLNDataset, which is dedicated to lymph nodes within the pelvis, is constructed to solve this issue. A two-level annotation calibration method is proposed to guarantee the quality and correctness of pelvic lymph node annotation. Moreover, a novel system composed of a keyframe localization network and a lymph node detection network is proposed to detect pelvic lymph nodes in CT scanning images. The proposed method makes full use of two kinds of prior knowledge: spatial prior knowledge for keyframe localization and anchor prior knowledge for lymph node detection. A series of experiments are carried out to evaluate the proposed method, including ablation experiments, comparing other state-of-the-art methods, and visualization of results. The experimental results demonstrate that our proposed method outperforms other methods on PLNDataset and CTLNDataset. This system is expected to be applied in future clinical practice.
Han Wang 0025, Jingling Wang, Mingtian Wei, Zhang Yi 0001, Haixian Zhang
Int. J. Intell. Syst.5
2021 Adaptive sparse dropout: Learning the certainty and uncertainty in deep neural networks
Yuanyuan Chen 0006, Zhang Yi 0001
Neurocomputing2
2021 The effect of coefficients on the continuous attractors in coupled Highway Neural Networks
Wenshuang Chen, Zhang Yi 0001, Hong Qu 0002
Neurocomputing3
2021 Cephalometric landmark detection by considering translational invariance in the two-stage framework
Tao He 0016, Zhang Yi 0001, Jixiang Guo
Neurocomputing4
2021 Multi-scale attention U-net for segmenting clinical target volume in graves' ophthalmopathy
Junjie Hu 0004, Lei Zhang 0005, Sen Bai, Zhang Yi 0001
Neurocomputing5
2021 A multi-instance networks with multiple views for classification of mammograms
Lei Zhang 0005, Lizhang Xie, Zhang Yi 0001
Neurocomputing4
2021 Coexistence of continuous attractors with different dimensions for neural networks
Wanyu Xiang, Zhang Yi 0001
Neurocomputing3
2021 Missed diagnoses detection by adversarial learning
Xiaofeng Qi, Junjie Hu 0004, Zhang Yi 0001
Knowl. Based Syst.3
2021 Incorporating historical sub-optimal deep neural networks for dose prediction in radiotherapy
Junjie Hu 0004, Sen Bai, Zhang Yi 0001
Medical Image Anal.5
2021 Deep adversarial domain adaptation for breast cancer screening from mammograms
Yan Wang 0015, Yangqin Feng, Lei Zhang 0005, Zizhou Wang, Zhang Yi 0001
Medical Image Anal.6
2021 Learning Manifold Structures With Subspace Segmentations
abstract
Manifold learning has been widely used for dimensionality reduction and feature extraction of data recently. However, in the application of the related algorithms, it often suffers from noisy or unreliable data problems. For example, when the sample data have complex background, occlusions, and/or illuminations, the clustering of data is still a challenging task. To address these issues, we propose a family of novel algorithms for manifold regularized non-negative matrix factorization in this paper. In the algorithms, based on the alpha-beta-divergences, graph regularization with multiple segments is utilized to constrain the data transitivity in data decomposition. By adjusting two tuning parameters, we show that the proposed algorithms can significantly improve the robustness with respect to the images with complex background. The efficiency of the proposed algorithms is confirmed by the experiments on four different datasets. For different initializations and datasets, variations of cost functions and decomposition data elements in the learning are presented to show the convergent properties of the algorithms.
Shangming Yang, Lei Zhang 0005, Xiaofei He 0001, Zhang Yi 0001
IEEE Trans. Cybern.4
2021 A Knee-Guided Evolutionary Algorithm for Compressing Deep Neural Networks
abstract
Deep neural networks (DNNs) have been regarded as fundamental tools for many disciplines. Meanwhile, they are known for their large-scale parameters, high redundancy in weights, and extensive computing resource consumptions, which pose a tremendous challenge to the deployment in real-time applications or on resource-constrained devices. To cope with this issue, compressing DNNs for accelerating its inference has drawn extensive interest recently. The basic idea is to prune parameters with little performance degradation. However, the overparameterized nature and the conflict between parameters reduction and performance maintenance make it prohibitive to manually search the pruning parameter space. In this paper, we formally establish filter pruning as a multiobjective optimization problem, and propose a knee-guided evolutionary algorithm (KGEA) that can automatically search for the solution with quality tradeoff between the scale of parameters and performance, in which both conflicting objectives can be optimized simultaneously. In particular, by incorporating a minimum Manhattan distance approach, the search effort in the proposed KGEA is explicitly guided toward the knee area, which greatly facilitates the manual search for a good tradeoff solution. Moreover, the parameter importance is directly estimated on the criterion of performance loss, which can robustly identify the redundancy. In addition to the knee solution, a performance-improved model can also be found in a fine-tuning-free fashion. The experiments on compressing fully convolutional LeNet and VGG-19 networks validate the superiority of the proposed algorithm over the state-of-the-art competing methods.
Yao Zhou 0002, Gary G. Yen, Zhang Yi 0001
IEEE Trans. Cybern.3
2021 DeepUWF: An Automated Ultra-Wide-Field Fundus Screening System via Deep Learning
abstract
The emerging ultra-wide field of view (UWF) fundus color imaging is a powerful tool for fundus screening. However, manual screening is labor-intensive and subjective. Based on 2644 UWF images, a set of early fundus abnormal screening system named DeepUWF is developed. DeepUWF includes an abnormal fundus screening subsystem and a disease diagnosis subsystem for three kinds of fundus diseases (retinal tear & retinal detachment, diabetic retinopathy and pathological myopia). The components in the system are composed of a set of excellent convolutional neural networks and two custom classifiers. However, the contrast of UWF images used in the research is low, which seriously limits the extraction of fine features of UWF images by depth model. Therefore, the high specificity and low sensitivity of prediction results have always been difficult problems in research. In order to solve this problem, six kinds of image preprocessing techniques are adopted, and their effects on the prediction performance of fundus abnormal and three kinds of fundus diseases models are studied. A variety of experimental indicators are used to evaluate the algorithms for validity and reliability. The experimental results show that these preprocessing methods are helpful to improve the learning ability of the networks and achieve good sensitivity and specificity. Without ophthalmologists, DeepUWF has potential application value, which is helpful for fundus health screening and workflow improvement.
Yuanyuan Chen 0006, Jie Zhong 0004, Zhang Yi 0001
IEEE J. Biomed. Health Informatics5
2020 Discriminative globality and locality preserving graph embedding for dimensionality reduction
Jianping Gou, Zhang Yi 0001, Jiancheng Lv 0001, Qirong Mao, Yongzhao Zhan 0001
Expert Syst. Appl.3
2020 Multilabel classification by exploiting data-driven pair-wise label dependence
abstract
Exploiting label dependence is a widely used approach to boost classification performance for multilabel classification problems. However, most of the traditional label dependence methods have high time complexity, especially when combined with deep neural networks (DNNs). Thus they usually can not be efficiently applied in large-scale data sets. Recent advances in large-scale multilabel classification widely developed pair-wise ranking and structure-driven methods, but label dependence was little exploited. In most of the structure-driven methods, binary relevance (BR) with multiple binary cross-entropy (BCE) loss functions, a simple but effective method, is still the prior solution incorporation with DNNs in large-scale data sets. In this paper, we propose a novel loss function called label dependent cross-entropy (LDCE), which directly introduces label dependence to BCE loss function by data-driven conditional probability. Combined with deep convolutional neural networks (DCNNs), LDCE introduces no extra parameters and induces very little extra computational complexity. Moreover, we develop its tiny variant with sparse label dependence and its learnable version for automatic learning pair-wise label dependence. Within the BR scheme, LDCE outperforms BCE on seven widely used benchmark datasets. We also perform two large-scale multilabel image classification tasks (VOC 2007 and ChestX-ray14) with DCNNs, and LDCE outperforms BCE and achieves comparable results to the state-of-the-art.
Tao He 0016, Lei Zhang 0005, Jixiang Guo, Zhang Yi 0001
Int. J. Intell. Syst.4
2020 DeepEC: An error correction framework for dose prediction and organ segmentation using deep neural networks
abstract
Radiotherapy is an indispensable part of adjuvant therapy for cancer that improves local control, overall survival, and the opportunity for good quality of life. Organ delineation and dose plan design are the key steps in the treatment. Organ delineation controls the area of radiotherapy and dose planning controls its intensity. However, both tasks are time-consuming, exhausting, and subjective, and automated methods are desirable. Although automated methods have been studied, the previous studies either focus on organ segmentation or dose prediction, without considering them from a holistic perspective. In this paper, we treat organ segmentation and dose prediction as similar tasks, and propose an error correction framework to improve their performance based on the same mechanism. The proposed error correction framework consists of a prediction network and a calibration network. The biggest difference between our framework and previous studies is that the state-of-the-art networks can be used as a prediction network or calibration network, and then the performance can be improved by the error correction mechanism. To evaluate the framework, we conducted a series of experiments on dose prediction and organ segmentation. These experimental results show that the framework is superior to other state-of-the-art methods in both tasks.
Han Wang 0025, Haixian Zhang, Junjie Hu 0004, Sen Bai, Zhang Yi 0001
Int. J. Intell. Syst.6
2020 A novel method to compute the weights of neural networks
Zhentao Gao, Yuanyuan Chen 0006, Zhang Yi 0001
Neurocomputing3
2020 Automated diagnosis of multi-plane breast ultrasonography images using deep neural networks
Yong Pi, Dan Deng, Xiaofeng Qi, Jilan Li, Zhang Yi 0001
Neurocomputing7
2020 Automated diagnosis of neonatal encephalopathy on aEEG using deep neural networks
Jingling Wang, Rong Ju, Yuanyuan Chen 0006, Guijun Liu, Zhang Yi 0001
Neurocomputing5
2020 Surrogate dropout: Learning optimal drop rate through proxy
Junjie Hu 0004, Yuanyuan Chen 0006, Lei Zhang 0005, Zhang Yi 0001
Knowl. Based Syst.4
2020 Automated detection of kidney abnormalities using multi-feature fusion convolutional neural networks
Zhang Yi 0001
Knowl. Based Syst.2
2020 Neural networks model based on an automated multi-scale method for mammogram classification
Lizhang Xie, Lei Zhang 0005, Haiying Huang 0004, Zhang Yi 0001
Knowl. Based Syst.5
2020 DeepLN: A framework for automatic lung nodule detection using multi-resolution CT screening images
Xiuyuan Xu, Chengdi Wang, Jixiang Guo, Hongli Bai, Weimin Li 0003, Zhang Yi 0001
Knowl. Based Syst.7
2020 Multi-task learning for the segmentation of organs at risk with label dependence
Tao He 0016, Junjie Hu 0004, Jixiang Guo, Zhang Yi 0001
Medical Image Anal.5
2020 Automated diagnosis of bone metastasis based on multi-view bone scans using attention-augmented deep neural networks
Yong Pi, Yongzhao Xiang, Huawei Cai, Zhang Yi 0001
Medical Image Anal.6
2020 Automatic diagnosis for thyroid nodules in ultrasound images by deep neural networks
Lituan Wang, Lei Zhang 0005, Minjuan Zhu, Xiaofeng Qi, Zhang Yi 0001
Medical Image Anal.5
2020 MSCS-DeepLN: Evaluating lung nodule malignancy using multi-scale cost-sensitive neural networks
Xiuyuan Xu, Chengdi Wang, Jixiang Guo, Yuncui Gan, Jianyong Wang 0002, Hongli Bai, Lei Zhang 0005, Weimin Li 0003, Zhang Yi 0001
Medical Image Anal.9
2020 Weighted discriminative collaborative competitive representation for robust image classification
Jianping Gou, Lei Wang 0095, Zhang Yi 0001, Yun-Hao Yuan 0001, Weihua Ou, Qirong Mao
Neural Networks3
2020 Outlier Detection Using Structural Scores in a High-Dimensional Space
abstract
Outlier detection has drawn significant interest from both academia and industry, such as network intrusion detection. Most existing methods implicitly or explicitly rely on distances in Euclidean space. However, the Euclidean distance may be incapable of measuring the similarity among high-dimensional data due to the curse of dimensionality, thus leading to inferior performance in practice. This paper presents an innovative approach for outlier detection from the view of meaningful structure scores. If two points have similar features, the difference between their structural scores is small and vice versa. The scores are calculated by measuring the variance of angles weighted by data representation, which takes the global data structure into the measurement. Thus, it could consistently rank more similar points. Compared with existing methods, our structural scores could be better to reflect the characteristics of data in a high-dimensional space. The proposed method consistently ranks more similar points. Experiments on synthetic and several real-world datasets have demonstrated the effectiveness and efficiency of our proposed methods.
Xiaojie Li 0001, Jiancheng Lv 0001, Zhang Yi 0001
IEEE Trans. Cybern.3
2020 MediMLP: Using Grad-CAM to Extract Crucial Variables for Lung Cancer Postoperative Complication Prediction
abstract
Lung cancer postoperative complication prediction (PCP) is significant for decreasing the perioperative mortality rate after lung cancer surgery. In this paper we concentrate on two PCP tasks: (1) the binary classification for predicting whether a patient will have postoperative complications; and (2) the three-class multi-label classification for predicting which postoperative complication a patient will experience. Furthermore, an important clinical requirement of PCP is the extraction of crucial variables from electronic medical records. We propose a novel multi-layer perceptron (MLP) model called medical MLP (MediMLP) together with the gradient-weighted class activation mapping (Grad-CAM) algorithm for lung cancer PCP. The proposed MediMLP, which involves one locally connected layer and fully connected layers with a shortcut connection, simultaneously extracts crucial variables and performs PCP tasks. The experimental results indicated that MediMLP outperformed normal MLP on two PCP tasks and had comparable performance with existing feature selection methods. Using MediMLP and further experimental analysis, we found that the variable of "time of indwelling drainage tube" was very relevant to lung cancer postoperative complications.
Tao He 0016, Jixiang Guo, Xiuyuan Xu, Zihuai Wang, Kaiyu Fu, Lunxu Liu, Zhang Yi 0001
IEEE J. Biomed. Health Informatics8
2020 Deep Neural Networks With Region-Based Pooling Structures for Mammographic Image Classification
abstract
Breast cancer is one of the most frequently diagnosed solid cancers. Mammography is the most commonly used screening technology for detecting breast cancer. Traditional machine learning methods of mammographic image classification or segmentation using manual features require a great quantity of manual segmentation annotation data to train the model and test the results. But manual labeling is expensive, time-consuming, and laborious, and greatly increases the cost of system construction. To reduce this cost and the workload of radiologists, an end-to-end full-image mammogram classification method based on deep neural networks was proposed for classifier building, which can be constructed without bounding boxes or mask ground truth label of training data. The only label required in this method is the classification of mammographic images, which can be relatively easy to collect from diagnostic reports. Because breast lesions usually take up a fraction of the total area visualized in the mammographic image, we propose different pooling structures for convolutional neural networks(CNNs) instead of the common pooling methods, which divide the image into regions and select the few with high probability of malignancy as the representation of the whole mammographic image. The proposed pooling structures can be applied on most CNN-based models, which may greatly improve the models' performance on mammographic image data with the same input. Experimental results on the publicly available INbreast dataset and CBIS dataset indicate that the proposed pooling structures perform satisfactorily on mammographic image data compared with previous state-of-the-art mammographic image classifiers and detection algorithm using segmentation annotations.
Xin Shu 0005, Lei Zhang 0005, Zizhou Wang, Zhang Yi 0001
IEEE Trans. Medical Imaging5
2020 Evolutionary Compression of Deep Neural Networks for Biomedical Image Segmentation
abstract
Biomedical image segmentation is lately dominated by deep neural networks (DNNs) due to their surpassing expert-level performance. However, the existing DNN models for biomedical image segmentation are generally highly parameterized, which severely impede their deployment on real-time platforms and portable devices. To tackle this difficulty, we propose an evolutionary compression method (ECDNN) to automatically discover efficient DNN architectures for biomedical image segmentation. Different from the existing studies, ECDNN can optimize network loss and number of parameters simultaneously during the evolution, and search for a set of Pareto-optimal solutions in a single run, which is useful for quantifying the tradeoff in satisfying different objectives, and flexible for compressing DNN when preference information is uncertain. In particular, a set of novel genetic operators is proposed for automatically identifying less important filters over the whole network. Moreover, a pruning operator is designed for eliminating convolutional filters from layers involved in feature map concatenation, which is commonly adopted in DNN architectures for capturing multi-level features from biomedical images. Experiments carried out on compressing DNN for retinal vessel and neuronal membrane segmentation tasks show that ECDNN can not only improve the performance without any retraining but also discover efficient network architectures that well maintain the performance. The superiority of the proposed method is further validated by comparison with the state-of-the-art methods.
Yao Zhou 0002, Gary G. Yen, Zhang Yi 0001
IEEE Trans. Neural Networks Learn. Syst.3
2019 Discriminative Group Collaborative Competitive Representation for Visual Classification
abstract
In pattern recognition, the representation-based classification (RBC) has attracted much attention recently. As a representative one of RBC, collaborative representation-based classification (CRC) and its variants have achieved promising classification performance in many visual classification tasks. However, most of the CRC methods cannot directly consider the class discrimination information of data that is very important for classification. To fully use the class discrimination information, we propose a novel discriminative group collaborative competitive representation-based classification method (DGCCR) in this paper. In the designed DGCCR model, the discriminative competitive relationships of classes, the discriminative decorrelations among classes and the weighted class-specific group constraints are simultaneously taken into account for strengthening the power of pattern discrimination. Experiments on three visual classification data sets demonstrate that the proposed DGCCR out-performs state-of-the-art RBC methods.
Jianping Gou, Lei Wang 0095, Zhang Yi 0001, Yun-Hao Yuan 0001, Weihua Ou, Qirong Mao
ICME3
2019 A New Sparse Restricted Boltzmann Machine
abstract
Although existing sparse restricted Boltzmann machine (SRBM) can make some hidden units activated, the major disadvantage is that the sparseness of data distribution is usually overlooked and the reconstruction error becomes very large after the hidden unit variables become sparse. Different from the SRBMs which only incorporate a sparse constraint term in the energy function formula from the original restricted Boltzmann machine (RBM), an energy function constraint SRBM (ESRBM) is proposed in this paper. The proposed ESRBM takes into account the sparseness of the data distribution so that the learned features can better reflect the intrinsic features of data. Simulations show that compared with SRBM, ESRBM has smaller reconstruction error and lower computational complexity, and that for supervised learning classification, ESRBM obtains higher accuracy rates than SRBM, classification RBM, and Softmax classifier.
Jiangshu Wei, Jiancheng Lv 0001, Zhang Yi 0001
Int. J. Pattern Recognit. Artif. Intell.3
2019 Locality-constrained least squares regression for subspace clustering
Yuanyuan Chen 0006, Zhang Yi 0001
Knowl. Based Syst.2
2019 Locality constrained representation-based K-nearest neighbor classification
Jianping Gou, Wenmo Qiu, Zhang Yi 0001, Xiangjun Shen, Yongzhao Zhan 0001, Weihua Ou
Knowl. Based Syst.3
2019 A multitask multiview clustering algorithm in heterogeneous situations based on LLE and LE
Zhang Yi 0001, Yan Yang 0001, Tianrui Li 0001, Hamido Fujita
Knowl. Based Syst.1
2019 Automated identification and grading system of diabetic retinopathy using deep neural networks
Jie Zhong 0004, Shijun Yang, Zhentao Gao, Junjie Hu 0004, Yuanyuan Chen 0006, Zhang Yi 0001
Knowl. Based Syst.7
2019 Automated segmentation of macular edema in OCT using deep neural networks
Junjie Hu 0004, Yuanyuan Chen 0006, Zhang Yi 0001
Medical Image Anal.3
2019 Automated diagnosis of breast ultrasonography images using deep neural networks
Xiaofeng Qi, Lei Zhang 0005, Yong Pi, Yi Chen 0034, Zhang Yi 0001
Medical Image Anal.7
2019 Spectrogram based multi-task audio classification
Yuni Zeng, Hua Mao 0001, Dezhong Peng, Zhang Yi 0001
Multim. Tools Appl.4
2019 Stem cell motion-tracking by using deep neural networks with multi-output
Yangxu Wang, Hua Mao 0001, Zhang Yi 0001
Neural Comput. Appl.3
2019 Predicting movie box-office revenues using deep neural networks
Yao Zhou 0002, Lei Zhang 0005, Zhang Yi 0001
Neural Comput. Appl.3
2019 A Novel Deep Learning-Based Collaborative Filtering Model for Recommendation System
abstract
The collaborative filtering (CF) based models are capable of grasping the interaction or correlation of users and items under consideration. However, existing CF-based methods can only grasp single type of relation, such as restricted Boltzmann machine which distinctly seize the correlation of user-user or item-item relation. On the other hand, matrix factorization explicitly captures the interaction between them. To overcome these setbacks in CF-based methods, we propose a novel deep learning method which imitates an effective intelligent recommendation by understanding the users and items beforehand. In the initial stage, corresponding low-dimensional vectors of users and items are learned separately, which embeds the semantic information reflecting the user-user and item-item correlation. During the prediction stage, a feed-forward neural networks is employed to simulate the interaction between user and item, where the corresponding pretrained representational vectors are taken as inputs of the neural networks. Several experiments based on two benchmark datasets (MovieLens 1M and MovieLens 10M) are carried out to verify the effectiveness of the proposed method, and the result shows that our model outperforms previous methods that used feed-forward neural networks by a significant margin and performs very comparably with state-of-the-art methods on both datasets.
Mingsheng Fu, Hong Qu 0002, Zhang Yi 0001, Li Lu 0001
IEEE Trans. Cybern.3
2019 Robust Multiobjective Optimization via Evolutionary Algorithms
abstract
Uncertainty inadvertently exists in most real-world applications. In the optimization process, uncertainty poses a very important issue and it directly affects the optimization performance. Nowadays, evolutionary algorithms (EAs) have been successfully applied to various multiobjective optimization problems (MOPs). However, current researches on EAs rarely consider uncertainty in the optimization process and existing algorithms often fail to handle the uncertainty, which have limited EAs' applications in real-world problems. When MOPs come with uncertainty, they are referred to as robust MOPs (RMOPs). In this paper, we aim at solving RMOPs using EA-based optimization search. We propose a novel robust multiobjective optimization EA (RMOEA) with two distinct, yet complement, parts: 1) multiobjective optimization finding global Pareto optimal front ignoring disturbance at first and 2) robust optimization searching for the robust optimal front afterward. Furthermore, a comprehensive performance evaluation method is proposed to quantify the performance of RMOEA in solving RMOPs. Experimental results on a group of benchmark functions demonstrate the superiority of the proposed design in terms of both solutions' quality under the disturbance and computational efficiency in solving RMOPs.
Zhenan He 0001, Gary G. Yen, Zhang Yi 0001
IEEE Trans. Evol. Comput.3
2019 Evolving Unsupervised Deep Neural Networks for Learning Meaningful Representations
abstract
Deep learning (DL) aims at learning the meaningful representations. A meaningful representation gives rise to significant performance improvement of associated machine learning (ML) tasks by replacing the raw data as the input. However, optimal architecture design and model parameter estimation in DL algorithms are widely considered to be intractable. Evolutionary algorithms are much preferable for complex and nonconvex problems due to its inherent characteristics of gradient-free and insensitivity to the local optimal. In this paper, we propose a computationally economical algorithm for evolving unsupervised deep neural networks to efficiently learn meaningful representations, which is very suitable in the current big data era where sufficient labeled data for training is often expensive to acquire. In the proposed algorithm, finding an appropriate architecture and the initialized parameter values for an ML task at hand is modeled by one computational efficient gene encoding approach, which is employed to effectively model the task with a large number of parameters. In addition, a local search strategy is incorporated to facilitate the exploitation search for further improving the performance. Furthermore, a small proportion labeled data is utilized during evolution search to guarantee the learned representations to be meaningful. The performance of the proposed algorithm has been thoroughly investigated over classification tasks. Specifically, error classification rate on MNIST with 1.15% is reached by the proposed algorithm consistently, which is considered a very promising result against state-of-the-art unsupervised DL algorithms.
Yanan Sun 0001, Gary G. Yen, Zhang Yi 0001
IEEE Trans. Evol. Comput.3
2019 IGD Indicator-Based Evolutionary Algorithm for Many-Objective Optimization Problems
abstract
Inverted generational distance (IGD) has been widely considered as a reliable performance indicator to concurrently quantify the convergence and diversity of multiobjective and many-objective evolutionary algorithms. In this paper, an IGD indicator-based evolutionary algorithm for solving many-objective optimization problems (MaOPs) has been proposed. Specifically, the IGD indicator is employed in each generation to select the solutions with favorable convergence and diversity. In addition, a computationally efficient dominance comparison method is designed to assign the rank values of solutions along with three newly proposed proximity distance assignments. Based on these two designs, the solutions are selected from a global view by linear assignment mechanism to concern the convergence and diversity simultaneously. In order to facilitate the accuracy of the sampled reference points for the calculation of IGD indicator, we also propose an efficient decomposition-based nadir point estimation method for constructing the Utopian Pareto front (PF) which is regarded as the best approximate PF for real-world MaOPs at the early stage of the evolution. To evaluate the performance, a series of experiments is performed on the proposed algorithm against a group of selected state-of-the-art many-objective optimization algorithms over optimization problems with 8-, 15-, and 20-objective. Experimental results measured by the chosen performance metrics indicate that the proposed algorithm is very competitive in addressing MaOPs.
Yanan Sun 0001, Gary G. Yen, Zhang Yi 0001
IEEE Trans. Evol. Comput.3
2019 A Local Mean Representation-based K-Nearest Neighbor Classifier
abstract
K -nearest neighbor classification method (KNN), as one of the top 10 algorithms in data mining, is a very simple and yet effective nonparametric technique for pattern recognition. However, due to the selective sensitiveness of the neighborhood size k , the simple majority vote, and the conventional metric measure, the KNN-based classification performance can be easily degraded, especially in the small training sample size cases. In this article, to further improve the classification performance and overcome the main issues in the KNN-based classification, we propose a local mean representation-based k -nearest neighbor classifier (LMRKNN). In the LMRKNN, the categorical k -nearest neighbors of a query sample are first chosen to calculate the corresponding categorical k -local mean vectors, and then the query sample is represented by the linear combination of the categorical k -local mean vectors; finally, the class-specific representation-based distances between the query sample and the categorical k -local mean vectors are adopted to determine the class of the query sample. Extensive experiments on many UCI and KEEL datasets and three popular face databases are carried out by comparing LMRKNN to the state-of-art KNN-based methods. The experimental results demonstrate that the proposed LMRKNN outperforms the related competitive KNN-based methods with more robustness and effectiveness.
Jianping Gou, Wenmo Qiu, Zhang Yi 0001, Yong Xu 0001, Qirong Mao, Yongzhao Zhan 0001
ACM Trans. Intell. Syst. Technol.3
2019 Automated Analysis for Retinopathy of Prematurity by Deep Neural Networks
abstract
Retinopathy of Prematurity (ROP) is a retinal vasproliferative disorder disease principally observed in infants born prematurely with low birth weight. ROP is an important cause of childhood blindness. Although automatic or semi-automatic diagnosis of ROP has been conducted, most previous studies have focused on "plus" disease, which is indicated by abnormalities of retinal vasculature. Few studies have reported methods for identifying the "stage" of the ROP disease. Deep neural networks have achieved impressive results in many computer vision and medical image analysis problems, raising expectations that it might be a promising tool in the automatic diagnosis of ROP. In this paper, convolutional neural networks with a novel architecture are proposed to recognize the existence and severity of ROP disease per-examination. The severity of ROP is divided into mild and severe cases according to the disease progression. The proposed architecture consists of two sub-networks connected by a feature aggregate operator. The first sub-network is designed to extract high-level features from images of the fundus. These features from different images in an examination are fused by the aggregate operator, then used as the input for the second sub-network to predict its class. A large data set imaged by RetCam 3 is used to train and evaluate the model. The high classification accuracy in the experiment demonstrates the effectiveness of the proposed architecture for recognizing the ROP disease.
Junjie Hu 0004, Yuanyuan Chen 0006, Jie Zhong 0004, Rong Ju, Zhang Yi 0001
IEEE Trans. Medical Imaging5
2019 A Highly Effective and Robust Membrane Potential-Driven Supervised Learning Method for Spiking Neurons
abstract
Spiking neurons are becoming increasingly popular owing to their biological plausibility and promising computational properties. Unlike traditional rate-based neural models, spiking neurons encode information in the temporal patterns of the transmitted spike trains, which makes them more suitable for processing spatiotemporal information. One of the fundamental computations of spiking neurons is to transform streams of input spike trains into precisely timed firing activity. However, the existing learning methods, used to realize such computation, often result in relatively low accuracy performance and poor robustness to noise. In order to address these limitations, we propose a novel highly effective and robust membrane potential-driven supervised learning (MemPo-Learn) method, which enables the trained neurons to generate desired spike trains with higher precision, higher efficiency, and better noise robustness than the current state-of-the-art spiking neuron learning methods. While the traditional spike-driven learning methods use an error function based on the difference between the actual and desired output spike trains, the proposed MemPo-Learn method employs an error function based on the difference between the output neuron membrane potential and its firing threshold. The efficiency of the proposed learning method is further improved through the introduction of an adaptive strategy, called skip scan training strategy, that selectively identifies the time steps when to apply weight adjustment. The proposed strategy enables the MemPo-Learn method to effectively and efficiently learn the desired output spike train even when much smaller time steps are used. In addition, the learning rule of MemPo-Learn is improved further to help mitigate the impact of the input noise on the timing accuracy and reliability of the neuron firing dynamics. The proposed learning method is thoroughly evaluated on synthetic data and is further demonstrated on real-world classification tasks. Experimental results show that the proposed method can achieve high learning accuracy with a significant improvement in learning time and better robustness to different types of noise.
Malu Zhang, Hong Qu 0002, Ammar Belatreche, Yi Chen 0034, Zhang Yi 0001
IEEE Trans. Neural Networks Learn. Syst.5
2018 A Supervised Learning Framework for Prediction of Incompatible Herb Pair in Traditional Chinese Medicine
abstract
Adverse drug-drug interaction has been a critical issue for the development of drugs. In Traditional Chinese Medicine, adverse herb-herb interaction is a negative reaction in patients after the absorption of decoction of Incompatible Herb Pair (IHP). Recently, many methods have been proposed for IHP research, but most of them focused on revealing and analyzing the adverse reaction of some known IHPs, despite that a number of new IHPs have been discovered by accidents. Up to now, IHPs have been a serious threat to public health in the TCM medication. In this paper, we propose a novel supervised learning framework for potential IHP prediction. In this framework, we model the prediction task as a non-negative matrix tri-factorization problem, in which two important herb attributes (efficacy and flavor) and their correlation are incorporated to characterize the incompatible relationship among herbs. A hypothetical test method is adopted to evaluate the statistical significance of dissimilar characteristics of two attributes and the results are used as a regularization term to improve the accuracy of IHP prediction. Experiments on the real-world IHP dataset demonstrate that the proposed framework is very effective for prediction of potential IHPs.
Jiajing Zhu, Yongguo Liu, Shangming Yang, Shuangqing Zhai, Zhang Yi 0001, Chuanbiao Wen
CIKM5
2018 Continuous Attractors of Nonlinear Neural Networks with Asymmetric Connection Weights
Zhang Yi 0001, Zhixin Pang
ICONIP (2)2
2018 Continuous Attractors of 3-D Discrete-Time Ring Networks with Circulant Weight Matrix
Zhang Yi 0001, De-An Wu, Xiong Dai
ISNN2
2018 A New Delay Connection for Long Short-Term Memory Networks
abstract
Connections play a crucial role in neural network (NN) learning because they determine how information flows in NNs. Suitable connection mechanisms may extensively enlarge the learning capability and reduce the negative effect of gradient problems. In this paper, a new delay connection is proposed for Long Short-Term Memory (LSTM) unit to develop a more sophisticated recurrent unit, called Delay Connected LSTM (DCLSTM). The proposed delay connection brings two main merits to DCLSTM with introducing no extra parameters. First, it allows the output of the DCLSTM unit to maintain LSTM, which is absent in the LSTM unit. Second, the proposed delay connection helps to bridge the error signals to previous time steps and allows it to be back-propagated across several layers without vanishing too quickly. To evaluate the performance of the proposed delay connections, the DCLSTM model with and without peephole connections was compared with four state-of-the-art recurrent model on two sequence classification tasks. DCLSTM model outperformed the other models with higher accuracy and F1[Formula: see text]score. Furthermore, the networks with multiple stacked DCLSTM layers and the standard LSTM layer were evaluated on Penn Treebank (PTB) language modeling. The DCLSTM model achieved lower perplexity (PPL)/bit-per-character (BPC) than the standard LSTM model. The experiments demonstrate that the learning of the DCLSTM models is more stable and efficient.
Jianyong Wang 0002, Lei Zhang 0005, Yuanyuan Chen 0006, Zhang Yi 0001
Int. J. Neural Syst.4
2018 Symmetric low-rank preserving projections for subspace learning
abstract
Graph construction plays an important role in graph-oriented subspace learning. However, most existing approaches cannot simultaneously consider the global and local structures of high-dimensional data. In order to solve this deficiency, we propose a symmetric low-rank preserving projection (SLPP) framework incorporating a symmetric constraint and a local regularization into low-rank representation learning for subspace learning. Under this framework, SLPP-M is incorporated with manifold regularization as its local regularization while SLPP-S uses sparsity regularization. Besides characterizing the global structure of high-dimensional data by a symmetric low-rank representation, both SLPP-M and SLPP-S effectively exploit the local manifold and geometric structure by incorporating manifold and sparsity regularization, respectively. The similarity matrix is successfully learned by solving the nuclear-norm minimization optimization problem . Combined with graph embedding techniques, a transformation matrix effectively preserves the low-dimensional structure features of high-dimensional data. In order to facilitate classification by exploiting available labels of training samples , we also develop a supervised version of SLPP-M and SLPP-S under the SLPP framework, named S-SLPP-M and S-SLPP-S, respectively. Experimental results in face, handwriting and object recognition applications demonstrate the efficiency of the proposed algorithm for subspace learning.
Jie Chen 0065, Hua Mao 0001, Haixian Zhang, Zhang Yi 0001
Neurocomputing4
2018 Subspace clustering using a low-rank constrained autoencoder
Yuanyuan Chen 0006, Lei Zhang 0005, Zhang Yi 0001
Inf. Sci.3
2018 Incremental rough set approach for hierarchical multicriteria classification
Chuan Luo 0001, Tianrui Li 0001, Hongmei Chen 0001, Hamido Fujita, Zhang Yi 0001
Inf. Sci.5
2018 Audio classification using attention-augmented convolutional neural network
abstract
Audio classification, as a set of important and challenging tasks, groups speech signals according to speakers’ identities, accents, and emotional states . Due to the high dimensionality of the audio data, task-specific hand-crafted features extraction is always required and regarded cumbersome for various audio classification tasks . More importantly, the inherent relationship among features has not been fully exploited. In this paper, the original speech signal is first represented as spectrogram and later be split along the frequency domain to form frequency-distributed spectrogram . This paper proposes a task-independent model, called FreqCNN, to automaticly extract distinctive features from each frequency band by using convolutional kernels. Further more, an attention mechanism is introduced to systematically enhance the features from certain frequency bands. The proposed FreqCNN is evaluated on three publicly available speech databases thorough three independent classification tasks . The obtained results demonstrate superior performance over the state-of-the-art.
Hua Mao 0001, Zhang Yi 0001
Knowl. Based Syst.3
2018 Symmetric convolutional neural network for mandible segmentation
Ming Yan 0007, Jixiang Guo, Zhang Yi 0001
Knowl. Based Syst.4
2018 Tracking topology structure adaptively with deep neural networks
Xueying Shi, Guangyong Chen, Pheng-Ann Heng, Zhang Yi 0001
Neural Comput. Appl.4
2018 Subnormal Distribution Derived From Evolving Networks With Variable Elements
abstract
During the past decades, power-law distributions have played a significant role in analyzing the topology of scale-free networks. However, in the observation of degree distributions in practical networks and other nonuniform distributions such as the wealth distribution, we discover that, there exists a peak at the beginning of most real distributions, which cannot be accurately described by a monotonic decreasing power-law distribution. To better describe the real distributions, in this paper, we propose a subnormal distribution derived from evolving networks with variable elements and study its statistical properties for the first time. By utilizing this distribution, we can precisely describe those distributions commonly existing in the real world, e.g., distributions of degree in social networks and personal wealth. Additionally, we fit connectivity in evolving networks and the data observed in the real world by the proposed subnormal distribution, resulting in a better performance of fitness.
Minyu Feng, Hong Qu 0002, Zhang Yi 0001, Jürgen Kurths
IEEE Trans. Cybern.3
2018 Improved Regularity Model-Based EDA for Many-Objective Optimization
abstract
The performance of multiobjective evolutionary algorithms deteriorates appreciably in solving many-objective optimization problems (MaOPs) which encompass more than three objectives. One of the known rationales is the loss of selection pressure which leads to the selected parents not generating promising offspring toward Pareto-optimal front (PF) with diversity. Estimation of distribution algorithms sample new solutions with a probabilistic model built from the statistics extracting over the existing solutions so as to mitigate the adverse impact of genetic operators. In this paper, an improved regularity-based estimation of distribution algorithm is proposed to effectively tackle unconstrained MaOPs. In the proposed algorithm, diversity repairing mechanism is utilized to mend the areas, where need nondominated solutions with a closer proximity to the PF. Then favorable solutions are generated by the model built from the regularity of the solutions surrounding a group of representatives. These two steps collectively enhance the selection pressure which gives rise to the superior convergence of the proposed algorithm. In addition, dimension reduction technique is employed in the decision space to speed up the estimation search of the proposed algorithm. Finally, by assigning the Pareto-optimal solutions to the uniformly distributed reference vectors, a set of solutions with excellent diversity and convergence is obtained. To measure the performance, NSGA-III, GrEA, MOEA/D, HypE, MBN-EDA, and RM-MEDA are selected to perform comparison experiments over DTLZ and DTLZ-test suites with 3-, 5-, 8-, 10-, and 15-objective. Experimental results quantified by the selected performance metrics reveal that the proposed algorithm shows considerable competitiveness in addressing unconstrained MaOPs.
Yanan Sun 0001, Gary G. Yen, Zhang Yi 0001
IEEE Trans. Evol. Comput.3
2018 A Fractional-Order Variational Framework for Retinex: Fractional-Order Partial Differential Equation-Based Formulation for Multi-Scale Nonlocal Contrast Enhancement with Texture Preserving
abstract
This paper discusses a novel conceptual formulation of the fractional-order variational framework for retinex, which is a fractional-order partial differential equation (FPDE) formulation of retinex for the multi-scale nonlocal contrast enhancement with texture preserving. The well-known shortcomings of traditional integer-order computation-based contrast-enhancement algorithms, such as ringing artefacts and staircase effects, are still in great need of special research attention. Fractional calculus has potentially received prominence in applications in the domain of signal processing and image processing mainly because of its strengths like long-term memory, nonlocality, and weak singularity, and because of the ability of a fractional differential to enhance the complex textural details of an image in a nonlinear manner. Therefore, in an attempt to address the aforementioned problems associated with traditional integer-order computation-based contrast-enhancement algorithms, we have studied here, as an interesting theoretical problem, whether it will be possible to hybridize the capabilities of preserving the edges and the textural details of fractional calculus with texture image multi-scale nonlocal contrast enhancement. Motivated by this need, in this paper, we introduce a novel conceptual formulation of the fractional-order variational framework for retinex. First, we implement the FPDE by means of the fractional-order steepest descent method. Second, we discuss the implementation of the restrictive fractional-order optimization algorithm and the fractional-order Courant-Friedrichs-Lewy condition. Third, we perform experiments to analyze the capability of the FPDE to preserve edges and textural details, while enhancing the contrast. The capability of the FPDE to preserve edges and textural details is a fundamental important advantage, which makes our proposed algorithm superior to the traditional integer-order computation-based contrast enhancement algorithms, especially for images rich in textural details.
Yi-Fei Pu, Patrick Siarry, Amitava Chatterjee, Zhengning Wang, Zhang Yi 0001, Yiguang Liu, Jiliu Zhou, Yan Wang 0015
IEEE Trans. Image Process.5
2018 Graph Regularized Restricted Boltzmann Machine
abstract
The restricted Boltzmann machine (RBM) has received an increasing amount of interest in recent years. It determines good mapping weights that capture useful latent features in an unsupervised manner. The RBM and its generalizations have been successfully applied to a variety of image classification and speech recognition tasks. However, most of the existing RBM-based models disregard the preservation of the data manifold structure. In many real applications, the data generally reside on a low-dimensional manifold embedded in high-dimensional ambient space. In this brief, we propose a novel graph regularized RBM to capture features and learning representations, explicitly considering the local manifold structure of the data. By imposing manifold-based locality that preserves constraints on the hidden layer of the RBM, the model ultimately learns sparse and discriminative representations. The representations can reflect data distributions while simultaneously preserving the local manifold structure of data. We test our model using several benchmark image data sets for unsupervised clustering and supervised classification problem. The results demonstrate that the performance of our method exceeds the state-of-the-art alternatives.
Dongdong Chen 0004, Jiancheng Lv 0001, Zhang Yi 0001
IEEE Trans. Neural Networks Learn. Syst.3
2018 An Efficient Representation-Based Method for Boundary Point and Outlier Detection
abstract
Detecting boundary points (including outliers) is often more interesting than detecting normal observations, since they represent valid, interesting, and potentially valuable patterns. Since data representation can uncover the intrinsic data structure, we present an efficient representation-based method for detecting such points, which are generally located around the margin of densely distributed data, such as a cluster. For each point, the negative components in its representation generally correspond to the boundary points among its affine combination of points. In the presented method, the reverse unreachability of a point is proposed to evaluate to what degree this observation is a boundary point. The reverse unreachability can be calculated by counting the number of zero and negative components in the representation. The reverse unreachability explicitly takes into account the global data structure and reveals the disconnectivity between a data point and other points. This paper reveals that the reverse unreachability of points with lower density has a higher score. Note that the score of reverse unreachability of an outlier is greater than that of a boundary point. The top- ranked points can thus be identified as outliers. The greater the value of the reverse unreachability, the more likely the point is a boundary point. Compared with related methods, our method better reflects the characteristics of the data, and simultaneously detects outliers and boundary points regardless of their distribution and the dimensionality of the space. Experimental results obtained for a number of synthetic and real-world data sets demonstrate the effectiveness and efficiency of our method.
Xiaojie Li 0001, Jiancheng Lv 0001, Zhang Yi 0001
IEEE Trans. Neural Networks Learn. Syst.3
2018 Connections Between Nuclear-Norm and Frobenius-Norm-Based Representations
abstract
A lot of works have shown that frobenius-norm-based representation (FNR) is competitive to sparse representation and nuclear-norm-based representation (NNR) in numerous tasks such as subspace clustering. Despite the success of FNR in experimental studies, less theoretical analysis is provided to understand its working mechanism. In this brief, we fill this gap by building the theoretical connections between FNR and NNR. More specially, we prove that: 1) when the dictionary can provide enough representative capacity, FNR is exactly NNR even though the data set contains the Gaussian noise, Laplacian noise, or sample-specified corruption and 2) otherwise, FNR and NNR are two solutions on the column space of the dictionary.
Xi Peng 0001, Canyi Lu, Zhang Yi 0001, Huajin Tang
IEEE Trans. Neural Networks Learn. Syst.3
2018 Recurrent Neural Networks With Auxiliary Memory Units
abstract
Memory is one of the most important mechanisms in recurrent neural networks (RNNs) learning. It plays a crucial role in practical applications, such as sequence learning. With a good memory mechanism, long term history can be fused with current information, and can thus improve RNNs learning. Developing a suitable memory mechanism is always desirable in the field of RNNs. This paper proposes a novel memory mechanism for RNNs. The main contributions of this paper are: 1) an auxiliary memory unit (AMU) is proposed, which results in a new special RNN model (AMU-RNN), separating the memory and output explicitly and 2) an efficient learning algorithm is developed by employing the technique of error flow truncation. The proposed AMU-RNN model, together with the developed learning algorithm, can learn and maintain stable memory over a long time range. This method overcomes both the learning conflict problem and gradient vanishing problem. Unlike the traditional method, which mixes the memory and output with a single neuron in a recurrent unit, the AMU provides an auxiliary memory neuron to maintain memory in particular. By separating the memory and output in a recurrent unit, the problem of learning conflicts can be eliminated easily. Moreover, by using the technique of error flow truncation, each auxiliary memory neuron ensures constant error flow during the learning process. The experiments demonstrate good performance of the proposed AMU-RNNs and the developed learning algorithm. The method exhibits quite efficient learning performance with stable convergence in the AMU-RNN learning and outperforms the state-of-the-art RNN models in sequence generation and sequence classification tasks.
Jianyong Wang 0002, Lei Zhang 0005, Quan Guo, Zhang Yi 0001
IEEE Trans. Neural Networks Learn. Syst.4
2018 Theoretical Study of Oscillator Neurons in Recurrent Neural Networks
abstract
Neurons in a network can be both active or inactive. Given a subset of neurons in a network, is it possible for the subset of neurons to evolve to form an active oscillator by applying some external periodic stimulus? Furthermore, can these oscillator neurons be observable, that is, is it a stable oscillator? This paper explores such possibility, finding that an important property: any subset of neurons can be intermittently co-activated to form a stable oscillator by applying some external periodic input without any condition. Thus, the existing of intermittently active oscillator neurons is an essential property possessed by the networks. Moreover, this paper shows that, under some conditions, a subset of neurons can be fully co-activated to form a stable oscillator. Such neurons are called selectable oscillator neurons. Necessary and sufficient conditions are established for a subset of neurons to be selectable oscillator neurons in linear threshold recurrent neuron networks. It is proved that a subset of neurons forms selectable oscillator neurons if and only if the real part of each eigenvalue of the associated synaptic connection weight submatrix of the network is not larger than one. This simple condition makes the concept of selectable oscillator neurons tractable. The selectable oscillator neurons can be regarded as memories stored in the synaptic connections of networks, which enables to find a new perspective of memories in neural networks, different from the equilibrium-type attractors.
Lei Zhang 0005, Zhang Yi 0001, Shun-ichi Amari
IEEE Trans. Neural Networks Learn. Syst.2
2017 Cascade Subspace Clustering
abstract
In this paper, we recast the subspace clustering as a verification problem. Our idea comes from an assumption that the distribution between a given sample x and cluster centers Omega is invariant to different distance metrics on the manifold, where each distribution is defined as a probability map (i.e. soft-assignment) between x and Omega. To verify this so-called invariance of distribution, we propose a deep learning based subspace clustering method which simultaneously learns a compact representation using a neural network and a clustering assignment by minimizing the discrepancy between pair-wise sample-centers distributions. To the best of our knowledge, this is the first work to reformulate clustering as a verification problem. Moreover, the proposed method is also one of the first several cascade clustering models which jointly learn representation and clustering in end-to-end manner. Extensive experimental results show the effectiveness of our algorithm comparing with 11 state-of-the-art clustering approaches on four data sets regarding to four evaluation metrics.
Xi Peng 0001, Jiashi Feng, Jiwen Lu, Weiyun Yau, Zhang Yi 0001
AAAI5
2017 Global view-based selection mechanism for many-objective evolutionary algorithms
abstract
In traditional many-objective evolutionary algorithms (MaOEAs), solutions survived to the next generation are individually selected which leads to the favorable quality upon the population composed of these selected solutions not necessarily to be gained. However, MaOEAs which are widely used in solving many-objective optimization problems (MaOPs) are considerably preferred due to their population-based nature. In this paper, a global view-based selection mechanism has been proposed to concern the quality of the selected solutions from a global prospective, which is capable of simultaneously facilitating the performance of the entire selected solutions. Indeed, the proposed selection mechanism is equivalent to solve a linear assignment problem whose cost matrix is constructed by the entries concurrently measuring the convergence and diversity of each solution. In addition, this design principle is also utilized for the mating selection to guarantee the convergence and diversity of the selected parents to generate promising offspring. As a case study, the proposed global view-based selection mechanism is integrated into NSGA-III (i.e., GS-NSGA-III). To validate the proposed selection mechanism, extensive experiments are performed by GS-NSGA-III against four state-of-the-art MaOEAs over 8-, 10-, and 15-objective DTLZ1-DTLZ7 test problems. The results measured by the selected performance metric reveal that GS-NSGA-III shows considerable competitiveness in addressing MaOPs.
Yanan Sun 0001, Gary G. Yen, Zhang Yi 0001
CEC3
2017 Cell tracking using deep neural networks with multi-task learning
Tao He 0016, Hua Mao 0001, Jixiang Guo, Zhang Yi 0001
Image Vis. Comput.4
2017 Subspace clustering using a symmetric low-rank representation
Jie Chen 0065, Hua Mao 0001, Yongsheng Sang, Zhang Yi 0001
Knowl. Based Syst.4
2017 Reference line-based Estimation of Distribution Algorithm for many-objective optimization
Yanan Sun 0001, Gary G. Yen, Zhang Yi 0001
Knowl. Based Syst.3
2017 Protein secondary structure prediction by using deep learning method
Yangxu Wang, Hua Mao 0001, Zhang Yi 0001
Knowl. Based Syst.3
2017 Cell mitosis detection using deep neural networks
Yao Zhou 0002, Hua Mao 0001, Zhang Yi 0001
Knowl. Based Syst.3
2017 Moving object recognition using multi-view three-dimensional convolutional neural networks
Tao He 0016, Hua Mao 0001, Zhang Yi 0001
Neural Comput. Appl.3
2017 Explicit guiding auto-encoders for learning meaningful representation
Yanan Sun 0001, Hua Mao 0001, Yongsheng Sang, Zhang Yi 0001
Neural Comput. Appl.4
2017 Automatic Subspace Learning via Principal Coefficients Embedding
abstract
In this paper, we address two challenging problems in unsupervised subspace learning: 1) how to automatically identify the feature dimension of the learned subspace (i.e., automatic subspace learning) and 2) how to learn the underlying subspace in the presence of Gaussian noise (i.e., robust subspace learning). We show that these two problems can be simultaneously solved by proposing a new method [(called principal coefficients embedding (PCE)]. For a given data set , PCE recovers a clean data set from and simultaneously learns a global reconstruction relation of . By preserving into an -dimensional space, the proposed method obtains a projection matrix that can capture the latent manifold structure of , where is automatically determined by the rank of with theoretical guarantees. PCE has three advantages: 1) it can automatically determine the feature dimension even though data are sampled from a union of multiple linear subspaces in presence of the Gaussian noise; 2) although the objective function of PCE only considers the Gaussian noise, experimental results show that it is robust to the non-Gaussian noise (e.g., random pixel corruption) and real disguises; and 3) our method has a closed-form solution and can be calculated very fast. Extensive experimental results show the superiority of PCE on a range of databases with respect to the classification accuracy, robustness, and efficiency.
Xi Peng 0001, Jiwen Lu, Zhang Yi 0001, Rui Yan 0005
IEEE Trans. Cybern.3
2017 Constructing the L2-Graph for Robust Subspace Learning and Subspace Clustering
abstract
Under the framework of graph-based learning, the key to robust subspace clustering and subspace learning is to obtain a good similarity graph that eliminates the effects of errors and retains only connections between the data points from the same subspace (i.e., intrasubspace data points). Recent works achieve good performance by modeling errors into their objective functions to remove the errors from the inputs. However, these approaches face the limitations that the structure of errors should be known prior and a complex convex problem must be solved. In this paper, we present a novel method to eliminate the effects of the errors from the projection space (representation) rather than from the input space. We first prove that ℓ1-, ℓ2-, ℓ∞-, and nuclear-norm-based linear projection spaces share the property of intrasubspace projection dominance, i.e., the coefficients over intrasubspace data points are larger than those over intersubspace data points. Based on this property, we introduce a method to construct a sparse similarity graph, called L2-graph. The subspace clustering and subspace learning algorithms are developed upon L2-graph. We conduct comprehensive experiment on subspace learning, image clustering, and motion segmentation and consider several quantitative benchmarks classification/clustering accuracy, normalized mutual information, and running time. Results show that L2-graph outperforms many state-of-the-art methods in our experiments, including L1-graph, low rank representation (LRR), and latent LRR, least square regression, sparse subspace clustering, and locally linear representation.
Xi Peng 0001, Zhiding Yu, Zhang Yi 0001, Huajin Tang
IEEE Trans. Cybern.3
2017 Trajectory Predictor by Using Recurrent Neural Networks in Visual Tracking
abstract
Motion models have been proved to be a crucial part in the visual tracking process. In recent trackers, particle filter and sliding windows-based motion models have been widely used. Treating motion models as a sequence prediction problem, we can estimate the motion of objects using their trajectories. Moreover, it is possible to transfer the learned knowledge from annotated trajectories to new objects. Inspired by recent advance in deep learning for visual feature extraction and sequence prediction, we propose a trajectory predictor to learn prior knowledge from annotated trajectories and transfer it to predict the motion of target objects. In this predictor, convolutional neural networks extract the visual features of target objects. Long short-term memory model leverages the annotated trajectory priors as well as sequential visual information, which includes the tracked features and center locations of the target object, to predict the motion. Furthermore, to extend this method to videos in which it is difficult to obtain annotated trajectories, a dynamic weighted motion model that combines the proposed trajectory predictor with a random sampler is proposed. To evaluate the transfer performance of the proposed trajectory predictor, we annotated a real-world vehicle dataset. Experiment results on both this real-world vehicle dataset and an online tracker benchmark dataset indicate that the proposed method outperforms several state-of-the-art trackers.
Lituan Wang, Lei Zhang 0005, Zhang Yi 0001
IEEE Trans. Cybern.3
2017 Contextual Noise Reduction for Domain Adaptive Near-Duplicate Retrieval on Merchandize Images
abstract
In this paper, we have proposed a novel method which utilizes the contextual relationship among visual words for reducing the Quantization errors in near-duplicate image retrieval (NDR). Instead of following the track of conventional NDR techniques which usually search new solutions by borrowing ideas from the text domain, we propose to model the problem back to image domain, which results in a more natural way of solution search. The idea of the proposed method is to construct a context graph that encapsulates the contextual relationship within an image and treat the graph as a pseudo-image, so that classical image filters can be adopted to reduce the mismapped visual words which are contextually inconsistent with others.With these contextual noises reduced, the method provides purified inputs to the subsequent processes in NDR, and improves the overall accuracy. More importantly, the purification further increases the sparsity of the image feature vectors, which thus speeds up the conventional methods by 1662% times and makes NDR practical to online applications on merchandize images where the requirement of response time is critical. The way of considering contextual noise reduction in image domain also makes the problem open to all sophisticated filters. Our study shows the classic anisotropic diffusion filter can be employed to address the cross-domain issue, resulting in the superiority of the method to conventional ones in both effectiveness and efficiency.
Zhen-Qun Yang, Xiaoyong Wei, Zhang Yi 0001, Gerald Friedland
IEEE Trans. Image Process.3
2017 High-Order Measurements for Residual Classifiers
abstract
Residual classifiers are common in dictionary-based multiclass classification. This paper proposes the concept of performance functions for residual classifiers. A performance function for multiclass classifications is a conceptual measurement function that combines local and global measurements. In general, the performance function is nonlinear. To explore the properties of the performance function, we employ the Taylor series expansion technique and derive a family of measurement functions. Specifically, the linear measurement and the quadratic measurement (QM) are derived. By exploiting the effect of the higher order terms in the performance function as well as the fundamental nondecreasing constrain, we derive the normalized QM (NQM). We present the classifier for multiclass classification using the proposed measurements. The proposed algorithms are tested against frontal faces and handwritten digit recognition tasks. Our tests show that the QM classifier achieves competitive classification results compared with baseline methods. NQM shows better stability with different parameter configurations.
Quan Guo, Haixian Zhang, Zhang Yi 0001
IEEE Trans. Neural Networks Learn. Syst.3
2017 Bag of Events: An Efficient Probability-Based Feature Extraction Method for AER Image Sensors
abstract
Address event representation (AER) image sensors represent the visual information as a sequence of events that denotes the luminance changes of the scene. In this paper, we introduce a feature extraction method for AER image sensors based on the probability theory, namely, bag of events (BOE). The proposed approach represents each object as the joint probability distribution of the concurrent events, and each event corresponds to a unique activated pixel of the AER sensor. The advantages of BOE include: 1) it is a statistical learning method and has a good interpretability in mathematics; 2) BOE can significantly reduce the effort to tune parameters for different data sets, because it only has one hyperparameter and is robust to the value of the parameter; 3) BOE is an online learning algorithm, which does not require the training data to be collected in advance; 4) BOE can achieve competitive results in real time for feature extraction (>275 frames/s and >120,000 events/s); and 5) the implementation complexity of BOE only involves some basic operations, e.g., addition and multiplication. This guarantees the hardware friendliness of our method. The experimental results on three popular AER databases (i.e., MNIST-dynamic vision sensor, Poker Card, and Posture) show that our method is remarkably faster than two recently proposed AER categorization systems while preserving a good classification accuracy.
Xi Peng 0001, Bo Zhao 0018, Rui Yan 0005, Huajin Tang, Zhang Yi 0001
IEEE Trans. Neural Networks Learn. Syst.5
2017 Fractional Hopfield Neural Networks: Fractional Dynamic Associative Recurrent Neural Networks
abstract
This paper mainly discusses a novel conceptual framework: fractional Hopfield neural networks (FHNN). As is commonly known, fractional calculus has been incorporated into artificial neural networks, mainly because of its long-term memory and nonlocality. Some researchers have made interesting attempts at fractional neural networks and gained competitive advantages over integer-order neural networks. Therefore, it is naturally makes one ponder how to generalize the first-order Hopfield neural networks to the fractional-order ones, and how to implement FHNN by means of fractional calculus. We propose to introduce a novel mathematical method: fractional calculus to implement FHNN. First, we implement fractor in the form of an analog circuit. Second, we implement FHNN by utilizing fractor and the fractional steepest descent approach, construct its Lyapunov function, and further analyze its attractors. Third, we perform experiments to analyze the stability and convergence of FHNN, and further discuss its applications to the defense against chip cloning attacks for anticounterfeiting. The main contribution of our work is to propose FHNN in the form of an analog circuit by utilizing a fractor and the fractional steepest descent approach, construct its Lyapunov function, prove its Lyapunov stability, analyze its attractors, and apply FHNN to the defense against chip cloning attacks for anticounterfeiting. A significant advantage of FHNN is that its attractors essentially relate to the neuron's fractional order. FHNN possesses the fractional-order-stability and fractional-order-sensitivity characteristics.
Yi-Fei Pu, Zhang Yi 0001, Jiliu Zhou
IEEE Trans. Neural Networks Learn. Syst.2
2017 Efficient Training of Supervised Spiking Neural Network via Accurate Synaptic-Efficiency Adjustment Method
abstract
The spiking neural network (SNN) is the third generation of neural networks and performs remarkably well in cognitive tasks, such as pattern recognition. The temporal neural encode mechanism found in biological hippocampus enables SNN to possess more powerful computation capability than networks with other encoding schemes. However, this temporal encoding approach requires neurons to process information serially on time, which reduces learning efficiency significantly. To keep the powerful computation capability of the temporal encoding mechanism and to overcome its low efficiency in the training of SNNs, a new training algorithm, the accurate synaptic-efficiency adjustment method is proposed in this paper. Inspired by the selective attention mechanism of the primate visual system, our algorithm selects only the target spike time as attention areas, and ignores voltage states of the untarget ones, resulting in a significant reduction of training time. Besides, our algorithm employs a cost function based on the voltage difference between the potential of the output neuron and the firing threshold of the SNN, instead of the traditional precise firing time distance. A normalized spike-timing-dependent-plasticity learning window is applied to assigning this error to different synapses for instructing their training. Comprehensive simulations are conducted to investigate the learning properties of our algorithm, with input neurons emitting both single spike and multiple spikes. Simulation results indicate that our algorithm possesses higher learning performance than the existing other methods and achieves the state-of-the-art efficiency in the training of SNN.
Xiurui Xie, Hong Qu 0002, Zhang Yi 0001, Jürgen Kurths
IEEE Trans. Neural Networks Learn. Syst.3
2017 Parameter as a Switch Between Dynamical States of a Network in Population Decoding
abstract
Population coding is a method to represent stimuli using the collective activities of a number of neurons. Nevertheless, it is difficult to extract information from these population codes with the noise inherent in neuronal responses. Moreover, it is a challenge to identify the right parameter of the decoding model, which plays a key role for convergence. To address the problem, a population decoding model is proposed for parameter selection. Our method successfully identified the key conditions for a nonzero continuous attractor. Both the theoretical analysis and the application studies demonstrate the correctness and effectiveness of this strategy.
Hua Mao 0001, Zhang Yi 0001
IEEE Trans. Neural Networks Learn. Syst.3
2017 Underdetermined Blind Source Separation Using Sparse Coding
abstract
In an underdetermined mixture system with unknown sources, it is a challenging task to separate these sources from their observed mixture signals, where . By exploiting the technique of sparse coding, we propose an effective approach to discover some 1-D subspaces from the set consisting of all the time-frequency (TF) representation vectors of observed mixture signals. We show that these 1-D subspaces are associated with TF points where only single source possesses dominant energy. By grouping the vectors in these subspaces via hierarchical clustering algorithm, we obtain the estimation of the mixing matrix. Finally, the source signals could be recovered by solving a series of least squares problems. Since the sparse coding strategy considers the linear representation relations among all the TF representation vectors of mixing signals, the proposed algorithm can provide an accurate estimation of the mixing matrix and is robust to the noises compared with the existing underdetermined blind source separation approaches. Theoretical analysis and experimental results demonstrate the effectiveness of the proposed method.
Liangli Zhen, Dezhong Peng, Zhang Yi 0001, Yong Xiang 0001, Peng Chen 0007
IEEE Trans. Neural Networks Learn. Syst.3
2016 A fuzzy c-means clustering based tournament selection for multiobjective optimization
abstract
This paper proposes a fuzzy c-means clustering based evolutionary algorithm called FCEA to optimize multiobjective optimization problems. FCEA firstly employs a fuzzy c-means clustering method (FCM) to discover the population distribution structure and to obtain a membership matrix of the population at each generation. Afterward, a membership based tournament selection (MBTS) operator is designed to select parents for recombination and to guide search. Comparison experiments show that the proposed FCEA outperforms MOEA/D-DE, NSGAII, SPEA2, RM-MEDA and SMS-EMOA on solving multiobjective optimization problems with complicated PF shapes. The experiments also present that MBTS significantly contributes to the performance of FCEA.
Zhang Yi 0001, Yu Zhen, Zi-mu Lu, Tong-Tong Lu
CEC1
2016 Manifold dimension reduction based clustering for multi-objective evolutionary algorithm
abstract
Real world optimization problems always possess multiple objectives which are conflict in nature. Multi-objective evolutionary algorithms (MOEAs), which provide a group of solutions in region of Pareto front, increasingly draw researchers attention for their excellent performance. In this regard, solutions with a wide diversity would be more favored as they give decision makers more choices to evaluate upon their problems. Based on the insight of investigating the evolution, the Pareto front often lies in a manifold space, not Euclidian space. However, most MOEAs utilize Euclidian distance as a sole mechanism to keep a wide range of diversity for solutions, which is not suitable somewhat from this aspect. To this end, manifold dimension reduction algorithm which has the ability to map solutions in the same front of objective space into Euclidian space is adapted in further. And then, general clustering algorithm are utilized. At the end, we use this technology to replace the crowding distance technology in NSGA-II to choose individuals when there is not enough slots in mating selection process. Based on a range of experiments over benchmark problems against state-of-the-art, it is fully expected benefit of performance improvement will be more significant when applied in many objectives optimization problems. This will be pursuit in our future study.
Yanan Sun 0001, Gary G. Yen, Hua Mao 0001, Zhang Yi 0001
CEC4
2016 Deep Subspace Clustering with Sparsity Prior
Xi Peng 0001, Shijie Xiao, Jiashi Feng, Weiyun Yau, Zhang Yi 0001
IJCAI5
2016 Symmetric low-rank representation for subspace clustering
Jie Chen 0065, Haixian Zhang, Hua Mao 0001, Yongsheng Sang, Zhang Yi 0001
Neurocomputing5
2016 A Novel Low Rank Representation Algorithm for Subspace Clustering
abstract
Low rank representation (LRR) is widely used to construct a good affinity matrix to cluster data drawn from the union of multiple linear subspaces. However, it is not easy to solve the LRR problem in a closed form, and augmented Lagrange multiplier method (ALM) is usually applied. ALM takes a relative long time dealing with the real-world data. To solve the LRR problem efficiently, we propose an efficient low rank representation (eLRR) algorithm. Given a contaminated data set, we propose a novel way to solve the LRR of the data. We establish a useful theorem which directly gives an approximate solution to our LRR optimization problem. Thus, we can construct a good affinity matrix for subspace clustering. Experimental results with several public databases verify the efficiency and effectiveness of our method.
Yuanyuan Chen 0006, Lei Zhang 0005, Zhang Yi 0001
Int. J. Pattern Recognit. Artif. Intell.3
2016 Double Gaussian mixture model for image segmentation with spatial relationships
Taisong Xiong, Lei Zhang 0005, Zhang Yi 0001
J. Vis. Commun. Image Represent.3
2016 Learning robust uniform features for cross-media social data by using cross autoencoders
Quan Guo, Jia Jia 0001, Guangyao Shen, Lei Zhang 0005, Lianhong Cai, Zhang Yi 0001
Knowl. Based Syst.6
2016 Matrix approach to decision-theoretic rough sets for evolving data
Chuan Luo 0001, Tianrui Li 0001, Zhang Yi 0001, Hamido Fujita
Knowl. Based Syst.3
2016 Efficient updating of probabilistic approximations with incremental objects
Chuan Luo 0001, Tianrui Li 0001, Hongmei Chen 0001, Hamido Fujita, Zhang Yi 0001
Knowl. Based Syst.5
2016 Shortest path computation using pulse-coupled neural networks with restricted autowave
Yongsheng Sang, Jiancheng Lv 0001, Hong Qu 0002, Zhang Yi 0001
Knowl. Based Syst.4
2016 A non-negative representation learning algorithm for selecting neighbors
Jiancheng Lv 0001, Zhang Yi 0001
Mach. Learn.3
2016 Learning a good representation with unsymmetrical auto-encoder
Yanan Sun 0001, Hua Mao 0001, Quan Guo, Zhang Yi 0001
Neural Comput. Appl.4
2016 Evolving Scale-Free Networks by Poisson Process: Modeling and Degree Distribution
abstract
Since the great mathematician Leonhard Euler initiated the study of graph theory, the network has been one of the most significant research subject in multidisciplinary. In recent years, the proposition of the small-world and scale-free properties of complex networks in statistical physics made the network science intriguing again for many researchers. One of the challenges of the network science is to propose rational models for complex networks. In this paper, in order to reveal the influence of the vertex generating mechanism of complex networks, we propose three novel models based on the homogeneous Poisson, nonhomogeneous Poisson and birth death process, respectively, which can be regarded as typical scale-free networks and utilized to simulate practical networks. The degree distribution and exponent are analyzed and explained in mathematics by different approaches. In the simulation, we display the modeling process, the degree distribution of empirical data by statistical methods, and reliability of proposed networks, results show our models follow the features of typical complex networks. Finally, some future challenges for complex systems are discussed.
Minyu Feng, Hong Qu 0002, Zhang Yi 0001, Xiurui Xie, Jürgen Kurths
IEEE Trans. Cybern.3
2016 A Unified Framework for Representation-Based Subspace Clustering of Out-of-Sample and Large-Scale Data
abstract
-norm-based representation, and have achieved the state-of-the-art performance. However, these methods have suffered from the following two limitations. First, the time complexities of these methods are at least proportional to the cube of the data size, which make those methods inefficient for solving the large-scale problems. Second, they cannot cope with the out-of-sample data that are not used to construct the similarity graph. To cluster each out-of-sample datum, the methods have to recalculate the similarity graph and the cluster membership of the whole data set. In this paper, we propose a unified framework that makes the representation-based subspace clustering algorithms feasible to cluster both the out-of-sample and the large-scale data. Under our framework, the large-scale problem is tackled by converting it as the out-of-sample problem in the manner of sampling, clustering, coding, and classifying. Furthermore, we give an estimation for the error bounds by treating each subspace as a point in a hyperspace. Extensive experimental results on various benchmark data sets show that our methods outperform several recently proposed scalable methods in clustering a large-scale data set.
Xi Peng 0001, Huajin Tang, Lei Zhang 0005, Zhang Yi 0001, Shijie Xiao
IEEE Trans. Neural Networks Learn. Syst.4
2015 Robust Subspace Clustering via Thresholding Ridge Regression
abstract
Given a data set from a union of multiple linear subspaces, a robust subspace clustering algorithm fits each group of data points with a low-dimensional subspace and then clusters these data even though they are grossly corrupted or sampled from the union of dependent subspaces. Under the framework of spectral clustering, recent works using sparse representation, low rank representation and their extensions achieve robust clustering results by formulating the errors (e.g., corruptions) into their objective functions so that the errors can be removed from the inputs. However, these approaches have suffered from the limitation that the structure of the errors should be known as the prior knowledge. In this paper, we present a new method of robust subspace clustering by eliminating the effect of the errors from the projection space (representation) rather than from the input space. We firstly prove that ell_1-, ell_2-, and ell_infty-norm-based linear projection spaces share the property of intra-subspace projection dominance, i.e., the coefficients over intra-subspace data points are larger than those over inter-subspace data points. Based on this property, we propose a robust and efficient subspace clustering algorithm, called Thresholding Ridge Regression (TRR). TRR calculates the ell2-norm-based coefficients of a given data set and performs a hard thresholding operator; and then the coefficients are used to build a similarity graph for clustering. Experimental studies show that TRR outperforms the state-of-the-art methods with respect to clustering quality, robustness, and time-saving.
Xi Peng 0001, Zhang Yi 0001, Huajin Tang
AAAI2
2015 Fast low rank representation based spatial pyramid matching for image classification
Xi Peng 0001, Rui Yan 0005, Bo Zhao 0018, Huajin Tang, Zhang Yi 0001
Knowl. Based Syst.5
2015 Constructing L1-graphs for subspace learning via recurrent neural networks
Yin Kuang, Lei Zhang 0005, Zhang Yi 0001
Pattern Anal. Appl.3
2015 Robust classifier using distance-based representation with square weights
Jiangshu Wei, Jiancheng Lv 0001, Zhang Yi 0001
Soft Comput.3
2015 A Class of Manifold Regularized Multiplicative Update Algorithms for Image Clustering
abstract
Multiplicative update algorithms are important tools for information retrieval, image processing, and pattern recognition. However, when the graph regularization is added to the cost function, different classes of sample data may be mapped to the same subspace, which leads to the increase of data clustering error rate. In this paper, an improved nonnegative matrix factorization (NMF) cost function is introduced. Based on the cost function, a class of novel graph regularized NMF algorithms is developed, which results in a class of extended multiplicative update algorithms with manifold structure regularization. Analysis shows that in the learning, the proposed algorithms can efficiently minimize the rank of the data representation matrix. Theoretical results presented in this paper are confirmed by simulations. For different initializations and data sets, variation curves of cost functions and decomposition data are presented to show the convergence features of the proposed update rules. Basis images, reconstructed images, and clustering results are utilized to present the efficiency of the new algorithms. Last, the clustering accuracies of different algorithms are also investigated, which shows that the proposed algorithms can achieve state-of-the-art performance in applications of image clustering.
Shangming Yang, Zhang Yi 0001, Xiaofei He 0001, Xuelong Li 0001
IEEE Trans. Image Process.2
2015 Non-Divergence of Stochastic Discrete Time Algorithms for PCA Neural Networks
abstract
Learning algorithms play an important role in the practical application of neural networks based on principal component analysis, often determining the success, or otherwise, of these applications. These algorithms cannot be divergent, but it is very difficult to directly study their convergence properties, because they are described by stochastic discrete time (SDT) algorithms. This brief analyzes the original SDT algorithms directly, and derives some invariant sets that guarantee the nondivergence of these algorithms in a stochastic environment by selecting proper learning parameters. Our theoretical results are verified by a series of simulation examples.
Jiancheng Lv 0001, Zhang Yi 0001, Yunxia Li
IEEE Trans. Neural Networks Learn. Syst.2
2014 A Local Non-Negative Pursuit Method for Intrinsic Manifold Structure Preservation
abstract
The local neighborhood selection plays a crucial role for most representation based manifold learning algorithms. This paper reveals that an improper selection of neighborhood for learning representation will introduce negative components in the learnt representations. Importantly, the representations with negative components will affect the intrinsic manifold structure preservation. In this paper, a local non-negative pursuit (LNP) method is proposed for neighborhood selection and non-negative representations are learnt. Moreover, it is proved that the learnt representations are sparse and convex. Theoretical analysis and experimental results show that the proposed method achieves or outperforms the state-of-the-art results on various manifold learning problems.
Dongdong Chen 0004, Jiancheng Lv 0001, Zhang Yi 0001
AAAI3
2014 A computational cognition model of perception, memory, and judgment
Xiaolan Fu, Lianhong Cai, Ye Liu 0010, Jia Jia 0001, Zhang Yi 0001, Guozhen Zhao, Yong-Jin Liu 0001, Changxu Wu
Sci. China Inf. Sci.6
2014 Nearest convex hull classification by using Lotka-Volterra recurrent neural networks
Yuanyuan Chen 0006, Lei Zhang 0005, Zhang Yi 0001
Neurocomputing3
2014 Sparse representation for face recognition by discriminative low-rank matrix recovery
Jie Chen 0065, Zhang Yi 0001
J. Vis. Commun. Image Represent.2
2014 Grayscale image segmentation by spatially variant mixture model with student's t-distribution
Taisong Xiong, Zhang Yi 0001, Lei Zhang 0005
Multim. Tools Appl.2
2014 Support vector set selection using pulse-coupled neural networks
Yunxia Li, Zhang Yi 0001, Jiancheng Lv 0001
Neural Comput. Appl.2
2014 Robust t-distribution mixture modeling via spatially directional information
Taisong Xiong, Lei Zhang 0005, Zhang Yi 0001
Neural Comput. Appl.3
2014 Learning locality-constrained collaborative representation for robust face recognition
Xi Peng 0001, Lei Zhang 0005, Zhang Yi 0001, Kok Kiong Tan
Pattern Recognit.3
2014 An adaptive rank-sparsity K-SVD algorithm for image sequence denoising
Yin Kuang, Lei Zhang 0005, Zhang Yi 0001
Pattern Recognit. Lett.3
2014 Convergence Analysis of Graph Regularized Non-Negative Matrix Factorization
abstract
Graph regularized non-negative matrix factorization (NMF) algorithms can be applied to information retrieval, image processing, and pattern recognition. However, challenge that still remains is to prove the convergence of this class of learning algorithms since the geometrical structure of the data space is considered. This paper presents the convergence properties of the graph regularized NMF learning algorithms. In the analysis, we focus on the study of Euclidian distance based algorithms. The structures of the fixed points are presented. The non-divergence of the learning algorithms is analyzed by constructing invariant sets for update rules. Based on Lyapunov indirect method, the stability of the algorithms is discussed in detail. The analysis shows that this class of NMF algorithms can converge to their fixed points under some given conditions. In the simulations, theoretical results presented in the paper are confirmed. For different initializations and data sets, variations of cost functions and decomposition data in the learning are presented to show the convergence features of the discussed NMF update rules, and the convergence speed of the algorithms is also investigated.
Shangming Yang, Zhang Yi 0001, Mao Ye 0001, Xiaofei He 0001
IEEE Trans. Knowl. Data Eng.2
2013 Scalable Sparse Subspace Clustering
abstract
In this paper, we address two problems in Sparse Subspace Clustering algorithm (SSC), i.e., scalability issue and out-of-sample problem. SSC constructs a sparse similarity graph for spectral clustering by using l1-minimization based coefficients, has achieved state-of-the-art results for image clustering and motion segmentation. However, the time complexity of SSC is proportion to the cubic of problem size such that it is inefficient to apply SSC into large scale setting. Moreover, SSC does not handle with out-of-sample data that are not used to construct the similarity graph. For each new datum, SSC needs recalculating the cluster membership of the whole data set, which makes SSC is not competitive in fast online clustering. To address the problems, this paper proposes out-of-sample extension of SSC, named as Scalable Sparse Subspace Clustering (SSSC), which makes SSC feasible to cluster large scale data sets. The solution of SSSC adopts a "sampling, clustering, coding, and classifying" strategy. Extensive experimental results on several popular data sets demonstrate the effectiveness and efficiency of our method comparing with the state-of-the-art algorithms.
Xi Peng 0001, Lei Zhang 0005, Zhang Yi 0001
CVPR3
2013 Locality-Based Discriminant Neighborhood Embedding
abstract
In this article, we develop a linear supervised subspace learning method called locality-based discriminant neighborhood embedding (LDNE), which can take advantage of the underlying submanifold-based structures of the data for classification. Our LDNE method can simultaneously consider both ‘locality ’ of locality preserving projection (LPP) and ‘discrimination ’ of discriminant neighborhood embedding (DNE) in manifold learning. It can find an embedding that not only preserves local information to explore the intrinsic submanifold structure of data from the same class, but also enhances the discrimination among submanifolds from different classes. To investigate the performance of LDNE, we compare it with the state-of-the-art dimensionality reduction techniques such as LPP and DNE on publicly available datasets. Experimental results show that our LDNE can be an effective and robust method for classification.
Jianping Gou, Zhang Yi 0001
Comput. J.2
2013 Free-gram phrase identification for modeling Chinese text
Xi Peng 0001, Zhang Yi 0001, Xiaoyong Wei, Dezhong Peng, Yongsheng Sang
Inf. Process. Lett.2
2013 An improved code selection algorithm for fault prediction
Yin Kuang, Zhang Yi 0001, Lei Zhang 0005
Neural Comput. Appl.2
2013 Using competitive layer model implemented by Lotka-Volterra recurrent neural networks for detecting brain activated regions from fMRI data
Bochuan Zheng, Zhang Yi 0001
Neural Comput. Appl.2
2013 Collaborative neighbor representation based classification using l2-minimization approach
Waqas Jadoon, Zhang Yi 0001, Lei Zhang 0005
Pattern Recognit. Lett.2
2013 Efficient Shortest-Path-Tree Computation in Network Routing Based on Pulse-Coupled Neural Networks
abstract
Shortest path tree (SPT) computation is a critical issue for routers using link-state routing protocols, such as the most commonly used open shortest path first and intermediate system to intermediate system. Each router needs to recompute a new SPT rooted from itself whenever a change happens in the link state. Most commercial routers do this computation by deleting the current SPT and building a new one using static algorithms such as the Dijkstra algorithm at the beginning. Such recomputation of an entire SPT is inefficient, which may consume a considerable amount of CPU time and result in a time delay in the network. Some dynamic updating methods using the information in the updated SPT have been proposed in recent years. However, there are still many limitations in those dynamic algorithms. In this paper, a new modified model of pulse-coupled neural networks (M-PCNNs) is proposed for the SPT computation. It is rigorously proved that the proposed model is capable of solving some optimization problems, such as the SPT. A static algorithm is proposed based on the M-PCNNs to compute the SPT efficiently for large-scale problems. In addition, a dynamic algorithm that makes use of the structure of the previously computed SPT is proposed, which significantly improves the efficiency of the algorithm. Simulation results demonstrate the effective and efficient performance of the proposed approach.
Hong Qu 0002, Zhang Yi 0001, Simon X. Yang
IEEE Trans. Cybern.2
2012 A Local Mean-Based k-Nearest Centroid Neighbor Classifier
abstract
K-nearest neighbor (KNN) rule is a simple and effective algorithm in pattern classification. In this article, we propose a local mean-based k-nearest centroid neighbor classifier that assigns to each query pattern a class label with nearest local centroid mean vector so as to improve the classification performance. The proposed scheme not only takes into account the proximity and spatial distribution of k neighbors, but also utilizes the local mean vector of k neighbors from each class in making classification decision. In the proposed classifier, a local mean vector of k nearest centroid neighbors from each class for a query pattern is well positioned to sufficiently capture the class distribution information. In order to investigate the classification behavior of the proposed classifier, we conduct extensive experiments on the real and synthetic data sets in terms of the classification error. Experimental results demonstrate that our proposed method performs significantly well, particularly in the small sample size cases, compared with the state-of-the-art KNN-based algorithms.
Jianping Gou, Zhang Yi 0001, Lan Du 0002, Taisong Xiong
Comput. J.2
2012 A Globally Convergent MC Algorithm With an Adaptive Learning Rate
abstract
This brief deals with the problem of minor component analysis (MCA). Artificial neural networks can be exploited to achieve the task of MCA. Recent research works show that convergence of neural networks based MCA algorithms can be guaranteed if the learning rates are less than certain thresholds. However, the computation of these thresholds needs information about the eigenvalues of the autocorrelation matrix of data set, which is unavailable in online extraction of minor component from input data stream. In this correspondence, we introduce an adaptive learning rate into the OJAn MCA algorithm, such that its convergence condition does not depend on any unobtainable information, and can be easily satisfied in practical applications.
Dezhong Peng, Zhang Yi 0001, Yong Xiang 0001, Haixian Zhang
IEEE Trans. Neural Networks Learn. Syst.2
2011 Selectable and Unselectable Sets of Neurons in Recurrent Neural Networks With Saturated Piecewise Linear Transfer Function
abstract
The concepts of selectable and unselectable sets are proposed to describe some interesting dynamical properties of a class of recurrent neural networks (RNNs) with saturated piecewise linear transfer function. A set of neurons is said to be selectable if it can be co-unsaturated at a stable equilibrium point by some external input. A set of neurons is said to be unselectable if it is not selectable, i.e., such set of neurons can never be co-unsaturated at any stable equilibrium point regardless of what the input is. The importance of such concepts is that they enable a new perspective of the memory in RNNs. Necessary and sufficient conditions for the existence of selectable and unselectable sets of neurons are obtained. As an application, the problem of group selection is discussed by using such concepts. It shows that, under some conditions, each group is a selectable set, and each selectable set is contained in some group. Thus, groups are indicated by selectable sets of the RNNs and can be selected by external inputs. Simulations are carried out to further illustrate the theory.
Lei Zhang 0005, Zhang Yi 0001
IEEE Trans. Neural Networks2
2010 A Method for MRI Segmentation of Brain Tissue
Bochuan Zheng, Zhang Yi 0001
ISNN (2)2
2010 SCALE: a scalable framework for efficiently clustering transactional data
Keke Chen, Ling Liu 0001, Zhang Yi 0001
Data Min. Knowl. Discov.4
2010 Spatial Point-Data Reduction Using Pulse Coupled Neural Network
Yongsheng Sang, Zhang Yi 0001, Jiliu Zhou
Neural Process. Lett.2
2010 Convergence Analysis of Non-Negative Matrix Factorization for BSS Algorithm
Shangming Yang, Zhang Yi 0001
Neural Process. Lett.2
2010 Adaptive multiple minor directions extraction in parallel using a PCA neural network
Kok Kiong Tan, Jiancheng Lv 0001, Zhang Yi 0001, Sunan Huang 0001
Theor. Comput. Sci.3
2010 A Family of Fuzzy Learning Algorithms for Robust Principal Component Analysis Neural Networks
abstract
In this paper, we analyze Xu and Yuille’s robust principal component analysis (RPCA) learning algorithms by means of the distance measurement in space. Based on the analysis, a family of fuzzy RPCA learning algorithms is proposed, which is robust against outliers. These algorithms can explicitly be understood from the viewpoint of fuzzy set theory, though Xu and Yuille’s algorithms were proposed based on a statistical physics approach. In the proposed algorithms, an adaptive learning procedure overcomes the difficulty of selection of learning parameters in Xu and Yuille’s algorithms. Furthermore, the robustness of proposed algorithms is investigated by using the theory of influence functions. Simulations are carried out to illustrate the robustness of these algorithms.
Jiancheng Lv 0001, Kok Kiong Tan, Zhang Yi 0001, Sunan Huang 0001
IEEE Trans. Fuzzy Syst.3
2010 A discrete-time neural network for optimization problems with hybrid constraints
abstract
Recurrent neural networks have become a prominent tool for optimizations including linear or nonlinear variational inequalities and programming, due to its regular mathematical properties and well-defined parallel structure. This brief presents a general discrete-time recurrent network for linear variational inequalities and related optimization problems with hybrid constraints. In contrary to the existing discrete-time networks, this general model can operate not only on bound constraints, but also on hybrid constraints comprised of inequality, equality and bound constraints. The model has dynamical properties of global convergence, asymptotical and exponential convergences under some weaker conditions. Numerical examples demonstrate its efficacy and performance.
Huajin Tang, Haizhou Li 0001, Zhang Yi 0001
IEEE Trans. Neural Networks3
2010 Continuous attractors of Lotka-Volterra recurrent neural networks with infinite neurons
abstract
Continuous attractors of Lotka-Volterra recurrent neural networks (LV RNNs) with infinite neurons are studied in this brief. A continuous attractor is a collection of connected equilibria, and it has been recognized as a suitable model for describing the encoding of continuous stimuli in neural networks. The existence of the continuous attractors depends on many factors such as the connectivity and the external inputs of the network. A continuous attractor can be stable or unstable. It is shown in this brief that a LV RNN can possess multiple continuous attractors if the synaptic connections and the external inputs are Gussian-like in shape. Moreover, both stable and unstable continuous attractors can coexist in a network. Explicit expressions of the continuous attractors are calculated. Simulations are employed to illustrate the theory.
Zhang Yi 0001, Jiliu Zhou
IEEE Trans. Neural Networks2
2009 Solving the CLM Problem by Discrete-Time Linear Threshold Recurrent Neural Networks
Lei Zhang 0005, Pheng-Ann Heng, Zhang Yi 0001
ICANN (1)3
2009 Continuous Attractors of Lotka-Volterra Recurrent Neural Networks
Haixian Zhang, Zhang Yi 0001
ICANN (1)3
2009 Constrained ZIP code segmentation by a PCNN-based thinning algorithm
Lifeng Shang, Zhang Yi 0001, Luping Ji
Neurocomputing2
2009 A modified pulse coupled neural network for shortest-path problem
XiaoBin Wang, Hong Qu 0002, Zhang Yi 0001
Neurocomputing3
2009 Some multistability properties of bidirectional associative memory recurrent neural networks with unsaturating piecewise linear transfer functions
Lei Zhang 0005, Zhang Yi 0001, Pheng-Ann Heng
Neurocomputing2
2009 A New Incremental PCA Algorithm With Application to Visual Learning and Recognition
Zhang Yi 0001, Xiaorong Pu
Neural Process. Lett.2
2009 A Winner-Take-All Neural Networks of N Linear Threshold Neurons without Self-Excitatory Connections
Hong Qu 0002, Zhang Yi 0001, XiaoBin Wang
Neural Process. Lett.2
2009 Real-Time Robot Path Planning Based on a Modified Pulse-Coupled Neural Network Model
abstract
This paper presents a modified pulse-coupled neural network (MPCNN) model for real-time collision-free path planning of mobile robots in nonstationary environments. The proposed neural network for robots is topologically organized with only local lateral connections among neurons. It works in dynamic environments and requires no prior knowledge of target or barrier movements. The target neuron fires first, and then the firing event spreads out, through the lateral connections among the neurons, like the propagation of a wave. Obstacles have no connections to their neighbors. Each neuron records its parent, that is, the neighbor that caused it to fire. The real-time optimal path is then the sequence of parents from the robot to the target. In a static case where the barriers and targets are stationary, this paper proves that the generated wave in the network spreads outward with travel times proportional to the linking strength among neurons. Thus, the generated path is always the global shortest path from the robot to the target. In addition, each neuron in the proposed model can propagate a firing event to its neighboring neuron without any comparing computations. The proposed model is applied to generate collision-free paths for a mobile robot to solve a maze-type problem, to circumvent concave U-shaped obstacles, and to track a moving target in an environment with varying obstacles. The effectiveness and efficiency of the proposed approach is demonstrated through simulation and comparison studies.
Hong Qu 0002, Simon X. Yang, Allan R. Willms, Zhang Yi 0001
IEEE Trans. Neural Networks4
2009 Permitted and Forbidden Sets in Discrete-Time Linear Threshold Recurrent Neural Networks
abstract
The concepts of permitted and forbidden sets enable a new perspective of the memory in neural networks. Such concepts exhibit interesting dynamics in recurrent neural networks. This paper studies the basic theories of permitted and forbidden sets of the linear threshold discrete-time recurrent neural networks. The linear threshold transfer function has been regarded as an adequate transfer function for recurrent neural networks. Networks with this transfer function form a class of hybrid analog and digital networks which are especially useful for perceptual computations. Networks in discrete time can directly provide algorithms for efficient implementation in digital hardware. The main contribution of this paper is to establish foundations of permitted and forbidden sets. Necessary and sufficient conditions for the linear threshold discrete-time recurrent neural networks are obtained for complete convergence, existence of permitted and forbidden sets, as well as conditionally multiattractivity, respectively. Simulation studies explore some possible interesting practical applications.
Zhang Yi 0001, Lei Zhang 0005, Kok Kiong Tan
IEEE Trans. Neural Networks1
2009 Representations of Continuous Attractors of Recurrent Neural Networks
abstract
A continuous attractor of a recurrent neural network (RNN) is a set of connected stable equilibrium points. Continuous attractors have been used to describe the encoding of continuous stimuli in neural networks. Dynamic behaviors of continuous attractors of RNNs exhibit interesting properties. This brief desires to derive explicit representations of continuous attractors of RNNs. Representations of continuous attractors of linear RNNs as well as linear-threshold (LT) RNNs are obtained under some conditions. These representations could be looked at as solutions of continuous attractors of the networks. Such results provide clear and complete descriptions to the continuous attractors.
Zhang Yi 0001, Lei Zhang 0005
IEEE Trans. Neural Networks2
2009 Manifold-Based Learning and Synthesis
abstract
This paper proposes a new approach to analyze high-dimensional data set using low-dimensional manifold. This manifold-based approach provides a unified formulation for both learning from and synthesis back to the input space. The manifold learning method desires to solve two problems in many existing algorithms. The first problem is the local manifold distortion caused by the cost averaging of the global cost optimization during the manifold learning. The second problem results from the unit variance constraint generally used in those spectral embedding methods where global metric information is lost. For the out-of-sample data points, the proposed approach gives simple solutions to transverse between the input space and the feature space. In addition, this method can be used to estimate the underlying dimension and is robust to the number of neighbors. Experiments on both low-dimensional data and real image data are performed to illustrate the theory.
Zhang Yi 0001, Xiaorong Pu
IEEE Trans. Syst. Man Cybern. Part B2
2008 Shape recovery by a generalized topology preserving SOM
Zhang Yi 0001
Neurocomputing2
2008 A new local PCA-SOM algorithm
Zhang Yi 0001, Xiaorong Pu
Neurocomputing2
2008 A mixed noise image filtering method using weighted-linking PCNNs
Luping Ji, Zhang Yi 0001
Neurocomputing2
2008 An algorithm for extracting fetal electrocardiogram
Yunxia Li, Zhang Yi 0001
Neurocomputing2
2008 A neural networks learning algorithm for minor component analysis and its convergence analysis
Dezhong Peng, Zhang Yi 0001, Jiancheng Lv 0001, Yong Xiang 0001
Neurocomputing2
2008 Switching analysis of 2-D neural networks with nonsaturating linear threshold transfer functions
Hong Qu 0002, Zhang Yi 0001, XiaoBin Wang
Neurocomputing2
2008 A tabu search approach for the minimum sum-of-squares clustering problem
Yongguo Liu, Zhang Yi 0001, Mao Ye 0001, Kefei Chen
Inf. Sci.2
2008 An improved pulse coupled neural network for image processing
Luping Ji, Zhang Yi 0001, Lifeng Shang
Neural Comput. Appl.2
2008 Holistic and partial facial features fusion by binary particle swarm optimization
Xiaorong Pu, Zhang Yi 0001, Zhongjie Fang
Neural Comput. Appl.2
2008 Stability and Chaos of a Class of Learning Algorithms for ICA Neural Networks
Jiancheng Lv 0001, Kok Kiong Tan, Zhang Yi 0001, Sunan Huang 0001
Neural Process. Lett.3
2008 Fingerprint orientation field estimation using ridge projection
Luping Ji, Zhang Yi 0001
Pattern Recognit.2
2008 Document clustering using locality preserving indexing and support vector machines
Chengfu Yang, Zhang Yi 0001
Soft Comput.2
2008 Multiperiodicity and Attractivity of Delayed Recurrent Neural Networks With Unsaturating Piecewise Linear Transfer Functions
abstract
This paper studies multiperiodicity and attractivity for a class of recurrent neural networks (RNNs) with unsaturating piecewise linear transfer functions and variable delays. Using local inhibition, conditions for boundedness and global attractivity are established. These conditions allow coexistence of stable and unstable trajectories. Moreover, multiperiodicity of the network is investigated by using local invariant sets. It shows that under some interesting conditions, there exists one periodic trajectory in each invariant set which exponentially attracts all trajectories in that region correspondingly. Simulations are carried out to illustrate the theories.
Lei Zhang 0005, Zhang Yi 0001
IEEE Trans. Neural Networks2
2007 Sequential Blind Signal Extraction with the Linear Predictor
Yunxia Li, Zhang Yi 0001
ISNN (3)2
2007 Global Exponential Convergence of Time-Varying Delayed Neural Networks with High Gain
Lei Zhang 0005, Zhang Yi 0001
ISNN (1)2
2007 A hierarchical intrusion detection model based on the PCA neural networks
Guisong Liu, Zhang Yi 0001, Shangming Yang
Neurocomputing2
2007 A class of binary images thinning using two PCNNs
Lifeng Shang, Zhang Yi 0001
Neurocomputing2
2007 Convergence analysis of a simple minor component analysis algorithm
Dezhong Peng, Zhang Yi 0001, Wenjing Luo
Neural Networks2
2007 Binary Image Thinning Using Autowaves Generated by PCNN
Lifeng Shang, Zhang Yi 0001, Luping Ji
Neural Process. Lett.2
2007 Convergence analysis of the OJAn MCA learning algorithm by the deterministic discrete time method
Dezhong Peng, Zhang Yi 0001
Theor. Comput. Sci.2
2007 Determination of the Number of Principal Directions in a Biologically Plausible PCA Model
abstract
Adaptively determining an appropriate number of principal directions for principal component analysis (PCA) neural networks is an important problem to address when one uses PCA neural networks for online feature extraction. In this letter, inspired from biological neural networks, a single-layer neural network model with lateral connections is proposed which uses an improved generalized Hebbian algorithm (GHA) to address this problem. In the proposed model, the number of principal directions can be adaptively determined to approximate the intrinsic dimensionality of the given data set so that the dimensionality of the data set can be reduced to approach the intrinsic dimensionality to any required precision through the network.
Jiancheng Lv 0001, Zhang Yi 0001, Kok Kiong Tan
IEEE Trans. Neural Networks2
2007 Global Convergence of GHA Learning Algorithm With Nonzero-Approaching Adaptive Learning Rates
abstract
The generalized Hebbian algorithm (GHA) is one of the most widely used principal component analysis (PCA) neural network (NN) learning algorithms. Learning rates of GHA play important roles in convergence of the algorithm for applications. Traditionally, the learning rates of GHA are required to converge to zero so that its convergence can be analyzed by studying the corresponding deterministic continuous-time (DCT) equations. However, the requirement for learning rates to approach zero is not a practical one in applications due to computational roundoff limitations and tracking requirements. In this paper, nonzero-approaching adaptive learning rates are proposed to overcome this problem. These proposed adaptive learning rates converge to some positive constants, which not only speed up the algorithm evolution considerably, but also guarantee global convergence of the GHA algorithm. The convergence is studied in detail by analyzing the corresponding deterministic discrete-time (DDT) equations. Extensive simulations are carried out to illustrate the theory.
Jiancheng Lv 0001, Zhang Yi 0001, Kok Kiong Tan
IEEE Trans. Neural Networks2
2007 Dynamics of Generalized PCA and MCA Learning Algorithms
abstract
Principal component analysis (PCA) and minor component analysis (MCA) are two important statistical tools which have many applications in the fields of signal processing and data analysis. PCA and MCA neural networks (NNs) can be used to online extract principal component and minor component from input data. It is interesting to develop generalized learning algorithms of PCA and MCA NNs. Some novel generalized PCA and MCA learning algorithms are proposed in this paper. Convergence of PCA and MCA learning algorithms is an essential issue in practical applications. Traditionally, the convergence is studied via deterministic continuous-time (DCT) method. The DCT method requires the learning rate of the algorithms to approach to zero, which is not realistic in many practical applications. In this paper, deterministic discrete-time (DDT) method is used to study the dynamical behaviors of the proposed algorithms. The DDT method is more reasonable for the convergence analysis since it does not require constraints as that of the DCT method. It is proven that under some mild conditions, the weight vector in these proposed algorithms will converge exponentially to principal or minor component. Simulation results are further used to illustrate the theoretical results.
Dezhong Peng, Zhang Yi 0001
IEEE Trans. Neural Networks2
2007 Binary Fingerprint Image Thinning Using Template-Based PCNNs
abstract
This correspondence presents a coarse-to-fine binary-image-thinning algorithm by proposing a template-based pulse-coupled neural-network model. Under the control of coupled templates, this algorithm iteratively skeletonizes a binary image by changing the load signals of pulse neurons. A direction-constraining scheme for avoiding fingerprint ridge spikes has been discussed. Experiments show that this algorithm is effective for fingerprint thinning, as well as other common images. Moreover, this algorithm can be coupled with a fingerprint identification system to improve the recognition performance.
Luping Ji, Zhang Yi 0001, Lifeng Shang, Xiaorong Pu
IEEE Trans. Syst. Man Cybern. Part B2
2006 Intrusion Detection Using PCASOM Neural Networks
Guisong Liu, Zhang Yi 0001
ISNN (2)2
2006 Recognizing Partially Damaged Facial Images by Subspace Auto-associative Memories
Xiaorong Pu, Zhang Yi 0001
ISNN (2)2
2006 Convergence and Periodicity of Solutions for a Class of Discrete-Time Recurrent Neural Network with Two Neurons
Hong Qu 0002, Zhang Yi 0001
ISNN (1)2
2006 Growing Hierarchical Principal Components Analysis Self-Organizing Map
Stones Lei Zhang, Zhang Yi 0001, Jiancheng Lv 0001
ISNN (1)2
2006 Convergence analysis of Xu's LMSER learning algorithm via deterministic discrete time system method
Jiancheng Lv 0001, Zhang Yi 0001, Kok Kiong Tan
Neurocomputing2
2006 Rigid medical image registration using PCA neural network
Lifeng Shang, Jiancheng Lv 0001, Zhang Yi 0001
Neurocomputing3
2006 Fuzzy SVM with a new fuzzy membership function
Xiufeng Jiang, Zhang Yi 0001, Jiancheng Lv 0001
Neural Comput. Appl.2
2006 Global convergence of Oja's PCA learning algorithm with a non-zero-approaching adaptive learning rate
Jiancheng Lv 0001, Zhang Yi 0001, Kok Kiong Tan
Theor. Comput. Sci.2
2006 Output convergence analysis for a class of delayed recurrent neural networks with time-varying inputs
abstract
This paper studies the output convergence of a class of recurrent neural networks with time-varying inputs. The model of the studied neural networks has different dynamic structure from that in the well known Hopfield model, it does not contain linear terms. Since different structures of differential equations usually result in quite different dynamic behaviors, the convergence of this model is quite different from that of Hopfield model. This class of neural networks has been found many successful applications in solving some optimization problems. Some sufficient conditions to guarantee output convergence of the networks are derived.
Zhang Yi 0001, Jiancheng Lv 0001, Lei Zhang 0005
IEEE Trans. Syst. Man Cybern. Part B1
2005 Clustering Categorical Data Using Coverage Density
Lei Zhang 0005, Zhang Yi 0001
ADMA3
2005 An Improved Backpropagation Algorithm Using Absolute Error Function
Jiancheng Lv 0001, Zhang Yi 0001
ISNN (1)2
2005 A Modified MCA EXIN Algorithm and Its Convergence Analysis
Dezhong Peng, Zhang Yi 0001, XiaoLin Xiang
ISNN (1)2
2005 Face Recognition Using Fisher Non-negative Matrix Factorization with Sparseness Constraints
Xiaorong Pu, Zhang Yi 0001, Ziming Zheng, Mao Ye 0001
ISNN (2)2
2005 Theoretical Analysis and Parameter Setting of Hopfield Neural Networks
Hong Qu 0002, Zhang Yi 0001, XiaoLin Xiang
ISNN (1)2
2005 Fast ICA for Online Cashflow Analysis
Shangming Yang, Zhang Yi 0001
ISNN (2)2
2005 A globally convergent learning algorithm for PCA neural networks
Mao Ye 0001, Zhang Yi 0001, Jiancheng Lv 0001
Neural Comput. Appl.2
2005 Complete Convergence of Competitive Neural Networks with Different Time Scales
Mao Ye 0001, Zhang Yi 0001
Neural Process. Lett.2
2005 Convergence analysis of a deterministic discrete time system of Oja's PCA learning algorithm
abstract
The convergence of Oja's principal component analysis (PCA) learning algorithms is a difficult topic for direct study and analysis. Traditionally, the convergence of these algorithms is indirectly analyzed via certain deterministic continuous time (DCT) systems. Such a method will require the learning rate to converge to zero, which is not a reasonable requirement to impose in many practical applications. Recently, deterministic discrete time (DDT) systems have been proposed instead to indirectly interpret the dynamics of the learning algorithms. Unlike DCT systems, DDT systems allow learning rates to be constant (which can be a nonzero). This paper will provide some important results relating to the convergence of a DDT system of Oja's PCA learning algorithm. It has the following contributions: 1) A number of invariant sets are obtained, based on which we can show that any trajectory starting from a point in the invariant set will remain in the set forever. Thus, the nondivergence of the trajectories is guaranteed. 2) The convergence of the DDT system is analyzed rigorously. It is proven, in the paper, that almost all trajectories of the system starting from points in an invariant set will converge exponentially to the unit eigenvector associated with the largest eigenvalue of the correlation matrix. In addition, exponential convergence rate are obtained, providing useful guidelines for the selection of fast convergence learning rate. 3) Since the trajectories may diverge, the careful choice of initial vectors is an important issue. This paper suggests to use the domain of unit hyper sphere as initial vectors to guarantee convergence. 4) Simulation results will be furnished to illustrate the theoretical results achieved.
Zhang Yi 0001, Mao Ye 0001, Jiancheng Lv 0001, Kok Kiong Tan
IEEE Trans. Neural Networks1
2004 Convergence Analysis for Oja+ MCA Learning Algorithm
Jiancheng Lv 0001, Mao Ye 0001, Zhang Yi 0001
ISNN (1)3
2004 Self-Organizing Feature Map Based Data Mining
Shangming Yang, Zhang Yi 0001
ISNN (1)2
2004 On the Discrete Time Dynamics of the MCA Neural Networks
Mao Ye 0001, Zhang Yi 0001
ISNN (1)2
2004 Authors' reply to comment on "stability of fuzzy control systems with bounded uncertain delays"
abstract
For comment see ibid., p.285-6 (2004). The authors reply to a comment on their paper (Z. Yi and P.A. Heng, see ibid., vol.10, p.92-7, 2002) by agreeing with the comment.
Zhang Yi 0001, Pheng-Ann Heng
IEEE Trans. Fuzzy Syst.1
2004 A columnar competitive model for solving combinatorial optimization problems
abstract
The major drawbacks of the Hopfield network when it is applied to some combinatorial problems, e.g., the traveling salesman problem (TSP), are invalidity of the obtained solutions, trial-and-error setting value process of the network parameters and low-computation efficiency. This letter presents a columnar competitive model (CCM) which incorporates winner-takes-all (WTA) learning rule for solving the TSP. Theoretical analysis for the convergence of the CCM shows that the competitive computational neural network guarantees the convergence to valid states and avoids the onerous procedures of determining the penalty parameters. In addition, its intrinsic competitive learning mechanism enables a fast and effective evolving of the network. The simulation results illustrate that the competitive model offers more and better valid solutions as compared to the original Hopfield network.
Huajin Tang, Kay Chen Tan, Zhang Yi 0001
IEEE Trans. Neural Networks3
2004 Multistability of discrete-time recurrent neural networks with unsaturating piecewise linear activation functions
abstract
This paper studies the multistability of a class of discrete-time recurrent neural networks with unsaturating piecewise linear activation functions. It addresses the nondivergence, global attractivity, and complete stability of the networks. Using the local inhibition, conditions for nondivergence are derived, which not only guarantee nondivergence, but also allow for the existence of multiequilibrium points. Under these nondivergence conditions, global attractive compact sets are obtained. Complete stability is studied via constructing novel energy functions and using the well-known Cauchy Convergence Principle. Examples and simulation results are used to illustrate the theory.
Zhang Yi 0001, Kok Kiong Tan
IEEE Trans. Neural Networks1
2003 Multistability Analysis for Recurrent Neural Networks with Unsaturating Piecewise Linear Transfer Functions
abstract
Multistability is a property necessary in neural networks in order to enable certain applications (e.g., decision making), where monostable networks can be computationally restrictive. This article focuses on the analysis of multistability for a class of recurrent neural networks with unsaturating piecewise linear transfer functions. It deals fully with the three basic properties of a multistable network: boundedness, global attractivity, and complete convergence. This article makes the following contributions: conditions based on local inhibition are derived that guarantee boundedness of some multistable networks, conditions are established for global attractivity, bounds on global attractive sets are obtained, complete convergence conditions for the network are developed using novel energy-like functions, and simulation examples are employed to illustrate the theory thus developed.
Zhang Yi 0001, Kok Kiong Tan, Tong Heng Lee
Neural Comput.1
2002 Stability of fuzzy control systems with bounded uncertain delays
abstract
Global exponential stability of fuzzy control systems with delays is studied. These delays in the fuzzy control systems are assumed to be any uncertain bounded continuous functions. Stability of systems with uncertain delays is interesting since in practical applications it is not easy to know the delays exactly. Conditions for global exponential stability of free fuzzy systems with uncertain delays are derived. Criteria for design of nonlinear fuzzy controllers to feedback control the stability of global nonlinear fuzzy systems are given. Theorems are proved via the method of functional differential inequalities analysis.
Zhang Yi 0001, Pheng-Ann Heng
IEEE Trans. Fuzzy Syst.1
2000 Clustering Categorical Data
abstract
Clustering has typically been a problem related to numerical data. However, in databases, oftentimes the data values are categorical and cannot be assigned meaningful numerical substitutes. With the recent interest in data mining, we begin to question the possibility of clustering numerical data. Following some recent work in this area, we propose an algorithm based on dynamical systems. To our knowledge, this is the first such algorithm that can guarantee the convergence of the dynamical system, which is a very important property for successful application. We demonstrated the effectiveness of the proposed method on both real data and synthetic data. We also propose a second method based on a graph partitioning approach, for which a new definition of similarity between two nodes is tailored for categorical data. 1 Introduction Mining numerical data has received much attention in recent research in data mining. One important form of knowledge that can be derived from such da...
Zhang Yi 0001, Ada Wai-Chee Fu, Chun Hing Cai, Pheng-Ann Heng
ICDE1
1999 Estimate of exponential convergence rate and exponential stability for neural networks
abstract
Estimate of exponential convergence rate and exponential stability are studied for a class of neural networks which includes the Hopfield neural networks and the cellular neural networks. Both local and global exponential convergence is discussed. Theorems for estimate of exponential convergence rate are established and the bounds on the rate of convergence are given. The domains of attraction in the case of local exponential convergence are obtained. Simple conditions are presented for checking exponential stability of the neural networks.
Zhang Yi 0001, Pheng-Ann Heng, Ada Wai-Chee Fu
IEEE Trans. Neural Networks1