Jianxin Zhang 0001

dblp:86/1463-1 · DBLP profile ↗
← Back
48ranked-venue papers
3as first author
38since 2021 · last 2026
0000-0001-6076-5433ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 10 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Gracefully Air-Written: Enhancing the Legibility and Style Consistency of In-Air Handwriting
abstract
Space computing devices expand handwritten input from two-dimensional screens into three-dimensional space, providing an unrestricted interactive experience. Due to the high degree of freedom and lack of tactile feedback in in-air handwriting, handwritten characters not only become less legible but also lose the writer's personal style. This paper proposes a method for reconstructing discrete in-air handwriting using continuous diffusion models, capturing the writing process and style from a small number of user-provided handwritten tracks and images, to restore the legibility of characters and mimics the writer's style. We represent handwritten track data in binary form and model it with continuous diffusion models, recovering discrete handwritten track data through threshold processing. Our approach reconstructs in-air handwritten characters in two stages. During the content preservation phase, we propose a partial noise injection strategy based on reference sequence modeling, using the content information of the original character as a guiding condition to maintain content consistency in handwritten character. In the style aggregation phase, we adaptively fuse the visual style of the handwritten in the image modality with the dynamic writing process in the sequence modality, overcoming issues of insufficient style capture due to noise interference in the backward process. Qualitative and quantitative experiments demonstrate the superiority of our method.
Yu Liu 0072, Cunrui Wang, Jianxin Zhang 0001, Bo Lu 0005
AAAI4
2026 A Multi-Style Yi Script Font Generation Method Based on a Diffusion Model Incorporating Multi-Channel Attention Mechanism
abstract
The core of font generation technology lies in preserving character structure and accurately transferring stylistic features. While extensive research has been conducted on Chinese character generation, studies on Yi script font generation remain in their infancy. This paper, therefore, proposes a framework that employs a multi-scale conditional diffusion model to generate Yi script fonts. A Multi-scale Content Perception (MCP) module is designed. This module employs a constructed channel-space dual-domain multi-scale attention coordination mechanism to hierarchically capture multi-scale spatial and content information of characters within images. To achieve cross-modal feature fusion for Yi characters, an Efficient Style Insertion (ESI) module was developed. This module employs a cross-attention mechanism integrating multi-head mechanisms and Efficient Channel Attention (ECA) to enable cross-modal interaction of font style features. It further introduces block-wise attention to overcome computational redundancy issues inherent in traditional cross-attention. Finally, comparative experiments across Yi font datasets — including the Public Standardized Yi Font File (PSFF-Yi) dataset, Self-Built Handwritten Yi (SBHW-Yi) dataset, and Ancient Manuscript Handwritten Yi (AMHW-Yi) dataset — demonstrate the feasibility of this approach.
Zedong Li, Cunrui Wang, Jianxin Zhang 0001
Int. J. Pattern Recognit. Artif. Intell.5
2026 QHSP-Net: query-aware higher-order statistical pooling network for referring image segmentation
Qiule Sun, Jianxin Zhang 0001, Bingbing Zhang 0001, Peihua Li
Multim. Syst.2
2026 A Novel Asynchronous Intermittent Communication Methodology for Multi-Agent Systems With Unmodeled Disturbances
abstract
This paper investigates the consensus control problem of multi-agent systems under intermittent communication and unmodeled external disturbances. The main contribution is to overcome the limitation of current synchronous intermittent framework and propose a novel asynchronous intermittent communication methodology on multi-agent systems. In this intermittent methodology, the state space is divided into three distinct regions by introducing both safety and intermittent boundaries, which enables effective monitoring of agent error dynamics.Furthermore, an asynchronous intermittent communication protocol is designed, where the activation and rest intervals are adjusted based on the real-time error states of the agents.By utilizing the distributed extended observer to observe the relative output information and unmodeled disturbances, the novel asynchronous intermittent consensus protocol with disturbance rejection is designed to realize the overall consensus of the multi-agent systems. The proposed spatial-segmentation-dependent intermittent communication methodology can adjust communication and non-communication time of each agent asynchronously according to the communication requirements, under which the multi-agent systems can tolerate more non-communication time and reduce the communication frequency. Finally, numerical simulations are performed to verify our results.
Ruotong Wang, Lei Liu 0006, Yang Yang 0052, Qi-He Shan, Jianxin Zhang 0001
IEEE Trans Autom. Sci. Eng.6
2026 MS2M: Multi-Granularity Self-Supervised Second-Order Multiple Instance Learning for Breast Cancer Pathology Image
abstract
Combining big data and deep learning can analyze large-scale breast cancer pathology images for auxiliary diagnosis. Furthermore, Whole Slide Images (WSIs) of breast cancer pathology offer detailed tissue feature information, which supports the accurate identification of malignant lesions. Current approaches combine Self-Supervised Learning (SSL) and Multiple Instance Learning (MIL) for WSI analysis, aiming to address the issues of billion-level pixels in a single WSI and the lack of precise annotations. However, pseudo-labels produced by SSL frequently lack accuracy, and MIL fails to effectively integrate global information at the WSI level, resulting in performance bottlenecks. This paper proposes the Multi-granularity Self-supervised Second-order MIL (MS2M) to tackle these issues. MS2M first achieves instance-level fine-grained feature learning through multi-granularity SSL and optimizes instance-level representations using bag-level labels within the MIL framework. Then, the transformer captures long-range dependencies between instances. When combined with second-order (covariance) pooling, it also captures high-order relational information. This process generates a robust bag-level representation. MS2M achieves accuracies of 0.9845 and 0.9719 on the CAMELYON16 and private breast cancer WSI datasets, respectively, outperforming existing methods.
Zhenwei Wang 0005, Haitao Yao, Guangjie Han, Bingcai Chen, Pengfei Wang 0013, Jianxin Zhang 0001
IEEE Trans. Big Data7
2026 Inverse Feature Consistency Federated Unlearning for Vision-Language Model
abstract
Vision-Language Models (VLMs), with their advantages in vision and language processing, exhibit immense potential in mobile intelligent systems. Integrating federated learning with parameter-efficient fine-tuning of VLMs helps address data heterogeneity challenges. However, existing methods mainly focus on task-specific patterns, neglecting the impact of general features, such as background information and low-quality data, which weakens the model's ability to generalize when handling data from different sources and dealing with fluctuations in quality. To tackle these challenges, we propose Inverse Feature Consistency Federated Unlearning (IFCFU) for VLM, comprising three components: 1) Feature Consistency Federated Learning (FCFL) aligns fine-tuned features with pre-trained features through constraints to ensure the preservation of general features; 2) Pseudo-label Low-quality Data Detection (PLDD) identifies potential low-quality data through model quality assessment and pseudo-label generation; 3) Inverse Feature Consistency Unlearning (IFCU) distances low-quality data features from optimal model features to eliminate the negative impact and restores training with pseudo-labels. Evaluations on StanfordCars show that FCFL increased accuracy by 4.97% and 29.88% under normal data and low-quality data configurations, respectively. PLDD identified over 90.00% of low-quality data, while IFCU improved the global model's accuracy by 4.43% with 80% low-quality data.
Zhenwei Wang 0005, Pengfei Wang 0013, Guangjie Han, Jianxin Zhang 0001, Muhammed Ameen, Qiang Zhang 0008
IEEE Trans. Mob. Comput.4
2025 EProtoSeg: An Explainable Prototype-Based Network with Multi-Scale Context for Brain Tumor Segmentation
abstract
Automatic segmentation of brain tumors in magnetic resonance imaging (MRI) is essential for clinical decision support, yet it remains a challenging task due to the pronounced heterogeneity of gliomas and the often indistinct boundaries between their subregions. Moreover, the opaque, “black-box” nature of conventional deep learning models limits their adoption in clinical workflows, where interpretability is critical. To address these challenges, we propose EProtoSeg, a novel 3D segmentation network that synergistically integrates explainable prototypebased feature learning with adaptive multi-scale context aggregation. Specifically, EProtoSeg incorporates an Explainable Prototype Fusion (EPF) module in the decoder, which learns class-specific prototypes to guide voxel-wise classification and improve boundary precision. In parallel, an Adaptive Multi-Scale Context (AMSC) module is embedded in the skip connections to dynamically fuse fine spatial details and high-level semantic information across scales. A deep supervision strategy is further employed to enhance discriminability in ambiguous regions and ensure stable optimization. Extensive experiments on the BraTS 2020 and 2021 benchmarks demonstrate that EProtoSeg achieves state-of-the-art segmentation performance while offering interpretable feature representations, thereby improving both accuracy and clinical trust. The code is available at https://github.com/chenbn266/EProtoSeg
Bonian Chen, Qiule Sun, Jianxin Zhang 0001, Bin Liu 0040, Qiang Zhang 0008
BIBM4
2025 Dual-Prompt Learning with Cross-Modal Decoders for Few-Shot Whole Slide Image Classification
abstract
Few-shot learning offers a promising solution for computational pathology by alleviating the reliance on large an-notated datasets, but faces challenges from the high redundancy in whole slide images and underutilized cross-modal knowledge. Existing methods typically use foundation models only for pre-liminary feature extraction while employing fixed or single-level prompts that lack multi-scale pathological representation. To address these limitations, we propose a Hierarchical Vision-Text Prompt (H- VTP) framework that enables multi-level cross-modal interaction through GPT-4 generated Local Instance Prompts for patch-level morphological details and Global Semantic Prompts for slide-level diagnostic context. A dual-branch decoding mech-anism with Text-Guided-Patch Decoder and Patch-Augmented-Text Decoder facilitates closed-loop vision-text fusion, while a parameter-efficient adaptation strategy trains only lightweight prompts and adapters. Extensive experiments on three cancer subtype datasets demonstrate the superiority of H-VTP in few-shot WSI classification, confirming its effectiveness for clinical applications.
Bingbing Zhang 0001, Wen Zhu, Bin Liu 0040, Jianxin Zhang 0001, Qiang Zhang 0008
BIBM6
2025 Volumetric Axial-Shift Mamba U-Net for 3D Brain Tumor Segmentation
abstract
Accurate MRI brain tumor segmentation is significant for disease diagnosis and treatment. U-Nets are widely explored in automatic segmentation due to high accuracy and efficiency. Additionally, global semantic information establishes effectiveness in improving MRI brain tumor segmentation accuracy, and corresponding models have gained increasing attention. Inspired by state-space models excelling in long-range dependency modeling, we propose the Volumetric Axial Shift Mamba network (VASMambaU-Net) - a novel brain tumor segmentation method more suitable for 3D images. It integrates an improved Mamba module with large kernel convolution in U-Net to capture global tumor features. In VASMambaU-Net, we propose VASMamba that using a volumetric axial shift mechanism as the bottleneck to capture long-range dependencies of 3D brain tumor images. In addition, along with the conventional convolutional encoder, a large kernel convolutional encoder is supplemented to enhance the image receptive field, further improving the capability to capture global information, thereby enhancing the brain tumor segmentation performance. VASMambaU-Net achieves mean DSC values of$84.37 \%, 84.31 \%$and 91.07 % in the BraTS2019-2021 datasets, respectively. The corresponding mean HD95 values obtained in these three datasets are 4.67 mm, 13.51 mm, and 5.77 mm, respectively. These results demonstrate the competitiveness and effectiveness of VASMambaU-Net compared to state-of-the-art methods.
Muqing Zhang, Bonian Chen, Yutong Han, Bin Liu 0040, Jianxin Zhang 0001, Qiang Zhang 0008
BIBM5
2025 SFPA-Net: Second-order Feature Enhancement Prototype Aggregation Network for MRI Brain Tumor Segmentation
abstract
Accurate segmentation of brain tumors faces challenges of difficult edge voxel recognition and lack of interpretability. Prototype network demonstrates significant potential for brain tumor segmentation owing to its high interpretability. However, accurately extracting typical tumor features using the sample mean as a class prototype is challenging, and directly extracting prototypes from backbone-processed features ignores feature dependencies. To address these challenges, we propose the Second-order Feature Prototype Aggregation Network (SFPANet) for brain tumor segmentation. SFPA-Net introduces Covariance Feature Enhancement Module (CFEM) and Weighted Prototype Integration Module (WPIM). CFEM enhances feature representation by computing the covariance (second-order) matrix and applying power normalization. WPIM leverages enhanced features to extract tumor category characteristics and aggregates these features to represent tumor prototypes. Additionally, to enhance gradient propagation, SFPA-Net introduces deep supervision signal to optimize the training process. The Dice Similarity Coefficients of SFPA-Net on the BraTS2020 and BraTS2021 training datasets for the whole tumor, tumor core, and enhanced tumor are 91.62% / 93.50%, 89.58% / 92.55%, and 77.76% / 86.26%, respectively. Furthermore, the 95% Hausdorff Distances are 6.81 / 4.98, 5.98 / 5.86, and 13.1 / 10.68, respectively, surpassing existing brain tumor segmentation methods.
Zhenwei Wang 0005, Jianxin Zhang 0001, Qiang Zhang 0008
IJCNN5
2025 Cross Attention Guided Multimodal Network for Video Action Recognition
Bingbing Zhang 0001, Yongqi Li 0014, Jianxin Zhang 0001, Qiang Zhang 0008
PRCV (7)4
2025 High-Order Multimodal Multi-task Video Action Recognition
Bingbing Zhang 0001, Yongqi Li 0014, Jianxin Zhang 0001, Qiang Zhang 0008
PRCV (7)4
2025 Few-Shot Action Recognition Based on Visual-Language Prototype Hierarchical Temporal Enhancement
Bingbing Zhang 0001, Yuanchen Ma, Jianxin Zhang 0001, Qiang Zhang 0008
PRCV (7)4
2025 Instance-aware context with mutually guided vision-language attention for referring image segmentation
Qiule Sun, Jianxin Zhang 0001, Bingbing Zhang 0001, Peihua Li
Appl. Intell.2
2025 A diffusion model based on multi-scale spatial Mamba for medical image segmentation
Qiule Sun, Muqing Zhang, Jianxin Zhang 0001
Eng. Appl. Artif. Intell.4
2025 MSPL: Multimodal Statistical Prompt Learning for New Energy Equipment Defect Recognition
abstract
Monitoring renewable energy devices is crucial for the timely detection of faults and the improvement of system stability, making it a key component of the Internet of Things (IoT) ecosystem. However, existing intelligent algorithms and IoT methods rely on edge devices with limited computational capacity, which requires balancing real-time performance and accuracy. Additionally, most methods require initial training on large-scale data using cloud platforms or high-performance servers, leading to high initial costs. To address these challenges, we propose multimodal statistical prompt learning (MSPL) for new energy equipment defect recognition on IoT edge devices. This method enables rapid learning with only a few samples, avoiding the need for large-scale centralized training. Specifically, MSPL uses the pretrained contrastive language-image pretraining model as its backbone, leveraging text-based conceptual information to enhance the understanding of visual inputs. A statistical query module is implemented at the end of the backbone to extract distinctive features from the outputs, integrating these features with soft prompts to customize them for the defect recognition task. Since the learnable parameters are limited to soft prompts added at the end of the backbone, MSPL restricts gradient backpropagation to this point. This improves parameter and memory efficiency, making it more suitable for scenarios with limited computational capacity on IoT edge devices. Experimental results of MSPL on two renewable energy equipment defect datasets and edge devices indicate that it meets real-time processing requirements while maintaining high accuracy, outperforming other methods.
Zhenwei Wang 0005, Pengfei Wang 0013, Guangjie Han, Jianxin Zhang 0001, Guangjie Fan, Qiang Zhang 0008
IEEE Internet Things J.6
2025 Few-Shot Defect Recognition for New Energy Equipment via Multimodal Harmony
abstract
Internet of Things (IoT) technologies have been applied to fault detection in new energy equipment, which is crucial for ensuring the stable operation of energy systems. However, existing approaches typically rely on large amounts of labeled data to train intelligent algorithms, which are difficult to obtain in real-world. To address this challenge, this paper proposes a novel multi-modal few-shot defect recognition framework for new energy equipment, enabling data-efficient defect recognition in real-world scenarios through multi-modal harmony. Unlike previous multi-modal few-shot methods, it eliminates the necessity of pairwise similarity calculations, thereby simplifying the training and inference processes. Our approach trains a shared classifier through text and visual modality features integration. Initially, we extract feature vectors using text and visual encoders and then map them to common feature space. Subsequently, the vectors of these two modalities are used to train a shared classifier, which aids in the simultaneous learning of corresponding visual representations and conceptual information. Furthermore, to enhance interaction between modalities, the feature vectors of the text prompts are used to initialize the classifier weights, thereby promoting cross-modal consistency. Experiments on datasets of wind turbine blades and solar cells show that combining the two modalities improves recognition accuracy by up to 19.13% and 28.52% compared to other multi-modal few-shot methods. Despite its reliance on prompt quality, the approach provides an effective and scalable solution for defect recognition.
Zhenwei Wang 0005, Pengfei Wang 0013, Mohammad S. Obaidat, Tianbao Yang, Jintao Zheng, Jianxin Zhang 0001, Qiang Zhang 0008
IEEE Internet Things J.7
2025 Model Recovery in Federated Unlearning With Restricted Server Data Resources
abstract
Recent model recovery methods in federated unlearning (FUL) either rely on additional communication with the remaining clients or require large amounts of high-quality data from the server for training, overlooking scenarios with limited data resources. Currently, contrastive language-image pretraining (CLIP) has demonstrated remarkable performance across a wide range of tasks, particularly excelling in few-shot learning scenarios. In this article, inspired by CLIP, we explore the scenario of few-shot knowledge distillation and propose CLIP-guided few-shot knowledge distillation (CGKD) for model recovery in FUL. CGKD mainly consists of three components: 1) the unlearning module constructs the unlearning model by erasing all historical contributions of the target client, and this model is treated as the student model; 2) fine-tuning the pretrained CLIP model using few-shot data from the server side to obtain a more robust teacher model (CLIP$^{\mathbf {*}}$); and 3) model recovery is achieved through knowledge distillation, leveraging the rich visual and semantic knowledge of CLIP$^{\mathbf {*}}$to enhance the student model’s understanding of image semantic context, thereby improving the performance of the unlearning model. Extensive experimental results demonstrate that CGKD outperforms the compared FUL method in recovery performance across four standard datasets, validating the effectiveness of our approach.
Jianxin Zhang 0001, Mengda Zhao, Zhenwei Wang 0005, Weijian Su, Pengfei Wang 0013
IEEE Internet Things J.1
2025 Cycle generative adversarial Transformer network for MRI brain tumor segmentation
Muqing Zhang, Qiule Sun, Yutong Han, Bin Liu 0040, Paule-J. Toussaint, Jianxin Zhang 0001, Alan C. Evans
Neural Comput. Appl.8
2024 APRNet: Cardinality Estimation Method Based on Attention Mechanism
Zhengxuan Yang, Yutong Han, Jianxin Zhang 0001
ADMA (1)3
2024 PM2: A New Prompting Multi-modal Model Paradigm for Few-shot Medical Image Classification
abstract
Few-shot learning has become a key technical solution for addressing the challenges of limited data and difficult annotation acquisition in medical image classification. However, relying solely on a single image modality proves inadequate for capture conceptual categories. This paper proposes a novel medical image classification paradigm based on a multi-modal foundation model, called PM2. In addition to the image modality, PM2introduces supplementary text input (prompt) to further describe images or conceptual categories and facilitate cross-modal few-shot learning. We empirically studied five different prompting schemes under this new paradigm. Furthermore, linear probing in multi-modal models only takes class token as input, ignoring the rich statistical data contained in high-level visual tokens. Therefore, we alternately perform linear classification on the feature distributions of visual tokens and class token. To effectively extract statistical information, we use global covariance pool with efficient matrix power normalization to aggregate the visual tokens. We then combine two classification heads: one for handling image class token and prompt representations encoded by the text encoder, and the other for classifying the feature distributions of visual tokens. Experiments on two medical datasets demonstrate that regardless of the prompting scheme, our method PM2outperforms its counterparts, achieving state-of-the-art performance.
Zhenwei Wang 0005, Qiule Sun, Bingbing Zhang 0001, Weijian Su, Pengfei Wang 0013, Jianxin Zhang 0001, Qiang Zhang 0008
BIBM6
2024 Dynamic Temporal Shift Feature Enhancement for Few-Shot Action Recognition
Bingbing Zhang 0001, Yuanchen Ma, Jianxin Zhang 0001, Qiang Zhang 0008
PRCV (10)5
2024 Visual-guided hierarchical iterative fusion for multi-modal video action recognition
Bingbing Zhang 0001, Jianxin Zhang 0001, Qiule Sun, Qiang Zhang 0008
Pattern Recognit. Lett.3
2024 Output Consensus Control of Multi-Agent Systems With Switching Networks and Incomplete Leader Measurement
abstract
This article investigates an output consensus control problem for heterogeneous multi-agent systems with switching disconnected networks. As compared to similar works, each follower can measure only part information of the leader’s output in this paper, which lightens the measurement burden of simple agents when the dimension of leader’s output is large-scaled. In this case, due to the coexistence of incomplete measurements of leader’s output and disconnected networks, the outputs of some agents can deviate from the leader though there exists the observer-based control on them. In order to overcome this difficulty, we utilize the theory of switching unstable systems and propose a novel segmented time unit method. With the aid of this method, the switching intervals are segmented into some time units. Then by analyzing the cooperative control rule within the time units, the stabilizing characteristics of switching behaviors can be obtained to offset the divergence during the switching intervals. On this basis, a novel segmented time-varying Lyapunov function is developed to analyze the error states and sufficient criteria for the output consensus are derived. At last, a numerical simulation is shown to verify the theoretical results.Note to Practitioners—Most existing works on switching disconnected networks require that the leader (or the exosystem) is critically stable (or stable). However, unstable high-dimensional leader widely exists in the fields of multi-agent systems, such as the formation control of MASs where the agents are affine functions of time. On this account, this paper studies multi-agent systems with unstable high-dimensional leader under switching disconnected networks. To solve this problem, a novel segmented time unit method is proposed in this paper to study multi-agent systems with switching disconnected networks and incomplete leader measurement. Based on the segmented time unit approach, the observer-based control protocols and switching signals are given to realize the overall consensus. Numerical simulations suggest that this approach is feasible but it has not been tested in production. Future works will consider the formation control problem with switching disconnected networks and incomplete leader measurement.
Jianxin Zhang 0001, Lei Liu 0006, Yanming Wu 0002, Qi-He Shan
IEEE Trans Autom. Sci. Eng.2
2024 Adaptive Consensus Control of Multiagent Systems With an Unstable High-Dimensional Leader and Switching Topologies
abstract
This article addresses an adaptive consensus control problem for heterogeneous multiagent systems (MASs) with switching disconnected topologies. Unlike the existing works on switching disconnected topologies, the unstable high-dimensional leader is first considered in this work. To tackle this problem, we propose a novel blockwise energy descent approach. This approach divides the switching periods into several time blocks and mines the operation laws of agents within these blocks. Then, the descent phenomenon at switching time can be obtained, which can be used to counteract the divergence within the switching periods. Building upon this, we develop a time-varying Lyapunov function to describe the system's dynamics and establish conditions for achieving the output consensus. Finally, we develop a simulation example to confirm the validity of our theoretical results.
Hongbo Lei, Jianxin Zhang 0001, Lei Liu 0006, Qi-He Shan
IEEE Trans. Ind. Informatics3
2024 A personalized insertion centers preoperative positioning method for minimally invasive surgery of cruciate ligament reconstruction
Pengxi Li, Dongpei Liu, Bocheng Zhang, Jieshu Ren, Jianxin Zhang 0001, Bin Liu 0040
Vis. Comput.8
2023 Breast Cancer Histopathology Image Classification Using Frequency Attention Convolution Network
Ruidong Lu, Qiule Sun, Xueyan Ding, Jianxin Zhang 0001
ADMA (2)4
2023 SCAU-net: 3D self-calibrated attention U-Net for brain tumor segmentation
Ning Sheng, Yutong Han, Yaqing Hou, Bin Liu 0040, Jianxin Zhang 0001, Qiang Zhang 0008
Neural Comput. Appl.6
2023 GSoANet: Group Second-Order Aggregation Network for Video Action Recognition
Zhenwei Wang 0005, Bingbing Zhang 0001, Jianxin Zhang 0001, Bin Liu 0040, Qiang Zhang 0008
Neural Process. Lett.4
2022 Attention multiple instance learning with Transformer aggregation for breast cancer whole slide image classification
abstract
Recently, attention-based multiple instance learning (MIL) methods have received more concentration in histopathology whole slide image (WSI) applications. However, existing attention-based MIL methods rarely consider the cross-channel information interaction of pathology images when identifying discriminant patches. Additionally, they also have limitations on capturing the correlation between different discriminant instances for the bag-level classification. To address these challenges, we present a novel attention-based MIL model (AMIL-Trans) for breast cancer WSI classification. AMIL-Trans first embeds the efficient channel attention to realize the cross-channel interaction of pathology images, thus computing more robust features for instance selection without introducing too much computation cost. Then, it leverages vision Transformer encoder to directly aggregate selected instance features for better bag-level prediction, which effectively considers the correlation between different discriminant instances. Experiment results illustrate that AMIL-Trans respectively achieves its optimal AUC of 94.27% and 84.22% on the Camelyon-16 dataset and MSK external validation dataset, demonstrating the competitive performance compared with state-of-the-art MIL methods on breast cancer WSI classification task. The code will be available at https://github.con CunqiaoHou/AMIL-Trans.
Jianxin Zhang 0001, Cunqiao Hou, Wen Zhu, Ying Zou 0015, Qiang Zhang 0008
BIBM1
2022 Shuffle Attention Multiple Instances Learning for Breast Cancer Whole Slide Image Classification
abstract
Multiple instance learning (MIL) has recently become a powerful tool to solve the weakly supervised classification problem on whole slide image (WSI) pathological diagnosis. However, current MIL methods lie in two drawbacks: 1) seldom depicting feature dependencies in multiple dimensions, 2) limitation on capturing dependencies of selected instances for predicting bag-level results. To address these issues, this work presents a novel two-stage shuffle attention MIL (SAMIL) model for breast cancer WSI classification. SAMIL first introduces shuffle attention to extract important features from both spatial and channel dimensions, which well includes pixel-level pairwise relationships and channel dependencies, thus helping select more discriminant breast cancer instances for bag-level prediction. Additionally, it stacks multi-head attention with long short-term memory (LSTM) to construct an aggregator, and this adaptively highlights the most distinctive instance features while exploring the correlation between selected breast cancer instances more effectively. Experiment results on the Camelyon-16 dataset demonstrate its superior performance compared with the state-of-the-art MIL methods. The code is available at https://github.com/CunqiaoHou/SAMIL.
Cunqiao Hou, Qiule Sun, Jianxin Zhang 0001
ICIP4
2022 High-order Correlation Network for Video Recognition
abstract
How to model global video representation is an important research content of video recognition. Among current convolutional neural network(CNN) based methods, only using first-order representations (i.e. global average pooling) has limitations in capturing spatiotemporal features of videos. Recent studies have shown that high-order statistics are more suitable to model complex feature distributions. To better characterize the spatiotemporal structure for video recognition, we propose a novel High-order Correlation Network (HoCNet) in this work, the core of which is to explore high-order video representations through correlation computation and covariance pooling. HoC-Net leverages the correlation module to obtain complex temporal dynamic information of frames via computing dot product of features in the fixed sliding window of two adjacent frames. As an approximate high-order calculation, the correlation module can be inserted into any stage of the deep network to model high-order representations in various spatial resolutions. Additionally, a robust high-order pooling module, i.e., iterative matrix square root normalization of covariance pooling (iSQRT-COV), is also introduced at the end of the network, and this further boosts modeling complex spatiotemporal distributions of video features. Experiments conducted on four widely used video benchmarks demonstrate the effectiveness of HoCNet, which achieves the comparable performance with the state-of-the-art models.
Zhenwei Wang 0005, Bingbing Zhang 0001, Jianxin Zhang 0001, Qiang Zhang 0008
IJCNN4
2021 DCET-Net: Dual-Stream Convolution Expanded Transformer for Breast Cancer Histopathological Image Classification
abstract
Researches on breast cancer histopathological image classification have achieved a great breakthrough using deep backbones of Convolutional Neural Networks (CNNs) in recent years. However, due to the inductive bias of locality, CNNs are unable to effectively extract the global feature information of breast cancer histopathological images, limiting the improvement of the classification results. To overcome this shortcoming, this paper reasonably introduces an extra backbone stream of a pure transformer, which consists of a self-attention mechanism to capture global receptive fields of histopathological images, thereby compensating the locality characteristic of CNNs backbone. Based on two backbone streams of CNN and transformer, a dual-stream network called DCET-Net is proposed, which considers local features and global ones simultaneously, and progressively combines them from these two streams to form the final representations for classification. DCET-Net is extensively evaluated on the representative BreakHis histopathological image dataset, and experimental results demonstrate that it is highly competitive with the state-of-the-art CNN methods in breast cancer histopathological image classification task.
Ying Zou 0015, Shannan Chen, Qiule Sun, Bin Liu 0040, Jianxin Zhang 0001
BIBM5
2021 AutoEncoder for Neuroimage
Fan Zhang 0045, Jianxin Zhang 0001, Ahmad Chaddad, Fenghua Guo, Wenbin Zhang 0002, Ji Zhang 0001, Alan C. Evans
DEXA (2)3
2021 Attention-Guided Second-Order Pooling Convolutional Networks
abstract
Recently, channel attention-guided convolutional networks (ConvNets) have shown great advance on visual recognition tasks. However, they mainly exploit coarse first-order statistics to characterize holistic image and rarely focus on long-range feature dependencies, which limits the representation power in a certain. To handle above limitations, this paper proposes a novel attention-guided second-order pooling convolutional network (ASP-Net). ASP-Net introduces bilinear pooling that captures pairwise feature interactions to model second-order statistics. Meanwhile, it explicitly collects long-range dependencies via non-local operations, thus providing a global view in lower layers. Then, the second-order statistics and non-local context features are fused to obtain the enhanced representation for predicting channel-wise attention map and scaling convolution features. Experiment results on three commonly used datasets illuminate that ASP-Net outperforms its counterparts and achieves competitive performance. The source code is available at https://github.com/ShannanChen/ASPNet.
Shannan Chen, Qiule Sun, Cunhua Li, Jianxin Zhang 0001, Qiang Zhang 0008
ICASSP4
2021 Weakly Supervised Gleason Grading of Prostate Cancer Slides using Graph Neural Network
Yaqing Hou, Pengfei Wang 0013, Jianxin Zhang 0001, Qiang Zhang 0008
ICPRAM5
2021 A Fragment Fracture Surface Segmentation Method Based on Learning of Local Geometric Features on Margins Used for Automatic Utensil Reassembly
Bin Liu 0040, Xiaolei Niu, Shengfa Wang, Jianxin Zhang 0001
Comput. Aided Des.6
2021 Local-aware spatio-temporal attention network with multi-stage feature fusion for human action recognition
abstract
Abstract In the study of human action recognition, two-stream networks have made excellent progress recently. However, there remain challenges in distinguishing similar human actions in videos. This paper proposes a novel local-aware spatio-temporal attention network with multi-stage feature fusion based on compact bilinear pooling for human action recognition. To elaborate, taking two-stream networks as our essential backbones, the spatial network first employs multiple spatial transformer networks in a parallel manner to locate the discriminative regions related to human actions. Then, we perform feature fusion between the local and global features to enhance the human action representation. Furthermore, the output of the spatial network and the temporal information are fused at a particular layer to learn the pixel-wise correspondences. After that, we bring together three outputs to generate the global descriptors of human actions. To verify the efficacy of the proposed approach, comparison experiments are conducted with the traditional hand-engineered IDT algorithms, the classical machine learning methods (i.e., SVM) and the state-of-the-art deep learning methods (i.e., spatio-temporal multiplier networks). According to the results, our approach is reported to obtain the best performance among existing works, with the accuracy of 95.3% and 72.9% on UCF101 and HMDB51, respectively. The experimental results thus demonstrate the superiority and significance of the proposed architecture in solving the task of human action recognition.
Yaqing Hou, Hua Yu 0006, Pengfei Wang 0013, Hong-Wei Ge, Jianxin Zhang 0001, Qiang Zhang 0008
Neural Comput. Appl.6
2020 Second-order Attention Guided Convolutional Activations for Visual Recognition
abstract
Recently, modeling deep convolutional activations by the global second-order pooling has shown great advance on visual recognition tasks. However, most of the existing deep second-order statistical models mainly compute second-order statistics of activations of the last convolutional layer as image representations, and they seldom introduce second-order statistics into earlier layers to better fit network topology, thus limiting the representational ability to a certain extent. Motivated by the flexibility of attention blocks that are commonly plugged into intermediate layers of deep convolutional networks (ConvNets), this work makes an attempt to combine deep second-order statistics with attention mechanisms in ConvNets, and further proposes a novel Second-order Attention Guided Network (SoAG-Net) for visual recognition. More specifically, SoAG-Net involves several SoAG modules seemingly inserted into intermediate layers of the network, in which SoAG collects second-order statistics of convolutional activations by polynomial kernel approximation to predict channel-wise attention maps utilized for guiding the learning of convolutional activations through tensor scaling along channel dimension. SoAG improves the nonlinearity of ConvNets and enables ConvNets to fit more complicated distribution of convolutional activations. Experiment results on three commonly used datasets illuminate that SoAG-Net outperforms its counterparts and achieves competitive performance with state-of-the-art models under the same backbone.
Shannan Chen, Qiule Sun, Bin Liu 0040, Jianxin Zhang 0001, Qiang Zhang 0008
ICPR5
2020 Breast Cancer Histopathological Image Classification Based on Deep Second-order Pooling Network
abstract
With the breakthrough performance in a variety of computer vision and medical image analysis problems, convolutional neural networks (CNNs) have been successfully introduced for the classification task of breast cancer histopathological images in recent years. Nevertheless, existing breast cancer histopathological image classification networks mainly utilize the first-order statistic information of deep features to represent histopathological images, failing to characterize the complex global feature distribution of breast cancer histopathological images. To address the problem, this work makes a first attempt to explore global second-order statistics of deep features for the above task. More specifically, we propose a novel deep second-order pooling network (DSoPN) for breast cancer histopatho-logical image classification, in which a robust global covariance pooling module based on matrix power normalization (MPN) is embedded into a simple yet effective CNN architecture. The given DSoPN model can capture richer second-order statistical information of deep convolutional features and produce more informative global representations for breast cancer histopatho-logical images. Experimental results on the public BreakHis dataset illuminate the promising performance of the second-order pooling for breast cancer histopathological image classification. Besides, our DSoPN achieves very competitive performance compared to the state-of-the-art methods.
Jiasen Li, Jianxin Zhang 0001, Qiule Sun, Hengbo Zhang, Jing Dong 0009, Chao Che, Qiang Zhang 0008
IJCNN2
2020 Deep High-order Asymmetric Supervised Hashing for Image Retrieval
abstract
Deep hashing has recently been attracting more and more attentions for large-scale image retrieval task owing to its superior performance of search efficiency and less storage space requirements. Among deep hashing models, asymmetric deep hashing performs feature learning on query dataset and directly generates hash code on database images, significantly improving the retrieval performance of deep hashing models. Meanwhile, recently works also establish that high-order statistic of deep features are helpful to obtain more discriminant representations of images. Therefore, to boost the retrieval capability of deep hashing, this work tries to integrate merits of the high-order statistic module and the asymmetric deep hashing architecture, and it further proposes a novel deep high-order asymmetric supervised hashing (DHoASH) for image retrieval. More specifically, we utilize a powerful global covariance pooling module based on matrix power normalization to compute the second-order statistic features of input images, which is fluently embedded into an asymmetric hashing architecture in an end-to-end manner, leading to the generation of more discriminant binary hashing code. Experiment results on two benchmarks illuminates the effectiveness of the proposed DHoASH, which also achieves very competitive retrieval accuracy compared to the state-of-the-art methods.
Yongchao Yang, Jianxin Zhang 0001, Bin Liu 0040
IJCNN2
2019 Deep Covariance Estimation Hashing for Image Retrieval
abstract
Recently, combination of advanced convolutional neural networks and efficient hashing, deep hashing have achieved impressive performance for image retrieval. However, state-of-the-art deep hashing methods mainly focus on constructing hash function, loss function and training strategies to preserve semantic similarity. For the fundamental image characteristics, they depend heavily on the first-order convolutional feature statistics, failing to take their global structure into consideration. To address this problem, we present a deep covariance estimation hashing (DCEH) method with robust covariance form to improve hash code quality. The core of DCEH involves covariance pooling as deep hashing representation performing global pairwise feature interactions. Due to convolutional features are usually high dimension and small sample size, we estimate robust covariance with matrix power normalization and then insert it into deep hashing paradigm in an end-to-end learning manner. Extensive experiments on three benchmarks show that the proposed DCEH outperforms its counterparts and achieves superior performance.
Qiule Sun, Jianxin Zhang 0001, Jingdong Cheng, Bin Liu 0040, Qiang Zhang 0008
ICIP3
2018 Deep High-order Supervised Hashing for Image Retrieval
abstract
Recently, deep hashing has achieved excellent performances in large-scale image retrieval by simultaneously learning deep features and hashing function. However, state-of-the-art works have so far failed to explore the feature statistics higher than first-order. In this paper, to take a step towards addressing this problem, we propose two novel Deep High-order Supervised Hashing architectures (DHoSH), i.e., point-wise labels based DHoSH (DHoSH-PO) and pair-wise labels based DHoSH (DHoSH-PA). The core of DHoSH is that a trainable layer of bilinear pooling incorporates into deep convolutional neural networks (CNNs) for end-to-end learning. This layer captures the local feature interactions of the image by outer product, employing the autocorrelation information and cross-correlation information of deep features. Furthermore, our DHoSH method systematically exploits the high-order statistics of features of multiple layers. Extensive experiments on commonly used benchmarks illuminate that both DHoSH-PO and DHoSH-PA can obtain competitive improvements over its first-order counterparts, and achieve state-of-the-art performance for image retrieval task.
Jingdong Cheng, Qiule Sun, Jianxin Zhang 0001, Xiaopeng Wei, Qiang Zhang 0008
ICPR3
2018 Hyperlayer Bilinear Pooling with application to fine-grained categorization and image retrieval
Qiule Sun, Qilong Wang 0001, Jianxin Zhang 0001, Peihua Li
Neurocomputing3
2017 Classification of ECG signals based on 1D convolution neural network
abstract
Recently, with the obvious increasing number of cardiovascular disease, the automatic classification research of Electrocardiogram signals (ECG) has been playing a significantly important part in the clinical diagnosis of cardiovascular disease. In this paper, a 1D convolution neural network (CNN) based method is proposed to classify ECG signals. The proposed CNN model consists of five layers in addition to the input layer and the output layer, i.e., two convolution layers, two down sampling layers and one full connection layer, extracting the effective features from the original data and classifying the features automatically. This model realizes the classification of 5 typical kinds of arrhythmia signals, i.e., normal, left bundle branch block, right bundle branch block, atrial premature contraction and ventricular premature contraction. The experimental results on the public MIT-BIH arrhythmia database show that the proposed method achieves a promising classification accuracy of 97.5%, significantly outperforming several typical ECG classification methods.
Jianxin Zhang 0001, Qiang Zhang 0008, Xiaopeng Wei
Healthcom2
2017 Exploring risk factors and predicting UPDRS score based on Parkinson's speech signals
abstract
The unified Parkinson's disease rating scale (UPDRS) is the most widely employed scale for tracking Parkinson's disease (PD) symptom progression. However, conventional way to achieve UPDRS, mainly based on the physical examinations of clinic patients performed by the trained medical staffs, involves the disadvantages of inconvenience and high medical expense. Hence, in this study, we try to explore some risk factors and accurately predict the UPDRS for PD, using the speech signals of PD patients published on UCI machine-learning archive. More specifically, inspired by the idea of ensemble learning, we firstly construct a framework of ensemble feature selection (EFS) to select a suitable subset of features among numerous speech signals. Subsequently, a personalized predictive model, trained by adopting information from similar patients, is developed to be customized for an individual PD patient. Finally, we employ the personalized predictive model to predict UPDRS score combined with various classical regression algorithms. Compared to conventional models, our study has a potential to capture more relevant risk factors and produces more accurate UPDRS score for individual patient. Experimental results on real-world dataset from UCI machine-learning archive show that our personalized predictive model gets a promising performance.
Jianxin Zhang 0001, Qiang Zhang 0008, Bo Jin 0001, Xiaopeng Wei
Healthcom1
2014 Smart Partitioning for Product DSM Model Based on Improved Genetic Algorithm
Yangjie Zhou 0002, Chao Che, Jianxin Zhang 0001, Qiang Zhang 0008, Xiaopeng Wei
ADMA3
2012 A novel color image encryption algorithm based on DNA sequence operation and hyper-chaotic system
Xiaopeng Wei, Qiang Zhang 0008, Jianxin Zhang 0001, Shiguo Lian
J. Syst. Softw.4