Xinfeng Zhang 0002

dblp:12/7627-2 · DBLP profile ↗
← Back
14ranked-venue papers
8as first author
9since 2021 · last 2025
0000-0001-6304-7189ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Enhanced Prediction of Intracranial Aneurysm Rupture Risk via Multimodal Fusion
abstract
ABSTRACT It is well known that subarachnoid haemorrhage caused by intracranial aneurysm rupture has a high fatality rate. Therefore, the prediction of rupture risk can help doctors make targeted diagnoses and treatments in advance. In this study, an image‐text‐based hierarchical prediction model for the rupture risk of intracranial aneurysms (IAIT) is proposed, which combines medical images with structured texts to improve the prediction accuracy. This model captures the detailed features in the images, explores the interaction of multimodal features, and achieves better performance. Specifically, ternary‐view partial attention (TPA) is introduced into the image encoder to improve the model's attention to small lesions. With two symmetric local paths and one global path, local features can be better extracted, and Kronecker product is used for full fusion of image‐text features. Experiments on a private dataset show that the proposed model substantially outperforms both unimodal and existing multimodal baselines. It achieves over 12% higher accuracy than the best unimodal text model and over 45% higher than the best unimodal image model. Moreover, the TPA module further improves classification performance, validating the model's effectiveness. Overall, this study demonstrates the potential of multimodal fusion for accurate and interpretable prediction of intracranial aneurysm rupture risk.
Xinfeng Zhang 0002, Wei Guo 0020, Xiangsheng Li, Mao-shen Jia
IET Image Process.1
2025 Cross-corpus speech emotion recognition using semi-supervised domain adaptation network
Mao-shen Jia, Xuan Cao, Jiawei Ru, Xinfeng Zhang 0002
Speech Commun.5
2025 Deep deterministic policy gradients with a self-adaptive reward mechanism for image retrieval
abstract
Abstract Traditional image retrieval methods often face challenges in adapting to varying user preferences and dynamic datasets. To address these limitations, this research introduces a novel image retrieval framework utilizing deep deterministic policy gradients (DDPG) augmented with a self-adaptive reward mechanism (SARM). The DDPG-SARM framework dynamically adjusts rewards based on user feedback and retrieval context, enhancing the learning efficiency and retrieval accuracy of the agent. Key innovations include dynamic reward adjustment based on user feedback, context-aware reward structuring that considers the specific characteristics of each retrieval task, and an adaptive learning rate strategy to ensure robust and efficient model convergence. Extensive experimentation with the three distinct datasets demonstrates that the proposed framework significantly outperforms traditional methods, achieving the highest retrieval accuracy having 3.38%, 5.26%, and 0.21% improvement overall as compared to the mainstream models over DermaMNIST, PneumoniaMNIST, and OrganMNIST datasets, respectively. The findings contribute to the advancement of reinforcement learning applications in image retrieval, providing a user-centric solution adaptable to various dynamic environments. The proposed method also offers a promising direction for future developments in intelligent image retrieval systems.
Farooq Ahmad, Xinfeng Zhang 0002, Zifang Tang, Fahad Sabah, Muhammad Azam 0006, Raheem Sarwar
J. Supercomput.2
2024 TIM-Net: A multi-label classification network for TCM tongue images fusing global-local features
abstract
Abstract Combining the extracted tongue features with other medical indicators can effectively judge the diseases of patients. The previous work usually only analyzes a certain feature of the tongue body and is unable to extract multiple features simultaneously. In this study, a multi‐label classification network named TIM‐Net is proposed, which integrates global and local features to achieve multi‐label intelligent diagnosis of Chinese medicine tongue images. First, a feature extraction network based on ResNet is proposed to capture the features of tongue images more sufficiently. Then, a multi‐label classification algorithm fusing global and local features is proposed, and targeted screening operations are carried out on the class‐related feature maps based on global confidence. In addition, a logical masking algorithm is proposed to ensure that the local features can only correct the feature labels they represent, and do not interfere with other feature labels. The classification accuracy is further improved by using local feature confidence and correcting the global classification results. Finally, the experimental results indicate that the classification accuracy of the tongue images is gradually improved through optimizing the feature extraction network and fusing local features, and it exceeds other state‐of‐the‐art multi‐label classification networks.
Xinfeng Zhang 0002, Haonan Bian, Mao-shen Jia
IET Image Process.1
2024 A semi-supervised segmentation network fusing pseudo-label with multi-level feature consistency correction for hard exudates
abstract
Abstract Timely detection of hard exudates in fundus images can effectively avoid the severity of the disease, but the labelling of small and numerous lesion areas requires a lot of labour costs. This paper proposes a semi‐supervised segmentation network, which integrates pseudo‐labels and multi‐level features consistency correction. It achieves accurate segmentation of hard exudates by making full use of a small amount of labelled data and a large amount of unlabelled data. The network effectively extracts features from the unlabelled data through knowledge transfer of the teacher‐student model, and incorporates a Transformer network for auxiliary training to promote the quality of transfer. In addition, three unsupervised losses are introduced to improve the performance: the perturbation loss improves the robustness of the model to noise by adding different noises to the same input; the multi‐level feature consistency correction loss ensures the consistency of features of the student model at different scales; and the pseudo‐labelling cross‐supervision loss utilizes the generated pseudo‐labels for supervision between CNN and Transformer. By comparing the segmentation results with different proportion of the labelled data, it has better segmentation performance compared to other methods. The proposed methods can totally increase dice by 16.56% and mean intersection over union (MIoU) by 25.11%.
Xinfeng Zhang 0002, Mao-shen Jia
IET Image Process.1
2023 Adaptive learning Unet-based adversarial network with CNN and transformer for segmentation of hard exudates in diabetes retinopathy
abstract
Abstract Accurate segmentation of hard exudates in early non‐proliferative diabetic retinopathy can assist physicians in taking appropriate treatment in a more targeted manner, in order to avoid more serious damage to vision caused by the deterioration of the disease in the later stages. Here, an Adaptive Learning Unet‐based adversarial network with Convolutional neural network and Transformer (CT‐ALUnet) is proposed for automatic segmentation of hard exudates, combining the excellent local modelling ability of Unet with the global attention mechanism of transformer. Firstly, multi‐scale features are extracted through a CNN dual‐branch encoder. Then, the information fusion of features at adjacent scale is realized and the fused features are selected adaptively to maintain the overall consistency of features by attention‐guided multi‐scale fusion blocks (AGMFB). After that, the high‐level encoded features are input to transformer blocks to extract global contexts. Finally, these features are fused layer‐by‐layer to achieve accurate segmentation of hard exudates. In addition, adversarial training is incorporated into the above segmentation model, which improves Dice scores and MIoU scores by 7.5% and 3%, respectively. Experiments demonstrate that CT‐ALUnet shows more reliable segmentation and stronger generalization ability than other SOTA methods, which lays a good foundation for computer‐assisted diagnosis and assessment of efficacy.
Xinfeng Zhang 0002, Mao-shen Jia
IET Image Process.1
2022 An improved tongue image segmentation algorithm based on Deeplabv3+ framework
abstract
Abstract Tongue image segmentation is the key step of traditional Chinese medicine (TCM) intelligent tongue image analysis. The subsequent tongue image quality analysis is directly affected by the precision of segmentation. Deeplabv3+ network has become an excellent algorithm in the field of tongue images segmentation by virtue of its ability to extract multi‐scale information and its codec structure. However, there are unclear edge segmentation of the tongue body and missegmentation of small areas in some tongue images. In view of the above phenomenon, an improved algorithm is proposed. Firstly, the network structure is optimized, so that the ability of the network to extract multi‐scale information and low‐level information is improved. Secondly, a loss function based on edge information is proposed, which makes the network pay more attention to the separation of tongue edges in the process of training. Finally, the segmentation results are post‐processed by using the prior knowledge of tongue image, so as to eliminate the phenomenon of misjudgement. The experimental results show that the algorithm significantly improves the ambiguity of image segmentation, and the MIOU value is still increased to 99.13% when the MIOU value has reached 98.77%.
Xinfeng Zhang 0002, Haonan Bian, Yiheng Cai, Keye Zhang
IET Image Process.1
2021 Prediction of Railway Freight Customer Churn Based on Deep Forest
Xinfeng Zhang 0002, Yongle Shi
ICIC (2)2
2021 Person Re-identification Based on Hash
Xinfeng Zhang 0002, Bowen Ren, Mao-shen Jia
ICIC (1)2
2019 An Algorithm of Bidirectional RNN for Offline Handwritten Chinese Text Recognition
Xinfeng Zhang 0002, Kunpeng Yan
ICIC (3)1
2018 Application of an Improved Grab Cut Method in Tongue Image Segmentation
Guangqin Hu, Xinfeng Zhang 0002, Yiheng Cai
ICIC (3)3
2018 Optical Character Detection and Recognition for Image-Based in Natural Scene
Bochao Wang, Xinfeng Zhang 0002, Yiheng Cai, Mao-shen Jia
ICIC (3)2
2015 An Assessment Method of Tongue Image Quality Based on Random Forest in Traditional Chinese Medicine
Xinfeng Zhang 0002, Yazhen Wang, Guangqin Hu, Jing Zhang 0023
ICIC (3)1
2015 Preliminary Study of Tongue Image Classification Based on Multi-label Learning
Xinfeng Zhang 0002, Jing Zhang 0023, Guangqin Hu, Yazhen Wang
ICIC (3)1