Enguang Zuo

dblp:260/3714 · DBLP profile ↗
← Back
34ranked-venue papers
1as first author
34since 2021 · last 2026
0000-0003-0029-2018ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 1 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Mixture-of-experts-based hierarchical dynamic multimodal fusion network for dermatological diagnosis
Min Li 0093, Enguang Zuo, Xiaoyi Lv, Shumei Bao, Chengwei Rao, Chen Chen 0078
Neurocomputing4
2026 Adaptive wavelet-mixed network (AWM): An efficient time series forecasting framework with wavelet-guided period attention for dimension transformation
Wenhan Song, Yuchen Ni, Fanghua Ren, Enguang Zuo, Junyu Zhu, Binglu Hu
Neurocomputing4
2026 FSC-MAE: Feature structure coordinated mask autoencoder
Enguang Zuo, Chen Chen 0078, Xiaoyi Lv, Ruishuang Sun, Yinhong Li, Hongbing Ma
Neurocomputing2
2026 PatchFusionMLP: A scalable multi-resolution MLP framework for time series prediction
Xinyu Bi, Xiaoyi Lv, Junyu Zhu, Hongbing Ma, Enguang Zuo
Pattern Recognit.7
2025 MBGNet: Mamba-Based Boundary-Guided Multimodal Medical Image Segmentation Network
Ke Xu 0005, Guangjian Liu, Enguang Zuo, Xiaoyi Lv
CVM (1)6
2025 KCGAFormer: When Large-Kernel ConvFormer Meets KAN in Semantic Segmentation
abstract
Vision Transformer, with the distinctive architecture and self-attention mechanisms, had profoundly influenced the field of computer vision, establishing Transformer-based models as benchmarks for semantic segmentation. In this study, we propose a pioneering hybrid model that fuses Kolmogorov-Arnold convolutions with ViT architecture to tackle the intrinsic challenges of semantic segmentation. By leveraging the unique attributes of Kolmogorov-Arnold convolutions, our approach introduces a convolutional attention mechanism within the Vision Transformer framework, effectively alleviating the quadratic complexity associated with self-attention. Furthermore, we integrate large-kernel convolutions and an upsampling module into the decoder, which is designed to enhance feature resolution, capture fine details, and maintain robust performance in complex scenarios for dense prediction tasks. Comprehensive experiments conducted on the ADE20K, Cityscapes, and COCO-Stuff datasets reveal that our method achieves mean Intersection over Union (mIoU) scores of 55.52%, 83.6%, and 51.8%, respectively.
Zhengxing Huang, Enguang Zuo, Alimjan Aysa, Kurban Ubul
ICASSP3
2025 SCGRL: Graph representation learning based on edge structure contrastive self-supervised framework
abstract
In recent years, significant advancements have been made in contrastive self-supervised learning for graphs. However, most of the existing methods start from the feature level and ignore the structural information. In this work, we propose a graph representation learning based on edge structure contrastive self-supervised framework (SCGRL), which leverages a novel edge-structure-based "masked edges vs. complementary edges" instance pairs to fully utilize the topological information of the graph, and attempts to reconstruct the original graph using the visible graph structure. In the feature processing, normal coded features are constrained with the coded features without gradient updating to enhance the encoder’s prediction ability for the masked representation. In addition, boundary losses are designed to ensure that the model can accurately distinguish between different instance pairs. We conduct extensive experiments on various benchmark datasets to demonstrate that SCGRL outperforms the state-of-the-art in different downstream tasks, especially link prediction.
Ruishuang Sun, Ruiting Wang, Enguang Zuo, Junyu Zhu, Chen Chen 0078, Xiaoyi Lv
ICME3
2025 Frequency-Spatial Domain Fusion for Graph Anomaly Detection
abstract
Graph anomaly detection (GAD) remains a major challenge in artificial intelligence security applications. Existing graph neural networks (GNN) face two key issues: (1) the neighborhood smoothing in spatial domain methods can mask the distribution differences between normal and abnormal graphs, filtering out important high-frequency signals; (2) their excessive reliance on local structural patterns fails to capture global frequency responses, limiting the effectiveness of anomaly detection. To address these issues, we propose FSGAD (Frequency-Spatial Graph Anomaly Detection), an unsupervised framework that combines spectral analysis and contrastive learning. Our main innovations include spectral fingerprint extraction, which extracts node-level spectral features through graph Fourier transforms, capturing global distribution patterns to accurately detect abnormal nodes; and neighborhood contrastive learning and dual-channel feature reconstruction, which use spatial and frequency domain information for precise anomaly pattern detection. Our experimental results on eight benchmark datasets show that FSGAD surpasses existing methods with an AUC improvement of 2.9% and a time consumption reduction of 21%. Our code and data are available at https://github.com/Senbao/FSGAD.
Senbao Hou, Enguang Zuo, Ruiting Wang, Xiaoyi Lv
IJCNN2
2025 Clinical Experience-inspired Multimodal Fusion Networks for Dermatological Classification
abstract
Multimodal fusion algorithms that integrate clinical images and metadata are important for dermatological classification. However, existing fusion algorithms are overly focused on model optimization at the technical level, neglecting the role of prior knowledge in the medical domain to guide the model. Therefore, this paper proposes a clinical experience-inspired multimodal fusion network (CEMF-Net), which embeds clinical experience into the multimodal fusion model and effectively improves the interpretability and clinical applicability of the model. Specifically, first, based on inspiration from a medical perspective, we have designed a clinical image amplification (CIA) block, through which the operation of doctors’ amplification observation is effectively simulated using bicubic interpolation. Meanwhile, considering the interdependence between the local detail analysis of the lesion area and the global localization demand, we design the metadata-guided lesion localization (MLL) block for precise localization operation. Secondly, an adaptive multi-scale fusion strategy is proposed at the decision-making level, which dynamically integrates the discriminative information by autonomously learning the prediction weight parameters of the features at different scales, to better simulate the clinician’s comprehensive trade-offs of the multi-scale visual features in the diagnosis process. Experimental results on the PAD-UFES-20 and Derm7pt datasets show that CEMF-Net outperforms existing representative dermatological classification algorithms in evaluation metrics such as accuracy, precision, and balanced accuracy. The experimental results demonstrate the effectiveness of embedding clinical perspectives into dermatological classification algorithms and provide new research ideas in the field of dermatological multimodal fusion.
Enguang Zuo, Guangjian Liu, Xiaoyi Lv
IJCNN5
2025 Beyond Local Features: A Metadata-Driven Image Frequency Modulation Network for Skin Disease Classification
abstract
Metadata (e.g., age, genetic history, etc.) and clinical images provide a multidimensional perspective of patients with skin diseases. Effective integration of the complementary information from these two modalities is critical. However, existing fusion techniques often combine convolutional neural networks and transformers to learn local and global features. The conventional matrix multiplication operation in transformers not only fails to focus on the frequency components of different modalities, but also leads to high spatial and temporal complexity. Therefore, this study introduces Fourier transform to capture the global information of images through point product operations in the frequency domain, thus achieving efficient integration of multi-frequency features of images driven by metadata. Specifically, we designed a Metadata-Driven Image Frequency Modulation Network (MDFM Net). The network consists of multiple cascaded frequency domain multidimensional collaborative layers (FDMC layers), which can refine the image frequency domain features step by step. The FDMC layer is composed of parallel Channel-Wise Frequency Domain Fusion Block (CWF Block) and Patch-Wise Frequency Domain Fusion block (PWF Block), which are driven by metadata and dynamically adjust the multi-frequency features of the image from different dimensions. Extensive experiments were conducted on the PAD-UFES-20 and Derm7pt datasets, and our method achieved an accuracy of 84.9% and 80.2%, respectively, improving the SOTA methods by 2.9% and 2.6%. The code will be released at: https://github.com/wwy8/MDFM.
Enguang Zuo, Xiaoyi Lv
IJCNN4
2025 FreTime:Dual-Branch Frequency-Time Representation Learning for Time Series
abstract
Time series analysis plays a fundamental role in revealing data evolution, trends, and cyclical patterns. However, existing studies often fail to effectively address the dynamic dependencies between variables in multidimensional time series and the temporal evolution patterns within variables, thereby limiting the effectiveness of complex time series feature analysis. In this paper, we propose a dual-branch frequency-time interactive representation learning model (FreqTime) that captures the correlations between variables and the temporal dependencies within variables through a collaborative architecture in the time domain and frequency domain. The time domain branch uses an inverse Transformer architecture to model cross-variable interactions, while the frequency domain branch utilizes multi-scale gated convolutions to capture features and map them back to the time domain. Finally, global representations are obtained by interactively fusing the representations learned from the two branches in the time domain. Experiments demonstrate that FreqTime achieves state-of-the-art performance on long sequence prediction, classification, and anomaly detection tasks, and exhibits strong robustness in noisy environments.
Junyu Zhu, Enguang Zuo, Ruishuang Sun, Ziwei Yan, Chen Chen 0078, Xiaoyi Lv
SMC2
2025 TDMFS: Tucker decomposition multimodal fusion model for pan-cancer survival prediction
Jinchao Chen, Enguang Zuo, Ziwei Yan, Xinya Chen, Xiaoyi Lv
Artif. Intell. Medicine5
2025 Disentangled global and local features of multi-source data variational autoencoder: An interpretable model for diagnosing IgAN via multi-source Raman spectral fusion techniques
Wei Shuai, Xuecong Tian, Enguang Zuo, Jin Gu, Chen Chen 0078, Xiaoyi Lv
Artif. Intell. Medicine3
2025 High-order graph convolutional networks for circular Ribonucleic Acid and disease association prediction incorporating multiple biological relationships
Xiaoyi Lv, Jin Gu, Enguang Zuo, Chenjie Chang
Eng. Appl. Artif. Intell.5
2025 WIGNN: An adaptive graph-structured reasoning model for credit default prediction
abstract
In credit default prediction, the main challenge is handling complex data structures and addressing data class imbalance . Given class imbalance and multi-dimensional data, general models find it difficult to fully explore the deep interdependencies within the data and the interaction effects between local and global. To overcome these challenges, this study proposes a Weighted Imbalanced Graph Neural Network (WIGNN) model that integrates adaptive graph structure inference with differential weight connectivity strategy, and the model solves the existing problems from the perspective of differential weight connectivity and graph balancing. Here, the weight connection uses the Gaussian kernel function to refine calculations and an adaptive percentile method to adjust sparsity , improving the understanding and efficiency of mining data connections. The weighted graph generated by this method can reflect the interaction between nodes and improve the model’s ability to analyse complex data structures. Based on this weighted graph, the graph imbalance module adopts a reinforcement learning-driven neighbour sampling strategy to adjust the sampling threshold automatically, optimizes the node embedding through message aggregation, and combines with a cost-sensitive matrix to improve classification accuracy and cost-effectiveness of the model on diverse credit datasets. We applied the WIGNN model to six real and class-imbalanced credit datasets, comparing it with 11 mainstream credit default prediction models. Evaluated using metrics Area Under the Curve (AUC), Geometric Mean (G-mean), and Accuracy. The results show that WIGNN significantly outperforms other models in handling class imbalance and graph sparsity , demonstrating its potential in financial credit applications.
Zhipeng Yan, Hanwen Qu, Chen Chen 0078, Xiaoyi Lv, Enguang Zuo, Xulun Cai
Eng. Appl. Artif. Intell.5
2025 The MLSE-SCAM architecture combines with the improved DRSN-TIC model for Raman spectroscopy small-sample data learning
Enguang Zuo, Zhongcheng Gong, Xiaoyi Lv
Expert Syst. Appl.3
2025 TreeXformer: Extracting tabular feature-context information using tree-structured semantics
Yinhong Li, Hanwen Qu, Chen Chen 0078, Xiaoyi Lv, Enguang Zuo, Xulun Cai
Inf. Process. Manag.5
2025 DCFusion: Difference correlation-driven fusion mechanism of infrared and visible images
Min Li 0093, Enguang Zuo, Chaoxun Guo, Yunling Wang, Xiaoyi Lv, Chen Chen 0078
Pattern Recognit.2
2025 Efficient time series adaptive representation learning via Dynamic Routing Sparse Attention
Enguang Zuo, Chen Chen 0078, Ziwei Yan, Xiaoyi Lv
Pattern Recognit.2
2025 Address Anomalies at Critical Crossroads for Graph Anomaly Detection
abstract
Graph anomaly detection (GAD) on attributed networks aims to capture abnormal nodes whose attributes or structures differ significantly from most nodes. The existing GAD models amplify the representation differences between normal and abnormal nodes to identify anomalies via carefully designed feature extraction modules. However, these models ignore the bottlenecks encountered by abnormal nodes in message passing. In particular, when the anomalies occurs at critical crossroads, the information of multiple nodes is compressed into a fixed-length representation, and the resulting over-squashing weakens the abnormal information. To address this, we propose an unsupervisedSTructural optimization model guided by sIMilarity reconstruction (STIM). Specifically, we define redundant edges that cause over-squashing, design the Neighbor-Structure Optimization module to filter redundant edges through the edge-dropping strategy based on critical crossroads, and optimize the graph structure to alleviate over-squashing. In addition, to alleviate the over-smoothing caused by the high inter-class node similarity of the data itself and the edge-dropping strategy, we design the Neighbor-Similarity Reconstruction module based on similarity calculation, which guides the model to expand inter-class variation. Extensive experiments on benchmark datasets show that STIM can effectively optimize message passing and improve anomaly detection performance. The source code is available athttps://github.com/Junyi-Yan/STIM.
Junyi Yan, Enguang Zuo, Ke Liang 0006, Meng Liu 0014, Miaomiao Li 0001, Xinwang Liu 0002, Xiaoyi Lv, Kai Lu 0001
IEEE Trans. Knowl. Data Eng.2
2024 MDKFusion: Medical Domain Knowledge-Inspired Area Amplification Network for Multi-Sequence MRI Image Fusion in Ischemic Stroke
abstract
Multi-sequence MRI image fusion technology aids radiologists in quickly and accurately assessing ischemic lesions and their surrounding areas by combining DWI and FLAIR images to generate information-rich fusion images. Despite the rapid development of medical image fusion techniques, existing methods are predominantly focused on technical-level model optimization and fail to effectively integrate medical domain knowledge. This limitation reduces their clinical applicability and model interpretability. Inspired by radiologists' diagnostic pattern, which involves focusing on and enlarging lesion areas, we propose a medical domain knowledge-inspired area amplification network for multi-sequence MRI image fusion in ischemic stroke, named MDKFusion. Specifically, we design the Lesion Area Amplification (LAA) module, which uses bicubic interpolation for adaptive amplification and incorporates crosslevel and neighboring-level feature mapping with high-level feature co-guidance. This design emulates radiologists' practice of zooming in to examine lesions, thereby enhancing interpretability. Additionally, we employ the Feature Guidance Module (FGM) to achieve progressive guidance and feature integration. We further introduce the ℒSCDloss function to minimize pixel discrepancies between source and fused images, improving fusion quality. Compared to various mainstream fusion methods, MDKFusion achieves state-of-the-art (SOTA) performance across eight objective evaluation metrics. To confirm its practical value in clinical diagnosis, we invited five radiologists to perform a subjective evaluation of the fused images. Our code will be available at https://github.com/MinLila/MDKFusion.
Min Li 0093, Pahati Tuxunjiang, Enguang Zuo, Xiaoyi Lv, Yunling Wang, Chen Chen 0078
BIBM4
2024 Improving Retrieval-Based Dialogue Systems: Fine-Grained Post-training Prompt Adaptation and Pairwise Optimization Fine-Tuning Strategy
Tianqing Zhang, Alimjan Aysa, Kurban Ubul, Enguang Zuo
ICDAR (6)5
2024 SMAE: A Split Masked Graph Autoencoder
abstract
Autoencoders, as a generative self-supervised learning, have received more and more attention in recent years in image, video, and other media-related information processing. However, Graph AutoEncoder (GAE) has yet to achieve the capability demonstrated by contrastive learning in the task-centered on attribute networks. The main limitation lies in the fact that traditional autoencoder architectures require pretext tasks that align with downstream tasks, resulting in limited expressive power of the encoder. In this paper, we propose a novel separable-task generative self-supervised learning framework capable of providing high-quality representations, Split Masked AutoEncoder (SMAE), which unleashes the encoder’s ability to extract representations through an intelligent design. Our approach focuses on unlocking the potential of the encoder by introducing encoding transfer and feature replacement strategies, thereby enabling self-supervised pretext tasks to achieve atomic separation and fully unleash the encoder’s feature representation potential. We conducted extensive experiments on widely-used graph classification datasets, and the results demonstrate that SMAE outperforms state-of-the-art baselines in terms of graph classification accuracy and generation quality. Furthermore, our experimental findings show that prediction at the representation layer is more effective than original graph layer reconstruction in the field of masked graph autoencoders.
Ruiting Wang, Enguang Zuo, Chen Chen 0078, Junyi Yan, Ziwei Yan, Xiaoyi Lv
ICME2
2024 Rethinking the Necessity of Learnable Modal Alignment for Medical Image Fusion
Min Li 0093, Enguang Zuo, Xiaoyi Lv, Chen Chen 0078
PRCV (5)3
2024 A prospective study: Advances in chaotic characteristics of serum Raman spectroscopy in the field of assisted diagnosis of disease
Chen Chen 0078, Xuecong Tian, Enguang Zuo, Chenjie Chang, Min Li 0093, Xiaoyi Lv
Expert Syst. Appl.4
2024 CMACF: Transformer-based cross-modal attention cross-fusion model for systemic lupus erythematosus diagnosis combining Raman spectroscopy, FTIR spectroscopy, and metabolomics
Xuguang Zhou, Chen Chen 0078, Xiaoyi Lv, Enguang Zuo, Min Li 0093
Inf. Process. Manag.4
2024 Self-contrastive Feature Guidance Based Multidimensional Collaborative Network of metadata and image features for skin disease classification
Min Li 0093, Enguang Zuo, Chen Chen 0078, Xiaoyi Lv
Pattern Recognit.3
2024 DSFusion: Infrared and visible image fusion method combining detail and scene information
Kuizhuang Liu, Min Li 0093, Chengwei Rao, Enguang Zuo, Yunling Wang, Ziwei Yan, Chen Chen 0078, Xiaoyi Lv
Pattern Recognit.5
2024 ASFFuse: Infrared and visible image fusion model based on adaptive selection feature maps
Kuizhuang Liu, Min Li 0093, Enguang Zuo, Chen Chen 0078, Yunling Wang, Xiaoyi Lv
Pattern Recognit.3
2023 Rethinking graph anomaly detection: A self-supervised Group Discrimination paradigm with Structure-Aware
abstract
Structural anomalies are the core problem in graph anomaly detection. However, the current mainstream self-supervised graph anomaly detection models do not directly model structural anomalies and their expensive time consumption limits the efficiency of graph anomaly detection. For this reason, we rethink graph anomaly detection and propose a self-supervised Group Discrimination paradigm with Structure-Aware (GDSA). Our model can be explicitly aware of the graph topology changes by multi-view structure disturbance. Moreover, GDSA transforms graph anomaly detection into discriminating the scalar summaries of positive and negative group nodes. The results of extensive experiments on four benchmark datasets show that GDSA outperforms current state-of-the-art methods, with the most significant AUC performance improvement of 28.7%. Notably, in scalability testing on a large-scale dataset, the training time and testing time of GDSA are 1181.0× and 5064.7× faster than the baseline, respectively, with 61.9% savings in memory usage.
Junyi Yan, Enguang Zuo, Chen Chen 0078, Tianle Li, Xiaoyi Lv
ICME2
2023 A Masked Attention Network with Query Sparsity Measurement for Time Series Anomaly Detection
abstract
Time series aomaly detection has been widely studied in recent years. Previous research focuses on point-wise features and pairwise associations for feature learning or designed anomaly scores based on prior knowledge. However, these methods cannot fully learn the intricate abnormal dynamic information and can only identify a limited class of anomalies. We propose a Masked Attention Network with Query Sparsity Measurement (MAN-QSM) to address the above challenges. This model uses two kinds of prior knowledge to fully exploit the differences between normal and abnormal points from two perspectives: pairwise association and sequence-level information. We designs the anomaly mask mechanism to collaborate with the training strategy to amplify the difference between normal and abnormal points. In experiments, we compare the model with classical methods, reconstruction-based models, autoregressive-based models, and state-of-the-art models, and the MAN-QSM achieves state-of-the-art results on SMD, PSM, and MSL datasets with an average of 16% reduction in error rate.
Enguang Zuo, Chen Chen 0078, Junyi Yan, Tianle Li, Xiaoyi Lv
ICME2
2023 MLDF-Net: Metadata Based Multi-level Dynamic Fusion Network
Enguang Zuo, Chen Chen 0078, Yunling Wang, Xiaoyi Lv, Min Li 0093
PRCV (1)2
2023 SUCOLA: Self-adaptive structure refinement unsupervised contrastive learning framework for food safety risk early warning
Enguang Zuo, Junyi Yan, Alimjan Aysa, Chen Chen 0078, Hongbing Ma, Xiaoyi Lv, Kurban Ubul
Eng. Appl. Artif. Intell.1
2022 Person re-identification based on deep learning - An overview
Wenyu Wei, Wenzhong Yang, Enguang Zuo, Yunyun Qian
J. Vis. Commun. Image Represent.3