Zhenfang Zhu

dblp:24/8073 · also Zhen-fang Zhu · DBLP profile ↗
← Back
56ranked-venue papers
3as first author
53since 2021 · last 2026
0000-0002-7217-3109ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 1 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 10 since 2021Systems, architecture and hardware · 8 · 8 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 7 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Semantic structure fusion graph for abstractive dialogue summarization
Furui Wang, Zhenfang Zhu, Qiang Lu 0006, Shuai Gong, Hongli Pei, Zhenrui Fu, Dawei Zhao 0001
Neurocomputing2
2026 TECL: Time-Equivariant Contrastive Learning for weakly-supervised Grounded Video Question Answering
Zhenfang Zhu, Shengtai Zhang, Dawei Zhao 0001, Menglin Zhu
Knowl. Based Syst.2
2026 A Prompt-Driven framework for compensation and fusion in multimodal sentiment analysis with missing modalities
Zhenfang Zhu, Qiang Lu 0006, Hongli Pei, Kefeng Li 0003, Yuzhi Ren, Meng Li 0049, Xiaowen Sun, Dawei Zhao 0001
Knowl. Based Syst.2
2026 Mgsc: multimodal generation and self-supervised contrast learning for mitigating language bias in visual question answering
Zhenfang Zhu, Jiangtao Qi, Yanhan Sun, Dawei Zhao 0001, Xuejuan Wang
Multim. Syst.2
2026 MDSF-Net: Mamba-driven spatial-spectral dual-stream fusion and hierarchical semantic linking for medical image segmentation
Jiayi Yu, Guangyuan Zhang, Kefeng Li 0003, Dianxin Chen, Zhenfang Zhu, Guoying Pang, Yufei Peng
Multim. Syst.5
2025 Semantic-Aware Prompt Learning for Multimodal Sarcasm Detection
abstract
Multimodal sarcasm detection aims to identify whether utterances express sarcastic intentions contrary to their literal meaning based on multimodal information. However, existing methods fail to explore the model’s "ability to understand" the semantics expressed by sentences in the image context from semantic diversity perspectives. In this paper, we propose a multi-view semantic awareness method, which concretizes semantics from multiple perspectives to improve the model’s ability to capture different semantic features. Specifically, two learnable prefixes are attached to the text representation respectively to construct semantic representations from both the literal meaning and sarcastic intention perspectives. Then, image-text information is further fused through cross-attention to guide the semantic representation of different perspectives in the image context. Finally, the semantics expressed by prefixes are strengthened through KL divergence, thereby encouraging the model to capture two distinctive semantic features. Experiments on benchmark datasets demonstrate the effectiveness of our method.
Guangjin Wang, Bao Wang 0005, Fuyong Xu, Zhenfang Zhu, Peipei Wang 0001, Ru Wang 0001, Peiyu Liu 0001
ICASSP4
2025 StruGS: Structurally consistent 3D Gaussian Splatting with targeted optimization strategies
Guoying Pang, Kefeng Li 0003, Guangyuan Zhang, Yufei Peng, Jiayi Yu, Zhenfang Zhu, Peng Wang 0109, Zhenfei Wang
Comput. Graph.7
2025 SenticNet and Abstract Meaning Representation driven Attention-Gate semantic framework for aspect sentiment triplet extraction
Xiaowen Sun, Jiangtao Qi, Zhenfang Zhu, Meng Li 0049, Hongli Pei
Eng. Appl. Artif. Intell.3
2025 Spatio-temporal memory-driven CT organ segmentation with hybrid CNN-Mamba encoding and frequency-domain decoding
Guangyuan Zhang, Kefeng Li 0003, Zhenfang Zhu, Jiayi Yu, Yongshuo Zhang, Zhiming Fan
Neurocomputing4
2025 Bias-guided margin loss for robust Visual Question Answering
Yanhan Sun, Jiangtao Qi, Zhenfang Zhu, Kefeng Li 0003, Lei Lv
Inf. Process. Manag.3
2025 Question guided multimodal receptive field reasoning network for fact-based visual question answering
Zicheng Zuo, Yanhan Sun, Zhenfang Zhu
Multim. Tools Appl.3
2025 SREGS: Sparse-view Gaussian radiance fields with geometric regularization and region exploration
abstract
Recent advances in few-shot novel-view synthesis based on 3D Gaussian Splatting (3DGS) have shown remarkable progress. Existing methods usually rely on carefully designed geometric regularizers to reinforce geometric supervision; however, applying multiple regularizers consistently across scenes is hard to tune and often degrades robustness. Consequently, generating reliable geometry from extremely sparse viewpoints remains a key challenge. To overcome this limitation, we introduce SREGS, a framework tailored for few-shot reconstruction whose contributions focus on two aspects: explicitly consistent geometry and multi-scale depth-guided optimization. Specifically, to explicitly optimize reconstruction consistency, we initialize the point cloud with 2D Gaussians, thereby enhancing depth consistency for the same Gaussian observed from different views. Secondly, we employ region-adaptive rapid densificationn to fill under-covered regions with additional representations, while an opacity-aware noise term injects stochasticity into each Gaussian to boost exploration in under-observed areas. In addition, to strengthen geometric refinement of the radiance field, we impose multi-scale depth constraints based on a monocular depth prior, performing geometric refinement from global to local scales and ensuring highly accurate reconstruction. Extensive experiments on LLFF, MipNeRF360, and Blender show that SREGS achieves higher synthesis quality with lower computational cost and demonstrates robust performance. The code is available at:https://github.com/LeeXiaoTong1/SREGS.
Kefeng Li 0003, Guangyuan Zhang, Zhenfang Zhu, Peng Wang 0109, Zhenfei Wang, Yongshuo Zhang, Zhiming Fan
Neural Networks4
2025 Diversity and Balance: Multimodal Sentiment Analysis Using Multimodal-Prefixed and Cross-Modal Attention
abstract
Multimodal Sentiment Analysis (MSA) is the technology of intelligently recognizing and assessing human sentiments using various data forms such as text, image, and audio. Despite current mainstream methods have made significant progress, MSA still faces the following issues: 1) most current methods train models based on pre-extracted features, lacking a sufficient understanding of sentiment diversity in multimodal data and may even lead to the loss of critical information in the raw data; and 2) textual modality, which possesses high-level semantic features, should typically dominate the fusion process, yet current methods fail to fully leverage this characteristic to balance modality information. To address the aforementioned issues, we propose a novel Multimodal Sentiment Analysis framework using Multimodal-Prefixed and Cross-Modal Attention (DB-MPCA). For the first issue, DB-MPCA employs multimodal raw data for pre-training, which not only allows for in-depth exploration of multimodal information but also significantly enhances the model’s learning capabilities and generalization, while reducing the substantial costs associated with manual annotation. Regarding the second issue, DB-MPCA introduces two prefix encoders designed to convert acoustic and visual features into prefix tokens. These tokens are then embedded into a pre-trained language model, where they are encoded together with textual tokens. Through this approach, DB-MPCA effectively learns cross-modal attention while maintaining the dominance of the textual modality, thereby optimizing the fusion of modalities. Comprehensive experiments conducted on the widely utilized dataset (CMU-MOSI) demonstrate the effectiveness of our model, highlighting its superiority over baseline models.
Meng Li 0049, Zhenfang Zhu, Kefeng Li 0003, Hongli Pei
IEEE Trans. Affect. Comput.2
2025 LFVGS: lightweight Gaussian splatting method for few-shot view synthesis
Kefeng Li 0003, Guangyuan Zhang, Zhenfang Zhu, Peng Wang 0109, Zhenfei Wang, Yongshuo Zhang, Zhiming Fan
J. Supercomput.4
2025 DRSS: a multimodal sentiment analysis approach based on dual representation and self-supervised learning strategy
Zhenfang Zhu, Jiangtao Qi
J. Supercomput.2
2025 Text-to-image person retrieval with implicit relation alignment and contrastive learning
Xiangyu Shui, Zhenfang Zhu, Hongli Pei, Kefeng Li 0003
J. Supercomput.2
2025 MFADU-Net: an enhanced DoubleU-Net with multi-level feature fusion and atrous decoder for medical image segmentation
Guangyuan Zhang, Kefeng Li 0003, Zhenfang Zhu, Yongshuo Zhang, Zhiming Fan
Vis. Comput.4
2024 HyperMR: Hyperbolic Hypergraph Multi-hop Reasoning for Knowledge-based Visual Question Answering
abstract
Knowledge-based Visual Question Answering (KBVQA) is a challenging task, which aims to answer an image related question based on external knowledge. Most of the works describe the semantic distance using the actual Euclidean distance between two nodes, which leads to distortion in modeling knowledge graphs with hierarchical and scale-free structure in KBVQA, and limits the multi-hop reasoning capability of the model. In contrast, the hyperbolic space shows exciting prospects for low-distortion embedding of graphs with hierarchical and free-scale structure. In addition, we map the different stages of reasoning into multiple adjustable hyperbolic spaces, achieving low-distortion, fine-grained reasoning. Extensive experiments on the KVQA, PQ and PQL datasets demonstrate the effectiveness of HyperMR for strong-hierarchy knowledge graphs.
Fuyong Xu, Peiyu Liu 0001, Zhenfang Zhu
LREC/COLING4
2024 DSAMR: Dual-Stream Attention Multi-hop Reasoning for knowledge-based visual question answering
Yanhan Sun, Zhenfang Zhu, Zicheng Zuo, Kefeng Li 0003, Shuai Gong, Jiangtao Qi
Expert Syst. Appl.2
2024 Domain-consistent syntactic representation for cross-domain aspect sentiment triplet extraction
Guangjin Wang, Bao Wang 0005, Fuyong Xu, Ru Wang 0001, Zhenfang Zhu, Peiyu Liu 0001
Expert Syst. Appl.5
2024 HTPosum:Heterogeneous Tree Structure augmented with Triplet Positions for extractive Summarization of scientific papers
Zhenfang Zhu, Shuai Gong, Jiangtao Qi, Chunling Tong
Expert Syst. Appl.1
2024 Collaborative denoised graph contrastive learning for multi-modal recommendation
Fuyong Xu, Zhenfang Zhu, Yixin Fu, Ru Wang 0001, Peiyu Liu 0001
Inf. Sci.2
2024 Joint training strategy of unimodal and multimodal for multimodal sentiment analysis
Meng Li 0049, Zhenfang Zhu, Kefeng Li 0003, Lihua Zhou, Hongli Pei
Image Vis. Comput.2
2024 Meta-learning triplet contrast network for few-shot text classification
Kaifang Dong, Baoxing Jiang, Hongye Li, Zhenfang Zhu, Peiyu Liu 0001
Knowl. Based Syst.4
2024 Robust Visual Question Answering utilizing Bias Instances and Label Imbalance
Kefeng Li 0003, Jiangtao Qi, Yanhan Sun, Zhenfang Zhu
Knowl. Based Syst.5
2024 Bridging the gap: dual perception attention and local-global similarity fusion for cross-modal image-text matching
Xiangyu Shui, Zhenfang Zhu, Hongli Pei, Kefeng Li 0003
Multim. Tools Appl.2
2024 A Relation Embedding Assistance Networks for Multi-hop Question Answering
abstract
Multi-hop Knowledge Graph Question Answering aims at finding an entity to answer natural language questions from knowledge graphs. When humans perform multi-hop reasoning, people tend to focus on specific relations across different hops and confirm the next entity. Therefore, most algorithms choose the wrong specific relation, which makes the system deviate from the correct reasoning path. The specific relation at each hop plays an important role in multi-hop question answering. Existing work mainly relies on the question representation as relation information, which cannot accurately calculate the specific relation distribution. In this article, we propose an interpretable assistance framework that fully utilizes the relation embeddings to assist in calculating relation distributions at each hop. Moreover, we employ the fusion attention mechanism to ensure the integrity of relation information and hence to enrich the relation embeddings. The experimental results on three English datasets and one Chinese dataset demonstrate that our method significantly outperforms all baselines. The source code of REAN will be available at https://github.com/2399240664/REAN
Songlin Jiao, Zhenfang Zhu, Jiangtao Qi, Fuyong Xu, Hongli Pei, Wenling Wang, Peiyu Liu 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2024 Gtpsum: guided tensor product framework for abstractive summarization
Jingan Lu, Zhenfang Zhu, Kefeng Li 0003, Shuai Gong, Hongli Pei, Wenling Wang
J. Supercomput.2
2024 Affective Commonsense Knowledge Enhanced Dependency Graph for aspect sentiment triplet extraction
Xiaowen Sun, Zhenfang Zhu, Jiangtao Qi, Hongli Pei
J. Supercomput.2
2023 Knowledge-Grounded Dialogue Generation with Contrastive Knowledge Selection
Fuyong Xu, Zhenfang Zhu, Peiyu Liu 0001
WISE3
2023 A knowledge inference model for question answering on an incomplete knowledge graph
Qimeng Guo, Zhenfang Zhu, Peiyu Liu 0001, Liancheng Xu
Appl. Intell.3
2023 An improving reasoning network for complex question answering over temporal knowledge graphs
Songlin Jiao, Zhenfang Zhu, Wenqing Wu 0002, Zicheng Zuo, Jiangtao Qi, Wenling Wang, Guangyuan Zhang, Peiyu Liu 0001
Appl. Intell.2
2023 Exploring semantic awareness via graph representation for text classification
Zhenfang Zhu, Peiyu Liu 0001
Appl. Intell.3
2023 Reducing noise-triplets via differentiable sampling for knowledge-enhanced recommendation with collaborative signal guidance
Huajuan Duan, Xiufang Liang, Yingzheng Zhu, Zhenfang Zhu, Peiyu Liu 0001
Neurocomputing4
2023 A dynamic graph expansion network for multi-hop knowledge base question answering
Wenqing Wu 0002, Zhenfang Zhu, Jiangtao Qi, Wenling Wang, Guangyuan Zhang, Peiyu Liu 0001
Neurocomputing2
2023 Semantic Decision Internal-Attention Graph Convolutional Network for End-to-End Emotion-Cause Pair Extraction
abstract
Emotion-cause pair extraction is an emergent natural language processing task; the target is to extract all pairs of emotion clauses and corresponding cause clauses from unannotated emotion text. Previous studies have employed two-step approaches. However, this research may lead to error propagation across stages. In addition, previous studies did not correctly handle the situation where emotion clauses and cause clauses are the same clauses. To overcome these issues, the authors first use a multitask learning model that is based on graph from the perspective of sorting, which can simultaneously extract emotion clauses, cause clauses and emotion-cause pairs via an end-to-end strategy. Then the authors propose to convert text into graph structured data, and process this scenario through a unique graph convolutional neural network. Finally, the authors design a semantic decision mechanism to address the scenario in which there are multiple emotion-cause pairs in a text.
Dianyuan Zhang, Zhenfang Zhu, Jiangtao Qi, Guangyuan Zhang, Linghui Zhong
Int. J. Semantic Web Inf. Syst.2
2023 Multi-feature fused collaborative attention network for sequential recommendation with semantic-enriched contrastive learning
Huajuan Duan, Yingzheng Zhu, Xiufang Liang, Zhenfang Zhu, Peiyu Liu 0001
Inf. Process. Manag.4
2023 Knowledge-guided multi-granularity GCN for ABSA
Zhenfang Zhu, Dianyuan Zhang, Lin Li 0001, Kefeng Li 0003, Jiangtao Qi, Wenling Wang, Guangyuan Zhang, Peiyu Liu 0001
Inf. Process. Manag.1
2023 SeburSum: a novel set-based summary ranking strategy for summary-level extractive summarization
Shuai Gong, Zhenfang Zhu, Jiangtao Qi, Wenqing Wu 0002, Chunling Tong
J. Supercomput.2
2023 Exploring implicit persona knowledge for personalized dialogue generation
Fuyong Xu, Zhaoxin Ding, Zhenfang Zhu, Peiyu Liu 0001
J. Supercomput.3
2022 Aspect-based Sentiment Analysis with Graph Convolutional Networks over Dependency Awareness
abstract
Aspect-based sentiment analysis (ABSA) aims to predict the sentiment polarities of specific aspects in a comment sentence. Nowadays, models based on graph neural networks enhance semantic perception by using dependency relations on dependency graphs to analyze context and aspect words. However, these models ignore the importance of the dependency type information contained in word relations and do not utilize the dependency types to pay attention to semantic information and the noise problem caused by dependency tree parsing error. To solve the above problems, we propose a novel deep dependency-aware graph convolutional networks (DA-GCN) model in this paper. Among them, the DA-GCN establishes interactive relations with multi-head attention and makes use of the grammar information of dependency perception jointly to effectively learn related information from the generated graphs. We introduce multiple conditional random fields fusing structured attention to better capture specific aspect opinion words. Experimental results on five datasets prove the effectiveness and advancement of our proposed model.
Peiyu Liu 0001, Zhenfang Zhu
ICPR3
2022 Improving Persona Understanding for Persona-based Dialogue Generation with Diverse Knowledge Selection
abstract
A significant goal in an open-domain dialogue system is to make chatbots generate more persona coherent responses given a context. To achieve this goal, some researchers attempt to introduce persona information into neural dialogue models. However, these neural dialogue models describe excessively persona traits during the conversation, which still suffer from the problem of generating boring and meaningful responses. In this paper, we divide the open-domain personalized dialogue generation task into two processes, persona recognition and persona fusion. In the persona recognition process, we use the pre-training model to encode the personas and conversation history independently, which is beneficial to the persona information fusion. Then, we design a dynamic persona fusion mechanism to effectively mine the relevance of dialogue context and persona information, and dynamically predict whether to incorporate persona features in the process of the dialogues. Our model outperforms with 1.23% in Acc., 0.36% in BLEU, 0.92% in F1, and 0.036% in Distinct than baseline models. The experimental results on the ConvAI2 dataset illustrate that the proposed model is superior to baseline approaches for generating more coherent and persona consistent responses.
Yuanying Wang, Fuyong Xu, Ru Wang 0001, Zhenfang Zhu, Peiyu Liu 0001
ICPR4
2022 IMCN: Identifying Modal Contribution Network for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis (MSA) aims to obtain the emotional polarity of language by analyzing multiple forms of human language, facial expressions, and vocal intonation. The traditional MSA model focuses on the fusion between modalities, ignoring the different contributions of language, visual, and acoustic. Thus different information of modality possesses different importance. To further explore the contributions of different modalities, we propose a highly generalized identifying modal contributions network(IMCN), which contains modality interaction module, modality fusion, and modality joint learning units in the framework. Specifically, we first designed a language modality gain detection module to make reasonable use of visual and acoustic information and reduce the noise of modal information. Secondly, crossmodal attention is used to enrich modal information. Finally, we perform joint learning of unimodal and multimodal modalities to explore the optimal solution for multimodal output. We compared with other popular multimodal sentiment analysis models and obtained better sentiment classification results on CMU-MOSI and CMU-MOSEI benchmark datasets. We also further validated the effectiveness of different modules of IMCN through ablation experiments and discussed the ideas of IMCN design.
Qiongan Zhang, Peiyu Liu 0001, Zhenfang Zhu, Liancheng Xu
ICPR4
2022 Interactive Double Graph Convolutional Networks for Aspect-based Sentiment Analysis
abstract
Aspect-based sentiment analysis aims to judge the sentiment polarity of specific aspects in comments. Recent meth-ods use graph neural networks based on dependency trees to obtain the relationship between aspects and opinion words by using syntactic information. However, these models ignore the situation that the results of dependency tree parsing are incorrect and the sentences without significant syntactic structure. To solve these problems, in this paper, we propose an interactive double graph convolution networks (Inter-DGCN) model. We reconstruct the dependency tree according to the syntactic in-formation and construct the syntactic graph convolution module. Moreover, we construct the semantic graph convolution module, which uses multi-head self-attention to represent the semantic correlation between words. In addition, we add interactive feature fusion to learn the dual graph convolutional network interactively. Experimental results on three datasets demonstrate the effectiveness and advancement of our proposed model.
Peiyu Liu 0001, Zhenfang Zhu
IJCNN3
2022 Automatically Generating Code Comment Using Heterogeneous Graph Neural Networks
abstract
Code summarization aims to generate readable summaries that describe the functionality of source code pieces. The main purpose of the code summarization is to help software developers understand the code and save their precious time. However, since programming languages are highly structured, it is challenging to generate high-quality code summaries. For this reason, this paper proposes a new approach named CCHG to automatically generate code comments. Compared to recent models that use additional information such as Abstract Syntax Trees as input, our proposed method only uses the most original code as input. We believe that programming languages are the same as natural languages. Each line of code is equivalent to a sentence, representing an independent meaning. Therefore, we split the entire code snippet into several sentence-level code. Coupled with token-level code, there are two types of code that need to be processed. So we propose heterogeneous graph networks to process the sentence-level and token-level code. Even though we do not introduce additional structural knowledge, the experimental results show that our model has a considerable performance, which indicates that our model can fully learn structural information and sequence information from code snippets.
Dun Jin, Peiyu Liu 0001, Zhenfang Zhu
SANER3
2022 Diverse dialogue generation by fusing mutual persona-aware and self-transferrer
Fuyong Xu, Guangtao Xu, Yuanying Wang, Ru Wang 0001, Peiyu Liu 0001, Zhenfang Zhu
Appl. Intell.7
2022 Enhancing aspect and opinion terms semantic relation for aspect sentiment triplet extraction
Zhenfang Zhu, Peiyu Liu 0001, Fu Xie
J. Intell. Inf. Syst.3
2022 Graph convolutional networks with hierarchical multi-head attention for aspect-level sentiment classification
Peiyu Liu 0001, Zhenfang Zhu
J. Supercomput.4
2021 Triple Tag Network for Aspect-Level Sentiment Classification
Guangtao Xu, Peiyu Liu 0001, Zhenfang Zhu, Ru Wang 0001, Fuyong Xu, Dun Jin
ICONIP (6)3
2021 Question Answering over Knowledge Base Embeddings with Triples Representation Learning
Zicheng Zuo, Zhenfang Zhu, Wenqing Wu 0002, Qiang Lu 0006, Dianyuan Zhang, Wenling Wang, Guangyuan Zhang
ICONIP (5)2
2021 Aspect-gated graph convolutional networks for aspect-based sentiment analysis
Qiang Lu 0006, Zhenfang Zhu, Guangyuan Zhang, Shiyong Kang, Peiyu Liu 0001
Appl. Intell.2
2021 A reasoning enhance network for muti-relation question answering
Wenqing Wu 0002, Zhenfang Zhu, Guangyuan Zhang, Shiyong Kang, Peiyu Liu 0001
Appl. Intell.2
2021 Syntactic and semantic analysis network for aspect-level sentiment classification
Dianyuan Zhang, Zhenfang Zhu, Shiyong Kang, Guangyuan Zhang, Peiyu Liu 0001
Appl. Intell.2
2019 A embedding model for text classification
abstract
Abstract Existing word embeddings learning algorithms only employ the contexts of words, but different text documents use words and their relevant parts of speech very differently. Based on the preceding assumption, in order to obtain appropriate word embeddings and further improve the effect of text classification, this paper studies in depth a representation of words combined with their parts of speech. First, using the parts of speech and context of words, a more expressive word embeddings can be obtained. Further, to improve the efficiency of look‐up tables, we construct a two‐dimensional table that is in the format to represent words in text documents. Finally, the two‐dimensional table and a Bayesian theorem are used for text classification. Experimental results show that our model has achieved more desirable results on standard data sets. And it has more preferable versatility and portability than alternative models.
Peiyu Liu 0001, Yuzhen Yang, Jing Yi, Zhenfang Zhu
Expert Syst. J. Knowl. Eng.5
2018 Research of Social Network Information Transmission Based on User Influence
Zhenfang Zhu, Peipei Wang 0001, Peiyu Liu 0001
ICIC (3)1
2011 Authentications and Key Management in 3G-WLAN Interworking
Xinghua Li 0001, Xiang Lu 0004, Jianfeng Ma 0001, Zhenfang Zhu, Li Xu 0002, Youngho Park 0005
Mob. Networks Appl.4