Pengzhou Zhang

dblp:46/7697 · DBLP profile ↗
← Back
22ranked-venue papers
0as first author
14since 2021 · last 2025
0000-0001-5721-5692ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 8Databases, data management, data science and information retrieval · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 THGNets: Constrained Temporal Hypergraphs and Graph Neural Networks in Hyperbolic Space for Information Diffusion Prediction
abstract
Information diffusion prediction aims to predict the next infected user in the information diffusion, which is a critical task to understand how information spreads on social platforms. Existing methods mainly focus on the sequences or topology structure in euclidean space. However, they fail to sufficiently consider the hierarchical structure or power-law structure of the underlying topology of information cascade graphs and social networks, resulting in distortion of user features. To tackle above issue, we propose an innovative Constrained Temporal Hypergraphs and Graph Neural Networks (THGNets) framework that is tailored for information diffusion prediction. Specifically, we introduce hyperbolic temporal hypergraphs neural network to alleviate the distortion of user features by hyperbolic hierarchical learning in information cascades. Additionally, it also captures high-order dynamic interaction patterns between users and further integrates the time-consistency constraint mechanism to mitigate the instability and non-smoothness of user features in latent space. In parallel, we apply the hyperbolic graph neural network to investigate the hierarchical structure and user homogeneity on social networks, enhancing our understanding of social relationships. Moreover, hyperbolic gated recurrent units are employed to capture the potential dependency relationships between contextual users. Experiments conducted on four public datasets demonstrate that the proposed THGNets significantly outperform the existing methods, thereby validating the superiority and rationality of our approach.
Pengzhou Zhang, Wenchao Song 0001, Junpeng Gong
AAAI2
2025 Enhancing Multi-source Localization via Tailored Feature Representation Framework
Wenchao Song 0001, Guowei Chen, Junpeng Gong, Pengzhou Zhang
KSEM (3)6
2025 Multimodal Hyperbolic Embedding and Hyperbolic Hypergraph Fusion for Emotion Recognition in Conversation
abstract
Multimodal Emotion Recognition in Conversation (MERC) is a critical task for comprehending complex emotions from multiple sources in human-computer interaction. Despite existing methods having made progress, they still struggle to capture hierarchy characteristics and high-order interactions between modalities. Recent research adopts Euclidean space geometry for multimodal representation learning and fusion, resulting in representation and distance distortions. To tackle these challenges, we propose a multimodal hyperbolic embedding and hyperbolic hypergraph fusion framework. A hyperbolic embedding learning module projects modality representations into the Poincaré ball model to perceive hierarchies. To enhance representation discriminative and semantic consistency, it incorporates radial and angular contrastive learning objectives. A hyperbolic hypergraph fusion module further exploits high-order latent relationships and improves fusion results. Experimental results demonstrate that the proposed framework outperforms state-of-the-art methods and augments stability on IEMOCAP and MELD datasets.
Guowei Chen, Wenchao Song 0001, Pengzhou Zhang
MMAsia5
2025 Joint Fine-grained Disentanglement and Modal-agnostic Fusion Multi-task Framework for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis (MSA) aims to comprehend human sentiment by integrating heterogeneous multisource information on social media. The challenge of MSA lies in representation learning and multimodal fusion. Existing disentangled methods rely on global information while neglecting fine-grained critical cues, amplifying representation distribution bias in disentangled spaces. Meanwhile, current fusion approaches heavily depend on text modality while disregarding informative cues from other modalities, exacerbating the contribution imbalance across modalities. To tackle these limitations, we propose a multi-task framework that collaborates with fine-grained disentanglement and modal-agnostic fusion. It extracts sentiment attributes at a fine granularity based on the multi-head property of the Transformer, and dynamically exploits the contribution of each modality while enhancing the reliability of cross-modal representations through fuzzy logic. We evaluate FDMA on two renowned multimodal benchmark datasets, CMU-MOSI and CMU-MOSEI. Experiments show that our method surpasses state-of-the-art methods on CMU-MOSI and CMU-MOSEI datasets. Further analysis demonstrates that our model also exhibits superior reliability and stability.
Yujun Wen, Pengzhou Zhang
MMAsia5
2025 MFGCN: Multimodal fusion graph convolutional network for speech emotion recognition
abstract
Speech emotion recognition (SER) is challenging owing to the complexity of emotional representation. Hence, this article focuses on multimodal speech emotion recognition that analyzes the speaker’s sentiment state via audio signals and textual content. Existing multimodal approaches utilize sequential networks to capture the temporal dependency in various feature sequences, ignoring the underlying relations in acoustic and textual modalities. Moreover, current feature-level and decision-level fusion methods have unresolved limitations. Therefore, this paper develops a novel multimodal fusion graph convolutional network that comprehensively executes information interactions within and between the two modalities. Specifically, we construct the intra-modal relations to excavate exclusive intrinsic characteristics in each modality. For the inter-modal fusion, a multi-perspective fusion mechanism is devised to integrate the complementary information between the two modalities. Substantial experiments on the IEMOCAP and RAVDESS datasets and experimental results demonstrate that our approach achieves superior performance. • Develop a multimodal fusion graph convolutional network to execute the intra- and inter-modal interactions. • Excavate the sentiment, semantic, and temporal dependency to construct the intra-modal relations. • Devise a multi-perspective fusion mechanism for inter-modal fusion. • Adopt a multi-angle loss to optimize the model.
Yujun Wen, Pengzhou Zhang, Heyan Huang
Neurocomputing3
2025 GTHNN: Graph Transformer and Sequential Hypergraph Neural Network with dynamic aggregation mechanism for multi-task prediction
Wenchao Song 0001, Pengzhou Zhang
Inf. Process. Manag.4
2025 Multimodal disentanglement implicit distillation for speech emotion recognition
Yujun Wen, Junpeng Gong, Pengzhou Zhang
Inf. Process. Manag.4
2025 MFKD: Multi-dimensional feature alignment for knowledge distillation
Pengzhou Zhang, Peng Liang 0014
Image Vis. Comput.2
2024 A Deep Learning Method for Automatic News Frame Identification
abstract
News framing plays a crucial role in shaping public opinion and understanding of events. Traditional news frame analysis relies heavily on manual annotation, which is timeconsuming and labor-intensive. This paper proposes a novel deep learning-based method for automatic news frame identification (ANFI). By leveraging a pre-trained Transformer encoder and frame-specific semantic features, ANFI captures the latent semantic structures of news text. It then classifies the text into one of the seven commonly used generic frames: Fact, Conflict, Responsibility, Economic Consequences, Human Interest, Morality, and Leadership. Experimental results on the CNN/Daily Mail dataset demonstrate the effectiveness of our proposed method in identifying news frames automatically. The ANFI model achieves high accuracy and outperforms traditional machine learning baselines. This research provides a new perspective on the application of deep learning techniques in computational journalism and facilitates large-scale news frame analysis.
Qiyi Wei, Pengzhou Zhang
SNPD4
2024 DJMF: A discriminative joint multi-task framework for multimodal sentiment analysis based on intra- and inter-task dynamics
abstract
Multimodal sentiment analysis (MSA) is a considerable research area within the domain of social networks where users continuously share a vast amount of content. This content includes text, images, and audio, all of which are accompanied with sentiment or emotion expressions. MSA encompasses multiple tasks, such as sentiment prediction and emotion recognition, which have multi-task dynamic relationships owing to their close correlation and complementary nature. However, previous works using a single-task framework have been limited in exploring these dynamics. In this work, we propose a discriminative joint multi-task framework (DJMF) realizing sentiment prediction and emotion recognition synchronously. The discriminative learning strategy takes into account intra-task dynamics, and customized fusion schemes are applied to the subtasks. Furthermore, we utilize a joint training strategy to capture the inter-task dynamics by leveraging their interrelatedness, enhancing the performance of each task. Additionally, in the fusion process, we incorporate two categories of regularizers to enable the model to obtain distinct fusion results by reducing redundancy and task-irrelevant information. We conducted extensive evaluations on three well-known multimodal baseline datasets: CMU-MOSI, CMU-MOSEI, and UR-FUNNY. The results demonstrate that our approach offers better reliability and stability compared to the single-task framework and outperforms state-of-the-art methods.
Junpeng Gong, Yujun Wen, Pengzhou Zhang
Expert Syst. Appl.4
2024 SAKD: Sparse attention knowledge distillation
Pengzhou Zhang, Peng Liang 0014
Image Vis. Comput.2
2024 FDR-MSA: Enhancing multimodal sentiment analysis through feature disentanglement and reconstruction
abstract
Multimodal sentiment analysis (MSA) is crucial as it integrates textual, visual, and audio information from videos to accurately identify human emotional states. This study proposes an innovative multimodal feature decoupling strategy that categorizes sentiment features into common and private features. The private features aim to accurately capture the uniqueness of each modality, thereby increasing feature diversity. In contrast, the common features seek to identify and capture commonalities among different modalities, thus reducing potential information loss during decoupling. To achieve this, we designed dedicated encoders and loss function constraints for both types of features. Additionally, to mitigate information redundancy and prevent key information loss during decoupled representation learning, we introduce a dual feature reconstruction mechanism comprising unimodal feature reconstruction (UFR) and multimodal feature reconstruction (MFR). These mechanisms preserve vital information from the decoupling process and mitigate the impact of redundant data. Our extensive experiments on three datasets demonstrate that our method achieves a significant margin of approximately 1%–3% in accuracy, illustrating that our approach outperforms existing advanced techniques significantly, resulting in noteworthy performance enhancements.
Yao Fu 0009, Biao Huang 0013, Yujun Wen, Pengzhou Zhang
Knowl. Based Syst.4
2023 Design and Implementation of Digital Consulting Capability Platform based on Knowledge Sharing
abstract
The vigorous development of the digital economy brings new opportunities for enterprise digital transformation. This article proposes a knowledge-sharing-based digital consulting capacity platform, focusing on important digital transformation concerns in the consulting profession. By sorting out the practical problems and challenges, the optimization path of constructing the digital consulting capability-sharing platform for the intelligent city field is explored. Seven business centers are built to gather digital resources such as policies, industry trends, and market information. Knowledge precipitation is more convenient and comprehensive, and business management is more scientific. This paper analyzes the framework of the digital consulting capability platform in detail. Furthermore, this paper carries out the actual deployment and verifies the effectiveness of the digital consulting capability platform scheme based on knowledge sharing.
Pengzhou Zhang, Lexi Xu, Peng Liang 0014, Shuwei Yao
TrustCom2
2022 Multi-channel Live Video Processing Method Based on Cloud-edge Collaboration
abstract
The rise of 5G edge computing technology brings new development opportunities for video applications. In this paper, a multi-channel video analysis method based on cloud-edge collaboration is designed. Considering edge computing resources and network transmission time, a cloud-edge collaborative multi-channel live video processing framework is proposed. The proposed method makes full use of the processing advantages of edge computing and the cloud-edge collaboration, which can effectively meet the needs of efficient computing and network transmission for live video streaming. This paper analyzes the cloud-edge collaborative multi-channel video processing process in details. Finally, this paper conducts the test comparison between 4G and 5G live network environment, and the deployment verification in 5G network environment. The effectiveness of the cloud-edge collaborative multi-channel video solution in the 5G environment is verified through the test.
Pengzhou Zhang, Junjie Xia, Juan Cao 0005
TrustCom2
2019 Research on News Keyword Extraction Technology Based on TF-IDF and TextRank
abstract
With the rapid development of information technology and the widespread use of the Internet, the Internet, as an information carrier, has gradually replaced the traditional media such as newspapers and television, and become the main channel for people to obtain information. This paper takes English news text as the research object of keyword extraction method. We combine TF-IDF and the TextRank algorithm to extract keywords from text by constructing word graph model, counting word frequency and inverse document frequency, and considering the weight of the positioning of headlines. A large number of experiments have been carried out with Sina News Corpus, and the performance of the algorithm is evaluated by recall rate, precision rate and macro average value. The results show that the integration of TF-IDF and the TextRank algorithm significantly outperforms the traditional algorithm in performance parameters and extraction effect.
Pengzhou Zhang
ICIS2
2018 Relation Extraction Based on Deep Learning
abstract
Extracting entities and their relations is to detect entity mentions and recognize the semantic relation between them. Most of traditional methods are feature-based and treat this task as a pipeline of two separated tasks, i.e., named entity recognition and relation classification. Due to deep neural network models can learn effective entities and relations features from the given sentences without complicated manual feature engineering. Deep learning methods have been viewed as one of the major driving force in the recent development of natural language processing. We will briefly review the previous works on deep learning and give a brief overview of recent progresses on relation extraction. We implement the models based on deep neural networks both in pipeline methods and joint methods and come to conclusion about the comparison.
Pengzhou Zhang
ICIS3
2018 Cn-MAKG: China Meteorology and Agriculture Knowledge Graph Construction Based on Semi-structured Data
abstract
In this paper, a method of constructing China meteorology and agriculture knowledge graph based on semi-structured data is proposed. Firstly, demand analysis, determine the boundary of knowledge, according to which design the schema of Cn-MAKG. Then the semi-structured knowledge data is preprocessed in accordance with the standard indexes in the domain of meteorology and agriculture. Finally, the graph database Neo4j is used to store the knowledge graph, so as to realize the construction of Cn-MAKG. At present, the knowledge graph has been successfully applied to the automatic generation of crop meteorological reports.
Chenglin Qi, Pengzhou Zhang
ICIS3
2018 K-means Text Dynamic Clustering Algorithm Based on KL Divergence
abstract
The random selection of the initial cluster center and distance measure function selection in the classical k-means algorithm have a great influence on the time and final precision of the clustering. Based on the above two problems, this paper proposes the classical VSM (vector space model) to represent textual materials. Based on the maximum distance method, K data points with large distribution difference are selected as the initial cluster centers, and the similarity between the cluster centers and the sample data is obtained through KL divergence. And then put what share the similarity in a cluster, form the cluster calculation formula and the distance measure function of the iterative center, and calculate the iteration until the sample data set is empty. The experiment proves that the improved text clustering algorithm proposed in this paper not only reduces the total consumption time of clustering, but also improves the accuracy of clustering at the same time.
Huan Zhu, Pengzhou Zhang, Zeyang Gao
ICIS2
2017 An automatic generation method of sports news based on knowledge rules
abstract
Nowadays, with the massive demand for sports news, automatic generation systems based on the template technology has been deployed, which could generate massive sports news quickly and effectively. However, by using one simple template for one scenario, the pattern of text generated by such system is single. In this paper, we propose an automatic generation method based on knowledge rules to select the template dynamically from a template set. The text generated by the system is flexible and the format is varied, which improves the quality of the generated news.
Junpeng Gong, Wen Ren, Pengzhou Zhang
ICIS3
2017 Research review on key techniques of topic-based news elements extraction
abstract
With the development of computer and network techniques, and the digital Chinese news texts explosion, facing a massive unstructured news data, a better way for knowledge extraction and storage, on the one hand, can help readers understand the core content of news ,on the other hand ,completed news knowledge accumulation will support the reportage. In recent years, information extraction technology of Chinese text has developed rapidly, and has big progress on Named Entity recognition, Entity Relation Extraction and Event Extraction. In this paper, we propose a topic-based Elements Extraction and storage of news method that based on thematic event frame, and the relationship between the event elements is stored in the form of element expressions to organize the knowledge of news. Expressions can be used to discover and extract event elements, relational instances in the same thematic news text, realize topicbased knowledge of news extraction and storage. This paper uses a variety of Natural Language Processing technologies, including document filtering, classify, cluster, dependency parsing, etc. Based on these theories we designed and realized the topic-based Chinese news texts Event Elements Automatic Extraction and Expressions Automatic Generation System.
Pengzhou Zhang
ICIS3
2016 Chinese text categorization based on deep belief networks
abstract
With the rapid development of Internet, text categorization becomes a mission-critical technology that organizes and processes large amounts of data in document. Deep belief networks have powerful abilities of learning and can extract highly distinguishable features from the high-dimensional original feature space. So a new Chinese text categorization algorithm based on deep learning structure and semi-supervised deep belief networks is presented in this paper. We extract original feature with TFIDF-ICF, construct the text classification model based on DBN, and select the number of hidden layers and hidden units. Our experimental results indicated that the performance of text categorization algorithm based on deep belief networks is better than support vector machine.
Sijun Qin, Pengzhou Zhang
ICIS3
2016 The research on event extraction of Chinese news based on subject elements
abstract
Nowadays, with the rapid economic development, the amount of social information is also going up. Facing the daily explosive growth of the news quantity, the audience can difficultly get important information. To this end, the paper puts forward a method of Chinese news event extraction based on subject elements, which mixes the study of news topic sentence extraction and the research of event extraction together. According to the characteristics of news sentence, use dependency parsing to analyze the syntax. With the result of syntactic parsing to be a feature, distinguishably use Conditional Random Field algorithm (CRF) and Manual rules to identify the triggers in complex and simple sentences. Finally, Semantic Role Labeling algorithm (SRL) is used to identify the key elements of news events. What's more, the method will help readers quickly get the key elements from the long news, improving the efficiency of getting message.
Songhong Hong, Pengzhou Zhang
ICIS3