Qingmeng Zhu

dblp:151/5095 · DBLP profile ↗
← Back
22ranked-venue papers
8as first author
21since 2021 · last 2025
0000-0002-9099-6782ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Heterogeneous Packet Translation for Cross-Technology Communication
abstract
Recent advances in cross-technology communication (CTC) enable heterogeneous wireless devices (e.g., WiFi, Zig-Bee, and BLE) operating in the ISM band to communicate and understand each other. However, due to the limitation of standards and devices, existing CTC techniques need to design specific scheme for each heterogeneous wireless pair, which limit the practical applications of CTC. A key insight of this work is that heterogeneous packets could share same semantics but have different pattern features and representations, therefore, we explore the feasibility of packet-level heterogeneous communication translation. We first build a heterogeneous parallel frame dataset for evaluation. Then, we propose HPT, a heterogeneous packet translation approach. A wingman coding mechanism for wireless packet, dual attention module and synchronous interactive decoding method are designed as advanced designs to improve accuracy. Experiments between different translation directions show the proposed HPT achieves high accuracy of heterogeneous wireless packet translation.1
Qihuan Wu, Ziyin Gu, Qingmeng Zhu
ICASSP6
2025 MAP: Supporting Multimodal Knowledge Graph Completion via Augmented Modality Alignment and Instance Preserving
abstract
Multimodal knowledge graphs (KGs) have found widespread applications in data integration and processing, yet existing multimodal knowledge graphs are often highly incomplete, which impedes their wide adoption. Thereby multimodal knowledge graph completion (MKGC) has attracted widespread attention. However, the heterogeneity of multiple modalities degenerates the representations’ capacity to model modalityshared discriminative knowledge. The state-of-the-art approach addresses this challenge by aligning the modality distributions by adopting a Sinkhorn-based approach, but such an approach is computationally expensive and the practical sampling strategy largely degrades the model performance. Therefore we propose the augmented modality distribution alignment module, which imposes the generalized Radon transform-based approach to perform efficient and accurate distribution alignment. Yet the alignment may result in undesirable instance-level feature structure disorder. We thus propose the relation-aware instance preserving module. Empirical comparisons on well-established MKGC benchmarks demonstrate the effectiveness of proposed method.
Qingmeng Zhu, Changwen Zheng, Jiangmeng Li
ICASSP2
2025 Domain-Aware Knowledge Debiasing for Generalizable Video Understanding in CLIP
abstract
The pre-trained models contain multitudinous knowledge from huge amount of data. However, when applying these models to downstream tasks, they may mis-locate to wrong knowledge distribution due to a lack of domain or contextual knowledge. To address the distribution bias between the pre-trained model and the downstream domain, an innovative domain-aware knowledge de-biasing strategy, DKD, is introduced to improve the generalization performances on downstream tasks. Specifically, we use a CLIP-based video understanding framework to demonstrate the proposed approach, which dynamically adjusts the model’s representation space using the knowledge distribution of the target domain, effectively mitigating bias. Experimental results show that the method significantly improves model accuracy in action recognition tasks on standard datasets such as UCF101 and HMDB51, while also demonstrating superior generalization in cross-domain tasks. A comparison with state-of-the-art algorithms further validates the method’s remarkable advantages in the field of video understanding.
Qingmeng Zhu, Qihuan Wu, Ziyin Gu
ICASSP1
2024 MENTALER: Toward Professional Mental Health Support with LLMs via Multi-Role Collaboration
abstract
As mental health issues such as anxiety and depression are increasingly prevalent nowadays, we introduce MENTALER, an advanced multi-role collaboration framework specifically designed to enhance large language models (LLMs) in the diagnosis and treatment of mental health issues. In the MENTALER framework, the mental health support process comprises three specialized roles: Analyzer, Knowledge-Collector, and Strategy-Planner. Analyzer guides LLMs to achieve a in-depth diagnosis via multi-stage chain-of-thought prompting. Knowledge-Collector focus on involving domain-specific knowledge from exemplars retrieval. Strategy-Planner integrates professional support strategies into the generation process to further improve the professionalism among the generated texts. Through extensive automatic and human evaluations, we have validated that the mental health support counseling texts generated by MENTALER demonstrate a high degree of fluency and professionalism, closely aligning with real professional counseling texts. Our research advances the application of LLMs in the field of mental health support, providing an innovative and effective tool for those with psychological problems.
Ziyin Gu, Qingmeng Zhu
BIBM2
2024 MentalBlend: Enhancing Online Mental Health Support through the Integration of LLMs with Psychological Counseling Theories
Ziyin Gu, Qingmeng Zhu
CogSci2
2024 Analysis of Emotional Cognitive Capacity of Artificial Intelligence
abstract
Emotional cognitive capacity is crucial for AI applications. Existing methods generally rely on indirect evaluation approaches to measure AI’s emotional cognitive capacity, lacking effective means for comprehensive evaluation. Therefore, in this work, we introduce Emo-Lens, a more intuitive, accurate, and interpretable method for assessing AI’s emotional cognitive capacity. Emo-Lens calculates the relevant influence weights of entities for emotional judgment by using an XAI (Explainable Artificial Intelligence) algorithm, which are then used to calculate what we term ’significance redundancy.’ By calculating significance redundancy, Emo-Lens assesses the AI model’s ability to perform human-like emotional cognition, achieving a more intuitive, accurate, and interpretable assessment of AI’s emotional cognitive capacity.We also validate the accuracy and sensitivity of our Emo-Lens approach through several knowledge-enhancement experiments on the sentiment analysis task. We explore whether AI models with knowledge enhancement can identify correct emotional cues as humans do when making emotional judgments. Our results demonstrate that Emo-Lens is an accurate method for evaluating AI’s emotional cognitive capacity, as it reveals that AI’s emotional cognition can be improved through knowledge-enhancement. This enhancement process mirrors human behavior in terms of focusing attention and integrating background knowledge during emotional understanding. Emo-Lens has deepened our understanding of AI’s emotional cognitive capacity to some extent.
Ziyin Gu, Qingmeng Zhu, Hao He 0003, Tianxing Lan
CSCWD2
2024 Multi-Level Knowledge-Enhanced Prompting for Empathetic Dialogue Generation
abstract
Empathetic dialogue systems can recognize users’ emotions and provide appropriate responses, which are crucial for enhancing the user experience. However, existing empathetic dialogue systems often fall short in understanding some complex implicit emotions. To address this problem, we propose a multi-level knowledge-enhanced prompting approach to achieve more effective empathetic dialogue generation effect. We first acquire topic words and emotional keywords as low-level emotional knowledge. Next, we retrieve dialogue samples that are most similar in topic and emotional attributes, forming mid-level emotional knowledge. Subsequently, we guide a large language model (LLM) to generate high-level comprehensive emotional knowledge based on the information from the previous two levels and the dialogue context. Finally, based on the emotional knowledge, we further guide LLM to generate empathetic responses. The research results indicate that our multi-level knowledge-enhanced prompting approach outperforms other baselines.
Ziyin Gu, Qingmeng Zhu, Hao He 0003, Tianxing Lan
CSCWD2
2024 TraitsPrompt: Does Personality Traits Influence the Performance of a Large Language Model?
abstract
Large Language Models (LLMs) have showcased astounding capabilities across numerous tasks. Given that an LLM ingests vast amounts of data from human society, its behavior exhibits striking human-like characteristics. However, the ramifications of these anthropomorphic features on the model’s proficiency in specialized tasks remain an unresolved query. This paper delves into the impact of personality traits on model performance. Specifically, we employ the Big Five Personality Theory for model analysis, guiding questions in diverse datasets using prompts constructed around specific personality traits to study LLM outputs. In comparison to the baseline test questions, tasks aligned with the Big Five personality traits manifested superior performance, with enhancements ranging from 0.6% to 3.6%, mirroring human behavior closely. This underscores that for optimal execution of specific professional tasks, incorporating corresponding personality traits in prompts can be an effective strategy when querying the LLM.
Qingmeng Zhu, Tianxing Lan, Xiaoguang Xue, Hao He 0003
CSCWD1
2024 Multi-modal Knowledge Graph Link Prediction via Neural Optimal Transport
Tianxing Lan, Qingmeng Zhu, Yanan He, Hao He 0003
ICONIP (9)2
2024 Multimodal Feature Enhancement Plugin for Action Recognition
abstract
The advanced action recognition technologies have attempted to adopt multimodal architecture based on powerful pretrained models (e.g., visual-text) to achieve higher accuracy. However, visual encoders treat all frames equally, which may lead to (i) loss of key information and (ii) poor robustness, as different video frames may contribute differently to the visual representation. The preliminary experiments show that the assessment of frame inherent features could be highly correlated with the classification accuracy of the visual encoder. Based on this key insight, this paper proposes MEP (Multimodal-feature Enhanced Plugin), a frame selection method based on visual inherent features to guide the model in selecting video frames, thereby enhancing the accuracy of multimodal action recognition. Specifically, MEP first leverages a pretrained text encoder and a pretrained visual encoder to extract different modality features. For the pretrained visual encoder, we design a plug-and-play module for frame selection, which can use visual features as well as clustering methods to achieve efficient frame selection. Additionally, we introduce a module to extract textual inherent feature (i.e., emotional features) to enhance text representation. Experimental results show that the proposed method achieves performance comparable to or better than existing state-of-the-art video action recognition models on multiple datasets.
Qingmeng Zhu, Yanan He, Yukai Lu
IJCNN1
2024 Feature Alignment and Reconstruction Constraints for Multimodal Sentiment Analysis
abstract
Sentiment is complex feedback of human perception of the outside world, which contains important potential information. Traditional sentiment analysis methods often use unimodal processing, which cannot accurately recognize complex emotional expressions. The existing multimodal sentiment analysis (MSA) methods based on feature fusion and post-fusion paradigms do not pay attention to the accuracy of the semantic expression of unimodal features and do not mine the correlation relationship between heterogeneous features, which leads to serious deviation of the fused multimodal sentiment. In this paper, a novel MSA method based on feature alignment and reconstruction constraints is proposed. The multimodal feature alignment module utilizes the cross-attention mechanism to establish the intrinsic connection between multimodal features and reduce the differences between multimodal heterogeneous features with the same sentiment. The multimodal feature reconstruction module is used to retain the unique semantics of unimodal features and reduce the loss of key information during heterogeneous feature alignment. Experimental results on the CH-SIMS v2.0 dataset show that the proposed method can significantly improve the model’s ability to cope with the recognition of complex sentiments.
Qingmeng Zhu, Tianxing Lan, Deliang Xiang, Hao He 0003
IJCNN1
2024 MSI: Multi-modal Recommendation via Superfluous Semantics Discarding and Interaction Preserving
abstract
Multi-modal recommendation aims at leveraging data of auxiliary modalities (e.g., linguistic descriptions and images) to enhance the representations of items, thereby accurately recommending items that users prefer from the vast expanse of Web-based data. Current multi-modal recommendation methods typically utilize multi-modal features to assist in learning item representations in a direct manner. However, the superfluous semantics in multi-modal features are ignored, resulting in the inclusion of excessive redundancy within the representations of items. Moreover, we disclose that multi-modal features of items rarely contain user-item interaction information. Hence, during the interaction among different item features, the user-item interaction information in ID-based representations diminishes, leading to the degeneration of recommendation performance. To this end, we propose a novel multi-modal recommendation approach, which compresses representations of extra modalities under the guidance of solid theoretical analysis and leverages two auxiliary multi-modal graphs to integrate user-item interaction information into multi-modal features. Empirical experiments on three multi-modal recommendation datasets demonstrate that our method outperforms benchmarks.
Qingmeng Zhu, Changwen Zheng, Jiangmeng Li
ICMR2
2024 A Decoupling Video Frame Selection Method for Action Recognition
Qingmeng Zhu, Yanan He, Tianxing Lan, Ziyin Gu, Qihuan Wu, Hao He 0003
PRICAI (3)1
2024 Precise Knowledge Enhancement via CBR Framework for Empathetic Dialogue Generation
abstract
Empathetic dialogue systems are designed to capture emotions in conversations and provide appropriate emotional responses. Previous researches have indicated that integrating specific knowledge into empathetic dialogue systems can enhance the overall effectiveness of generating empathetic responses. Nevertheless, existing methods for knowledge-enhanced empathetic dialogue generation lack a focus on the precise selection of knowledge enhancement configurations for this specific task. To address this, we propose a Case-Based Reasoning (CBR) framework called CBR-KNOWLEDGE for autonomously select precise knowledge enhancement configurations tailored to specific empathetic dialogue contexts. Firstly, CBR-KNOWLEDGE establishes a case base that mirrors the overall quality of empathetic dialogues generated under various knowledge enhancement configurations. Subsequently, CBR-KNOWLEDGE employs an innovative text representation method, integrating an additional representation for words with noteworthy emotional impact. This approach facilitates the retrieval of analogous empathetic dialogues, enabling the reuse of their knowledge enhancement configurations to determine a new knowledge enhancement configuration. Ultimately, CBR-KNOWLEDGE employs this precise knowledge enhancement configuration for the purpose of empathetic dialogue generation. Experimental results demonstrate that CBR-KNOWLEDGE effectively enhances the performance of empathetic dialogue generation task.
Qingmeng Zhu, Ziyin Gu, Hao He 0003
SMC1
2023 Generative Event Extraction via Internal Knowledge-Enhanced Prompt Learning
Hetian Song, Qingmeng Zhu
ICANN (5)2
2023 3D Point Cloud Completion Based on Multi-Scale Degradation
abstract
Recent advances in 3D point cloud completion adopt unsupervised deep learning-based methods, which does not rely on labeled data and improves generalization ability. However, existing methods tend to focus more on the generation overall shape rather than detailed structure. To explore unsupervised 3D point cloud completion methods that give attention to both, we propose a multi resolution completion net (MRC-Net) which introduces a multi-scale degradation (KM- mask) and multi-discriminator into GAN inversion paradigm. First, degrade the point clouds completed by the generator under three different resolutions. Then, the losses of multi-stage reconstruction and feature matching by multi-scale discriminator are used to jointly optimize the generator. The experimental results demonstrate that MRC-Net outperforms existing unsupervised point cloud completion methods and has better completion performances on virtual scanning dataset.
Jianing Long, Qingmeng Zhu, Hao He 0003
ICASSP2
2023 MOC: Multi-modal Sentiment Analysis via Optimal Transport and Contrastive Interactions
Qingmeng Zhu, Hao He 0003, Ziyin Gu, Changwen Zheng
ICONIP (2)2
2023 Multi-Feature Enhanced Multimodal Action Recognition
abstract
Action recognition applications achieve impressive success in various fields, yet existing approaches cannot make good use of different information flows. To tackle such an issue, multimodal action recognition exhibit remarkable potential in strengthening the video representation. However, the direct fusion may not enable the model to learn appropriate knowledge. To leverage extra feature information flow to help improve model performance, this paper proposes MFE, a multi-feature enhanced multimodal action recognition model. MFE introduces a guiding tag to explicit guide different extra feature information flow and designs a feature fusion module to fuse different information flows, to enhance hidden representation through more semantic supervision. The experimental results show that MFE achieves better or comparable accuracies with some advanced video action recognition models on several action recognition datasets.
Qingmeng Zhu, Ziyin Gu, Hao He 0003, Tianci Zhao
SMC1
2022 PromptFusion: A Low-Cost Prompt-Based Task Composition for Multi-task Learning
Hetian Song, Hao He 0003, Qingmeng Zhu, Xiaoguang Xue
ICONIP (1)3
2022 Implicit and Explicit Emotion Enhanced Empathetic Dialogue Generation
abstract
Empathetic conversation systems identify the users' emotions and give appropriate responses, which is crucial to improve users' experiences. However, existing empathetic dialogue models (especially to the dominant pre-trained language model-based systems) did not focus on modelling the holistic properties of implicit and explicit emotions. In this paper, we propose an Implicit and Explicit Emotion Enhanced (IEEE) empathetic dialogue generation model to handle such challenges. Specifically, we first propose a prompt tuning-based approach to mine emotional words as additional information to obtain the users' explicit emotion. A variational auto-encoder is then introduced to extract the topic words of the input sequence as additional priori knowledge to get the implicit emotion related information. Finally, a pre-trained language model is utilized as the auto-regressive decoder to generate empathetic responses related to the content of the topics and user emotions. To demonstrate the effectiveness of the proposed approach, IEEE has been tested on empathic dialogue dataset. The experimental results show that our method achieves better performance than some competitive models.
Qingmeng Zhu, Hao He 0003, Hetian Song, Ziyin Gu, Wenjing Ying
ICTAI1
2022 Rotating Target Detection Based on Lightweight Network
Yunxu Jiao, Qingmeng Zhu, Hao He 0003, Tianci Zhao, Haihui Wang
PRICAI (3)2
2014 Ship Damage Control as a Service Based on Spatio-temporal Database
abstract
To solve the increasingly prominent contradiction between the traditional damage control and the demand of high efficiency and reliability of ship system, a ship damage control system based on spatio-temporal database is presented and accomplished with cloud solution. A path planning algorithm based on Dijkstra is proposed to meet the dynamic road network as in the fire rescue scenario. The binary group is adopted to describe the weight of the path, and path network is pruning to reduce nodes accessed in advance. The simulation results show that the proposed algorithm improves the efficiency, capacity, intelligence and user experience, and provide efficient support for assistant decision-making.
Yicheng Zheng, Qingmeng Zhu
IC2E3