VLDB 2026 Research / reviewers in the wild / expert
Ziyin Gu
dblp:307/6575
· DBLP profile ↗
14ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Group Causal Policy Optimization for Post-Training Large Language ModelsabstractRecent advances in large language models (LLMs) have broadened their applicability across diverse tasks, yet specialized domains still require targeted post-training. Among existing methods, Group Relative Policy Optimization (GRPO) stands out for its efficiency, leveraging groupwise relative rewards while avoiding costly value function learning. However, GRPO treats candidate responses as independent, overlooking semantic interactions such as complementarity and contradiction. To address this challenge, we first introduce a Structural Causal Model (SCM) that reveals hidden dependencies among candidate responses induced by conditioning on a final integrated output, forming a collider structure. Then, our causal analysis leads to two insights: (1) projecting responses onto a causally-informed subspace improves prediction quality, and (2) this projection yields a better baseline than query-only conditioning. Building on these insights, we propose Group Causal Policy Optimization (GCPO), which integrates causal structure into optimization through two key components: a causally-informed reward adjustment and a novel KL-regularization term that aligns the policy with a causally-projected reference distribution. Comprehensive experimental evaluations on various benchmarks demonstrate that GCPO consistently surpasses existing methods. Ziyin Gu, Ran Zuo, Chuxiong Sun, Zeen Song, Changwen Zheng, Wenwen Qiang |
AAAI | 1 |
| 2026 | On the Transferability and Discriminability of Representation Learning in Unsupervised Domain AdaptationabstractIn this paper, we addressed the limitation of relying solely on distribution alignment and source-domain empirical risk minimization in Unsupervised Domain Adaptation (UDA). Our information-theoretic analysis showed that this standard adversarial-based framework neglects the discriminability of target-domain features, leading to suboptimal performance. To bridge this theoretical-practical gap, we defined "good representation learning" as guaranteeing both transferability and discriminability, and proved that an additional loss term targeting target-domain discriminability is necessary. Building on these insights, we proposed a novel adversarial-based UDA framework that explicitly integrates a domain alignment objective with a discriminability-enhancing constraint. Instantiated as Domain-Invariant Representation Learning with Global and Local Consistency (RLGLC), our method leverages Asymmetrically-Relaxed Wasserstein of Wasserstein Distance (AR-WWD) to address class imbalance and semantic dimension weighting, and employs a local consistency mechanism to preserve fine-grained target-domain discriminative information. Extensive experiments across multiple benchmark datasets demonstrate that RLGLC consistently surpasses state-of-the-art methods, confirming the value of our theoretical perspective and underscoring the necessity of enforcing both transferability and discriminability in adversarial-based UDA. Wenwen Qiang, Ziyin Gu, Lingyu Si, Jiangmeng Li, Changwen Zheng, Fuchun Sun 0001, Hui Xiong 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Heterogeneous Packet Translation for Cross-Technology CommunicationabstractRecent advances in cross-technology communication (CTC) enable heterogeneous wireless devices (e.g., WiFi, Zig-Bee, and BLE) operating in the ISM band to communicate and understand each other. However, due to the limitation of standards and devices, existing CTC techniques need to design specific scheme for each heterogeneous wireless pair, which limit the practical applications of CTC. A key insight of this work is that heterogeneous packets could share same semantics but have different pattern features and representations, therefore, we explore the feasibility of packet-level heterogeneous communication translation. We first build a heterogeneous parallel frame dataset for evaluation. Then, we propose HPT, a heterogeneous packet translation approach. A wingman coding mechanism for wireless packet, dual attention module and synchronous interactive decoding method are designed as advanced designs to improve accuracy. Experiments between different translation directions show the proposed HPT achieves high accuracy of heterogeneous wireless packet translation.1 Qihuan Wu, Ziyin Gu, Qingmeng Zhu |
ICASSP | 4 |
| 2025 | Domain-Aware Knowledge Debiasing for Generalizable Video Understanding in CLIPabstractThe pre-trained models contain multitudinous knowledge from huge amount of data. However, when applying these models to downstream tasks, they may mis-locate to wrong knowledge distribution due to a lack of domain or contextual knowledge. To address the distribution bias between the pre-trained model and the downstream domain, an innovative domain-aware knowledge de-biasing strategy, DKD, is introduced to improve the generalization performances on downstream tasks. Specifically, we use a CLIP-based video understanding framework to demonstrate the proposed approach, which dynamically adjusts the model’s representation space using the knowledge distribution of the target domain, effectively mitigating bias. Experimental results show that the method significantly improves model accuracy in action recognition tasks on standard datasets such as UCF101 and HMDB51, while also demonstrating superior generalization in cross-domain tasks. A comparison with state-of-the-art algorithms further validates the method’s remarkable advantages in the field of video understanding. Qingmeng Zhu, Qihuan Wu, Ziyin Gu |
ICASSP | 5 |
| 2025 | On the Generalization and Causal Explanation in Self-Supervised Learning
Wenwen Qiang, Zeen Song, Ziyin Gu, Jiangmeng Li, Changwen Zheng, Fuchun Sun 0001, Hui Xiong 0001 |
Int. J. Comput. Vis. | 3 |
| 2024 | MENTALER: Toward Professional Mental Health Support with LLMs via Multi-Role CollaborationabstractAs mental health issues such as anxiety and depression are increasingly prevalent nowadays, we introduce MENTALER, an advanced multi-role collaboration framework specifically designed to enhance large language models (LLMs) in the diagnosis and treatment of mental health issues. In the MENTALER framework, the mental health support process comprises three specialized roles: Analyzer, Knowledge-Collector, and Strategy-Planner. Analyzer guides LLMs to achieve a in-depth diagnosis via multi-stage chain-of-thought prompting. Knowledge-Collector focus on involving domain-specific knowledge from exemplars retrieval. Strategy-Planner integrates professional support strategies into the generation process to further improve the professionalism among the generated texts. Through extensive automatic and human evaluations, we have validated that the mental health support counseling texts generated by MENTALER demonstrate a high degree of fluency and professionalism, closely aligning with real professional counseling texts. Our research advances the application of LLMs in the field of mental health support, providing an innovative and effective tool for those with psychological problems. Ziyin Gu, Qingmeng Zhu |
BIBM | 1 |
| 2024 | MentalBlend: Enhancing Online Mental Health Support through the Integration of LLMs with Psychological Counseling Theories
Ziyin Gu, Qingmeng Zhu |
CogSci | 1 |
| 2024 | Analysis of Emotional Cognitive Capacity of Artificial IntelligenceabstractEmotional cognitive capacity is crucial for AI applications. Existing methods generally rely on indirect evaluation approaches to measure AI’s emotional cognitive capacity, lacking effective means for comprehensive evaluation. Therefore, in this work, we introduce Emo-Lens, a more intuitive, accurate, and interpretable method for assessing AI’s emotional cognitive capacity. Emo-Lens calculates the relevant influence weights of entities for emotional judgment by using an XAI (Explainable Artificial Intelligence) algorithm, which are then used to calculate what we term ’significance redundancy.’ By calculating significance redundancy, Emo-Lens assesses the AI model’s ability to perform human-like emotional cognition, achieving a more intuitive, accurate, and interpretable assessment of AI’s emotional cognitive capacity.We also validate the accuracy and sensitivity of our Emo-Lens approach through several knowledge-enhancement experiments on the sentiment analysis task. We explore whether AI models with knowledge enhancement can identify correct emotional cues as humans do when making emotional judgments. Our results demonstrate that Emo-Lens is an accurate method for evaluating AI’s emotional cognitive capacity, as it reveals that AI’s emotional cognition can be improved through knowledge-enhancement. This enhancement process mirrors human behavior in terms of focusing attention and integrating background knowledge during emotional understanding. Emo-Lens has deepened our understanding of AI’s emotional cognitive capacity to some extent. Ziyin Gu, Qingmeng Zhu, Hao He 0003, Tianxing Lan |
CSCWD | 1 |
| 2024 | Multi-Level Knowledge-Enhanced Prompting for Empathetic Dialogue GenerationabstractEmpathetic dialogue systems can recognize users’ emotions and provide appropriate responses, which are crucial for enhancing the user experience. However, existing empathetic dialogue systems often fall short in understanding some complex implicit emotions. To address this problem, we propose a multi-level knowledge-enhanced prompting approach to achieve more effective empathetic dialogue generation effect. We first acquire topic words and emotional keywords as low-level emotional knowledge. Next, we retrieve dialogue samples that are most similar in topic and emotional attributes, forming mid-level emotional knowledge. Subsequently, we guide a large language model (LLM) to generate high-level comprehensive emotional knowledge based on the information from the previous two levels and the dialogue context. Finally, based on the emotional knowledge, we further guide LLM to generate empathetic responses. The research results indicate that our multi-level knowledge-enhanced prompting approach outperforms other baselines. Ziyin Gu, Qingmeng Zhu, Hao He 0003, Tianxing Lan |
CSCWD | 1 |
| 2024 | A Decoupling Video Frame Selection Method for Action Recognition
Qingmeng Zhu, Yanan He, Tianxing Lan, Ziyin Gu, Qihuan Wu, Hao He 0003 |
PRICAI (3) | 4 |
| 2024 | Precise Knowledge Enhancement via CBR Framework for Empathetic Dialogue GenerationabstractEmpathetic dialogue systems are designed to capture emotions in conversations and provide appropriate emotional responses. Previous researches have indicated that integrating specific knowledge into empathetic dialogue systems can enhance the overall effectiveness of generating empathetic responses. Nevertheless, existing methods for knowledge-enhanced empathetic dialogue generation lack a focus on the precise selection of knowledge enhancement configurations for this specific task. To address this, we propose a Case-Based Reasoning (CBR) framework called CBR-KNOWLEDGE for autonomously select precise knowledge enhancement configurations tailored to specific empathetic dialogue contexts. Firstly, CBR-KNOWLEDGE establishes a case base that mirrors the overall quality of empathetic dialogues generated under various knowledge enhancement configurations. Subsequently, CBR-KNOWLEDGE employs an innovative text representation method, integrating an additional representation for words with noteworthy emotional impact. This approach facilitates the retrieval of analogous empathetic dialogues, enabling the reuse of their knowledge enhancement configurations to determine a new knowledge enhancement configuration. Ultimately, CBR-KNOWLEDGE employs this precise knowledge enhancement configuration for the purpose of empathetic dialogue generation. Experimental results demonstrate that CBR-KNOWLEDGE effectively enhances the performance of empathetic dialogue generation task. Qingmeng Zhu, Ziyin Gu, Hao He 0003 |
SMC | 2 |
| 2023 | MOC: Multi-modal Sentiment Analysis via Optimal Transport and Contrastive Interactions
Qingmeng Zhu, Hao He 0003, Ziyin Gu, Changwen Zheng |
ICONIP (2) | 4 |
| 2023 | Multi-Feature Enhanced Multimodal Action RecognitionabstractAction recognition applications achieve impressive success in various fields, yet existing approaches cannot make good use of different information flows. To tackle such an issue, multimodal action recognition exhibit remarkable potential in strengthening the video representation. However, the direct fusion may not enable the model to learn appropriate knowledge. To leverage extra feature information flow to help improve model performance, this paper proposes MFE, a multi-feature enhanced multimodal action recognition model. MFE introduces a guiding tag to explicit guide different extra feature information flow and designs a feature fusion module to fuse different information flows, to enhance hidden representation through more semantic supervision. The experimental results show that MFE achieves better or comparable accuracies with some advanced video action recognition models on several action recognition datasets. Qingmeng Zhu, Ziyin Gu, Hao He 0003, Tianci Zhao |
SMC | 2 |
| 2022 | Implicit and Explicit Emotion Enhanced Empathetic Dialogue GenerationabstractEmpathetic conversation systems identify the users' emotions and give appropriate responses, which is crucial to improve users' experiences. However, existing empathetic dialogue models (especially to the dominant pre-trained language model-based systems) did not focus on modelling the holistic properties of implicit and explicit emotions. In this paper, we propose an Implicit and Explicit Emotion Enhanced (IEEE) empathetic dialogue generation model to handle such challenges. Specifically, we first propose a prompt tuning-based approach to mine emotional words as additional information to obtain the users' explicit emotion. A variational auto-encoder is then introduced to extract the topic words of the input sequence as additional priori knowledge to get the implicit emotion related information. Finally, a pre-trained language model is utilized as the auto-regressive decoder to generate empathetic responses related to the content of the topics and user emotions. To demonstrate the effectiveness of the proposed approach, IEEE has been tested on empathic dialogue dataset. The experimental results show that our method achieves better performance than some competitive models. Qingmeng Zhu, Hao He 0003, Hetian Song, Ziyin Gu, Wenjing Ying |
ICTAI | 5 |