Zhuang Chen 0002

dblp:54/8492-2 · DBLP profile ↗
← Back
23ranked-venue papers
10as first author
18since 2021 · last 2026
0000-0002-7048-7833ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 9 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Unveiling the Landscape of Clinical Depression Assessment: From Behavioral Signatures to Psychiatric Reasoning
abstract
Depression is a widespread mental disorder that affects millions worldwide. While automated depression assessment shows promise, most studies rely on limited or non-clinically validated data, and often prioritize complex model design over real-world effectiveness. In this paper, we aim to unveil the landscape of clinical depression assessment. We introduce C-MIND, a clinical multimodal neuropsychiatric diagnosis dataset collected over two years from real hospital visits. Each participant completes three structured psychiatric tasks and receives a final diagnosis from expert clinicians, with informative audio, video, transcript, and functional near-infrared spectroscopy (fNIRS) signals recorded. Using C-MIND, we first analyze behavioral signatures relevant to diagnosis. We train a range of classical models to quantify how different tasks and modalities contribute to diagnostic performance, and dissect the effectiveness of their combinations. We then explore whether LLMs can perform psychiatric reasoning like clinicians and identify their clear limitations in realistic clinical settings. In response, we propose to guide the reasoning process with clinical expertise and consistently improve LLM diagnostic performance by up to 10% in Macro-F1 score. We aim to build an infrastructure for clinical depression assessment from both data and algorithmic perspectives, enabling C-MIND to facilitate grounded and reliable research for mental healthcare.
Zhuang Chen 0002, Guanqun Bi, Aoyun Wang, Xiyao Xiao, Minlie Huang
AAAI1
2026 Ψ-Arena: Interactive Assessment and Optimization of LLM-based Psychological Counselors with Tripartite Feedback
abstract
Large language models (LLMs) have shown promise in providing scalable mental health support, while evaluating their counseling capability remains crucial to ensure both efficacy and safety. Existing evaluations are limited by the static assessment that focuses on knowledge tests, the single perspective that centers on user experience, and the open-loop framework that lacks actionable feedback. To address these issues, we propose Ψ-Arena, an interactive framework for comprehensive assessment and optimization of LLM-based counselors, featuring three key characteristics: (1) Realistic arena interactions that simulate real-world counseling through multi-stage dialogues with psychologically profiled NPC clients; (2) Tripartite evaluation that integrates assessments from the client, supervisor, and counselor perspectives; (3) Closed-loop optimization that iteratively improves LLM counselors using diagnostic feedback. Experiments across eight state-of-the-art LLMs show significant performance variations in different real-world scenarios and evaluation perspectives. Moreover, reflection-based optimization results in up to a 141% improvement in counseling performance. We hope Ψ-Arena provides a foundational resource for advancing reliable and human-aligned LLM applications in mental healthcare.
Shijing Zhu, Zhuang Chen 0002, Guanqun Bi, Binghang Li, Yaxi Deng, Dazhen Wan, Libiao Peng, Xiyao Xiao, Tangjie Lv, Zhipeng Hu, Minlie Huang
AAAI2
2026 S⌃4: Operationalizing Speech Act Theory for Strategic Semi-Structured Psychiatric Interview
abstract
Psychiatric interviewing is a strategic, goaloriented interaction that requires proactively steering the conversation to elicit latent information.However, existing methods often degenerate into rigid interrogation or aimless chitchat due to a lack of strategic planning.In this work, we introduce S 4 , a comprehensive framework grounded in Speech Act Theory, modeling the interview as a unified process of internal strategy (Illocution and Perlocution) and external realization (Locution).We synthesize a large-scale dataset with fine-grained psychiatric speech act annotations.Trained on this data, S 4 employs reinforcement learning driven by long-term therapeutic effects to optimize the strategic chaining of atomic acts, aiming to maximally elicit information and maintain patient engagement.Experiments demonstrate that S 4 significantly outperforms baselines, validating the effectiveness of our effectdriven strategic modeling.Action (A) Sample Locution (L) Definition & Intended Perlocution (P) I. Information Seeking (Directives: Eliciting Disclosure) Explore "How have you been sleeping?"Solicit Narrative: Ask open-ended questions to elicit detailed disclosure and expand symptom scope.Probe "Could you tell me more about that?" Deepen Inquiry: Follow up on ambiguity to clarify details and deepen focus.Confirm "Do you feel this way every day?" Pinpoint Fact: Ask closed-ended questions to verify diagnostic criteria.Clarify "By 'fatigue', I mean tiredness."Resolve Confusion: Provide explanations to align cognition and correct misunderstandings.II.Affective Regulation (Expressives: Modifying State) Validate "That sounds incredibly hard."Affirm Emotion: Acknowledge patient distress to lower defensiveness and build trust.Support "I understand.Please go on."Maintain Flow: Use back-channeling to demonstrate active listening and boost efficacy.Ease "Do you have any hobbies?"Reduce Tension: Engage in non-clinical conversation to de-escalate anxiety and humanize the agent. III. Interview Management (Representatives: Setting Frame)Initiate "Hi, I'm your AI counselor."Set Frame: Establish professional boundaries and the purpose of the session.Conclude "Thanks for sharing.Take care."Ensure Closure: Formally end the session to provide a safe exit and consolidation.
Guanqun Bi, Zhoufu Liu, Zhuang Chen 0002, Dazhen Wan, Xiyao Xiao, Minlie Huang
ACL (1)3
2025 SocialSim: Towards Socialized Simulation of Emotional Support Conversation
abstract
Emotional support conversation (ESC) helps reduce people's psychological stress and provide emotional value through interactive dialogues. Due to the high cost of crowdsourcing a large ESC corpus, recent attempts use large language models for dialogue augmentation. However, existing approaches largely overlook the social dynamics inherent in ESC, leading to less effective simulations. In this paper, we introduce SocialSim, a novel framework that simulates ESC by integrating key aspects of social interactions: social disclosure and social awareness. On the seeker side, we facilitate social disclosure by constructing a comprehensive persona bank that captures diverse and authentic help-seeking scenarios. On the supporter side, we enhance social awareness by eliciting cognitive reasoning to generate logical and supportive responses. Building upon SocialSim, we construct SSConv, a large-scale synthetic ESC corpus of which quality can even surpass crowdsourced ESC data. We further train a chatbot on SSConv and demonstrate its state-of-the-art performance in both automatic and human evaluations. We believe SocialSim offers a scalable way to synthesize ESC, making emotional care more accessible and practical.
Zhuang Chen 0002, Yaru Cao, Guanqun Bi, Jincenzi Wu, Jinfeng Zhou, Xiyao Xiao, Hongning Wang, Minlie Huang
AAAI1
2025 SS-GEN: A Social Story Generation Framework with Large Language Models
abstract
Children with Autism Spectrum Disorder (ASD) often misunderstand social situations and struggle to participate in daily routines. Social Stories™ are traditionally crafted by psychology experts under strict constraints to address these challenges but are costly and limited in diversity. As Large Language Models (LLMs) advance, there's an opportunity to develop more automated, affordable, and accessible methods to generate Social Stories in real-time with broad coverage. However, adapting LLMs to meet the unique and strict constraints of Social Stories is a challenging issue. To this end, we propose SS-GEN, a Social Story GENeration framework with LLMs. Firstly, we develop a constraint-driven sophisticated strategy named StarSow to hierarchically prompt LLMs to generate Social Stories at scale, followed by rigorous human filtering to build a high-quality dataset. Additionally, we introduce quality assessment criteria to evaluate the effectiveness of these generated stories. Considering that powerful closed-source large models require very complex instructions and expensive API fees, we finally fine-tune smaller language models with our curated high-quality dataset, achieving comparable results at lower costs and with simpler instruction and deployment. This work marks a significant step in leveraging AI to personalize Social Stories cost-effectively for autistic children at scale, which we hope can encourage future research on special groups.
Jiaqi Wang 0006, Zhuang Chen 0002, Guanqun Bi, Minlie Huang, Liping Jing, Jian Yu 0001
AAAI4
2025 CharacterBench: Benchmarking Character Customization of Large Language Models
abstract
Character-based dialogue (aka role-playing) enables users to freely customize characters for interaction, which often relies on LLMs, raising the need to evaluate LLMs’ character customization capability. However, existing benchmarks fail to ensure a robust evaluation as they often only involve a single character category or evaluate limited dimensions. Moreover, the sparsity of character features in responses makes feature-focused generative evaluation both ineffective and inefficient. To address these issues, we propose CharacterBench, the largest bilingual generative benchmark, with 22,859 human-annotated samples covering 3,956 characters from 25 detailed character categories. We define 11 dimensions of 6 aspects, classified as sparse and dense dimensions based on whether character features evaluated by specific dimensions manifest in each response. We enable effective and efficient evaluation by crafting tailored queries for each dimension to induce characters’ responses related to specific dimensions. Further, we develop CharacterJudge model for cost-effective and stable evaluations. Experiments show its superiority over SOTA automatic judges (e.g., GPT-4) and our benchmark’s potential to optimize LLMs’ character customization.
Jinfeng Zhou, Yongkang Huang, Bosi Wen, Guanqun Bi, Pei Ke, Zhuang Chen 0002, Xiyao Xiao, Libiao Peng, Kuntian Tang, Tangjie Lv, Zhipeng Hu, Hongning Wang, Minlie Huang
AAAI7
2025 Reframe Your Life Story: Interactive Narrative Therapist and Innovative Moment Assessment with Large Language Models
abstract
Yi Feng, Jiaqi Wang, Wenxuan Zhang, Zhuang Chen, Shen Yutong, Xiyao Xiao, Minlie Huang, Liping Jing, Jian Yu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jiaqi Wang 0006, Zhuang Chen 0002, Yutong Shen, Xiyao Xiao, Minlie Huang, Liping Jing, Jian Yu 0001
EMNLP4
2025 Generative Meta-Learning for Zero-Shot Relation Triplet Extraction
abstract
Zero-shot Relation Triplet Extraction (ZeroRTE) aims to extract relation triplets from texts containing unseen relation types. This capability benefits various downstream information retrieval (IR) tasks. The primary challenge lies in enabling models to generalize effectively to unseen relation categories. Existing approaches typically leverage the knowledge embedded in pre-trained language models to accomplish the generalization process. However, these methods focus solely on fitting the training data during training, without specifically improving the model's generalization performance, resulting in limited generalization capability. For this reason, we explore the integration of bi-level optimization (BLO) with pre-trained language models for learning generalized knowledge directly from the training data, and propose a generative meta-learning framework which exploits the 'learning-to-learn' ability of meta-learning to boost the generalization capability of generative models.
Wanli Li 0002, Tieyun Qian, Zeyu Zhang 0004, Jiawei Li 0008, Zhuang Chen 0002, Lixin Zou
SIGIR6
2024 ToMBench: Benchmarking Theory of Mind in Large Language Models
abstract
Zhuang Chen, Jincenzi Wu, Jinfeng Zhou, Bosi Wen, Guanqun Bi, Gongyao Jiang, Yaru Cao, Mengting Hu, Yunghwei Lai, Zexuan Xiong, Minlie Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zhuang Chen 0002, Jincenzi Wu, Jinfeng Zhou, Bosi Wen, Guanqun Bi, Gongyao Jiang, Yaru Cao, Mengting Hu 0002, Yunghwei Lai, Zexuan Xiong, Minlie Huang
ACL (1)1
2024 COKE: A Cognitive Knowledge Graph for Machine Theory of Mind
abstract
Theory of mind (ToM) refers to humans' ability to understand and infer the desires, beliefs, and intentions of others.The acquisition of ToM plays a key role in humans' social cognition and interpersonal relations.Though indispensable for social intelligence, ToM is still lacking for modern AI and NLP systems since they cannot access the human mental state and cognitive process beneath the training corpus.To empower AI systems with the ToM ability and narrow the gap between them and humans, in this paper, we propose COKE: the first cognitive knowledge graph for machine theory of mind, formalizing cognitive processes as a chained structure.Specifically, COKE formalizes ToM as a collection of 45k+ manually verified cognitive chains that characterize human mental activities and subsequent behavioral/affective responses when facing specific social circumstances.In addition, we further generalize COKE using LLMs and build a powerful generation model COLM tailored for cognitive reasoning.Experimental results in both automatic and human evaluation demonstrate the high quality of COKE, the superior ToM ability of COLM, and its potential to significantly enhance social applications.We release our code and data at https://github.com/jincenziwu/COKE.
Jincenzi Wu, Zhuang Chen 0002, Jiawen Deng 0006, Sahand Sabour, Helen M. Meng, Minlie Huang
ACL (1)2
2024 Depression Detection in Clinical Interviews with LLM-Empowered Structural Element Graph
abstract
Zhuang Chen, Jiawen Deng, Jinfeng Zhou, Jincenzi Wu, Tieyun Qian, Minlie Huang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zhuang Chen 0002, Jiawen Deng 0006, Jinfeng Zhou, Jincenzi Wu, Tieyun Qian, Minlie Huang
NAACL-HLT1
2024 Enhancing Emotional Support Conversation with Cognitive Chain-of-Thought Reasoning
Yaru Cao, Zhuang Chen 0002, Guanqun Bi, Yulin Feng, Fucheng Wan, Minlie Huang, Hongzhi Yu
NLPCC (1)2
2023 Facilitating Multi-turn Emotional Support Conversation with Positive Emotion Elicitation: A Reinforcement Learning Approach
abstract
Emotional support conversation (ESC) aims to provide emotional support (ES) to improve one's mental state.Existing works stay at fitting grounded responses and responding strategies (e.g., question), which ignore the effect on ES and lack explicit goals to guide emotional positive transition.To this end, we introduce a new paradigm to formalize multi-turn ESC as a process of positive emotion elicitation.Addressing this task requires finely adjusting the elicitation intensity in ES as the conversation progresses while maintaining conversational goals like coherence.In this paper, we propose SUPPORTER, a mixture-of-expert-based reinforcement learning model, and well design ES and dialogue coherence rewards to guide policy's learning for responding.Experiments verify the superiority of SUPPORTER in achieving positive emotion elicitation during responding while maintaining conversational goals including coherence.
Jinfeng Zhou, Zhuang Chen 0002, Minlie Huang
ACL (1)2
2023 Mimicking the Thinking Process for Emotion Recognition in Conversation with Prompts and Paraphrasing
abstract
Emotion recognition in conversation, which aims to predict the emotion for all utterances, has attracted considerable research attention in recent years. It is a challenging task since the recognition of the emotion in one utterance involves many complex factors, such as the conversational context, the speaker's background, and the subtle difference between emotion labels. In this paper, we propose a novel framework which mimics the thinking process when modeling these factors. Specifically, we first comprehend the conversational context with a history-oriented prompt to selectively gather information from predecessors of the target utterance. We then model the speaker's background with an experience-oriented prompt to retrieve the similar utterances from all conversations. We finally differentiate the subtle label semantics with a paraphrasing mechanism to elicit the intrinsic label related knowledge. We conducted extensive experiments on three benchmarks. The empirical results demonstrate the superiority of our proposed framework over the state-of-the-art baselines.
Zhuang Chen 0002, Ming Zhong 0002, Tieyun Qian
IJCAI2
2022 Retrieve-and-Edit Domain Adaptation for End2End Aspect Based Sentiment Analysis
abstract
End-to-end aspect based sentiment analysis (E2E-ABSA) aims to jointly extract aspect terms and predict aspect-level sentiment for opinion reviews. Though supervised methods show effectiveness for E2E-ABSA tasks, the annotation cost is extremely high due to the necessity of fine-grained labels. Recent attempts alleviate this problem using the domain adaptation technique to transfer the word-level common knowledge across domains. However, the biggest issue in domain adaptation, i.e., how to transfer the domain-specific words likepizzaanddeliciousin the source “Restaurant” to the target “Laptop” domain, has not been resolved. In this paper, we propose a novel domain adaptation method to address this issue by enhancing the transferability of domain-specific source words in a retrieve-and-edit way. Specifically, for all source words, we first retrieve the transferable prototypes from unlabeled target data via their syntactic and semantic roles. We then edit the source words to enhance their transferability by absorbing the knowledge carried in prototypes. Finally, we design an end-to-end framework to jointly accomplish cross-domain aspect term extraction and aspect-level sentiment classification. We conduct extensive experiments on four real-world datasets. The results prove that, by introducing transferable prototypes, our method significantly outperforms the state-of-the-art methods, achieving an absolute 3.95% F1 increase over the best baseline.
Zhuang Chen 0002, Tieyun Qian
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 Description and demonstration guided data augmentation for sequence tagging
Zhuang Chen 0002, Tieyun Qian
World Wide Web1
2021 Bridge-Based Active Domain Adaptation for Aspect Term Extraction
abstract
Zhuang Chen, Tieyun Qian. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Zhuang Chen 0002, Tieyun Qian
ACL/IJCNLP (1)1
2021 Generating Pseudo Connectives with MLMs for Implicit Discourse Relation Recognition
Congcong Jiang, Tieyun Qian, Zhuang Chen 0002, Kejian Tang, Shaohui Zhan
PRICAI (2)3
2020 Relation-Aware Collaborative Learning for Unified Aspect-Based Sentiment Analysis
abstract
Aspect-based sentiment analysis (ABSA) involves three subtasks, i.e., aspect term extraction, opinion term extraction, and aspect-level sentiment classification.Most existing studies focused on one of these subtasks only.Several recent researches made successful attempts to solve the complete ABSA problem with a unified framework.However, the interactive relations among three subtasks are still underexploited.We argue that such relations encode collaborative signals between different subtasks.For example, when the opinion term is "delicious", the aspect term must be "food" rather than "place".In order to fully exploit these relations, we propose a Relation-Aware Collaborative Learning (RACL) framework which allows the subtasks to work coordinately via the multi-task learning and relation propagation mechanisms in a stacked multi-layer network.Extensive experiments on three real-world datasets demonstrate that RACL significantly outperforms the state-ofthe-art methods for the complete ABSA task.
Zhuang Chen 0002, Tieyun Qian
ACL1
2020 Enhancing Aspect Term Extraction with Soft Prototypes
abstract
Aspect term extraction (ATE) aims to extract aspect terms from a review sentence that users have expressed opinions on.Existing studies mostly focus on designing neural sequence taggers to extract linguistic features from the token level.However, since the aspect terms and context words usually exhibit long-tail distributions, these taggers often converge to an inferior state without enough sample exposure.In this paper, we propose to tackle this problem by correlating words with each other through soft prototypes.These prototypes, generated by a soft retrieval process, can introduce global knowledge from internal or external data and serve as the supporting evidence for discovering the aspect terms.Our proposed model is a general framework and can be combined with almost all sequence taggers.Experiments on four SemEval datasets show that our model boosts the performance of three typical ATE methods by a large margin.
Zhuang Chen 0002, Tieyun Qian
EMNLP (1)1
2019 Transfer Capsule Network for Aspect Level Sentiment Classification
abstract
Aspect-level sentiment classification aims to determine the sentiment polarity of a sentence towards an aspect.Due to the high cost in annotation, the lack of aspect-level labeled data becomes a major obstacle in this area.On the other hand, document-level labeled data like reviews are easily accessible from online websites.These reviews encode sentiment knowledge in abundant contexts.In this paper, we propose a Transfer Capsule Network (Tran-sCap) model for transferring document-level knowledge to aspect-level sentiment classification.To this end, we first develop an aspect routing approach to encapsulate the sentence-level semantic representations into semantic capsules from both aspect-level and document-level data.We then extend the dynamic routing approach to adaptively couple the semantic capsules with the class capsules under the transfer learning framework.Experiments on SemEval datasets demonstrate the effectiveness of TransCap.
Zhuang Chen 0002, Tieyun Qian
ACL (1)1
2019 Aspect-Level Sentiment Classification with Dependency Rules and Dual Attention
Yunkai Yang, Tieyun Qian, Zhuang Chen 0002
ICONIP (2)3
2019 Aspect Aware Learning for Aspect Category Sentiment Analysis
abstract
Aspect category sentiment analysis (ACSA) is an underexploited subtask in aspect level sentiment analysis. It aims to identify the sentiment of predefined aspect categories. The main challenge in ACSA comes from the fact that the aspect category may not occur in the sentence in most of the cases. For example, the review “ they have delicious sandwiches ” positively talks about the aspect category “ food ” in an implicit manner. In this article, we propose a novel aspect aware learning (AAL) framework for ACSA tasks. Our key idea is to exploit the interaction between the aspect category and the contents under the guidance of both sentiment polarity and predefined categories. To this end, we design a two-way memory network for integrating AAL into the framework of sentiment classification. We further present two algorithms to incorporate the potential impacts of aspect categories. One is to capture the correlations between aspect terms and the aspect category like “sandwiches” and “food.” The other is to recognize the aspect category for sentiment representations like “food” for “delicious.” We conduct extensive experiments on four SemEval datasets. The results reveal the essential role of AAL in ACSA by achieving the state-of-the-art performance.
Peisong Zhu, Zhuang Chen 0002, Haojie Zheng, Tieyun Qian
ACM Trans. Knowl. Discov. Data2