EDBT 2026 Demo / reviewers in the wild / expert
Zhuang Chen 0002
dblp:54/8492-2
· DBLP profile ↗
23ranked-venue papers
10as first author
18since 2021 · last 2026
0000-0002-7048-7833ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 9 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unveiling the Landscape of Clinical Depression Assessment: From Behavioral Signatures to Psychiatric ReasoningabstractDepression is a widespread mental disorder that affects millions worldwide. While automated depression assessment shows promise, most studies rely on limited or non-clinically validated data, and often prioritize complex model design over real-world effectiveness. In this paper, we aim to unveil the landscape of clinical depression assessment. We introduce C-MIND, a clinical multimodal neuropsychiatric diagnosis dataset collected over two years from real hospital visits. Each participant completes three structured psychiatric tasks and receives a final diagnosis from expert clinicians, with informative audio, video, transcript, and functional near-infrared spectroscopy (fNIRS) signals recorded. Using C-MIND, we first analyze behavioral signatures relevant to diagnosis. We train a range of classical models to quantify how different tasks and modalities contribute to diagnostic performance, and dissect the effectiveness of their combinations. We then explore whether LLMs can perform psychiatric reasoning like clinicians and identify their clear limitations in realistic clinical settings. In response, we propose to guide the reasoning process with clinical expertise and consistently improve LLM diagnostic performance by up to 10% in Macro-F1 score. We aim to build an infrastructure for clinical depression assessment from both data and algorithmic perspectives, enabling C-MIND to facilitate grounded and reliable research for mental healthcare. Zhuang Chen 0002, Guanqun Bi, Aoyun Wang, Xiyao Xiao, Minlie Huang |
AAAI | 1 |
| 2026 | Ψ-Arena: Interactive Assessment and Optimization of LLM-based Psychological Counselors with Tripartite FeedbackabstractLarge language models (LLMs) have shown promise in providing scalable mental health support, while evaluating their counseling capability remains crucial to ensure both efficacy and safety. Existing evaluations are limited by the static assessment that focuses on knowledge tests, the single perspective that centers on user experience, and the open-loop framework that lacks actionable feedback. To address these issues, we propose Ψ-Arena, an interactive framework for comprehensive assessment and optimization of LLM-based counselors, featuring three key characteristics: (1) Realistic arena interactions that simulate real-world counseling through multi-stage dialogues with psychologically profiled NPC clients; (2) Tripartite evaluation that integrates assessments from the client, supervisor, and counselor perspectives; (3) Closed-loop optimization that iteratively improves LLM counselors using diagnostic feedback. Experiments across eight state-of-the-art LLMs show significant performance variations in different real-world scenarios and evaluation perspectives. Moreover, reflection-based optimization results in up to a 141% improvement in counseling performance. We hope Ψ-Arena provides a foundational resource for advancing reliable and human-aligned LLM applications in mental healthcare. Shijing Zhu, Zhuang Chen 0002, Guanqun Bi, Binghang Li, Yaxi Deng, Dazhen Wan, Libiao Peng, Xiyao Xiao, Tangjie Lv, Zhipeng Hu, Minlie Huang |
AAAI | 2 |
| 2026 | S⌃4: Operationalizing Speech Act Theory for Strategic Semi-Structured Psychiatric InterviewabstractPsychiatric interviewing is a strategic, goaloriented interaction that requires proactively steering the conversation to elicit latent information.However, existing methods often degenerate into rigid interrogation or aimless chitchat due to a lack of strategic planning.In this work, we introduce S 4 , a comprehensive framework grounded in Speech Act Theory, modeling the interview as a unified process of internal strategy (Illocution and Perlocution) and external realization (Locution).We synthesize a large-scale dataset with fine-grained psychiatric speech act annotations.Trained on this data, S 4 employs reinforcement learning driven by long-term therapeutic effects to optimize the strategic chaining of atomic acts, aiming to maximally elicit information and maintain patient engagement.Experiments demonstrate that S 4 significantly outperforms baselines, validating the effectiveness of our effectdriven strategic modeling.Action (A) Sample Locution (L) Definition & Intended Perlocution (P) I. Information Seeking (Directives: Eliciting Disclosure) Explore "How have you been sleeping?"Solicit Narrative: Ask open-ended questions to elicit detailed disclosure and expand symptom scope.Probe "Could you tell me more about that?" Deepen Inquiry: Follow up on ambiguity to clarify details and deepen focus.Confirm "Do you feel this way every day?" Pinpoint Fact: Ask closed-ended questions to verify diagnostic criteria.Clarify "By 'fatigue', I mean tiredness."Resolve Confusion: Provide explanations to align cognition and correct misunderstandings.II.Affective Regulation (Expressives: Modifying State) Validate "That sounds incredibly hard."Affirm Emotion: Acknowledge patient distress to lower defensiveness and build trust.Support "I understand.Please go on."Maintain Flow: Use back-channeling to demonstrate active listening and boost efficacy.Ease "Do you have any hobbies?"Reduce Tension: Engage in non-clinical conversation to de-escalate anxiety and humanize the agent. III. Interview Management (Representatives: Setting Frame)Initiate "Hi, I'm your AI counselor."Set Frame: Establish professional boundaries and the purpose of the session.Conclude "Thanks for sharing.Take care."Ensure Closure: Formally end the session to provide a safe exit and consolidation. Guanqun Bi, Zhoufu Liu, Zhuang Chen 0002, Dazhen Wan, Xiyao Xiao, Minlie Huang |
ACL (1) | 3 |
| 2025 | SocialSim: Towards Socialized Simulation of Emotional Support ConversationabstractEmotional support conversation (ESC) helps reduce people's psychological stress and provide emotional value through interactive dialogues. Due to the high cost of crowdsourcing a large ESC corpus, recent attempts use large language models for dialogue augmentation. However, existing approaches largely overlook the social dynamics inherent in ESC, leading to less effective simulations. In this paper, we introduce SocialSim, a novel framework that simulates ESC by integrating key aspects of social interactions: social disclosure and social awareness. On the seeker side, we facilitate social disclosure by constructing a comprehensive persona bank that captures diverse and authentic help-seeking scenarios. On the supporter side, we enhance social awareness by eliciting cognitive reasoning to generate logical and supportive responses. Building upon SocialSim, we construct SSConv, a large-scale synthetic ESC corpus of which quality can even surpass crowdsourced ESC data. We further train a chatbot on SSConv and demonstrate its state-of-the-art performance in both automatic and human evaluations. We believe SocialSim offers a scalable way to synthesize ESC, making emotional care more accessible and practical. Zhuang Chen 0002, Yaru Cao, Guanqun Bi, Jincenzi Wu, Jinfeng Zhou, Xiyao Xiao, Hongning Wang, Minlie Huang |
AAAI | 1 |
| 2025 | SS-GEN: A Social Story Generation Framework with Large Language ModelsabstractChildren with Autism Spectrum Disorder (ASD) often misunderstand social situations and struggle to participate in daily routines. Social Stories™ are traditionally crafted by psychology experts under strict constraints to address these challenges but are costly and limited in diversity. As Large Language Models (LLMs) advance, there's an opportunity to develop more automated, affordable, and accessible methods to generate Social Stories in real-time with broad coverage. However, adapting LLMs to meet the unique and strict constraints of Social Stories is a challenging issue. To this end, we propose SS-GEN, a Social Story GENeration framework with LLMs. Firstly, we develop a constraint-driven sophisticated strategy named StarSow to hierarchically prompt LLMs to generate Social Stories at scale, followed by rigorous human filtering to build a high-quality dataset. Additionally, we introduce quality assessment criteria to evaluate the effectiveness of these generated stories. Considering that powerful closed-source large models require very complex instructions and expensive API fees, we finally fine-tune smaller language models with our curated high-quality dataset, achieving comparable results at lower costs and with simpler instruction and deployment. This work marks a significant step in leveraging AI to personalize Social Stories cost-effectively for autistic children at scale, which we hope can encourage future research on special groups. Jiaqi Wang 0006, Zhuang Chen 0002, Guanqun Bi, Minlie Huang, Liping Jing, Jian Yu 0001 |
AAAI | 4 |
| 2025 | CharacterBench: Benchmarking Character Customization of Large Language ModelsabstractCharacter-based dialogue (aka role-playing) enables users to freely customize characters for interaction, which often relies on LLMs, raising the need to evaluate LLMs’ character customization capability. However, existing benchmarks fail to ensure a robust evaluation as they often only involve a single character category or evaluate limited dimensions. Moreover, the sparsity of character features in responses makes feature-focused generative evaluation both ineffective and inefficient. To address these issues, we propose CharacterBench, the largest bilingual generative benchmark, with 22,859 human-annotated samples covering 3,956 characters from 25 detailed character categories. We define 11 dimensions of 6 aspects, classified as sparse and dense dimensions based on whether character features evaluated by specific dimensions manifest in each response. We enable effective and efficient evaluation by crafting tailored queries for each dimension to induce characters’ responses related to specific dimensions. Further, we develop CharacterJudge model for cost-effective and stable evaluations. Experiments show its superiority over SOTA automatic judges (e.g., GPT-4) and our benchmark’s potential to optimize LLMs’ character customization. Jinfeng Zhou, Yongkang Huang, Bosi Wen, Guanqun Bi, Pei Ke, Zhuang Chen 0002, Xiyao Xiao, Libiao Peng, Kuntian Tang, Tangjie Lv, Zhipeng Hu, Hongning Wang, Minlie Huang |
AAAI | 7 |
| 2025 | Reframe Your Life Story: Interactive Narrative Therapist and Innovative Moment Assessment with Large Language ModelsabstractYi Feng, Jiaqi Wang, Wenxuan Zhang, Zhuang Chen, Shen Yutong, Xiyao Xiao, Minlie Huang, Liping Jing, Jian Yu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jiaqi Wang 0006, Zhuang Chen 0002, Yutong Shen, Xiyao Xiao, Minlie Huang, Liping Jing, Jian Yu 0001 |
EMNLP | 4 |
| 2025 | Generative Meta-Learning for Zero-Shot Relation Triplet ExtractionabstractZero-shot Relation Triplet Extraction (ZeroRTE) aims to extract relation triplets from texts containing unseen relation types. This capability benefits various downstream information retrieval (IR) tasks. The primary challenge lies in enabling models to generalize effectively to unseen relation categories. Existing approaches typically leverage the knowledge embedded in pre-trained language models to accomplish the generalization process. However, these methods focus solely on fitting the training data during training, without specifically improving the model's generalization performance, resulting in limited generalization capability. For this reason, we explore the integration of bi-level optimization (BLO) with pre-trained language models for learning generalized knowledge directly from the training data, and propose a generative meta-learning framework which exploits the 'learning-to-learn' ability of meta-learning to boost the generalization capability of generative models. Wanli Li 0002, Tieyun Qian, Zeyu Zhang 0004, Jiawei Li 0008, Zhuang Chen 0002, Lixin Zou |
SIGIR | 6 |
| 2024 | ToMBench: Benchmarking Theory of Mind in Large Language ModelsabstractZhuang Chen, Jincenzi Wu, Jinfeng Zhou, Bosi Wen, Guanqun Bi, Gongyao Jiang, Yaru Cao, Mengting Hu, Yunghwei Lai, Zexuan Xiong, Minlie Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhuang Chen 0002, Jincenzi Wu, Jinfeng Zhou, Bosi Wen, Guanqun Bi, Gongyao Jiang, Yaru Cao, Mengting Hu 0002, Yunghwei Lai, Zexuan Xiong, Minlie Huang |
ACL (1) | 1 |
| 2024 | COKE: A Cognitive Knowledge Graph for Machine Theory of MindabstractTheory of mind (ToM) refers to humans' ability to understand and infer the desires, beliefs, and intentions of others.The acquisition of ToM plays a key role in humans' social cognition and interpersonal relations.Though indispensable for social intelligence, ToM is still lacking for modern AI and NLP systems since they cannot access the human mental state and cognitive process beneath the training corpus.To empower AI systems with the ToM ability and narrow the gap between them and humans, in this paper, we propose COKE: the first cognitive knowledge graph for machine theory of mind, formalizing cognitive processes as a chained structure.Specifically, COKE formalizes ToM as a collection of 45k+ manually verified cognitive chains that characterize human mental activities and subsequent behavioral/affective responses when facing specific social circumstances.In addition, we further generalize COKE using LLMs and build a powerful generation model COLM tailored for cognitive reasoning.Experimental results in both automatic and human evaluation demonstrate the high quality of COKE, the superior ToM ability of COLM, and its potential to significantly enhance social applications.We release our code and data at https://github.com/jincenziwu/COKE. Jincenzi Wu, Zhuang Chen 0002, Jiawen Deng 0006, Sahand Sabour, Helen M. Meng, Minlie Huang |
ACL (1) | 2 |
| 2024 | Depression Detection in Clinical Interviews with LLM-Empowered Structural Element GraphabstractZhuang Chen, Jiawen Deng, Jinfeng Zhou, Jincenzi Wu, Tieyun Qian, Minlie Huang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zhuang Chen 0002, Jiawen Deng 0006, Jinfeng Zhou, Jincenzi Wu, Tieyun Qian, Minlie Huang |
NAACL-HLT | 1 |
| 2024 | Enhancing Emotional Support Conversation with Cognitive Chain-of-Thought Reasoning
Yaru Cao, Zhuang Chen 0002, Guanqun Bi, Yulin Feng, Fucheng Wan, Minlie Huang, Hongzhi Yu |
NLPCC (1) | 2 |
| 2023 | Facilitating Multi-turn Emotional Support Conversation with Positive Emotion Elicitation: A Reinforcement Learning ApproachabstractEmotional support conversation (ESC) aims to provide emotional support (ES) to improve one's mental state.Existing works stay at fitting grounded responses and responding strategies (e.g., question), which ignore the effect on ES and lack explicit goals to guide emotional positive transition.To this end, we introduce a new paradigm to formalize multi-turn ESC as a process of positive emotion elicitation.Addressing this task requires finely adjusting the elicitation intensity in ES as the conversation progresses while maintaining conversational goals like coherence.In this paper, we propose SUPPORTER, a mixture-of-expert-based reinforcement learning model, and well design ES and dialogue coherence rewards to guide policy's learning for responding.Experiments verify the superiority of SUPPORTER in achieving positive emotion elicitation during responding while maintaining conversational goals including coherence. Jinfeng Zhou, Zhuang Chen 0002, Minlie Huang |
ACL (1) | 2 |
| 2023 | Mimicking the Thinking Process for Emotion Recognition in Conversation with Prompts and ParaphrasingabstractEmotion recognition in conversation, which aims to predict the emotion for all utterances, has attracted considerable research attention in recent years. It is a challenging task since the recognition of the emotion in one utterance involves many complex factors, such as the conversational context, the speaker's background, and the subtle difference between emotion labels. In this paper, we propose a novel framework which mimics the thinking process when modeling these factors. Specifically, we first comprehend the conversational context with a history-oriented prompt to selectively gather information from predecessors of the target utterance. We then model the speaker's background with an experience-oriented prompt to retrieve the similar utterances from all conversations. We finally differentiate the subtle label semantics with a paraphrasing mechanism to elicit the intrinsic label related knowledge. We conducted extensive experiments on three benchmarks. The empirical results demonstrate the superiority of our proposed framework over the state-of-the-art baselines. Zhuang Chen 0002, Ming Zhong 0002, Tieyun Qian |
IJCAI | 2 |
| 2022 | Retrieve-and-Edit Domain Adaptation for End2End Aspect Based Sentiment AnalysisabstractEnd-to-end aspect based sentiment analysis (E2E-ABSA) aims to jointly extract aspect terms and predict aspect-level sentiment for opinion reviews. Though supervised methods show effectiveness for E2E-ABSA tasks, the annotation cost is extremely high due to the necessity of fine-grained labels. Recent attempts alleviate this problem using the domain adaptation technique to transfer the word-level common knowledge across domains. However, the biggest issue in domain adaptation, i.e., how to transfer the domain-specific words likepizzaanddeliciousin the source “Restaurant” to the target “Laptop” domain, has not been resolved. In this paper, we propose a novel domain adaptation method to address this issue by enhancing the transferability of domain-specific source words in a retrieve-and-edit way. Specifically, for all source words, we first retrieve the transferable prototypes from unlabeled target data via their syntactic and semantic roles. We then edit the source words to enhance their transferability by absorbing the knowledge carried in prototypes. Finally, we design an end-to-end framework to jointly accomplish cross-domain aspect term extraction and aspect-level sentiment classification. We conduct extensive experiments on four real-world datasets. The results prove that, by introducing transferable prototypes, our method significantly outperforms the state-of-the-art methods, achieving an absolute 3.95% F1 increase over the best baseline. Zhuang Chen 0002, Tieyun Qian |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Description and demonstration guided data augmentation for sequence tagging
Zhuang Chen 0002, Tieyun Qian |
World Wide Web | 1 |
| 2021 | Bridge-Based Active Domain Adaptation for Aspect Term ExtractionabstractZhuang Chen, Tieyun Qian. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zhuang Chen 0002, Tieyun Qian |
ACL/IJCNLP (1) | 1 |
| 2021 | Generating Pseudo Connectives with MLMs for Implicit Discourse Relation Recognition
Congcong Jiang, Tieyun Qian, Zhuang Chen 0002, Kejian Tang, Shaohui Zhan |
PRICAI (2) | 3 |
| 2020 | Relation-Aware Collaborative Learning for Unified Aspect-Based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA) involves three subtasks, i.e., aspect term extraction, opinion term extraction, and aspect-level sentiment classification.Most existing studies focused on one of these subtasks only.Several recent researches made successful attempts to solve the complete ABSA problem with a unified framework.However, the interactive relations among three subtasks are still underexploited.We argue that such relations encode collaborative signals between different subtasks.For example, when the opinion term is "delicious", the aspect term must be "food" rather than "place".In order to fully exploit these relations, we propose a Relation-Aware Collaborative Learning (RACL) framework which allows the subtasks to work coordinately via the multi-task learning and relation propagation mechanisms in a stacked multi-layer network.Extensive experiments on three real-world datasets demonstrate that RACL significantly outperforms the state-ofthe-art methods for the complete ABSA task. Zhuang Chen 0002, Tieyun Qian |
ACL | 1 |
| 2020 | Enhancing Aspect Term Extraction with Soft PrototypesabstractAspect term extraction (ATE) aims to extract aspect terms from a review sentence that users have expressed opinions on.Existing studies mostly focus on designing neural sequence taggers to extract linguistic features from the token level.However, since the aspect terms and context words usually exhibit long-tail distributions, these taggers often converge to an inferior state without enough sample exposure.In this paper, we propose to tackle this problem by correlating words with each other through soft prototypes.These prototypes, generated by a soft retrieval process, can introduce global knowledge from internal or external data and serve as the supporting evidence for discovering the aspect terms.Our proposed model is a general framework and can be combined with almost all sequence taggers.Experiments on four SemEval datasets show that our model boosts the performance of three typical ATE methods by a large margin. Zhuang Chen 0002, Tieyun Qian |
EMNLP (1) | 1 |
| 2019 | Transfer Capsule Network for Aspect Level Sentiment ClassificationabstractAspect-level sentiment classification aims to determine the sentiment polarity of a sentence towards an aspect.Due to the high cost in annotation, the lack of aspect-level labeled data becomes a major obstacle in this area.On the other hand, document-level labeled data like reviews are easily accessible from online websites.These reviews encode sentiment knowledge in abundant contexts.In this paper, we propose a Transfer Capsule Network (Tran-sCap) model for transferring document-level knowledge to aspect-level sentiment classification.To this end, we first develop an aspect routing approach to encapsulate the sentence-level semantic representations into semantic capsules from both aspect-level and document-level data.We then extend the dynamic routing approach to adaptively couple the semantic capsules with the class capsules under the transfer learning framework.Experiments on SemEval datasets demonstrate the effectiveness of TransCap. Zhuang Chen 0002, Tieyun Qian |
ACL (1) | 1 |
| 2019 | Aspect-Level Sentiment Classification with Dependency Rules and Dual Attention
Yunkai Yang, Tieyun Qian, Zhuang Chen 0002 |
ICONIP (2) | 3 |
| 2019 | Aspect Aware Learning for Aspect Category Sentiment AnalysisabstractAspect category sentiment analysis (ACSA) is an underexploited subtask in aspect level sentiment analysis. It aims to identify the sentiment of predefined aspect categories. The main challenge in ACSA comes from the fact that the aspect category may not occur in the sentence in most of the cases. For example, the review “ they have delicious sandwiches ” positively talks about the aspect category “ food ” in an implicit manner. In this article, we propose a novel aspect aware learning (AAL) framework for ACSA tasks. Our key idea is to exploit the interaction between the aspect category and the contents under the guidance of both sentiment polarity and predefined categories. To this end, we design a two-way memory network for integrating AAL into the framework of sentiment classification. We further present two algorithms to incorporate the potential impacts of aspect categories. One is to capture the correlations between aspect terms and the aspect category like “sandwiches” and “food.” The other is to recognize the aspect category for sentiment representations like “food” for “delicious.” We conduct extensive experiments on four SemEval datasets. The results reveal the essential role of AAL in ACSA by achieving the state-of-the-art performance. Peisong Zhu, Zhuang Chen 0002, Haojie Zheng, Tieyun Qian |
ACM Trans. Knowl. Discov. Data | 2 |