VLDB 2026 Research / reviewers in the wild / expert
Guanqun Bi
dblp:302/4016
· DBLP profile ↗
15ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0001-8829-9489ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unveiling the Landscape of Clinical Depression Assessment: From Behavioral Signatures to Psychiatric ReasoningabstractDepression is a widespread mental disorder that affects millions worldwide. While automated depression assessment shows promise, most studies rely on limited or non-clinically validated data, and often prioritize complex model design over real-world effectiveness. In this paper, we aim to unveil the landscape of clinical depression assessment. We introduce C-MIND, a clinical multimodal neuropsychiatric diagnosis dataset collected over two years from real hospital visits. Each participant completes three structured psychiatric tasks and receives a final diagnosis from expert clinicians, with informative audio, video, transcript, and functional near-infrared spectroscopy (fNIRS) signals recorded. Using C-MIND, we first analyze behavioral signatures relevant to diagnosis. We train a range of classical models to quantify how different tasks and modalities contribute to diagnostic performance, and dissect the effectiveness of their combinations. We then explore whether LLMs can perform psychiatric reasoning like clinicians and identify their clear limitations in realistic clinical settings. In response, we propose to guide the reasoning process with clinical expertise and consistently improve LLM diagnostic performance by up to 10% in Macro-F1 score. We aim to build an infrastructure for clinical depression assessment from both data and algorithmic perspectives, enabling C-MIND to facilitate grounded and reliable research for mental healthcare. Zhuang Chen 0002, Guanqun Bi, Aoyun Wang, Xiyao Xiao, Minlie Huang |
AAAI | 2 |
| 2026 | Ψ-Arena: Interactive Assessment and Optimization of LLM-based Psychological Counselors with Tripartite FeedbackabstractLarge language models (LLMs) have shown promise in providing scalable mental health support, while evaluating their counseling capability remains crucial to ensure both efficacy and safety. Existing evaluations are limited by the static assessment that focuses on knowledge tests, the single perspective that centers on user experience, and the open-loop framework that lacks actionable feedback. To address these issues, we propose Ψ-Arena, an interactive framework for comprehensive assessment and optimization of LLM-based counselors, featuring three key characteristics: (1) Realistic arena interactions that simulate real-world counseling through multi-stage dialogues with psychologically profiled NPC clients; (2) Tripartite evaluation that integrates assessments from the client, supervisor, and counselor perspectives; (3) Closed-loop optimization that iteratively improves LLM counselors using diagnostic feedback. Experiments across eight state-of-the-art LLMs show significant performance variations in different real-world scenarios and evaluation perspectives. Moreover, reflection-based optimization results in up to a 141% improvement in counseling performance. We hope Ψ-Arena provides a foundational resource for advancing reliable and human-aligned LLM applications in mental healthcare. Shijing Zhu, Zhuang Chen 0002, Guanqun Bi, Binghang Li, Yaxi Deng, Dazhen Wan, Libiao Peng, Xiyao Xiao, Tangjie Lv, Zhipeng Hu, Minlie Huang |
AAAI | 3 |
| 2026 | S⌃4: Operationalizing Speech Act Theory for Strategic Semi-Structured Psychiatric InterviewabstractPsychiatric interviewing is a strategic, goaloriented interaction that requires proactively steering the conversation to elicit latent information.However, existing methods often degenerate into rigid interrogation or aimless chitchat due to a lack of strategic planning.In this work, we introduce S 4 , a comprehensive framework grounded in Speech Act Theory, modeling the interview as a unified process of internal strategy (Illocution and Perlocution) and external realization (Locution).We synthesize a large-scale dataset with fine-grained psychiatric speech act annotations.Trained on this data, S 4 employs reinforcement learning driven by long-term therapeutic effects to optimize the strategic chaining of atomic acts, aiming to maximally elicit information and maintain patient engagement.Experiments demonstrate that S 4 significantly outperforms baselines, validating the effectiveness of our effectdriven strategic modeling.Action (A) Sample Locution (L) Definition & Intended Perlocution (P) I. Information Seeking (Directives: Eliciting Disclosure) Explore "How have you been sleeping?"Solicit Narrative: Ask open-ended questions to elicit detailed disclosure and expand symptom scope.Probe "Could you tell me more about that?" Deepen Inquiry: Follow up on ambiguity to clarify details and deepen focus.Confirm "Do you feel this way every day?" Pinpoint Fact: Ask closed-ended questions to verify diagnostic criteria.Clarify "By 'fatigue', I mean tiredness."Resolve Confusion: Provide explanations to align cognition and correct misunderstandings.II.Affective Regulation (Expressives: Modifying State) Validate "That sounds incredibly hard."Affirm Emotion: Acknowledge patient distress to lower defensiveness and build trust.Support "I understand.Please go on."Maintain Flow: Use back-channeling to demonstrate active listening and boost efficacy.Ease "Do you have any hobbies?"Reduce Tension: Engage in non-clinical conversation to de-escalate anxiety and humanize the agent. III. Interview Management (Representatives: Setting Frame)Initiate "Hi, I'm your AI counselor."Set Frame: Establish professional boundaries and the purpose of the session.Conclude "Thanks for sharing.Take care."Ensure Closure: Formally end the session to provide a safe exit and consolidation. Guanqun Bi, Zhoufu Liu, Zhuang Chen 0002, Dazhen Wan, Xiyao Xiao, Minlie Huang |
ACL (1) | 1 |
| 2025 | SocialSim: Towards Socialized Simulation of Emotional Support ConversationabstractEmotional support conversation (ESC) helps reduce people's psychological stress and provide emotional value through interactive dialogues. Due to the high cost of crowdsourcing a large ESC corpus, recent attempts use large language models for dialogue augmentation. However, existing approaches largely overlook the social dynamics inherent in ESC, leading to less effective simulations. In this paper, we introduce SocialSim, a novel framework that simulates ESC by integrating key aspects of social interactions: social disclosure and social awareness. On the seeker side, we facilitate social disclosure by constructing a comprehensive persona bank that captures diverse and authentic help-seeking scenarios. On the supporter side, we enhance social awareness by eliciting cognitive reasoning to generate logical and supportive responses. Building upon SocialSim, we construct SSConv, a large-scale synthetic ESC corpus of which quality can even surpass crowdsourced ESC data. We further train a chatbot on SSConv and demonstrate its state-of-the-art performance in both automatic and human evaluations. We believe SocialSim offers a scalable way to synthesize ESC, making emotional care more accessible and practical. Zhuang Chen 0002, Yaru Cao, Guanqun Bi, Jincenzi Wu, Jinfeng Zhou, Xiyao Xiao, Hongning Wang, Minlie Huang |
AAAI | 3 |
| 2025 | SS-GEN: A Social Story Generation Framework with Large Language ModelsabstractChildren with Autism Spectrum Disorder (ASD) often misunderstand social situations and struggle to participate in daily routines. Social Stories™ are traditionally crafted by psychology experts under strict constraints to address these challenges but are costly and limited in diversity. As Large Language Models (LLMs) advance, there's an opportunity to develop more automated, affordable, and accessible methods to generate Social Stories in real-time with broad coverage. However, adapting LLMs to meet the unique and strict constraints of Social Stories is a challenging issue. To this end, we propose SS-GEN, a Social Story GENeration framework with LLMs. Firstly, we develop a constraint-driven sophisticated strategy named StarSow to hierarchically prompt LLMs to generate Social Stories at scale, followed by rigorous human filtering to build a high-quality dataset. Additionally, we introduce quality assessment criteria to evaluate the effectiveness of these generated stories. Considering that powerful closed-source large models require very complex instructions and expensive API fees, we finally fine-tune smaller language models with our curated high-quality dataset, achieving comparable results at lower costs and with simpler instruction and deployment. This work marks a significant step in leveraging AI to personalize Social Stories cost-effectively for autistic children at scale, which we hope can encourage future research on special groups. Jiaqi Wang 0006, Zhuang Chen 0002, Guanqun Bi, Minlie Huang, Liping Jing, Jian Yu 0001 |
AAAI | 5 |
| 2025 | CharacterBench: Benchmarking Character Customization of Large Language ModelsabstractCharacter-based dialogue (aka role-playing) enables users to freely customize characters for interaction, which often relies on LLMs, raising the need to evaluate LLMs’ character customization capability. However, existing benchmarks fail to ensure a robust evaluation as they often only involve a single character category or evaluate limited dimensions. Moreover, the sparsity of character features in responses makes feature-focused generative evaluation both ineffective and inefficient. To address these issues, we propose CharacterBench, the largest bilingual generative benchmark, with 22,859 human-annotated samples covering 3,956 characters from 25 detailed character categories. We define 11 dimensions of 6 aspects, classified as sparse and dense dimensions based on whether character features evaluated by specific dimensions manifest in each response. We enable effective and efficient evaluation by crafting tailored queries for each dimension to induce characters’ responses related to specific dimensions. Further, we develop CharacterJudge model for cost-effective and stable evaluations. Experiments show its superiority over SOTA automatic judges (e.g., GPT-4) and our benchmark’s potential to optimize LLMs’ character customization. Jinfeng Zhou, Yongkang Huang, Bosi Wen, Guanqun Bi, Pei Ke, Zhuang Chen 0002, Xiyao Xiao, Libiao Peng, Kuntian Tang, Tangjie Lv, Zhipeng Hu, Hongning Wang, Minlie Huang |
AAAI | 4 |
| 2024 | ToMBench: Benchmarking Theory of Mind in Large Language ModelsabstractZhuang Chen, Jincenzi Wu, Jinfeng Zhou, Bosi Wen, Guanqun Bi, Gongyao Jiang, Yaru Cao, Mengting Hu, Yunghwei Lai, Zexuan Xiong, Minlie Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhuang Chen 0002, Jincenzi Wu, Jinfeng Zhou, Bosi Wen, Guanqun Bi, Gongyao Jiang, Yaru Cao, Mengting Hu 0002, Yunghwei Lai, Zexuan Xiong, Minlie Huang |
ACL (1) | 5 |
| 2024 | Enhancing Emotional Support Conversation with Cognitive Chain-of-Thought Reasoning
Yaru Cao, Zhuang Chen 0002, Guanqun Bi, Yulin Feng, Fucheng Wan, Minlie Huang, Hongzhi Yu |
NLPCC (1) | 3 |
| 2024 | Generative Models for Complex Logical Reasoning over Knowledge GraphsabstractAnswering complex logical queries over knowledge graphs (KGs) is a fundamental yet challenging task. Recently, query representation has been a mainstream approach to complex logical reasoning, making the target answer and query closer in the embedding space. However, there are still two limitations. First, prior methods model the query as a fixed vector, but ignore the uncertainty of relations on KGs. In fact, different relations may contain different semantic distributions. Second, traditional representation frameworks fail to capture the joint distribution of queries and answers, which can be learned by generative models that have the potential to produce more coherent answers. To alleviate these limitations, we propose a novel generative model, named DiffCLR, which exploits the diffusion model for complex logical reasoning to approximate query distributions. Specifically, we first devise a query transformation to convert logical queries into input sequences by dynamically constructing contextual subgraphs. Then, we integrate them into the diffusion model to execute a multi-step generative process, and a structure-enhanced self-attention is further designed for incorporating the structural features embodied in KGs. Experimental results on two benchmark datasets show our model effectively outperforms state-of-the-art methods, particularly in multi-hop chain queries with significant improvement. Yu Liu 0118, Yanan Cao 0001, Shi Wang 0002, Qingyue Wang, Guanqun Bi |
WSDM | 5 |
| 2023 | DiffusEmp: A Diffusion Model-Based Framework with Multi-Grained Control for Empathetic Response GenerationabstractEmpathy is a crucial factor in open-domain conversations, which naturally shows one's caring and understanding to others.Though several methods have been proposed to generate empathetic responses, existing works often lead to monotonous empathy that refers to generic and safe expressions.In this paper, we propose to use explicit control to guide the empathy expression and design a framework DIFFUSEMP based on conditional diffusion language model to unify the utilization of dialogue context and attribute-oriented control signals.Specifically, communication mechanism, intent, and semantic frame are imported as multi-grained signals that control the empathy realization from coarse to fine levels.We then design a specific masking strategy to reflect the relationship between multi-grained signals and response tokens, and integrate it into the diffusion model to influence the generative process.Experimental results on a benchmark dataset EMPA-THETICDIALOGUE show that our framework outperforms competitive baselines in terms of controllability, informativeness, and diversity without the loss of context-relatedness. Guanqun Bi, Lei Shen 0001, Yanan Cao 0001, Meng Chen 0006, Yuqiang Xie, Zheng Lin 0001, Xiaodong He 0001 |
ACL (1) | 1 |
| 2023 | Seri: Sketching-Reasoning-Integrating Progressive Workflow for Empathetic Response GenerationabstractEmpathy is a key ability for a human-like dialogue system. Inspired by social psychology, empathy includes both affective and cognitive aspects. Previous works on this topic have merely focused on recognizing emotions or modeling cognition with commonsense knowledge. Nevertheless, the generated results of these works still have a big gap with human-like empathetic responses. In this paper, we propose Seri, a SkEtching-Reasoning-Integrating framework for empathetic response generation. In particular, we define an empathy planner to capture and reason about multi-source information that considers cognition and affection. Further, we introduce a dynamic integrator module that allows the model dynamically select the appropriate information to generate empathetic responses. Experimental results on EmpatheticDialogue show that our method outperforms competitive baselines and generates responses with higher diversity and cognitive empathy levels. Guanqun Bi, Yanan Cao 0001, Piji Li, Yuqiang Xie, Fang Fang 0009, Zheng Lin 0001 |
ICASSP | 1 |
| 2022 | How Does Knowledge Graph Embedding Extrapolate to Unseen Data: A Semantic Evidence ViewabstractKnowledge Graph Embedding (KGE) aims to learn representations for entities and relations. Most KGE models have gained great success, especially on extrapolation scenarios. Specifically, given an unseen triple (h, r, t), a trained model can still correctly predict t from (h, r, ?), or h from (?, r, t), such extrapolation ability is impressive. However, most existing KGE works focus on the design of delicate triple modeling function, which mainly tells us how to measure the plausibility of observed triples, but offers limited explanation of why the methods can extrapolate to unseen data, and what are the important factors to help KGE extrapolate. Therefore in this work, we attempt to study the KGE extrapolation of two problems: 1. How does KGE extrapolate to unseen data? 2. How to design the KGE model with better extrapolation ability? For the problem 1, we first discuss the impact factors for extrapolation and from relation, entity and triple level respectively, propose three Semantic Evidences (SEs), which can be observed from train set and provide important semantic information for extrapolation. Then we verify the effectiveness of SEs through extensive experiments on several typical KGE methods. For the problem 2, to make better use of the three levels of SE, we propose a novel GNN-based KGE model, called Semantic Evidence aware Graph Neural Network (SE-GNN). In SE-GNN, each level of SE is modeled explicitly by the corresponding neighbor pattern, and merged sufficiently by the multi-layer aggregation, which contributes to obtaining more extrapolative knowledge representation. Finally, through extensive experiments on FB15k-237 and WN18RR datasets, we show that SE-GNN achieves state-of-the-art performance on Knowledge Graph Completion task and performs a better extrapolation ability. Our code is available at https://github.com/renli1024/SE-GNN. Yanan Cao 0001, Qiannan Zhu, Guanqun Bi, Fang Fang 0009, Yi Liu 0067, Qian Li 0003 |
AAAI | 4 |
| 2022 | COMMA: Modeling Relationship among Motivations, Emotions and Actions in Language-based Human ActivitiesabstractMotivations, emotions, and actions are inter-related essential factors in human activities. While motivations and emotions have long been considered at the core of exploring how people take actions in human activities, there has been relatively little research supporting analyzing the relationship between human mental states and actions. We present the first study that investigates the viability of modeling motivations, emotions, and actions in language-based human activities, named COMMA (Cognitive Framework of Human Activities). Guided by COMMA, we define three natural language processing tasks (emotion understanding, motivation understanding and conditioned action generation), and build a challenging dataset Hail through automatically extracting samples from Story Commonsense. Experimental results on NLP applications prove the effectiveness of modeling the relationship. Furthermore, our models inspired by COMMA can better reveal the essential relationship among motivations, emotions and actions than existing methods. Yuqiang Xie, Yue Hu 0002, Wei Peng 0008, Guanqun Bi, Luxi Xing |
COLING | 4 |
| 2022 | Psychology-guided Controllable Story GenerationabstractControllable story generation is a challenging task in the field of NLP, which has attracted increasing research interest in recent years. However, most existing works generate a whole story conditioned on the appointed keywords or emotions, ignoring the psychological changes of the protagonist. Inspired by psychology theories, we introduce global psychological state chains, which include the needs and emotions of the protagonists, to help a story generation system create more controllable and well-planned stories. In this paper, we propose a Psychology-guided Controllable Story Generation System (PICS) to generate stories that adhere to the given leading context and desired psychological state chains for the protagonist. Specifically, psychological state trackers are employed to memorize the protagonist’s local psychological states to capture their inner temporal relationships. In addition, psychological state planners are adopted to gain the protagonist’s global psychological states for story planning. Eventually, a psychology controller is designed to integrate the local and global psychological states into the story context representation for composing psychology-guided stories. Automatic and manual evaluations demonstrate that PICS outperforms baselines, and each part of PICS shows effectiveness for writing stories with more consistent psychological changes. Yuqiang Xie, Yue Hu 0002, Yunpeng Li 0006, Guanqun Bi, Luxi Xing, Wei Peng 0008 |
COLING | 4 |
| 2021 | No news is an island: Joint heterogeneous graph network for news classificationabstractNews classification is important for people to organize web information. There are many connections between daily news, for example, they may be different reports about the same event or the same person. However, previous works classify news mainly based on single news, they ignore relationships between multiple news. This paper innovatively utilizes relationships between multiple news, such as their relevance in time, place and people, to classify news. To take full advantage of these relationships and integrate various information of multiple news, we propose News Classification Graph (NCG), a heterogeneous graph with different types of nodes and edges. Furthermore, we propose Joint Heterogeneous graph Network (JHN) to properly embed the NCG. It fully utilizes the information of heterogeneous nodes and heterogeneous edges in NCG. Extensive experiments carried on four news datasets demonstrate the effectiveness of our work in solving the news classification problem. Zhezhou Kang, Yi Liu 0067, Guanqun Bi, Fang Fang 0009, Pengfei Yin |
ISCC | 3 |