Jinpeng Hu

dblp:238/1225 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AgentMental: An Interactive Multi-Agent Framework for Explainable and Adaptive Mental Health Assessment
abstract
Mental health assessment is crucial for early intervention and effective treatment, yet traditional clinician-based approaches are limited by the shortage of qualified professionals. Recent advances in artificial intelligence have sparked growing interest in automated psychological assessment, yet most existing approaches are constrained by their reliance on static text analysis, limiting their ability to capture deeper and more informative insights that emerge through dynamic interaction and iterative questioning. Therefore, in this paper, we propose a multi-agent framework for mental health evaluation that simulates clinical doctor-patient dialogues, with specialized agents assigned to questioning, adequacy evaluation, scoring, and updating. In detail, we introduce an adaptive questioning mechanism in which an evaluation agent assesses the adequacy of user responses to determine the necessity of generating targeted follow-up queries to address ambiguity and missing information. Additionally, we employ a tree-structured memory in which the root node encodes the user's basic information, while child nodes (e.g., topic and statement) organize key information according to distinct symptom categories and interaction turns. This memory is dynamically updated throughout the interaction to reduce redundant questioning and enhance the information extraction and contextual tracking capabilities. Experimental results on the DAIC-WOZ dataset illustrate the effectiveness of our proposed method, which achieves better performance than existing approaches. Our code is released at \url{https://github.com/MindIntLab-HFUT/AgentMental}.
Jinpeng Hu, Qianqian Xie, Hui Ma 0011, Dan Guo 0001
AAAI1
2026 Psyche-R1: Towards Reliable Psychological LLMs through Unified Empathy, Expertise, and Reasoning
abstract
Amidst a shortage of qualified mental health professionals, the integration of large language models (LLMs) into psychological applications offers a promising way to alleviate the growing burden of mental health disorders.Recent reasoning-augmented LLMs have achieved remarkable performance in mathematics and programming, while research in the psychological domain has predominantly emphasized emotional support and empathetic dialogue, with limited attention to reasoning mechanisms that are beneficial to generating accurate responses.Therefore, in this paper, we propose Psyche-R1, the first Chinese psychological LLM that jointly integrates empathy, psychological expertise, and reasoning, built upon a novel data curation pipeline.Specifically, we design a comprehensive data synthesis pipeline that produces over 75k high-quality psychological questions paired with detailed rationales, generated through an iterative prompt-rationale optimization procedure, along with 73k empathetic dialogues.Subsequently, we employ a hybrid training strategy wherein challenging samples are identified through a multi-LLM cross-selection strategy for group relative policy optimization (GRPO) to improve reasoning ability, while the remaining data are used for supervised fine-tuning (SFT) to enhance empathetic response generation and psychological domain knowledge.Extensive experiment results demonstrate the effectiveness of Psyche-R1 across several psychological benchmarks, where our 7B Psyche-R1 achieves comparable results to 671B DeepSeek-R1.
Chongyuan Dai, Jinpeng Hu, Hongchang Shi, Dan Guo 0001, Xun Yang 0001, Meng Wang 0001
ACL (1)2
2026 Tears or Cheers? Benchmarking LLMs via Culturally Elicited Distinct Affective Responses
abstract
Chongyuan Dai, Yaling Shen, Zihan Gao, Jia Li, Yishun Jiang, Yaxiong Wang, Liu Liu, Zongyuan Ge, Jinpeng Hu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Chongyuan Dai, Yaling Shen, Jia Li 0057, Yishun Jiang, Yaxiong Wang, ZongYuan Ge, Jinpeng Hu
ACL (1)9
2026 From Social Media to Psychological Scale: An Adaptive Framework with Two-Hop Retrieval for Depression Screening
abstract
Depressive disorders represent a major global public health challenge. As an increasing number of individuals share their emotional experiences and concerns on social media, researchers have shown growing interest in leveraging such data for early depression screening. However, most existing methods rely on a fixed model and a singular reasoning paradigm, which constrains their adaptability to depression detection. The limited availability of mental health-related data and variability in training data distributions across different LLMs hinder their consistent and comprehensive understanding of diverse psychological symptoms. In this paper, we propose AdaDepression, a framework that enables explainable depression screening through a two-hop retrieval algorithm to identify symptom-relevant posts and a two-stage adaptive routing mechanism for selecting appropriate reasoning strategies and LLMs. Specifically, we first collect representative posts from the training dataset to capture the real-world symptom expressions, and then utilize these posts to retrieve symptom-relevant posts from the user's posting history. Subsequently, we employ the Mixture of Routers (MoR), which integrates the Mixture of Experts (MoE) into the routing mechanism to select the optimal reasoning strategies and LLMs in a cascaded manner. Finally, we complete the standardized psychological questionnaire using the selected LLMs and reasoning strategies. Experimental results on the Reddit-based benchmarks demonstrate the effectiveness of the proposed method, outperforming existing studies on various metrics. Our code is released at https://github.com/MindIntLab-HFUT/AdaDepression.
Yangyang Xu 0002, Jinpeng Hu, Peipei Song, Zhangling Duan, Xun Yang 0001
WWW2
2026 CLAIP-Emo: Parameter-Efficient Adaptation of Language-Supervised Models for In-the-Wild Audiovisual Emotion Recognition
abstract
Audiovisual emotion recognition (AVER) in the wild is still hindered by pose variation, occlusion, and background noise. Prevailing methods primarily rely on large-scale domain-specific pre-training, which is costly and often mismatched to real-world affective data. To address this, we present CLAIP-Emo, a modular framework that reframes in-the-wild AVER as a parameter-efficient adaptation of language-supervised foundation models (CLIP/CLAP). Specifically, it (i) preserves language-supervised priors by freezing CLIP/CLAP backbones and performing emotion-oriented adaptation via LoRA (updating \ensuremath{\le}4.0\% of the total parameters), (ii) allocates temporal modeling asymmetrically, employing a lightweight Transformer for visual dynamics while applying mean pooling for audio prosody, and (iii) applies a simple fusion head for prediction. On DFEW and MAFW, CLAIP-Emo (ViT-L/14) achieves 80.14\% and 61.18\% weighted average recall with only 8M training parameters, setting a new state of the art. Our findings suggest that parameter-efficient adaptation of language-supervised foundation models provides a scalable alternative to domain-specific pre-training for real-world AVER. The code and models will be available at \href{https://github.com/MSA-LMC/CLAIP-Emo}{https://github.com/MSA-LMC/CLAIP-Emo}.
Jia Li 0013, Jinpeng Hu, Zhenzhen Hu 0004, Richang Hong
IEEE Signal Process. Lett.3
2026 Bridging Subjectivity in Affective Explanation Captioning via Consensus-Prompted Emotion Reasoning
abstract
Affective Explanation Captioning (AEC) aims to perform viewer-centered visual emotion analysis by not only identifying the emotions evoked by an image but also explaining their underlying causes. Prior efforts have achieved promising results by fine-tuning LLMs on affective data; however, two key challenges remain: 1) the inherent subjectivity of human emotion leads to diverse interpretations of the same image, making it difficult for models to catch dominant emotions; and 2) the affective gap between abstract emotions and concrete visual content hinders models from capturing both semantic and emotional aspects effectively. To tackle these challenges, we propose Consensus-Prompted Emotion Reasoning (CPER), a new framework that explicitly models emotional diversity and enforces emotional-semantic alignment. Inspired by psychological studies, we observe that common emotional patterns often emerge within certain groups, which we refer to as affective consensus. Capturing this consensus across varying levels is helpful for bridging the subjectivity in AEC. Specifically, we introduce a consensus-based bucket prompt, which depicts the consensus level of each emotional perspective, serving as a control signal to adjust the emotion reasoning. To reconcile abstract emotion understanding and concrete visual grounding, we design a dual-space representation, where a CLIP encoder extracts objective semantic evidence and an emotion encoder captures abstract affective cues for AEC. Furthermore, an emotion consistency learning strategy is devised, which explicitly aligns the generated explanation with the input image and the emotion label, ensuring both emotionally and semantically grounded explanations. Extensive experiments on three benchmark datasets, ranging from visual arts (ArtEmis v1.0 and ArtEmis v2.0) and real-world images (Affection), demonstrate the effectiveness of our CPER in terms of emotional diversity and semantic coherence compared to state-of-the-art methods. Our code is publicly available at https://github.com/songpipi/CPER.
Peipei Song, Zhiyan Zhang, Weidong Chen 0013, Jinpeng Hu, Xun Yang 0001, Xiaojun Chang
IEEE Trans. Image Process.4
2025 Self-Instructed Derived Prompt Generation Meets In-Context Learning: Unlocking New Potential of Black-Box LLMs
abstract
Improving prompt quality is crucial for enhancing the performance of large language models (LLMs), particularly for Black-Box models like GPT4.Existing prompt refinement methods, while effective, often suffer from semantic inconsistencies between refined and original prompts, and fail to maintain users' real intent.To address these challenges, we propose a selfinstructed in-context learning framework that generates reliable derived prompts, keeping semantic consistency with the original prompts.Specifically, our framework incorporates a reinforcement learning mechanism, enabling direct interaction with the response model during prompt generation to better align with human preferences.We then formulate the querying as an in-context learning task, combining responses from LLMs with derived prompts to create a contextual demonstration for the original prompt.This approach effectively enhances alignment, reduces semantic discrepancies, and activates the LLM's in-context learning ability for generating more beneficial responses.Extensive experiments demonstrate that the proposed method not only generates better derived prompts but also significantly enhances LLMs' ability to deliver more effective responses, particularly for Black-Box models like GPT4.
Jinpeng Hu, Anningzhe Gao
ACL (1)3
2025 Add-One-In: Incremental Sample Selection for Large Language Models via a Choice-Based Greedy Paradigm
abstract
Selecting high-quality and diverse training samples from extensive datasets plays a crucial role in reducing training overhead and enhancing the performance of Large Language Models (LLMs).However, existing studies fall short in assessing the overall value of selected data, focusing primarily on individual quality, and struggle to strike an effective balance between ensuring diversity and minimizing data point traversals.Therefore, this paper introduces a novel choice-based sample selection framework that shifts the focus from evaluating individual sample quality to comparing the contribution value of different samples when incorporated into the subset.Thanks to the advanced language understanding capabilities of LLMs, we utilize LLMs to evaluate the value of each option during the selection process.Furthermore, we design a greedy sampling process where samples are incrementally added to the subset, thereby improving efficiency by eliminating the need for exhaustive traversal of the entire dataset with the limited budget.Extensive experiments demonstrate that selected data from our method not only surpasses the performance of the full dataset but also achieves competitive results with recent powerful studies, while requiring fewer selections.Moreover, we validate our approach on a larger medical dataset, highlighting its practical applicability in real-world applications.Our code and data are available at https://github.com/BIRlz/ comperative_sample_selection.
Xiaoqi Jiao, Steven Y. Guo, Yuege Feng, Anningzhe Gao, Jinpeng Hu
EMNLP8
2025 APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport
abstract
The reward model (RM) plays a crucial role in aligning Large Language Models (LLMs) with human preferences through Reinforcement Learning, where the Bradley-Terry (BT) objective has been recognized as simple yet powerful, specifically for pairwise preference learning.However, BT-based RMs often struggle to effectively distinguish between similar preference responses, leading to insufficient separation between preferred and non-preferred outputs.Consequently, they may easily overfit easy samples and cannot generalize well to Out-Of-Distribution (OOD) samples, resulting in suboptimal performance.To address these challenges, this paper introduces an effective enhancement to BT-based RMs through an adaptive margin mechanism.Specifically, we design to dynamically adjust the RM focus on more challenging samples through margins, based on both semantic similarity and model-predicted reward differences, which is approached from a distributional perspective solvable with Optimal Transport (OT).By incorporating these factors into a principled OT cost matrix design, our adaptive margin enables the RM to better capture distributional differences between chosen and rejected responses, yielding significant improvements in performance, convergence speed, and generalization capabilities.Experimental results across multiple benchmarks demonstrate that our method outperforms several existing RM techniques, showcasing enhanced performance in both In-Distribution (ID) and OOD settings.Moreover, RLHF experiments support our practical effectiveness in better aligning LLMs with human preferences.
Yuege Feng, Dandan Guo, Jinpeng Hu, Anningzhe Gao
EMNLP4
2025 MultiAgentESC: A LLM-based Multi-Agent Collaboration Framework for Emotional Support Conversation
abstract
The development of Emotional Support Conversation (ESC) systems is critical for delivering mental health support tailored to the needs of help-seekers.Recent advances in large language models (LLMs) have contributed to progress in this domain, while most existing studies focus on generating responses directly and overlook the integration of domain-specific reasoning and expert interaction.Therefore, in this paper, we propose a training-free Multi-Agent collaboration framework for ESC (Mul-tiAgentESC).The framework is designed to emulate the human-like process of providing emotional support through three stages: dialogue analysis, strategy deliberation, and response generation.At each stage, a multi-agent system is employed to iteratively enhance information understanding and reasoning, simulating real-world decision-making processes by incorporating diverse interactions among these expert agents.Additionally, we introduce a novel response-centered approach to handle the one-to-many problem on strategy selection, where multiple valid strategies are initially employed to generate diverse responses, followed by the selection of the optimal response through multi-agent collaboration.Experiments on the ESConv dataset reveal that our proposed framework excels at providing emotional support as well as diversifying support strategy selection 1 .
Yangyang Xu 0002, Jinpeng Hu, Zhuoer Zhao, Zhangling Duan, Xiao Sun 0003, Xun Yang 0001
EMNLP2
2025 Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark
abstract
Multimodal large language models (MLLMs) have been widely applied across various fields due to their powerful perceptual and reasoning capabilities. In the realm of psychology, these models hold promise for a deeper understanding of human emotions and behaviors. However, recent research primarily focuses on enhancing their emotion recognition abilities, leaving the substantial potential in emotion reasoning, which is crucial for improving the naturalness and effectiveness of human-machine interactions. Therefore, in this paper, we introduce a multi-turn multimodal emotion understanding and reasoning (MTMEUR) benchmark, which encompasses 1,451 video data from real-life scenarios, along with 5,101 progressive questions. These questions cover various aspects, including emotion recognition, potential causes of emotions, future action prediction, etc. Besides, we propose a multi-agent framework, where each agent specializes in a specific aspect, such as background context, character dynamics, and event details, to improve the system's reasoning capabilities. Furthermore, we conduct experiments with existing MLLMs and our agent-based method on the proposed benchmark, revealing that most models face significant challenges with this task.
Jinpeng Hu, Hongchang Shi, Chongyuan Dai, Peipei Song, Meng Wang 0001
ACM Multimedia1
2025 SIEP-YOLO: Small Target Cluster Detection in Aerial Images
abstract
Accurate detection of small object clusters in modern security surveillance systems is critical for threat prevention and response, directly impacting monitoring efficiency and risk management efficacy. However, prevailing object detection algorithms struggle with high-density small target scenarios, suffering from slow inference speeds, insufficient precision, frequent false positives, and high miss rates. These limitations undermine real-time reliability and responsiveness to emergencies. To address these challenges, we propose SIEP-YOLO, an enhanced YOLOv11-based model integrating three novel components: the SDI-iAFF feature fusion module, EUCB upsampling module, and a specialized small object detection layer P2. This architecture boosts representational capacity and detection performance while simplifying computational complexity, achieving a balance between high accuracy and lightweight design. Experimental results demonstrate that SIEP-YOLO outperforms the original YOLOv11 by 3.2% and 2.0% in [email protected] and mAP@(0.5:0.95), respectively, on benchmark datasets. The proposed model thus emerges as a superior solution for small object cluster detection, enabling more efficient and reliable surveillance in complex environments.
Jinpeng Hu, Yuming Bo
SMC3
2025 PsycoLLM: Enhancing LLM for Psychological Understanding and Evaluation
abstract
Mental health has attracted substantial attention in recent years and large language model (LLM) can be an effective technology for alleviating this problem owing to its capability in text understanding and dialogue. However, existing research in this domain often suffers from limitations, such as training on datasets lacking crucial prior knowledge and evidence, and the absence of comprehensive evaluation methods. In this article, we propose a specialized psychological LLM, named PsycoLLM, trained on a proposed high-quality psychological dataset, including single-turn QA, multiturn dialogues, and knowledge-based QA. Specifically, we construct multi-turn dialogues through a three-step pipeline comprising multiturn QA generation, evidence judgment, and dialogue refinement. We augment this process with real-world psychological case backgrounds extracted from online platforms, enhancing the relevance and applicability of the generated data. Additionally, to compare the performance of PsycoLLM with other LLMs, we develop a comprehensive psychological benchmark based on authoritative psychological counseling examinations in China, which includes assessments of professional ethics, theoretical proficiency, and case analysis. The experimental results on the benchmark illustrate the effectiveness of PsycoLLM, which demonstrates superior performance compared with other LLMs.
Jinpeng Hu, Tengteng Dong, Hui Ma 0011, Xiao Sun 0003, Dan Guo 0001, Xun Yang 0001, Meng Wang 0001
IEEE Trans. Comput. Soc. Syst.1
2024 Mapping medical image-text to a joint space via masked modeling
Jinpeng Hu, Yang Liu 0258, Guanbin Li, Tsung-Hui Chang
Medical Image Anal.3
2023 A Simple Yet Effective Subsequence-Enhanced Approach for Cross-Domain NER
abstract
Cross-domain named entity recognition (NER), aiming to address the limitation of labeled resources in the target domain, is a challenging yet important task. Most existing studies alleviate the data discrepancy across different domains at the coarse level via combing NER with language modelings or introducing domain-adaptive pre-training (DAPT). Notably, source and target domains tend to share more fine-grained local information within denser subsequences than global information within the whole sequence, such that subsequence features are easier to transfer, which has not been explored well. Besides, compared to token-level representation, subsequence-level information can help the model distinguish different meanings of the same word in different domains. In this paper, we propose to incorporate subsequence-level features for promoting the cross-domain NER. In detail, we first utilize a pre-trained encoder to extract the global information. Then, we re-express each sentence as a group of subsequences and propose a novel bidirectional memory recurrent unit (BMRU) to capture features from the subsequences. Finally, an adaptive coupling unit (ACU) is proposed to combine global information and subsequence features for predicting entity labels. Experimental results on several benchmark datasets illustrate the effectiveness of our model, which achieves considerable improvements.
Jinpeng Hu, Dandan Guo, Yang Liu 0258, Tsung-Hui Chang
AAAI1
2023 EASAL: Entity-Aware Subsequence-Based Active Learning for Named Entity Recognition
abstract
Active learning is a critical technique for reducing labelling load by selecting the most informative data. Most previous works applied active learning on Named Entity Recognition (token-level task) similar to the text classification (sentence-level task). They failed to consider the heterogeneity of uncertainty within each sentence and required access to the entire sentence for the annotator when labelling. To overcome the mentioned limitations, in this paper, we allow the active learning algorithm to query subsequences within sentences and propose an Entity-Aware Subsequences-based Active Learning (EASAL) that utilizes an effective Head-Tail pointer to query one entity-aware subsequence for each sentence based on BERT. For other tokens outside this subsequence, we randomly select 30% of these tokens to be pseudo-labelled for training together where the model directly predicts their pseudo-labels. Experimental results on both news and biomedical datasets demonstrate the effectiveness of our proposed method. The code is released at https://github.com/lylylylylyly/EASAL.
Yang Liu 0258, Jinpeng Hu, Tsung-Hui Chang
AAAI2
2022 Graph Enhanced Contrastive Learning for Radiology Findings Summarization
abstract
The impression section of a radiology report summarizes the most prominent observation from the findings section and is the most important section for radiologists to communicate to physicians.Summarizing findings is timeconsuming and can be prone to error for inexperienced radiologists, and thus automatic impression generation has attracted substantial attention.With the encoder-decoder framework, most previous studies explore incorporating extra knowledge (e.g., static pre-defined clinical ontologies or extra background information).Yet, they encode such knowledge by a separate encoder to treat it as an extra input to their models, which is limited in leveraging their relations with the original findings.To address the limitation, we propose a unified framework for exploiting both extra knowledge and the original findings in an integrated way so that the critical information (i.e., key words and their relations) can be extracted in an appropriate way to facilitate impression generation.In detail, for each input findings, it is encoded by a text encoder, and a graph is constructed through its entities and dependency tree.Then, a graph encoder (e.g., graph neural networks (GNNs)) is adopted to model relation information in the constructed graph.Finally, to emphasize the key words in the findings, contrastive learning is introduced to map positive samples (constructed by masking non-key words) closer and push apart negative ones (constructed by masking key words).The experimental results on OpenI and MIMIC-CXR confirm the effectiveness of our proposed method. 1
Jinpeng Hu, Zhen Li 0026, Tsung-Hui Chang
ACL (1)1
2022 Multi-modal Masked Autoencoders for Medical Vision-and-Language Pre-training
Jinpeng Hu, Yang Liu 0258, Guanbin Li, Tsung-Hui Chang
MICCAI (5)3
2022 Hero-Gang Neural Model For Named Entity Recognition
abstract
Jinpeng Hu, Yaling Shen, Yang Liu, Xiang Wan, Tsung-Hui Chang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Jinpeng Hu, Yaling Shen, Yang Liu 0258, Tsung-Hui Chang
NAACL-HLT1