VLDB 2026 Research / reviewers in the wild / expert
Hui Ma 0011
dblp:46/5778-11
· DBLP profile ↗
14ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0002-7146-8036ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LoRA in LoRA: Towards Parameter-Efficient Architecture Expansion for Continual Visual Instruction TuningabstractContinual Visual Instruction Tuning (CVIT) enables Multimodal Large Language Models (MLLMs) to incrementally learn new tasks over time. However, this process is challenged by catastrophic forgetting, where performance on previously learned tasks deteriorates as the model adapts to new ones. A common approach to mitigate forgetting is architecture expansion, which introduces task-specific modules to prevent interference. Yet, existing methods often expand entire layers for each task, leading to significant parameter overhead and poor scalability. To overcome these issues, we introduce LoRA in LoRA (LiLoRA), a highly efficient architecture expansion method tailored for CVIT in MLLMs. LiLoRA shares the LoRA matrix A across tasks to reduce redundancy, applies an additional low-rank decomposition to matrix B to minimize task-specific parameters, and incorporates a cosine-regularized stability loss to preserve consistency in shared representations over time. Extensive experiments on a diverse CVIT benchmark show that LiLoRA consistently achieves superior performance in sequential task learning while significantly improving parameter efficiency compared to existing approaches. Chang Che, Pengwan Yang, Cheems Wang, Hui Ma 0011, Zenglin Shi |
AAAI | 5 |
| 2026 | AgentMental: An Interactive Multi-Agent Framework for Explainable and Adaptive Mental Health AssessmentabstractMental health assessment is crucial for early intervention and effective treatment, yet traditional clinician-based approaches are limited by the shortage of qualified professionals. Recent advances in artificial intelligence have sparked growing interest in automated psychological assessment, yet most existing approaches are constrained by their reliance on static text analysis, limiting their ability to capture deeper and more informative insights that emerge through dynamic interaction and iterative questioning. Therefore, in this paper, we propose a multi-agent framework for mental health evaluation that simulates clinical doctor-patient dialogues, with specialized agents assigned to questioning, adequacy evaluation, scoring, and updating. In detail, we introduce an adaptive questioning mechanism in which an evaluation agent assesses the adequacy of user responses to determine the necessity of generating targeted follow-up queries to address ambiguity and missing information. Additionally, we employ a tree-structured memory in which the root node encodes the user's basic information, while child nodes (e.g., topic and statement) organize key information according to distinct symptom categories and interaction turns. This memory is dynamically updated throughout the interaction to reduce redundant questioning and enhance the information extraction and contextual tracking capabilities. Experimental results on the DAIC-WOZ dataset illustrate the effectiveness of our proposed method, which achieves better performance than existing approaches. Our code is released at \url{https://github.com/MindIntLab-HFUT/AgentMental}. Jinpeng Hu, Qianqian Xie, Hui Ma 0011, Dan Guo 0001 |
AAAI | 5 |
| 2025 | Efficient Tuning of Large Language Models for Knowledge-Grounded Dialogue GenerationabstractAbstract Large language models (LLMs) demonstrate remarkable text comprehension and generation capabilities but often lack the ability to utilize up-to-date or domain-specific knowledge not included in their training data. To address this gap, we introduce KEDiT, an efficient method for fine-tuning LLMs for knowledge-grounded dialogue generation. KEDiT operates in two main phases. First, it employs an information bottleneck to compress retrieved knowledge into learnable parameters, retaining essential information while minimizing computational overhead. Second, a lightweight knowledge-aware adapter integrates these compressed knowledge vectors into the LLM during fine-tuning, updating less than 2% of the model parameters. The experimental results on the Wizard of Wikipedia and a newly constructed PubMed-Dialog dataset demonstrate that KEDiT excels in generating contextually relevant and informative responses, outperforming competitive baselines in automatic, LLM-based, and human evaluations. This approach effectively combines the strengths of pretrained LLMs with the adaptability needed for incorporating dynamic knowledge, presenting a scalable solution for fields such as medicine.1 Bo Zhang 0121, Hui Ma 0011, Dailin Li, Jian Wang 0021, Bo Xu 0009, Hongfei Lin |
Trans. Assoc. Comput. Linguistics | 2 |
| 2025 | Empathy Level Alignment via Reinforcement Learning for Empathetic Response GenerationabstractEmpathetic response generation, aiming to understand the user’s situation and feelings and respond empathically, is crucial in building human-like dialogue systems. Traditional approaches typically employ maximum likelihood estimation as the optimization objective during training, yet fail to align the empathy levels between generated and target responses. To this end, we propose an empathetic response generation framework using reinforcement learning (EmpRL). The framework develops an effective empathy reward function and generates empathetic responses by maximizing the expected reward through reinforcement learning. EmpRL utilizes the pre-trained T5 model as the generator and further fine-tunes it to initialize the policy. To align the empathy levels between generated and target responses within a given context, an empathy reward function containing three empathy communication mechanisms—emotional reaction, interpretation, and exploration—is constructed using pre-designed and pre-trained empathy identifiers. During reinforcement learning training, the proximal policy optimization algorithm is used to fine-tune the policy, enabling the generation of empathetic responses. Both automatic and human evaluations demonstrate that the proposed EmpRL framework significantly improves the quality of generated responses, enhances the similarity in empathy levels between generated and target responses, and produces empathetic responses covering both affective and cognitive aspects. Hui Ma 0011, Bo Zhang 0121, Bo Xu 0009, Jian Wang 0021, Hongfei Lin, Xiao Sun 0003 |
IEEE Trans. Affect. Comput. | 1 |
| 2025 | PsycoLLM: Enhancing LLM for Psychological Understanding and EvaluationabstractMental health has attracted substantial attention in recent years and large language model (LLM) can be an effective technology for alleviating this problem owing to its capability in text understanding and dialogue. However, existing research in this domain often suffers from limitations, such as training on datasets lacking crucial prior knowledge and evidence, and the absence of comprehensive evaluation methods. In this article, we propose a specialized psychological LLM, named PsycoLLM, trained on a proposed high-quality psychological dataset, including single-turn QA, multiturn dialogues, and knowledge-based QA. Specifically, we construct multi-turn dialogues through a three-step pipeline comprising multiturn QA generation, evidence judgment, and dialogue refinement. We augment this process with real-world psychological case backgrounds extracted from online platforms, enhancing the relevance and applicability of the generated data. Additionally, to compare the performance of PsycoLLM with other LLMs, we develop a comprehensive psychological benchmark based on authoritative psychological counseling examinations in China, which includes assessments of professional ethics, theoretical proficiency, and case analysis. The experimental results on the benchmark illustrate the effectiveness of PsycoLLM, which demonstrates superior performance compared with other LLMs. Jinpeng Hu, Tengteng Dong, Hui Ma 0011, Xiao Sun 0003, Dan Guo 0001, Xun Yang 0001, Meng Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | A Transformer-Based Model With Self-Distillation for Multimodal Emotion Recognition in ConversationsabstractEmotion recognition in conversations (ERC), the task of recognizing the emotion of each utterance in a conversation, is crucial for building empathetic machines. Existing studies focus mainly on capturing context- and speaker-sensitive dependencies on the textual modality but ignore the significance of multimodal information. Different from emotion recognition in textual conversations, capturing intra- and inter-modal interactions between utterances, learning weights between different modalities, and enhancing modal representations play important roles in multimodal ERC. In this paper, we propose a transformer-based model with self-distillation (SDT)11The code is available athttps://github.com/butterfliesss/SDT.for the task. The transformer-based model captures intra- and inter-modal interactions by utilizing intra- and inter-modal transformers, and learns weights between modalities dynamically by designing a hierarchical gated fusion strategy. Furthermore, to learn more expressive modal representations, we treat soft labels of the proposed model as extra training supervision. Specifically, we introduce self-distillation to transfer knowledge of hard and soft labels from the proposed model to each modality. Experiments on IEMOCAP and MELD datasets demonstrate that SDT outperforms previous state-of-the-art baselines. Hui Ma 0011, Jian Wang 0021, Hongfei Lin, Bo Zhang 0121, Yi-Jia Zhang 0001, Bo Xu 0009 |
IEEE Trans. Multim. | 1 |
| 2023 | ZRIGF: An Innovative Multimodal Framework for Zero-Resource Image-Grounded Dialogue GenerationabstractImage-grounded dialogue systems benefit greatly from integrating visual information, resulting in high-quality response generation. However, current models struggle to effectively utilize such information in zero-resource scenarios, mainly due to the disparity between image and text modalities. To overcome this challenge, we propose an innovative multimodal framework, called ZRIGF, which assimilates image-grounded information for dialogue generation in zero-resource situations. ZRIGF implements a two-stage learning strategy, comprising contrastive pre-training and generative pre-training. Contrastive pre-training includes a text-image matching module that maps images and texts into a unified encoded vector space, along with a text-assisted masked image modeling module that preserves pre-training visual features and fosters further multimodal feature alignment. Generative pre-training employs a multimodal fusion module and an information transfer module to produce insightful responses based on harmonized multimodal representations. Comprehensive experiments conducted on both text-based and image-grounded dialogue datasets demonstrate ZRIGF's efficacy in generating contextually pertinent and informative responses. Furthermore, we adopt a fully zero-resource scenario in the image-grounded dialogue dataset to demonstrate our framework's robust generalization capabilities in novel domains. Bo Zhang 0121, Jian Wang 0021, Hui Ma 0011, Bo Xu 0009, Hongfei Lin |
ACM Multimedia | 3 |
| 2023 | Graph augmented sequence-to-sequence model for neural question generation
Hui Ma 0011, Jian Wang 0021, Hongfei Lin, Bo Xu 0009 |
Appl. Intell. | 1 |
| 2022 | Global and local interaction matching model for knowledge-grounded response selection in retrieval-based chatbots
Hui Ma 0011, Jian Wang 0021, Hongfei Lin, Liang Yang 0003 |
Neurocomputing | 1 |
| 2022 | A multi-view network for real-time emotion recognition in conversations
Hui Ma 0011, Jian Wang 0021, Hongfei Lin, Xuejun Pan, Yi-Jia Zhang 0001 |
Knowl. Based Syst. | 1 |
| 2022 | Exploiting Pairwise Mutual Information for Knowledge-Grounded DialogueabstractExternal document knowledge is helpful for dialogue systems to generate high-quality responses. Although several knowledge-grounded dialogue models have been designed, external knowledge cannot be comprehensively exploited due to the complex relationships among dialogue context, knowledge, and responses. To this end, we propose a novel transformer-based model, named TransIKG, which incorporates external document knowledge for dialogue generation. TransIKG comprises a two-step integration mechanism, including correlation integration and overall integration. Correlation integration is designed to fully exploit the pairwise mutual information among dialogue context, knowledge, and responses, while overall integration adopts an integration gate to capture global information. Furthermore, we utilize the positional information of dialogue turns to better represent the dialogue context and enhance the generalization ability of our model on out-of-domain documents. Finally, we propose a novel knowledge-aware pointer network to generate knowledge-enhanced response tokens. Experimental results on two benchmark datasets demonstrate that our model outperforms state-of-the-art models on both open-domain and domain-specific dialogues. Bo Zhang 0121, Jian Wang 0021, Hongfei Lin, Hui Ma 0011, Bo Xu 0009 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Co-Attentive Span Network with Multi-task learning for Biomedical Named Entity RecognitionabstractBiomedical Named Entity Recognition (BioNER) is often modeled as a sequence labeling task, which assigns the predefined label to each token in given input sequence. Although these sequential labeling models achieve significant achievements, they often fail to give precise boundaries of the named entity. In addition, a vast amount of work focuses more on textual sequence representation but ignores label information. To tackle these problems, in this paper, we directly model span-level named entity recognition, specifically, we treat the BioNER as a joint task of boundary detection and span classification under a multitask framework. In order to enhance boundary supervision, we introduce an entity type label as an additional guide and propose a co-attentive interactive mechanism to improve the span representation. Extensive experiments1on four benchmark datasets demonstrate that our proposed method obtains competitive results, achieving 90.26%, 78.04%, 90.21%, and 86.58% on BC5CDR, JNLPBA, NCBI, and BC2GM datasets, respectively, in terms of F1 score. Jian Wang 0021, Hongfei Lin, Yi-Jia Zhang 0001, Di Zhao 0003, Hui Ma 0011 |
BIBM | 7 |
| 2021 | HAN-ReGRU: hierarchical attention network with residual gated recurrent unit for emotion recognition in conversation
Hui Ma 0011, Jian Wang 0021, Lingfei Qian, Hongfei Lin |
Neural Comput. Appl. | 1 |
| 2021 | Hierarchical matching network for multi-turn response selection in retrieval-based chatbots
Hui Ma 0011, Jian Wang 0021, Hongfei Lin, Yi-Jia Zhang 0001 |
Soft Comput. | 1 |