Ronghan Li

dblp:227/4510 · DBLP profile ↗
← Back
25ranked-venue papers
8as first author
24since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 7 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Sticker-Enriched Empathetic Response Generation: A Role-Aware Pairing Benchmark and A Novel Multimodal Framework
abstract
Stickers significantly enhance emotional expression in empathetic dialogue, yet research remains predominantly text-centric. Despite the emergence of multimodal empathetic dialogue datasets, two critical deficiencies persist: (1) a stylistic misalignment between the subtle cues essential for empathy and the unconstrained nature of open-domain stickers; and (2) role-agnostic sticker–utterance pairing that neglects the asymmetric communicative objectives of interlocutors. To bridge these gaps, we introduce StickerED, a large-scale sticker-enriched empathetic dialogue dataset that ensures stylistic harmony and preserves the functional divergence in sticker selection between speaker intent and listener feedback. Building upon this, we propose SEERG, a Sticker-Enhanced Empathetic Response Generation framework. Notably, SEERG augments standard visual features with explicit intent and style guidance, employing dual utterance-level cross-modal fusion mechanisms to enrich contextual representations. Through multi-task learning, SEERG generates contextually and emotionally resonant responses. Extensive experiments demonstrate that our multi-dimensional sticker enrichment yields substantial gains over text-only baselines, underscoring the pivotal role of stickers in advancing empathetic dialogue systems.
Wei Zhang 0267, Changhong Jiang, Mingting Yu, Ronghan Li
ICMR6
2026 Ubuntu-IntAct : A dialogue-act annotated multi-party conversation dataset and graph-based intention modeling
Zhongtian Hu, Changhong Jiang, Yiwen Cui, Ronghan Li, Wei Zhang 0267, Jiashi Lin
Inf. Process. Manag.4
2026 PDPA: A prompt-based dual persona-aware approach for empathetic response generation
Wei Zhang 0267, Changhong Jiang, Ming Xia 0002, Zhongtian Hu, Jiashi Lin, Ronghan Li
Knowl. Based Syst.7
2026 Query expansion with topic-aware in-context learning and vocabulary projection for open-domain dense retrieval
Ronghan Li, Mingze Cui, Benben Wang, Yu Wang 0313, Qiguang Miao
Pattern Recognit.1
2025 Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models
abstract
The rapid advancements in Vision Language Models (VLMs) have prompted the development of multi-modal medical assistant systems. Despite this progress, current models still have inherent probabilistic uncertainties, often producing erroneous or unverified responses-an issue with serious implications in medical applications. Existing methods aim to enhance the performance of Medical Vision Language Model (MedVLM) by adjusting model structure, fine-tuning with high-quality data, or through preference fine-tuning. However, these training-dependent strategies are costly and still lack sufficient alignment with clinical expertise. To address these issues, we propose an expert-in-the-loop framework named Expert-Controlled Classifier-Free Guidance (Expert-CFG) to align MedVLM with clinical expertise without additional training. This framework introduces an uncertainty estimation strategy to identify unreliable outputs. It then retrieves relevant references to assist experts in highlighting key terms and applies classifier-free guidance to refine the token embeddings of MedVLM, ensuring that the adjusted outputs are correct and align with expert highlights. Evaluations across three medical visual question answering benchmarks demonstrate that the proposed Expert-CFG, with 4.2B parameters and limited expert annotations, outperforms state-of-the-art models with 13B parameters. The results demonstrate the feasibility of deploying such a system in resource-limited settings for clinical use.
Di Wang 0011, Zhicheng Jiao, Ronghan Li, Pengfei Yang 0001, Quan Wang 0006, Tat-Seng Chua
ICCV4
2025 CheXPO: Preference Optimization for Chest X-ray VLMs with Counterfactual Rationale
abstract
Vision-language models (VLMs) are prone to hallucinations that critically compromise reliability in medical applications. While preference optimization can mitigate these hallucinations through clinical feedback, its implementation faces challenges such as clinically irrelevant training samples, imbalanced data distributions, and prohibitive expert annotation costs. To address these challenges, we introduce CheXPO, a Chest X-ray Preference Optimization strategy that combines confidence-similarity joint mining with counterfactual rationale. Our approach begins by synthesizing a unified, fine-grained multi-task chest X-ray visual instruction dataset across different question types for supervised fine-tuning (SFT). We then identify hard examples through token-level confidence analysis of SFT failures and use similarity-based retrieval to expand hard examples for balancing preference sample distributions, while synthetic counterfactual rationales provide fine-grained clinical preferences, eliminating the need for additional expert input. Experiments show that CheXPO achieves 8.93% relative performance gain using only 5% of SFT samples, reaching state-of-the-art performance across diverse clinical tasks. 1Code: https://github.com/ResearchGroup-MedVLLM/CheX-Phi35V
Di Wang 0011, Lin Zhao 0003, Ronghan Li, Bo Wan 0002, Quan Wang 0006
ACM Multimedia6
2025 UniRQR: A Unified Model for Retrieval Decision, Query, and Response Generation in internet-based knowledge dialogue systems
Zhongtian Hu, Yangqi Chen, Meng Zhao 0004, Ronghan Li
Expert Syst. Appl.4
2025 RECoT: Relation-enhanced Chains-of-Thoughts for knowledge-intensive multi-hop questions answering
Ronghan Li, Haoxiang Jin, Rongcheng Pu, Qiguang Miao
Neurocomputing1
2025 RoleCF: Role-oriented coarse-to-fine emotion cause recognition for empathetic response generation
Wei Zhang 0267, Ming Xia 0002, Ronghan Li, Zhongtian Hu, Jiashi Lin
Neurocomputing4
2025 Different paths to the same destination: Diversifying LLMs generation for multi-hop open-domain question answering
Ronghan Li, Yu Wang 0313, Zijian Wen, Mingze Cui, Qiguang Miao
Knowl. Based Syst.1
2024 Divide and Conquer: Isolating Normal-Abnormal Attributes in Knowledge Graph-Enhanced Radiology Report Generation
abstract
Radiology report generation aims to automatically generate clinical descriptions for radiology images, reducing the workload of radiologists. Compared to general image captioning tasks, the subtle differences in medical images and the specialized, complex nature of medical terminology limit the performance of data-driven radiology report generation. Previous research has attempted to leverage prior knowledge, such as organ-disease graphs, to enhance models' abilities to identify specific diseases and generate corresponding medical terminology. However, these methods cover only a limited number of disease types, focusing solely on disease terms mentioned in reports but ignoring their normal or abnormal attributes, which are critical to generating accurate reports. To address this issue, we propose a Divide-and-Conquer approach, named DCG, which separately constructs disease-free and disease-specific nodes within the knowledge graphs. Specifically, we extracted more comprehensive organ-disease entities from reports than previous methods and constructed disease-free and disease-specific nodes by rigorously distinguishing between normal conditions and specific diseases. This enables our model to consciously focus on abnormal information and mitigate the impact of excessively common diseases on report generation. Subsequently, the constructed graph is utilized to enhance the correlation between visual representations and disease terminology, thereby guiding the decoder in report generation. Extensive experiments conducted on benchmark datasets IU-Xray and MIMIC-CXR demonstrate the superiority of our proposed method. Code is available at https://github.com/ecoxial2007/DCG_Enhanced_distilGPT2.
Yanlei Zhang, Di Wang 0011, Haodi Zhong, Ronghan Li, Quan Wang 0006
ACM Multimedia5
2024 Dynamically retrieving knowledge via query generation for informative dialogue generation
Zhongtian Hu, Yangqi Chen, Yushuang Liu, Ronghan Li, Meng Zhao 0004, Ze-Jun Jiang
Neurocomputing5
2024 Candidate-Heuristic In-Context Learning: A new framework for enhancing medical visual question answering with LLMs
Di Wang 0011, Haodi Zhong, Quan Wang 0006, Ronghan Li, Rui Jia, Bo Wan 0002
Inf. Process. Manag.5
2024 Dialogue summarization enhanced response generation for multi-domain task-oriented dialogue systems
Meng Zhao 0004, Hongru Ji, Ze-Jun Jiang, Ronghan Li, Zhongtian Hu
Inf. Process. Manag.5
2024 From easy to hard: Improving personalized response generation of task-oriented dialogue systems by leveraging capacity in open-domain dialogues
Meng Zhao 0004, Ze-Jun Jiang, Yushuang Liu, Ronghan Li, Zhongtian Hu
Knowl. Based Syst.5
2024 Transformer with convolution and graph-node co-embedding: An accurate and interpretable vision backbone for predicting gene expressions from local histopathological image
abstract
Inferring gene expressions from histopathological images has long been a fascinating yet challenging task, primarily due to the substantial disparities between the two modality. Existing strategies using local or global features of histological images are suffering model complexity, GPU consumption, low interpretability, insufficient encoding of local features, and over-smooth prediction of gene expressions among neighboring sites. In this paper, we develop TCGN (Transformer with Convolution and Graph-Node co-embedding method) for gene expression estimation from H&E-stained pathological slide images. TCGN comprises a combination of convolutional layers, transformer encoders, and graph neural networks, and is the first to integrate these blocks in a general and interpretable computer vision backbone. Notably, TCGN uniquely operates with just a single spot image as input for histopathological image analysis, simplifying the process while maintaining interpretability. We validate TCGN on three publicly available spatial transcriptomic datasets. TCGN consistently exhibited the best performance (with median PCC 0.232). TCGN offers superior accuracy while keeping parameters to a minimum (just 86.241 million), and it consumes minimal memory, allowing it to run smoothly even on personal computers. Moreover, TCGN can be extended to handle bulk RNA-seq data while providing the interpretability. Enhancing the accuracy of omics information prediction from pathological images not only establishes a connection between genotype and phenotype, enabling the prediction of costly-to-measure biomarkers from affordable histopathological images, but also lays the groundwork for future multi-modal data modeling. Our results confirm that TCGN is a powerful tool for inferring gene expressions from histopathological images in precision health applications.
Yan Kong, Ronghan Li, Zuoheng Wang, Hui Lu 0004
Medical Image Anal.3
2023 Mutually improved response generation and dialogue summarization for multi-domain task-oriented dialogue systems
Meng Zhao 0004, Hongru Ji, Ze-Jun Jiang, Ronghan Li, Zhongtian Hu
Knowl. Based Syst.5
2023 Multi-task learning with graph attention networks for multi-domain task-oriented dialogue systems
Meng Zhao 0004, Ze-Jun Jiang, Ronghan Li, Zhongtian Hu
Knowl. Based Syst.4
2022 An effective context-focused hierarchical mechanism for task-oriented dialogue response generation
abstract
Abstract Task‐oriented dialogue system (TOD) is one kind of application of artificial intelligence (AI). The response generation module is a key component of TOD for replying to user's questions and concerns in sequential natural words. In the past few years, the works on response generation have attracted increasing research attention and have seen much progress. However, existing works ignore the fact that not each turn of dialogue history contributes to the dialogue response generation and give little consideration to the different weights of utterances in a dialogue history. In this article, we propose a hierarchical memory network mechanism with two steps to filter out unnecessary information of dialogue history. First, an utterance‐level memory network distributes various weights to each utterance (coarse‐grained). Second, a token‐level memory network assigns higher weights to keywords based on the former's output (fine‐grained). Furthermore, the output of the token‐level memory network will be employed to query the knowledge base (KB) to capture the dialogue‐related information. In the decoding stage, we take a gated‐mechanism to generate response word by word from dialogue history, vocabulary, or KB. Experiments show that the proposed model achieves superior results compared with state‐of‐the‐art models on several public datasets. Further analysis demonstrates the effectiveness of the proposed method and the robustness of the model in the case of an incomplete training set.
Meng Zhao 0004, Ze-Jun Jiang, Ronghan Li, Zhongtian Hu, Da-Qing Chen 0001
Comput. Intell.4
2022 A multiturn complementary generative framework for conversational emotion recognition
abstract
Conversational emotion recognition (CER) is a significant task due to its application in human–computer interaction. Existing work treats CER as an utterance-level classification task without considering that empathic response also reflects contextual emotion understanding. Previous work has proven that accurate recognition of emotions in the dialogue history is helpful to generate high-fit responses. In this paper, we investigate whether this conclusion is a sufficient and necessary condition. Specifically, we define an auxiliary empathic multiturn dialogue generation (MDG) task to enhance emotion understanding. Correspondingly, we present a Sequence-to-Sequence oriented framework that combines CER and MDG in a multitask learning manner to verify the complementarity between the two tasks. First, we use alternate recurrent neural networks to encode the content of historical utterances and represent the states of multiparty emotions, which are used for emotion classification. Second, since most MDG methods ignore the emotional coherence of the dialogue context itself, we use affine transformation to fuse hidden states of content and emotions to initialize the decoder. Finally, at each step of generation, an attention mechanism is used to fuse information from the dialogue history to ensure emotional coherence. The CER results of our models outperform the state-of-the-art on three prevalent emotional dialogue data sets. Further analysis demonstrates the mutual promotion and empathy interpretability between MDG and CER. Furthermore, our framework is scalable for different coding strategies and multimodal fusion. To the best of our knowledge, this is the first work to explore CER from the perspective of empathy through multitask learning with dialogue generation.
Ronghan Li, Ze-Jun Jiang
Int. J. Intell. Syst.2
2022 Mutually improved dense retriever and GNN-based reader for arbitrary-hop open-domain question answering
Ronghan Li, Ze-Jun Jiang, Zhongtian Hu, Meng Zhao 0004
Neural Comput. Appl.1
2021 Asynchronous Multi-grained Graph Network For Interpretable Multi-hop Reading Comprehension
abstract
Multi-hop machine reading comprehension (MRC) task aims to enable models to answer the compound question according to the bridging information. Existing methods that use graph neural networks to represent multiple granularities such as entities and sentences in documents update all nodes synchronously, ignoring the fact that multi-hop reasoning has a certain logical order across granular levels. In this paper, we introduce an Asynchronous Multi-grained Graph Network (AMGN) for multi-hop MRC. First, we construct a multigrained graph containing entity and sentence nodes. Particularly, we use independent parameters to represent relationship groups defined according to the level of granularity. Second, an asynchronous update mechanism based on multi-grained relationships is proposed to mimic human multi-hop reading logic. Besides, we present a question reformulation mechanism to update the latent representation of the compound question with updated graph nodes. We evaluate the proposed model on the HotpotQA dataset and achieve top competitive performance in distractor setting compared with other published models. Further analysis shows that the asynchronous update mechanism can effectively form interpretable reasoning chains at different granularity levels.
Ronghan Li, Shengli Wang, Ze-Jun Jiang
IJCAI1
2021 Enhancing Transformer-based language models with commonsense representations for knowledge-driven machine comprehension
Ronghan Li, Ze-Jun Jiang, Meng Zhao 0004, Da-Qing Chen 0001
Knowl. Based Syst.1
2021 Incremental BERT with commonsense representations for multi-choice reading comprehension
Ronghan Li, Ze-Jun Jiang, Meng Zhao 0004
Multim. Tools Appl.1
2020 Directional attention weaving for text-grounded conversational question answering
Ronghan Li, Ze-Jun Jiang, Meng Zhao 0004
Neurocomputing1