Qingqiu Li

dblp:283/3034 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
13since 2021 · last 2025
0009-0000-4848-4583ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Fine-Grained Knowledge-Guided Alignment for Medical Vision-Language Pre-Training
abstract
Medical contrastive Vision-Language Pre-training (VLP) has emerged as a promising approach, enabling models to learn joint representations from paired medical images and radiology reports. Despite existing methods exploring local visual representation learning techniques, they often fall short in local alignment and knowledge infusion, e.g., uniform token treatment and isolated knowledge assignment. To address these issues, we propose a novel Fine-grained Knowledge-Guided Alignment (FKGA) framework for medical VLP. Specifically, we propose a Fine-grained Disease Knowledge Integration (FDKI) module to inject detailed disease descriptions into corresponding disease tokens in reports. Based on these semantic-enriched tokens, we introduce global instance-wise and local token-wise contrastive learning to further align the semantically related visual and textual modalities. In contrast to previous local visual representation learning methods, our design of semantic-enriched token alignment and context-preserved knowledge infusion enhances the semantic understanding of diseases. Extensive experimental results on five downstream tasks demonstrate that our proposed method outperforms other state-of-the-art methods across seven datasets.
Yaning Pan, Ying Cheng 0005, Qingqiu Li, Runtian Yuan, Rui Feng 0001
BIBM4
2025 Human Simulacra: Benchmarking the Personification of Large Language Models
abstract
Large Language Models (LLMs) are recognized as systems that closely mimic aspects of human intelligence. This capability has attracted the attention of the social science community, who see the potential in leveraging LLMs to replace human participants in experiments, thereby reducing research costs and complexity. In this paper, we introduce a benchmark for LLMs personification, including a strategy for constructing virtual characters' life stories from the ground up, a Multi-Agent Cognitive Mechanism capable of simulating human cognitive processes, and a psychology-guided evaluation method to assess human simulations from both self and observational perspectives. Experimental results demonstrate that our constructed simulacra can produce personified responses that align with their target characters. We hope this work will serve as a benchmark in the field of human simulation, paving the way for future research.
Qiujie Xie, Qiming Feng, Qingqiu Li, Linyi Yang, Yuejie Zhang, Rui Feng 0001, Shang Gao 0003, Yue Zhang 0004
ICLR4
2025 An Empirical Analysis of Uncertainty in Large Language Model Evaluations
abstract
As LLM-as-a-Judge emerges as a new paradigm for assessing large language models (LLMs), concerns have been raised regarding the alignment, bias, and stability of LLM evaluators. While substantial work has focused on alignment and bias, little research has concentrated on the stability of LLM evaluators. In this paper, we conduct extensive experiments involving 9 widely used LLM evaluators across 2 different evaluation settings to investigate the uncertainty in model-based LLM evaluations. We pinpoint that LLM evaluators exhibit varying uncertainty based on model families and sizes. With careful comparative analyses, we find that employing special prompting strategies, whether during inference or post-training, can alleviate evaluation uncertainty to some extent. By utilizing uncertainty to enhance LLM's reliability and detection capability in Out-Of-Distribution (OOD) data, we further fine-tune an uncertainty-aware LLM evaluator named ConfiLM using a human-annotated fine-tuning set and assess ConfiLM's OOD evaluation ability on a manually designed test set sourced from the 2024 Olympics. Experimental results demonstrate that incorporating uncertainty as additional information during the fine-tuning phase can largely improve the model's evaluation performance in OOD scenarios. The code and data are released at: https://github.com/hasakiXie123/LLM-Evaluator-Uncertainty.
Qiujie Xie, Qingqiu Li, Zhuohao Yu 0001, Yuejie Zhang, Yue Zhang 0004, Linyi Yang
ICLR2
2025 EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
abstract
Generating videos in the first-person perspective has broad application prospects in the field of augmented reality and embodied intelligence. In this work, we explore the cross-view video prediction task, where given an exo-centric video, the first frame of the corresponding ego-centric video, and textual instructions, the goal is to generate future frames of the ego-centric video. Inspired by the notion that hand-object interactions (HOI) in ego-centric videos represent the primary intentions and actions of the current actor, we present EgoExo-Gen that explicitly models the hand-object dynamics for cross-view video prediction. EgoExo-Gen consists of two stages. First, we design a cross-view HOI mask prediction model that anticipates the HOI masks in future ego-frames by modeling the spatio-temporal ego-exo correspondence. Next, we employ a video diffusion model to predict future ego-frames using the first ego-frame and textual instructions, while incorporating the HOI masks as structural guidance to enhance prediction quality. To facilitate training, we develop a fully automated pipeline to generate pseudo HOI masks for both ego- and exo-videos by exploiting vision foundation models. Extensive experiments demonstrate that our proposed EgoExo-Gen achieves better prediction performance compared to previous video prediction models on the public Ego-Exo4D and H2O benchmark datasets, with the HOI masks significantly improving the generation of hands and interactive objects in the ego-centric videos.
Jilan Xu, Yifei Huang 0002, Baoqi Pei, Junlin Hou, Qingqiu Li, Guo Chen 0006, Yuejie Zhang, Rui Feng 0001, Weidi Xie
ICLR5
2025 TGSAM-2: Text-Guided Medical Image Segmentation Using Segment Anything Model 2
Runtian Yuan, Ling Zhou 0002, Jilan Xu, Qingqiu Li, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
MICCAI (10)4
2025 Text-Promptable Propagation for Referring Medical Image Sequence Segmentation
abstract
Referring Medical Image Sequence Segmentation (Ref-MISS) is a novel and challenging task that aims to segment anatomical structures in medical image sequences (e.g., endoscopy, ultrasound, CT, and MRI) based on natural language descriptions. Existing 2D and 3D segmentation models struggle to explicitly track objects of interest across medical image sequences, and lack support for interactive, text-driven guidance. To address these limitations, we propose Text-Promptable Propagation (TPP), which enables the recognition of referred objects through cross-modal referring interaction, and maintains continuous tracking across the sequence via Transformer-based triple propagation, using text embeddings as queries. To support this task, we curate a large-scale benchmark, Ref-MISS-Bench, which covers 4 imaging modalities and 20 different organs and lesions. Experimental results on this benchmark demonstrate that TPP consistently outperforms state-of-the-art methods in both medical segmentation and referring video object segmentation. Code and data are available at https://github.com/yuanruntian/TPP.
Runtian Yuan, Mohan Chen 0001, Jilan Xu, Ling Zhou 0002, Qingqiu Li, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ACM Multimedia5
2025 EmoCharacter: Evaluating the Emotional Fidelity of Role-Playing Agents in Dialogues
abstract
Qiming Feng, Qiujie Xie, Xiaolong Wang, Qingqiu Li, Yuejie Zhang, Rui Feng, Tao Zhang, Shang Gao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Qiming Feng, Qiujie Xie, Qingqiu Li, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
NAACL (Long Papers)4
2025 AOR: Anatomical Ontology-Guided Reasoning for Medical Large Multimodal Model in Chest X-Ray Interpretation
abstract
Chest X-rays (CXRs) are the most frequently performed imaging examinations in clinical settings. Recent advancements in Medical Large Multimodal Models (MLMMs) have enabled automated CXR interpretation, improving diagnostic accuracy and efficiency. However, despite their strong visual understanding, current MLMMs still face two major challenges: (1) insufficient region-level understanding and interaction, and (2) limited accuracy and interpretability due to single-step prediction. In this paper, we address these challenges by empowering MLMMs with anatomy-centric reasoning capabilities to enhance their interactivity and explainability. Specifically, we propose an Anatomical Ontology-Guided Reasoning (AOR) framework that accommodates both textual and optional visual prompts, centered on region-level information to enable multimodal multi-step reasoning. We also develop AOR-Instruction, a large instruction dataset for MLMMs training, under the guidance of expert physicians. Our experiments demonstrate AOR's superior performance in both Visual Question Answering (VQA) and report generation tasks. Code and data are available at: https://github.com/Liqq1/AOR.
Qingqiu Li, Zihang Cui, Seongsu Bae, Jilan Xu, Runtian Yuan, Yuejie Zhang, Rui Feng 0001, Quanli Shen, Shang Gao 0003, Junjun He
NeurIPS1
2024 Anatomical Structure-Guided Medical Vision-Language Pre-training
Qingqiu Li, Xiaohan Yan, Jilan Xu, Runtian Yuan, Yuejie Zhang, Rui Feng 0001, Quanli Shen
MICCAI (11)1
2023 Enhanced Knowledge Injection for Radiology Report Generation
abstract
Automatic generation of radiology reports holds crucial clinical value, as it can alleviate substantial workload on radiologists and remind less experienced ones of potential anomalies. Despite the remarkable performance of various image captioning methods in the natural image field, generating accurate reports for medical images still faces challenges, i.e., disparities in visual and textual data, and lack of accurate domain knowledge. To address these issues, we propose an enhanced knowledge injection framework, which utilizes two branches to extract different types of knowledge. The Weighted Concept Knowledge (WCK) branch is responsible for introducing clinical medical concepts weighted by TF-IDF scores. The Multimodal Retrieval Knowledge (MRK) branch extracts triplets from similar reports, emphasizing crucial clinical information related to entity positions and existence. By integrating this finer-grained and well-structured knowledge with the current image, we are able to leverage the multi-source knowledge gain to ultimately facilitate more accurate report generation. Extensive experiments have been conducted on two public benchmarks, demonstrating that our method achieves superior performance over other state-of-the-art methods. Ablation studies further validate the effectiveness of two extracted knowledge sources.
Qingqiu Li, Jilan Xu, Runtian Yuan, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Shang Gao 0003
BIBM1
2023 Semi-MedSeq: Semi-supervised Semantic Segmentation for Medical Image Sequences
abstract
In clinical practice, medical imaging techniques include 2D video-based examinations that capture sequential scans, and 3D volumetric imaging that forms a comprehensive 3D representation from a stack of 2D slices. The medical image sequences produced by the above techniques provide valuable spatio-temporal characteristics for analysis and segmentation, but the annotation of image sequences is extremely time-consuming and labor-intensive. To exploit the coherence and address the scarcity of labeled data, we propose a novel semi-supervised semantic segmentation framework for medical image sequences, which consists of a conditional network and a denoising network. Specifically, we embed a Sequential Feature Reconstruction module into both networks. This module reconstructs the target frame from contiguous frames and captures their shared visual features. Guided by the context-enhancing information from the conditioning network, the denoising network suppresses background noise via a Diffusion-based Noise Elimination module. Extensive experiments are conducted on 2D and 3D tasks, including cardiac segmentation, polyp segmentation, placenta vessel segmentation and abdomen multi-organ segmentation. The results show our method is superior to existing semi-supervised methods and exhibits advantages over fully-supervised medical image segmentation methods with only 1/2 labeled data, validating its effectiveness and generalization ability.
Runtian Yuan, Jilan Xu, Qingqiu Li, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
BIBM3
2023 SCSGNet: Spatial-Correlated and Shape-Guided Network for Breast Mass Segmentation
abstract
Automatic and accurate breast mass segmentation plays a crucial role in the early diagnosis of breast cancer. However, it has been a challenging task for two main reasons: (1) Breast masses are diverse; and (2) The boundaries of masses are ambiguous. To address these problems, we propose a Spatial-Correlated and Shape-Guided Network (SCSGNet), which combines global context extraction with local boundary refinement. Specifically, the high-level features are aggregated to produce a global map as the initial guidance area, and a Series-Parallel Feature Fusion (SPFF) module is added to capture masses of different shapes and sizes. Besides, we design a Dynamic Long-range Correlation Capture (DLCC) module to capture the spatial correlation of masses at different positions. Finally, we devise a Triplet Attention Guide (TAG) module to iteratively update the feature map and refine the boundary. Experiments on two public datasets demonstrate that our method achieves superior performance over other state-of-the-art methods.
Qingqiu Li, Jilan Xu, Runtian Yuan, Yuejie Zhang, Rui Feng 0001
ICASSP1
2021 Parallel Deep Learning Algorithms With Hybrid Attention Mechanism for Image Segmentation of Lung Tumors
abstract
At present, medical images have played a more and more important role in clinical treatment. Lung images provide an important reference for doctors to make a diagnosis. Especially for surgical patients, a tumor can be accurately removed based on the full cognition about its size, position, and quantity. Therefore, computer-aided diagnosis for the analysis and treatment of a lot of lung tumor images is very important. Aiming at complexity and self-adaption of image segmentation in lung tumors, this article proposed a parallel deep learning algorithm with hybrid attention mechanism for image segmentation. First, lung parenchyma was extracted via preprocessing images. Then, images were input into hybrid attention mechanism and densely connected convolutional networks (DenseNet) module, respectively, where hybrid attention mechanism consisted of a spatial attention mechanism and a channel attention mechanism. Finally, four feasible solutions were proposed for the verification through changing the convolution quantity of dense block in DenseNet. The network structure with the better performance was achieved. The experimental results prove the parallel deep learning algorithm with hybrid attention mechanism performed well in image segmentation of lung tumors, and its accuracy can reach 94.61%.
Hexuan Hu 0001, Qingqiu Li, Ye Zhang 0010
IEEE Trans. Ind. Informatics2