Qingqing Zhu

dblp:288/7124 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
16since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 How well do multimodal LLMs interpret CT scans? An auto-evaluation framework for analyses
abstract
OBJECTIVE: This study introduces a novel evaluation framework, GPTRadScore, to systematically assess the performance of multimodal large language models (MLLMs) in generating clinically accurate findings from CT imaging. Specifically, GPTRadScore leverages LLMs as an evaluation metric, aiming to provide a more accurate and clinically informed assessment than traditional language-specific methods. Using this framework, we evaluate the capability of several MLLMs, including GPT-4 with Vision (GPT-4V), Gemini Pro Vision, LLaVA-Med, and RadFM, to interpret findings in CT scans. METHODS: This retrospective study leverages a subset of the public DeepLesion dataset to evaluate the performance of several multimodal LLMs in describing findings in CT slices. GPTRadScore was developed to assess the generated descriptions (location, body part, and type) using GPT-4, alongside traditional metrics. RadFM was fine-tuned using a subset of the DeepLesion dataset with additional labeled examples targeting complex findings. Post fine-tuning, performance was reassessed using GPTRadScore to measure accuracy improvements. RESULTS: Evaluations demonstrated a high correlation of GPTRadScore with clinician assessments, with Pearson's correlation coefficients of 0.87, 0.91, 0.75, 0.90, and 0.89. These results highlight its superiority over traditional metrics, such as BLEU, METEOR, and ROUGE, and indicate that GPTRadScore can serve as a reliable evaluation metric. Using GPTRadScore, it was observed that while GPT-4V and Gemini Pro Vision outperformed other models, significant areas for improvement remain, primarily due to limitations in the datasets used for training. Fine-tuning RadFM resulted in substantial accuracy gains: location accuracy increased from 3.41% to 12.8%, body part accuracy improved from 29.12% to 53%, and type accuracy rose from 9.24% to 30%. These findings reinforce the hypothesis that fine-tuning RadFM can significantly enhance its performance. CONCLUSION: GPT-4 effectively correlates with expert assessments, validating its use as a reliable metric for evaluating multimodal LLMs in radiological diagnostics. Additionally, the results underscore the efficacy of fine-tuning approaches in improving the descriptive accuracy of LLM-generated medical imaging findings.
Qingqing Zhu, Benjamin Hou, Tejas Sudharshan Mathai, Pritam Mukherjee, Qiao Jin 0001, Xiuying Chen, Zhizheng Wang, Ruida Cheng, Ronald M. Summers, Zhiyong Lu
J. Biomed. Informatics1
2025 Modified PUMA/EPUMA Based on Forward and Backward Linear Prediction for DOA Estimation
abstract
The principal-singular-vector utilization for modal analysis (PUMA) and its modification (Mod-PUMA), which utilize forward linear prediction (FLP) to process the signal subspace, experience significant performance degradation if there are multiple coherent sources and such a performance degradation will be further aggravated in low-SNR regions, which is primarily attributed to the outliers arising from inaccurate estimations of the signal subspace. To address these issues, we propose an extension version of PUMA-related algorithms, called FBLP-Mod-PUMA/enhanced-PUMA (EPUMA). The proposed algorithms improve the threshold performance by refining the signal subspace through forward and backward linear prediction (FBLP), effectively mitigating subspace leakage when dealing with coherent sources. The number of resolvable coherent sources has been theoretically derived and simulation results are provided to show the performance of the proposed algorithms.
Biyun Ma, Fu Zhu, Yide Wang, Qingqing Zhu
IEEE Geosci. Remote. Sens. Lett.4
2025 DOA Estimation of Multibeam Frequency Beam Scanning LWAs Based on Sparse Bayesian Learning
abstract
Multi-beam frequency beam scanning leaky-wave antennas (FBS-LWAs) enable a compact system with reduced bandwidth but introduce parasitic interference that impairs direction-of-arrival (DOA) estimation. To solve this issue, this letter proposes a novel DOA estimation method tailored for multi-beam FBS-LWAs. Theoretical analysis reveals that the true DOAs align with multiple peaks in the radiation pattern. Based on this insight, a cropped hypercomplete set is constructed via grid refinement in the frequency-space domain. The sparse Bayesian learning algorithm, enhanced by singular value decomposition (SVD-SBL), is then employed to suppress parasitic peaks and accurately recover the true DOAs. Simulation results demonstrate the proposed method’s effectiveness and robustness.
Qingqing Zhu, Biyun Ma, Yide Wang, Yuehui Cui
IEEE Geosci. Remote. Sens. Lett.1
2025 New Paradigm for Evaluating Scholar Summaries: A Facet-aware Metric and a Meta-evaluation Benchmark
abstract
Evaluation of summary quality is particularly crucial within the scientific domain, because it facilitates efficient knowledge dissemination and automated scientific information retrieval. This article presents conceptual and experimental analyses of scientific summarization, highlighting the inadequacies of traditional evaluation methods. These methods, including \( n \) -gram overlap calculations, embedding comparisons, verification, and QA-based approaches, often fall short in providing explanations, grasping scientific concepts, or identifying key content. Correspondingly, we introduce the Facet-aware Metric (FM), employing LLMs for advanced semantic matching to evaluate summaries based on different facets. The facet granularity is tailored to the structure of scientific abstracts, offering an integrated evaluation approach that is not fragmented, while also providing fine-grained interpretability. Recognizing the absence of an evaluation benchmark in the scientific domain, we curate a Scientific abstract summary evaluation Dataset (ScholarSum) with facet-level annotations. Our findings confirm that FM offers a more logical approach to evaluating scientific summaries. In addition, fine-tuned smaller models can compete with LLMs in scientific contexts, while LLMs have limitations in learning from in-context information in scientific domains. We hope our benchmark inspires better evaluation metrics and future enhancements to LLMs: https://github.com/iriscxy/ScholarSum .
Tairan Wang, Xiuying Chen, Qingqing Zhu, Taicheng Guo, Shen Gao, Zhiyong Lu, Xin Gao 0001, Xiangliang Zhang 0001
ACM Trans. Inf. Syst.3
2024 Flexible and Adaptable Summarization via Expertise Separation
abstract
A proficient summarization model should exhibit both flexibility -- the capacity to handle a range of in-domain summarization tasks, and adaptability -- the competence to acquire new knowledge and adjust to unseen out-of-domain tasks. Unlike large language models (LLMs) that achieve this through parameter scaling, we propose a more parameter-efficient approach in this study. Our motivation rests on the principle that the general summarization ability to capture salient information can be shared across different tasks, while the domain-specific summarization abilities need to be distinct and tailored. Concretely, we propose MoeSumm, a Mixture-of-Expert Summarization architecture, which utilizes a main expert for gaining the general summarization capability and deputy experts that selectively collaborate to meet specific summarization task requirements. We further propose a max-margin loss to stimulate the separation of these abilities. Our model's distinct separation of general and domain-specific summarization abilities grants it with notable flexibility and adaptability, all while maintaining parameter efficiency. MoeSumm achieves flexibility by managing summarization across multiple domains with a single model, utilizing a shared main expert and selected deputy experts. It exhibits adaptability by tailoring deputy experts to cater to out-of-domain few-shot and zero-shot scenarios. Experimental results on 11 datasets show the superiority of our model compared with recent baselines and LLMs. We also provide statistical and visual evidence of the distinct separation of the two abilities in MoeSumm https://github.com/iriscxy/MoE_Summ
Xiuying Chen, Mingzhe Li 0001, Shen Gao, Xin Cheng 0002, Qingqing Zhu, Rui Yan 0001, Xin Gao 0001, Xiangliang Zhang 0001
SIGIR5
2024 DKAF: Diffusion Kolmogorov-Arnold Fourier Hard Sample Mining for CTR
Hailong Luo, Qingqing Zhu, Dazhao Ding
WISE (3)3
2024 Opportunities and challenges for ChatGPT and large language models in biomedicine and health
abstract
ChatGPT has drawn considerable attention from both the general public and domain experts with its remarkable text generation capabilities. This has subsequently led to the emergence of diverse applications in the field of biomedicine and health. In this work, we examine the diverse applications of large language models (LLMs), such as ChatGPT, in biomedicine and health. Specifically we explore the areas of biomedical information retrieval, question answering, medical text summarization, information extraction, and medical education, and investigate whether LLMs possess the transformative power to revolutionize these tasks or whether the distinct complexities of biomedical domain presents unique challenges. Following an extensive literature survey, we find that significant advances have been made in the field of text generation tasks, surpassing the previous state-of-the-art methods. For other applications, the advances have been modest. Overall, LLMs have not yet revolutionized biomedicine, but recent rapid progress indicates that such methods hold great potential to provide valuable means for accelerating discovery and improving health. We also find that the use of LLMs, like ChatGPT, in the fields of biomedicine and health entails various risks and challenges, including fabricated information in its generated responses, as well as legal and privacy concerns associated with sensitive patient data. We believe this survey can provide a comprehensive and timely overview to biomedical researchers and healthcare practitioners on the opportunities and challenges associated with using ChatGPT and other LLMs for transforming biomedicine and health.
Shubo Tian, Qiao Jin 0001, Lana Yeganova, Po-Ting Lai, Qingqing Zhu, Xiuying Chen, Yifan Yang 0006, Qingyu Chen 0001, Won Kim 0003, Donald C. Comeau, Rezarta Islamaj Dogan, Aadit Kapoor, Xin Gao 0001, Zhiyong Lu
Briefings Bioinform.5
2024 Write Summary Step-by-Step: A Pilot Study of Stepwise Summarization
abstract
Nowadays, neural text generation has made tremendous progress in abstractive summarization tasks. However, most of the existing summarization models take in the whole document all at once, which sometimes cannot meet the needs in practice. Practically, social text streams such as news events and tweets keep growing from time to time, and can only be fed to the summarization system step by step. Hence, in this paper, we propose the task ofStepwise Summarization, which aims to generate a new appended summary each time a new document is proposed. The appended summary should not only summarizes the newly added content but is also coherent with the previous summary, to form an up-to-date complete summary. To tackle this challenge, we design an adversarial learning model, namedStepwise Summary Generator(SSG). First, SSG selectively processes the new document under the guidance of the previous summary, obtaining polished document representation. Next, SSG generates the summary considering both the previous summary and the document. Finally, a convolutional-based discriminator is employed to determine whether the newly generated summary is coherent with the previous summary. For the experiment, we extend the traditional two-step update summarization setting to multi-step stepwise setting, and re-propose a large-scale stepwise summarization dataset based on a public story generation dataset. Extensive experiments on this dataset show that SSG achieves state-of-the-art performance in terms of both automatic metrics and human evaluations. Ablation studies demonstrate the effectiveness of each module in our framework. We also discuss the benefits and limitations of recent large language models on this task.
Xiuying Chen, Shen Gao, Mingzhe Li 0001, Qingqing Zhu, Xin Gao 0001, Xiangliang Zhang 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Utilizing Longitudinal Chest X-Rays and Reports to Pre-fill Radiology Reports
Qingqing Zhu, Tejas Sudharshan Mathai, Pritam Mukherjee, Yifan Peng 0002, Ronald M. Summers, Zhiyong Lu
MICCAI (5)1
2023 A scoping review on multimodal deep learning in biomedical images and texts
Zhaoyi Sun, Mingquan Lin, Qingqing Zhu, Qianqian Xie, Fei Wang 0001, Zhiyong Lu, Yifan Peng 0002
J. Biomed. Informatics3
2023 Graph Enhanced Transformer for Aspect Category Detection
Houfeng Wang, Qingqing Zhu
J. Comput. Sci. Technol.3
2022 A Gaussian Mixture Model for Dialogue Generation with Dynamic Parameter Sharing Strategy
abstract
Existing dialog models are trained with data in an encoder-decoder framework with the same parameters, ignoring the multinomial distribution nature in the dataset. In fact, model improvement and development commonly requires fine-grained modeling on individual data subsets. However, collecting a labeled fine-grained dialogue dataset often requires expert-level domain knowledge and therefore is difficult to scale in the real world. As we focus on better modeling multinomial data for dialog generation, we study an approach that combines the unsupervised clustering and generative model together with a GMM (Gaussian Mixture Model) based encoder-decoder framework. Specifically, our model samples from the prior and recognition distributions over the latent variables by a Gaussian mixture network and the latent layer with the capability to form multiple clusters. We also introduce knowledge distillation to guide and improve the clustering results. Finally, we use a dynamic parameter sharing strategy conditioned on different labels to train different decoders. Experimental results on a widely used dialogue dataset verify the effectiveness of the proposed method.
Qingqing Zhu, Zhouxing Tan, Jiaxin Duan, Fengyu Lu
ICASSP1
2022 Densely-connected neural networks for aspect term extraction
Houfeng Wang, Qingqing Zhu
Sci. China Inf. Sci.3
2021 Dynamic Curriculum Learning with Co-training for Medical Dialogue Generation
abstract
The purpose of medical dialogue generation is to provide automatic and accurate responses to help doctors provide diagnosis and treatment advice in an efficient w ay. However, due to the subjectivity and open-ended nature of human conversations, the complexity of training dataset varies greatly. Some methods try to solve this problem by using curriculum learning that organizes training data from easy to hard and optimizes the model with an “easy-to-difficult” scheme, but they ignore the intrinsic nature of conversation dataset that has different categories. To solve this problem, in this article, we learn a Dynamic Curriculum with Co-training (DCC) for generative dialogue systems considering both the quality and category of the data. Under our framework, we first t rain two models (query generation model and reply generation model) from dual view by leveraging end-to-end deep Gaussian Mixture Variational Autoencoders (GMVAE) architectures that combine generative and unsupervised clustering tasks together. Then we promote the clustering performance via an ensemble method by combining the clustering distribution of the two networks. Finally, we let the two models determine the training order for each other in each category according to their loss and category confidence dynamically through single-task and multitask curriculum learning. Evaluation results on the widely used medical dialogue generation dataset indicate that our proposed learning approach makes significant improvements compared to strong baselines. varies greatly.
Qingqing Zhu, Zhouxing Tan, Jiaxin Duan, Pengfei Wu 0003, Dongyan Zhao 0001
BIBM1
2021 Knowledge Distillation with Metric Learning for Medical Dialogue Generation
abstract
In recent years, the research of the medical dialogue system has attracted much attention. Considering that in the dialogue system, queries with similar meanings tend to have similar replies. In the medical field, this phenomenon is even more prevalent. For queries of the same class, their corresponding replies typically have similar meanings and can be classified into the same category. Having observed that, we propose to improve the neural sequence-to-sequence (Seq2Seq) based medical dialogue system by utilizing this internal relationship of category information between queries and replies. In our model, we first cluster similar queries into the same category according to their query vectors obtained from the encoder. Then we put forward the indirect and direct distillation learning approach to transfer the category information and category center distance from the queries to the replies. In the indirect distillation process, we employ metric learning to learn better representations of replies, in which replies of corresponding queries in the same category are closely grouped together, whereas those with different categories are far apart. In the direct distillation, to transfer the inter-class relationship, we minimize the Kullback-Leibler (KL) divergence between the category center distance distribution of queries and replies. A large number of experimental results on medical datasets have proved that our method is superior to the most advanced one.
Qingqing Zhu, Pengfei Wu 0003, Zhouxing Tan, Jiaxin Duan, Dongyan Zhao 0001
BIBM1
2021 Bidirectional Distillation for Multi-Guidance Medical Dialogue Generation
abstract
Although researches on the dialogue system with deep learning methods have achieved good performance, medical dialogue generation confronts particular difficulties against other domains. As requiring highly accurate replies, many different types of external guidance signals are provided in previous studies to control the output and increase faithfulness. However, how these strategies compare and combine to each other is not known. In light of these challenges, we propose a multi-guidance model with bidirectional distillation for medical dialogue generation. Firstly, we fuse different guidance signals (keywords, categories and summaries) with neural sequence-to-sequence (Seq2Seq) model as teacher models. Meanwhile, we consider a simplified model without guidance as the student model. We also propose an attention mechanism to ensemble for the fusion of knowledge from multiple teachers. We further develop a bidirectional distillation module to exchange the knowledge between the teachers and a student from both sides during the training process. Through extensive experiments on medical dataset, we demonstrate the superiority of our proposed approach over state-of-the-art ones.
Qingqing Zhu, Pengfei Wu 0003, Dongyan Zhao 0001
BIBM1