Yiwen Shi

dblp:84/3407 · DBLP profile ↗
← Back
3ranked-venue papers in the field
2as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2 (2 first)Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2025 Dynamic Retrieval Strategy for Summarizing Doctor-Patient Dialogues with RAG
Yiwen Shi
IEEE Big Data1
2024 Automatic Prompt Generation and Optimization by Leveraging Large Language Models to Enhance Few-Shot Learning in Biomedical Tasks
abstract
Recent advancements in scaling large language models (LLMs) have enhanced various natural language processing (NLP) tasks. However, open-source moderately sized models, such as BERT, are still needed because of the high computational cost and concerns regarding data privacy from the LLMs, especially in the biomedical area. Prompt-based fine-tuning of BERT has demonstrated good performance in a few-shot setting. However, the prompt selection can result in substantial differences in final accuracy. This study introduces a simple yet effective approach that leverages LLMs, such as GPT-4 Turbo, to automatically generate and optimize task-specific prompts for BERT. Our approach includes two steps: automatic prompt generation and optimization. Initially, we design a framework to generate prompts for LLMs to infer a task-specific candidate prompt set. Subsequently, we employ a dialog with a chatbot to optimize the prompt iteratively. We conduct extensive evaluations and analyses on three different types of biomedical benchmarks. Our method demonstrates superior 5-shot learning performance, outperforming manual prompts by a substantial margin in low-resource settings, achieving up to a 7% absolute accuracy improvement. These results highlight that our method is a task-agnostic approach to utilizing LLMs and automatically enhancing performance on relatively small open-source models with limited resources and human effort.
Yiwen Shi
IEEE Big Data1
2023 Partisan US News Media Representations of Syrian Refugees
abstract
We investigate how representations of Syrian refugees (2011-2021) differ across US partisan news outlets. We analyze 47,388 articles from the online US media about Syrian refugees to detail differences in reporting between left- and right-leaning media. We use various NLP techniques to understand these differences. Our polarization and question answering results indicated that left-leaning media tended to represent refugees as child victims, welcome in the US, and right-leaning media cast refugees as Islamic terrorists. We noted similar results with our sentiment and offensive speech scores over time, which detail possibly unfavorable representations of refugees in right-leaning media. A strength of our work is how the different techniques we have applied validate each other. Based on our results, we provide several recommendations. Stakeholders may utilize our findings to intervene around refugee representations, and design communications campaigns that improve the way society sees refugees and possibly aid refugee outcomes.
Marzieh Babaeianjelodar, Yiwen Shi, Kamila Janmohamed, Rupak Sarkar, Ingmar Weber, Thomas Davidson, Munmun De Choudhury, Jonathan Huang, Shweta Yadav 0001, Ashiqur R. KhudaBukhsh, Chris T. Bauch, Preslav Nakov, Orestis Papakyriakopoulos, Koustuv Saha, Kaveh Khoshnood, Navin Kumar 0004
ICWSM3