EDBT 2026 Demo / reviewers in the wild / expert
Yiwen Shi
dblp:84/3407
· DBLP profile ↗
3ranked-venue papers in the field
2as first author
3since 2021 · last 2025
—ORCID · conflict
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 2 (2 first)Information Retrieval & Web Search · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dynamic Retrieval Strategy for Summarizing Doctor-Patient Dialogues with RAG
Yiwen Shi |
IEEE Big Data | 1 |
| 2024 | Automatic Prompt Generation and Optimization by Leveraging Large Language Models to Enhance Few-Shot Learning in Biomedical TasksabstractRecent advancements in scaling large language models (LLMs) have enhanced various natural language processing (NLP) tasks. However, open-source moderately sized models, such as BERT, are still needed because of the high computational cost and concerns regarding data privacy from the LLMs, especially in the biomedical area. Prompt-based fine-tuning of BERT has demonstrated good performance in a few-shot setting. However, the prompt selection can result in substantial differences in final accuracy. This study introduces a simple yet effective approach that leverages LLMs, such as GPT-4 Turbo, to automatically generate and optimize task-specific prompts for BERT. Our approach includes two steps: automatic prompt generation and optimization. Initially, we design a framework to generate prompts for LLMs to infer a task-specific candidate prompt set. Subsequently, we employ a dialog with a chatbot to optimize the prompt iteratively. We conduct extensive evaluations and analyses on three different types of biomedical benchmarks. Our method demonstrates superior 5-shot learning performance, outperforming manual prompts by a substantial margin in low-resource settings, achieving up to a 7% absolute accuracy improvement. These results highlight that our method is a task-agnostic approach to utilizing LLMs and automatically enhancing performance on relatively small open-source models with limited resources and human effort. Yiwen Shi |
IEEE Big Data | 1 |
| 2023 | Partisan US News Media Representations of Syrian RefugeesabstractWe investigate how representations of Syrian refugees (2011-2021) differ across US partisan news outlets. We analyze 47,388 articles from the online US media about Syrian refugees to detail differences in reporting between left- and right-leaning media. We use various NLP techniques to understand these differences. Our polarization and question answering results indicated that left-leaning media tended to represent refugees as child victims, welcome in the US, and right-leaning media cast refugees as Islamic terrorists. We noted similar results with our sentiment and offensive speech scores over time, which detail possibly unfavorable representations of refugees in right-leaning media. A strength of our work is how the different techniques we have applied validate each other. Based on our results, we provide several recommendations. Stakeholders may utilize our findings to intervene around refugee representations, and design communications campaigns that improve the way society sees refugees and possibly aid refugee outcomes. Marzieh Babaeianjelodar, Yiwen Shi, Kamila Janmohamed, Rupak Sarkar, Ingmar Weber, Thomas Davidson, Munmun De Choudhury, Jonathan Huang, Shweta Yadav 0001, Ashiqur R. KhudaBukhsh, Chris T. Bauch, Preslav Nakov, Orestis Papakyriakopoulos, Koustuv Saha, Kaveh Khoshnood, Navin Kumar 0004 |
ICWSM | 3 |