Yiwen Shi

dblp:84/3407 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Advances in Deep Learning Applications for Choroidal Images Analysis: A Narrative Review
abstract
ABSTRACT The choroid, a critical vascular structure in the eye, is closely associated with a wide range of ocular diseases, and its early and accurate evaluation is fundamental to effective diagnosis and management. Advances in imaging modalities like optical coherence tomography (OCT) have revolutionised choroidal visualisation, yet manual analysis remains labour‐intensive and subjective. This review synthesizes recent developments in deep learning (DL)‐driven solutions for choroidal imaging. DL architectures, including convolutional neural networks (CNNs), U‐Net variants and vision transformers, demonstrate high accuracy in automating choroidal sublayer segmentation, 3D vascular network reconstruction and biomarker extraction, outperforming traditional methods in speed and reproducibility. Innovations such as hybrid loss functions, attention mechanisms, and generative adversarial networks address challenges like class imbalance and boundary ambiguity. At the same time, multimodal DL models integrating OCT, fundus photography and angiography enhance diagnostic precision. However, limitations persist, including data dependency, annotation variability and the lack of standardised protocols for sublayer demarcation. Future directions emphasise image enhancement techniques, self‐supervised learning to reduce annotation burdens and ethical frameworks for clinical translation.
Yiwen Shi, Zixuan Jiang, Houjin Chen, Xuemin Li
IET Image Process.1
2025 Dynamic Retrieval Strategy for Summarizing Doctor-Patient Dialogues with RAG
Yiwen Shi
IEEE Big Data1
2024 Automatic Prompt Generation and Optimization by Leveraging Large Language Models to Enhance Few-Shot Learning in Biomedical Tasks
abstract
Recent advancements in scaling large language models (LLMs) have enhanced various natural language processing (NLP) tasks. However, open-source moderately sized models, such as BERT, are still needed because of the high computational cost and concerns regarding data privacy from the LLMs, especially in the biomedical area. Prompt-based fine-tuning of BERT has demonstrated good performance in a few-shot setting. However, the prompt selection can result in substantial differences in final accuracy. This study introduces a simple yet effective approach that leverages LLMs, such as GPT-4 Turbo, to automatically generate and optimize task-specific prompts for BERT. Our approach includes two steps: automatic prompt generation and optimization. Initially, we design a framework to generate prompts for LLMs to infer a task-specific candidate prompt set. Subsequently, we employ a dialog with a chatbot to optimize the prompt iteratively. We conduct extensive evaluations and analyses on three different types of biomedical benchmarks. Our method demonstrates superior 5-shot learning performance, outperforming manual prompts by a substantial margin in low-resource settings, achieving up to a 7% absolute accuracy improvement. These results highlight that our method is a task-agnostic approach to utilizing LLMs and automatically enhancing performance on relatively small open-source models with limited resources and human effort.
Yiwen Shi
IEEE Big Data1
2024 Two-stage fine-tuning with ChatGPT data augmentation for learning class-imbalanced data
abstract
Classification of long-tailed distributed data is a challenging problem, which suffers from serious class imbalance and hence poor performance on tail classes, which have only a few samples. Owing to this paucity of samples, learning on the tail classes is especially challenging for fine-tuning when transferring a pretrained model to a downstream task. In this work, we present a simple modification of standard fine-tuning to cope with these challenges. Specifically, we propose a two-stage fine-tuning. In Stage 1, we fine-tune the final layer of the pretrained model with class-balanced augmented data, generated using ChatGPT. As a large generative language model, ChatGPT is capable of generating novel and contextually similar responses to a given prompt, which makes it an excellent candidate for data augmentation. In Stage 2, we perform the standard fine-tuning. Our modification has several benefits: (1) it leverages pretrained representations by only fine-tuning a small portion of the model parameters while keeping the rest untouched; (2) it allows the model to learn an initial representation of the specific task; and importantly (3) it protects the learning of tail classes from being at a disadvantage during the model updating. We conduct extensive experiments on synthetic datasets of both two-class and multi-class tasks of text classification as well as a real-world application to ADME (i.e., absorption, distribution, metabolism, and excretion) semantic drug labeling. The experimental results show that the proposed two-stage fine-tuning outperforms vanilla fine-tuning and state-of-the-art methods on the above datasets.
Taha ValizadehAslani, Yiwen Shi, Hualou Liang
Neurocomputing2
2023 Partisan US News Media Representations of Syrian Refugees
abstract
We investigate how representations of Syrian refugees (2011-2021) differ across US partisan news outlets. We analyze 47,388 articles from the online US media about Syrian refugees to detail differences in reporting between left- and right-leaning media. We use various NLP techniques to understand these differences. Our polarization and question answering results indicated that left-leaning media tended to represent refugees as child victims, welcome in the US, and right-leaning media cast refugees as Islamic terrorists. We noted similar results with our sentiment and offensive speech scores over time, which detail possibly unfavorable representations of refugees in right-leaning media. A strength of our work is how the different techniques we have applied validate each other. Based on our results, we provide several recommendations. Stakeholders may utilize our findings to intervene around refugee representations, and design communications campaigns that improve the way society sees refugees and possibly aid refugee outcomes.
Marzieh Babaeianjelodar, Yiwen Shi, Kamila Janmohamed, Rupak Sarkar, Ingmar Weber, Thomas Davidson, Munmun De Choudhury, Jonathan Huang, Shweta Yadav 0001, Ashiqur R. KhudaBukhsh, Chris T. Bauch, Preslav Nakov, Orestis Papakyriakopoulos, Koustuv Saha, Kaveh Khoshnood, Navin Kumar 0004
ICWSM3
2023 PharmBERT: a domain-specific BERT model for drug labels
abstract
Human prescription drug labeling contains a summary of the essential scientific information needed for the safe and effective use of the drug and includes the Prescribing Information, FDA-approved patient labeling (Medication Guides, Patient Package Inserts and/or Instructions for Use), and/or carton and container labeling. Drug labeling contains critical information about drug products, such as pharmacokinetics and adverse events. Automatic information extraction from drug labels may facilitate finding the adverse reaction of the drugs or finding the interaction of one drug with another drug. Natural language processing (NLP) techniques, especially recently developed Bidirectional Encoder Representations from Transformers (BERT), have exhibited exceptional merits in text-based information extraction. A common paradigm in training BERT is to pretrain the model on large unlabeled generic language corpora, so that the model learns the distribution of the words in the language, and then fine-tune on a downstream task. In this paper, first, we show the uniqueness of language used in drug labels, which therefore cannot be optimally handled by other BERT models. Then, we present the developed PharmBERT, which is a BERT model specifically pretrained on the drug labels (publicly available at Hugging Face). We demonstrate that our model outperforms the vanilla BERT, ClinicalBERT and BioBERT in multiple NLP tasks in the drug label domain. Moreover, how the domain-specific pretraining has contributed to the superior performance of PharmBERT is demonstrated by analyzing different layers of PharmBERT, and more insight into how it understands different linguistic aspects of the data is gained.
Taha ValizadehAslani, Yiwen Shi, Hualou Liang
Briefings Bioinform.2
2023 Leveraging GPT-4 for food effect summarization to enhance product-specific guidance development via iterative prompting
Yiwen Shi, Taha ValizadehAslani, Felix Agbavor, Hualou Liang
J. Biomed. Informatics1
2023 Fine-tuning BERT for automatic ADME semantic labeling in FDA drug labeling to enhance product-specific guidance assessment
Yiwen Shi, Taha ValizadehAslani, Hualou Liang
J. Biomed. Informatics1
2021 Dual Stream Fusion Network for Multi-spectral High Resolution Remote Sensing Image Segmentation
Yiwen Shi, Chunlei Huo, Shiming Xiang, Chunhong Pan
PRCV (2)2
2021 Relational Attention with Textual Enhanced Transformer for Image Captioning
Lifei Song, Yiwen Shi, Xinyu Xiao, Chunxia Zhang 0001, Shiming Xiang
PRCV (3)2
2013 A Simulated Annealing Inspired Test Optimization Method for Enhanced Detection of Highly Critical Faults and Defects
Yiwen Shi, Jennifer Dworak
J. Electron. Test.1
2012 Using implications to choose tests through suspect fault identification
abstract
As circuits continue to scale to smaller feature sizes, wearout and latent defects are expected to cause an increasing number of errors in the field. Online error detection techniques, including logic implication-based checker hardware, are capable of detecting at least some of these errors as they occur. However, recovery may be expensive, and the underlying problem may lead to multiple failures of a core over time. In this article, we will investigate the diagnostic capability of logic implications to identify possible failure locations when an error is detected online. We will then utilize this information to select a highly efficient test set that can be used to effectively test the identified suspect locations in both the failing core and in other identical cores in the system.
Jennifer Dworak, Kundan Nepal, Nuno Alves, Yiwen Shi, Nicholas Imbriglia, R. Iris Bahar
ACM Trans. Design Autom. Electr. Syst.4
2011 Dynamic Test Set Selection Using Implication-Based On-Chip Diagnosis
abstract
We propose using logic implications as a source of online diagnostic data for on-chip test set selection by taking advantage of their ability to automatically identify a restricted set of faults as the potential cause of an observed error. This information will be used to dynamically choose a test set to detect systematic latent defects or wear out in a multi core system.
Nuno Alves, Yiwen Shi, Nicholas Imbriglia, Jennifer Dworak, Kundan Nepal, R. Iris Bahar
ETS2
2011 Partial state monitoring for fault detection estimation
abstract
Obtaining fault coverage information for functional input sequences is often very difficult. Although many simulation-based techniques have been proposed, they are generally computationally expensive, and if the input sequence changes, new expensive simulations must be run. In this paper, we propose a new type of hardware monitor for the probabilistic determination of how many times a fault was likely to have been covered during functional test or program execution. In addition to providing coverage information for functional test sequences - even those that have never been simulated - these monitors can also be used to determine the relative criticality of faults for the applications a user is running in real time. Thus, in the future, this method has the potential to provide new dynamic optimization capabilities for on-chip field testing.
Yiwen Shi, Kantapon Kaewtip, Wan-Chan Hu, Jennifer Dworak
ITC1
2011 Enhancing online error detection through area-efficient multi-site implications
abstract
We present a new method to identify multi-site implications that can significantly increase the fault coverage of error-detecting hardware without increasing the area overhead. This method intelligently divides the input space about the functions of internal circuit sites and finds new valuable implications that can share gates in checker logic.
Nuno Alves, Yiwen Shi, Jennifer Dworak, R. Iris Bahar, Kundan Nepal
VTS2
2010 Too many faults, too little time on creating test sets for enhanced detection of highly critical faults and defects
abstract
When testing resources are severely limited, special attention must be paid to critical faults so that important or frequent field failures arising from test escapes can be minimized. We present a new algorithm to optimize test sets that considers the criticality of potential undetected defects throughout the testing process and dramatically reduces the criticality of test escapes.
Yiwen Shi, Wan-Chan Hu, Jennifer Dworak
VTS1