Shawn Chen

dblp:136/2173 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
2since 2021 · last 2025
0000-0002-6678-3293ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 50% Trustworthy machine learning · 50%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning · NeurIPS 2025
Natural language and speech › Language models and text generation
large language model reasoning
0.912025
ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning · NeurIPS 2025
Machine learning › Trustworthy machine learning
out-of-distribution evaluation
0.912025
ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning · NeurIPS 2025
Machine learning › Trustworthy machine learning
robustness evaluation
0.912025
ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

dynamic data generation · 0.9
YearPublicationVenuePosition
2025 ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning
abstract
Evaluating large language models (LLMs) poses significant challenges, particularly due to issues of data contamination and the leakage of correct answers. To address these challenges, we introduce ThinkBench, a novel evaluation framework designed to robustly evaluate the reasoning capability of LLMs. ThinkBench proposes a dynamic data generation method for constructing out-of-distribution (OOD) datasets and offers an OOD dataset that contains 2,912 samples drawn from reasoning tasks. ThinkBench unifies the evaluation of reasoning models and non-reasoning models. We evaluate 16 LLMs and 4 PRMs under identical experimental conditions and show that most of the LLMs' performance are far from robust and they face a certain level of data leakage. By dynamically generating OOD datasets, ThinkBench effectively provides a reliable evaluation of LLMs and reduces data contamination impact. Our data and codes are available at https://github.com/huangshulin123/ThinkBench.
Shulin Huang, Linyi Yang, Yan Song 0003, Shawn Chen, Leyang Cui, Ziyu Wan, Qingcheng Zeng, Ying Wen 0001, Kun Shao, Weinan Zhang 0001, Jun Wang 0012, Yue Zhang 0004
NeurIPS4
2025 Increasing adherence and collecting symptom-specific biometric signals in remote monitoring of heart failure patients: a randomized controlled trial
abstract
OBJECTIVES: Mobile health (mHealth) regimens can improve health through the continuous monitoring of biometric parameters paired with appropriate interventions. However, adherence to monitoring tends to decay over time. Our randomized controlled trial sought to determine: (1) if a mobile app with gamification and financial incentives significantly increases adherence to mHealth monitoring in a population of heart failure patients; and (2) if activity data correlate with disease-specific symptoms. MATERIALS AND METHODS: We recruited individuals with heart failure into a prospective 180-day monitoring study with 3 arms. All 3 arms included monitoring with a connected weight scale and an activity tracker. The second arm included an additional mobile app with gamification, and the third arm included the mobile app and a financial incentive awarded based on adherence to mobile monitoring. RESULTS: We recruited 111 heart failure patients into the study. We found that the arm including the financial incentive led to significantly higher adherence to activity tracker (95% vs 72.2%, P = .01) and weight (87.5% vs 69.4%, P = .002) monitoring compared to the arm that included the monitoring devices alone. Furthermore, we found a significant correlation between daily steps and daily symptom severity. DISCUSSION AND CONCLUSION: Our findings indicate that mobile apps with added engagement features can be useful tools for improving adherence over time and may thus increase the impact of mHealth-driven interventions. Additionally, activity tracker data can provide passive monitoring of disease burden that may be used to predict future events.
Sukanya Mohapatra, Mirna Issa, Vedrana Ivezic, Rose Doherty, Leonard S. Marks, Esther Lan, Shawn Chen, Keith Rozett, Lauren Cullen, Wren Reynolds, Rose Rocchio, Gregg C. Fonarow, Michael K. Ong, William Speier, Corey W. Arnold
J. Am. Medical Informatics Assoc.7
2013 Research and applications: Imaging informatics for consumer health: towards a radiology patient portal
abstract
OBJECTIVE: With the increased routine use of advanced imaging in clinical diagnosis and treatment, it has become imperative to provide patients with a means to view and understand their imaging studies. We illustrate the feasibility of a patient portal that automatically structures and integrates radiology reports with corresponding imaging studies according to several information orientations tailored for the layperson. METHODS: The imaging patient portal is composed of an image processing module for the creation of a timeline that illustrates the progression of disease, a natural language processing module to extract salient concepts from radiology reports (73% accuracy, F1 score of 0.67), and an interactive user interface navigable by an imaging findings list. The portal was developed as a Java-based web application and is demonstrated for patients with brain cancer. RESULTS AND DISCUSSION: The system was exhibited at an international radiology conference to solicit feedback from a diverse group of healthcare professionals. There was wide support for educating patients about their imaging studies, and an appreciation for the informatics tools used to simplify images and reports for consumer interpretation. Primary concerns included the possibility of patients misunderstanding their results, as well as worries regarding accidental improper disclosure of medical information. CONCLUSIONS: Radiologic imaging composes a significant amount of the evidence used to make diagnostic and treatment decisions, yet there are few tools for explaining this information to patients. The proposed radiology patient portal provides a framework for organizing radiologic results into several information orientations to support patient education.
Corey W. Arnold, Mary McNamara, Suzie El-Saden, Shawn Chen, Ricky K. Taira, Alex Bui
J. Am. Medical Informatics Assoc.4