VLDB 2026 Research / reviewers in the wild / expert
Shan Chen 0004
dblp:23/6979-4
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0001-7999-7410ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 53% Trustworthy machine learning · 44% Information extraction and text analysis · 3% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model |
2.5 | 3 | 2025 | KScope: A Framework for Characterizing the Knowledge Status of Language Models · NeurIPS 2025 Sparse Autoencoder Features for Classifications and Transferability · EMNLP 2025 Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | Sparse Autoencoder Features for Classifications and Transferability · EMNLP 2025 |
Natural language and speech › Language models and text generation › retrieval-augmented generation
knowledge conflict |
0.9 | 1 | 2025 | KScope: A Framework for Characterizing the Knowledge Status of Language Models · NeurIPS 2025 |
Natural language and speech › Language models and text generation › large language model › knowledge in language models
parametric vs contextual knowledge |
0.9 | 1 | 2025 | KScope: A Framework for Characterizing the Knowledge Status of Language Models · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
sparse autoencoder |
0.9 | 1 | 2025 | Sparse Autoencoder Features for Classifications and Transferability · EMNLP 2025 |
Machine learning › Trustworthy machine learning › fairness
demographic bias |
0.8 | 1 | 2024 | Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
fairness |
0.8 | 1 | 2024 | Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias · NeurIPS 2024 |
Natural language and speech › Information extraction and text analysis
abusive language detection |
0.3 | 1 | 2025 | Sparse Autoencoder Features for Classifications and Transferability · EMNLP 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.3 | 1 | 2025 | Sparse Autoencoder Features for Classifications and Transferability · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
benchmark framework · 1.5alignment method · 1.5sparse autoencoder · 0.9hierarchical statistical testing · 0.9feature extraction · 0.9context summarization · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Sparse Autoencoder Features for Classifications and TransferabilityabstractSparse Autoencoders (SAEs) provide potentials for uncovering structured, human-interpretable representations in Large Language Models (LLMs), making them a crucial tool for transparent and controllable AI systems.We systematically analyze SAE for interpretable feature extraction from LLMs in safety-critical classification tasks 1 .Our framework evaluates (1) model-layer selection and scaling properties, (2) SAE architectural configurations, including width and pooling strategies, and (3) the effect of binarizing continuous SAE activations.SAE-derived features achieve macro F1 > 0.8, outperforming hidden-state and BoW baselines while demonstrating cross-model transfer from Gemma 2 2B to 9B-IT models.These features generalize in a zero-shot manner to cross-lingual toxicity detection and visual classification tasks.Our analysis highlights the significant impact of pooling strategies and binarization thresholds, showing that binarization offers an efficient alternative to traditional feature selection while maintaining or improving performance.These findings establish new best practices for SAE-based interpretability and enable scalable, transparent deployment of LLMs in real-world applications. Jack Gallifant, Shan Chen 0004, Kuleen Sasse, Hugo J. W. L. Aerts, Thomas Hartvigsen, Danielle S. Bitterman |
EMNLP | 2 |
| 2025 | KScope: A Framework for Characterizing the Knowledge Status of Language ModelsabstractCharacterizing a large language model's (LLM's) knowledge of a given question is challenging.
As a result, prior work has primarily examined LLM behavior under knowledge conflicts, where the model's internal parametric memory contradicts information in the external context.
However, this does not fully reflect how well the model knows the answer to the question.
In this paper, we first introduce a taxonomy of five knowledge statuses based on the consistency and correctness of LLM knowledge modes.
We then propose KScope, a hierarchical framework of statistical tests that progressively refines hypotheses about knowledge modes and characterizes LLM knowledge into one of these five statuses.
We apply KScope to nine LLMs across four datasets and systematically establish:
(1) Supporting context narrows knowledge gaps across models.
(2) Context features related to difficulty, relevance, and familiarity drive successful knowledge updates.
(3) LLMs exhibit similar feature preferences when partially correct or conflicted, but diverge sharply when consistently wrong.
(4) Context summarization constrained by our feature analysis, together with enhanced credibility, further improves update effectiveness and generalizes across LLMs. Yuxin Xiao, Shan Chen 0004, Jack Gallifant, Danielle S. Bitterman, Thomas Hartvigsen, Marzyeh Ghassemi |
NeurIPS | 2 |
| 2025 | LCD benchmark: long clinical document benchmark on mortality prediction for language modelsabstractOBJECTIVES: The application of natural language processing (NLP) in the clinical domain is important due to the rich unstructured information in clinical documents, which often remains inaccessible in structured data. When applying NLP methods to a certain domain, the role of benchmark datasets is crucial as benchmark datasets not only guide the selection of best-performing models but also enable the assessment of the reliability of the generated outputs. Despite the recent availability of language models capable of longer context, benchmark datasets targeting long clinical document classification tasks are absent. MATERIALS AND METHODS: To address this issue, we propose Long Clinical Document (LCD) benchmark, a benchmark for the task of predicting 30-day out-of-hospital mortality using discharge notes of Medical Information Mart for Intensive Care IV and statewide death data. We evaluated this benchmark dataset using baseline models, from bag-of-words and convolutional neural network to instruction-tuned large language models. Additionally, we provide a comprehensive analysis of the model outputs, including manual review and visualization of model weights, to offer insights into their predictive capabilities and limitations. RESULTS: Baseline models showed 28.9% for best-performing supervised models and 32.2% for GPT-4 in F1 metrics. Notes in our dataset have a median word count of 1687. DISCUSSION: Our analysis of the model outputs showed that our dataset is challenging for both models and human experts, but the models can find meaningful signals from the text. CONCLUSION: We expect our LCD benchmark to be a resource for the development of advanced supervised models, or prompting methods, tailored for clinical text. Wonjin Yoon, Shan Chen 0004, Yanjun Gao, Zhanzhan Zhao, Dmitriy Dligach, Danielle S. Bitterman, Majid Afshar, Timothy A. Miller |
J. Am. Medical Informatics Assoc. | 2 |
| 2024 | Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model BiasabstractLarge language models (LLMs) are increasingly essential in processing natural languages, yet their application is frequently compromised by biases and inaccuracies originating in their training data.In this study, we introduce \textbf{Cross-Care}, the first benchmark framework dedicated to assessing biases and real world knowledge in LLMs, specifically focusing on the representation of disease prevalence across diverse demographic groups.We systematically evaluate how demographic biases embedded in pre-training corpora like $ThePile$ influence the outputs of LLMs.We expose and quantify discrepancies by juxtaposing these biases against actual disease prevalences in various U.S. demographic groups.Our results highlight substantial misalignment between LLM representation of disease prevalence and real disease prevalence rates across demographic subgroups, indicating a pronounced risk of bias propagation and a lack of real-world grounding for medical applications of LLMs.Furthermore, we observe that various alignment methods minimally resolve inconsistencies in the models' representation of disease prevalence across different languages.For further exploration and analysis, we make all data and a data visualization tool available at: \url{www.crosscare.net}. Shan Chen 0004, Jack Gallifant, Mingye Gao, Nikolaj Munch, Ajay Muthukkumar, Arvind Rajan, Jaya Kolluri, Amelia Fiske, Janna Hastings, Hugo J. W. L. Aerts, Brian Anthony 0001, Leo A. Celi, William G. La Cava, Danielle S. Bitterman |
NeurIPS | 1 |
| 2024 | Evaluating the ChatGPT family of models for biomedical reasoning and classificationabstractOBJECTIVE: Large language models (LLMs) have shown impressive ability in biomedical question-answering, but have not been adequately investigated for more specific biomedical applications. This study investigates ChatGPT family of models (GPT-3.5, GPT-4) in biomedical tasks beyond question-answering. MATERIALS AND METHODS: We evaluated model performance with 11 122 samples for two fundamental tasks in the biomedical domain-classification (n = 8676) and reasoning (n = 2446). The first task involves classifying health advice in scientific literature, while the second task is detecting causal relations in biomedical literature. We used 20% of the dataset for prompt development, including zero- and few-shot settings with and without chain-of-thought (CoT). We then evaluated the best prompts from each setting on the remaining dataset, comparing them to models using simple features (BoW with logistic regression) and fine-tuned BioBERT models. RESULTS: Fine-tuning BioBERT produced the best classification (F1: 0.800-0.902) and reasoning (F1: 0.851) results. Among LLM approaches, few-shot CoT achieved the best classification (F1: 0.671-0.770) and reasoning (F1: 0.682) results, comparable to the BoW model (F1: 0.602-0.753 and 0.675 for classification and reasoning, respectively). It took 78 h to obtain the best LLM results, compared to 0.078 and 0.008 h for the top-performing BioBERT and BoW models, respectively. DISCUSSION: The simple BoW model performed similarly to the most complex LLM prompting. Prompt engineering required significant investment. CONCLUSION: Despite the excitement around viral ChatGPT, fine-tuning for two fundamental biomedical natural language processing tasks remained the best strategy. Shan Chen 0004, Yingya Li, Hoang Van, Hugo J. W. L. Aerts, Guergana K. Savova, Danielle S. Bitterman |
J. Am. Medical Informatics Assoc. | 1 |