VLDB 2026 Research / reviewers in the wild / expert
Shuyue Stella Li
dblp:312/6501
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 48% Trustworthy machine learning · 18% Question answering and dialogue systems · 11% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 19 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › uncertainty estimation
selective classification |
1.5 | 2 | 2024 | MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning · NeurIPS 2024 Teaching LLMs to Abstain across Languages via Multilingual Feedback · EMNLP 2024 |
Natural language and speech › Language models and text generation › large language model evaluation › domain-specific benchmark
cultural benchmark |
0.9 | 1 | 2025 | CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs' Cultural Knowledge Through Human-AI Red-Teaming · ACL (1) 2025 |
Natural language and speech › Language models and text generation › large language model › knowledge in language models
cultural knowledge |
0.9 | 1 | 2025 | CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs' Cultural Knowledge Through Human-AI Red-Teaming · ACL (1) 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge management › knowledge evaluation
cultural knowledge evaluation |
0.9 | 1 | 2025 | CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs' Cultural Knowledge Through Human-AI Red-Teaming · ACL (1) 2025 |
Natural language and speech › Language models and text generation
faithful generation |
0.9 | 1 | 2025 | Precise Information Control in Long-Form Text Generation · NeurIPS 2025 |
Natural language and speech › Language models and text generation
hallucination mitigation |
0.9 | 1 | 2025 | Precise Information Control in Long-Form Text Generation · NeurIPS 2025 |
Natural language and speech › Language models and text generation › text generation
long-form text generation |
0.9 | 1 | 2025 | Precise Information Control in Long-Form Text Generation · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
robustness evaluation |
0.9 | 1 | 2025 | CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs' Cultural Knowledge Through Human-AI Red-Teaming · ACL (1) 2025 |
Natural language and speech › Question answering and dialogue systems
interactive question answering |
0.8 | 1 | 2024 | MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning · NeurIPS 2024 |
Natural language and speech › Language models and text generation › trustworthy language model
large language model reliability |
0.8 | 1 | 2024 | MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning · NeurIPS 2024 |
Natural language and speech › Question answering and dialogue systems
question asking |
0.8 | 1 | 2024 | MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.7 | 1 | 2023 | Condensing Multilingual Knowledge with Lightweight Language-Specific Modules · EMNLP 2023 |
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation |
0.7 | 1 | 2023 | Condensing Multilingual Knowledge with Lightweight Language-Specific Modules · EMNLP 2023 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.7 | 1 | 2023 | Condensing Multilingual Knowledge with Lightweight Language-Specific Modules · EMNLP 2023 |
Natural language and speech › Language models and text generation › text generation
factual text generation |
0.3 | 1 | 2025 | Precise Information Control in Long-Form Text Generation · NeurIPS 2025 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.3 | 1 | 2025 | Precise Information Control in Long-Form Text Generation · NeurIPS 2025 |
Medical and health informatics › clinical decision support
diagnostic decision support |
0.2 | 1 | 2024 | MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning · NeurIPS 2024 |
Medical and health informatics › clinical decision-making
medical reasoning |
0.2 | 1 | 2024 | MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.2 | 1 | 2023 | Condensing Multilingual Knowledge with Lightweight Language-Specific Modules · EMNLP 2023 |
Methods — techniques the papers use, named apart from their topics
prompting · 1.5confidence estimation · 1.5abstention strategies · 1.5weakly supervised preference learning · 0.9red teaming · 0.9post-training · 0.9multilingual feedback · 0.8mixture of experts · 0.7knowledge distillation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs' Cultural Knowledge Through Human-AI Red-TeamingabstractYu Ying Chiu, Liwei Jiang, Bill Yuchen Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov, Vered Shwartz, Yejin Choi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yu Ying Chiu, Bill Y. Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov, Vered Shwartz, Yejin Choi 0001 |
ACL (1) | 5 |
| 2025 | Precise Information Control in Long-Form Text GenerationabstractA central challenge in language models (LMs) is faithfulness hallucination: the generation of information unsubstantiated by input context. To study this problem, we propose Precise Information Control (PIC), a new task formulation that requires models to generate long-form outputs grounded in a provided set of short self-contained statements, without adding any unsupported ones. PIC includes a full setting that tests a model’s ability to include exactly all input claims, and a partial setting that requires the model to selectively incorporate only relevant claims. We present PIC-Bench, a benchmark of eight long-form generation tasks (e.g., summarization, biography generation) adapted to the PIC setting, where LMs are supplied with well-formed, verifiable input claims. Our evaluation of a range of open and proprietary LMs on PIC-Bench reveals that, surprisingly, state-of-the-art LMs still hallucinate against user-provided input in over 70% of generations. To alleviate this lack of faithfulness, we introduce a post-training framework that uses a weakly supervised preference data construction method to train an 8B PIC-LM with stronger PIC ability—improving from 69.1% to 91.0% F1 in the full PIC setting. When integrated into end-to-end factual generation pipelines, PIC-LM improves exact match recall by 17.1% on ambiguous QA with retrieval, and factual precision by 30.5% on a birthplace fact-checking task, underscoring the potential of precisely grounded generation. Jacqueline He, Howard Yen, Margaret Li, Shuyue Stella Li, Yulia Tsvetkov, Danqi Chen 0001, Pang Wei W. Koh, Luke Zettlemoyer |
NeurIPS | 4 |
| 2024 | Teaching LLMs to Abstain across Languages via Multilingual FeedbackabstractShangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding, Orevaoghene Ahia, Shuyue Stella Li, Vidhisha Balachandran, Sunayana Sitaram, Yulia Tsvetkov. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Shangbin Feng, Yike Wang 0002, Wenxuan Ding 0001, Orevaoghene Ahia, Shuyue Stella Li, Vidhisha Balachandran, Sunayana Sitaram, Yulia Tsvetkov |
EMNLP | 6 |
| 2024 | MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical ReasoningabstractUsers typically engage with LLMs interactively, yet most existing benchmarks evaluate them in a static, single-turn format, posing reliability concerns in interactive scenarios. We identify a key obstacle towards reliability: LLMs are trained to answer any question, even with incomplete context or insufficient knowledge. In this paper, we propose to change the static paradigm to an interactive one, develop systems that proactively ask questions to gather more information and respond reliably, and introduce an benchmark—MEDIQ—to evaluate question-asking ability in LLMs. MEDIQ simulates clinical interactions consisting of a Patient System and an adaptive Expert System; with potentially incomplete initial information, the Expert refrains from making diagnostic decisions when unconfident, and instead elicits missing details via follow-up questions. We provide a pipeline to convert single-turn medical benchmarks into an interactive format. Our results show that directly prompting state-of-the-art LLMs to ask questions degrades performance, indicating that adapting LLMs to proactive information-seeking settings is nontrivial. We experiment with abstention strategies to better estimate model confidence and decide when to ask questions, improving diagnostic accuracy by 22.3%; however, performance still lags compared to an (unrealistic in practice) upper bound with complete information upfront. Further analyses show improved interactive performance with filtering irrelevant contexts and reformatting conversations. Overall, we introduce a novel problem towards LLM reliability, an interactive MEDIQ benchmark and a novel question-asking system, and highlight directions to extend LLMs’ information-seeking abilities in critical domains. Shuyue Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan Ilgen, Emma Pierson, Pang Wei W. Koh, Yulia Tsvetkov |
NeurIPS | 1 |
| 2023 | Condensing Multilingual Knowledge with Lightweight Language-Specific ModulesabstractIncorporating language-specific (LS) modules or Mixture-of-Experts (MoE) are proven methods to boost performance in multilingual model performance, but the scalability of these approaches to hundreds of languages or experts tends to be hard to manage.We present Language-specific Matrix Synthesis (LMS), a novel method that addresses the issue.LMS utilizes parameter-efficient and lightweight modules, reducing the number of parameters while outperforming existing methods, e.g., +1.73 BLEU over Switch Transformer on OPUS-100 multilingual translation.Additionally, we introduce Fuse Distillation (FD) to condense multilingual knowledge from multiple LS modules into a single shared module, improving model inference and storage efficiency.Our approach demonstrates superior scalability and performance compared to state-of-the-art methods. 1 * Equal contribution computational cost may only come from communication among devices (such as ALLToALL) or gate routing. Weiting Tan, Shuyue Stella Li, Yunmo Chen, Benjamin Van Durme, Philipp Koehn, Kenton Murray |
EMNLP | 3 |
| 2023 | PQLM - Multilingual Decentralized Portable Quantum Language ModelabstractWith careful manipulation, malicious agents can reverse engineer private information encoded in pre-trained language models. Security concerns motivate the development of quantum pre-training. In this work, we propose a highly portable quantum language model (PQLM) that can easily transmit information to downstream tasks on classical machines. The framework consists of a cloud PQLM built with random Variational Quantum Classifiers (VQC) and local models for downstream applications. We demonstrate the ad hoc portability of the quantum model by extracting only the word embeddings and effectively applying them to downstream tasks on classical machines. Our PQLM exhibits comparable performance to its classical counterpart on both intrinsic evaluation (loss, perplexity) and extrinsic evaluation (multilingual sentiment analysis accuracy) metrics. We also perform ablation studies on the factors affecting PQLM performance to analyze model stability. Our work establishes a theoretical foundation for a portable quantum pre-trained language model that could be trained on private data and made available for public use with privacy protection guarantees. Shuyue Stella Li, Xiangyu Zhang 0005, Hongchao Shu, Ruixing Liang, Hexin Liu, L. Paola García-Perera |
ICASSP | 1 |
| 2023 | A New Approach to Extract Fetal Electrocardiogram Using Affine Combination of Adaptive FiltersabstractThe detection of abnormal fetal heartbeats during pregnancy is important for monitoring the health conditions of the fetus. While adult ECG has made several advances in modern medicine, noninvasive fetal electrocardiography (FECG) remains a great challenge. In this paper, we introduce a new method based on affine combinations of adaptive filters to extract FECG signals. The affine combination of multiple filters is able to precisely fit the reference signal, and thus obtain more accurate FECGs. We proposed a method to combine the Least Mean Square (LMS) and Recursive Least Squares (RLS) filters. Our approach found that the Combined Recursive Least Squares (CRLS) filter achieves the best performance among all proposed combinations. In addition, we found that CRLS is more advantageous in extracting FECG from abdominal electrocardiograms (AECG) with a small signal-to-noise ratio (SNR). Compared with the state-of-the-art Multiple Sub-Filter Adaptive Noise Canceller (MSF-ANC) method, CRLS shows improved performance. The sensitivity, accuracy and F1 score are improved by 3.58%, 2.39% and 1.36%, respectively. Yu Xuan, Xiangyu Zhang 0005, Shuyue Stella Li, Zihan Shen, L. Paola García-Perera, Roberto Togneri |
ICASSP | 3 |
| 2023 | Learning from Mistakes: Towards Robust Neural Machine Translation for Disfluent L2 SentencesabstractWe study the sentences written by second-language (L2) learners to improve the robustness of current neural machine translation (NMT) models on this type of data. Current large datasets used to train NMT systems are mostly Wikipedia or government documents written by highly competent speakers of that language, especially English. However, given that English is the most common second language, it is crucial that machine translation systems are robust against the large number of sentences written by L2 learners of English. By studying the difficulties faced by humans in their L2 acquisition process, we are able to transfer such insights to machine translation systems to recover from source-side fluency variations. In this work, we create additional training data with artificial errors similar to mistakes made by L2 learners of various fluency levels to improve the quality of the machine translation system. We test our method in zero-shot settings on the JFLEG-es (English-Spanish) dataset. The quality of our machine translation system on disfluent sentences outperforms the baseline by 1.8 BLEU scores. Shuyue Stella Li, Philipp Koehn |
MTSummit (1) | 1 |