Shuyue Stella Li

dblp:312/6501 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 48% Trustworthy machine learning · 18% Question answering and dialogue systems · 11%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › uncertainty estimation
selective classification
1.522024
MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning · NeurIPS 2024
Teaching LLMs to Abstain across Languages via Multilingual Feedback · EMNLP 2024
Natural language and speech › Language models and text generation › large language model evaluation › domain-specific benchmark
cultural benchmark
0.912025
CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs' Cultural Knowledge Through Human-AI Red-Teaming · ACL (1) 2025
Natural language and speech › Language models and text generation › large language model › knowledge in language models
cultural knowledge
0.912025
CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs' Cultural Knowledge Through Human-AI Red-Teaming · ACL (1) 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge management › knowledge evaluation
cultural knowledge evaluation
0.912025
CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs' Cultural Knowledge Through Human-AI Red-Teaming · ACL (1) 2025
Natural language and speech › Language models and text generation
faithful generation
0.912025
Precise Information Control in Long-Form Text Generation · NeurIPS 2025
Natural language and speech › Language models and text generation
hallucination mitigation
0.912025
Precise Information Control in Long-Form Text Generation · NeurIPS 2025
Natural language and speech › Language models and text generation › text generation
long-form text generation
0.912025
Precise Information Control in Long-Form Text Generation · NeurIPS 2025
Machine learning › Trustworthy machine learning
robustness evaluation
0.912025
CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs' Cultural Knowledge Through Human-AI Red-Teaming · ACL (1) 2025
Natural language and speech › Question answering and dialogue systems
interactive question answering
0.812024
MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning · NeurIPS 2024
Natural language and speech › Language models and text generation › trustworthy language model
large language model reliability
0.812024
MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning · NeurIPS 2024
Natural language and speech › Question answering and dialogue systems
question asking
0.812024
MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning · NeurIPS 2024
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.712023
Condensing Multilingual Knowledge with Lightweight Language-Specific Modules · EMNLP 2023
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation
0.712023
Condensing Multilingual Knowledge with Lightweight Language-Specific Modules · EMNLP 2023
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.712023
Condensing Multilingual Knowledge with Lightweight Language-Specific Modules · EMNLP 2023
Natural language and speech › Language models and text generation › text generation
factual text generation
0.312025
Precise Information Control in Long-Form Text Generation · NeurIPS 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.312025
Precise Information Control in Long-Form Text Generation · NeurIPS 2025
Medical and health informatics › clinical decision support
diagnostic decision support
0.212024
MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning · NeurIPS 2024
Medical and health informatics › clinical decision-making
medical reasoning
0.212024
MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning · NeurIPS 2024
Machine learning › Deep learning architectures and training
mixture of experts
0.212023
Condensing Multilingual Knowledge with Lightweight Language-Specific Modules · EMNLP 2023

Methods — techniques the papers use, named apart from their topics

prompting · 1.5confidence estimation · 1.5abstention strategies · 1.5weakly supervised preference learning · 0.9red teaming · 0.9post-training · 0.9multilingual feedback · 0.8mixture of experts · 0.7knowledge distillation · 0.7
YearPublicationVenuePosition
2025 CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs' Cultural Knowledge Through Human-AI Red-Teaming
abstract
Yu Ying Chiu, Liwei Jiang, Bill Yuchen Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov, Vered Shwartz, Yejin Choi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yu Ying Chiu, Bill Y. Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov, Vered Shwartz, Yejin Choi 0001
ACL (1)5
2025 Precise Information Control in Long-Form Text Generation
abstract
A central challenge in language models (LMs) is faithfulness hallucination: the generation of information unsubstantiated by input context. To study this problem, we propose Precise Information Control (PIC), a new task formulation that requires models to generate long-form outputs grounded in a provided set of short self-contained statements, without adding any unsupported ones. PIC includes a full setting that tests a model’s ability to include exactly all input claims, and a partial setting that requires the model to selectively incorporate only relevant claims. We present PIC-Bench, a benchmark of eight long-form generation tasks (e.g., summarization, biography generation) adapted to the PIC setting, where LMs are supplied with well-formed, verifiable input claims. Our evaluation of a range of open and proprietary LMs on PIC-Bench reveals that, surprisingly, state-of-the-art LMs still hallucinate against user-provided input in over 70% of generations. To alleviate this lack of faithfulness, we introduce a post-training framework that uses a weakly supervised preference data construction method to train an 8B PIC-LM with stronger PIC ability—improving from 69.1% to 91.0% F1 in the full PIC setting. When integrated into end-to-end factual generation pipelines, PIC-LM improves exact match recall by 17.1% on ambiguous QA with retrieval, and factual precision by 30.5% on a birthplace fact-checking task, underscoring the potential of precisely grounded generation.
Jacqueline He, Howard Yen, Margaret Li, Shuyue Stella Li, Yulia Tsvetkov, Danqi Chen 0001, Pang Wei W. Koh, Luke Zettlemoyer
NeurIPS4
2024 Teaching LLMs to Abstain across Languages via Multilingual Feedback
abstract
Shangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding, Orevaoghene Ahia, Shuyue Stella Li, Vidhisha Balachandran, Sunayana Sitaram, Yulia Tsvetkov. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Shangbin Feng, Yike Wang 0002, Wenxuan Ding 0001, Orevaoghene Ahia, Shuyue Stella Li, Vidhisha Balachandran, Sunayana Sitaram, Yulia Tsvetkov
EMNLP6
2024 MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning
abstract
Users typically engage with LLMs interactively, yet most existing benchmarks evaluate them in a static, single-turn format, posing reliability concerns in interactive scenarios. We identify a key obstacle towards reliability: LLMs are trained to answer any question, even with incomplete context or insufficient knowledge. In this paper, we propose to change the static paradigm to an interactive one, develop systems that proactively ask questions to gather more information and respond reliably, and introduce an benchmark—MEDIQ—to evaluate question-asking ability in LLMs. MEDIQ simulates clinical interactions consisting of a Patient System and an adaptive Expert System; with potentially incomplete initial information, the Expert refrains from making diagnostic decisions when unconfident, and instead elicits missing details via follow-up questions. We provide a pipeline to convert single-turn medical benchmarks into an interactive format. Our results show that directly prompting state-of-the-art LLMs to ask questions degrades performance, indicating that adapting LLMs to proactive information-seeking settings is nontrivial. We experiment with abstention strategies to better estimate model confidence and decide when to ask questions, improving diagnostic accuracy by 22.3%; however, performance still lags compared to an (unrealistic in practice) upper bound with complete information upfront. Further analyses show improved interactive performance with filtering irrelevant contexts and reformatting conversations. Overall, we introduce a novel problem towards LLM reliability, an interactive MEDIQ benchmark and a novel question-asking system, and highlight directions to extend LLMs’ information-seeking abilities in critical domains.
Shuyue Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan Ilgen, Emma Pierson, Pang Wei W. Koh, Yulia Tsvetkov
NeurIPS1
2023 Condensing Multilingual Knowledge with Lightweight Language-Specific Modules
abstract
Incorporating language-specific (LS) modules or Mixture-of-Experts (MoE) are proven methods to boost performance in multilingual model performance, but the scalability of these approaches to hundreds of languages or experts tends to be hard to manage.We present Language-specific Matrix Synthesis (LMS), a novel method that addresses the issue.LMS utilizes parameter-efficient and lightweight modules, reducing the number of parameters while outperforming existing methods, e.g., +1.73 BLEU over Switch Transformer on OPUS-100 multilingual translation.Additionally, we introduce Fuse Distillation (FD) to condense multilingual knowledge from multiple LS modules into a single shared module, improving model inference and storage efficiency.Our approach demonstrates superior scalability and performance compared to state-of-the-art methods. 1 * Equal contribution computational cost may only come from communication among devices (such as ALLToALL) or gate routing.
Weiting Tan, Shuyue Stella Li, Yunmo Chen, Benjamin Van Durme, Philipp Koehn, Kenton Murray
EMNLP3
2023 PQLM - Multilingual Decentralized Portable Quantum Language Model
abstract
With careful manipulation, malicious agents can reverse engineer private information encoded in pre-trained language models. Security concerns motivate the development of quantum pre-training. In this work, we propose a highly portable quantum language model (PQLM) that can easily transmit information to downstream tasks on classical machines. The framework consists of a cloud PQLM built with random Variational Quantum Classifiers (VQC) and local models for downstream applications. We demonstrate the ad hoc portability of the quantum model by extracting only the word embeddings and effectively applying them to downstream tasks on classical machines. Our PQLM exhibits comparable performance to its classical counterpart on both intrinsic evaluation (loss, perplexity) and extrinsic evaluation (multilingual sentiment analysis accuracy) metrics. We also perform ablation studies on the factors affecting PQLM performance to analyze model stability. Our work establishes a theoretical foundation for a portable quantum pre-trained language model that could be trained on private data and made available for public use with privacy protection guarantees.
Shuyue Stella Li, Xiangyu Zhang 0005, Hongchao Shu, Ruixing Liang, Hexin Liu, L. Paola García-Perera
ICASSP1
2023 A New Approach to Extract Fetal Electrocardiogram Using Affine Combination of Adaptive Filters
abstract
The detection of abnormal fetal heartbeats during pregnancy is important for monitoring the health conditions of the fetus. While adult ECG has made several advances in modern medicine, noninvasive fetal electrocardiography (FECG) remains a great challenge. In this paper, we introduce a new method based on affine combinations of adaptive filters to extract FECG signals. The affine combination of multiple filters is able to precisely fit the reference signal, and thus obtain more accurate FECGs. We proposed a method to combine the Least Mean Square (LMS) and Recursive Least Squares (RLS) filters. Our approach found that the Combined Recursive Least Squares (CRLS) filter achieves the best performance among all proposed combinations. In addition, we found that CRLS is more advantageous in extracting FECG from abdominal electrocardiograms (AECG) with a small signal-to-noise ratio (SNR). Compared with the state-of-the-art Multiple Sub-Filter Adaptive Noise Canceller (MSF-ANC) method, CRLS shows improved performance. The sensitivity, accuracy and F1 score are improved by 3.58%, 2.39% and 1.36%, respectively.
Yu Xuan, Xiangyu Zhang 0005, Shuyue Stella Li, Zihan Shen, L. Paola García-Perera, Roberto Togneri
ICASSP3
2023 Learning from Mistakes: Towards Robust Neural Machine Translation for Disfluent L2 Sentences
abstract
We study the sentences written by second-language (L2) learners to improve the robustness of current neural machine translation (NMT) models on this type of data. Current large datasets used to train NMT systems are mostly Wikipedia or government documents written by highly competent speakers of that language, especially English. However, given that English is the most common second language, it is crucial that machine translation systems are robust against the large number of sentences written by L2 learners of English. By studying the difficulties faced by humans in their L2 acquisition process, we are able to transfer such insights to machine translation systems to recover from source-side fluency variations. In this work, we create additional training data with artificial errors similar to mistakes made by L2 learners of various fluency levels to improve the quality of the machine translation system. We test our method in zero-shot settings on the JFLEG-es (English-Spanish) dataset. The quality of our machine translation system on disfluent sentences outperforms the baseline by 1.8 BLEU scores.
Shuyue Stella Li, Philipp Koehn
MTSummit (1)1