Scarlett Li

dblp:382/3945 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2026
0009-0002-8912-4861ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 60% Vision and language · 17% Question answering and dialogue systems · 17%
Software engineering, system software, and programming languages
3 papers
Program synthesis and code generation · 46% Software maintenance and evolution · 23% Compilers and program optimization · 23%
Human-computer interaction and pervasive computing
1 paper
Design research and methods · 77% Human-AI interaction · 23%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program synthesis and code generation › code generation with language models
repository-level code generation
1.722025
EpiCoder: Encompassing Diversity and Complexity in Code Generation · ICML 2025
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation · ACL (1) 2025
Natural language and speech › Language models and text generation
large language model training
1.012026
Demystifying Data Organization for Enhanced LLM Training · ACL (1) 2026
Natural language and speech › Language models and text generation › large language model evaluation
benchmark contamination
0.912025
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark · ACL (1) 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark · ACL (1) 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
PEACE: Empowering Geologic Map Holistic Understanding with MLLMs · CVPR 2025
Natural language and speech › Question answering and dialogue systems
multimodal question answering
0.912025
PEACE: Empowering Geologic Map Holistic Understanding with MLLMs · CVPR 2025
Design research and methods › design process
design ideation
0.912025
ProductMeta: An Interactive System for Metaphorical Product Design Ideation with Multimodal Large Language Models · CHI 2025
Software maintenance and evolution › program comprehension
code comprehension
0.912025
Teaching Your Models to Understand Code via Focal Preference Alignment · EMNLP 2025
Compilers and program optimization
code generation
0.912025
EpiCoder: Encompassing Diversity and Complexity in Code Generation · ICML 2025
Natural language and speech › Language models and text generation
code language models
0.312025
EpiCoder: Encompassing Diversity and Complexity in Code Generation · ICML 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge engineering › knowledge integration
domain knowledge integration
0.312025
PEACE: Empowering Geologic Map Holistic Understanding with MLLMs · CVPR 2025
Human-AI interaction › AI-assisted creativity
AI-assisted design ideation
0.312025
ProductMeta: An Interactive System for Metaphorical Product Design Ideation with Multimodal Large Language Models · CHI 2025

Methods — techniques the papers use, named apart from their topics

fine-tuning · 1.7feature tree synthesis · 1.7validation-test split · 0.9prompt-enhanced question answering · 0.9multimodal large language model · 0.9large language model · 0.9hierarchical information extraction · 0.9focal preference alignment · 0.9decontamination rules · 0.9AI expert group · 0.9
YearPublicationVenuePosition
2026 Demystifying Data Organization for Enhanced LLM Training
abstract
Yalun Dai, Yangyu Huang, Tongshen Yang, Yonghan Wang, Xin Zhang, Wenshan Wu, Qihao Zhao, Hao Li, Yuanyuan Gao, Kim-Hui Yap, Scarlett Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yalun Dai, Yangyu Huang, Tongshen Yang, Yonghan Wang, Wenshan Wu, Qihao Zhao, Hao Li 0069, Kim-Hui Yap, Scarlett Li
ACL (1)11
2025 FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
abstract
Wei Li, Xin Zhang, Zhongxin Guo, Shaoguang Mao, Wen Luo, Guangyue Peng, Yangyu Huang, Houfeng Wang, Scarlett Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Wei Li 0232, Xin Zhang 0099, Zhongxin Guo, Shaoguang Mao, Wen Luo 0001, Guangyue Peng, Yangyu Huang, Houfeng Wang, Scarlett Li
ACL (1)9
2025 MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
abstract
Multiple-choice question (MCQ) datasets like Massive Multitask Language Understanding (MMLU) are widely used to evaluate the commonsense, understanding, and problem-solving abilities of large language models (LLMs). However, the open-source nature of these benchmarks and the broad sources of training data for LLMs have inevitably led to benchmark contamination, resulting in unreliable evaluation. To alleviate this issue, we propose the contamination-free MCQ benchmark called MMLU-CF, which reassesses LLMs’ understanding of world knowledge by averting both unintentional and malicious data contamination. To mitigate unintentional data contamination, we source questions from a broader domain of over 200 billion webpages and apply three specifically designed decontamination rules. To prevent malicious data contamination, we divide the benchmark into validation and test sets with similar difficulty and subject distributions. The test set remains closed-source to ensure reliable results, while the validation set is publicly available to promote transparency and facilitate independent evaluation. The performance gap between these two sets of LLMs will indicate the contamination degree on the validation set in the future. We evaluated over 40 mainstream LLMs on the MMLU-CF. Compared to the original MMLU, not only LLMs’ performances significantly dropped but also the performance rankings of them changed considerably. This indicates the effectiveness of our approach in establishing a contamination-free and fairer evaluation standard.
Qihao Zhao, Yangyu Huang, Tengchao Lv, Lei Cui 0001, Qinzheng Sun, Shaoguang Mao, Qiufeng Yin, Scarlett Li, Furu Wei
ACL (1)10
2025 ProductMeta: An Interactive System for Metaphorical Product Design Ideation with Multimodal Large Language Models
Qinyi Zhou, Jie Deng 0001, Yun Wang 0012, Zhicong Lu, Scarlett Li, Ying-Qing Xu
CHI9
2025 PEACE: Empowering Geologic Map Holistic Understanding with MLLMs
abstract
Geologic map, as a fundamental diagram in geology science, provides critical insights into the structure and composition of Earth’s subsurface and surface. These maps are indispensable in various fields, including disaster assessment, resource exploration, and civil engineering. Despite their significance, current Multimodal Large Language Models (MLLMs) often fall short in geologic map understanding. This gap is primarily due to the challenging nature of cartographic generalization, which involves handling high-resolution map, managing multiple associated components, and requiring domain-specific knowledge. To quantify this gap, we construct GeoMap-Bench, the first-ever benchmark for evaluating MLLMs in geologic map understanding, which assesses the full-scale abilities in extracting, referring, grounding, reasoning, and analyzing. To bridge this gap, we introduce GeoMap-Agent, the inaugural agent designed for geologic map understanding, which features three modules: Hierarchical Information Extraction (HIE), Domain Knowledge Injection (DKI), and Prompt-enhanced Question Answering (PEQA). Inspired by the interdisciplinary collaboration among human scientists, an AI expert group acts as consultants, utilizing a diverse tool pool to comprehensively analyze questions. Through comprehensive experiments, GeoMap-Agent achieves an overall score of 0.811 on GeoMap-Bench, significantly outperforming 0.369 of GPT-4o. Our work, emPowering gEologic mAp holistiC undErstanding (PEACE) with MLLMs, paves the way for advanced AI applications in geology, enhancing the efficiency and accuracy of geological investigations. The code and data are available at https://github.com/microsoft/PEACE.
Yangyu Huang, Qihao Zhao, Zhipeng Gui, Tengchao Lv, Lei Cui 0001, Scarlett Li, Furu Wei
CVPR10
2025 Teaching Your Models to Understand Code via Focal Preference Alignment
abstract
Jie Wu, Haoling Li, Xin Zhang, Xiao Liu, Yangyu Huang, Jianwen Luo, Yizhen Zhang, Zuchao Li, Ruihang Chu, Yujiu Yang, Scarlett Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jie Wu 0001, Haoling Li, Xin Zhang 0099, Xiao Liu 0029, Yangyu Huang, Zuchao Li, Ruihang Chu, Yujiu Yang 0001, Scarlett Li
EMNLP11
2025 EpiCoder: Encompassing Diversity and Complexity in Code Generation
abstract
Existing methods for code generation use code snippets as seed data, restricting the complexity and diversity of the synthesized data. In this paper, we introduce a novel feature tree-based synthesis framework, which revolves around hierarchical code features derived from high-level abstractions of code. The feature tree is constructed from raw data and refined iteratively to increase the quantity and diversity of the extracted features, which captures and recognizes more complex patterns and relationships within the code. By adjusting the depth and breadth of the sampled subtrees, our framework provides precise control over the complexity of the generated code, enabling functionalities that range from function-level operations to multi-file scenarios. We fine-tuned widely-used base models to obtain EpiCoder series, achieving state-of-the-art performance on multiple benchmarks at both the function and file levels. In particular, empirical evidence indicates that our approach shows significant potential in the synthesizing of repository-level code data. Our code and data are publicly available.
Yaoxiang Wang, Haoling Li, Xin Zhang 0099, Jie Wu 0001, Xiao Liu 0029, Wenxiang Hu, Zhongxin Guo, Yangyu Huang, Yujiu Yang 0001, Jinsong Su, Qi Chen 0009, Scarlett Li
ICML13
2024 Significant ASR Error Detection for Conversational Voice Assistants
abstract
Modern Automatic Speech Recognition (ASR) systems are evaluated with respect to Word Error Rate (WER). While WER is a useful metric for training and evaluation of speech models, it does not fully reflect the difference in semantics between predicted and ground truth transcriptions. In conversational voice assistants, the ability to sufficiently understand semantic meaning of the user request is often more important than recognizing all words correctly. In this work, we propose a system that can determine, to a high degree of accuracy, whether the semantics of a predicted and reference transcript are significantly different. This knowledge is used to identify ASR errors that can result in downstream failure in conversational voice assistants. Reliable identification of these errors can be used to inform design choices for ASR systems targeting improvement on the most harmful errors.
John Harvill, Rinat Khaziev, Scarlett Li, Randy Cogill, Gopinath Chennupati, Hari Thadakamalla
ICASSP3