VLDB 2026 Research / reviewers in the wild / expert
Yun Wang 0030
dblp:36/3235-30
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0004-6611-0752ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AutoSCORE: Enhancing Automated Scoring with Multi-Agent Large Language Models via Structured Component RecognitionabstractAutomated scoring plays a crucial role in education by reducing the reliance on human raters and offering scalable and immediate evaluation of student work. While large language models (LLMs) have shown strong potential in this task, their use as end-to-end raters faces challenges such as low accuracy, prompt sensitivity, limited interpretability, and rubric misalignment, which hinder practical implementation. To address the limitations, we propose AutoSCORE, a multi-agent LLM framework enhancing automated scoring via rubric-aligned Structured COmponent REcognition. With two agents, AutoSCORE first extracts rubric-relevant components from student responses and encodes them into a structured representation (i.e., Scoring Rubric Component Extraction Agent), which is then used to assign final scores (i.e., Scoring Agent). This design ensures that model reasoning follows a human-like grading process, enhancing interpretability and robustness. We evaluate AutoSCORE on four benchmark datasets from the ASAP benchmark, using both proprietary and open-source LLMs (GPT-4o, LLaMA-3.1-8B, LLaMA-3.1-70B). Across diverse tasks and rubrics, AutoSCORE predominantly improves scoring accuracy, human-machine agreement (QWK, correlations), and reduces error metrics (MAE, RMSE) compared to single-agent baselines, with particularly strong benefits on complex, multidimensional rubrics, and especially large relative gains on smaller LLMs. These results demonstrate that structured component recognition combined with multi-agent design offers a scalable, reliable, and interpretable solution for automated scoring. Yun Wang 0030, Zhaojun Ding, Xuansheng Wu, Siyue Sun, Ninghao Liu 0001, Xiaoming Zhai |
AAAI | 1 |
| 2026 | BRIDGE the Gap: Mitigating Bias Amplification in Automated Scoring of English Language Learners via Inter-group Data Augmentation
Yun Wang 0030, Xuansheng Wu, Lei Liu 0057, Xiaoming Zhai, Ninghao Liu 0001 |
AIED (6) | 1 |
| 2026 | AI Evaluation and Feedback to Support Middle-School Students' Scientific Argumentation and Reasoning
Field M. Watts, Lei Liu 0057, Teresa M. Ober, Euvelisse Jusino-Del Valle, Yun Wang 0030, Xiaoming Zhai |
AIED (5) | 6 |
| 2026 | Using Learning Progressions to Guide AI Feedback for Science Learning
Nejla Yuruk, Yun Wang 0030, Xiaoming Zhai |
AIED (5) | 3 |
| 2025 | Artificial Intelligence Bias on English Language Learners in Automatic Scoring
Shuchen Guo, Yun Wang 0030, Jichao Yu, Xuansheng Wu, Bilgehan Ayik, Field M. Watts, Ehsan Latif, Ninghao Liu 0001, Lei Liu 0057, Xiaoming Zhai |
AIED (5) | 2 |