EDBT 2026 Demo / reviewers in the wild / expert
Xiaoming Zhai
dblp:218/4494
· DBLP profile ↗
17ranked-venue papers
0as first author
17since 2021 · last 2026
0000-0003-4519-1931ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 11 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 11 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generalizable and Efficient Automated Scoring with a Knowledge-Distilled Multi-Task Mixture-of-ExpertsabstractAutomated scoring of written constructed responses typically relies on separate models per task, straining computational resources, storage, and maintenance in real-world education settings. We propose UniMoE-Guided, a knowledge-distilled multi-task Mixture-of-Experts (MoE) approach that transfers expertise from multiple task-specific large models (teachers) into a single compact, deployable model (student). The student combines (i) a shared encoder for cross-task representations, (ii) a gated MoE block that balances shared and task-specific processing, and (iii) lightweight task heads. Trained with both ground-truth labels and teacher guidance, the student matches strong task-specific models while being far more efficient to train, store, and deploy. Beyond efficiency, the MoE layer improves transfer and generalization: experts develop reusable skills that boost cross-task performance and enable rapid adaptation to new tasks with minimal additions and tuning. On nine NGSS-aligned science-reasoning tasks (seven for training/evaluation and two held out for adaptation), UniMoE-Guided attains performance comparable to per-task models while using 6x less storage than maintaining separate students, and 87x less than the 20B-parameter teacher. The method offers a practical path toward scalable, reliable, and resource-efficient automated scoring for classroom and large-scale assessment systems. Luyang Fang, Ping Ma 0001, Xiaoming Zhai |
AAAI | 4 |
| 2026 | AutoSCORE: Enhancing Automated Scoring with Multi-Agent Large Language Models via Structured Component RecognitionabstractAutomated scoring plays a crucial role in education by reducing the reliance on human raters and offering scalable and immediate evaluation of student work. While large language models (LLMs) have shown strong potential in this task, their use as end-to-end raters faces challenges such as low accuracy, prompt sensitivity, limited interpretability, and rubric misalignment, which hinder practical implementation. To address the limitations, we propose AutoSCORE, a multi-agent LLM framework enhancing automated scoring via rubric-aligned Structured COmponent REcognition. With two agents, AutoSCORE first extracts rubric-relevant components from student responses and encodes them into a structured representation (i.e., Scoring Rubric Component Extraction Agent), which is then used to assign final scores (i.e., Scoring Agent). This design ensures that model reasoning follows a human-like grading process, enhancing interpretability and robustness. We evaluate AutoSCORE on four benchmark datasets from the ASAP benchmark, using both proprietary and open-source LLMs (GPT-4o, LLaMA-3.1-8B, LLaMA-3.1-70B). Across diverse tasks and rubrics, AutoSCORE predominantly improves scoring accuracy, human-machine agreement (QWK, correlations), and reduces error metrics (MAE, RMSE) compared to single-agent baselines, with particularly strong benefits on complex, multidimensional rubrics, and especially large relative gains on smaller LLMs. These results demonstrate that structured component recognition combined with multi-agent design offers a scalable, reliable, and interpretable solution for automated scoring. Yun Wang 0030, Zhaojun Ding, Xuansheng Wu, Siyue Sun, Ninghao Liu 0001, Xiaoming Zhai |
AAAI | 6 |
| 2026 | Simulating Validity: Modal Decoupling in MLLM Generated Feedback on Science Drawings
Arne Bewersdorff, Nejla Yuruk, Xiaoming Zhai |
AIED (3) | 3 |
| 2026 | Confidence-Aware Automated Assessment of Student-Drawn Scientific Models
Luyang Fang, Yingchuan Zhang, Zhaoji Wang 0001, Ping Ma 0001, Xiaoming Zhai |
AIED (3) | 6 |
| 2026 | ArguAgent: AI-Supported Real-Time Grouping for Productive Argumentation in STEM Classrooms
Jennifer Kleiman, Yizhu Gao, Zhaoji Wang 0001, Zipei Zhu, Xiaoming Zhai |
AIED (3) | 7 |
| 2026 | BRIDGE the Gap: Mitigating Bias Amplification in Automated Scoring of English Language Learners via Inter-group Data Augmentation
Yun Wang 0030, Xuansheng Wu, Lei Liu 0057, Xiaoming Zhai, Ninghao Liu 0001 |
AIED (6) | 5 |
| 2026 | AI Evaluation and Feedback to Support Middle-School Students' Scientific Argumentation and Reasoning
Field M. Watts, Lei Liu 0057, Teresa M. Ober, Euvelisse Jusino-Del Valle, Yun Wang 0030, Xiaoming Zhai |
AIED (5) | 7 |
| 2026 | Using Learning Progressions to Guide AI Feedback for Science Learning
Nejla Yuruk, Yun Wang 0030, Xiaoming Zhai |
AIED (5) | 4 |
| 2026 | A Framework for Human-AI Q-Matrix Refinement: A NeuralCDM Evaluation
Ningxi Cheng, Yizhu Gao, Lehong Shi, Nicholas Young, Geng Yuan, Xiaoming Zhai |
AIED (3) | 8 |
| 2026 | Rethinking the Potential of Layer Freezing for DNN Training EfficiencyabstractWith the growing scale of deep neural networks and datasets, training has become increasingly expensive. Layer freezing reduces this cost by stopping updates to selected layers, but frozen layers still require forward propagation to generate activations for later layers. Caching these activations as a surrogate dataset can eliminate this redundant computation, but it faces two key challenges: effectively augmenting cached features and reducing the storage overhead of high-dimensional activations. This paper provides the first systematic study of these challenges and proposes practical solutions. We introduce Similarity-Aware Channel Augmentation to preserve accuracy by caching transformation-sensitive channels with limited overhead. We further incorporate lossy compression and design a progressive compression strategy that exploits the higher compressibility of deeper-layer activations. Our method reduces computation cost, memory usage, and training time while maintaining accuracy. Experiments on NVIDIA Orin Edge GPU further demonstrate training acceleration and significant power savings, highlighting its practicality for resource-constrained training. Chence Yang, Ningxi Cheng, Ci Zhang, Qitao Tan, Sheng Li 0019, Ao Li 0004, Xulong Tang, Shaoyi Huang, Jinzhen Wang, Jundong Li, Xiaoming Zhai, Jin Lu 0001, Geng Yuan |
ACM Great Lakes Symposium on VLSI | 13 |
| 2026 | Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM EraabstractExplainable AI (XAI) refers to techniques that provide human-understandable insights into the workings of AI models. Recently, the focus of XAI has been extended toward explaining Large Language Models (LLMs). This extension calls for a significant transformation in the XAI methodologies for two reasons. First, many existing XAI methods cannot be directly applied to LLMs due to their complexity and advanced capabilities. Second, as LLMs are increasingly deployed in diverse applications, the role of XAI shifts from merely opening the “black box” to actively enhancing the productivity and applicability of LLMs in real-world settings. Meanwhile, the conversation and generation abilities of LLMs can reciprocally enhance XAI. Therefore, in this article, we introduce Usable XAI in the context of LLMs by analyzing (1) how XAI can explain and improve LLM-based AI systems and (2) how XAI techniques can be improved by using LLMs. We introduce 10 strategies, introducing the key techniques for each and discussing their associated challenges. We also provide case studies to demonstrate how to obtain and leverage explanations. Xuansheng Wu, Haiyan Zhao 0003, Yaochen Zhu, Fan Yang 0023, Lijie Hu, Tianming Liu 0001, Xiaoming Zhai, Wenlin Yao, Jundong Li, Mengnan Du, Ninghao Liu 0001 |
ACM Trans. Knowl. Discov. Data | 8 |
| 2025 | Understanding University Students' Use of Generative AI: The Roles of Demographics and Personality Traits
Newnew Deng, Edward Jiusi Liu, Xiaoming Zhai |
AIED (1) | 3 |
| 2025 | Artificial Intelligence Bias on English Language Learners in Automatic Scoring
Shuchen Guo, Yun Wang 0030, Jichao Yu, Xuansheng Wu, Bilgehan Ayik, Field M. Watts, Ehsan Latif, Ninghao Liu 0001, Lei Liu 0057, Xiaoming Zhai |
AIED (5) | 10 |
| 2025 | Self-Regularization with Sparse Autoencoders for Controllable LLM-based ClassificationabstractModern text classification methods heavily rely on contextual embeddings from large language models (LLMs). Compared to human-engineered features, these embeddings provide automatic and effective representations for classification model training. However, they also introduce a challenge: we lose the ability to manually remove unintended features, such as sensitive or task-irrelevant features, to guarantee regulatory compliance or improve the generalizability of classification models. This limitation arises because LLM embeddings are opaque and difficult to interpret. In this paper, we propose a novel framework to identify and regularize unintended features in the LLM latent space. Specifically, we first pre-train a sparse autoencoder (SAE) to extract interpretable features from LLM latent spaces. To ensure the SAE can capture task-specific features, we further fine-tune it on task-specific datasets. In training the classification model, we propose a simple and effective regularizer, by minimizing the similarity between the classifier weights and the identified unintended feature, to remove the impact of these unintended features on classification. We evaluate the proposed framework on three real-world tasks, including toxic chat detection, reward modeling, and disease diagnosis. Results show that the proposed self-regularization framework can improve the classifier's generalizability by regularizing those features that are not semantically correlated to the task. This work pioneers controllable text classification on LLM latent spaces by leveraging interpreted features to address generalizability, fairness, and privacy challenges. The code and data are publicly available at https://github.com/JacksonWuxs/Controllable_LLM_Classifier. Xuansheng Wu, Wenhao Yu 0002, Xiaoming Zhai, Ninghao Liu 0001 |
KDD (2) | 3 |
| 2025 | SketchMind: A Multi-Agent Cognitive Framework for Assessing Student-Drawn Scientific SketchesabstractScientific sketches (e.g., models) offer a powerful lens into students' conceptual understanding, yet AI-powered automated assessment of such free-form, visually diverse artifacts remains a critical challenge. Existing solutions often treat sketch evaluation as either an image classification task or monolithic vision-language models, which lack interpretability, pedagogical alignment, and adaptability across cognitive levels. To address these limitations, we present SketchMind, a cognitively grounded, multi-agent framework for evaluating and improving student-drawn scientific sketches. SketchMind introduces Sketch Reasoning Graphs (SRGs), semantic graph representations that embed domain concepts and Bloom's taxonomy-based cognitive labels. The system comprises modular agents responsible for rubric parsing, sketch perception, cognitive alignment, and iterative feedback with sketch modification, enabling personalized and transparent evaluation. We evaluate SketchMind on a curated dataset of 3,575 student-generated sketches across six science assessment items with different highest order of Bloom's level that require students to draw models to explain phenomena. Compared to baseline GPT-4o performance without SRG (average accuracy: 55.6%), the model with SRG integration achieves 77.1% average accuracy (+21.4% average absolute gain). We also demonstrate that multi-agent orchestration with SRG enhances SketchMind performance, for example, SketchMind with GPT-4.1 gains an average 8.9% increase in sketch prediction accuracy, outperforming single-agent pipelines across all items. Human evaluators rated the feedback and co-created sketches generated by SketchMind with GPT-4.1, which achieved an average of 4.1 out of 5, significantly higher than those of baseline models (e.g., 2.3 for GPT-4o). Experts noted the system’s potential to meaningfully support conceptual growth through guided revision. Our code and (pending approval) dataset will be released to support reproducibility and future research in AI-driven education. Ehsan Latif, Zirak Khan, Xiaoming Zhai |
NeurIPS | 3 |
| 2024 | PhysicsAssistant: An LLM-Powered Interactive Learning Robot for Physics Lab InvestigationsabstractRobot systems in education can leverage Large language models’ (LLMs) natural language understanding capabilities to provide assistance and facilitate learning. This paper proposes a multimodal interactive robot (PhysicsAssistant) built on YOLOv8 object detection, cameras, speech recognition, and chatbot using LLM to provide assistance to students’ physics labs. We conduct a user study on ten 8th-grade students to empirically evaluate the performance of PhysicsAssistant with a human expert. The Expert rates the assistants’ responses to student queries on a 0-4 scale based on Bloom’s taxonomy to provide educational support. We have compared the performance of PhysicsAssistant (YOLOv8+GPT-3.5-turbo) with GPT-4 and found that the human expert rating of both systems for factual understanding is same. However, the rating of GPT-4 for conceptual and procedural knowledge (3 and 3.2 vs 2.2 and 2.6, respectively) is significantly higher than PhysicsAssistant (p < 0.05). However, the response time of GPT-4 is significantly higher than PhysicsAssistant (3.54 vs 1.64 sec, p < 0.05). Hence, despite the relatively lower response quality of PhysicsAssistant than GPT-4, it has shown potential for being used as a real-time lab assistant to provide timely responses and can offload teachers’ labor to assist with repetitive tasks. To the best of our knowledge, this is the first attempt to build such an interactive multimodal robotic assistant for K-12 science (physics) education. Ehsan Latif, Ramviyas Parasuraman, Xiaoming Zhai |
RO-MAN | 3 |
| 2023 | Matching Exemplar as Next Sentence Prediction (MeNSP): Zero-Shot Prompt Learning for Automatic Scoring in Science Education
Xuansheng Wu, Tianming Liu 0001, Ninghao Liu 0001, Xiaoming Zhai |
AIED | 5 |