Youheng Bai

dblp:322/8976 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-1261-7708ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Beyond Next-Response Prediction: Evaluating Knowledge State Transition Consistency in Deep Learning Based Knowledge Tracing Models
Youheng Bai, Shen Han, Gangyi Tan, Jiahao Chen 0006, Zitao Liu 0001, Weiqi Luo 0002
AIED (1)1
2026 Improving Interpretability of Cognitive Diagnosis Models with LLM-based Semantic Augmentation
abstract
Cognitive diagnosis aims to infer students' mastery levels over knowledge components from their learning interactions, supporting personalized education applications. However, existing models encode responses as binary correctness labels, discarding information about which specific option a student selected. This input-level information loss limits their ability to distinguish qualitatively different error types. As a result, providing interpretable diagnostic outputs becomes challenging. To address these limitations, we propose SACD, a semantic-augmented cognitive diagnosis framework that integrates LLM-based semantic analysis with student behavioral modeling. SACD comprises the following key components. First, an LLM-based exercise diagnostic generator analyzes exercise content and produces structured semantic annotations for each answer choice, capturing the specific misconceptions each option represents. Second, a kernel-based alignment mechanism projects semantic embeddings and behavioral representations into a unified kernel space, enabling effective fusion of heterogeneous information. Third, an interpretable diagnosis layer predicts student performance and generates fine-grained mastery estimates, which LLMs further process to produce actionable learning plans. Extensive experiments on three real-world datasets demonstrate that SACD achieves superior prediction accuracy while enabling interpretable, actionable diagnostics.
Youheng Bai, Jiaqi Zheng 0012, Mingliang Hou, Teng Guo 0002, Mi Tian 0008, Xiangyu Zhao 0001, Zitao Liu 0001, Weiqi Luo 0002
SIGIR1
2026 A Frequency-Aware Mixture of Heterogeneous Experts Framework for Knowledge Tracing
abstract
Knowledge tracing (KT) aims to personalize online education on large-scale web-based platforms by modeling students' evolving knowledge states from their interaction sequences. However, most KT models rely on a single encoder architecture (e.g., self-attention or RNN), with fixed inductive biases that fails to capture the diversity of learning behaviors. Specifically, student learning unfolds across multiple timescales, and interaction sequences contain diverse frequency components ranging from short-term variations to long-term trends. Our data-driven analysis reveals that existing encoders exhibit characteristic frequency biases (e.g., self-attention tends to emphasize low-frequency patterns), highlighting the limitations of any single architecture. To address this problem, we propose FA-KT, a frequency-aware mixture of heterogeneous experts framework. FA-KT combines self-attention, Mamba, CNN, and LSTM experts, each with complementary frequency biases. A frequency-aware router analyzes each sequence's frequency characteristics and adaptively combines experts to create dynamic, personalized encoders for individual students. Across five benchmark datasets, FA-KT consistently outperforms 20 strong KT baselines in predicting future performance. Code is available at https://pykt.org/.
Youheng Bai, Mingliang Hou, Teng Guo 0002, Zitao Liu 0001, Weiqi Luo 0002
WWW1
2025 Rethinking and Improving Student Learning and Forgetting Processes for Attention based Knowledge Tracing Models
abstract
Knowledge tracing (KT) models students' knowledge states and predicts their future performance based on their historical interaction data. However, attention based KT models struggle to accurately capture diverse forgetting behaviors in ever-growing interaction sequences. First, existing models use uniform time decay matrices, conflating forgetting representations with problem relevance. Second, the fixed-length window prediction paradigm fails to model continuous forgetting processes in expanding sequences. To address these challenges, this paper introduces LefoKT, a unified architecture that enhances attention based KT models by incorporating proposed relative forgetting attention. LefoKT improves forgetting modeling through relative forgetting attention to decouple forgetting patterns from problem relevance. It also enhances attention based KT models' length extrapolation capability for capturing continuous forgetting processes in ever-growing interaction sequences. Extensive experimental results on three datasets validate the effectiveness of LefoKT.
Youheng Bai, Xueyi Li 0005, Zitao Liu 0001, Yaying Huang, Mi Tian 0008, Weiqi Luo 0002
AAAI1
2025 Denoised Attention and Question-Augmented Representations for Knowledge Tracing
abstract
Knowledge tracing (KT) is an essential task in online education systems. It aims to predict the future performance of students based on their historical learning interaction data. Despite significant advancements in attention-based KT models, they still face some limitations: inaccurate input representation and excessive student forgetting modeling. These limitations often lead to the attention noise problem: the model assigns non-negligible attention weight to some information that is cognitively irrelevant in nature, thereby generating interference signals. To address this problem, we propose a novel KT model, i.e., DenoiseKT. DenoiseKT effectively models the difficulty of the questions and utilizes graph neural network to capture the complex relationship between questions, thereby refining the representations of input features. Additionally, the denoised attention mechanism introduces a weight factor to reduce the model's attention weight distribution on irrelevant information. We extensively compare DenoiseKT with 22 state-of-the-art KT models on 4 widely-used public datasets. Experimental results show that DenoiseKT can effectively solve the attention noise problem and outperform other models. The source code of DenoiseKT is available at https://pykt.org.
Jiwei Deng, Youheng Bai, Mingliang Hou, Teng Guo 0002, Zitao Liu 0001, Weiqi Luo 0002
IJCAI2
2025 csKT: Addressing cold-start problem in knowledge tracing via kernel bias and cone attention
Youheng Bai, Xueyi Li 0005, Zitao Liu 0001, Yaying Huang, Teng Guo 0002, Mingliang Hou, Feng Xia 0001, Weiqi Luo 0002
Expert Syst. Appl.1
2025 Learning multi-granularity temporal characteristics for attention based knowledge tracing
Youheng Bai, Xueyi Li 0005, Zitao Liu 0001, Yaying Huang, Teng Guo 0002, Mingliang Hou, Weiqi Luo 0002
Neurocomputing1
2024 Extending Context Window of Attention Based Knowledge Tracing Models via Length Extrapolation
abstract
Knowledge tracing (KT) is a prediction task that aims to predict students’ future performance based on their past learning data. The rapid progress in attention mechanisms has led to the emergence of various high-performing attention based KT models. However, in online or personalized education settings, students’ varying learning paths result in different lengths of student interaction sequences, which poses a significant challenge for attention based KT models as their context window sizes are fixed during both training and prediction stages. We refer to this as the length extrapolation of KT model. In this paper, we propose extraKT to facilitate better extrapolation that learn from student interactions with a short context window and continue to perform well across various longer context window sizes at prediction stage. Specifically, we negatively bias attention scores with linearly decreasing penalties that are proportional to query-key distance, which efficiently represents short-term forgetting characteristics of student knowledge states. We conduct comprehensive and rigorous experiments on three real-world educational datasets. The results show that our extraKT model exhibits robust length extrapolation capability and outperforms state-of-the-art baseline models in terms of AUC and accuracy. To encourage reproducible research, we merge our data and code to the publicly available pyKT benchmark at https://github.com/pykt-team/pykt-toolkit.
Xueyi Li 0005, Youheng Bai, Teng Guo 0002, Ying Zheng 0010, Mingliang Hou, Bojun Zhan, Yaying Huang, Zitao Liu 0001, Boyu Gao 0003, Weiqi Luo 0002
ECAI2
2024 Enhancing Length Generalization for Attention Based Knowledge Tracing Models with Linear Biases
Xueyi Li 0005, Youheng Bai, Teng Guo 0002, Zitao Liu 0001, Yaying Huang, Xiangyu Zhao 0001, Feng Xia 0001, Weiqi Luo 0002, Jian Weng 0001
IJCAI2
2022 Extracting Precedence Relations between Video Lectures in MOOCs
abstract
Nowadays, the high dropout rate has become a widespread phenomenon in various MOOC platforms. When learning a MOOC, many learners are reluctant to spend time learning from the first video lecture to the last one. If we can recommend a learning path based on learners' individual needs and ignore irrelevant video lectures in the MOOC, it will help them learn more efficiently. The premise of learning path recommendation is to understand the precedence relations between learning resources. In this paper, we propose a novel approach for extracting precedence relations between video lectures in a MOOC. According to "knowledge depth" of concepts, we extract the core concepts from the video captions accurately. Transformer-based models are used to discover concept prerequisite relations, which help us identify the precedence relations between video lectures in MOOCs. Experiments show that the proposed method outperforms the state-of-the-art methods.
Kui Xiao, Youheng Bai
ICMR2
2022 Extracting Prerequisite Relations among Concepts From the Course Descriptions (SEKEEO-RN)
abstract
Nowadays, online learning is becoming more and more popular. Various online learning platforms provide a huge amount of learning resources for learners around the world. When choosing or sorting learning resources, learners often need to know what important knowledge concepts are addressed in each learning resource. Exploring the prerequisite relations among concepts is of great significance to educational planning. In this paper, we extracted concepts from the content of course descriptions and proposed a new approach that uses both course-based features and Wikipedia-based features to discover the prerequisite relations between knowledge concepts. Experiments on both English and Chinese datasets show that the proposed method outperforms existing baselines.
Kui Xiao, Youheng Bai
Int. J. Softw. Eng. Knowl. Eng.2