Haotian Zhang 0007

dblp:83/4184-7 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0003-0133-9762ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Look as You Think: Unifying Reasoning and Visual Evidence Attribution for Verifiable Document RAG via Reinforcement Learning
abstract
Aiming to identify precise evidence sources from visual documents, visual evidence attribution for visual document retrieval–augmented generation (VD-RAG) ensures reliable and verifiable predictions from vision-language models (VLMs) in multimodal question answering. Most existing methods adopt end-to-end training to facilitate intuitive answer verification. However, they lack fine-grained supervision and progressive traceability throughout the reasoning process. In this paper, we introduce the Chain-of-Evidence (CoE) paradigm for VD-RAG. CoE unifies Chain-of-Thought (CoT) reasoning and visual evidence attribution by grounding reference elements in reasoning steps to specific regions with bounding boxes and page indexes. To enable VLMs to generate such evidence-grounded reasoning, we propose Look As You Think (LAT), a reinforcement learning framework that trains models to produce verifiable reasoning paths with consistent attribution. During training, LAT evaluates the attribution consistency of each evidence region and provides rewards only when the CoE trajectory yields correct answers, encouraging process-level self-verification. Experiments on vanilla Qwen2.5-VL-7B-Instruct with Paper‑ and Wiki‑VISA benchmarks show that LAT consistently improves the vanilla model in both single- and multi-image settings, yielding average gains of 8.23% in soft exact match (EM) and 47.0% in [email protected]. Meanwhile, LAT not only outperforms the supervised fine-tuning baseline, which is trained to directly produce answers with attribution, but also exhibits stronger generalization across domains.
Shuochen Liu, Pengfei Luo, Chao Zhang 0096, Haotian Zhang 0007, Qi Liu 0003, Xin Kou, Tong Xu 0001, Enhong Chen
AAAI5
2025 GenAL: Generative Agent for Adaptive Learning
abstract
Adaptive learning, also known as adaptive teaching, relies on learning path recommendations that sequentially suggest personalized learning items (such as lectures and exercises) to meet the unique needs of each learner. Despite the extensive research in this field, previous approaches have primarily modeled the interaction sequences between learners and items using simple indexing, leading to three issues: (1) The utilization of information from both learners and items is not sufficient. For instance, these models are unable to leverage the semantic information contained within the textual content of the items. (2) Models need to be retrained on different datasets separately, which makes it difficult to adapt to the continuously expanding item pool in online educational scenarios. (3) The existing recommendation paradigm based on trained reinforcement learning frameworks, suffers from unstable recommendation performance in sparse learning logs. To address these challenges, we propose a generalized Generative Agent for Adaptive Learning (GenAL), which integrates educational tools with LLMs' semantic understanding to enable effective and generalizable learning path recommendations across diverse data distributions. Specifically, our framework consists of two components: the Global Thinking Agent, which updates the learner profile and reflects on recommendation outcomes based on the learner's historical learning records. The other is the Local Teaching Agent, which recommends items using educational prior knowledge. Leveraging the LLM's robust semantic understanding, our framework does not rely on item indexing but instead extracts relevant information from the textual content. We evaluated our approach on three real-world datasets, and the experimental results demonstrate that our GenAL not only consistently outperforms all baselines but also exhibits strong generalization ability.
Rui Lv, Qi Liu 0003, Weibo Gao, Haotian Zhang 0007, Junyu Lu 0003, Linbo Zhu
AAAI4
2025 Continuous Dynamic Modeling via Neural ODEs for Popularity Trajectory Prediction
Songbo Yang, Ziwei Zhao 0002, Haotian Zhang 0007, Tong Xu 0001, Mengxiao Zhu 0001
DASFAA (2)4
2025 SA-MBKT: Surrogate Model-assisted Multi-skills Bayesian Knowledge Tracing
abstract
Knowledge Tracing (KT) is a fundamental task in educational data mining that mainly focuses on tracing students’ dynamic knowledge states of skills. Bayesian Knowledge Tracing (BKT) has been widely researched and applied due to its good interpretability, using the hidden Markov model to model students’ question–answering process. Standard BKT considers only one skill in each question. To address this limitation, we proposed a Multi-skills Bayesian Knowledge Tracing (MBKT) method based on evolutionary algorithms in our previous work. MBKT employs evolutionary algorithms as the optimization method for BKT, enabling it to trace changes in students’ mastery of multiple skills simultaneously. However, MBKT has the drawback of taking too long for a single individual evaluation, and a large number of valueless individuals invoke the real evaluation process, especially when dealing with a large amount of data to be evaluated. This hinders its application in real online education scenarios. Therefore, the Surrogate Model-assisted Multi-skills Bayesian Knowledge Tracing (SA-MBKT) method is proposed to address these issues by introducing a window strategy and a surrogate model method. Extensive experiments on real-world datasets demonstrate that SA-MBKT significantly enhances temporal performance without affecting the predictive performance of the model.
Chenyang Bu, Haotian Zhang 0007, Lei Li 0002, Wenjian Luo
ACM Trans. Evol. Learn. Optim.3
2024 Item-Difficulty-Aware Learning Path Recommendation: From a Real Walking Perspective
abstract
Learning path recommendation aims to provide learners with a reasonable order of items to achieve their learning goals. Intuitively, the learning process on the learning path can be metaphorically likened to walking. Despite extensive efforts in this area, most previous methods mainly focus on the relationship among items but overlook the difficulty of items, which may raise two issues from a real walking perspective: (1) The path may be rough: When learners tread the path without considering item difficulty, it's akin to walking a dark, uneven road, making learning harder and dampening interest. (2) The path may be inefficient: Allowing learners only a few attempts on very challenging items before switching, or persisting with a difficult item despite numerous attempts without mastery, can result in inefficiencies in the learning journey. To conquer the above limitations, we propose a novel method named Difficulty-constrained Learning Path Recommendation (DLPR), which is aware of item difficulty. Specifically, we first explicitly categorize items into learning items and practice items, then construct a hierarchical graph to model and leverage item difficulty adequately. Then we design a Difficulty-driven Hierarchical Reinforcement Learning (DHRL) framework to facilitate learning paths with efficiency and smoothness. Finally, extensive experiments on three different simulators demonstrate our framework achieves state-of-the-art performance.
Haotian Zhang 0007, Shuanghong Shen, Bihan Xu, Zhenya Huang, Jing Sha, Shijin Wang 0001
KDD1
2024 Graph-based Student Knowledge Profile for Online Intelligent Education
abstract
Student knowledge profile is the basis for adaptive learning applications in online learning resulting from modeling the student mastery of knowledge concepts. In recent years, typical works based on knowledge tracing (KT) expect to profile students and have achieved significant success for the next performance prediction. However, in practical online learning scenarios, current methods tend to suffer from the following challenges: 1) Prediction inconsistency: The accuracy of the next performance prediction is inconsistent with the accuracy of student knowledge profile prediction, which is the more required result. 2) Cold start of knowledge: In online learning scenarios, it is often necessary to profile some knowledge concepts without learning records in advance. In this paper, we propose a novel Graph-based Student Knowledge Profile Model (GSKPM), along with a new end-to-end training objective, to tackle these challenges. We first define a new training objective to ensure the model is capable of inferring consistent student knowledge profiles. Then in this model, a two-stage hyper-aggregation process is employed to make full use of the topological relations between knowledge concepts and knowledge domains to provide information during profiling, especially for cold start knowledge concepts. Finally, through extensive experiments on real-world datasets, we will show that GSKPM achieves better prediction performances on student knowledge profiles and well deals with the cold start problem.
Haotian Zhang 0007, Zhenya Huang, Qi Liu 0003, Jing Sha, Enhong Chen, Shijin Wang 0001
SDM2
2024 FDKT: Towards an Interpretable Deep Knowledge Tracing via Fuzzy Reasoning
abstract
In educational data mining, knowledge tracing (KT) aims to model learning performance based on student knowledge mastery. Deep-learning-based KT models perform remarkably better than traditional KT and have attracted considerable attention. However, most of them lack interpretability, making it challenging to explain why the model performed well in the prediction. In this paper, we propose an interpretable deep KT model, referred to as fuzzy deep knowledge tracing (FDKT) via fuzzy reasoning. Specifically, we formalize continuous scores into several fuzzy scores using the fuzzification module. Then, we input the fuzzy scores into the fuzzy reasoning module (FRM). FRM is designed to deduce the current cognitive ability, based on which the future performance was predicted. FDKT greatly enhanced the intrinsic interpretability of deep-learning-based KT through the interpretation of the deduction of student cognition. Furthermore, it broadened the application of KT to continuous scores. Improved performance with regard to both the advantages of FDKT was demonstrated through comparisons with the state-of-the-art models.
Fei Liu 0038, Chenyang Bu, Haotian Zhang 0007, Le Wu 0001, Kui Yu, Xuegang Hu
ACM Trans. Inf. Syst.3
2022 APGKT: Exploiting Associative Path on Skills Graph for Knowledge Tracing
Haotian Zhang 0007, Chenyang Bu, Fei Liu 0038, Shuochen Liu, Yuhong Zhang 0002, Xuegang Hu
PRICAI (1)1