Shuaishuai Zu

dblp:353/7408 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0003-4273-4240ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing the Knowledge Tracing via a Plug-In Guided Diffusion Model
abstract
Knowledge tracing (KT) refers to the problem of predicting students' future performance given their past performance. Scrutinizing previous studies, we can summarize a common learn-to-predict paradigm: a KT model first learns the student's latent knowledge states from historical question-solving learning interactions and then directly predicts whether the student could correctly answer new questions. Alongside the paradigm, existing KT models are dedicated to tailoring refinements for improving predictive performance. However, this has led to increasing model complexity and reduced usability. Inspired by the diagnosis process of human teachers, they conduct correctness prediction based on the students' responses, which are further derived from their latent knowledge states. To achieve this, we propose a novel plug-in Guided diffusiOn mODule (GOOD), which reframes the KT problem as a learn-generate-to-predict paradigm. Specifically, we first employ an existing KT backbone to learn the student's evolving latent knowledge states, subsequently feeding these into our GOOD. Next, GOOD employs a person-wise noise scheduling strategy to add noise to the target responses in the diffusion process, thereby exploring the underlying distribution of response space. Then, GOOD designs a flexible transformer-modulated denoising network to generate target responses utilizing the latent knowledge states as conditional guidance in the reverse process. Finally, the generated responses can explicitly reflect the student's performance, thereby facilitating the correctness prediction. Extensive experiments on four datasets have verified the effectiveness of GOOD in boosting existing KT models to achieve state-of-the-art performance, as well as its generalizability as a flexible plugin.
Shuaishuai Zu, Jihao Zhao, Biao Qin
AAAI1
2026 QChunker: Learning Question-Aware Text Chunking for Domain RAG via Multi-Agent Debate
abstract
The effectiveness upper bound of retrieval-augmented generation (RAG) is fundamentally constrained by the semantic integrity and information granularity of text chunks in its knowledge base. Moreover, domain documents are characterized by dense terminology and strong contextual dependencies, which exacerbate the semantic fragmentation of text chunks, thereby making it difficult to efficiently utilize their key information. To address these challenges, this paper proposes QChunker, which restructures the RAG paradigm from retrieval-augmentation to understanding-retrieval-augmentation. Firstly, QChunker models the text chunking as a composite task of text segmentation and knowledge completion to ensure the logical coherence and integrity of text chunks. Drawing inspiration from Hal Gregersen's ''Questions Are the Answer'' theory, we design a multi-agent debate framework comprising four specialized components: a question outline generator, text segmenter, integrity reviewer, and knowledge completer. This framework operates on the principle that questions serve as catalysts for profound insights. Through this pipeline, we successfully construct a high-quality dataset of 45K entries and transfer this capability to small language models. Additionally, to handle long evaluation chains and low efficiency in existing chunking evaluation methods, which overly rely on downstream QA tasks, we introduce a novel direct evaluation metric, ChunkScore. Both theoretical and experimental validations demonstrate that ChunkScore can directly and efficiently discriminate the quality of text chunks. Furthermore, during the text segmentation phase, we utilize document outlines for multi-path sampling to generate multiple candidate chunks and select the optimal solution employing ChunkScore. Extensive experimental results across four heterogeneous domains exhibit that QChunker effectively resolves aforementioned issues by providing RAG with more logically coherent and information-rich text chunks. Notably, this study also establishes a small-domain QA dataset concerning hazardous chemical safety, which fully reveals the significant value of RAG in specialized domains and the generalization capability of the QChunker framework.
Jihao Zhao, Daixuan Li, Shuaishuai Zu, Biao Qin, Hongyan Liu 0002
WWW4
2025 Multi-Modal Large Language Model with RAG Strategies in Soccer Commentary Generation
abstract
As a globally celebrated sport, soccer has seen its appeal greatly amplified by engaging and vivid commentary. Recently, Multi-Modal Large Language Models (MLLMs) have attracted attention in generating soccer commentaries due to their remarkable capacities of understanding different modalities of the input videos. Most of these methods have shown that the use of multiple modalities can enhance the commentary quality, which includes video, audio, and structured meta-data. However, delivering precise and rich commentary requires the ability to accurately discern sub-tle differences in similar backgrounds, events, and players. This presents a significant challenge for existing MLLMs. So we propose SoccerComment, a framework for generating soccer commentary that integrates MLLMs with Retrieval-Augmented Generation (RAG) strategies. This framework enhances inference efficiency and reduces the need for continuous training through a multimodal clustering memory unit and retrieval-augmented in-context learning mechanisms, ultimately improving the accuracy and diversity of the commentary. Based on similar retrieved scenarios, SoccerComment demonstrates outstanding zero-shot performance, offering a new direction and scalable solution for future research in soccer commentary generation.
Yangfan He, Shuaishuai Zu, Yiting Xie
WACV3
2025 HSEKT-GS: Hypergraph structure-enhanced for knowledge tracing with gumbel-softmax sampling
abstract
Knowledge tracing is a core task in intelligent education, aiming to model students’ knowledge states based on their learning behaviors and dynamically predict their mastery of specific concepts. In practical applications, online learning platforms typically provide a large number of questions, but each student can only interact with a small subset. This limited interaction results in extreme data sparsity for certain questions, hindering the model’s ability to build effective representations and make accurate predictions for them. Although recent hypergraph neural network-based knowledge tracing methods can model high-order heterogeneous relationships between questions, improving question representation to some extent, the hypergraph structure often overlooks latent global structural information. This limitation weakens the comprehensive semantic representation of questions, thereby affecting prediction performance. To address these challenges, we propose a Hypergraph Structure-Enhanced for Knowledge Tracing with Gumbel-Softmax sampling (HSEKT-GS). First, we construct a question–concept hypergraph and its dual graph, and incorporate a structural embedding mechanism to capture local high-order relational information between questions and between concepts. Second, to further enhance question representation, we introduce a hypergraph star expansion and use Gumbel-Softmax sampling to generate multiple perturbed embeddings per node to explore structural uncertainty. Finally, the updated representations of all sampled paths are averaged to reveal latent structural links and mitigate the over-smoothing issue in fully connected graphs. In addition, we incorporate a hypergraph structure regularization term as structural supervision to improve the robustness and interpretability of the framework. Experimental results on four publicly available datasets demonstrate that HSEKT-GS outperforms existing baseline methods.
Mingjing Tang, Jun Shen 0001, Shuaishuai Zu, Wei Gao 0012
Knowl. Based Syst.4
2025 Enhancing knowledge-aware recommendation with a cross-view contrastive learning
Shuaishuai Zu, Zhisheng Yang, Li Li 0006
Neural Comput. Appl.2
2025 Contrastive graph auto-encoder for graph embedding
Shuaishuai Zu, Li Li 0006, Jun Shen 0001, Weitao Tang
Neural Networks1
2024 CIKT: Causality Inspired Knowledge Tracing
Shuaishuai Zu, Li Li 0006, Songtao Cai, Jun Shen 0001
DASFAA (4)1
2024 GuessKT: Improving Knowledge Tracing via Considering Guess Behaviors
abstract
Knowledge tracing (KT) aims to predict students’ responses to given questions based on their historical question-answering interactions. Recent studies have proposed multiple types of KT models, mainly relying on learners’ feedback to capture the evolution of their knowledge states. However, these models leave the influence of learners’ guess behaviors out of consideration. These behaviors would cause unreliable feedback and mislead the inference process of the KT models. Exploring the guess behaviors is important since it can potentially help us understand realistic learning interactions for more accurate knowledge state estimations. In this paper, we propose GuessKT to better trace the learning progress by introducing the guess behaviors into learning feedback attribution, leading to improved prediction performance. Specifically, a windowed attention network is proposed to capture rich information from local interactions, which enhances the reliability of extracted historical feedback. Furthermore, a recovery network is proposed to recover the students’ responses after additional mask processing, which bolsters the model’s ability to recognize guess behaviors. Experiments on five datasets show that our model GuessKT advanced predictive performance over other baselines.
Shuaishuai Zu, Songtao Cai, Weitao Tang, Li Li 0006, Jun Shen 0001
ICASSP1
2023 CKGE: Improving Distance Based Knowledge Graph Embedding via Contrastive Learning
Shuaishuai Zu
ADMA (2)2
2023 Multi-level Noise Filtering and Preference Propagation Enhanced Knowledge Graph Recommendation
Shuaishuai Zu, Zhisheng Yang
ADMA (1)2
2023 Contrastive Learning Augmented Graph Auto-Encoder
Shuaishuai Zu, Jun Shen 0001, Li Li 0006
ICONIP (11)1
2023 CAKT: Coupling contrastive learning with attention networks for interpretable knowledge tracing
abstract
In intelligent systems, knowledge tracing (KT) plays a vital role in providing personalized education. Existing KT methods often rely on students' learning interactions to trace their knowledge states by predicting future performance on the given questions. While deep learning-based KT models have achieved improved predictive performance compared with traditional KT models, they often lack interpretability into the captured knowledge states. Furthermore, previous works generally neglect the multiple semantic information contained in knowledge states and sparse learning interactions. In this paper, we propose a novel model named CAKT that couples contrastive learning with attention networks for interpretable knowledge tracing. Specifically, we use three attention-based encoders to model three dynamic factors of the Item Response Theory (IRT) model, based on designed learning sequences. Then, we identify two key properties related to the knowledge states and learning interactions: consistency and separability. We utilize contrastive learning to incorporate the semantic information of the above properties into the representations of knowledge states and learning interactions. With the training goal of contrastive learning, we can obtain more representative representations of them. Extensive experiments demonstrate the excellent predictive performance of CAKT and the positive effects of considering the two properties. Additionally, CAKT can exhibit high interpretability for captured knowledge states.
Shuaishuai Zu, Li Li 0006, Jun Shen 0001
IJCNN1