EDBT 2026 Demo / reviewers in the wild / expert
Shengyingjie Liu
dblp:333/2798
· DBLP profile ↗
10ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0003-0998-6825ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Invariant Representation Learning for Memory Behavior Modeling via Adaptive Environment SeparationabstractMemory behavior modeling seeks to predict individual recall performance and understand its underlying cognitive mechanisms. However, the dynamic and heterogeneous nature of memory data poses significant challenges to the generalization ability of models under unseen conditions. To address this challenge, we propose an invariant representation learning framework I-Mem that integrates self-supervised contrastive learning with decorrelation constraints, enabling the adaptive identification and suppression of environment-related factors in sequential behavioral data, thereby mitigating the influence of spurious features and enhancing the modeling of stable cognitive structures. Importantly, the method does not rely on explicit environment partitioning or predefined environment labels, while our theoretical analysis demonstrates that it can effectively resist environmental perturbations and facilitate the extraction of invariant structural representations, thereby ensuring adaptability and generalization. Empirical evaluations on both synthetic and real-world datasets further confirm its superiority over mainstream methods in terms of generalization performance and stable feature identification. Feature attribution analysis reveals that I-Mem extracts invariant features aligned with classical cognitive effects, and reflects short-term behavioral patterns that may indicate latent cognitive mechanisms beyond existing theories, highlighting both interpretability and discovery potential. Xiaoxuan Shen, Zhihai Hu, Fuqing Li, Shengyingjie Liu |
AAAI | 4 |
| 2026 | Information Processing Dynamic Graph For Knowledge TracingabstractKnowledge tracking (KT) aims to predict students’ performances by inferring students’ implicit knowledge mastery. Existing methods have limitations in capturing the implicit knowledge construction process of students, which often relies on the automatic integration of knowledge through information processing rather than explicit knowledge association analysis. To address this issue, we propose an Information Processing Dynamic Graph for Knowledge Tracing (IPDGKT) model based on information processing theory. IPDGKT consists of three main components, namely Short-term memory Activation (SA), Long-term memory Perception (LP), and Memory Application (MA) modules. SA extracts higher-order features of exercises and mines long-term dependencies in exercise sequences and contextual information to capture students’ short-term memory. LP characterizes the associative structure of students’ short-term and long-term memory using a Memory Interaction Graph (MIG) and designs a Memory Processing Network (MPN) to simulate the learning process to capture students’ long-term memory. MA uses students’ long-term memory to predict their future performance. We select three publicly available datasets for our experiments, and the results show that IPDGKT produces better performance predictions than existing methods. We demonstrate the value of simulated information processing for KT tasks through visualization experiments. Our code is available at https://github.com/xxiongGG/IPDGKT-main . • We simulated students’ problem-solving processes during response exercises through the deconstruction of information processing theory. • We proposed a novel DLKT model named IPDGKT. It can fully simulate the process of students from exercise to knowledge absorption. • We conducted extensive experiments on three publicly available datasets, and the results demonstrated the effectiveness of IPDGKT. Shengyingjie Liu, Yue Li 0043, Xiuling He |
Inf. Process. Manag. | 2 |
| 2025 | VCR: A "Cone of Experience" Driven Synthetic Data Generation Framework for Mathematical ReasoningabstractLarge language models (LLMs) have shown excellent performance in natural language processing but struggle with mathematical reasoning. As the training mode gradually solidifies, researchers propose a data-centric concept of artificial intelligence, emphasizing the development of higher-quality data to empower LLMs. Existing studies construct synthetic data for mathematical reasoning by expanding public datasets, thereby performing supervised fine-tuning of LLMs. However, these methods mostly focus on quantity while neglecting quality. The challenging samples fail to receive adequate consideration during data synthesis process, resulting in high construction costs, low-quality density, and serious data homogenization. This paper proposes a multi-agent environment called Virtual ClassRoom (VCR), which leverages various agents driven by LLM to construct high-quality diversified synthetic data. Inspired by the "Cone of Experience" educational theory, VCR introduces three experience levels (direct, iconic, and symbolic) into data synthesis process by analogy with human learning. A user-friendly instruction set and role-playing system are carefully designed, enabling VCR to autonomously plan the scale of synthetic data. This system covers various educational scenarios, including lecture, discussion, problem design and problem-solving. The Adaboost idea embodied in the global iterative process further promotes steady performance improvement. Extensive experiments show that the synthetic data generated by VCR possess higher quality density and generalization capability, which can give LLMs superior mathematical reasoning performance with the same scale. Sannyuya Liu, Jintian Feng, Xiaoxuan Shen, Shengyingjie Liu, Qian Wan 0007 |
AAAI | 4 |
| 2025 | ROKAN: Toward Interpretable and Domain-Robust Memory Behavior ModelingabstractMemory behavior modeling aims to predict individual performance over time and uncover underlying cognitive mechanisms. However, existing approaches often struggle to balance predictive accuracy, domain generalization, and model interpretability. To address this, we propose ROKAN, a cognitively inspired and symbolically interpretable memory modeling framework. Based on the Multiscale Context Model, ROKAN formalizes the evolution of memory traces as a differentiable Ordinary Differential Equation system, implemented via Kolmogorov-Arnold Networks to derive human-readable symbolic expressions. To enhance generalization across heterogeneous learning domains, we design an Adaptive Domain-Aware loss function, which integrates Empirical Risk Minimization with Distributionally Robust Optimization through dynamic domain-aware weighting. Our experiments demonstrate that ROKAN significantly outperforms existing mainstream methods in both predictive accuracy and domain generalization. The symbolic expressions were found to exhibit formal consistency with classical memory theories, which lends support to the model's theoretical assumptions and empirical performance, and provides a new pathway toward theoretically grounded white-box memory modeling. Our code is available at https://github.com/hellowads/ROKAN. Xiaoxuan Shen, Zhihai Hu, Shengyingjie Liu |
CIKM | 5 |
| 2025 | Empowering Math Problem Generation and Reasoning for Large Language Model via Synthetic Data based Continual Learning FrameworkabstractThe large language models (LLMs) learning framework for math problem generation (MPG) mostly performs homogeneous training in different epochs on small-scale manually annotated data.This pattern struggles to provide large-scale new quality data to support continual improvement, and fails to stimulate the mutual promotion reaction between generation and reasoning ability of math problem, resulting in the lack of reliable solving process.This paper proposes a synthetic data based continual learning framework to improve LLMs ability for MPG and math reasoning.The framework cycles through three stages, "supervised fine-tuning, data synthesis, direct preference optimization", continuously and steadily improve performance.We propose a synthetic data method with dual mechanism of model self-play and multi-agent cooperation is proposed, which ensures the consistency and validity of synthetic data through sample filtering and rewriting strategies, and overcomes the dependence of continual learning on manually annotated data.A data replay strategy that assesses sample importance via loss differentials is designed to mitigate catastrophic forgetting.Experimental analysis on abundant authoritative math datasets demonstrates the superiority and effectiveness of our framework. Qian Wan 0007, Wangzi Shi, Jintian Feng, Shengyingjie Liu, Luona Wei, Zhicheng Dai |
EMNLP | 4 |
| 2025 | HKT: Hierarchical structure-based knowledge tracing
Qing Li 0045, Zhijun Huang, Shengyingjie Liu, Zhonghua Yan |
Inf. Process. Manag. | 5 |
| 2024 | Hyperbolic embedding of discrete evolution graphs for intelligent tutoring systems
Shengyingjie Liu, Zongkai Yang, Sannyuya Liu, Ruxia Liang, Qing Li 0045, Xiaoxuan Shen |
Expert Syst. Appl. | 1 |
| 2024 | Autobalanced Multitask Node Embedding Framework for Intelligent EducationabstractRecently, online education has become popular. Many e-learning platforms have been launched with various intelligent services aimed at improving the learning efficiency and effectiveness of learners. Graphs are used to describe the pairwise relations between entities, and the node embedding technique is the foundation of many intelligent services, which have received increasing attention from researchers. However, the graph in the intelligent education scenario has three noteworthy properties, namely, heterogeneity, evolution, and lopsidedness, which makes it challenging to implement ecumenical node embedding methods on it. In this article, an autobalanced multitask node embedding model is proposed, named MNE, and applied to the interaction graph, settling a few actual tasks in intelligent education. More specifically, MNE builds two purpose-built self-supervised node embedding learning tasks for heterogeneous evolutive graphs. Edge-specific reconstruction tasks are built according to the semantic information and properties of the heterogeneous edges, and an evolutive weight regression task is designed, aiding the model to perceive the evolution of learners' implicit cognitive states. Then, both aleatoric and epistemic uncertainty quantification techniques are introduced, achieving both task- and node-level weight estimation and instructing subtask autobalancing. Experimental results on real-world datasets indicate that the proposed model outperforms the state-of-the-art graph embedding methods on two assessment tasks and demonstrates the validity of the proposed multitask framework and subtask balancing mechanism. Our implementations are available at https://github.com/ccnu-mathits/MNE4HEN. Xiaoxuan Shen, Ruxia Liang, Qing Li 0045, Shengyingjie Liu, Shangheng Du, Sannyuya Liu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Heterogeneous Evolution Network Embedding with Temporal Extension for Intelligent Tutoring SystemsabstractGraph embedding (GE) aims to acquire low-dimensional node representations while maintaining the graph’s structural and semantic attributes. Intelligent tutoring systems (ITS) signify a noteworthy achievement in the fusion of AI and education. Utilizing GE to model ITS can elevate their performance in predictive and annotation tasks. Current GE techniques, whether applied to heterogeneous or dynamic graphs, struggle to efficiently model ITS data. The GEs within ITS should retain their semidynamic, independent, and smooth characteristics. This article introduces a heterogeneous evolution network (HEN) for illustrating entities and relations within an ITS. Additionally, we introduce a temporal extension graph neural network (TEGNN) to model both evolving and static nodes within the HEN. In the TEGNN framework, dynamic nodes are initially improved over time through temporal extension (TE), providing an accurate depiction of each learner’s implicit state at each time step. Subsequently, we propose a stochastic temporal pooling (STP) strategy to estimate the embedding sets of all evolving nodes. This effectively enhances model efficiency and usability. Following this, a heterogeneous aggregation network is devised to proficiently extract heterogeneous features from the HEN. This network employs both node-level and relation-level attention mechanisms to craft aggregated node features. To emphasize the superiority of TEGNN, we perform experiments on several real ITS datasets and show that our method significantly outperforms the state-of-the-art approaches. The experiments validate that TE serves as an efficient framework for modeling temporal information in GE, and STP not only accelerates the training process but also enhances the resultant accuracy. Sannyuya Liu, Shengyingjie Liu, Zongkai Yang, Xiaoxuan Shen, Qing Li 0045, Shangheng Du |
ACM Trans. Inf. Syst. | 2 |
| 2023 | Separated Graph Neural Networks for Recommendation SystemsabstractAutomatic recommendation has become an increasingly relevant problem for industries, which allows users to discover items that match their tastes and enables the system to target items at the right users. Graph neural networks have attracted many researchers' attention and have become a useful tool for recommendation. However, these models face two major challenges, which are heterogeneous information aggregation and aggregation weight estimation. In this article, we propose a graph neural networks-based recommendation model, i.e., a separated graph neural recommendation (SGNR) model, which achieves high-quality performance. SGNR separates BINs in recommendation systems into two weighted homogeneous networks for users and items, respectively, resolving the heterogeneous information aggregation problem. In addition, a propagation coefficient estimation method is proposed, which combines parametric and nonparametric estimation strategies. And, it is constructed with three characteristics, which are collaborative, side-information constrained, and adaptive. Thereinto, a three-hierarchy attention operator is contained for feature fusion, which optimizes the feature aggregation process via a more sensible and flexible propagation mechanism. Experimental results on four public databases indicate that the proposed methods perform better than the state-of-the-art recommendation algorithms on prediction accuracy in terms of quantitative assessments and achieve readability and interpretability to some extent. Xiaoxuan Shen, Sannyuya Liu, Ruxia Liang, Shangheng Du, Shengyingjie Liu |
IEEE Trans. Ind. Informatics | 7 |