EDBT 2026 Demo / reviewers in the wild / expert
Qian Wan 0007
dblp:25/3876-7
· DBLP profile ↗
18ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0003-4504-3912ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Leveraging Image as Compressed Visual Prompt and Hierarchical Visual Knowledge for Effective Image Utilization in MLLMsabstractMultimodal Large Language Models (MLLMs) integrate text and images for complex reasoning tasks, but efficiently utilizing image remains a challenge due to redundancy and noise. Traditional methods take the entire image features as visual prompt into the MLLMs, leading to excessive visual tokens that disrupt textual information expression. Thus, recent studies treat image features as visual knowledge, storing them in the feed-forward network for retrieval when needed. These methods, completely removing images from the input, may hinder the activation of image-related knowledge. Besides, current visual knowledge focuses on fine-grained details but overlooks the hierarchical process of visual perception. As described in feature integration theory, global structure is first processed before details are integrated. Ignoring this process may lead to a fragmented visual understanding, making it difficult to capture high-level semantic relationships. To overcome these issues, we propose a novel image utilization mechanism in MLLMs. We leverage a compression-based attention mechanism to generate the compressed visual prompt, which not only mitigates the interference of excessively long visual prompts but also preserves crucial visual information necessary for activating knowledge in the MLLM. Furthermore, we extract hierarchical visual features as visual knowledge using wavelet transforms, allowing the model to capture both global structures and fine-grained details. Experiments show that our method achieves state-of-the-art performance. Shezheng Song, Kangcheng Ding, Shan Zhao 0002, Shasha Li 0001, Xiaopeng Li 0006, Chengyu Wang 0008, Qian Wan 0007, Bin Ji 0002, Jie Yu 0008 |
AAAI | 7 |
| 2026 | RICA: Re-ranking with intra-modal and cross-modal alignment for text-based person search
Yu Bai 0022, Wentao Ma 0003, Shan Zhao 0002, Tianwei Yan 0001, Shezheng Song, Chengyu Wang 0008, Qian Wan 0007 |
Expert Syst. Appl. | 7 |
| 2025 | VCR: A "Cone of Experience" Driven Synthetic Data Generation Framework for Mathematical ReasoningabstractLarge language models (LLMs) have shown excellent performance in natural language processing but struggle with mathematical reasoning. As the training mode gradually solidifies, researchers propose a data-centric concept of artificial intelligence, emphasizing the development of higher-quality data to empower LLMs. Existing studies construct synthetic data for mathematical reasoning by expanding public datasets, thereby performing supervised fine-tuning of LLMs. However, these methods mostly focus on quantity while neglecting quality. The challenging samples fail to receive adequate consideration during data synthesis process, resulting in high construction costs, low-quality density, and serious data homogenization. This paper proposes a multi-agent environment called Virtual ClassRoom (VCR), which leverages various agents driven by LLM to construct high-quality diversified synthetic data. Inspired by the "Cone of Experience" educational theory, VCR introduces three experience levels (direct, iconic, and symbolic) into data synthesis process by analogy with human learning. A user-friendly instruction set and role-playing system are carefully designed, enabling VCR to autonomously plan the scale of synthetic data. This system covers various educational scenarios, including lecture, discussion, problem design and problem-solving. The Adaboost idea embodied in the global iterative process further promotes steady performance improvement. Extensive experiments show that the synthetic data generated by VCR possess higher quality density and generalization capability, which can give LLMs superior mathematical reasoning performance with the same scale. Sannyuya Liu, Jintian Feng, Xiaoxuan Shen, Shengyingjie Liu, Qian Wan 0007 |
AAAI | 5 |
| 2025 | Empowering Math Problem Generation and Reasoning for Large Language Model via Synthetic Data based Continual Learning FrameworkabstractThe large language models (LLMs) learning framework for math problem generation (MPG) mostly performs homogeneous training in different epochs on small-scale manually annotated data.This pattern struggles to provide large-scale new quality data to support continual improvement, and fails to stimulate the mutual promotion reaction between generation and reasoning ability of math problem, resulting in the lack of reliable solving process.This paper proposes a synthetic data based continual learning framework to improve LLMs ability for MPG and math reasoning.The framework cycles through three stages, "supervised fine-tuning, data synthesis, direct preference optimization", continuously and steadily improve performance.We propose a synthetic data method with dual mechanism of model self-play and multi-agent cooperation is proposed, which ensures the consistency and validity of synthetic data through sample filtering and rewriting strategies, and overcomes the dependence of continual learning on manually annotated data.A data replay strategy that assesses sample importance via loss differentials is designed to mitigate catastrophic forgetting.Experimental analysis on abundant authoritative math datasets demonstrates the superiority and effectiveness of our framework. Qian Wan 0007, Wangzi Shi, Jintian Feng, Shengyingjie Liu, Luona Wei, Zhicheng Dai |
EMNLP | 1 |
| 2025 | DiffuQKT: A Diffusion-Based Approach for Improved Question Representation in Knowledge TracingabstractThe rapid advancement of multimedia technologies and their increasing integration in education have underscored the importance of multimedia learning. Knowledge Tracing (KT) plays a crucial role in enabling adaptive multimedia learning by continuously monitoring students' progress and forecasting their performance throughout the learning process. Question lies at the heart of the KT process, making its representation crucial for building efficient KT models. However, the sparsity and complexity of question data pose significant challenges for existing methods to capture the underlying features of questions, thereby affecting the accuracy of knowledge state predictions. To address this issue, this paper attempts to introduce the diffusion model to the KT field, proposing a novel knowledge tracing model, DiffuQKT. The model presents a diffusion-based generative approach for question representation and enhances the stability of knowledge states through contrastive learning. Specifically, DiffuQKT first constructs question representations based on their concepts, difficulty, and variations, and then, during the forward phase, progressively adds noise to the question representations, disrupting them into a Gaussian distribution. In the reverse phase, DiffuQKT gradually recovers the representations from noise, generating higher-quality question representations for knowledge tracing. Furthermore, to guide more meaningful question generation, we incorporate question concepts and difficulty as conditions during the denoising process. In addition, to improve the robustness of knowledge states against subtle variations in question representations, we employ contrastive learning to stabilize knowledge states across both original and denoised question representations. We conduct extensive experiments on four public datasets, comparing DiffuQKT with 15 baseline methods. The results demonstrate that DiffuQKT significantly outperforms existing models. Moreover, we find that the diffusion-based generative approach for question representation proposed in this paper has the ability to significantly improve the performance of baseline models. The code can be found at https://github.com/lilstrawberry/DiffuQKT. Fenghua Yu, Qian Wan 0007, Meicheng Chen, Xiaoxuan Shen, Qing Li 0045 |
ACM Multimedia | 3 |
| 2025 | Enhancing knowledge tracing with question-based contrastive learning
Xiaoxuan Shen, Fenghua Yu, Ruxia Liang, Qian Wan 0007, Tianhao Yang, Mengtian Shi |
Knowl. Based Syst. | 5 |
| 2025 | Multi-level Contrastive Learning for Knowledge TracingabstractKnowledge Tracing (KT) is the task of predicting students’ future performance based on their past interactions with educational resources. A key aspect of KT is representation learning, which aims to capture meaningful features from students’ learning behaviors to improve prediction performance. Recently, contrastive learning methods have shown great promise in representation learning. As a result, KT models based on contrastive learning have been introduced to enhance representation learning for KT. However, these models have posed several challenges. Firstly, most of these models adopt the contrastive learning approach used in other fields, which involves data augmentation followed by contrastive learning, yet effectively applying data augmentation in KT remains an open challenge. Secondly, these models typically apply contrastive learning to only one of the fundamental components of KT: questions, interactions, or knowledge states, thereby limiting their overall performance. To address these issues, this article proposes a Multi-level Contrastive learning model for Knowledge Tracing (MCKT). MCKT (The code can be found at https://github.com/lilstrawberry/MCKT .) does not rely on data augmentation strategies; instead, it deeply integrates domain knowledge and performs contrastive learning at three levels: questions, interactions, and knowledge states. Experimental results on four publicly available datasets, compared against a total of 20 state-of-the-art KT models, demonstrate that MCKT consistently outperforms other models. Subsequent experiments further validate the effectiveness of the multi-level contrastive learning approach. Xiaoxuan Shen, Fenghua Yu, Qian Wan 0007, Ruxia Liang |
ACM Trans. Knowl. Discov. Data | 3 |
| 2024 | Revisiting Knowledge Tracing: A Simple and Powerful ModelabstractAdvances in multimedia technology and its widespread application in education have made multimedia learning increasingly important. Knowledge Tracing (KT) is the key technology for achieving adaptive multimedia learning, aiming to monitor the degree of knowledge acquisition and predict students' performance during the learning process. Current KT research is dedicated to enhancing the performance of KT problems by integrating the most advanced deep learning techniques. However, this has led to increasingly complex models, which reduce model usability and divert researchers' attention away from exploring the core issues of KT. This paper aims to tackle the fundamental challenges of KT tasks, including the knowledge state representation and the core architecture design, and investigate a novel KT model that is both simple and powerful. We have revisited the KT task and propose the ReKT model. First, taking inspiration from the decision-making process of human teachers, we model the knowledge state of students from three distinct perspectives: questions, concepts, and domains. Second, building upon human cognitive development models, such as constructivism, we have designed a Forget-Response-Update (FRU) framework to serve as the core architecture for the KT task. The FRU is composed of just two linear regression units, making it an extremely lightweight framework. Extensive comparisons were conducted with 22 state-of-the-art KT models on 7 publicly available datasets. The experimental results demonstrate that ReKT outperforms all the comparative methods in question-based KT tasks, and consistently achieves the best (in most cases) or near-best performance in concept-based KT tasks. Furthermore, in comparison to other KT core architectures like Transformers or LSTMs, the FRU achieves superior prediction performance with approximately only 38% computing resources. Through an exploration of the ReKT model that is both simple and powerful, is able to offer new insights to future KT research. The code can be found at https://github.com/lilstrawberry/ReKT. Xiaoxuan Shen, Fenghua Yu, Ruxia Liang, Qian Wan 0007, Kai Yang 0043 |
ACM Multimedia | 5 |
| 2024 | Learning to Solve Quadratic Unconstrained Binary Optimization in a Classification WayabstractThe quadratic unconstrained binary optimization (QUBO) is a well-known NP-hard problem that takes an $n\times n$ matrix $Q$ as input and decides an $n$-dimensional 0-1 vector $x$, to optimize a quadratic function. Existing learning-based models that always formulate the solution process as sequential decisions suffer from high computational overload. To overcome this issue, we propose a neural solver called the Value Classification Model (VCM) that formulates the solution process from a classification perspective. It applies a Depth Value Network (DVN) based on graph convolution that exploits the symmetry property in $Q$ to auto-grasp value features. These features are then fed into a Value Classification Network (VCN) which directly generates classification solutions. Trained by a highly efficient model-tailored Greedy-guided Self Trainer (GST) which does not require any priori optimal labels, VCM significantly outperforms competitors in both computational efficiency and solution quality with a remarkable generalization ability. It can achieve near-optimal solutions in milliseconds with an average optimality gap of just 0.362\% on benchmarks with up to 2500 variables. Notably, a VCM trained at a specific DVN depth can steadily find better solutions by simply extending the testing depth, which narrows the gap to 0.034\% on benchmarks. To our knowledge, this is the first learning-based model to reach such a performance. Jie Chun, Shang Xiang, Luona Wei, Yonghao Du, Qian Wan 0007, Yuning Chen |
NeurIPS | 6 |
| 2024 | Interpretable Knowledge Tracing with Multiscale State RepresentationabstractKnowledge Tracing (KT) is vital for education, continuously monitoring students' knowledge states (mastery of knowledge) as they interact with online education materials. Despite significant advancements in deep learning-based KT models, existing approaches often struggle to strike the right balance in granularity, leading to either overly coarse or excessively fine tracing and representation of students' knowledge states, thereby limiting their performance. Additionally, achieving a high-performing model while ensuring interpretability presents a challenge. Therefore, in this paper, we propose a novel approach called Multiscale-state-based Interpretable Knowledge Tracing (MIKT). Specifically, MIKT traces students' knowledge states on two scales: a coarse-grained representation to trace students' domain knowledge state, and a fine-grained representation to monitor their conceptual knowledge state. Furthermore, the classical psychological measurement model, IRT (Item Response Theory), is introduced to explain the prediction process of MIKT, enhancing its interpretability without sacrificing performance. Additionally, we extended the Rasch representation method to effectively handle scenarios where questions are associated with multiple concepts, making it more applicable to real-world situations. We extensively compared MIKT with 20 state-of-the-art KT models on four widely-used public datasets. Experimental results demonstrate that MIKT outperforms other models while maintaining its interpretability. Moreover, experimental observations have revealed that our proposed extended Rasch representation method not only benefits MIKT but also significantly improves the performance of other KT baseline models. The code can be found at https://github.com/lilstrawberry/MIKT. Fenghua Yu, Qian Wan 0007, Qing Li 0045, Sannyuya Liu, Xiaoxuan Shen |
WWW | 3 |
| 2024 | COMET : "cone of experience" enhanced large multimodal model for mathematical problem generation
Sannyuya Liu, Jintian Feng, Zongkai Yang, Yawei Luo, Qian Wan 0007, Xiaoxuan Shen |
Sci. China Inf. Sci. | 5 |
| 2024 | Entity-relation triple extraction based on relation sequence information
Zhanjun Zhang, Qian Wan 0007, Jie Liu 0002 |
Expert Syst. Appl. | 3 |
| 2024 | Document-level relation extraction with three channels
Zhanjun Zhang, Shan Zhao 0002, Qian Wan 0007, Jie Liu 0002 |
Knowl. Based Syst. | 4 |
| 2023 | Document-level relation extraction with hierarchical dependency tree and bridge path
Qian Wan 0007, Shangheng Du, Luona Wei, Sannyuya Liu |
Knowl. Based Syst. | 1 |
| 2023 | A Span-based Multi-Modal Attention Network for joint entity-relation extraction
Qian Wan 0007, Luona Wei, Shan Zhao 0002, Jie Liu 0002 |
Knowl. Based Syst. | 1 |
| 2022 | LELNER: A Lightweight and Effective Low-resource Named Entity Recognition model
Zhanjun Zhang, Qian Wan 0007, Jie Liu 0002 |
Knowl. Based Syst. | 3 |
| 2022 | Ancient poetry generation with an unsupervised method
Zhanjun Zhang, Qian Wan 0007, Xiangyu Jia, Jie Liu 0002 |
Neural Comput. Appl. | 3 |
| 2021 | A region-based hypergraph network for joint entity-relation extraction
Qian Wan 0007, Luona Wei, Xinhai Chen 0001, Jie Liu 0002 |
Knowl. Based Syst. | 1 |