Moyu Zhang

dblp:286/9517 · DBLP profile ↗
← Back
11ranked-venue papers in the field
9as first author
11since 2021 · last 2026
0000-0002-9104-1881ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 8 (7 first)Data Mining & Knowledge Discovery · 3 (2 first)
YearPublicationVenuePosition
2026 MMQ: Multimodal Mixture-of-Quantization Tokenization for Semantic ID Generation and User Behavioral Adaptation
abstract
Recommender systems traditionally represent items using unique identifiers (ItemIDs), but this approach struggles with large, dynamic item corpora and sparse long-tail data, limiting scalability and generalization. Semantic IDs, derived from multimodal content such as text and images, offer a promising alternative by mapping items into a shared semantic space, enabling knowledge transfer and improving recommendations for new or rare items. However, existing methods face two key challenges: (1) balancing cross-modal synergy with modality-specific uniqueness, and (2) bridging the semantic-behavioral gap, where semantic representations may misalign with actual user preferences. To address these challenges, we propose Multimodal Mixture-of-Quantization (MMQ), a two-stage framework that trains a novel multimodal tokenizer. First, a shared-specific tokenizer leverages a multi-expert architecture with modality-specific and modality-shared experts, using orthogonal regularization to capture comprehensive multimodal information. Second, behavior-aware fine-tuning dynamically adapts semantic IDs to downstream recommendation objectives while preserving modality information through a multimodal reconstruction loss. Extensive offline experiments and online A/B tests demonstrate that MMQ effectively unifies multimodal synergy, specificity, and behavioral adaptation, providing a scalable and versatile solution for both generative retrieval and discriminative ranking tasks.
Moyu Zhang, Chenxuan Li 0003, Zhihao Liao 0001, Haibo Xing, Hao Deng 0011, Jinxin Hu, Yu Zhang 0206, Xiaoyi Zeng, Jing Zhang 0037
WSDM2
2026 Infer As You Train: A Symmetric Paradigm of Masked Generative for Click-Through Rate Prediction
Moyu Zhang, Yujun Jin, Jinxin Hu, Yu Zhang 0206, Xiaoyi Zeng
WWW1
2025 Global-Distribution Aware Scenario-Specific Variational Representation Learning Framework
abstract
Current recommendation methods typically use a unified framework to offer personalized recommendations for different scenarios provided by commercial platforms. However, they often employ shared bottom representations, which partially hinders the model's capacity to capture scenario uniqueness. Ideally, users and items should exhibit specific characteristics in different scenarios, prompting the need to learn scenario-specific representations to differentiate scenarios. Yet, variations in user and item interactions across scenarios lead to data sparsity issues, impeding the acquisition of scenario-specific representations. To learn robust scenario-specific representations, we introduce a Global-Distribution Aware Scenario-Specific Variational Representation Learning Framework (GSVR) that can be directly applied to existing multi-scenario methods. Specifically, considering the uncertainty stemming from limited samples, our approach employs a probabilistic model to generate scenario-specific distributions for each user and item in each scenario, estimated through variational inference (VI). Additionally, we introduce the global knowledge-aware multinomial distributions as prior knowledge to regulate the learning of the posterior user and item distributions, ensuring similarities among distributions for users with akin interests and items with similar side information. This mitigates the risk of users or items with fewer records being overwhelmed in sparse scenarios. Extensive experimental results affirm the efficacy of GSVR in learning more robust representations.
Moyu Zhang, Yujun Jin, Jinxin Hu, Yu Zhang 0206
CIKM1
2025 Distribution-Guided Auto-Encoder for User Multimodal Interest Cross Fusion
abstract
Traditional recommendation methods model a user's interest in a target item by correlating its embedding with the embeddings of items from the user's interaction history, thereby capturing implicit collaborative filtering signals. Consequently, traditional ID-based methods often encounter data sparsity problems stemming from the sparse nature of ID features. To mitigate this issue, recommendation models incorporate multimodal item information to enhance recommendation accuracy. However, existing multimodal recommendation methods typically rely on early fusion approaches, which focus primarily on combining text and image features, while neglecting the dynamic context provided by user behavior sequences. This oversight precludes the dynamic adaptation of multimodal interest representations to behavioral patterns, thereby hindering the model's ability to effectively capture user multimodal interests. Therefore, this paper proposes the Distribution-Guided Multimodal-Interest Auto-Encoder (DMAE), which achieves the cross fusion of user multimodal interest at the behavioral level. Specifically, DMAE comprises three key components: 1) Multimodal Interest Encoding Unit (MIEU), which encodes the similarity scores between the target item and historically clicked items as the corresponding representation vectors of user interest across different modalities. 2) Multimodal Interest Fusion Unit (MIFU), which dynamically adapts these interest representations through both intra- and inter-modal fusion, a process contextualized by the user's behavioral sequence to achieve a fine-grained and behavior-aware representation of interest. 3) Interest-Distribution Decoding Unit (IDDU), which employs a decoder to reconstruct the encoded user interest representations into true similarity distributions for each modality. The similarity distributions serve as a guide for model learning, aiming to retain as much multimodal information as possible. Ultimately, extensive experiments demonstrate the superiority of DMAE.
Moyu Zhang, Yongxiang Tang 0001, Yujun Jin, Jinxin Hu, Yu Zhang 0206
CIKM1
2025 A Dual-Channel Heterogeneous Hypergraph Convolutional Network for Dual-target Cross-domain Recommendation
Moyu Zhang, Zhe Yang 0005
ECML/PKDD (5)1
2025 CSMF: Cascaded Selective Mask Fine-Tuning for Multi-Objective Embedding-Based Retrieval
abstract
Multi-objective embedding-based retrieval (EBR) has become increasingly critical due to the growing complexity of user behaviors and commercial objectives. While traditional approaches often suffer from data sparsity and limited information sharing between objectives, recent methods utilizing a shared network alongside dedicated sub-networks for each objective partially address these limitations. However, such methods significantly increase the model parameters, leading to an increased retrieval latency and a limited ability to model causal relationships between objectives. To address these challenges, we propose the Cascaded Selective Mask Fine-Tuning (CSMF), a novel method that enhances both retrieval efficiency and serving performance for multi-objective EBR. The CSMF framework selectively masks model parameters to free up independent learning space for each objective, leveraging the cascading relationships between objectives during the sequential fine-tuning. Without increasing network parameters or online retrieval overhead, CSMF computes a linearly weighted fusion score for multiple objective probabilities while supporting flexible adjustment of each objective's weight across various recommendation scenarios. Experimental results on real-world datasets demonstrate the superior performance of CSMF, and online experiments validate its significant practical value.
Hao Deng 0011, Haibo Xing, Kanefumi Matsuyama, Moyu Zhang, Jinxin Hu, Hong Wen 0002, Yu Zhang 0206, Xiaoyi Zeng, Jing Zhang 0037
SIGIR4
2024 Scenario-Adaptive Fine-Grained Personalization Network: Tailoring User Behavior Representation to the Scenario Context
abstract
As e-commerce has evolved, commercial platforms accommodate various scenarios to cater to the diverse shopping preferences of users.To conserve resources, current methods utilize a unified framework to deliver personalized recommendations across various scenarios.Given the overlap of users and items in multiple scenarios, current methods typically employ shared bottom representations, capturing similarities and differences between scenarios through adaptive adjustments.However, they adjust representations adaptively after aggregating user behavior sequences.This coarse-grained approach to re-weighting the entire user sequence hampers the model's ability to model the user interest migration across different scenarios.To enhance the model's capacity to capture user interests across scenarios, we develop a ranking framework named the Scenario-Adaptive Fine-Grained Personalization Network (SFPNet), which designs a fine-grained method for multiscenario personalized recommendations.Specifically, SFPNet comprises a series of blocks, stacked sequentially.Each block initially deploys a parameter personalization unit to integrate scenario information into fundamental features at a coarse-grained level, where adjusted feature representations will serve as context information.By employing residual connection, we incorporate the context into the representation of each historical behavior, allowing for contextaware fine-grained customization of the behavior representations at the scenario-level, which supports scenario-aware user interest modeling.Ultimately, the effectiveness of our method is strongly substantiated by extensive experiments and online A/B testing.
Moyu Zhang, Yongxiang Tang 0001, Jinxin Hu, Yu Zhang 0206
SIGIR1
2023 No Length Left Behind: Enhancing Knowledge Tracing for Modeling Sequences of Excessive or Insufficient Lengths
abstract
Knowledge tracing (KT) aims to predict students' responses to practices based on their historical question-answering behaviors. However, most current KT methods focus on improving overall AUC, leaving ample room for optimization in modeling sequences of excessive or insufficient lengths. As sequences get longer, computational costs will increase exponentially. Therefore, KT methods usually truncate sequences to an acceptable length, which makes it difficult for models on online service systems to capture complete historical practice behaviors of students with too long sequences. Conversely, modeling students with short practice sequences using most KT methods may result in overfitting due to limited observation samples. To address the above limitations, we propose a model called Sequence-Flexible Knowledge Tracing (SFKT). Specifically, to flexibly handle long sequences, SFKT introduces a total-term encoder to effectively model complete historical practice behaviors of students at an affordable computational cost. Additionally, to improve the prediction accuracy of students with short practice sequences, we introduce a contrastive learning task and data augmentation schema to improve the generality of modeling short sequences by constructing more learning objectives. Extensive experimental results show that SFKT achieves significant improvements over multiple benchmarks, demonstrating the value of exploring the modeling of sequences of excessive or insufficient lengths. Our code is available at https://github.com/zmy-9/SFKT.
Moyu Zhang, Xinning Zhu, Chunhong Zhang, Feng Pan 0010, Wenchen Qian, Hui Zhao 0001
CIKM1
2023 Counterfactual Monotonic Knowledge Tracing for Assessing Students' Dynamic Mastery of Knowledge Concepts
abstract
As the core of the Knowledge Tracking (KT) task, assessing students' dynamic mastery of knowledge concepts is crucial for both offline teaching and online educational applications. Since students' mastery of knowledge concepts is often unlabeled, existing KT methods focus on predicting students' responses to practices. However, purely predicting student responses without imposing specific constraints on hidden concept mastery values does not guarantee the accuracy of these intermediate values as concept mastery values. To address this issue, we propose a principled approach called Counterfactual Monotonic Knowledge Tracing (CMKT), which builds on the implicit paradigm described above by using a counterfactual assumption to constrain the evolution of students' mastery of knowledge concepts. Specifically, CMKT first assesses students' knowledge concept mastery value based on their historical practice sequences. Then, CMKT sets the answer of the most recent practice as the opposite of the actual answer and, based on this counterfactual answer, assesses the student's corresponding counterfactual knowledge mastery value. During the model training process, CMKT constrains the update of the student's knowledge states by ensuring that the two types of knowledge mastery values of students satisfy a fundamental educational theory, the monotonicity theory, to provide specific semantics for the assessed mastery values by the model. Finally, extensive experiments on five datasets demonstrate the superiority of CMKT over baseline models.
Moyu Zhang, Xinning Zhu, Chunhong Zhang, Wenchen Qian, Feng Pan 0010, Hui Zhao 0001
CIKM1
2023 Cognition-Mode Aware Variational Representation Learning Framework for Knowledge Tracing
abstract
The Knowledge Tracing (KT) task plays a crucial role in personalized learning, and its purpose is to predict student responses based on their historical practice behavior sequence. However, the KT task suffers from data sparsity, which makes it challenging to learn robust representations for students with few practice records and increases the risk of model overfitting. Therefore, in this paper, we propose a Cognition-Mode Aware Variational Representation Learning Framework (CMVF) that can be directly applied to existing KT methods. Our framework uses a probabilistic model to generate a distribution for each student, accounting for uncertainty in those with limited practice records, and estimate the student’s distribution via variational inference (VI). In addition, we also introduce a cognition-mode aware multinomial distribution as prior knowledge that constrains the posterior student distributions learning, so as to ensure that students with similar cognition modes have similar distributions, avoiding overwhelming personalization for students with few practice records. At last, extensive experimental results confirm that CMVF can effectively aid existing KT methods in learning more robust student representations. Our code is available at https://github.com/zmy-9/CMVF.
Moyu Zhang, Xinning Zhu, Chunhong Zhang, Feng Pan 0010, Wenchen Qian, Hui Zhao 0001
ICDM1
2021 Multi-Factors Aware Dual-Attentional Knowledge Tracing
abstract
With the increasing demands of personalized learning, knowledge tracing has become important which traces students' knowledge states based on their historical practices. Factor analysis methods mainly use two kinds of factors which are separately related to students and questions to model students' knowledge states. These methods use the total number of attempts of students to model students' learning progress and hardly highlight the impact of the most recent relevant practices. Besides, current factor analysis methods ignore rich information contained in questions. In this paper, we propose Multi-Factors Aware Dual-Attentional model (MF-DAKT) which enriches question representations and utilizes multiple factors to model students' learning progress based on a dual-attentional mechanism. More specifically, we propose a novel student-related factor which records the most recent attempts on relevant concepts of students to highlight the impact of recent exercises. To enrich questions representations, we use a pre-training method to incorporate two kinds of question information including questions' relation and difficulty level. We also add a regularization term about questions' difficulty level to restrict pre-trained question representations to fine-tuning during the process of predicting students' performance. Moreover, we apply a dual-attentional mechanism to differentiate contributions of factors and factor interactions to final prediction in different practice records. At last, we conduct experiments on several real-world datasets and results show that MF-DAKT can outperform existing knowledge tracing methods. We also conduct several studies to validate the effects of each component of MF-DAKT.
Moyu Zhang, Xinning Zhu, Chunhong Zhang, Yang Ji 0001, Feng Pan 0010, Changchuan Yin
CIKM1