Ding Zou

dblp:149/7833 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Revisiting the Data Sampling in Multimodal Post-training from a Difficulty-Distinguish View
abstract
Recent advances in Multimodal Large Language Models (MLLMs) have spurred significant progress in Chain-of-Thought (CoT) reasoning. Building on the success of Deepseek-R1, researchers extended multimodal reasoning to post-training paradigms based on reinforcement learning (RL), focusing predominantly on mathematical datasets. However, existing post-training paradigms tend to neglect two critical aspects: (1) The lack of quantifiable difficulty metrics capable of strategically screening samples for post-training optimization. (2) Suboptimal post-training paradigms that fail to jointly optimize perception and reasoning capabilities. To address this gap, we propose two novel difficulty-aware sampling strategies: Progressive Image Semantic Masking (PISM) quantifies sample hardness through systematic image degradation, while Cross-Modality Attention Balance (CMAB) assesses cross-modal interaction complexity via attention distribution analysis. Leveraging these metrics, we design a hierarchical training framework that incorporates both GRPO-only and SFT+GRPO hybrid training paradigms, and evaluate them across six benchmark datasets. Experiments demonstrate consistent superiority of GRPO applied to difficulty-stratified samples compared to conventional SFT+GRPO pipelines, indicating that strategic data sampling can obviate the need for supervised fine-tuning while improving model accuracy.
Jianyu Qi, Ding Zou, Wenrui Yan, Rongchang Zhao
AAAI2
2024 An Equally-Split Bin Packing Problem
Ding Zou, Jiayi Lian, Wei Lu 0030, Yichao Duan, Yuchen Mao 0001, Guochuan Zhang
COCOA (1)1
2024 Exploring global information for session-based recommendation
Wei Wei 0002, Ding Zou, Yifan Liu 0004, Xiaoli Li 0001, Xianling Mao, Minghui Qiu
Pattern Recognit.3
2024 Towards Hierarchical Intent Disentanglement for Bundle Recommendation
abstract
Bundle recommendation aims to recommend a bundle of items for the user to purchase together, for which two scenarios (i.e.Next-bundle recommendation and Within-bundle recommendation) are explored to recommend a specific bundle of items for the user and a specific item to fill the user's current bundle, respectively. Previous works largely model the user's preference with a uniform intent, without considering the diversity of intents when adopting the items within the bundle. In the real scenario of bundle recommendation, user intents modeling actually needs to be considered from three hierarchical levels, for that: a user's intents may be naturally distributed in different bundles (user level), one bundle may contain multiple intents of a user (bundle level), and an item in different bundles may also present different user intents (item level). To this end, we develop a novel model,HierarchicalIntentDisentangleGraphNetworks (HIDGN) for bundle recommendation. HIDGN is capable of capturing the diversity of the user's intent precisely and comprehensively from the hierarchical structure with an cross-task intent contrastive learning, which is unified with the supervised next-/within-bundle recommendation sub-tasks as a multi-task framework. Extensive experiments on three benchmark datasets demonstrate that HIDGN outperforms the state-of-the-art methods by 43.0%, 13.2%, and 73.3%, respectively.
Ding Zou, Sen Zhao 0001, Wei Wei 0002, Xianling Mao, Ruixuan Li 0001, Dangyang Chen
IEEE Trans. Knowl. Data Eng.1
2022 Multi-View Intent Disentangle Graph Networks for Bundle Recommendation
abstract
Bundle recommendation aims to recommend the user a bundle of items as a whole. Previous models capture user’s preferences on both items and the association of items. Nevertheless, they usually neglect the diversity of user’s intents on adopting items and fail to disentangle user’s intents in representations. In the real scenario of bundle recommendation, a user’s intent may be naturally distributed in the different bundles of that user (Global view). And a bundle may contain multiple intents of a user (Local view). Each view has its advantages for intent disentangling: 1) In the global view, more items are involved to present each intent, which can demonstrate the user’s preference under each intent more clearly. 2) The local view can reveal the association between items under each intent since the items within the same bundle are highly correlated to each other. To this end, in this paper we propose a novel model named Multi-view Intent Disentangle Graph Networks (MIDGN), which is capable of precisely and comprehensively capturing the diversity of user intent and items’ associations at the finer granularity. Specifically, MIDGN disentangles user’s intents from two different perspectives, respectively: 1) taking the Global view, MIDGN disentangles the user’s intent coupled with inter-bundle items; 2) taking the Local view, MIDGN disentangles the user’s intent coupled with items within each bundle. Meanwhile, we compare user’s intents disentangled from different views by a contrast method to improve the learned intents. Extensive experiments are conducted on two benchmark datasets and MIDGN outperforms the state-of-the-art methods by over 10.7% and 26.8%, respectively.
Sen Zhao 0001, Wei Wei 0002, Ding Zou, Xianling Mao
AAAI3
2022 Improving Knowledge-aware Recommendation with Multi-level Interactive Contrastive Learning
abstract
Incorporating Knowledge Graphs (KG) into recommeder system as side information has attracted considerable attention. Recently, the technical trend of Knowledge-aware Recommendation (KGR) is to develop end-to-end models based on graph neural networks (GNNs). However, the extremely sparse user-item interactions significantly degrade the performance of the GNN-based models, from the following aspects: 1) the sparse interaction, itself, means inadequate supervision signals and limits the supervised GNN-based models; 2) the combination of sparse interactions (CF part) and redundant KG facts (KG part) further results in an unbalanced information utilization. Besides, the GNN paradigm aggregates local neighbors for node representation learning, while ignoring the non-local KG facts and making the knowledge extraction insufficient. Inspired by the recent success of contrastive learning in mining supervised signals from data itself, in this paper, we focus on exploring contrastive learning in KGR and propose a novel multi-level interactive contrastive learning mechanism, to alleviate the aforementioned challenges. Different from traditional contrastive learning methods which contrast nodes of two generated graph views, interactive contrastive mechanism conducts layer-wise self-supervised learning by contrasting layers of different parts within graphs, which is also an "interaction" action. Specifically, we first construct local and non-local graphs for user/item in KG, exploring more KG facts for KGR. Then an intra-graph level interactive contrastive learning is performed within each local/non-local graph, which contrasts layers of the CF and KG parts, for more consistent information leveraging. Besides, an inter-graph level interactive contrastive learning is performed between the local and non-local graphs, for sufficiently and coherently extracting non-local KG signals. Extensive experiments conducted on three benchmark datasets show the superior performance of our proposed method over the state-of-the-arts. The implementations are available at: https://github.com/CCIIPLab/KGIC.
Ding Zou, Wei Wei 0002, Xianling Mao, Feida Zhu 0001, Dangyang Chen
CIKM1
2022 Multi-level Cross-view Contrastive Learning for Knowledge-aware Recommender System
abstract
Knowledge graph (KG) plays an increasingly important role in recommender systems. Recently, graph neural networks (GNNs) based model has gradually become the theme of knowledge-aware recommendation (KGR). However, there is a natural deficiency for GNN-based KGR models, that is, the sparse supervised signal problem, which may make their actual performance drop to some extent. Inspired by the recent success of contrastive learning in mining supervised signals from data itself, in this paper, we focus on exploring the contrastive learning in KG-aware recommendation and propose a novel multi-level cross-view contrastive learning mechanism, named MCCLK. Different from traditional contrastive learning methods which generate two graph views by uniform data augmentation schemes such as corruption or dropping, we comprehensively consider three different graph views for KG-aware recommendation, including global-level structural view, local-level collaborative and semantic views. Specifically, we consider the user-item graph as a collaborative view, the item-entity graph as a semantic view, and the user-item-entity graph as a structural view. MCCLK hence performs contrastive learning across three views on both local and global levels, mining comprehensive graph feature and structure information in a self-supervised manner. Besides, in semantic view, a k-Nearest-Neighbor (k NN) item-item semantic graph construction module is proposed, to capture the important item-item semantic relation which is usually ignored by previous work. Extensive experiments conducted on three benchmark datasets show the superior performance of our proposed method over the state-of-the-arts. The implementations are available at: https://github.com/CCIIPLab/MCCLK.
Ding Zou, Wei Wei 0002, Xianling Mao, Minghui Qiu, Feida Zhu 0001, Xin Cao 0001
SIGIR1
2014 Mode-Multiplexed Multi-Tb/s Superchannel Transmission With Advanced Multidimensional Signaling in the Presence of Fiber Nonlinearities
abstract
We have analyzed the possibility of long-haul superchannel transmission with an aggregate serial bit rate exceeding 1 Tb/s by using the mode-multiplexed multidimensional signaling. We considered nonbinary quasi-cyclic LDPC-coded OFDM signals transmitted over few-mode fibers (FMFs). The optimum vector-form nonlinear Schrödinger equation is developed to evaluate the performance of the proposed system. Both the impacts of nonlinear effects and nonlinear interaction between spatial modes have been included through the modified nonlinear Schrödinger equation we applied for the FMF case. Both two-dimensional and optimized four-dimensional (4D) signal constellations have been considered. To overcome the constraints imposed by the linear and nonlinear impairments in FMF, we proposed the use of block-coded modulation with advanced channel estimation and compensation techniques. We verified by means of simulation that the transmission of an aggregate serial bit rate of 1.2 Tb/s over 3000 km is achievable with a proposed LDPC-coded QPSK-OFDM format, whereas superchannel transmission with an aggregate serial rate of 2.4 Tb/s over 1800 km is achievable with the 16-QAM format. When a 4D 16-ary optimized constellation is used, we can extend the transmission distance of mode-multiplexed QPSK-OFDM by an additional 300 km.
Changyu Lin, Ivan B. Djordjevic, Milorad Cvijetic, Ding Zou
IEEE Trans. Commun.4