VLDB 2026 Research / reviewers in the wild / expert
Mingsheng Fu
dblp:222/8830
· DBLP profile ↗
21ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0002-9257-126XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 5 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-Policy Fusion for Multitask Multiagent Reinforcement LearningabstractMultiagent reinforcement learning (MARL) has shown strong performance in cooperative tasks. However, most existing approaches are designed for single-task scenarios and struggle to adapt to complex and dynamic environments. Multitask MARL methods aim to improve adaptability by sharing policies across tasks, but they often suffer from negative transfer due to conflicting task-specific knowledge. To address this, we propose dual-policy fusion for multitask MARL (DPF-MTMARL), which explicitly integrates a shared policy for leveraging common knowledge and task-specific policies for capturing task-specific information. Specifically, in DPF-MTMARL, we propose a learning method to efficiently train the task-specific policies and provide corresponding theoretical analysis. Additionally, we derive the theoretical conditions for decentralizing the joint policy and enforce these conditions through a regularization term during training. Extensive experiments demonstrate that DPF-MTMARL significantly outperforms state-of-the-art baselines in both homogeneous and heterogeneous task sets, effectively mitigating negative transfer and enabling robust multitask learning. Naizhuo Zeng, Mingsheng Fu, Liwei Huang, Hong Qu 0002, Zhang Yi 0001 |
IEEE Trans. Cybern. | 4 |
| 2025 | A fully value distributional deep reinforcement learning framework for multi-agent cooperation
Mingsheng Fu, Liwei Huang, Hong Qu 0002, Cheng-Zhong Xu 0001 |
Neural Networks | 1 |
| 2025 | Incremental model-based reinforcement learning with model constraintabstractIn model-based reinforcement learning (RL) approaches, the estimated model of a real environment is learned with limited data and then utilized for policy optimization. As a result, the policy optimization process in model-based RL is influenced by both policy and estimated model updates. In practice, previous model-based RL methods only perform incremental policy constraint to policy updates, which cannot assure the complete incremental updates, thereby limiting the algorithm's performance. To address this issue, we propose an incremental model-based RL update scheme by analyzing the policy optimization procedure of model-based RL. This scheme includes both an incremental model constraint that guarantees incremental updates to the estimated model, and an incremental policy constraint that ensures incremental updates to the policy. Further, we establish a performance bound incorporating the incremental model-based RL update scheme between the real environment and the estimated model, which can assure non-decreasing policy performance improvement in the real environment. To implement the incremental model-based RL update scheme, we develop a simple and efficient model-based RL algorithm known as IMPO (Incremental Model-based Policy Optimization), which leverages previous knowledge to enhance stability during the learning process. Experimental results across various control benchmarks demonstrate that IMPO significantly outperforms previous state-of-the-art model-based RL methods in terms of overall performance and sample efficiency. Zhiyou Yang, Mingsheng Fu, Hong Qu 0002, Shuqing Shi, Wang Hu 0001 |
Neural Networks | 2 |
| 2025 | Q-ADER: An Effective Q-Learning for Recommendation With Diminishing Action SpaceabstractDeep reinforcement learning (RL) has been widely applied to personalized recommender systems (PRSs) as they can capture user preferences progressively. Among RL-based techniques, deep Q-network (DQN) stands out as the most popular choice due to its simple update strategy and superior performance. Typically, many recommendation scenarios are accompanied by the diminishing action space setting, where the available action space will gradually decrease to avoid recommending duplicate items. However, existing DQN-based recommender systems inherently grapple with a discrepancy between the fixed full action space inherent in the Q-network and the diminishing available action space during recommendation. This article elucidates how this discrepancy induces an issue termed action diminishing error in the vanilla temporal difference (TD) operator. Due to this discrepancy, standard DQN methods prove impractical for learning accurate value estimates, rendering them ineffective in the context of diminishing action space. To mitigate this issue, we propose the Q-learning-based action diminishing error reduction (Q-ADER) algorithm to modify the value estimate error at each step. In practice, Q-ADER augments the standard TD learning with an error reduction term which is straightforward to implement on top of the existing DQN algorithms. Experiments are conducted on four real-world datasets to verify the effectiveness of our proposed algorithm. Hong Qu 0002, Mingsheng Fu, Wenyu Chen 0001, Zhang Yi 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Dynamic training for handling textual label noise
Shaohuan Cheng, Wenyu Chen 0001, Wanlong Liu, Li Zhou 0010, Honglin Zhao, Weishan Kong, Hong Qu 0002, Mingsheng Fu |
Appl. Intell. | 8 |
| 2024 | Class-agnostic counting and localization with feature augmentation and scale-adaptive aggregation
Yuhui Du, Hong Qu 0002, Tianlei Wang, Fan Zhang 0068, Mingsheng Fu, Wenyu Chen 0001 |
Knowl. Based Syst. | 6 |
| 2024 | ChainFrame: A Chain Framework for Point Cloud ClassificationabstractPoint cloud analysis is challenging due to its data structure. To capture the 3-D geometries, prior works mainly rely on exploring local geometric extractors. However, the human visual system suggested that both global and local features should be considered. In this article, we introduce a novel framework for point cloud classification, called ChainFrame, which takes the pair-wise global-local correlations into consideration within the intermediate scales hierarchically. The ChainFrame captures the global features that characterize the entire outline of the object. Simultaneously, the ChainFrame organizes the local features that incorporate the point itself and its neighboring region. With such a framework, our practical implementations (ChainMLP and ChainGraph) perform on par or even better than other methods. Evaluations on two popular datasets show the effectiveness and efficiency of our ChainFrame. ChainMLP and ChainGraph achieve$ 87.2%$and$87.6%$overall point-wise accuracy scores, respectively, on the real-world ScanObjectNN benchmark. Besides, ChainMLP delivers comparable performance on ModelNet40 with only 0.47 M parameters and 0.33 G floating point operations (FLOPs), which are much smaller than the prior methods. Tianlei Wang, Mingsheng Fu, Hong Qu 0002, Ma Luo |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | A Distributional Perspective on Multiagent Cooperation With Deep Reinforcement LearningabstractAmong various value decomposition-based multiagent reinforcement learning (MARL) algorithms, the overall performance of the multiagent system is represented by a scalar global Q value and optimized by minimizing the temporal difference (TD) error with respect to that global Q value. However, the global Q value cannot accurately model the distributed dynamics of the multiagent system, since it is only a simplified representation for different individual Q values of agents. To explicitly consider the correlations between different cooperative agents, in this article, we propose a distributional framework and construct a practical model called distributional multiagent cooperation (DMAC) from a novel distributional perspective. Specifically, in DMAC, we view the individual Q value for the executed action of a random agent as a value distribution, whose expectation can further represent the overall performance. Then, we employ distributional RL to minimize the difference between the estimated distribution and its target for the optimization. The advantage of DMAC is that the distributed dynamics of agents can be explicitly modeled, and this results in better performance. To verify the effectiveness of DMAC, we conduct extensive experiments under nine different scenarios of the StarCraft Multiagent Challenge (SMAC). Experimental results show that the DMAC can significantly outperform the baselines with respect to the average median test win rate. Liwei Huang, Mingsheng Fu, Ananya Rao, Athirai Aravazhi Irissappane, Jie Zhang 0002, Cheng-Zhong Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Improving Exploration in Actor-Critic With Weakly Pessimistic Value Estimation and Optimistic Policy OptimizationabstractDeep off-policy actor-critic algorithms have been successfully applied to challenging tasks in continuous control. However, these methods typically suffer from the poor sample efficiency problem, limiting their widespread adoption in real-world domains. To mitigate this issue, we propose a novel actor-critic algorithm with weakly pessimistic value estimation and optimistic policy optimization (WPVOP) for continuous control. WPVOP integrates two key ingredients: 1) a weakly pessimistic value estimation, which compensates the pessimism of lower confidence bound in conventional value function (i.e., clipped double Q -learning) to trigger exploration in low-value state-action regions and 2) an optimistic policy optimization algorithm by sampling actions that could benefit the policy learning most toward optimal Q -values for efficient exploration. We theoretically analyze that the proposed weakly pessimistic value estimation method is lower and upper bounded, and empirically show that it could avoid extremely over-optimistic value estimates. We show that these two ideas are largely complementary, and can be fruitfully integrated to improve performance and promote sample efficiency of exploration. We evaluate WPVOP on the suite of continuous control tasks from MuJoCo, achieving state-of-the-art sample efficiency and performance. Mingsheng Fu, Wenyu Chen 0001, Fan Zhang 0068, Haixian Zhang, Hong Qu 0002, Zhang Yi 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | A Maximum Divergence Approach to Optimal Policy in Deep Reinforcement LearningabstractModel-free reinforcement learning algorithms based on entropy regularized have achieved good performance in control tasks. Those algorithms consider using the entropy-regularized term for the policy to learn a stochastic policy. This work provides a new perspective that aims to explicitly learn a representation of intrinsic information in state transition to obtain a multimodal stochastic policy, for dealing with the tradeoff between exploration and exploitation. We study a class of Markov decision processes (MDPs) with divergence maximization, called divergence MDPs. The goal of the divergence MDPs is to find an optimal stochastic policy that maximizes the sum of both the expected discounted total rewards and a divergence term, where the divergence function learns the implicit information of state transition. Thus, it can provide better-off stochastic policies to improve both in robustness and performance in a high-dimension continuous setting. Under this framework, the optimality equations can be obtained, and then a divergence actor-critic algorithm is developed based on the divergence policy iteration method to address large-scale continuous problems. The experimental results, compared to other methods, show that our approach achieved better performance and robustness in the complex environment particularly. The code of DivAC can be found in https://github.com/yzyvl/DivAC. Zhiyou Yang, Hong Qu 0002, Mingsheng Fu, Wang Hu 0001, Yongze Zhao |
IEEE Trans. Cybern. | 3 |
| 2023 | A Deep Reinforcement Learning Recommender System With Multiple Policies for RecommendationsabstractDeep reinforcement learning (DRL) based recommender systems are suitable for user cold-start problems as they can capture user preferences progressively. However, most existing DRL-based recommender systems are suboptimal, since they use the same policy to suit the dynamics of different users. We reformulate recommendation as a multitask Markov Decision Process, where each task represents a set of similar users. Since similar users have closer dynamics, a task-specific policy is more effective than a single universal policy for all users. To make recommendations for cold-start users, we use a default policy to collect some initial interactions to identify the user task, after which a task-specific policy is employed. We use Q-learning to optimize our framework and consider the task uncertainty by the mutual information regarding tasks. Experiments are conducted on three real-world datasets to verify the effectiveness of our proposed framework. Mingsheng Fu, Liwei Huang, Ananya Rao, Athirai Aravazhi Irissappane, Jie Zhang 0002, Hong Qu 0002 |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | Neural Reranking-Based Collaborative Filtering by Leveraging Listwise Relative Ranking InformationabstractReranking is a critical task used to refine the initial collaborative filtering (CF) recommendation by incorporating information from different viewpoints, such as the extra item side-information and user profile. In this article, a neural reranking-based CF (NRCF) model is proposed to leverage composite viewpoints from the basic CF model and user preference. More precisely, the predictive implicit user preference is first constructed from the initial top-$k$items. The implicit user preference is then aggregated with the explicit user embedding to enrich the user intent representation. Moreover, the traditional listwise loss functions for reranking optimization are suboptimal, due to the fact that they neglect the relative ranking information (ReinRank) between the unobserved and positive items. To address this issue, a novel listwise loss function that leverages relative ranking information, referred to as ReinRank, is proposed for reranking optimization. ReinRank assigns different values to the unobserved items, according to their relative ranking distances between the positive items. Extensive experiments are performed on three public benchmarks and different CF models, in order to demonstrate the effectiveness of NRCF and ReinRank. Hong Qu 0002, Mingsheng Fu, Fan Zhang 0068, Wenyu Chen 0001, Ruixuan Sun, Haixian Zhang |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2022 | Deep Reinforcement Learning Framework for Category-Based Item RecommendationabstractDeep reinforcement learning (DRL)-based recommender systems have recently come into the limelight due to their ability to optimize long-term user engagement. A significant challenge in DRL-based recommender systems is the large action space required to represent a variety of items. The large action space weakens the sampling efficiency and thereby, affects the recommendation accuracy. In this article, we propose a DRL-based method called deep hierarchical category-based recommender system (DHCRS) to handle the large action space problem. In DHCRS, categories of items are used to reconstruct the original flat action space into a two-level category-item hierarchy. DHCRS uses two deep Q -networks (DQNs): 1) a high-level DQN for selecting a category and 2) a low-level DQN to choose an item in this category for the recommendation. Hence, the action space of each DQN is significantly reduced. Furthermore, the categorization of items helps capture the users' preferences more effectively. We also propose a bidirectional category selection (BCS) technique, which explicitly considers the category-item relationships. The experiments show that DHCRS can significantly outperform state-of-the-art methods in terms of hit rate and normalized discounted cumulative gain for long-term recommendations. Mingsheng Fu, Anubha Agrawal, Athirai Aravazhi Irissappane, Jie Zhang 0002, Liwei Huang, Hong Qu 0002 |
IEEE Trans. Cybern. | 1 |
| 2022 | An Attention-Based Interactive Learning-to-Rank Model for Document RetrievalabstractThe core issue of learning-to-rank (LTR) for document retrieval lies in finding an optimal ranking policy to meet the search intent of the user. The majority of proposed LTR approaches treat the ranking as a static process, employing a fixed ranking policy to immediately assign scores to documents. By contrast, ranking is not a static but an interactive process where the user continues interacting with the document retrieval system through information exchange such as search intent (e.g., rating or clicking for the retrieved items). We model the interactive ranking process (IRP), and propose an Attention-Based Interactive LTR model (AIRank) to constitute an intent-aware flexible ranking policy to gratify the user’s need. To enhance the ranking quality, the inherent relations among documents are procured by the self-attention method to contribute to an enriched user intent representation. Furthermore, we mend the policy gradient learning method to train the AIRank in the IRP. Experiments demonstrate the effectiveness of AIRank compared to the state-of-the-art methods in terms of normalized discounted cumulative gain and expected reciprocal rank. Fan Zhang 0068, Wenyu Chen 0001, Mingsheng Fu, Hong Qu 0002, Zhang Yi 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | A deep reinforcement learning-based method applied for solving multi-agent defense and attack problems
Liwei Huang, Mingsheng Fu, Hong Qu 0002, Siying Wang 0002, Shangqian Hu |
Expert Syst. Appl. | 2 |
| 2021 | A deep reinforcement learning based long-term recommender system
Liwei Huang, Mingsheng Fu, Hong Qu 0002, Yangjun Liu, Wenyu Chen 0001 |
Knowl. Based Syst. | 2 |
| 2019 | A Novel Deep Learning-Based Collaborative Filtering Model for Recommendation SystemabstractThe collaborative filtering (CF) based models are capable of grasping the interaction or correlation of users and items under consideration. However, existing CF-based methods can only grasp single type of relation, such as restricted Boltzmann machine which distinctly seize the correlation of user-user or item-item relation. On the other hand, matrix factorization explicitly captures the interaction between them. To overcome these setbacks in CF-based methods, we propose a novel deep learning method which imitates an effective intelligent recommendation by understanding the users and items beforehand. In the initial stage, corresponding low-dimensional vectors of users and items are learned separately, which embeds the semantic information reflecting the user-user and item-item correlation. During the prediction stage, a feed-forward neural networks is employed to simulate the interaction between user and item, where the corresponding pretrained representational vectors are taken as inputs of the neural networks. Several experiments based on two benchmark datasets (MovieLens 1M and MovieLens 10M) are carried out to verify the effectiveness of the proposed method, and the result shows that our model outperforms previous methods that used feed-forward neural networks by a significant margin and performs very comparably with state-of-the-art methods on both datasets. Mingsheng Fu, Hong Qu 0002, Zhang Yi 0001, Li Lu 0001 |
IEEE Trans. Cybern. | 1 |
| 2018 | Deep Dilated Convolution on Multimodality Time Series for Human Activity RecognitionabstractConvolutional Neural Networks (CNNs) is capable of automatically learning feature representations, CNN-based recognition algorithm has been an alternative method for human activity recognition. Even though general convolution operation followed by pooling could expand the receptive fields for extracting features, it will bring about information loss in feature representation. Due to that dilated convolutions not only could expand receptive field exponentially without changing the size of field map or pooling, but it also will not cause information loss, hence, we propose D2CL, a novel deep learning framework for human activity recognition using multi-model wearable sensors. This framework consists of dilated convolutional neural networks and recurrent neural networks. At first, learning from previous works, we add a general convolutional layer to map inputs into a hidden space for improving the capability of nonlinear representations. Subsequently, a stacked dilated convolutional networks automatically learn feature representations for inter-sensors and intra-sensors from hidden space. Then, given these learned features, two RNNs are applied to model their latent temporal dependencies. Finally, a softmax classifier at the topmost layer is utilized to recognize activities. To evaluate the performance of D2CL on activity recognition, we select two open datasets OPPORTUNITY and PAMAP2 for training and testing. Results show that our proposed model achieves a higher classification performance than the state-of-the-art DeepConvLSTM. Mengshu Hou, Mingsheng Fu, Hong Qu 0002, Daibo Liu |
IJCNN | 3 |
| 2018 | Reinforcement Learning for Mobile Robot Obstacle Avoidance Under Dynamic Environments
Liwei Huang, Hong Qu 0002, Mingsheng Fu, Wu Deng 0004 |
PRICAI (1) | 3 |
| 2018 | Bag of meta-words: A novel method to represent document for the sentiment classification
Mingsheng Fu, Hong Qu 0002, Li Huang 0002, Li Lu 0001 |
Expert Syst. Appl. | 1 |
| 2018 | Attention based collaborative filtering
Mingsheng Fu, Hong Qu 0002, Alemu Dagmawi Moges, Li Lu 0001 |
Neurocomputing | 1 |