Mu Du

dblp:144/1327 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
2since 2021 · last 2025
0000-0002-2988-8897ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Transfer Reinforcement Learning for Mixed Observability Markov Decision Processes with Time-Varying Interval-Valued Parameters and Its Application in Pandemic Control
abstract
We investigate a novel type of online sequential decision problem under uncertainty, namely mixed observability Markov decision process with time-varying interval-valued parameters (MOMDP-TVIVP). Such data-driven optimization problems with online learning widely have real-world applications (e.g., coordinating surveillance and intervention activities under limited resources for pandemic control). Solving MOMDP-TVIVP is a great challenge as online system identification and reoptimization based on newly observational data are required considering the unobserved states and time-varying parameters. Moreover, for many practical problems, the action and state spaces are intractably large for online optimization. To address this challenge, we propose a novel transfer reinforcement learning (TRL)-based algorithmic approach that ingrates transfer learning (TL) into deep reinforcement learning (DRL) in an offline-online scheme. To accelerate the online reoptimization, we pretrain a collection of promising networks and fine-tune them with newly acquired observational data of the system. The hallmark of our approach comes from combining the strong approximation ability of neural networks with the high flexibility of TL through efficiently adapting the previously learned policy to changes in system dynamics. Computational study under different uncertainty configurations and problem scales shows that our approach outperforms existing methods in solution optimality, robustness, efficiency, and scalability. We also demonstrate the value of fine-tuning by comparing TRL with DRL, in which at least 21% solution improvement can be yielded by TRL with fine-tuning for no more than 0.62% of time spent on pretraining in each period for problem instances with a continuous state-action space of modest dimensionality. A retrospective study on a pandemic control use case in Shanghai, China shows improved decision making via TRL in several public health metrics. Our approach is the first-ever endeavor of employing intensive neural network training in solving Markov decision processes requiring online system identification and reoptimization. History: Accepted by Paul Brooks, Area Editor for Applications in Biology, Medicine, & Healthcare. Funding: This work was supported in part by the National Natural Science Foundation of China [Grants 72371051 and 72201047] to the first and second authors and in part by the National Science Foundation [Grant 1825725] to the third author. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2022.0236 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2022.0236 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ .
Mu Du, Nan Kong
INFORMS J. Comput.1
2021 App2Vec: Context-Aware Application Usage Prediction
abstract
Both app developers and service providers have strong motivations to understandwhenandwherecertain apps are used by users. However, it has been a challenging problem due to the highly skewed and noisy app usage data. Moreover, apps are regarded as independent items in existing studies, which fail to capture the hidden semantics in app usage traces. In this article, we propose App2Vec, a powerful representation learning model to learn the semantic embedding of apps with the consideration of spatio-temporal context. Based on the obtained semantic embeddings, we develop a probabilistic model based on the Bayesian mixture model and Dirichlet process to capturewhen,where, andwhatsemantics of apps are used to predict the future usage. We evaluate our model using two different app usage datasets, which involve over 1.7 million users and 2,000+ apps. Evaluation results show that our proposed App2Vec algorithm outperforms the state-of-the-art algorithms in app usage prediction with a performance gap of over 17.0%.
Huandong Wang, Yong Li 0008, Mu Du, Zhenhui Li, Depeng Jin
ACM Trans. Knowl. Discov. Data3
2019 A Model Predictive Control Approach for Cholera Outbreaks with Mobile Sensing
abstract
Spatially specific intervention is proven to be more cost-effective than the “one-fit-all” strategy in controlling infectious disease outbreaks. However, it presents decision challenges due to partially observable epidemic state information and imperfectly determined model parameters. With deployment of mobile sensor, additional information such as geo-tagged bacteria concentration can be acquired, which can help to improve the understanding of the epidemic state and model parameters. Thus, a mobile sensor dispatching problem deciding further sensing spots should be studied considering the surveillance capacity. To solve this problem, we develop a metapopulation model based predictive control approach for making intelligent public health intervention decisions and sensor spatial dispatch decisions for cholera outbreak control. This approach incorporates a mobile sensor deployment scheme considering the improvement on unmeasurable parameter estimation and epidemic state prediction. Then it optimizes spatially specific intervention strengths to mitigate the harm of the infected population with progressively gained understanding of the system via mobile sensing.
Mu Du, Aditya Sai, Lindu Zhao, Nan Kong
KES1
2018 Tactical Production and Distribution Planning in Urban Logistics under Vehicle Operational Restrictions
abstract
Significant uncertainty associated with Chinese urban logistics, caused by random vehicle operational restrictions due to severe weather (e.g., smog) in addiction to normal traffic variation makes the tactical production and distribution planning decisions quite challenging. In this paper, we propose a two-stage stochastic integer programming model for an optimal production distribution capacity planning problem under the aforementioned uncertainties. We aim to minimize both procurement spending and the expected operational cost under logistic uncertainty. Given the computational burden of solving the resultant stochastic integer program for real-world instances, we develop an improved stochastic branch-and-bound (SBB) algorithm embedding with Tabu search method. We conduct the numerical study to verify the superiority of the proposed algorithm. We also offer managerial insights to practitioners and policy recommendations to municipal governments based on the numerical study results.
Mu Du, Nan Kong, Xiangpei Hu, Lindu Zhao
KES1
2015 Optimal Prices and Associated Factors of Product with Substitution for One Supplier and Multiple Retailers Supply Chain
abstract
This paper models the prices decision process of a one supplier and multiple retailers supply chain considering different product substitution degrees and retailer amount. Then we analyze the optimal prices for the supplier and the retailers, as well as their influencing factors under both decentralized and centralized decision-making process. The results indicate that: For the wholesale price of raw materials under the decentralized mode, both of the rigid demand of retailing market and the retailer amount act positive effect while retailers’ sub-packaging cost acts negative one. For the retail prices of the products under both decentralized and centralized mode, the rigid demand and sub-packaging costs make positive effect. However, the sub-packaging costs of other competitors have negative effect on the retail price under decentralized mode but acts positive one under the centralized mode. This paper also reveals the uncertainty of the effect from the product substitution degree through numerical analysis. Finally, specific countermeasures for the government are provided according to the perspective of the analyzed factors in order to ensure the welfare of consumers and promote the development of the industry.
Wei Fei, Mu Du
KES2