Mingjun Xiao

dblp:94/3601 · DBLP profile ↗
← Back
12ranked-venue papers in the field
1as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 10 (1 first)Data Mining & Knowledge Discovery · 1Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2025 Online Federated Learning on Distributed Unknown Data Using UAVs
abstract
Along with the advance of low-altitude economy, a variety of applications based on Unmanned Aerial Vehicles (UAVs) have been developed to accomplish diverse tasks. In this paper, we focus on the scenario of multiple UAVs performing Federated Learning (FL) tasks. Specifically, a group of UAVs is scheduled to repeatedly visit some Points of Interest (PoIs), collect the data produced by these PoIs, and jointly train a machine learning model based on the collected data. The most challenging issue is how to schedule UAVs to collect data so as to optimize the generalization and convergence of model training under the case that the distributions of the data produced by PoIs have not been known in advance. To address this issue, we propose a novel framework for online FL on distributed unknown data, named OFL-UD2, which is dedicated to online decision-making for UAVs to optimize model training performance. Concretely, we formulate the optimization problem while considering the convergence and quality of trained models as well as energy constraints. Then, we define a utility metric for the data quality of different PoIs and conduct a rigorous convergence analysis for OFL-UD2. Based on the analysis results, we design a two-stage algorithm to determine the scheduling of UAVs. Extensive simulations demonstrate that OFL-UD2can improve model accuracy and speed up running time compared to existing benchmarks significantly.
Xichong Zhang, Yin Xu 0004, Mingjun Xiao, Jie Wu 0001, Jinrui Zhou
ICDE4
2025 Federated Graph Out-of-Distribution Generalization via Representation Propagation and Scattering
abstract
Federated Graph Learning (FGL) enables collab-orative model training across decentralized graph data while preserving privacy. However, FGL faces severe performance degradation under out-of-distribution (OOD) shifts due to both feature distribution divergence and structural heterogeneity among clients. To address this, we propose FGOOD, a lightweight and effective framework that improves OOD generalization in FGL. FGOOD integrates two key components: (1) representation propagation, which enhances structural robustness by aggregating multi-hop topology while preserving local features, and (2) representation scattering, which regularizes node embeddings toward a uniformly dispersed distribution on the hypersphere, improving inter-class separation without requiring contrastive pairs. The theoretical analysis provides an upper bound on the generalization error under distribution shifts. Extensive experiments on three real-world datasets demonstrate that FGOOD outperforms existing state-of-the-art baselines, improving OOD accuracy by up to 5% while remaining lightweight and scalable.
Yukai Zhu 0005, Mingjun Xiao
ICDM3
2024 MACRO: Incentivizing Multi-Leader Game-Based Pareto-Efficient Crowdsourcing for Video Analytics
abstract
In recent years, many crowdsourcing platforms have emerged, using the resources of recruited workers to perform diverse outsourcing tasks, where the video analytics attracts much attention due to its practical implications. For maximum profits, platforms carefully choose the workers and determine the video analytics configurations to ensure accuracy; meanwhile, workers possess the flexibility to tailor the configurations for their indivi-dual gains, which makes it hard for platforms to optimize their profits considering the platform-worker conflicts. In this paper, we design an incentive mechanism for Multi-leader game-based video Analytics upon CROwdsourcing, named MACRO, to over-come the above situation. Under that mechanism, we first formu-late the utility optimization problems for platforms and workers, respectively. We then propose a dual ascent-based method to op-timally determine the video analytics configurations for a multi-platform game, ensuring Pareto efficiency. Moreover, in the context of a multi-leader game involving platform-worker conflicts, we design an incentive function with its incentive factor update strategy and propose an ADMM-based approach for maximizing incentives that motivate workers to contribute to the platforms' profits. Rigorous proofs demonstrate the linear convergence of the MACRO to the multi-leader Stackelberg equilibrium. Trace-driven experiments show that MACRO improves the Pareto efficiency by 26.3%, outperforming other approaches.
Yu Chen 0038, Sheng Zhang 0001, Ziying Zhou, Xiaokun Wang 0002, Yu Liang 0001, Ning Chen 0010, Mingjun Xiao, Jie Wu 0001, Zhuzhong Qian, Guoqing Harry Xu
ICDE8
2024 Joint Mobile Edge Caching and Pricing: A Mean-Field Game Approach
abstract
In this paper, we investigate the competitive content placement problem in Mobile Edge Caching (MEC) systems, where Edge Data Providers (EDPs) cache appropriate contents and trade them with requesters at a suitable price. Most of the existing works ignore the complicated strategic and economic interplay between content caching, pricing, and content sharing. Therefore, we propose a joint Mean-Field Game framework for mobile edge Caching and Pricing (MFG-CP) in large-scale dynamic MEC systems, which can facilitate distributed optimal decision-making based on the mean-field game theory. Specifi-cally, we first formulate the competitive content placement issue among EDPs as a non-cooperative stochastic differential game. To significantly reduce the communication and computation complexity, we further devise a mean-field model to approximate the collective impact of all EDPs on caching, trading, and sharing, by which each EDP can quickly estimate some unknown information without considerable interactions. Then, we develop a distributed best response scheme based on iterative learning, enabling each EDP to solely customize its optimal caching strategy and pricing policy. Besides, we theoretically prove the existence of a unique MFG equilibrium. Finally, trace-driven simulations demonstrate the effectiveness of MFG-CP compared with some baselines.
Yin Xu 0004, Xichong Zhang, Mingjun Xiao, Jie Wu 0001, An Liu 0002, Sheng Zhang 0001
ICDE3
2021 Crowdsensing Data Trading based on Combinatorial Multi-Armed Bandit and Stackelberg Game
abstract
Crowdsensing Data Trading (CDT), through which a platform can aggregate some data collected by a group of mobile users with sensing devices (a.k.a., data sellers) and sell the corresponding statistics to data consumers, has been recognized as a promising paradigm for large-scale data trading in recent years. It is critical to select sellers with high sensing qualities and maximize all trading participants' profits simultaneously. However, most existing CDT systems either assume that sellers' sensing qualities are known in advance or cannot realize concurrent profit maximization. In this paper, we propose a data trading mechanism based on Combinatorial Multi-Armed Bandit and three-stage Hierarchical Stackelberg game, called CMAB-HS, to tackle the problem of quality unknown seller selection and incentive strategy design. Our objective is to select a group of sellers to maximize the total sensing quality within time budget, and determine the optimal incentive strategy for each participant to maximize individual profit simultaneously. We theoretically prove that CMAB-HS achieves Stackelberg Equilibrium and a tight bound on regret. Additionally, we demonstrate its significant performances through extensive simulations on real data traces.
Baoyi An 0002, Mingjun Xiao, An Liu 0002, Xike Xie, Xiaofang Zhou 0001
ICDE2
2021 On Efficient and Scalable Time-Continuous Spatial Crowdsourcing
abstract
The proliferation of advanced mobile terminals opened up a new crowdsourcing avenue, spatial crowdsourcing, to utilize the crowd potential to perform real-world tasks. In this work, we study a new type of spatial crowdsourcing, called time-continuous spatial crowdsourcing (TCSC in short). It supports broad applications for long-term continuous spatial data acquisition, ranging from environmental monitoring to traffic surveillance in citizen science and crowdsourcing projects. However, due to limited budgets and limited availability of workers in practice, the data collected is often incomplete, incurring data deficiency problem. To tackle that, in this work, we first propose an entropy-based quality metric, which captures the joint effects of incompletion in data acquisition and the imprecision in data interpolation. Based on that, we investigate quality-aware task assignment methods for both single- and multi-task scenarios. We show the NP-hardness of the single-task case, and design polynomial-time algorithms with guaranteed approximation ratios. We study novel indexing and pruning techniques for further enhancing the performance in practice. Then, we extend the solution to multi-task scenarios and devise a parallel framework for speeding up the process of optimization. We conduct extensive experiments on both real and synthetic datasets to show the effectiveness of our proposals.
Xike Xie, Xin Cao 0001, Torben Bach Pedersen, Yang Wang 0015, Mingjun Xiao
ICDE6
2020 Differentially Private Resource Auction in Distributed Spatial Crowdsourcing
Yin Xu 0004, Mingjun Xiao, An Liu 0002
DASFAA (2)2
2020 SRA: Secure Reverse Auction for Task Assignment in Spatial Crowdsourcing
abstract
In this paper, we study a new type of spatial crowdsourcing, namely competitive detour tasking, where workers can make detours from their original travel paths to perform multiple tasks, and each worker is allowed to compete for preferred tasks by strategically claiming his/her detour costs. The objective is to make suitable task assignment by maximizing the social welfare of crowdsourcing systems and protecting workers' private sensitive information. We first model the task assignment problem as a reverse auction process. We formalize the winning bid selection of reverse auction as an n-to-one weighted bipartite graph matching problem with multiple 0-1 knapsack constraints. Since this problem is NP-hard, we design an approximation algorithm to select winning bids and determine corresponding payments. Based on this, a Secure Reverse Auction (SRA) protocol is proposed for this novel spatial crowdsourcing. We analyze the approximation performance of the proposed protocol and prove that it has some desired properties, including truthfulness, individual rationality, computational efficiency, and security. To the best of our knowledge, this is the first theoretically provable secure auction protocol for spatial crowdsourcing systems. In addition, we also conduct extensive simulations on a real trace to verify the performance of the proposed protocol.
Mingjun Xiao, An Liu 0002, Hui Zhao 0003, Zhixu Li, Kai Zheng 0001, Xiaofang Zhou 0001
IEEE Trans. Knowl. Data Eng.1
2019 Truthful Crowdsensed Data Trading Based on Reverse Auction and Blockchain
Baoyi An 0002, Mingjun Xiao, An Liu 0002, Guoju Gao, Hui Zhao 0003
DASFAA (1)2
2019 Reverse-Auction-Based Competitive Order Assignment for Mobile Taxi-Hailing Systems
Hui Zhao 0003, Mingjun Xiao, Jie Wu 0001, An Liu 0002, Baoyi An 0002
DASFAA (2)2
2008 QoS-Aware Scheduling of Web Services
abstract
QoS-aware Web services composition has recently received much attention. While most work focused on service selection, we study QoS in another stage of the life cycle of composite services, namely, scheduling. An interesting problem is whether we can obtain better QoS via scheduling even when the component services have been fixed. In this paper, we propose an approach to find an optimal (near-optimal) schedule with the least cancellation cost, which can further improve the overall QoS of composite services. An approach to analyze the expected cancellation cost of a schedule of a composite service is proposed and QoS-Aware service scheduling is formalized as a Constraint Satisfaction Optimization Problem (CoSOP). Two algorithms - heuristic back tracking and genetic algorithm - are presented to find an optimal (near-optimal) schedule, and their performance is studied by simulations. Preliminary experimental results show that our approach is effective.
An Liu 0002, Qing Li 0001, Liusheng Huang, Mingjun Xiao, Hai Liu 0008
WAIM4
2006 Fault-Tolerant Orchestration of Transactional Web Services
An Liu 0002, Liusheng Huang, Qing Li 0001, Mingjun Xiao
WISE4