Xiaomao Zhou

dblp:207/4747 · DBLP profile ↗
← Back
12ranked-venue papers
10as first author
9since 2021 · last 2026
0009-0004-8608-0845ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 5 first-author · 2 since 2021
YearPublicationVenuePosition
2026 ProxyLLM: Augmenting LLMs With Proxy Models for Tool Utilization in Network Service Generation
abstract
This paper introduces ProxyLLM, a novel framework designed to enhance the tool utilization capabilities of Large Language Models (LLMs) by leveraging an ensemble of smaller, specialized proxy models. Specifically, instead of invoking tools directly, ProxyLLM delegates tasks to these proxy models, each of which is responsible for a distinct domain and equipped with a curated set of relevant tools. Meanwhile, ProxyLLM employs a two-step knowledge transfer mechanism, utilizing data generated by the LLM for knowledge distillation and LLM-guided Deep Reinforcement Learning (DRL) to enhance the decision-making abilities of the proxy models. During the data-driven knowledge distillation process, the introduction of rationales ensures that proxy models maintain a comprehensive understanding of tasks, thereby improving the learning effectiveness. In the DRL learning process, LLM guidance is separately integrated into both the actor and critic learning phases. This ensures consistency in strategy and uniformity in evaluating the action space, which enhances both the efficiency and effectiveness of the learning process. Extensive experiments, including real-world applications such as network service generation in a Computing Power Network (CPN) system, demonstrate that ProxyLLM significantly outperforms existing methods in terms of task accuracy and tool invocation efficiency. The proposed framework offers a promising solution for constructing generalizable, large-scale intelligent agents capable of effectively leveraging diverse tools to solve complex, cross-domain problems.
Xiaomao Zhou, Zihao Shao, Qingmin Jia, Renchao Xie
IEEE Trans. Netw. Serv. Manag.1
2025 Efficient and Adaptive Human Pose Estimation on Resource-Constrained Computing Devices via Knowledge Distillation and Temporal Propagation
abstract
Existing video-based human pose estimation (HPE) methods commonly rely on large networks to localize body joints across all frames, achieving remarkable accuracy but imposing high memory and computational demands that limit their applications on resource-constrained devices. Moreover, most models lack the capability to accommodate dynamic changes in available resources, which can negatively impact the performance of parallel tasks. To address these issues, this article proposes a novel yet effective framework for efficient and adaptive HPE on resource-constrained devices. Specifically, the proposed approach adopts the knowledge distillation (KD) strategy to train a light-weight pose estimator network, which is capable of executing rapidly with low computational cost. To further increase the overall efficiency, it exploits the temporal coherence between successive video frames and explicitly propagates body joints from previous frames rather than naively extracting them using a pose estimator. Furthermore, a prediction-based mechanism is adopted to facilitate adaptive key-frame selection, dynamically determining the optimal number of keyframes, thus enhancing the overall efficiency and adaptability. Experiments on Penn Action, Sub-JHMDB, and real-world systems demonstrate that the proposed method achieves comparative accuracy, superior efficiency, and robust flexibility in dynamic scenarios.
Xiaomao Zhou, Yujiao Hu, Qingmin Jia, Renchao Xie
IEEE Internet Things J.1
2025 NestFL: Enhancing Federated Learning Through Nested Multicapacity Model Pruning in Heterogeneous Edge Computing
abstract
Federated learning (FL) has emerged as a pivotal approach for edge-based distributed machine learning, yet it faces significant challenges due to the constrained capacities and heterogeneity of edge devices, including non-IID data distribution, communication constraints, and learning inefficiencies. Furthermore, a one-fits-all global model often fails to perform optimally across diverse participating devices. In this paper, we present NestFL, an efficient FL framework for edge computing that can jointly improve the training efficiency and achieve personalization. Specifically, NestFL innovates by incorporating distributed model pruning, creating a hierarchy of structured-sparse subnetworks tailored to the unique resource profiles of client devices. These subnetworks are integrated into a nested global model, ensuring parameter sharing without increasing the parameter space, thereby significantly reducing computational and communication burdens. Meanwhile, it implements a cross-training mechanism, allowing clients to train on a broader dataset and maintain consistent decision boundaries. Furthermore, a weighted aggregation mechanism is designed to improve training performance and maximally preserve personalization. Experimental results in different applications demonstrate the superiority of NestFL over the baseline approaches in terms of model accuracy, convergence speed, and personalization preservation.
Xiaomao Zhou, Yujiao Hu, Qingmin Jia, Renchao Xie
IEEE Internet Things J.1
2025 AdaHPE: Adaptive Human Pose Estimation on Resource-Constrained Edge Computing Devices via Temporal Propagation
abstract
This paper presents AdaHPE, an innovative and efficient framework for human pose estimation (HPE) designed specifically for edge computing devices with constrained and fluctuating resources. AdaHPE redefines the conventional HPE workflow by converting the resource-demanding pose regression into a sequence of computationally feasible pose propagation tasks. The framework incorporates a memory-augmented LSTM network with a global memory repository, allowing AdaHPE to adaptively choose keyframes based on real-time data and the device’s resource status, thereby optimizing the trade-off between accuracy and computational efficiency. A reinforcement learning component is further integrated to intelligently adjust the ratio of keyframes used, enhancing the framework’s adaptability. Utilizing policy gradient algorithms, AdaHPE is optimized to maximize a reward function that encourages both accurate and resource-efficient pose estimations, while respecting a given keyframe constraint. Extensive experiments on benchmarks including Penn Action, Sub-JHMDB, NTU RGB+D 120, and real-world datasets demonstrate that AdaHPE can significantly reduce computational overhead compared to per-frame HPE models while preserving high accuracy and robustness under varying resource limitations. Moreover, the seamless compatibility of our approach with various off-the-shelf HPE models highlights its versatility and potential for broad applications.
Xiaomao Zhou, Yujiao Hu, Qingmin Jia, Renchao Xie
IEEE Internet Things J.1
2024 Dynamic Staleness Control for Asynchronous Federated Learning in Decentralized Topology
Qianpiao Ma, Jianchun Liu, Qingmin Jia, Xiaomao Zhou, Yujiao Hu, Renchao Xie
WASA (2)4
2024 Industrial Internet of Things Intelligence Empowering Smart Manufacturing: A Literature Review
abstract
The fiercely competitive business environment and increasingly personalized customization needs are driving the digital transformation and upgrading of the manufacturing industry. IIoT intelligence, which can provide innovative and efficient solutions for various aspects of the manufacturing value chain, illuminates the path of transformation for the manufacturing industry. It’s time to provide a systematic vision of IIoT intelligence. However, existing surveys often focus on specific areas of IIoT intelligence, leading researchers and readers to have biases in their understanding of IIoT intelligence, that is, believing that research in one direction is the most important for the development of IIoT intelligence, while ignoring contributions from other directions. Therefore, this paper provides a comprehensive overview of IIoT intelligence. We first conduct an in-depth analysis of the inevitability of manufacturing transformation and study the successful experiences from the practices of Chinese enterprises. Then we give our definition of IIoT intelligence and demonstrate the value of IIoT intelligence for industries in fucntions, operations, deployments, and application. Afterwards, we propose a hierarchical development architecture for IIoT intelligence, which consists of five layers. The practical values of technical upgrades at each layer are illustrated by a close look on lighthouse factories. Following that, we identify seven kinds of technologies that accelerate the transformation of manufacturing, and clarify their contributions. The ethical implications and environmental impacts of adopting IIoT intelligence in manufacturing are analyzed as well. Finally, we explore the open challenges and development trends from four aspects to inspire future researches.
Yujiao Hu, Qingmin Jia, Yuan Yao 0004, Mengjie Lee, Xiaomao Zhou, Renchao Xie, F. Richard Yu
IEEE Internet Things J.7
2022 Fast and Accurate Pose Estimation in Videos based on Knowledge Distillation and Pose Propagation
abstract
Existing video-based human pose estimation methods typically adopt large networks to perform body joints localization on all frames. Despite of impressive accuracy performance, the relatively high memory and computation requirements significantly burden their applicability on resource-constraint systems (e.g., embedded devices). To solve this issue, this paper proposes a novel yet effective lightweight framework, called FVPE, for fast and accurate human pose estimation in videos. Specifically, FVPE adopts the knowledge distillation (KD) strategy to train a small pose estimator network, which is capable of executing rapidly with low computational cost. To increase the overall efficiency, FVPE exploits the temporal coherence between successive video frames and explicitly propagates body joints from previous frames rather than naively extracting them using a pose estimator. Furthermore, FVPE introduces an online key-frame selection scheme to decide whether the current pose should be calculated by the pose estimator or be propagated from the previous key-frame, being able to flexibly deal with video sequences with different length, frame rate, pose complexity, etc.. Experiments on Penn Action and Sub-JHMDB datasets demonstrate that the proposed method achieves comparative accuracy, but with substantial speed-up.
Xiaomao Zhou
IJCNN1
2022 NestFL: efficient federated learning through progressive model pruning in heterogeneous edge computing
abstract
In this paper, we present NestFL, a learning-efficient FL framework for edge computing, which can jointly improve the training efficiency and achieve personalization. Specifically, NestFL takes the runtime resources of the edge devices into consideration and assigns each device a sparse-structured subnetwork by progressively performing the structured pruning. During training, only the updates of these subnetworks are transmitted to the central server. Additionally, these generated subnetworks adopt a structure- and parameter-sharing mechanism, making themselves nested inside a multi-capacity global model. In doing so, the overall communication and computation costs can be significantly reduced, and each device can learn a personalized model without introducing extra parameters. Furthermore, a weighted aggregation mechanism is designed to improve the training performance and maximally preserve personalization.
Xiaomao Zhou, Qingmin Jia, Renchao Xie
MobiCom1
2021 PSG-GAN: Progressive Person Image Generation with Self-Guided Local Focuses
abstract
This paper proposes PSG-GAN, a novel Generative Adversarial Network for pose-guided person image synthesis, which can progressively generate realistic person images of desired poses together with corresponding semantic segmentation masks. Specifically, PSG-GAN consists of a sequence of Region-Focal Transfer Blocks (RFBs) where each contains two generation pathways: the appearance generation pathway and the semantic generation pathway. The former pathway is responsible for generating the target image by explicitly preserving appearance-related features within certain regions, where local region transformations are considered. The latter pathway is used to generate semantic masks which define the areas for the local transformations to attend to. These two learning pathways work together and reinforce each other to simultaneously generate the target image and semantic masks progressively. Qualitative and quantitative experimental results on two benchmark datasets demonstrate PSG-GAN’s superiority over other approaches in generating realistic person images in pose transfer tasks.
Xiaomao Zhou
ICTAI1
2018 A Hybrid Planning Strategy Through Learning from Vision for Target-Directed Navigation
Xiaomao Zhou, Cornelius Weber, Chandrakant Bothe, Stefan Wermter
ICANN (2)1
2018 A Self-organizing Method for Robot Navigation based on Learned Place and Head-Direction Cells
abstract
This paper describes a neural model for a robot learning spatial knowledge and navigating on learned place and head-direction (HD) cell representations. The place and HD cells, which are trained through unsupervised slow feature analysis (SFA) from sequences of visual stimuli, provide positional and directional information for navigation. Based on the ensemble activity of place cells, the robot learns a topological map of the environment through extracting the statistical distribution of the place cell activities covering the traversable areas and realizes self-localization based on the map. The robot's heading direction, which is encoded by the HD cells, works as a control signal to adjust its behavior. Action representations supporting state transitions are learned through memorizing the same movement from a previous phase where an experimenter drives a robot to explore an environment. Given reward signals spreading from a target location along the topological map, the robot can reach the goal in a reward-ascending way. This work intends to build a practical navigation system by simulating animals' hippocampal cell firing activities on a robot platform using its self-contained sensor. Experimental results from simulation demonstrate that our system navigates a robot to the desired position smoothly and effectively.
Xiaomao Zhou, Cornelius Weber, Stefan Wermter
IJCNN1
2017 Robot Localization and Orientation Detection Based on Place Cells and Head-Direction Cells
Xiaomao Zhou, Cornelius Weber, Stefan Wermter
ICANN (1)1