EDBT 2026 Demo / reviewers in the wild / expert
Zhen-guo Ma
dblp:53/9555 · also Zhen-Guo Ma, Zhenguo Ma
· DBLP profile ↗
15ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 4 first-author · 9 since 2021Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Caesar: Optimizing Federated Learning via Low-deviation CompressionabstractCompression is an efficient way to relieve the tremendous communication overhead of federated learning (FL) systems. However, for the existing works, the information loss under compression will lead to unexpected model/gradient deviation for the FL training, significantly degrading the training performance, especially under the challenges of data heterogeneity and model obsolescence. To strike a delicate trade-off between model accuracy and traffic cost, we propose Caesar, a novel FL framework with a low-deviation compression approach. For the global model download, we design a greedy method to optimize the compression ratio for each device based on the staleness of the local model, ensuring a precise initial model for local training. Regarding the local gradient upload, we utilize the device's local data properties (i.e., sample volume and label distribution) to quantify its local gradient's importance, which then guides the determination of the gradient compression ratio. We have implemented Caesar, on two physical platforms with 40 smartphones and 80 NVIDIA Jetson devices. Extensive results show that Caesar, can reduce the traffic costs by about 25.54%þicksim37.88% when achieving the same target accuracy compared to the compression-based baselines, while incurring only a 0.68% degradation in final test accuracy relative to the full-precision communication. Jiaming Yan, Jianchun Liu, Hongli Xu 0001, Zhen-guo Ma, Shilong Wang 0002 |
KDD (1) | 4 |
| 2026 | Enhancing federated unlearning using catastrophic forgetting in heterogeneous Industrial Internet of Things
Zhen-guo Ma, Yanjing Sun, Hongli Xu 0001, Jianchun Liu, Yang Xu 0020, Yafei Lu |
Comput. Commun. | 1 |
| 2025 | Diffusion Model-based Graph Reinforcement Learning for Task Offloading in Computation Reuse-Enabled Industrial Edge NetworksabstractCollaborative edge computing (CEC) enables multiple edge servers (ESs) to cooperatively process computation-intensive tasks generated by resource-constrained industrial internet of things devices. A challenge arises when offloading similar tasks from devices to ESs triggers duplicate computation, resulting in severe resource waste. To address this, computation reuse has recently been proposed to reduce response delay and energy consumption by caching task results at the edge. However, existing studies either overlook the time validity of task results or ignore the result retrieval cost, resulting in decreased practicality in real-world scenarios. In this paper, we propose a computation reuse-enabled CEC framework, where reusable task results are cached in edge networks and can be accessed if they are within the validity and the similarity of task inputs exceeds the preset thresholds. To minimize the long-term system cost comprising weighted task delay and energy consumption under the resource constraints, we formulate a joint task offloading, resource allocation, result retrieval and result caching (T3R) problem. Then, we propose a diffusion model-based graph reinforcement learning (DMGRL) algorithm to fully exploit the graph structural information of the edge network and optimize T3R policy. Extensive experimental results demonstrate that compared to baseline algorithms, the DMGRL algorithm achieves 37.4% maximum cost reduction and faster convergence speed. Yanjing Sun, Beibei Zhang 0001, Zhen-guo Ma, Bowen Wang 0004, Song Li 0001 |
GLOBECOM | 4 |
| 2025 | Hier-FUN: Hierarchical Federated Learning and Unlearning in Heterogeneous Edge ComputingabstractFederated learning (FL) has emerged as a pivotal paradigm for distributed model training in edge computing (EC), enabling cooperation among numerous Internet of Things devices while safeguarding their data privacy. Despite its successes in machine learning, concerns regarding data security and model fidelity necessitate the efficient unlearning of target device, i.e., federated unlearning (FUN). However, due to resource constraints, device heterogeneity, and non-independent and identically distributed (Non-IID) data, securely eliminating a device’s impact without retraining the model from scratch presents a complex challenge. In response to these challenges, we propose a hierarchical FUN framework, called Hier-FUN. Hier-FUN organizes edge devices into K clusters, each managed by a head device responsible for aggregating local models within the cluster. To expedite both the learning and unlearning processes of Hier-FUN, we design a heuristic algorithm to determine an appropriate value for K based on devices’ data distributions and available resources. In addition, Hier-FUN denies the communication between the server and cluster heads during training, which can constrain the influence sphere of target device and accelerate the unlearning process. We conduct extensive experiments using real-world datasets, and the experimental results illustrate that Hier-FUN can improve test accuracy by 3.19% during the learning phase and achieve a$6.8\times $speedup during unlearning compared with the baseline methods. Zhen-guo Ma, Huaqing Tu, Pengli Ji, Xiaoran Yan, Hongli Xu 0001, Zhiyuan Wang 0002, Suo Chen |
IEEE Internet Things J. | 1 |
| 2025 | Area-Efficient Pipeline Architecture for Serial Real-Valued Fast Fourier TransformabstractThis brief presents a novel pipeline architecture designed to compute the fast Fourier transform (FFT) on real input signals in a serial format. This architecture significantly improves resource efficiency by sharing adders between butterfly and rotator structures. In addition, a novel data management approach for N-point radix-2 serial real-valued FFT (RFFT) has been proposed, which not only simplifies the data reordering circuit between processing elements (PEs) but also achieves natural order data output. The real-valued 1024-point FFT has been implemented on a field-programmable gate array (FPGA). Compared with typical real-valued serial commutator (RSC) FFT architecture, the proposed architecture achieves substantial improvement, including a reduction of 10.3% in the number of lookup tables (LUTs) and 12.5% in flip-flops (FFs). Kun Li 0032, Hongji Fang, Zhen-guo Ma, Feng Yu 0003, Bo Zhang 0097, Qianjian Xing |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | A Fast Floating-Point Multiply-Accumulator Optimized for Sparse Linear Algebra on FPGAsabstractThis brief presents a pipelined floating-point Multiply–Accumulator (FPMAC) architecture designed to accelerate sparse linear algebra operations. By designing a lookup-table-based 5–3 carry-save adder (CSA) and combining it with a 3–2 CSA, the proposed design minimizes the critical path and boosts operational speed. Moreover, the proposed architecture takes advantage of data characteristics in sparse linear algebra to displace the shift unit in the critical accumulation loop, further increasing the throughput rate. In addition, the integration of a lookup-table-based leading-zero anticipator (LZA) enhances normalization efficiency. Experimental results show that, compared with reported FPMAC designs, the proposed architecture may achieve a significantly higher maximum clock frequency for single-precision floating-point operations. Kun Li 0032, Xiangyu Hao, Zhen-guo Ma, Feng Yu 0003, Bo Zhang 0097, Qianjian Xing |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | Enhancing Decentralized and Personalized Federated Learning With Topology ConstructionabstractThe emerging Federated Learning (FL) permits all workers (e.g., mobile devices) to cooperatively train a model using their local data at the network edge. In order to avoid the possible bottleneck of conventional parameter server architecture, the decentralized federated learning (DFL) is developed on the peer-to-peer (P2P) communication. Non-IID issue is a key challenge in FL and will significantly degrade the model training performance. To this end, we propose a personalized solution called TOPFL, in which only parts of the local models (not the entire models) are shared and aggregated. Moreover, considering the limited communication bandwidth on workers, we propose a topology construction algorithm to accelerate the training process. To verify the convergence of the decentralized training framework, we theoretically analyze the impact of the data heterogeneity and topology on the convergence upper bound. Extensive simulation results show that TOPFL can achieve 2.2× speedup when reaching convergence and 5.8% higher test accuracy under the same resource consumption, compared with the baseline solutions. Suo Chen, Yang Xu 0020, Hongli Xu 0001, Zhen-guo Ma, Zhiyuan Wang 0002 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Like Attracts Like: Personalized Federated Learning in Decentralized Edge ComputingabstractThe emerging Personalized Federated Learning (PFL) methods aim to produce personalized models for different users, so as to keep track of their individualized requirements in Edge Computing (EC). The centralized PFL methods may suffer from the communication bottleneck and single point of failure. As an alternative solution, the decentralized PFL (DPFL) methods are performed in a Peer-to-Peer (P2P) manner, and collaboratively train personalized models by model aggregation among congenial devices. However, these DPFL methods may incur high communication cost and low resource utilization induced by large-scale models. Herein, we take the communication constraint and heterogeneity into consideration and propose to realize communication-efficient DPFL with adaptive model pruning and neighbor selection. We theoretically analyze the convergence of the proposed DPFL method, and study the impacts of both model pruning and neighbor selection on training performance. Furthermore, we propose an efficient algorithm that combines model pruning and neighbor selection to achieve a trade-off between model quality and communication cost. Extensive simulation and testbed experiments on real-world datasets are conducted. The experimental results demonstrate that the proposed algorithm can improve the test accuracy by at most 13% and save the traffic consumption by 45.4% on average compared with the existing PFL methods. Zhen-guo Ma, Yang Xu 0020, Hongli Xu 0001, Jianchun Liu, Yinxing Xue |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | FedLC: Accelerating Asynchronous Federated Learning in Edge ComputingabstractFederated Learning (FL) has been widely adopted to process the enormous data in the application scenarios like Edge Computing (EC). However, the commonly-used synchronous mechanism in FL may incur unacceptable waiting time for heterogeneous devices, leading to a great strain on the devices' constrained resources. In addition, the alternative asynchronous FL is known to suffer from the model staleness, which will lead to performance degradation of the trained model, especially onnon-i.i.d.data. In this paper, we design a novel asynchronous FL mechanism, named FedLC, to handle thenon-i.i.d.issue in EC by enabling the local collaboration among edge devices. Specifically, apart from uploading the local model directly to the server, each device will transmit its gradient to the other devices with different data distributions for local collaboration, which can improve the model generality. We theoretically analyze the convergence rate of FedLC and obtain the quantitative relationship between convergence bound and local collaboration. We design an efficient algorithm utilizing demand-list to determine the set of devices receiving gradients from each device. To handle the model staleness, we further assign different learning rates for various devices according to their participation frequency. The extensive experimental results demonstrate the effectiveness of our proposed mechanism. Yang Xu 0020, Zhen-guo Ma, Hongli Xu 0001, Suo Chen, Jianchun Liu, Yinxing Xue |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | Adaptive Batch Size for Federated Learning in Resource-Constrained Edge ComputingabstractThe emerging Federated Learning (FL) enables IoT devices to collaboratively learn a shared model based on their local datasets. However, due to end devices’ heterogeneity, it will magnify the inherent synchronization barrier issue of FL and result in non-negligible waiting time when local models are trained with the identical batch size. Moreover, the useless waiting time will further lead to a great strain on devices’ limited battery life. Herein, we aim to alleviate the negative impact of synchronization barrier through adaptive batch size during model training. When using different batch sizes, stability and convergence of the global model should be enforced by assigning appropriate learning rates on different devices. Therefore, we first study the relationship between batch size and learning rate, and formulate a scaling rule to guide the setting of learning rate in terms of batch size. Then we theoretically analyze the convergence rate of global model and obtain a convergence upper bound. On these bases, we propose an efficient algorithm that adaptively adjusts batch size with scaled learning rate for heterogeneous devices to reduce the waiting time and save battery life. We conduct extensive simulations and testbed experiments, and the experimental results demonstrate the effectiveness of our method. Zhen-guo Ma, Yang Xu 0020, Hongli Xu 0001, Zeyu Meng, Liusheng Huang, Yinxing Xue |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | Adaptive Control of Local Updating and Model Compression for Efficient Federated LearningabstractData generated at the network edge can be processed locally by leveraging the paradigm of Edge Computing (EC). Aided by EC, Federated Learning (FL) has been becoming a practical and popular approach for distributed machine learning over locally distributed data. However, FL faces three critical challenges, i.e., resource constraint, system heterogeneity and context dynamics in EC. To address these challenges, we present a training-efficient FL method, termedFedLamp, by optimizing both theLocal updating frequency andmodel compression ratio in the resource-constrained EC systems. We theoretically analyze the model convergence rate and obtain a convergence upper bound related to the local updating frequency and model compression ratio. Upon the convergence bound, we propose a control algorithm, that adaptively determines diverse and appropriate local updating frequencies and model compression ratios for different edge nodes, so as to reduce the waiting time and enhance the training efficiency. We evaluate the performance ofFedLampthrough extensive simulation and testbed experiments. Evaluation results show thatFedLampcan reduce the traffic consumption by 63% and the completion time by about 52% for achieving the similar test accuracy, compared to the baselines. Yang Xu 0020, Yunming Liao, Hongli Xu 0001, Zhen-guo Ma, Lun Wang 0003, Jianchun Liu |
IEEE Trans. Mob. Comput. | 4 |
| 2021 | Joint Network Selection and Task Offloading in Mobile Edge ComputingabstractAs some delay-sensitive mobile services such as augmented reality and autonomous driving proliferate, users' demand for low latency access to computation resources increases dramatically, and existing centralized cloud computing paradigm is difficult to solve the current dilemma. As a emerging computing paradigm in which computational capabilities are pushed from the central cloud to the network edges, Mobile Edge Computing (MEC) is expected to be an effective solution. However, due to the limited capacity (e.g. computation and bandwidth) of MEC nodes, it is not easy to maintain satisfactory quality of service for user applications. Most of the previous work is limited to reducing the processing delay by dynamically adjusting the task offloading strategy, while ignoring the key impact of access network selection on network congestion. To fill this gap, we study the joint optimization of network selection and task offloading in MEC networks with multidimensional resources constraints. To address a number of key challenges in MEC systems, including spatial demand coupling and decentralized coordination, we propose an efficient online algorithm and achieve provable close-to-optimal performance. Extensive simulation results are presented to verify the performance of our algorithm. Hongli Xu 0001, Zhen-guo Ma, Suo Chen |
CCGRID | 3 |
| 2021 | Communication-efficient asynchronous federated learning in resource-constrained edge computing
Jianchun Liu, Hongli Xu 0001, Yang Xu 0020, Zhen-guo Ma, Zhiyuan Wang 0002, Chen Qian 0001, He Huang 0001 |
Comput. Networks | 4 |
| 2020 | MFCFSiam: A Correlation-Filter-Guided Siamese Network with Multifeature for Visual TrackingabstractWith the development of deep learning, trackers based on convolutional neural networks (CNNs) have made significant achievements in visual tracking over the years. The fully connected Siamese network (SiamFC) is a typical representation of those trackers. SiamFC designs a two-branch architecture of a CNN and models’ visual tracking as a general similarity-learning problem. However, the feature maps it uses for visual tracking are only from the last layer of the CNN. Those features contain high-level semantic information but lack sufficiently detailed texture information. This means that the SiamFC tracker tends to drift when there are other same-category objects or when the contrast between the target and the background is very low. Focusing on addressing this problem, we design a novel tracking algorithm that combines a correlation filter tracker and the SiamFC tracker into one framework. In this framework, the correlation filter tracker can use the Histograms of Oriented Gradients (HOG) and color name (CN) features to guide the SiamFC tracker. This framework also contains an evaluation criterion which we design to evaluate the tracking result of the two trackers. If this criterion finds the SiamFC tracker fails in some cases, our framework will use the tracking result from the correlation filter tracker to correct the SiamFC. In this way, the defects of SiamFC’s high-level semantic features are remedied by the HOG and CN features. So, our algorithm provides a framework which combines two trackers together and makes them complement each other in visual tracking. And to the best of our knowledge, our algorithm is also the first one which designs an evaluation criterion using correlation filter and zero padding to evaluate the tracking result. Comprehensive experiments are conducted on the Online Tracking Benchmark (OTB), Temple Color (TC128), Benchmark for UAV Tracking (UAV-123), and Visual Object Tracking (VOT) Benchmark. The results show that our algorithm achieves quite a competitive performance when compared with the baseline tracker and several other state-of-the-art trackers. Chenpu Li, Qianjian Xing, Zhen-guo Ma, Ke Zang |
Wirel. Commun. Mob. Comput. | 3 |
| 2011 | An efficient radix-2 fast Fourier transform processor with ganged butterfly engines on field programmable gate arraysabstractWe present a novel method to implement the radix-2 fast Fourier transform (FFT) algorithm on field programmable gate arrays (FPGA). The FFT architecture exploits parallelism by having more pipelined units in the stages, and more parallel units within a stage. It has the noticeable advantages of high speed and more efficient resource utilization by employing four ganged butterfly engines (GBEs), and can be well matched to the placement of the resources on the FPGA. We adopt the decimation-infrequency (DIF) radix-2 FFT algorithm and implement the FFT processor on a state-of-the-art FPGA. Experimental results show that the processor can compute 1024-point complex radix-2 FFT in about 11 µs with a clock frequency of 200 MHz. Zhen-guo Ma, Feng Yu 0003, Ruifeng Ge, Zeke Wang |
J. Zhejiang Univ. Sci. C | 1 |