Ningxin Su

dblp:324/4033 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0003-0132-7585ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 5 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Square: Towards Sharing Internet-Scale GPU Clouds Fairly and Efficiently
Ningxin Su, Freeman Cheng, Baochun Li
IWQoS1
2026 State of the Union: Toward Reproducible Performance Evaluations in Federated Learning
abstract
Federated learning (FL) is a privacy-motivated paradigm for distributed training of deep learning models, as it allows a large number of clients to collaboratively train a shared global model without centralizing their private data. Since its debut with Federated Averaging as the first server aggregation algorithm, many FL algorithms have been proposed to improve performance. Yet, these algorithms are rarely benchmarked and compared under the same open-source framework and controlled configurations, and their performance claims can be difficult to substantiate in fair and reproducible studies. In this paper, we evaluate a curated and representative collection of FL algorithms in the same open-source benchmarking framework so that they can be compared fairly at scale in a reproducible fashion. To achieve this objective, we presentPlato, an open-source FL research framework that we have designed and implemented from scratch. WithPlato, we evaluate and compare algorithms spanning (1) server aggregation; (2) client training customization; (3) client selection, in both synchronous and asynchronous settings; (4) personalized federated learning; and (5) communication efficiency and payload processing. Across diverse experimental scenarios (tasks, client populations, and data distributions), we report findings and practical insights, including pitfalls and confounding factors that can lead to misleading conclusions if not reported carefully. Under our unified experimental settings and time model, Federated Averaging with random client selection remains a strong baseline and is often competitive with more complex alternatives.
Ningxin Su, Baochun Li, Bo Li 0001
IEEE Trans. Knowl. Data Eng.1
2024 Pack: Towards Communication-Efficient Homomorphic Encryption in Federated Learning
abstract
Federated learning allows multiple clients to collaboratively train a shared model without sharing local private data. It is regarded as privacy-preserving since only model updates are communicated. Unfortunately, it has been shown in the recent literature that, model updates transmitted by participating clients can be used by a malicious server in gradient leakage attacks to obtain private training data. To prevent such potential leakage from occurring, it has widely been acknowledged that homomorphic encryption can be used to encrypt these model updates before sending them to the server, which performs computations directly on encrypted data. Although homomorphic encryption has a strong guarantee on privacy, its practical use increases communication overhead by around 17×, even with its most efficient implementation, called CKKS. In this paper, we present Pack, a novel communication-efficient mechanism over CKKS, designed specifically to reduce the communication overhead by a substantial margin. In addition, we propose new error correction and weight filtering mechanisms in Pack to improve the accuracy of the trained model. Compared to vanilla CKKS, Pack reduces the communication overhead by 3.1×, while increasing the accuracy by 5.5% and 2.5% under the i.i.d. and non-i.i.d. settings.
Zeyuan Zuo, Ningxin Su, Baochun Li
SoCC2
2024 Calibre: Towards Fair and Accurate Personalized Federated Learning with Self-Supervised Learning
abstract
In the context of personalized federated learning, existing approaches train a global model to extract transferable representations, based on which any client could train personalized models with a limited number of data samples. Self-supervised learning is considered a promising direction as the global model it produces is generic and facilitates personalization for all clients fairly. However, when data is heterogeneous across clients, the global model trained using SSL is unable to learn high-quality personalized models. In this paper, we show that when the global model is trained with SSL without modifications, its produced representations have fuzzy class boundaries. As a result, personalized learning within each client produces models with low accuracy. In order to improve SSL towards better accuracy without sacrificing its advantage in fairness, we propose Calibre, a new personalized federated learning framework designed to calibrate SSL representations by maintaining a suitable balance between more generic and more client-specific representations. Calibre is designed based on theoretically-sound properties, and introduces (1) a client-specific prototype loss as an auxiliary training objective; and (2) an aggregation algorithm guided by such prototypes across clients. Our experimental results in an extensive array of non-i.i.d. settings show that Calibre achieves state-of-the-art performance in terms of both mean accuracy and fairness across clients.
Ningxin Su, Baochun Li
ICDCS2
2024 Titanic: Towards Production Federated Learning with Large Language Models
abstract
With the recent surge of research interests in Large Language Models (LLMs), a natural question that arises is how pre-trained LLMs can be fine-tuned to tailor to specific needs of enterprises and individual users, while preserving the privacy of data used in the fine-tuning process. On the one hand, sending private data to cloud datacenters for fine-tuning is, without a doubt, unacceptable from a privacy perspective. On the other hand, conventional federated learning requires each client to perform local training, which is not feasible for LLMs with respect to both computation costs and communication overhead, since they involve billions of model parameters. In this paper, we present Titanic, a new distributed training paradigm that allows LLMs to be fine-tuned in a privacy-preserving fashion directly on the client devices where private data is produced, while operating within the resource constraints on computation and communication bandwidth. Titanic first optimally selects a subset of clients with an efficient solution to an integer optimization problem, then partitions an LLM across multiple client devices, and finally fine-tunes the model with no or minimal losses in training performance. A primary focus in the design of Titanic is its feasibility in real-world systems: it is first and foremost designed for production-quality systems, featuring a model-agnostic partitioning mechanism that is fully automated. Our experimental results show that Titanic achieves superior training performance as compared to conventional federated learning, while preserving data privacy and satisfying all constraints on local computation and bandwidth resources.
Ningxin Su, Chenghao Hu, Baochun Li, Bo Li 0001
INFOCOM1
2024 Relic: Federated Conditional Textual Inversion with Prototype Alignment
abstract
Text-to-image models can generate personalized images with unprecedented freedom by using a pseudo-word learned from a few images, using a novel technique called textual inversion. It is conceivable that, in the spirit of federated learning, multiple users wish to learn a pseudo-word based on their local images collaboratively. However, how a more effective pseudo-word can be trained in the context of federated learning remains unclear.In this paper, our experiments show that such federated textual inversion is neither secure nor feasible. First, once one client exposes its pseudo-word embedding to the server for aggregation, an attacker can directly generate similar images to this client. Second, training one shared pseudo-word without personalization hinders individuals from generating images that exhibit local characteristics. Finally, after global aggregation, the averaged pseudo-word embedding may lose learned concepts. Motivated by these insights, we propose Relic, a new framework that encompasses federated conditional textual inversion with prototype alignment. With privacy guarantees, Relic allows clients to learn personalized pseudo-words conditional on local samples while enforcing a globally consistent clustering of clients’ pseudo-words into discriminable prototypes instead of averaging. The experiments conducted on both i.i.d. and extreme non-i.i.d. data demonstrate that Relic is able to achieve state-of-the-art performance as compared to baseline approaches.
Ningxin Su, Baochun Li
IWQoS2
2024 MLOps in the Metaverse: Human-Centric Continuous Integration
abstract
The metaverse is a virtual world that exists entirely in a computer-generated environment, and it offers a new frontier for machine learning. One of the major challenges for using machine learning in the metaverse is MLOps (Machine Learning Operations), an emerging field that focuses on deploying and managing machine learning models in production. It has been widely acknowledged that machine learning models require a large amount of data to learn and make accurate predictions, and such data is generated progressively in real-time as human users interact with the metaverse. Due to the human-centric nature of the metaverse, it goes without saying that, once deployed, models need to be able to adapt to the constantly changing interactive environment and still make accurate predictions. Borrowing a page from software engineering, in this paper, we explore the design space of human-centric continuous integration in metaverse environments, where labeled data samples accumulated with explicit human interactive behavior (e.g., using virtual reality or augmented reality headsets) are used for fine-tuning a deployed deep learning model over a sustained period of time. We propose SPIN, a new mechanism that efficiently utilizes data samples collected from a large number of participating human users over time to fine-tune a deployed model that is shared across all the users. In an extensive array of experimental results using image classification and state-of-the-artYOLOv8object detection models as case studies, we show that SPIN outperforms FedBuff, a state-of-the-art asynchronous FL mechanism from conventional federated learning, by a substantial margin.
Ningxin Su, Baochun Li
IEEE J. Sel. Areas Commun.1
2023 Asynchronous Federated Unlearning
abstract
Thanks to regulatory policies such as the General Data Protection Regulation (GDPR), it is essential to provide users with the right to erasure regarding their own private data, even if such data has been used to train a neural network model. Such a machine unlearning problem becomes even more challenging in the context of federated learning, where clients collaborate to train a global model with their private data. When a client requests its data to be erased, its effects have already gradually permeated through a large number of clients, as the server aggregates client updates over multiple communication rounds. All of these affected clients need to participate in the retraining process, leading to prohibitive retraining costs with respect to the wall-clock training time.In this paper, we present the design and implementation of Knot, a new clustered aggregation mechanism custom-tailored to asynchronous federated learning. The design of Knot is based upon our intuition that, with asynchronous federated learning, clients can be divided into clusters, and aggregation can be performed within each cluster only so that retraining due to data erasure can be limited to within each cluster as well. To optimize client-cluster assignment, we formulated a lexicographical minimization problem that could be transformed into a linear programming problem and solved efficiently. Over a variety of datasets and tasks, we have shown clear evidence that Knot outperformed the state-of-the-art federated unlearning mechanisms by up to 85% in the context of asynchronous federated learning.
Ningxin Su, Baochun Li
INFOCOM1
2022 How Asynchronous can Federated Learning Be?
abstract
As a practical paradigm designed to involve large numbers of edge devices in distributed training of deep learning models, federated learning has witnessed a significant amount of research attention in the recent years. Yet, most existing mechanisms on federated learning assumed either fully synchronous or asynchronous communication strategies between clients and the federated learning server. Existing designs that were partially asynchronous in their communication were simple heuristics, and were evaluated using the number of communication rounds or updates required for convergence, rather than the wall-clock time in practice.In this paper, we seek to explore the entire design space between fully synchronous and asynchronous mechanisms of communication. Based on insights from our exploration, we propose Port, a new partially asynchronous mechanism designed to allow fast clients to aggregate asynchronously, yet without waiting excessively for the slower ones. In addition, Port is designed to adjust the aggregation weights based on both the staleness and divergence of model updates, with provable convergence guarantees. We have implemented Port and its leading competitors in Plato, an open-source scalable federated learning research framework designed from the ground up to emulate real-world scenarios. With respect to the wall-clock time it takes for converging to the target accuracy, Port outperformed its closest competitor, FedBuff, by up to 40% in our experiments.
Ningxin Su, Baochun Li
IWQoS1