EDBT 2026 Demo / reviewers in the wild / expert
Junxu Xia
dblp:229/7749
· DBLP profile ↗
19ranked-venue papers
11as first author
13since 2021 · last 2026
0000-0002-9522-3025ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 first-author · 6 since 2021Computer networks · 9 · 6 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Breaking Bucket Effect in In-Network Aggregation via Memory-Bandwidth Coordination
Junxu Xia, Geyao Cheng, Deke Guo, Lailong Luo, Wenfei Wu |
INFOCOM | 1 |
| 2026 | When Server Joins INA: The Resource-Aware Repair Acceleration for Erasure-Coded Storage Systems
Geyao Cheng, Junxu Xia, Hao Fan 0006, Fengzeng Liu, Haibo Mi, Deke Guo |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | SINA: A Server-Assisted In-Network Repair Acceleration for Erasure-Coded Storage SystemsabstractIn the erasure-coded storage systems, multiple related blocks have to be retrieved from other surviving nodes to repair a failed block. This incurs significant communication overhead with the surging scale of distributed storage systems. To mitigate the bandwidth bottleneck, in-network repair (INR) has emerged as a promising transport paradigm, which migrates the aggregation operations from the repair node to the programmable hardware, such as Intel Tofino switches. However, due to the limited on-chip memory size of these switches, the INR can degrade to the most primitive incast-type transmission, leading to massive traffic volume and hindered repair throughput. While we notice that, there are spare CPU cores in the storage servers that can be leveraged as alternative computing resources. With this intuition, we propose SINA, a Server-assisted In-Network repair Acceleration framework in this paper, which leverages the spare servers to assist aggregation operations when the programming switches' memory size is scarce for failure repair. We formulate this problem by adjusting the aggregation modes across the involved racks and solve this NP-hard problem using the Gurobi optimization solver. For all we know, this is the first work exploring spare servers for assisting the memory-scarce INR in erasure-coded storage systems. We have implemented SINA on an FPGA-based prototype system, and the experimental results show that SINA can ensure fault tolerance and accelerate failure repair by$5.0 \times$compared to the conventional methods. Geyao Cheng, Junxu Xia, Haibo Mi, Deke Guo, Kun Wang 0059 |
IWQoS | 2 |
| 2025 | HyperPart: A Hypergraph-Based Abstraction for Deduplicated Storage SystemsabstractCurrently, deduplication techniques are utilized to minimize the space overhead by deleting redundant data blocks across large-scale servers in data centers. However, such a process exacerbates the fragmentation of data blocks, causing more cross-server file retrievals with plummeting retrieval throughput. Some attempts prefer better file retrieval performance by confining all blocks of a file to one single server, resulting in non-trivial space consumption for more replicated blocks across servers. An ideal network storage system, in effect, should take both the deduplication and retrieval performance into account by implementing reasonable assignment of the detected unique blocks. Such a fine-grained assignment requires an accurate and comprehensive abstraction of the files, blocks, and the file-block affiliation relationships. To achieve this, we innovatively design the weighted hypergraph to profile the multivariate data correlations. With this delicate abstraction in place, we propose HyperPart, which elegantly transforms this complex block allocation problem into a hypergraph partition problem. For more general scenarios with dynamic file updates, we further propose a two-phase incremental hypergraph repartition scheme, which mitigates the performance degradation with minimal migration volume. We implement a prototype system of HyperPart, and the experiment results validate that it saves around 50% of the storage space and improves the retrieval throughput by approximately 30% of state-of-the-art methods under the balance constraints. Geyao Cheng, Junxu Xia, Lailong Luo, Haibo Mi, Deke Guo, Richard T. B. Ma |
IEEE Trans. Cloud Comput. | 2 |
| 2025 | Exploring Communication-Efficient Federated Learning via Stateless in-Network AggregationabstractAs an ambitious training paradigm, federated learning has garnered increasing attention in recent years, which enables collaborative training of a global model without accessing users’ private data. However, due to the simultaneous and constant model updates gathering from massive distributed clients, the central server generally becomes a performance bottleneck. Additionally, the stateful aggregation (retaining all the updates from each client) conducted by the central server further poses potential threats to privacy, since it may recover the raw data based on such model updates inversely. The state-of-the-art methodologies, however, fail to address these two problems concurrently and efficiently. To this end, we propose GAIN, a secure aggregation acceleration service for federated learning. At its core, GAIN leverages programmable switches deployed at the edge network to aggregate model updates in a stateless manner before transmitting them to the central server. Consequently, GAIN can accelerate the transmission and aggregation of model updates while eliminating the chance of recovering private data. We evaluate the performance of GAIN through FPGA-based experiments and large-scale simulations. The results show that GAIN can effectively reduce bandwidth overhead and achieve up to 4.11× training throughput acceleration while prioritizing privacy protection. Junxu Xia, Geyao Cheng, Wenfei Wu, Lailong Luo, Deke Guo |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | In-Network Aggregation as a Generic Service for Distributed ApplicationsabstractThe performance of distributed applications has long been hindered by network communication, which has emerged as a significant bottleneck. At the core of this issue, the many-to-one incast transfer stands out as one of the primary culprits. Existing works typically decompose the transmission into multiple concurrent sub-processes and utilize servers to aggregate relevant traffic, thus avoiding the incast transfer. However, limited by their theoretical bounds, these methods can only obtain limited performance improvement. In this paper, we discover that leveraging network devices for aggregating incast traffic proves highly effective in surpassing such limitations, while the advent of programmable switches further makes this envision practical. Based on this, we propose GISA as a solution for providing network acceleration across diverse distributed applications. GISA offers generic and uniform interfaces to various applications along with a switch resource sharing mechanism and policy for concurrent tasks. It also ensures correct and reliable transport while minimizing overhead through a low-overhead routing mechanism. Our FPGA-based prototype demonstrates that GISA can achieve line-rate processing when performing data aggregation with minor traffic overhead. Additionally, it supports a wide range of concurrent applications with little development effort. Junxu Xia, Wenfei Wu, Lailong Luo, Deke Guo, Geyao Cheng |
IEEE Trans. Netw. | 1 |
| 2024 | Accelerating and Securing Federated Learning with Stateless In-Network Aggregation at the EdgeabstractIn federated learning, sending the trained models (instead of raw data) from clients to the central server can surely decrease the volume of exchanged data and preserve data privacy to some extent. However, the central server can still be a system bottleneck due to the simultaneous and constant model gathering from massive distributed clients. Besides, the central server conducts stateful aggregation (retaining all the updates from each client), making it a potential threat to privacy, since it may recover the raw data based on such model updates inversely. The state-of-the-art methodologies, however, fail to address these two problems concurrently. To this end, we propose GAIN, a secure aggregation acceleration service for federated learning. At its core, GAIN aggregates the model updates at the programmable ingress switches in a stateless manner (storing the aggregated model parameters from the clients temporarily rather than permanently) before proceeding to the central server. Consequently, GAIN can accelerate the transmission and aggregation of model parameters while eliminating the chance of data recovery. We implemented a prototype of GAIN on an FPGA-based testbed to validate its performance. The results demonstrate that GAIN can achieve up to 4.11x speedup in training throughput and reduce up to 86.5% of traffic overhead. Furthermore, through theoretical analysis, we illustrate that GAIN can achieve even more substantial performance gains with a larger number of clients while guaranteeing privacy protection. Junxu Xia, Wenfei Wu, Lailong Luo, Geyao Cheng, Deke Guo, Qifeng Nian |
ICDCS | 1 |
| 2024 | Parallelized In-Network Aggregation for Failure Repair in Erasure-Coded Storage SystemsabstractTo repair a failed block in the erasure-coded storage system, multiple related blocks have to be retrieved from other storage nodes across the network. Such a process can lead to significant incast-type repair traffics and delays. The existing efforts mainly try to schedule the transmission of the requested blocks across different storage nodes to avoid network congestion. At their cores, they utilize part of the involved hosts to rely on or aggregate the file blocks from others. While we notice that, the programmability and capability of today’s network devices (i.e., routers and switches) bring a great opportunity to further speed up the repair progress by aggregating the file blocks with such devices. By mitigating the aggregation operations from the network edges to network cores, it is possible to save more time and bandwidth. With this intuition, we propose Paint, a parallelized in-network aggregation framework for failure repair. Paint utilizes programmable switches to aggregate relevant data and improves the repair performance by implementing multiple parallelized repair pipelines. We propose a series of novel and time-friendly algorithms to construct the routing paths for Paint and design the Aggregation Control Protocol to implement Paint in production clusters. For all we know, this is the first work to explore and implement parallelized in-network repair with programmable switches. The extensive experiments on the prototype system and real-world datasets indicate that Paint can significantly improve repair performance while effectively reducing bandwidth overhead. Junxu Xia, Lailong Luo, Geyao Cheng, Deke Guo |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | When Deduplication Meets Migration: An Efficient and Adaptive Strategy in Distributed Storage SystemsabstractThe traditional migration methods are confronted with formidable challenges when data deduplication technologies are incorporated. First, the deduplication creates data-sharing dependencies in the stored files; breaking such dependencies in migration may attach extra space overhead. Second, the redundancy elimination makes the storage system reserves only one copy for each storage file, and heightens the risk of data unavailability. The existing methods fail to tackle them in one shot. To this end, we propose Jingwei, an efficient and adaptive data migration strategy for deduplicated storage systems. To be specific, Jingwei tries to minimize the extra space cost in migration for space efficiency. Meanwhile, Jingwei realizes the service adaptability by encouraging replicas of hot files to spread out their data access requirements. We first model such a problem as an integer linear programming (ILP) and solve it with a commercial solver when only one empty migration target server is allowed. We then extend this problem to a scenario wherein multiple non-empty target servers are available for migration. We solve it by effective heuristic algorithms based on the Bloom Filter-based data sketches. The Jingwei strategy can suffer from performance degradation when the heat degree varies significantly. Therefore, we further present incremental adjustment strategies for the two scenarios, which adjust the number of block replicas and their locations in an incremental manner. The mathematical analyses and trace-driven experiments show the effectiveness of our Jingwei strategy. To be specific, Jingwei fortifies the file replicas by 25% with only 5.7% of the extra storage space, compared with the latest “Goseed” method. With the small extra space cost, the file retrieval throughput of Jingwei can reach up to 333.5 Mbps, which is 12.3% higher than that of the Random method. Geyao Cheng, Lailong Luo, Junxu Xia, Deke Guo, Yuchen Sun 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2023 | The Doctrine of MEAN: Realizing Deduplication Storage at Unreliable EdgeabstractPlacing popular data at the network edge helps reduce the retrieval latency, but it also brings challenges to the limited edge storage space. Currently, using available yet not necessarily reliable edge resources is common sense for edge space expansion, while deploying deduplication storage strategies is a general method for better space utilization. However, a contradiction arises when jointly implementing data deduplication with unreliable edge resources. On the one hand, the deduplication policy stipulates that any data chunk can be stored exactly once; on the other hand, the use of unreliable resources imposes that data should be backed up for the seek of file availability. To resolve such contradiction, we propose MEAN, a deduplication-enabled storage system using unreliable resources at the network edge. The core idea of MEAN is to place similar files together for better deduplication and maintain replicas of popular files for higher reliability. We first formulate this problem and prove its NP-hardness, then provide efficient heuristics based on similarity-aware hierarchical clustering. Three different reliability scenarios are comprehensively considered to develop our algorithms. We also implement a prototype system and evaluate the performance of MEAN with a real-world dataset. The results show that MEAN can fortify the file hit ratio under unreliable environments by 77% while reducing the file retrieval delay up to 71%, compared with the state-of-the-art approach. Junxu Xia, Geyao Cheng, Lailong Luo, Deke Guo |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | Jingwei: An Efficient and Adaptable Data Migration Strategy for Deduplicated Storage SystemsabstractThe traditional migration methods are confronted with formidable challenges when data deduplication technologies are incorporated. Firstly, the deduplication creates data-sharing dependencies in the stored files; breaking such dependencies in migration would attach extra space overhead. Secondly, the redundancy elimination heightens the risk of data unavailability during server crashes. The existing methods fail to tackle them at one shot. To this end, we propose Jingwei, an efficient and adaptable data migration strategy for deduplicated storage systems. To be specific, Jingwei tries to minimize the extra space cost in migration for space efficiency. Meanwhile, Jingwei realizes the service adaptability by encouraging replicas of hot data to spread out their data access requirements. We first model such a problem as an integer linear programming (ILP) and solve it with a commercial solver when only one empty migration target server is allowed. We then extend this problem to a scenario wherein multiple non-empty target servers are available for migration. We solve it by effective heuristic algorithms based on the Bloom Filter-based data sketches. Trace-driven experiments show that Jingwei fortifies the file replicas by 25%, while only 5.7% of the extra storage space is occupied compared with the latest "Goseed" method. Geyao Cheng, Deke Guo, Lailong Luo, Junxu Xia, Yuchen Sun 0001 |
INFOCOM | 4 |
| 2022 | LOFS: A Lightweight Online File Storage Strategy for Effective Data Deduplication at Network EdgeabstractEdge computing responds to users’ requests with low latency by storing the relevant files at the network edge. Various data deduplication technologies are currently employed at edge to eliminate redundant data chunks for space saving. However, the lookup for the global huge-volume fingerprint indexes imposed by detecting redundancies can significantly degrade the data processing performance. Besides, we envision a novel file storage strategy that realizes the following rationales simultaneously: 1) space efficiency, 2) access efficiency, and 3) load balance, while the existing methods fail to achieve them at one shot. To this end, we report LOFS, a Lightweight Online File Storage strategy, which aims at eliminating redundancies through maximizing the probability of successful data deduplication, while realizing the three design rationales simultaneously. LOFS leverages a lightweight three-layer hash mapping scheme to solve this problem with constant-time complexity. To be specific, LOFS employs the Bloom filter to generate a sketch for each file, and thereafter feeds the sketches to the Locality Sensitivity hash (LSH) such that similar files are likely to be projected nearby in LSH tablespace. At last, LOFS assigns the files to real-world edge servers with the joint consideration of the LSH load distribution and the edge server capacity. Trace-driven experiments show that LOFS closely tracks the global deduplication ratio and generates a relatively low load std compared with the comparison methods. Geyao Cheng, Deke Guo, Lailong Luo, Junxu Xia, Siyuan Gu |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | A Hybrid Framework for Class-Imbalanced Classification
Lailong Luo, Yingwen Chen 0001, Junxu Xia, Deke Guo |
WASA (1) | 4 |
| 2020 | Guaranteeing the response deadline for general aggregation trees
Jiangfan Li, Chendie Yao, Junxu Xia, Deke Guo |
Frontiers Comput. Sci. | 3 |
| 2020 | Efficient in-network aggregation mechanism for data block repairing in data centers
Junxu Xia, Deke Guo |
Future Gener. Comput. Syst. | 1 |
| 2020 | Secure and Trust-Oriented Edge Storage for Internet of ThingsabstractThe edge storage is a promising paradigm to support the Internet of Things (IoT) data storage, and is more efficient than the cloud storage in terms of the bandwidth overhead, the response latency, and so on. However, existing edge storage models cannot offer the security-aware data robustness and the adaptable data sharing among many uncertain users, due to the limitations of the utilized fault-tolerant storage technologies and the access control methods. In this article, we propose a secure and trust-oriented edge storage model, which would efficiently tackle the aforementioned two challenging issues in the IoT environment. More precisely, we first propose a robust and secure edge storage (RoSES) model using the totally local reconstruction code (TLRC) method presented in this article. It can achieve data robustness, high security, and lightweight computation at end devices. We further propose a trust-oriented data access (TODA) strategy for our RoSES model, which supports a wide and adaptable range of legitimate data accesses from uncertain requesters for IoT data sharing. We conduct extensive comparison and simulations to evaluate the performance of our new edge storage models. The results show that our model can efficiently realize the data storage, data recovery, and data sharing at the network edge, saving about 35% of storage cost and 76% degraded read latency. Besides, the data leakage probability is significantly reduced during the data storage and sharing processes. Junxu Xia, Geyao Cheng, Siyuan Gu, Deke Guo |
IEEE Internet Things J. | 1 |
| 2020 | A QoE-Aware Service-Enhancement Strategy for Edge Artificial Intelligence ApplicationsabstractDue to the high complexity of artificial intelligence (AI) algorithms, performing the AI tasks on the resource-limited Internet-of-Things (IoT) devices has been proved to be inadvisable. Edge computing provides an effective computing paradigm for executing AI tasks, where large numbers of AI tasks can be offloaded to the edge servers. Most of the existing works focus on achieving efficient computing offload through improving the Quality of Service (QoS), such as reducing the average server-side delay. However, we show that those efforts are inefficient due to the heterogeneous impact of delays on users' Quality of Experience (QoE). Inspired by the observations, in this article, we reconsider the scheduling method from an orthometric perspective, i.e., improving the QoE by designing a QoE-aware service-enhancement strategy for edge AI applications. Besides, multiple AI algorithms are utilized in our service model to execute the same type of tasks concurrently, thus meeting users' heterogeneity requirements of accuracy and delays. Specifically, for the online arriving AI tasks, we optimize the task allocation and scheduling strategy according to the QoE sensitivity of each task. The model can be formulated as the mixed-integer nonlinear programming problem, which is known to be NP-hard. Hence, we then propose an efficient two-phase scheduling strategy for this problem. The results of comprehensive emulations validate that our model can effectively improve the average QoE of users and achieve a higher task completion ratio. Junxu Xia, Geyao Cheng, Deke Guo, Xiaolei Zhou 0001 |
IEEE Internet Things J. | 1 |
| 2019 | In-network block repairing for erasure coding storage systemsabstractSummary In the erasure coding storage system, it is necessary to extract multiple data blocks from other remaining storage nodes to a new node when a storage node fails, which repairs the failed data block satisfactorily. However, this would incur the incast problem at this new node. The existing solutions for the repair process in the incast problem mainly rely on path planning and resource allocation. Although these solutions improve the performance of repairing the failed data blocks, they still waste a large amount of storage and bandwidth resources unavoidably. In this paper, we propose the incast problem to be resolved economically via the in‐network aggregation. Specifically, we assume that the switches in data centers have certain data processing capabilities and can aggregate data flows efficiently. Thereafter, we propose a set of in‐network methods to repair a failed data block in the erasure coding storage systems, taking the fat‐tree data center as an example. Thus, the incast problem can be solved effectively during the data transmission process. Compared with the prior methods, our approach effectively avoids the overhead of extra path computing, as well as significantly reduces the link cost of repairing data blocks, while promising similar or faster repair speed. Junxu Xia, Deke Guo, Geyao Cheng |
Concurr. Comput. Pract. Exp. | 1 |
| 2018 | Topology-Aware Efficient Storage Scheme for Fault-Tolerant Storage Systems in Data CentersabstractIn data centers, files are stored with the method of erasure code or replication to guarantee data reliability. However, both of the methods are not communication-friendly. An erasure coding system spends vast time to extract k data blocks across the racks during the decoding process. The transmission time contributes up to 94% of the total decoding time. In a multi-replica system, the frequent data writing and updating also lead to non-trivial bandwidth overhead. In this paper, we consider server-centric data centers (such as BCube), in which any pair of nodes are interconnected with multiple parallel paths. In such data centers, the transmissions for file storage can be significantly speed up by utilizing the parallel paths concurrently. With this insight, we first define the disjoint node in BCube. Based on this definition, we design the node-disjoint storage strategy (NDSS)and the nested node-disjoint storage strategy (N-NDSS)to improve the transmission between distributed storage nodes. Comprehensive simulations show the performance of our strategies in both erasure coding system and multi-replica system. Junxu Xia, Deke Guo, Lailong Luo, Jiangfan Li, Chendie Yao |
ICPADS | 1 |