Xuming Ye

dblp:294/5062 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
13since 2021 · last 2025
0009-0005-8860-9674ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Context-aware resemblance detection for data deduplication with neural network
Xuming Ye, Yaping Wan, Ruixuan Li 0001, Weijun Xiao, Zhiyong Xu 0003
Eng. Appl. Artif. Intell.1
2025 Sym-CS-HFL: A secure and efficient solution for privacy-preserving heterogeneous federated learning
Jinzhao Wang, Junwei Tang, Xuming Ye, Yaping Wan, Zhiyong Xu 0003, Lingna Chen
J. Inf. Secur. Appl.4
2025 Nis-PoW: Non-interactive secure Proof of Ownership for cloud storage
Zhihuan Yang, Ruixuan Li 0001, Xuming Ye, Zhiyong Xu 0003
J. Syst. Archit.4
2025 IBNR-RD: Intra-Block Neighborhood Relationship-Based Resemblance Detection for High-Performance Multi-Node Post-Deduplication
abstract
Post-deduplication in traditional cloud environments primarily focuses on single-node, where delta compression is performed on the same deduplication node located on server side. However, with data explosion, the multi-node post-deduplication, also called global deduplication, has become a hot issue in research communities, which aims to simultaneously execute delta compression on data distributed across all nodes. Simply setting up single-node deduplication systems on multi-node environments would significantly affect storage utilization and incur secondary overhead from file migration. Nevertheless, existing global deduplication solutions suffer from lower data compression ratios and high computational overhead due to their resemblance detection's inherent limitations and overly coarse granularities. Similar blocks typically have high correlations between sub-blocks; inspired by this observation, we propose IBNR (Intra-Block Neighborhood Relationship-Based Resemblance Detection for High-Performance Multi-Node Post-Deduplication), which introduces a novel resemblance detection based on relationships between sub-blocks and determines the ownership of blocks in entry stage to achieve efficient global deduplication. Furthermore, the by-products of IBNR have shown powerful scalability by replacing internal resemblance detection scheme with existing solutions on practical workloads. Experimental results indicate that IBNR outperforms state-of-the-art solutions, achieving an average 1.99× data reduction ratio and varying degrees of improvement across other key metrics.
Dewen Zeng, Ruixuan Li 0001, Xuming Ye, Zhiyong Xu 0003
IEEE Trans. Cloud Comput.5
2024 Multimodal Summarization with Modality-Aware Fusion and Summarization Ranking
Xuming Ye, Chaomurilige Wang, Haoyu Luo, Yingzhe Luo
ICA3PP (2)1
2024 Facet-Aware Multimodal Summarization via Cross-Modal Alignment
Xuming Ye, Tianjiao Xing, Chaomurilige Wang, Xuan Liu 0008
ICPR (19)2
2024 Length Controllable Model for Images and Texts Summarization
Xiangyu Qu, Xuan Liu 0008, Xuming Ye
ICWS4
2024 Who Owns the Cloud Data? Exploring a non-interactive way for secure proof of ownership
abstract
Cloud storage is widely used for flexible and efficient data management, yet redundant data across users requires storage optimization. Deduplication helps by storing only unique data, making data sharing and ownership verification essential post-deduplication. Current Proof of Ownership (PoW) methods rely on interactive communication, leading to delays and performance issues during intensive data operations, and often assume a level of trust that may not hold in practical scenarios. To overcome these issues, we propose ES-PoW, a non-interactive secure proof of ownership scheme for cloud storage. ES-PoW performs ownership verification in a single round, avoiding delays in block verification. Using modular exponentiation and the discrete logarithm problem, ES-PoW generates and verifies ownership proofs efficiently. Unlike previous schemes, ES-PoW is resilient against brute-force, replay, and Man-in-the-Middle attacks, without relying on trusted nodes, making it better suited to real-world applications. Experimental results show that ES-PoW reduces I/O operation time and accelerates computation, achieving up to 53.6× and 54.5× speed improvements over current methods.
Zhihuan Yang, Ruixuan Li 0001, Xuming Ye, Zhiyong Xu 0003
TrustCom4
2023 Sym-Fed: Unleashing the Power of Symmetric Encryption in Cross-Silo Federated Learning
abstract
With the increasing number of big data applications, large amounts of valuable data are distributed in different organizations or regions. Federated Learning (FL) enables collaborative model training without sharing sensitive data and is widely used in AI medical diagnosis, economy, and autonomous driving scenarios. However, it still leaks the privacy from the gradient exchange in federated learning. What’s worse, state-of-the-art work, such as Batchcrypt, still suffers from computational overhead due to a considerable amount of computation and communication costs caused by homomorphic encryption. Therefore, we propose a novel symmetric key-based homomorphic encryption scheme, Sym-Fed. To unleash the power of symmetric encryption in federated learning, we combine random masking with symmetric encryption and keep the homomorphic property during the gradient exchange in the federated learning process. Finally, the security analysis and experimental results on real workloads show that our design achieves performance improvement 6× to 668× and reduces the communication overhead 1.2× to 107× compared with the state-of-the-art work, BatchCrypt and FATE, without model accuracy degradation and security compromise.
Jinzhao Wang, Ruixuan Li 0001, Junwei Tang, Xuming Ye, Yaping Wan, Zhiyong Xu 0003
TrustCom5
2022 Context-aware Resemblance Detection based Deduplication Ratio Prediction for Cloud Storage
abstract
With the prevalence of cloud storage, people prefer to outsource their data to the cloud for flexibility and reliability. Undoubtedly, there are lots of redundancy among these data. However, high-end storage with deduplication costs heavy computation and increases the data management complexity. Potential customers need the redundancy proportion information of their outsourced data to decide whether high-end storage with deduplication is worthwhile. Thus, many researchers have previously attempted to predict the redundant ratio. However, existing mechanisms ignore the redundancy proportion among similar chunks containing many duplicate data. Although resemblance detection, detecting the duplicate parts among similar data, has become a hot issue, it is hardly applied to the conventional deduplication ratio estimation because of unacceptable calculation cost. Therefore, we analyze the limitations and challenges of deduplication ratio prediction in prediction scope and response time and further propose a novel prediction scheme. By leveraging the context-aware resemblance detection, and confidence interval theory, our method can achieve faster estimation speed with higher accuracy in deduplication ratio compared with the state-of-the-art work. Finally, the results show that our method can efficiently and effectively estimate the proportion of duplicate chunks and redundant data among similar chunks by conducting experiments on real workloads.
Yuqing Geng, Ruixuan Li 0001, Weijun Xiao, Chunping Ouyang, Qifei Liu, Xuming Ye, Zhiyong Xu 0003
BDCAT9
2022 Chunk Content is not Enough: Chunk-Context Aware Resemblance Detection for Deduplication Delta Compression
abstract
In this paper, we propose a novel chunk-context-aware resemblance detection al-gorithm called CARD. By introducing machine learning into deduplication, the chunk feature will embed the chunk-context information after the N-sub-chunk shingles based initial feature extraction and BP-Neural network training. In the predicting process, each chunk's initial feature corresponds to a chunk-context feature. Finally, the cloud calculates the different part among resemblance chunks based on these feature by delta encoding. Only the different part is stored. The basic workflow corresponds to Figure 1. For more detailed illustrations, please see our full paper here
Xuming Ye, Xiaoye Xue, Ruixuan Li 0001, Weijun Xiao, Zhiyong Xu 0003, Yaping Wan
DCC1
2022 Cross-domain Resemblance Detection based on Meta-learning for Cloud Storage
abstract
Recently, cloud storage has been widely used in our daily life. And there are lots of redundancy among these outsourced data. Conventional deduplication technology efficiently splits these data at the chunk level and removes the duplicate chunks to save the network bandwidth and improve the cloud storage utility. But it ignores the redundancy among similar chunks. Resemblance detection has recently become a hot issue with detecting these redundant parts among similar data. CARD, the state-of-the-art work, can efficiently and effectively remove these redundancies by introducing the neural network with resemblance detection. However, the source domain of the CARD model may have an explicitly different input distribution. The cloud cannot deal with the possible future domain data based on CARD design. This cross-domain setting may serials degrades the performance of CARD. To overcome this problem, we propose a cross-domain resemblance detection scheme called MetaContext. Integrating the chunk-context aware model and the learn-to-learn idea can produce a more robust chunk feature than CARD. As a byproduct, it also outperforms the CARD in speed. Finally, we implement the MetaContext and conduct serial experiments on real workloads. The results show that our method can efficiently and effectively detect and remove the redundancy among similar data.
Baisong Li, Ruixuan Li 0001, Weijun Xiao, Zhongming Fu, Xuming Ye, Renjiao Duan, Zhiyong Xu 0003
IPCCC6
2021 Fast Variable-Grained Resemblance Data Deduplication For Cloud Storage
abstract
With the prevalence of cloud storage, data deduplication has been a widely used technology by removing cross users’ duplicate data and saving network bandwidth. Nevertheless, traditional data deduplication hardly detects duplicate data among resemblance chunks. Currently, a resemblance data deduplication, called Finesse, has been proposed to detect and remove the duplicate data among similar chunks efficiently. However, we observe that the chunks following the similar chunk have a high chance of resembling data locality property, and vice versa. Processing these adjacent similar chunks in small average chunk size level increases the metadata, which deteriorates the deduplication system performance. Moreover, existing resemblance data deduplication schemes ignore the performance impact from metadata. Therefore, we propose a fast variable-grained resemblance data deduplication for cloud storage. It dynamically combines the adjacent resemblance chunks or unique chunks or breaks those chunks, located at the transition region between resemblance chunks and unique chunks. Finally, we implement a prototype and conduct a serial of experiments on real-world datasets. The results show that our method dramatically reduces the metadata size while achieving the high deduplication ratio.
Xuming Ye, Ruixuan Li 0001, Weijun Xiao, Yuqing Geng, Zhiyong Xu 0003
NAS1