EDBT 2026 Demo / reviewers in the wild / expert
Weijun Xiao
dblp:30/4252
· DBLP profile ↗
5ranked-venue papers in the field
0as first author
2since 2021 · last 2022
0000-0002-2147-7575ORCID · corroborated
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 2Big Data, Cloud & Distributed Data Systems · 2Information Retrieval & Web Search · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Context-aware Resemblance Detection based Deduplication Ratio Prediction for Cloud StorageabstractWith the prevalence of cloud storage, people prefer to outsource their data to the cloud for flexibility and reliability. Undoubtedly, there are lots of redundancy among these data. However, high-end storage with deduplication costs heavy computation and increases the data management complexity. Potential customers need the redundancy proportion information of their outsourced data to decide whether high-end storage with deduplication is worthwhile. Thus, many researchers have previously attempted to predict the redundant ratio. However, existing mechanisms ignore the redundancy proportion among similar chunks containing many duplicate data. Although resemblance detection, detecting the duplicate parts among similar data, has become a hot issue, it is hardly applied to the conventional deduplication ratio estimation because of unacceptable calculation cost. Therefore, we analyze the limitations and challenges of deduplication ratio prediction in prediction scope and response time and further propose a novel prediction scheme. By leveraging the context-aware resemblance detection, and confidence interval theory, our method can achieve faster estimation speed with higher accuracy in deduplication ratio compared with the state-of-the-art work. Finally, the results show that our method can efficiently and effectively estimate the proportion of duplicate chunks and redundant data among similar chunks by conducting experiments on real workloads. Yuqing Geng, Ruixuan Li 0001, Weijun Xiao, Chunping Ouyang, Qifei Liu, Xuming Ye, Zhiyong Xu 0003 |
BDCAT | 4 |
| 2022 | Chunk Content is not Enough: Chunk-Context Aware Resemblance Detection for Deduplication Delta CompressionabstractIn this paper, we propose a novel chunk-context-aware resemblance detection al-gorithm called CARD. By introducing machine learning into deduplication, the chunk feature will embed the chunk-context information after the N-sub-chunk shingles based initial feature extraction and BP-Neural network training. In the predicting process, each chunk's initial feature corresponds to a chunk-context feature. Finally, the cloud calculates the different part among resemblance chunks based on these feature by delta encoding. Only the different part is stored. The basic workflow corresponds to Figure 1. For more detailed illustrations, please see our full paper here Xuming Ye, Xiaoye Xue, Ruixuan Li 0001, Weijun Xiao, Zhiyong Xu 0003, Yaping Wan |
DCC | 5 |
| 2012 | A GPU-Based Accelerator for Chinese Word Segmentation
Xiwu Gu, Ruixuan Li 0001, Kunmei Wen, Bei Peng 0001, Weijun Xiao |
APWeb | 5 |
| 2011 | Measuring Social Tag Confidence: Is It a Good or Bad Tag?
Xiwu Gu, Xianbing Wang, Ruixuan Li 0001, Kunmei Wen, Weijun Xiao |
WAIM | 6 |
| 2011 | A New Vector Space Model Exploiting Semantic Correlations of Social Annotations for Web Page Clustering
Xiwu Gu, Xianbing Wang, Ruixuan Li 0001, Kunmei Wen, Weijun Xiao |
WAIM | 6 |