EDBT 2026 Demo / reviewers in the wild / expert
Zhiyong Xu 0003
dblp:54/3171-3
· DBLP profile ↗
5ranked-venue papers in the field
0as first author
3since 2021 · last 2026
0000-0001-9544-1500ORCID · conflict
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 2Data Mining & Knowledge Discovery · 1Information Retrieval & Web Search · 1Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Blockchain-Based Decentralized Trusted Cloud Resource Storage Pricing Incentive Mechanism
Yuxuan Chi, Qiong Tao, Jianfeng Lu 0002, Zhiyong Xu 0003, Yaping Wan, Wei Liang 0005, Meikang Qiu |
KSEM (4) | 5 |
| 2022 | Context-aware Resemblance Detection based Deduplication Ratio Prediction for Cloud StorageabstractWith the prevalence of cloud storage, people prefer to outsource their data to the cloud for flexibility and reliability. Undoubtedly, there are lots of redundancy among these data. However, high-end storage with deduplication costs heavy computation and increases the data management complexity. Potential customers need the redundancy proportion information of their outsourced data to decide whether high-end storage with deduplication is worthwhile. Thus, many researchers have previously attempted to predict the redundant ratio. However, existing mechanisms ignore the redundancy proportion among similar chunks containing many duplicate data. Although resemblance detection, detecting the duplicate parts among similar data, has become a hot issue, it is hardly applied to the conventional deduplication ratio estimation because of unacceptable calculation cost. Therefore, we analyze the limitations and challenges of deduplication ratio prediction in prediction scope and response time and further propose a novel prediction scheme. By leveraging the context-aware resemblance detection, and confidence interval theory, our method can achieve faster estimation speed with higher accuracy in deduplication ratio compared with the state-of-the-art work. Finally, the results show that our method can efficiently and effectively estimate the proportion of duplicate chunks and redundant data among similar chunks by conducting experiments on real workloads. Yuqing Geng, Ruixuan Li 0001, Weijun Xiao, Chunping Ouyang, Qifei Liu, Xuming Ye, Zhiyong Xu 0003 |
BDCAT | 10 |
| 2022 | Chunk Content is not Enough: Chunk-Context Aware Resemblance Detection for Deduplication Delta CompressionabstractIn this paper, we propose a novel chunk-context-aware resemblance detection al-gorithm called CARD. By introducing machine learning into deduplication, the chunk feature will embed the chunk-context information after the N-sub-chunk shingles based initial feature extraction and BP-Neural network training. In the predicting process, each chunk's initial feature corresponds to a chunk-context feature. Finally, the cloud calculates the different part among resemblance chunks based on these feature by delta encoding. Only the different part is stored. The basic workflow corresponds to Figure 1. For more detailed illustrations, please see our full paper here Xuming Ye, Xiaoye Xue, Ruixuan Li 0001, Weijun Xiao, Zhiyong Xu 0003, Yaping Wan |
DCC | 6 |
| 2014 | Online Anomaly Detection by Improved Grammar Compression of Log SequencesabstractNowadays, log sequences mining techniques are widely used in detecting anomalies for Internet services. The state-of-the-art anomaly detection methods either need significant computational costs, or require specific assumptions that the test logs are holding certain data distribution patterns in order to be effective. Therefore, it is very difficult to achieve real time responses and it greatly reduces the effectiveness of these mechanisms in reality. To address these issues, we propose an innovative anomaly detection strategy called CADM. In CADM, the relative entropy between test logs and normal logs is exploited to discover the anomalous levels. Instead of calculating the relative entropy based on certain predefined data distribution models, our solution inspects the relationship between relative entropy and compression size with an improved grammar-based compression method. No assumptions are needed. In addition, our mechanism has excellent scalability with only O(n) computational complexity. It can generate the detection results on the fly. Experimental analysis with both synthetic and real world logs proves that CADM is superior to the other methods. It can achieve very high anomaly detection accuracy with the minimal computational overhead. It is suitable for log mining tasks and can be applied on a broad variety of application fields. Wei Zhou 0019, Jizhong Han, Dan Meng 0002, Zhiyong Xu 0003 |
SDM | 6 |
| 2012 | ℓ1-Graph Based Community Detection in Online Social Networks
Ruixuan Li 0001, Yuhua Li 0003, Xiwu Gu, Kunmei Wen, Zhiyong Xu 0003 |
APWeb | 6 |