Wei Zhou 0077

dblp:69/5011-77 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0005-8205-1753ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AlignSketch: A Framework for Aligning Theoretical and Practical Estimation Errors
Hanyue Zheng, Jingwei Shi, Xinye Xu, Wei Zhou 0077, Tong Yang 0003, Zhenyu Guan 0002, Yong Cui 0001
ICDE5
2026 Stair Sketch: Towards Clearer Memory of Recent Events for Network Security Audits
abstract
In large-scale and complex network environments, frequent security events make security audits increasingly important. Effective audits rely on maintaining a clearer memory of recent security events to support reliable identification and analysis. However, in high-volume data streams, existing methods struggle to ensure more accurate recording of recent time periods under a fixed memory budget while keeping the overall error controllable. To address this, we propose a novel data stream processing structure, the Stair Sketch. The key idea is to organize limited memory into a staircase of atomic sketches, where newer time periods hold more sketches. When memory is exhausted, one sketch is reclaimed from each existing period for the new one, efficiently allocating more memory to recent periods without updating all sketches. To record longer time periods and handle load imbalance, we introduce vertical sampling and horizontal sharing techniques to optimize our method. We theoretically derive the error bounds of Stair Sketch. Extensive experiments show that Stair Sketch improves accuracy by at least one order of magnitude compared with state-of-the-art methods, while reducing memory accesses by up to 50%. All related codes are open-sourced at GitHub.
Wei Zhou 0077, Yikai Zhao 0001, Tong Yang 0003, Zhenyu Guan 0002, Bin Cui 0001
IEEE Trans. Dependable Secur. Comput.2
2025 Extendible RDMA-Based Remote Memory KV Store with Dynamic Perfect Hashing Index
abstract
Perfect hashing is a special hashing function that maps each item to a unique location without collision, which enables the creation of a KV store with small and constant lookup time. Recent dynamic perfect hashing attains high load factor by increasing associativity, which impacts bandwidth and throughput. This paper proposes a novel dynamic perfect hashing index without sacrificing associativity, and uses it to devise an RDMA-based remote memory KV store called CuckooDuo. CuckooDuo simultaneously achieves high load factor, fast speed, minimal bandwidth, and efficient expansion without item movement. We theoretically analyze the properties of CuckooDuo, and implement it in an RDMA-network based testbed. The results show CuckooDuo achieves 1.9~17.6x smaller insertion latency and 9.0~18.5x smaller insertion bandwidth than prior works.
Zirui Liu 0002, Xian Niu, Wei Zhou 0077, Yisen Hong, Zhouran Shi, Tong Yang 0003, Yuchao Zhang 0004, Yuhan Wu 0001, Yikai Zhao 0001, Zhuochen Fan, Bin Cui 0001
ICDE3
2025 Cooled-KLL: Enhancing Quantile Estimation by Filtering Hot Item
abstract
Quantile estimation is critical for diverse applications, including database management and network traffic monitoring. Probabilistic quantile sketches are widely employed in practice, with the KLL sketch (introduced in 2016) being particularly notable for its theoretically space-optimal properties. However, KLL overlooks the inherent repetition of elements often present in real-world data streams. Such streams are frequently highly skewed, characterized by ''hot items''-items that appear with high frequency. The KLL sketch processes these hot items without accounting for their prevalence, resulting in suboptimal space utilization due to redundant insertions and storage. To overcome this limitation, we propose Cooled-KLL, an enhanced KLL sketch. Cooled-KLL introduces a novel ''Hot Filter'' structure that efficiently identifies and stores hot items as compact key-value pairs. This mechanism ensures that only ''cold'' (less frequent) items are subsequently processed by the core KLL sketch. Our approach significantly reduces memory consumption without compromising processing speed. Extensive experiments demonstrate that Cooled-KLL consistently outperforms five other state-of-the-art algorithms, achieving up to 2.5 orders of magnitude higher accuracy compared to the standard KLL sketch.
Qilong Shi, Wei Zhou 0077, Yizhuo Zheng, Xinye Xu, Yuanyuan Zhang 0006, Long Yao, Yangyang Wang 0001, Mingwei Xu 0001
KDD (2)2
2025 Achieving Top-$K$K-fairness for Finding Global Top-$K$K Frequent Items
abstract
Finding top-$K$frequent items has been a hot topic in data stream processing with wide-ranging applications. However, most existing sketch algorithms focus on finding local top-$K$in a single data stream. In this paper, we tackle finding global top-$K$across multiple data streams. We find that using prior sketch algorithms directly is often unfair in global scenarios, degrading global top-$K$accuracy. We define top-$K$-fairness and show its importance for finding global top-$K$. To achieve this, we propose the Double-Anonymous (DA) sketch, where double-anonymity ensures fairness. We also propose two techniques, hot-filtering and early-freezing, to improve accuracy further. We theoretically prove that the DA sketch achieves top-$K$-fairness while maintaining high accuracy. Extensive experiments verify top-$K$-fairness in disjoint data streams, showing that the DA sketch's error is up to 129 times (60 times on average) smaller than the state-of-the-art. To enhance the applicability and technical depth, we also investigate how to extend the DA sketch to general distributed data stream scenarios and how to provide a fairer and more accurate global ranking for top-$K$items. The experimental results show that the extended version of the DA sketch can indeed compute better rankings and still has significant advantages in general data streams.
Yikai Zhao 0001, Wei Zhou 0077, Wenchen Han, Yinda Zhang 0002, Xiuqi Zheng, Tong Yang 0003, Bin Cui 0001
IEEE Trans. Knowl. Data Eng.2