EDBT 2026 Demo / reviewers in the wild / expert
Hengrui Wang
dblp:274/2152
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Storage systems · 96% Distributed systems · 4% | |
| Computer networks
1 paper |
Network measurement and analytics · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems
key-value storage |
1.6 | 2 | 2025 | Rethinking The Compaction Policies in LSM-trees · Proc. ACM Manag. Data 2025 GRF: A Global Range Filter for LSM-Trees with Shape Encoding · Proc. ACM Manag. Data 2024 |
Storage systems › key-value storage
LSM-tree |
1.6 | 2 | 2025 | Rethinking The Compaction Policies in LSM-trees · Proc. ACM Manag. Data 2025 GRF: A Global Range Filter for LSM-Trees with Shape Encoding · Proc. ACM Manag. Data 2024 |
Storage systems › key-value storage
compaction strategy |
0.9 | 1 | 2025 | Rethinking The Compaction Policies in LSM-trees · Proc. ACM Manag. Data 2025 |
Storage systems › flash and SSD › flash memory management › garbage collection
write amplification |
0.9 | 1 | 2025 | Rethinking The Compaction Policies in LSM-trees · Proc. ACM Manag. Data 2025 |
Storage systems › flash and SSD
read amplification |
0.8 | 1 | 2024 | GRF: A Global Range Filter for LSM-Trees with Shape Encoding · Proc. ACM Manag. Data 2024 |
Network measurement and analytics › traffic measurement › flow measurement
flow size estimation |
0.7 | 1 | 2023 | Enhanced Machine Learning Sketches for Network Measurements · IEEE Trans. Computers 2023 |
Network measurement and analytics
sketch-based measurement |
0.7 | 1 | 2023 | Enhanced Machine Learning Sketches for Network Measurements · IEEE Trans. Computers 2023 |
Network measurement and analytics
traffic measurement |
0.7 | 1 | 2023 | Enhanced Machine Learning Sketches for Network Measurements · IEEE Trans. Computers 2023 |
Distributed systems › concurrency control
multi-version concurrency control |
0.2 | 1 | 2024 | GRF: A Global Range Filter for LSM-Trees with Shape Encoding · Proc. ACM Manag. Data 2024 |
Methods — techniques the papers use, named apart from their topics
three-level model · 0.9dynamic programming · 0.9shape encoding · 0.8machine learning · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Vision-language joint modeling framework for rubber-tree planting-hole detection in unmanned aerial vehicle imagery
Pintian Lin, Wentao Peng, Yaowen Hu, Yujian Liu, Huaiqing Zhang, Hengrui Wang, Jiangquan Zeng, Shicong He, Zidi Wu, Amar Jain, Yingfang Zhu, Guoxiong Zhou |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | Rethinking The Compaction Policies in LSM-treesabstractLog-structured merge-trees (LSM-trees) are widely used to construct key-value stores. They periodically compact overlapping sorted runs to reduce the read amplification. Prior research on compaction policies has focused on the trade-off between write amplification (WA) and read amplification (RA). In this paper, we propose to treat the compaction operation in LSM-trees as a computational and I/O-bandwidth investment for improving the system's future query throughput, and thus rethink the compaction policy designs. A typical LSM-tree application handles a steady but moderate write stream and prioritizes resources for top-level flushes of small sorted runs to avoid data loss due to write stalls. The goal of the compaction policy, therefore, is to maintain an optimal number of sorted runs to maximize average query throughput. Because compaction and read operations compete for the CPU and I/O resources from the same pool, we must perform a joint optimization to determine the appropriate timing and aggressiveness of the compaction. We introduce a three-level model of an LSM-tree and propose EcoTune, an algorithm based on dynamic programming to find the optimal compaction policy according to workload characterizations. Our evaluation on RocksDB shows that EcoTune improves the average query throughput by 1.5x to 3x over the leveling policy and by up to 2.5x over the lazy-leveling policy on workloads with range/point query ratios. Hengrui Wang, Jiansheng Qiu, Fangzhou Yuan, Huanchen Zhang |
Proc. ACM Manag. Data | 1 |
| 2025 | A Fine-Scale Segmentation Method for Individual Rubber Trees Based on UAV LiDAR Point CloudabstractAs a key tropical economic crop, rubber trees play a vital role in both the global rubber industry and the health of ecological systems. Fine-grained segmentation of rubber tree point clouds is essential for accurately extracting structural parameters and achieving effective monitoring and management. However, existing unsupervised segmentation methods are often affected by ground noise and overlapping tree crowns, leading to suboptimal segmentation results and posing significant challenges for individual rubber tree segmentation. To address these issues, this study proposes a fine-grained segmentation network for rubber trees based on UAV LiDAR point clouds, termed RTreeNet. First, we designed a Multi-Scale Feature Aggregation (MSFA) module to tackle the issue of leaf overlap by capturing geometric features at the edges of tree crowns. Secondly, we proposed a Cosine-Space Cross Attention (CSCA) module, which calculates the cosine similarity of vertical and horizontal features for each point, effectively eliminating interference from ground noise. Additionally, an Adaptive Coati Particle Optimization Algorithm (ACPA) was proposed to determine the optimal learning rate for the network, further enhancing segmentation accuracy. Experimental evaluation demonstrates that the proposed RTreeNet outperforms seven state-of-the-art point cloud segmentation architectures and four conventional segmentation algorithms on our custom dataset, achieving a mean Intersection over Union (mIoU) of 86.3% and an F-score of 92.5%. In the generalization experiment, RTreeNet showed high accuracy and stability on three public datasets. The method also measured the specific structural parameters (tree height, crown diameter, and breast diameter) of rubber trees in the two regions, providing strong technical support for the refined management of rubber trees, agricultural planning, pest control, and rubber yield prediction. Zilin Ye, Miying Yan, Guoxiong Zhou, Hengrui Wang, Mingjie Lv |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Incremental Learning-Driven Segmentation and Clustering Optimization of UAV-LiDAR Rubber Tree Point CloudsabstractAs an important economic crop in tropical regions, rubber trees require precise individual-level identification and structural parameter extraction to support productivity assessment and fine-scale plantation management. However, existing point cloud segmentation methods face significant challenges in handling species diversity, overlapping canopies, and clustering parameter sensitivity. Therefore, we propose an incremental learning-driven network, termed RtSegNet, for UAV LiDAR-based rubber tree point cloud segmentation and clustering optimization. Firstly, we design a Tree-Structure Incremental Learning Sampling (TSILS) strategy, which significantly improves model adaptability across diverse forest structures and enhances training efficiency. Secondly, a Fourier Spectrum-enhanced Spatial Graph Clustering (FSGC) method is proposed to improve feature discrimination in overlapping canopy regions. Finally, a Gaussian Random Duck Swarm Algorithm (GR-DSA) is proposed to adaptively optimize clustering parameters, enabling robust segmentation across trees of varying scales. Experimental results on a self-constructed rubber tree dataset show that RtSegNet outperforms ten state-of-the-art methods, achieving a mean Intersection over Union (mIoU) of 76.37%, an F-score of 87.11%, and reducing training time to 25.04 hours. Generalization tests on the FOR-instance and NIBIO MLS public datasets further confirm the method’s cross-environment robustness.Based on the segmentation outputs, key structural parameters such as tree height, canopy diameter, and canopy volume are extracted. The estimated tree height achieves anR2of 0.94 with a root mean square error (RMSE) of 0.52, providing strong technical support for accurate monitoring, yield prediction, health assessment, and the modernization of tropical agriculture. The source code in this work is available at https://github.com/aaaaasleep/RtSegNet. Hengrui Wang, Guoxiong Zhou, Yongfei Xue |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Improving multiple sclerosis lesion segmentation across clinical sites: A federated learning approach with noise-resilient trainingabstractAccurately measuring the evolution of Multiple Sclerosis (MS) with magnetic resonance imaging (MRI) critically informs understanding of disease progression and helps to direct therapeutic strategy. Deep learning models have shown promise for automatically segmenting MS lesions, but the scarcity of accurately annotated data hinders progress in this area. Obtaining sufficient data from a single clinical site is challenging and does not address the heterogeneous need for model robustness. Conversely, the collection of data from multiple sites introduces data privacy concerns and potential label noise due to varying annotation standards. To address this dilemma, we explore the use of the federated learning framework while considering label noise. Our approach enables collaboration among multiple clinical sites without compromising data privacy under a federated learning paradigm that incorporates a noise-robust training strategy based on label correction. Specifically, we introduce a Decoupled Hard Label Correction (DHLC) strategy that considers the imbalanced distribution and fuzzy boundaries of MS lesions, enabling the correction of false annotations based on prediction confidence. We also introduce a Centrally Enhanced Label Correction (CELC) strategy, which leverages the aggregated central model as a correction teacher for all sites, enhancing the reliability of the correction process. Extensive experiments conducted on two multi-site datasets demonstrate the effectiveness and robustness of our proposed methods, indicating their potential for clinical applications in multi-site collaborations to train better deep learning models with lower cost in data collection and annotation. Lei Bai 0001, Dongang Wang, Hengrui Wang, Michael Barnett 0006, Mariano Cabezas, Tom Weidong Cai, Fernando Calamante, Kain Kyle, Dongnan Liu, Linda Ly, Aria Nguyen, Chun-Chien Shieh, Ryan Sullivan, Geng Zhan, Wanli Ouyang, Chenyu Wang 0001 |
Artif. Intell. Medicine | 3 |
| 2024 | GRF: A Global Range Filter for LSM-Trees with Shape EncodingabstractLog-structured merge-trees (LSM-trees) are widely used in key-value stores because of its excellent write performance. To reduce LSM-tree's read amplification due to overlapping sorted runs, each file (i.e., SSTable) in an LSM-tree is typically associated with a point or range filter to reduce unnecessary I/Os to the runs that do not contain the target key (range). However, as modern SSDs get faster, probing multiple in-memory filters per query often makes the system CPU bottlenecked, thus compromising the system's throughput. In this paper, we developed the Global Range Filter (GRF) for RocksDB that reduces the number of filter probes per query to one. We follow the pioneering Chucky's approach by storing the sorted run IDs within the filter. However, we identify two practical challenges in building a global range filter: correctness in multi-version concurrency control and efficiency in frequent updates. We solve both challenges by the novel Shape Encoding algorithm. With further optimizations, GRF achieves a dominating performance over the state-of-the-art filters under different workloads when integrated into RocksDB. Hengrui Wang, Te Guo 0001, Junzhao Yang, Huanchen Zhang |
Proc. ACM Manag. Data | 1 |
| 2023 | Enhanced Machine Learning Sketches for Network MeasurementsabstractNetwork monitoring and management require accurate statistics of a variety of flow-level metrics such as flow sizes, top-$k$flows, and number of flows. Arguably, the current best technique to measure these metrics is sketches. While a significant amount of work has already been done on sketching techniques, there is still a lot of room for improvement because the accuracy of existing sketches varies with changing characteristics of network traffic. In this paper, we propose the idea of using machine learning to improve the accuracy of sketches, and propose ageneric machine learning frameworkto reduce the dependence of accuracy of sketches on network traffic characteristics. We further present three case studies, where we applied our machine learning framework on sketches for measuring three flow-level network metrics, namely flow sizes, top-$k$flows, and number of flows. We implemented and extensively evaluated this framework for these three metrics using both real-world and synthetic traffic traces. To the best of our knowledge, this is the first work that uses machine learning to reduce the dependence of sketching techniques on the characteristics of network traffic. We have released all our traces and implementation codes at Github. Hengrui Wang, Tong Yang 0003, Muhammad Shahzad 0001 |
IEEE Trans. Computers | 1 |
| 2021 | Hard Disk Failure Prediction Based on Lightgbm with CIDabstractIn data centers, hard disks are the most prone to failure of IT equipment. Although there is data backup, data reliability still faces challenges due to hard disks failure. In recent years, many hard disk failure prediction approaches based on SMART data have been proposed. In this paper, we proposed a novel disk failure prediction approach based on Lightgbm algorithm with CID (complexity invariant distance). Our failure prediction model has been built and evaluated on SMART data of about 80,000 hard disks from two manufacturers. The experimental result shows that by adding CID features, the TPR is increased from 0.28 to 0.96, and the number of days that the model can predict failures in advance is extended by 1.2 days. Compared with the several existing failure prediction models, our model has better performance on AUC score, f1-score and TPR. Hengrui Wang, Yahui Yang, Hongzhang Yang |
ISCC | 1 |