Wenhai Li

dblp:39/3944 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Boost Lock-free Queue and Stack with Batching
abstract
Concurrent queues and stacks are important components of many software systems. They are well-known contended data structures due to intense contention on hotspots, resulting in limited performance and poor scalability. To achieve higher performance, several attempts applied the batching technique, which packs a group of standard operations into a single batch for execution, to lock-free queues and stacks built upon the linked list, but still inherited the contended compare-and-swap (CAS).
Wenhai Li, Lingfeng Deng
PPoPP2
2025 PAB-Road: A Patch-Wise Boundary for Road Network Extraction via Multitask UNet
abstract
Road network extraction from remote sensing images is a fundamental task for applications like autonomous driving and urban planning. Mainstream methods, however, face a critical trade-off: segmentation-based approaches provide high geometric detail but often yield fragmented roadmaps, while graph-based approaches ensure connectivity but can sacrifice fine-grained accuracy. While hybrid models have been explored to resolve this, effectively fusing pixel-level features with structural information remains a key challenge. To address this, we propose PAB-Road, a novel framework with a unique fusion mechanism. Its core novelty is a multitask UNet that learns a patch-wise boundary representation to explicitly model local connectivity. This learned information then guides the synthesis of a geometrically accurate and structurally coherent road network. Experimental results in the real dataset show that PAB-Road achieves a compelling F1 score of 79.29%, demonstrating the effectiveness of our proposed fusion strategy.
Wenhai Li, Xianhong Zhu, Xiaohui Huang 0003, Xiaofei Yang 0002
IEEE Geosci. Remote. Sens. Lett.1
2025 FB+-tree: A Memory-Optimized B+-tree with Latch-Free Update
abstract
B + -trees are prevalent in traditional database systems due to their versatility and balanced structure. While binary search is typically utilized for branch operations, it may lead to inefficient cache utilization in main-memory scenarios. In contrast, trie-based index structures drive branch operations through prefix matching. While these structures generally produce fewer cache misses and are thus increasingly popular, they may underperform in range scans because of frequent pointer chasing. This paper proposes a new high-performance B + -tree variant called Feature B + -tree (FB + -tree ). Similar to employing bit or byte for branch operation in tries, FB + -tree progressively considers several bytes following the common prefix on each level of its inner nodes—referred to as features, which allows FB + -tree to benefit from prefix skewness. FB + -tree blurs the lines between B + -trees and tries, while still retaining balance. In the best case, FB + -tree almost becomes a trie, whereas in the worst case, it continues to function as a B + -tree. Meanwhile, a crafted synchronization protocol that combines the link technique and optimistic lock is designed to support efficient concurrent index access. Distinctively, FB + -tree leverages subtle atomic operations seamlessly coordinated with optimistic lock to facilitate latch-free updates, which can be easily extended to other structures. Intensive experiments on multiple workload-dataset combinations demonstrate that FB + -tree shows comparable lookup performance to state-of-the-art trie-based indexes and outperforms popular B + -trees by 2.3x ~ 3.7x under 96 threads. FB + -tree also exhibits significant potential on other workloads, especially update workloads under contention and scan workloads.
Wenhai Li, Lingfeng Deng
Proc. VLDB Endow.3
2024 High-Dimensional Feature Fault Diagnosis Method Based on HEFS-LGBM
Wenhai Li, Tianzhu Wen, Weichao Sun
J. Electron. Test.2
2023 Lock-Free Bucketized Cuckoo Hashing
abstract
Abstract Concurrent hash tables are one of the fundamental building blocks for cloud computing. In this paper, we introduce lock-free modifications to in-memory bucketized cuckoo hashing. We present a novel concurrent strategy in designing a lock-free hash table, called LFBCH, that paves the way towards scalability and high space efficiency. To the best of our knowledge, this is the first attempt to incorporate lock-free technology into in-memory bucketized cuckoo hashing, while still providing worst-case constant-scale lookup time and extremely high load factor. All of the operations over LFBCH, such as get, put, “kick out” and rehash, are guaranteed to be lock-free, without introducing notorious problems like false miss and duplicated key. The experimental results indicate that under mixed workloads with up to 64 threads, the throughput of LFBCH is 14%–360% higher than other popular concurrent hash tables.
Wenhai Li, Zhiling Cheng, Lingfeng Deng
Euro-Par1
2023 Accelerating large-scale weighted similarity queries based on external storage
Wenhai Li, Zhiling Cheng, Lingfeng Deng
Inf. Syst.1
2022 Delta Debugging Microservice Systems with Parallel Optimization
abstract
Microservice systems are complicated due to their runtime environments and service communications. Debugging a failure involves the deployment and manipulation of microservice systems on a containerized environment and faces unique challenges due to the high complexity and dynamism of microservices. To address these challenges, we propose a debugging approach for microservice systems based on the delta debugging algorithm, which is to minimalize failure-inducing deltas of circumstances (e.g., deployment, environmental configurations). Our approach includes novel techniques for defining, deploying/manipulating, and executing deltas during delta debugging. In particular, to construct a (failing) circumstance space for delta debugging to minimalize, our approach defines a set of circumstance dimensions that can affect the execution of microservice systems. To automate the testing of deltas, our approach includes the design of an infrastructure layer for automating deployment and manipulation of microservice systems. To optimize the delta debugging process, our approach includes the design of parallel execution for delta testing tasks. Our evaluation shows that our approach is scalable and efficient with the provided infrastructure resources and the designed parallel execution for optimization. Our experimental study on a medium-size microservice benchmark system shows that our approach can effectively identify failure-inducing deltas that help diagnose the root causes.
Xin Peng 0001, Tao Xie 0001, Jun Sun 0001, Wenhai Li
IEEE Trans. Serv. Comput.6
2021 Fault Analysis and Debugging of Microservice Systems: Industrial Survey, Benchmark System, and Empirical Study
abstract
The complexity and dynamism of microservice systems pose unique challenges to a variety of software engineering tasks such as fault analysis and debugging. In spite of the prevalence and importance of microservices in industry, there is limited research on the fault analysis and debugging of microservice systems. To fill this gap, we conduct an industrial survey to learn typical faults of microservice systems, current practice of debugging, and the challenges faced by developers in practice. We then develop a medium-size benchmark microservice system (being the largest and most complex open source microservice system within our knowledge) and replicate 22 industrial fault cases on it. Based on the benchmark system and the replicated fault cases, we conduct an empirical study to investigate the effectiveness of existing industrial debugging practices and whether they can be further improved by introducing the state-of-the-art tracing and visualization techniques for distributed systems. The results show that the current industrial practices of microservice debugging can be improved by employing proper tracing and visualization techniques and strategies. Our findings also suggest that there is a strong need for more intelligent trace analysis and visualization, e.g., by combining trace visualization and improved fault localization, and employing data-driven and learning-based recommendation for guided visual exploration and comparison of traces.
Xin Peng 0001, Tao Xie 0001, Jun Sun 0001, Wenhai Li
IEEE Trans. Software Eng.6
2020 Similarity query support in big data management systems
Taewoo Kim 0001, Wenhai Li, Alexander Behm, Inci Cetindil, Rares Vernica, Vinayak R. Borkar, Michael J. Carey 0001, Chen Li 0001
Inf. Syst.2
2019 CORES: Towards Scan-Optimized Columnar Storage for Nested Records
abstract
The relatively high cost of record deserialization is increasingly becoming the bottleneck of column-based storage systems in tree-structured applications [58]. Due to record transformation in the storage layer, unnecessary processing costs derived from fields and rows irrelevant to queries may be very heavy in nested schemas, significantly wasting the computational resources in large-scale analytical workloads. This leads to the question of how to reduce both the deserialization and IO costs of queries with highly selective filters following arbitrary paths in a nested schema. We present CORES (Column-Oriented Regeneration Embedding Scheme) to push highly selective filters down into column-based storage engines, where each filter consists of several filtering conditions on a field. By applying highly selective filters in the storage layer, we demonstrate that both the deserialization and IO costs could be significantly reduced. We show how to introduce fine-grained composition on filtering results. We generalize this technique by two pair-wise operations, rollup and drilldown, such that a series of conjunctive filters can effectively deliver their payloads in nested schema. The proposed methods are implemented on an open-source platform. For practical purposes, we highlight how to build a column storage engine and how to drive a query efficiently based on a cost model. We apply this design to the nested relational model especially when hierarchical entities are frequently required by ad hoc queries. The experiments, including a real workload and the modified TPCH benchmark, demonstrate that CORES improves the performance by 0.7×--26.9× compared to state-of-the-art platforms in scan-intensive workloads.
Weidong Wen, Wenhai Li, Lingfeng Deng, Yanxiang He
ACM Trans. Storage3
2018 Supporting Similarity Queries in Apache AsterixDB
Taewoo Kim 0001, Wenhai Li, Alexander Behm, Inci Cetindil, Rares Vernica, Vinayak R. Borkar, Michael J. Carey 0001, Chen Li 0001
EDBT2
2018 Delta debugging microservice systems
abstract
Debugging microservice systems involves the deployment and manipulation of microservice systems on a containerized environment and faces unique challenges due to the high complexity and dynamism of microservices. To address these challenges, in this paper, we propose a debugging approach for microservice systems based on the delta debugging algorithm, which is to minimize failureinducing deltas of circumstances (e.g., deployment, environmental configurations) for effective debugging. Our approach includes novel techniques for defining, deploying/manipulating, and executing deltas following the idea of delta debugging. In particular, to construct a (failing) circumstance space for delta debugging to minimize, our approach defines a set of dimensions that can affect the execution of microservice systems. Our experimental study on a medium-size microservice benchmark system shows that our approach can effectively identify failure-inducing deltas that help diagnose the root causes.
Xin Peng 0001, Tao Xie 0001, Jun Sun 0001, Wenhai Li
ASE5
2018 ZigZag: Supporting Similarity Queries on Vector Space Models
abstract
In this paper we study the problem of supporting similarity queries on a large number of records using a vector space model, where each record is a bag of tokens. We consider similarity functions that incorporate non-negative global token weights as well as record-specific token degrees. We develop a family of algorithms based on an inverted index for large data sets, especially for the case of using external storage such as hard disks or flash drives, and present pruning techniques based on various bounds to improve their performance. We formally prove the correctness of these techniques, and show how to achieve better pruning power by iteratively tightening these bounds to exactly filter dissimilar records. We conduct an extensive experimental study using real, large-scale data sets based on different storage platforms, including memory, hard disks, and flash drives. The results show that these algorithms and techniques can efficiently support similarity queries on large data sets.
Wenhai Li, Lingfeng Deng, Chen Li 0001
SIGMOD Conference1
2017 Massive spatial query on the Kepler architecture
abstract
In this paper, we present an optimized framework that can efficiently perform massive spatial queries on the current GPUs. To benefit the widely adopted filter-and-verify paradigm from GPUs, the skewed workloads are first associated with certain cells in a scaled spatial grid, such that the following range verification cost against the massive spatial objects can be significantly reduced. Particularly on the Kepler architecture, we highlight a two-level scheduling method to exploit good data localities by developing a novel dynamic scheduling method. Based on this virtual warp-based scheduling method, groups of threads can compete for the unbalanced tasks to ensure good load balance. We conduct various of skewed workloads with different object positions and query distributions, to evaluate our optimized methods. Experimental results show that, as compared to the existing fixed-size allocation methods, the proposed adaptive scheduling strategies improve the query throughput by one order of magnitude.
Yili Gong, Wenhai Li, Zihui Ye
ASAP3
2014 Fault Models of CMOS Gates: An Empirical Study Based on Mutation Analysis
abstract
An empirical study is conducted to analyze the fault models of CMOS gates in VLSI. The mutation analysis methodology is used to ensure the simulation experiments are sufficient enough. For the primitive CMOS logic cells: inverter, 2-input NAND, NOR and XOR gates, a total of 335 mutant circuits are generated, simulated and analyzed. Comparing with the previous work, several new fault models are revealed, such as other-logic and the indetermination fault categorys sub-categories: 0-to-X, 1-to-X and other-to-X. Besides, the experimental results show that the classic static stuck-at fault cannot cover many practical defects in a circuit since it just averagely accounts for less than 15% of the total faults. Many other useful conclusions are drawn, which may benefit the related fault-based applications.
Xiaofeng Tang, Aiqiang Xu, Wenhai Li, Zhiyong Yang 0006
DASC3
2013 SHOE: A SPARQL Query Engine Using MapReduce
abstract
As the reasoning aspects and the knowledge based processing capabilities of RDF (Resource Description Framework) have been widely adopted in W3C Recommendation, the ontology layer and query languages of the Semantic Web stack achieve a certain level of maturity. There exists an increasing need for high performance, read-only semantic analysis for the massive RDF data. In this demo we will present a Map-Reduce based SPARQL processing engine, called SHOE (SPARQL on Hadoop with Optimization Encoding), to handle billions of RDF triples. SHOE consists of three major components (1) the RDF data loader, (2) the partition generator and (3) the query processor. While this demonstration mainly focuses on enhancing SPARQL processing in the Hadoop platform, the underlying encoding and partitioning optimization strategies can be utilized by the common Map-Reduce frameworks in the share-nothing environment.
Wenhai Li, Biren Chen, Ruijiang Yao, Weidong Wen, Chungwai Cheung, Wanghong Li
ICPADS1
2013 Towards Multi-way Join Evaluating with Indexing Partition Support in Map-Reduce
abstract
In the era of "big data", the emergence and increasing adoptions of the related enabling technologies make it possible for Map-Reduce to accommodate DSS (Decision Support Systems) load, which is commonly targeted for high-performance Data Warehouse analyses in the context of RDBMS. However, the non-predetermined mapping of the Map-Reduce tasks to the physical machines makes it difficult to utilize the pre-partitioned and indexing techniques of DBMS to improve the data locality. In this paper, towards multi-way join evaluating OLAP (Online Analysis Processing) workloads, we introduce table partitioning by reference to Map-Reduce. For avoiding the dispersion of the initial tuples that belong to the same segment keys, we present a detailed description of the data organization model that partitions the dominated tables by cascade reference constraints. In order to push multiple joins on these clustered partitions down to the map task, we design a one-pass multi-way join algorithm along with its optimization implementations for the major Map-Reduce stages. We conduct an empirically study with TPCH benchmark on different scales of clusters, and experimentally verify the high efficiency of the proposed optimization model.
Wenhai Li, Biren Chen, Weidong Wen, Wanghong Li
ICPADS2
2010 On Sliding Window Based Change Point Detection for Hybrid SIP DoS Attack
abstract
As the core infrastructure of the VoIP, IMS and IPTV, SIP based network is now increasingly been deployed throughout the world. Due mainly to the relatively high flow rate and the exorbitant session maintenance, SIP servers are similarly susceptible to the Denial-of-Service (DoS) Attacks above the IP stack, especially when the Distributed spoofing URI is considered. A hybrid SIP DoS detection method is proposed in this paper to handle three representative flooding attacks, and the CUSUM statistical detection founded on the adaptive sliding window based acquisition is adopted for achieving high accuracy and low latency. Experimental result shows that this method achieves a high rate, low latency and low false alarm rate of SIP flooding detection.
Wenhai Li, Xiaolei Luo
APSCC1
2006 Reduced Attribute Oriented Inconsistency Handling in Decision Generation
Yucai Feng, Wenhai Li, Zehua Lv, Xiaoming Ma
IDEAL2
1988 A new adaptive echo canceller used in the ISDN U-interface
abstract
A good adaptive echo canceller (EC), in terms of its performance, is an interesting problem for those who are engaged in the ISDN (integrated-services digital network) U-interface subscriber loop design. It is also important to implement an EC with low cost chips. An adaptive EC is proposed that makes its contribution in this respect. The filter combines some advantages possessed by common types of FIR (finite-impulse response) adaptive filters, transversal, and table lookup filters, together. In implementation, the filter design is mainly based on developing DSP (digital-signal-processor) chips, e.g. the Texas Instruments TMS320 family of microprocessors. The proposed architecture not only has a faster convergence speed than that of usual table lookup adaptive filters, but can also compensate the long period of echo tail caused by insufficient filter taps, especially when it is used in the full-duplex two-wire data transmission system with a net rate of 144 kb/s.>
Dali Yang, Wenhai Li
ICASSP2