EDBT 2026 Demo / reviewers in the wild / expert
Luwei Cheng
dblp:43/10334
· DBLP profile ↗
17ranked-venue papers
7as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-authorComputer networks · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 2Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
10 papers |
Cloud and datacenter computing · 58% Parallel and multicore computing · 33% GPUs and heterogeneous computing · 5% | |
| Network and information security
1 paper |
Authentication and access control · 100% | |
| Software engineering, system software, and programming languages
4 papers |
Operating systems · 100% | |
| Computer networks
3 papers |
Transport protocols and congestion control · 63% Datacenter networks · 37% | |
| Databases, data mining, and information retrieval
2 papers |
Data stream processing · 100% |
Topics — the 22 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing › virtualization › virtual machine management
virtual machine scheduling |
0.9 | 5 | 2018 | Effectively Mitigating I/O Inactivity in vCPU Scheduling · USENIX ATC 2018 Revisiting TCP Congestion Control in a Virtual Cluster Environment · IEEE/ACM Trans. Netw. 2016 vScale: automatic and efficient processor scaling for SMP virtual machines · EuroSys 2016 |
Authentication and access control
password authentication |
0.9 | 1 | 2025 | Using Parallel Techniques to Accelerate PCFG-Based Password Cracking Attacks · IEEE Trans. Dependable Secur. Comput. 2025 |
Authentication and access control
password guessing |
0.9 | 1 | 2025 | Using Parallel Techniques to Accelerate PCFG-Based Password Cracking Attacks · IEEE Trans. Dependable Secur. Comput. 2025 |
Cloud and datacenter computing
virtualization |
0.6 | 4 | 2016 | Offloading Interrupt Load Balancing from SMP Virtual Machines to the Hypervisor · IEEE Trans. Parallel Distributed Syst. 2016 Revisiting TCP Congestion Control in a Virtual Cluster Environment · IEEE/ACM Trans. Netw. 2016 PVTCP: Towards practical and effective congestion control in virtualized datacenters · ICNP 2013 |
Parallel and multicore computing › synchronization
NUMA-aware locking |
0.5 | 2 | 2017 | Scalable Adaptive NUMA-Aware Lock · IEEE Trans. Parallel Distributed Syst. 2017 Scalable adaptive NUMA-aware lock: combining local locking and remote locking for efficient concurrency · PPoPP 2016 |
Cloud and datacenter computing
autoscaling |
0.4 | 1 | 2020 | Turbine: Facebook's Service Management Platform for Stream Processing · ICDE 2020 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.4 | 1 | 2020 | Turbine: Facebook's Service Management Platform for Stream Processing · ICDE 2020 |
Parallel and multicore computing
task scheduling |
0.4 | 1 | 2020 | Turbine: Facebook's Service Management Platform for Stream Processing · ICDE 2020 |
Operating systems › resource management › process management
CPU scheduling |
0.4 | 2 | 2018 | Effectively Mitigating I/O Inactivity in vCPU Scheduling · USENIX ATC 2018 vScale: automatic and efficient processor scaling for SMP virtual machines · EuroSys 2016 |
Data stream processing
stream join |
0.3 | 1 | 2018 | Providing Streaming Joins as a Service at Facebook · Proc. VLDB Endow. 2018 |
Cloud and datacenter computing › cluster resource management and scheduling › resource scheduling
vCPU scheduling |
0.3 | 1 | 2018 | Effectively Mitigating I/O Inactivity in vCPU Scheduling · USENIX ATC 2018 |
Cloud and datacenter computing › datacenter architecture
virtualized datacenter |
0.3 | 2 | 2013 | PVTCP: Towards practical and effective congestion control in virtualized datacenters · ICNP 2013 Rethinking congestion control in virtualized datacenters · ICNP 2013 |
Parallel and multicore computing
synchronization |
0.3 | 1 | 2017 | Scalable Adaptive NUMA-Aware Lock · IEEE Trans. Parallel Distributed Syst. 2017 |
Transport protocols and congestion control
TCP congestion control |
0.2 | 1 | 2016 | Revisiting TCP Congestion Control in a Virtual Cluster Environment · IEEE/ACM Trans. Netw. 2016 |
Operating systems › kernel
interrupt handling |
0.2 | 1 | 2016 | Offloading Interrupt Load Balancing from SMP Virtual Machines to the Hypervisor · IEEE Trans. Parallel Distributed Syst. 2016 |
Parallel and multicore computing
concurrent data structures |
0.2 | 1 | 2016 | Scalable adaptive NUMA-aware lock: combining local locking and remote locking for efficient concurrency · PPoPP 2016 |
Memory systems
non-uniform memory access |
0.2 | 1 | 2016 | Scalable adaptive NUMA-aware lock: combining local locking and remote locking for efficient concurrency · PPoPP 2016 |
Cloud and datacenter computing › virtualization
virtual machine |
0.2 | 1 | 2016 | Offloading Interrupt Load Balancing from SMP Virtual Machines to the Hypervisor · IEEE Trans. Parallel Distributed Syst. 2016 |
Datacenter networks
incast |
0.2 | 2 | 2016 | PVTCP: Towards practical and effective congestion control in virtualized datacenters · ICNP 2013 Revisiting TCP Congestion Control in a Virtual Cluster Environment · IEEE/ACM Trans. Netw. 2016 |
Data stream processing
stream processing systems |
0.1 | 1 | 2020 | Turbine: Facebook's Service Management Platform for Stream Processing · ICDE 2020 |
Data stream processing › continuous query processing
streaming SQL |
0.1 | 1 | 2018 | Providing Streaming Joins as a Service at Facebook · Proc. VLDB Endow. 2018 |
Parallel and multicore computing › parallel architecture
NUMA multicore |
0.1 | 1 | 2017 | Scalable Adaptive NUMA-Aware Lock · IEEE Trans. Parallel Distributed Syst. 2017 |
Methods — techniques the papers use, named apart from their topics
parallelization · 1.7memory load balancing · 1.7paravirtualization · 1.2GPU kernels · 0.9GPU kernel · 0.9predictive auto scaling · 0.9fault tolerance · 0.9i/o-aware scheduling · 0.7adaptive lock switching · 0.6NUMA-aware server selection · 0.6dynamic vCPU resizing · 0.5streaming join operator · 0.3adaptive synchronization · 0.3testbed experimentation · 0.2physical CPU cycle monitoring · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Using Parallel Techniques to Accelerate PCFG-Based Password Cracking AttacksabstractTextual passwords play an important role among access-control mechanisms and are usually stored as ciphertext in the server. However, an attacker may attempt to hash a large number of candidate passwords to find the match of the target hash of a password database. To crack the password database, attackers in industry usually use the cracking software like Hashcat. Academic researchers recently proposed many data-driven probabilistic models, in which the Probabilistic Context-free Grammars (PCFG, for short) stand out. Despite the great cracking efficiency, the data-driven models are seldom used by industrial practice due to the significant slow generation speed of password candidates. To bridge the gap and promote the efficient data-driven models being practically used in industry, we propose that using parallel techniques to accelerate the candidate password generation, enabling the integration of PCFG models into the practically-used Hashcat tool. To this end, we mainly propose two algorithms to accelerate the password generation for PCFG-based models: first, we design a storage structure with the memory load balance strategy to more evenly store the data structures used to generate passwords; second, we design an algorithm to produce candidate passwords in parallel by different threads. Based on the two algorithms, we proposeParallel_PCFG, and implementParallel_PCFGupon Hashcat based on its built-in GPU kernel. We comprehensively evaluateParallel_PCFGagainst state-of-the-art data-driven models, and find thatParallel_PCFGonly takes 14.33% of time to achieve the same cracking rates compared with the best-performing models, paving a way about the integration between PCFG-based models and Hashcat. Ming Xu 0006, Kai Zhang 0006, Jitao Yu, Luwei Cheng, Weili Han |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2020 | Turbine: Facebook's Service Management Platform for Stream ProcessingabstractThe demand for stream processing at Facebook has grown as services increasingly rely on real-time signals to speed up decisions and actions. Emerging real-time applications require strict Service Level Objectives (SLOs) with low downtime and processing lag-even in the presence of failures and load variability. Addressing this challenge at Facebook scale led to the development of Turbine, a management platform designed to bridge the gap between the capabilities of the existing general-purpose cluster management frameworks and Facebook's stream processing requirements. Specifically, Turbine features a fast and scalable task scheduler; an efficient predictive auto scaler; and an application update mechanism that provides fault-tolerance, atomicity, consistency, isolation and durability. Turbine has been in production for over three years, and one of the core technologies that enabled a booming growth of stream processing at Facebook. It is currently deployed on clusters spanning tens of thousands of machines, managing several thousands of streaming pipelines processing terabytes of data per second in real time. Our production experience has validated Turbine's effectiveness: its task scheduler evenly balances workload fluctuation across clusters; its auto scaler effectively and predictively handles unplanned load spikes; and the application update mechanism consistently and efficiently completes high scale updates within minutes. This paper describes the Turbine architecture, discusses the design choices behind it, and shares several case studies demonstrating Turbine capabilities in production. Luwei Cheng, Vanish Talwar, Michael Y. Levin, Gabriela Jacques-Silva, Nikhil Simha, Anirban Banerjee, Tim Williamson, Serhat Yilmaz, Guoqiang Jerry Chen |
ICDE | 2 |
| 2018 | Effectively Mitigating I/O Inactivity in vCPU Scheduling
Weiwei Jia 0001, Cheng Wang 0021, Xusheng Chen, Jianchen Shan, Xiaowei Shang, Heming Cui, Xiaoning Ding, Luwei Cheng, Francis C. M. Lau 0001, Yuangang Wang |
USENIX ATC | 8 |
| 2018 | Providing Streaming Joins as a Service at FacebookabstractStream processing applications reduce the latency of batch data pipelines and enable engineers to quickly identify production issues. Many times, a service can log data to distinct streams, even if they relate to the same real-world event (e.g., a search on Facebook's search bar). Furthermore, the logging of related events can appear on the server side with different delay, causing one stream to be significantly behind the other in terms of logged event times for a given log entry. To be able to stitch this information together with low latency , we need to be able to join two different streams where each stream may have its own characteristics regarding the degree in which its data is out-of-order . Doing so in a streaming fashion is challenging as a join operator consumes lots of memory, especially with significant data volumes. This paper describes an end-to-end streaming join service that addresses the challenges above through a streaming join operator that uses an adaptive stream synchronization algorithm that is able to handle the different distributions we observe in real-world streams regarding their event times. This synchronization scheme paces the parsing of new data and reduces overall operator memory footprint while still providing high accuracy. We have integrated this into a streaming SQL system and have successfully reduced the latency of several batch pipelines using this approach. Gabriela Jacques-Silva, Ran Lei, Luwei Cheng, Guoqiang Jerry Chen, Kuen Ching, Tanji Hu, Kevin Wilfong, Rithin Shetty, Serhat Yilmaz, Anirban Banerjee, Benjamin Heintz, Shridhar Iyer, Anshul Jaiswal |
Proc. VLDB Endow. | 3 |
| 2017 | Preserving I/O prioritization in virtualized OSesabstractWhile virtualization helps to enable multi-tenancy in data centers, it introduces new challenges to the resource management in traditional OSes. We find that one important design in an OS, prioritizing interactive and I/O-bound workloads, can become ineffective in a virtualized OS. Resource multiplexing between multiple tenants breaks the assumption of continuous CPU availability in physical systems and causes two types of priority inversions in virtualized OSes. In this paper, we present xBalloon, a lightweight approach to preserving I/O prioritization. It uses a balloon process in the virtualized OS to avoid priority inversion in both short-term and long-term scheduling. Experiments in a local Xen environment and Amazon EC2 show that xBalloon improves I/O performance in a recent Linux kernel by as much as 136% on network throughput, 95% on disk throughput, and 125x on network tail latency. Kun Suo, Jia Rao, Luwei Cheng, Xiaobo Zhou 0002, Francis C. M. Lau 0001 |
SoCC | 4 |
| 2017 | Scheduler activations for interference-resilient SMP virtual machine schedulingabstractThe wide adoption of SMP virtual machines (VMs) and resource consolidation present challenges to efficiently executing multi-threaded programs in the cloud. An important problem is the semantic gaps between the guest OS and the hypervisor. The well-known lock-holder preemption (LHP) and lock-waiter preemption (LWP) problems are examples of such semantic gaps, in which the hypervisor is unaware of the activities in the guest OS and adversely deschedules virtual CPUs (vCPUs) that are executing in critical sections. Existing studies have focused on inferring a high-level semantic state of the guest OS to aid hypervisor-level scheduling so as to avoid the LHP and LWP problems. Kun Suo, Luwei Cheng, Jia Rao |
Middleware | 3 |
| 2017 | Scalable Adaptive NUMA-Aware LockabstractScalable locking is a key building block for scalable multi-threaded software. Its performance is especially critical in multi-socket, multi-core machines with non-uniform memory access (NUMA). Previous schemes such as in-place locks and delegation locks only perform well under a certain level of contention, and often require non-trivial tuning for a particular configuration. Besides, in large NUMA systems, current delegation locks cannot perform satisfactorily due to lack of optimized NUMA policies. In this work, we propose SANL, a locking scheme that can deliver high performance under various contention levels by adaptively switching between in-place locks and delegation locks. To optimize the performance of delegation locks, we introduce a new NUMA policy that jointly considers node distances and server utilization when choosing lock servers. We have implemented SANL and evaluated it with four popular multi-threaded applications (Memcached, Berkeley DB, Phoenix2 and SPLASH-2), on a 40-core Intel machine and a 64-core AMD machine. The comparison results with seven other representative locking schemes show that SANL outperforms them in most contention situations. For example, in one group test, SANL is 3.7 times faster than RCL lock and 17 times faster than POSIX mutex. Haibo Chen 0001, Luwei Cheng, Francis C. M. Lau 0001, Cho-Li Wang |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2016 | vScale: automatic and efficient processor scaling for SMP virtual machinesabstractSMP virtual machines (VMs) have been deployed extensively in clouds to host multithreaded applications. A widely known problem is that when CPUs are oversubscribed, the scheduling delays due to VM preemption give rise to many performance problems because of the impact of these delays on thread synchronization and I/O efficiency. Dynamically changing the number of virtual CPUs (vCPUs) by considering the available physical CPU (pCPU) cycles has been shown to be a promising approach. Unfortunately, there are currently no efficient mechanisms to support such vCPU-level elasticity. Luwei Cheng, Jia Rao, Francis C. M. Lau 0001 |
EuroSys | 1 |
| 2016 | Scalable adaptive NUMA-aware lock: combining local locking and remote locking for efficient concurrencyabstractScalable locking is a key building block for scalable multi-threaded software. Its performance is especially critical in multi-socket, multi-core machines with non-uniform memory access (NUMA). Previous schemes such as local locking and remote locking only perform well under a certain level of contention, and often require non-trivial tuning for a particular configuration. Besides, for large NUMA systems, because of unmanaged lock server's nomination, current distance-first NUMA policies cannot perform satisfactorily. Francis C. M. Lau 0001, Cho-Li Wang, Luwei Cheng, Haibo Chen 0001 |
PPoPP | 4 |
| 2016 | Revisiting TCP Congestion Control in a Virtual Cluster EnvironmentabstractVirtual machines (VMs) are widely adopted today to provide elastic computing services in datacenters, and they still heavily rely on TCP for congestion control. VM scheduling delays due to CPU sharing can cause frequent spurious retransmit timeouts (RTOs). Using current detection methods, we find that such spurious RTOs cannot be effectively identified because of the retransmission ambiguity caused by the delayed ACK (DelACK) mechanism. Disabling DelACK would add significant CPU overhead to the VMs and thus degrade the network's performance. In this paper, we first report our practical experience about TCP's reaction to VM scheduling delays. We then provide an analysis of the problem that has two components corresponding to VM preemption on the sender side and the receiver side, respectively. Finally, we propose PVTCP, a ParaVirtualized approach to counteract the distortion of congestion information caused by the hypervisor scheduler. PVTCP is completely embedded in the guest OS and requires no modification in the hypervisor. Taking incast congestion as an example, we evaluate our solution in a 21-node testbed. The results show that PVTCP has high adaptability in virtualized environments and deals satisfactorily with the throughput collapse problem. Luwei Cheng, Francis C. M. Lau 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2016 | Offloading Interrupt Load Balancing from SMP Virtual Machines to the HypervisorabstractCloud computing increasingly leverages SMP virtual machines (VMs) to host multi-threaded applications. Interrupt balancing as a problem becomes more challenging because VMs are subject to the hypervisor's scheduling. Since the scheduling delays are typically tens of milliseconds, when they are added to one VM's interrupt delivery, they can seriously degrade the VM's I/O performance. Traditional balancing techniques are designed for dedicated environments, which cannot work well in virtualized environments because VMs are disallowed to directly control the hardware in many cases. In this paper, we present hBalance, a very simple approach to offload interrupt load balancing from SMP-VMs to the hypervisor. To accelerate the interrupt processing, our approach does not require shortening the hypervisor's scheduling time slice, but dynamically redirects interrupts from preempted virtual CPUs to running ones in a balanced manner. hBalance supports both Fully Virtualiized (FV) guests and Para-Virtualized (PV) guests, and exhibits high portability among various hypervisors. With our prototype implementation in Xen, the experimental results with both micro-level and application-level benchmarks show that hBalance significantly improves SMP-VMs' I/O performance while introduces moderate overhead. Luwei Cheng, Francis C. M. Lau 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | Rethinking congestion control in virtualized datacentersabstractCloud datacenters are increasingly adopting virtual machines (VMs) to provide elastic cloud services, with TCP being prevalently used for congestion control. In virtualized datacenters, the delays from the hypervisor scheduler can heavily contaminate RTTs sensed by VM senders, preventing TCP from correctly learning the physical network condition. In this dissertation, my direction is to paravirtualize the transport-layer protocol in the guest OS, making it automatically tolerate the virtualized running environment. I then present a preliminary solution, PVTCP, to overcome the distorted congestion information caused by VM scheduling delays. Luwei Cheng |
ICNP | 1 |
| 2013 | PVTCP: Towards practical and effective congestion control in virtualized datacentersabstractWhile modern datacenters are increasingly adopting virtual machines (VMs) to provide elastic cloud services, they still rely on traditional TCP for congestion control. In virtualized datacenters, TCP endpoints are separated by a virtualization layer and subject to the intervention of the hypervisor's scheduling. Most previous attempts focused on tuning the hypervisor layer to try to improve the VMs' I/O performance, and there is very little work on how a VM's guest OS may help the transport layer to adapt to the virtualized environment. In this paper, we find that VM scheduling delays can heavily contaminate RTTs as sensed by VM senders, preventing TCP from correctly learning the physical network condition. After giving an account of the source of the problem, we propose PVTCP, a ParaVirtualized TCP to counter the distorted congestion information caused by VM scheduling on the sender side. PVTCP is self-contained, requiring no modification to the hypervisor. Experiments show that PVTCP is much more effective in addressing incast congestion in virtualized datacenters than standard TCP. Luwei Cheng, Cho-Li Wang, Francis C. M. Lau 0001 |
ICNP | 1 |
| 2013 | Network performance isolation for latency-sensitive cloud applications
Luwei Cheng, Cho-Li Wang |
Future Gener. Comput. Syst. | 1 |
| 2012 | vBalance: using interrupt load balance to improve I/O performance for SMP virtual machinesabstractA Symmetric MultiProcessing (SMP) virtual machine (VM) enables users to take advantage of a multiprocessor infrastructure in supporting scalable job throughput and request responsiveness. It is known that hypervisor scheduling activities can heavily degrade a VM's I/O performance, as the scheduling latencies of the virtual CPU (vCPU) eventually translates into the processing delays of the VM's I/O events. As for a UniProcessor (UP) VM, since all its interrupts are bound to the only vCPU, it completely relies on the hypervisor's help to shorten I/O processing delays, making the hypervisor increasingly complicated. Regarding SMP-VMs, most researches ignore the fact that the problem can be greatly mitigated at the level of guest OS, instead of imposing all scheduling pressure on the hypervisor. Luwei Cheng, Cho-Li Wang |
SoCC | 1 |
| 2011 | Probabilistic Best-Fit Multi-dimensional Range Query in Self-Organizing CloudabstractWith virtual machine (VM) technology being increasingly mature, computing resources in modern Cloud systems can be partitioned in fine granularity and allocated on demand with "pay-as-you-go" model. In this work, we study the resource query and allocation problems in a Self-Organizing Cloud (SOC), where host machines are connected by a peer-to-peer (P2P) overlay network on the Internet. To run a user task in SOC, the requester needs to perform a multi-dimensional range search over the P2P network for locating host machines that satisfy its minimal demand on each type of resources. The multi-dimensional range search problem is known to be challenging as contentions along multiple dimensions could happen in the presence of the uncoordinated analogous queries. Moreover, low resource matching rate may happen while restricting query delay and network traffic. We design a novel resource discovery protocol, namely Proactive Index Diffusion CAN (PID-CAN), which can proactively diffuse resource indexes over the nodes and randomly route query messages among them. Such a protocol is especially suitable for the range query that needs to maximize its best-fit resource shares under possible competition along multiple resource dimensions. Via simulation, we show that PID-CAN could keep stable and optimized searching performance with low query delay and traffic overhead, for various test cases under different distributions of query ranges and competition degrees. It also performs satisfactorily in dynamic node-churning situation. Sheng Di, Cho-Li Wang, Weida Zhang, Luwei Cheng |
ICPP | 4 |
| 2011 | WAVNet: Wide-Area Network Virtualization Technique for Virtual Private CloudabstractA Virtual Private Cloud (VPC) is a secure collection of computing, storage and network resources spanning multiple sites over Wide Area Network (WAN). With VPC, computation and services are no longer restricted to a fixed site but can be relocated dynamically across geographical sites to improve manageability, performance and fault tolerance. We propose WAVNet, a layer 2 virtual private network (VPN) which supports virtual machine live migration over WAN to realize mobility of execution environment across multiple security domains. WAVNet adopts a UDP hole punching technique to achieve direct network connection between two Internet hosts without special router configuration. We evaluate our design in an emulated WAN with 64 hosts and also in a real WAN environment with 10 machines located at seven different sites across the Asia-Pacific region. The experimental results show that WAVNet not only achieves close-to-native host-to-host network bandwidth and latency, but also guarantees more effective VM live migration than existing solutions. Zheming Xu, Sheng Di, Weida Zhang, Luwei Cheng, Cho-Li Wang |
ICPP | 4 |