VLDB 2026 Research / reviewers in the wild / expert
Weigang Li 0002
dblp:285/2485-2
· DBLP profile ↗
11ranked-venue papers
2as first author
5since 2021 · last 2024
0009-0004-8932-5801ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 since 2021Computer networks · 3 · 2 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | HD-IOV: SW-HW Co-designed I/O Virtualization with Scalability and Flexibility for Hyper-Density CloudabstractAs the resource density of cloud servers increases, cloud providers deploy hundreds of VMs concurrently on a single server, requiring a high-performance, scalable, flexible and high-density I/O virtualization method. Hardware assisted virtualization such as device pass-through with SR-IOV can achieve near-native performance, however, at the expense of flexibility and a limited device count. Traditional software-based I/O virtualization systems tend to dedicate additional computing cores for higher performance, but suffer from critical scalability problems especially in high-density cloud. Zongpu Zhang, Jiangtao Chen, Banghao Ying, Yahui Cao, Lingyu Liu, Jian Li 0021, Weigang Li 0002, Haibing Guan |
EuroSys | 9 |
| 2024 | vCrypto: a Unified Para-Virtualization Framework for Heterogeneous Cryptographic ResourcesabstractTransport Layer Security (TLS) connections involve costly cryptographic operations which incur significant resource consumption in the cloud. Hardware accelerators are affordable substitutes of expensive CPU cores to accommodate with the constantly increasing security requirements of datacenters. Existing accelerators virtualization mainly relies on passthrough of Single Root I/O Virtualization (SR-IOV) devices. However, deficiency of service accessibility, functionality and availability make device passthrough not an optimal solution for heterogeneous accelerators with different capabilities. To make up the gap, we propose vCrypto, a unified para-virtualization framework for heterogeneous cryptographic resources. vCrypto supports stateful crypto requests offloading and result retrieval with session lifecycle management and event driven notification. vCrypto transparently integrates virtual crypto device capabilities into the OpenSSL framework to benefit existing applications that are based on crypto library APIs without modification. Multiple physical resources can be partitioned flexibly and scheduled cooperatively to enhance the functionality, performance and robustness of virtual crypto service. Finally, vCrypto achieves an optimized performance with two layers polling and memory sharing mechanism. The comprehensive experiments show that with the same cryptographic resources used, vCrypto framework can provide 2.59x to 3.36x higher AES-CBC-HMAC-SHA1 throughput compared to passthrough SR-IOV device. Chao Zhang 0115, Zongpu Zhang, Hubin Zhang, Weigang Li 0002, Yibin Shen, Jian Li 0021, Haibing Guan |
INFOCOM | 6 |
| 2023 | QKPT: Securing Your Private Keys in Cloud With Performance, Scalability and TransparencyabstractPrivate key (e.g., RSA key) protection is a significant issue for cloud but existing keyless or keyguard solutions suffer from performance, elasticity or applicability limitations. Recently, represented by Intel KPT, a novel keyguard architecture emerges to combine trusted platform module and crypto accelerator for achieving both security and performance. However, the straight use of KPT for private key protection may not be a good fit in cloud as it incurs challenges on protection capacity, key provisioning latency and transparency. Based on KPT-like hardware, we propose QKPT, a comprehensive key management system to bring your own private keys (BYOPK) into multi-tenant clouds. QKPT introduces a carefully-designed key wrapping layer to overcome these challenges. A small symmetric wrapping key (SWK) is generated for each tenant as the master key to resolve the former two challenges, while a special private key wrapping scheme is adopted to resolve the transparency limitation. Additionally, QKPT incorporates certificate trust to enhance the security of the SWK lifecycle and provides a hardened key server solution without expensive HSM. The evaluation shows that QKPT has a low runtime overhead ($\leq$1.2% for SSL/TLS handshakes) and still greatly outperforms the software baseline (3.5x-17x) owing to the crypto offloading. Zongpu Zhang, Hubin Zhang, Xiaokang Hu, Jian Li 0021, Weigang Li 0002, Guodong Zhu, Kapil Sood, Brian Will, Haibing Guan |
IEEE Trans. Dependable Secur. Comput. | 8 |
| 2021 | Multi-Task Hierarchical Learning Based Network Traffic AnalyticsabstractClassifying network traffic is the basis for important network applications. Prior research in this area has faced challenges on the availability of representative datasets, and many of the results cannot be readily reproduced. Such a problem is exacerbated by emerging data-driven machine learning based approaches. To address this issue, we present (Net)2database with three open datasets containing nearly 1.3M labeled flows in total, with a comprehensive list of flow features, for the research community1. We focus on broad aspects in network traffic analysis, including both malware detection and application classification. As we continue to grow them, we expect the datasets to serve as a common ground for AI driven, reproducible research on network flow analytics. We release the datasets publicly and also introduce a Multi-Task Hierarchical Learning (MTHL) model to perform all tasks in a single model. Our results show that MTHL is capable of accurately performing multiple tasks with hierarchical labeling with a dramatic reduction in training time. Onur Barut, Yan Luo 0001, Weigang Li 0002 |
ICC | 4 |
| 2021 | STYX: A Hierarchical Key Management System for Elastic Content Delivery Networks on Public CloudsabstractHosting content delivery networks (CDNs) on clouds has the potential to improve the performance as resources and caches can be placed closer to subscribers. However, avoiding data leakage over an untrusted public cloud is critical, especially for sensitive data such as the SSL private key. The popular Keyless SSL solution allows content owners to retain on-premise custody of SSL private keys on their own key servers, but this solution likely causes performance bottlenecks and impedes the elasticity of CDNs. This paper describes a novel key management system, named STYX, for transmitting trusted data over untrusted channels and storing them on untrusted platforms. STYX accomplishes secure key provisioning for CDN scale-out and the key is securely protected with full revocation rights for CDN scale-in. STYX is implemented as a three-phase hierarchical key management scheme by leveraging Intel Software Guard Extensions (SGX) and QuickAssist Technology (QAT). Furthermore, STYX supports CDN services by integrating Nginx as the SSL termination proxy and the popular Redis/Memcached/Apache as backend caching engines. The performance evaluation shows that STYX significantly outperforms the native HTTPS servers on the CDN node due to QAT acceleration, providing up to a 5× enhancement in throughput and a 50 percent reduction in latency. Xiaokang Hu, Jian Li 0021, Changzheng Wei, Weigang Li 0002, Haibing Guan |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2020 | ACETA: Accelerating Encrypted Traffic Analytics on Network EdgeabstractApplying machine learning techniques to detect malicious encrypted network traffic has become a challenging research topic. Traditional approaches based on studying network patterns fail to operate on encrypted data, especially without compromising the integrity of encryption. In addition, the requirement of rendering network-wide intelligent protection in a timely manner further exacerbates the problem. In this paper, we propose to leverage ×86 multicore platforms provisioned at enterprises' network edge with the software accelerators to design an encrypted traffic analytics (ETA) system with accelerated speed. Specifically, we explore a suite of data features and machine learning models with an open dataset. Then we show that by using Intel DAAL and OpenVINO libraries in model training and inference, we are able to reduce the training and inference time by a maximum order of 31× and 46× respectively while retaining the model accuracy. Derek Manning, Xiaoban Wu, Yan Luo 0001, Weigang Li 0002 |
ICC | 6 |
| 2020 | QWEB: High-Performance Event-Driven Web Architecture With QAT AccelerationabstractHardware accelerators have been a promising solution to reduce the cost of cloud datacenters. This article investigates the acceleration of an important datacenter workload: the web server (or proxy) that faces high computational consumption originated from SSL/TLS processing and HTTP compression. Our study reveals that for the widely-deployed event-driven web architecture, the straight offloading of SSL/TLS or compression tasks suffers from frequent blockings in the offload I/O, leading to the underutilization of both CPU and accelerator resources. To achieve efficient acceleration, we propose QWEB, a comprehensive offload solution based on Intel QuickAssist Technology (QAT). QWEB introduces an asynchronous offload mode for SSL/TLS processing and a pipelining offload mode for HTTP compression, both allowing concurrent offload tasks from a single application process/thread. With these two novel offload modes, the blocking penalty is amortized or even eliminated, and the utilization rate of the parallel computation engines inside the QAT accelerator is greatly increased. The evaluation shows that QWEB provides up to 9x handshake performance with TLS-RSA (2048-bit) over the software baseline. Additionally, the secure data transfer throughput is enhanced by 2x for the SSL/TLS offloading only, 3.5x for the compression offloading only and 5x for the combined offloading. Jian Li 0021, Xiaokang Hu, David Qian, Changzheng Wei, Gordon McFadden, Brian Will, Weigang Li 0002, Haibing Guan |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2019 | QZFS: QAT Accelerated Compression in File System for Application Agnostic and Cost Efficient Data Storage
Xiaokang Hu, Fuzong Wang, Weigang Li 0002, Jian Li 0021, Haibing Guan |
USENIX ATC | 3 |
| 2017 | STYX: a trusted and accelerated hierarchical SSL key management and distribution system for cloud based CDN applicationabstractProtecting the customer's SSL private key is the paramount issue to persuade the website owners to migrate their contents onto the cloud infrastructure, besides the advantages of cloud infrastructure in terms of flexibility, efficiency, scalability and elasticity. The emerging Keyless SSL solution retains on-premise custody of customers' SSL private keys on their own servers. However, it suffers from significant performance degradation and limited scalability, caused by the long distance connection to Key Server for each new coming end-user request. The performance improvements using persistent session and key caching onto cloud will degrade the key invulnerability and discourage the website owners because of the cloud's security bugs. Changzheng Wei, Jian Li 0021, Weigang Li 0002, Haibing Guan |
SoCC | 3 |
| 2017 | Optimize Genomics Data Compression with Hardware AcceleratorabstractGenomics is a Big Data science, the rate of increase in DNA sequencing is significantly exceeding the rate of increase in storage capacity, study shows the genomics data generation will exceed Twitter, YouTube, and astrophysics data combined by the year 2025. Storage and data management have become one of the most challenging bottlenecks in genomics and life sciences research. Data compression is an important technique to improve the efficiency of genomics data analysis and storage, and is widely deployed in the IT infrastructure in the life science institutes. In this paper we analyze the data compression characterization in the genomics workflows, and evaluate the performance & cost for different compression algorithms in the genomics analysis tool, we present a new hardware acceleration method that efficiently compresses DNA sequences with the reduced computation time and CPU utilization. Weigang Li 0002 |
DCC | 1 |
| 2016 | Accelerate Data Compression in File SystemabstractBTRFS (B-tree filesystem) is a new copy on write (CoW) filesystem for Linux aimed at implementing advanced features while focusing on fault tolerance, repair and easy administration. One important feature for BTRFS is the transparent file compression. There are two compression algorithms available in BTRFS: ZLIB and LZO. Because the compression workload is very CPU intensive, normally we have to pay additional CPU power to enable data compression in the filesystem. In this work we study the nature of data compression in BTRFS, setup the test bench to measure the benefit and the cost when enable data compression in BTRFS, and develop a new hardware acceleration method to offload the compression and de-compression workloads from the BTRFS software stack. Our experiment proves the hardware offloading solution can provide BTRFS data compression with the least CPU overhead comparing with the software implementation of ZLIB and LZO, meanwhile it also reaches a very good compression ratio and disk write throughput. Weigang Li 0002 |
DCC | 1 |