VLDB 2026 Research / reviewers in the wild / expert
Guokuan Li
dblp:82/11194
· DBLP profile ↗
17ranked-venue papers
0as first author
17since 2021 · last 2026
0009-0005-7998-5520ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 9 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Vista: Scene-Aware Optimization for Streaming Video Question Answering Under Post-Hoc QueriesabstractStreaming video question answering (Streaming Video QA) poses distinct challenges for multimodal large language models (MLLMs), as video frames arrive sequentially and user queries can be issued at arbitrary timepoints. Existing solutions relying on fixed-size memory or naive compression often suffer from context loss or memory overflow, limiting their effectiveness in long-form, real-time scenarios.We present Vista, a novel framework for scene-aware streaming video QA that enables efficient and scalable reasoning over continuous video streams. The innovation of Vista can be summarized in three aspects: (1) Scene-aware segmentation. Vista dynamically clusters incoming frames into temporally and visually coherent scene units. (2) Scene-aware compression. Each scene is compressed into a compact token representation and stored in GPU memory for efficient index-based retrieval, while the full-resolution frames are offloaded to CPU memory. (3) Scene-aware recall. Upon receiving a question, relevant scenes are selectively recalled and reintegrated into the model’s input space, enabling both efficiency and completeness. Vista is model-agnostic and integrates seamlessly with a variety of vision-language backbones, enabling long-context reasoning without compromising latency or memory efficiency. Extensive experiments on StreamingBench demonstrate that Vista achieves state-of-the-art performance, establishing a strong baseline for real-world streaming video understanding. Haocheng Lu, Xiaoyang Qu, Guokuan Li, Jiguang Wan 0001, Jianzong Wang |
AAAI | 5 |
| 2025 | RUNA: Object-Level Out-of-Distribution Detection via Regional Uncertainty Alignment of Multimodal RepresentationsabstractEnabling object detectors to recognize out-of-distribution (OOD) objects is vital for building reliable systems. A primary obstacle stems from the fact that models frequently do not receive supervisory signals from unfamiliar data, leading to overly confident predictions regarding OOD objects. Despite previous progress that estimates OOD uncertainty based on the detection model and in-distribution (ID) samples, we explore using pre-trained vision-language representations for object-level OOD detection. We first discuss the limitations of applying image-level CLIP-based OOD detection methods to object-level scenarios. Building upon these insights, we propose RUNA, a novel framework that leverages a dual encoder architecture to capture rich contextual information and employs a regional uncertainty alignment mechanism to distinguish ID from OOD objects effectively. We introduce a few-shot fine-tuning approach that aligns region-level semantic representations to further improve the model's capability to discriminate between similar objects. Our experiments show that RUNA substantially surpasses state-of-the-art methods in object-level OOD detection, particularly in challenging scenarios with diverse and complex object instances. Jinggang Chen, Xiaoyang Qu, Guokuan Li, Kai Lu 0002, Jiguang Wan 0001, Jing Xiao 0006, Jianzong Wang |
AAAI | 4 |
| 2025 | VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD DetectionabstractAs object detectors are increasingly deployed as black-box cloud services or pre-trained models with restricted access to the original training data, the challenge of zero-shot object-level out-of-distribution (OOD) detection arises. This task becomes crucial in ensuring the reliability of detectors in open-world settings. While existing methods have demonstrated success in image-level OOD detection using pre-trained vision-language models like CLIP, directly applying such models to object-level OOD detection presents challenges due to the loss of contextual information and reliance on image-level alignment. To tackle these challenges, we introduce a new method that leverages visual prompts and text-augmented in-distribution (ID) space construction to adapt CLIP for zero-shot object-level OOD detection. Our method preserves critical contextual information and improves the ability to differentiate between ID and OOD objects, achieving competitive performance across different benchmarks. Xiaoyang Qu, Guokuan Li, Jiguang Wan 0001, Jianzong Wang |
ICASSP | 3 |
| 2025 | MADLLM: Multivariate Anomaly Detection via Pre-trained LLMsabstractWhen applying pre-trained large language models (LLMs) to address anomaly detection tasks, the multivariate time series (MTS) modality of anomaly detection does not align with the text modality of LLMs. Existing methods simply transform the MTS data into multiple univariate time series sequences, which can cause many problems. This paper introduces MADLLM, a novel multivariate anomaly detection method via pre-trained LLMs. We design a new triple encoding technique to align the MTS modality with the text modality of LLMs. Specifically, this technique integrates the traditional patch embedding method with two novel embedding approaches: (i) Skip Embedding, which alters the order of patch processing in traditional methods to help LLMs retain knowledge of previous features, and (ii) Feature Embedding, which leverages contrastive learning to allow the model to better understand the correlations between different features. Experimental results demonstrate that our method outperforms state-of-the-art methods in various public anomaly detection datasets. Xiaoyang Qu, Kai Lu 0002, Jiguang Wan 0001, Guokuan Li, Jianzong Wang |
ICME | 5 |
| 2025 | NStore: A High-Performance NUMA-Aware Key-Value Store for Hybrid MemoryabstractEmerging persistent memory (PM) promises near-DRAM performance, larger capacity, and data persistence, attracting researchers to design PM-based key-value stores. However, existing PM-based key-value stores lack awareness of the Non-Uniform Memory Access (NUMA) architecture on PM, where accessing PM on remote NUMA sockets is considerably slower than accessing local PM. This NUMA-unawareness results in sub-optimal performance when scaling on NUMA. Although DRAM caching alleviates this issue, existing cache policies ignore the performance disparity between remote and local PM accesses, keeping remote PM access as a performance bottleneck when scaling PM stores on NUMA. Furthermore, creating hot data views in each socket's PM fails to eliminate remote PM writes and, worse, induces additional local PM writes. This paper presents NStore, a high-performance NUMA-aware key-value store for the PM-DRAM hybrid memory. NStore introduces a NUMA-aware cache replacement strategy, called Remote Access First (RAF) cache in DRAM, to minimize remote PM accesses. In addition, NStore deploys Nlog, a write-optimized log-structured persistent storage, purposed to eliminate remote PM writes. NStore further mitigates the NUMA impacts through localized scan operations, efficient garbage collection, and multi-thread recovery for Nlog. Evaluations show that NStore outperforms state-of-the-art PM-based key-value stores, achieving up to 13.9$\times$and 11.2$\times$higher write and read throughput, respectively. Zhonghua Wang 0001, Kai Lu 0002, Jiguang Wan 0001, Hong Jiang 0001, Zeyang Zhao, Biliang Lai, Guokuan Li, Changsheng Xie 0001 |
IEEE Trans. Computers | 8 |
| 2024 | Value-Driven Mixed-Precision Quantization for Patch-Based Inference on MicrocontrollersabstractDeploying neural networks on microcontroller units (MCUs) presents substantial challenges due to their constrained computation and memory resources. Previous researches have explored patch-based inference as a strategy to conserve memory without sacrificing model accuracy. However, this technique suffers from severe redundant computation overhead, leading to a substantial increase in execution latency. A feasible solution to address this issue is mixed-precision quantization, but it faces the challenges of accuracy degradation and a time-consuming search time. In this paper, we propose QuantMCU, a novel patch-based inference method that utilizes value-driven mixed-precision quantization to reduce redundant computation. We first utilize value-driven patch classification (VDPC) to maintain the model accuracy. VDPC classifies patches into two classes based on whether they contain outlier values. For patches containing outlier values, we apply 8-bit quantization to the feature maps on the dataflow branches that follow. In addition, for patches without outlier values, we utilize value-driven quantization search (VDQS) on the feature maps of their following dataflow branches to reduce search time. Specifically, VDQS introduces a novel quantization search metric that takes into account both computation and accuracy, and it employs entropy as an accuracy representation to avoid additional training. VDQS also adopts an iterative approach to determine the bitwidth of each feature map to further accelerate the search process. Experimental results on real-world MCU devices show that QuantMCU can reduce computation by 2.2x on average while maintaining comparable model accuracy compared to the state-of-the-art patch-based inference methods. Shenglin He, Kai Lu 0002, Xiaoyang Qu, Guokuan Li, Jiguang Wan 0001, Jianzong Wang, Jing Xiao 0006 |
DATE | 5 |
| 2024 | Enhancing Anomalous Sound Detection with Multi-Level Memory BankabstractAbnormal sound detection (ASD) is crucial for the timely detection of machine faults in industrial scenarios and has emerged as a popular topic. However, exhaustively collecting all ever-changing anomalous samples is impractical for the associated time and cost. Under unsupervised conditions, identifying rare or even unseen abnormal sounds from a large set of normal samples is a notable challenge in the real-world setting. To address this, we propose a novel ASD method based on a multi-level memory bank to estimate the distribution of normal samples in the latent space. We employ a distance-based metric to distinguish inliers from outliers, leveraging high, mid, and low-level features to improve accuracy. We also propose an acoustic-aware farthest embedding sampling algorithm for inference acceleration and memory bank reduction. Experimental results demonstrate our method outperforms existing methods for anomaly detection. Additionally, we analyze the effect of multilevel and acoustic-aware farthest embedding sampling methods, respectively. Baoping Deng, Jinggang Chen, Zhenhou Hong, Xiaoyang Qu, Guokuan Li, Jiguang Wan 0001, Jianzong Wang |
IJCNN | 5 |
| 2024 | PRENet: A Plane-Fit Redundancy Encoding Point Cloud Sequence Network for Real-Time 3D Action RecognitionabstractRecognizing human actions from point cloud sequence has attracted tremendous attention from both academia and industry due to its wide applications. However, most previous studies on point cloud action recognition typically require complex networks to extract intra-frame spatial features and inter-frame temporal features, resulting in an excessive number of redundant computations. This leads to high latency, rendering them impractical for real-world applications. To address this problem, we propose a Plane-Fit Redundancy Encoding point cloud sequence network named PRENet. The primary concept of our approach involves the utilization of plane fitting to mitigate spatial redundancy within the sequence, concurrently encoding the temporal redundancy of the entire sequence to minimize redundant computations. Specifically, our network comprises two principal modules: a Plane-Fit Embedding module and a Spatio-Temporal Consistency Encoding module. The Plane-Fit Embedding module capitalizes on the observation that successive point cloud frames exhibit unique geometric features in physical space, allowing for the reuse of spatially encoded data for temporal stream encoding. The Spatio-Temporal Consistency Encoding module amalgamates the temporal structure of the temporally redundant part with its corresponding spatial arrangement, thereby enhancing recognition accuracy. We have done numerous experiments to verify the effectiveness of our network. The experimental results demonstrate that our method achieves almost identical recognition accuracy while being nearly four times faster than other state-of-the-art methods. Shenglin He, Xiaoyang Qu, Jiguang Wan 0001, Guokuan Li, Jianzong Wang |
IJCNN | 4 |
| 2024 | Gecko: Resource-Efficient and Accurate Queries in Real-Time Video Streams at the EdgeabstractSurveillance cameras are ubiquitous nowadays and users’ increasing needs for accessing real-world information (e.g., finding abandoned luggage) have urged object queries in real-time videos. While recent real-time video query processing systems exhibit excellent performance, they lack utility in deployment in practice as they overlook some crucial aspects, including multi-camera exploration, resource contention, and content awareness. Motivated by these issues, we propose a framework Gecko, to provide resource-efficient and accurate real-time object queries of massive videos on edge devices. Gecko (i) obtains optimal models from the model zoo and assigns them to edge devices for executing current queries, (ii) optimizes resource usage of the edge cluster at runtime by dynamically adjusting the frame query interval of each video stream and forking/joining running models on edge devices, and (iii) improves accuracy in changing video scenes by fine-grained stream transfer and continuous learning of models. Our evaluation with real-world video streams and queries shows that Gecko achieves up to 2x more resource efficiency gains and increases overall query accuracy by at least 12% compared with prior work, further delivering excellent scalability for practical deployment. Liang Wang 0057, Xiaoyang Qu, Jianzong Wang, Guokuan Li, Jiguang Wan 0001, Song Guo 0001, Jing Xiao 0006 |
INFOCOM | 4 |
| 2024 | Scythe: A Low-latency RDMA-enabled Distributed Transaction System for Disaggregated MemoryabstractDisaggregated memory separates compute and memory resources into independent pools connected by RDMA (Remote Direct Memory Access) networks, which can improve memory utilization, reduce cost, and enable elastic scaling of compute and memory resources. However, existing RDMA-based distributed transactions on disaggregated memory suffer from severe long-tail latency under high-contention workloads. In this article, we propose Scythe, a novel low-latency RDMA-enabled distributed transaction system for disaggregated memory. Scythe optimizes the latency of high-contention transactions in three approaches: (1) Scythe proposes a hot-aware concurrency control policy that uses optimistic concurrency control (OCC) to improve transaction processing efficiency in low-conflict scenarios. Under high conflicts, Scythe designs a timestamp-ordered OCC (TOCC) strategy based on fair locking to reduce the number of retries and cross-node communication overhead. (2) Scythe presents an RDMA-friendly timestamp service for improved timestamp management. And, (3) Scythe designs an RDMA-optimized RPC framework to improve RDMA bandwidth utilization. The evaluation results show that, compared with state-of-the-art distributed transaction systems, Scythe achieves more than 2.5× lower latency with 1.8× higher throughput under high-contention workloads. Kai Lu 0002, Siqi Zhao, Haikang Shan, Guokuan Li, Jiguang Wan 0001, Ting Yao 0001, Huatao Wu, Daohui Wang |
ACM Trans. Archit. Code Optim. | 5 |
| 2024 | WIPE: A Write-Optimized Learned Index for Persistent MemoryabstractLearned Index, which utilizes effective machine learning models to accelerate locating sorted data positions, has gained increasing attention in many big data scenarios. Using efficient learned models, the learned indexes build large nodes and flat structures, thereby greatly improving the performance. However, most of the state-of-the-art learned indexes are designed for DRAM, and there is hence an urgent need to enable high-performance learned indexes for emerging Non-Volatile Memory (NVM). In this article, we first evaluate and analyze the performance of the existing learned indexes on NVM. We discover that these learned indexes encounter severe write amplification and write performance degradation due to the requirements of maintaining large sorted/semi-sorted data nodes. To tackle the problems, we propose a novel three-tiered architecture of write-optimized persistent learned index, which is named WIPE , by adopting unsorted fine-granularity data nodes to achieve high write performance on NVM. Thereinto, we devise a new root node construction algorithm to accelerate searching numerous small data nodes. The algorithm ensures stable flat structure and high read performance in large-size datasets by introducing an intermediate layer (i.e., index nodes) and achieving accurate prediction of index node positions from the root node. Our extensive experiments on Intel DCPMM show that WIPE can improve write throughput and read throughput by up to 3.9× and 7×, respectively, compared to the state-of-the-art learned indexes. Also, WIPE can recover from a system crash in ∼ 18 ms. WIPE is free as an open-source software package. 1 Zhonghua Wang 0001, Chen Ding 0012, Fengguang Song, Kai Lu 0002, Jiguang Wan 0001, Zhihu Tan, Changsheng Xie 0001, Guokuan Li |
ACM Trans. Archit. Code Optim. | 8 |
| 2023 | DoW-KV: A DPU-offloaded and Write-optimized Key-Value Store on Disaggregated Persistent MemoryabstractDisaggregated Persistent Memory (DPM) is a promising technology offering elasticity, high resource utilization, persistent data storage, and lower power consumption. While building KV stores on the DPM benefits from these merits, achieving efficient writes also faces two primary challenges: 1) limited scalability caused by the underused PM bandwidth, and 2) limited CPU resources the persistent memory server (PMS) can provide. Integrating the SmartNIC such as the Data Processing Unit (DPU) into the DPM gives developers the chance to optimize writing to KV stores by utilizing both the memory and processor of DPU. However, simple offloading cannot make full use of the DPU’s potential capacity. To address these challenges, we propose DoW-KV, a persistent hash KV store on DPM. DoW-KV employs a two-tier hash index consisting of a DPU cache table in DPU memory and multiple PM persistent tables on the PM. It relocates small random writes to the DPU memory and consolidates them to the PM at a coarse granularity. Furthermore, DoW-KV uses DPU-offloaded step merge and a coroutine-based asynchronous processing framework to efficiently manage the PM persistent tables. DoW-KV also introduces a client-mixed read strategy to boost key searching on the two-tier hash index. Experimental results show that DoW-KV outperforms the state-of-the-art DINOMO by 2.1× and 1.3× in the Put and Get operations, respectively. Guokuan Li, Jiguang Wan 0001, Junyue Wang, Ting Yao 0001, Huatao Wu, Daohui Wang |
CLUSTER | 2 |
| 2023 | Shoggoth: Towards Efficient Edge-Cloud Collaborative Real-Time Video Inference via Adaptive Online LearningabstractThis paper proposes Shoggoth, an efficient edge-cloud collaborative architecture, for boosting inference performance on real-time video of changing scenes. Shoggoth uses online knowledge distillation to improve the accuracy of models suffering from data drift and offloads the labeling process to the cloud, alleviating constrained resources of edge devices. At the edge, we design adaptive training using small batches to adapt models under limited computing power, and adaptive sampling of training frames for robustness and reducing bandwidth. The evaluations on the realistic dataset show 15%–20% model accuracy improvement compared to the edge-only strategy and fewer network costs than the cloud-only strategy. Liang Wang 0057, Kai Lu 0002, Xiaoyang Qu, Jianzong Wang, Jiguang Wan 0001, Guokuan Li, Jing Xiao 0006 |
DAC | 7 |
| 2023 | EdgeMA: Model Adaptation System for Real-Time Video Analytics on Edge Devices
Liang Wang 0057, Xiaoyang Qu, Jianzong Wang, Jiguang Wan 0001, Guokuan Li, Kaiyu Hu, Guilin Jiang, Jing Xiao 0006 |
ICONIP (1) | 6 |
| 2022 | Accelerating range queries of primary and secondary indices for key-value separationabstractPrimary and secondary indices in LSM-tree-based key-value (KV) stores play significant roles for real-world applications, but they suffer severe I/O amplification due to compaction operations. Prior works show that KV separation can mitigate the I/O amplification under various workloads for either primary or secondary indices. However, range queries of primary and secondary indices only achieve suboptimal efficiency for two reasons: (1) KV separation improves insert/update performance by sacrificing the performance of range queries, (2) range queries of primary and secondary indices may conflict with each other. Chenlei Tang, Jiguang Wan 0001, Zhihu Tan, Guokuan Li |
SoCC | 4 |
| 2022 | RepKV: A Replicated Key-Value Store to Boost Multiple Indices for Key-Value SeparationabstractPrimary and secondary indices are demanded in real-world applications. Recent works show that key-value(KV) separation is efficient to improve queries of multiple secondary indices in LSM-based KV stores. It stores the value in a separate value log and only stores keys and the value address in primary and secondary indices. However, this share-value scheme on multiple indices results in suboptimal efficiency: (1) queries on secondary indices and range queries on all indices cannot fully exploit the bandwidth of SSD devices simultaneously (2) the put operation is inefficient to update secondary indices.To address the above inefficiency, we propose RepKV, a replicated KV store aiming to boost operations of multiple in-dices. Firstly, RepKV uses a primary-backup replication scheme. Each replication stores the same KV pairs but adopts different organizations to exploit the benefit of SSD devices. Secondly, RepKV proposes a lightweight replication scheme to mitigate the extra KV pairs synchronized in replications. Thirdly, RepKV uses a parallel parsing policy to boost the put operation. Experimental results show that RepKV can improve the query performance on secondary indices by up to 22.24%, the range query performance of all indices by up to 31.38%, and the put performance by up to 13.8%. Besides, RepKV can reduce the I/O amplification of replication by up to 3.05x via the lightweight replication scheme. Chenlei Tang, Jiguang Wan 0001, Zhihu Tan, Guokuan Li |
ICCD | 4 |
| 2022 | ADSTS: Automatic Distributed Storage Tuning System Using Deep Reinforcement LearningabstractModern distributed storage systems with the immense number of configurations, unpredictable workloads and difficult performance evaluation pose higher requirements to parameter tuning. Providing an automatic parameter tuning solution for distributed storage systems is in demand. Lots of researches have attempted to build automatic tuning systems based on deep reinforcement learning (RL). However, they have several limitations in the face of these requirements, including lack of parameter spaces processing, less advanced RL models and time-consuming and unstable training process. In this paper, we present and evaluate the ADSTS, which is an automatic distributed storage tuning system based on deep reinforcement learning. A general preprocessing guideline is first proposed to generate standardized tunable parameter domain. Thereinto, Recursive Stratified Sampling without the nonincremental nature is designed to sample huge parameter spaces and Lasso regression is adopted to identify important parameters. Besides, the twin-delayed deep deterministic policy gradient method is utilized to find the optimal values of tunable parameters. Finally, Multi-processing Training and Workload-directed Model Fine-tuning are adopted to accelerate the model convergence. ADSTS is implemented on Park and is used in the real-world system Ceph. The evaluation results show that ADSTS can recommend near-optimal configurations and improve system performance by 1.5 × ∼2.5 × with acceptable overheads. Kai Lu 0002, Guokuan Li, Jiguang Wan 0001, Ruixiang Ma |
ICPP | 2 |