Wenrui Lu

dblp:276/4832 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0001-6328-714XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning
abstract
Multimodal human action recognition (HAR) utilizes complementary data for activity classification. Built on traditional HAR tasks, recent advances in Large Language Models (LLMs) enable detailed descriptions and causal reasoning of human actions, advancing new tasks of human action understanding (HAU) and human action reasoning (HARn). However, most LLMs, especially multimodal Large Vision-Language Models (LVLMs), struggle with modalities other than RGB images, like depth, IMU, ormmWave, due to a lack of large-scale datasets in these task domains. Existing HAR datasets provide only coarse-grained annotations, in-sufficient for depicting the detailed action dynamics required in HAU and HARn tasks. Simply combining annotations and generating captions with LLMs often lacks necessary logical and spatiotemporal consistency. In this paper, we introduce CUHK-X, a large-scale multi-modal dataset and benchmarks for HAR, HAU, and HARn. It includes 64,267 samples of 40 actions performed by 30 participants across two indoor environments, covering diverse daily scenarios. To address the challenge of spatiotemporal inconsistencies in captions, we propose a prompt-based scene creation method that leverages LLMs to generate logically connected activity sequences. CUHK-X also includes three benchmarks with six tasks to evaluate state-of-the-art models. Experimental results show average accuracies of 76.52% for HAR, 40.76% for HAU, and 70.25% for HARn. This large-scale multimodal dataset aims to empower the research community to apply, develop, and adapt data-intensive learning techniques for a wide range of human activity-related tasks.
Siyang Jiang, Mu Yuan, Bufang Yang, Lilin Xu, Yang Li 0147, Yuting He 0006, Liran Dong, Wenrui Lu, Zhenyu Yan 0002, Xiaofan Jiang 0001, Wei Gao 0006, Hongkai Chen 0001, Guoliang Xing
MobiSys10
2026 Spiking Depth: Depth estimation from sparse events with spiking neural networks
Dongze Liu, Yimeng Fan, Wenrui Lu, Changsong Liu, Wei Zhang 0055
Expert Syst. Appl.3
2026 Time-Sensitive Multi-DNN Inference on CPU-GPU Edge Platforms
abstract
In recent years, Deep Neural Networks (DNNs) have been increasingly adopted in a wide range of time-critical applications running on edge platforms equipped with heterogeneous multiprocessors. Given the limited resources available on these platforms, efficiently utilizing both CPU and GPU resources for time-sensitive DNN inference is crucial. However, this cross-processor inference paradigm poses significant challenges due to inherent performance imbalances between different processors. In this paper, we introduce BlastNet, a system that leverages duo-blocks—a novel model inference abstraction designed to enable highly efficient cross-processor, time-sensitive DNN inference. Each duo-block features a dual model structure, facilitating fine-grained, alternate inference across different processors. Duo-blocks are optimized during design and dynamically scheduled at runtime to maximize the resource utilization of CPU and GPU. To address memory constraints on edge devices, we also propose a duo-block selection algorithm that selectively constructs duo-blocks based on performance gains. BlastNet is implemented on an indoor autonomous driving platform and three popular edge platforms. Extensive evaluations demonstrate that BlastNet reduces the deadline missing rate by$35.07\,\%$with only a mere$1.63 \%$loss in model accuracy.
Neiwen Ling, Wenrui Lu, Xuan Huang 0001, Nan Guan, Zhenyu Yan 0002, Guoliang Xing
IEEE Trans. Mob. Comput.2
2026 An Efficient Edge-Cloud Collaboration System With Foundational Models for Open-Set IoT Applications
abstract
Artificial intelligence (AI) models have been widely deployed on edge devices, enabling various IoT applications. However, lightweight on-device AI models on resource-limited edge devices hinder their adaptability to dynamic environments and tasks. Despite the superior generalization capabilities of recently developed Foundation Models (FMs), utilizing their extensive knowledge on the resource-constrained edge platforms remains unexplored. In this work, we introduce DeepEdgeFM, an edge-cloud collaborative system with FMs that enables open-set learning, simultaneously achieving generalizability and efficiency for IoT applications. DeepEdgeFM employs a spatiotemporalaware semantic customization approach that leverages spatial, temporal, and domain-specific knowledge from FMs to continuously customize edge models using unlabeled sensor data in emerging IoT environments. Meanwhile, DeepEdgeFM utilizes a dynamic model switching strategy to selectively query the knowledge of FMs based on sensor-data uncertainty and real-time network fluctuations. We implement DeepEdgeFM on five FMs and multi-modal large language models (MLLMs), covering four types of sensor data modalities. We evaluate DeepEdgeFM on two edge platforms, five public datasets, and two self-collected datasets covering both indoor and outdoor real-world environments. The results show that DeepEdgeFM outperforms state-ofthe- art baselines, achieving up to an 18.6% accuracy gain and a 38.6
Bufang Yang, Wenrui Lu, Lixing He, Neiwen Ling, Zhenyu Yan 0002, Guoliang Xing, Xian Shuai, Xiaozhe Ren, Xin Jiang 0002
IEEE Trans. Mob. Comput.2
2025 ContextAgent: Context-Aware Proactive LLM Agents with Open-world Sensory Perceptions
abstract
Recent advances in Large Language Models (LLMs) have propelled intelligent agents from reactive responses to proactive support. While promising, existing proactive agents either rely exclusively on observations from enclosed environments (e.g., desktop UIs) with direct LLM inference or employ rule-based proactive notifications, leading to suboptimal user intent understanding and limited functionality for proactive service. In this paper, we introduce ContextAgent, the first context-aware proactive agent that incorporates extensive sensory contexts surrounding humans to enhance the proactivity of LLM agents. ContextAgent first extracts multi-dimensional contexts from massive sensory perceptions on wearables (e.g., video and audio) to understand user intentions. ContextAgent then leverages the sensory contexts and personas from historical data to predict the necessity for proactive services. When proactive assistance is needed, ContextAgent further automatically calls the necessary tools to assist users unobtrusively. To evaluate this new task, we curate ContextAgentBench, the first benchmark for evaluating context-aware proactive LLM agents, covering 1,000 samples across nine daily scenarios and twenty tools. Experiments on ContextAgentBench show that ContextAgent outperforms baselines by achieving up to 8.5% and 6.0% higher accuracy in proactive predictions and tool calling, respectively. We hope our research can inspire the development of more advanced, human-centric, proactive AI assistants. The code and dataset are publicly available at https://github.com/openaiotlab/ContextAgent.
Bufang Yang, Lilin Xu, Liekang Zeng, Kaiwei Liu 0001, Siyang Jiang, Wenrui Lu, Hongkai Chen 0001, Xiaofan Jiang 0001, Guoliang Xing, Zhenyu Yan 0002
NeurIPS6
2024 SFOD: Spiking Fusion Object Detector
abstract
Event cameras, characterized by high temporal resolution, high dynamic range, low power consumption, and high pixel bandwidth, offer unique capabilities for object detection in specialized contexts. Despite these advantages, the inherent sparsity and asynchrony of event data pose challenges to existing object detection algorithms. Spiking Neural Networks (SNNs), inspired by the way the human brain codes and processes information, offer a potential solution to these difficulties. However, their performance in object detection using event cameras is limited in current imple-mentations. In this paper, we propose the Spiking Fusion Object Detector (SFOD), a simple and efficient approach to SNN-based object detection. Specifically, we design a Spiking Fusion Module, achieving the first-time fusion of feature maps from different scales in SNNs applied to event cameras. Additionally, through integrating our analysis and experiments conducted during the pretraining of the back-bone network on the NCAR dataset, we delve deeply into the impact of spiking decoding strategies and loss functions on model performance. Thereby, we establish state-of-the-art classification results based on SNNs, achieving 93.7% accuracy on the NCAR dataset. Experimental results on the GEN1 detection dataset demonstrate that the SFOD achieves a state-of-the-art mAP of 32.1%, outperforming existing SNN-based approaches. Our research not only underscores the potential of SNNs in object detection with event cameras but also propels the advancement of SNNs. Code is available at https://github.com/yimeng-fan/SFOD.
Yimeng Fan, Changsong Liu, Wenrui Lu
CVPR5
2023 Mozart: A Mobile ToF System for Sensing in the Dark through Phase Manipulation
abstract
Sensing in low-light and dark environments has a wide range of applications. However, existing sensing technologies suffer several major challenges, such as excessive noise and low resolution. This paper proposes Mozart - a new mobile sensing system that leverages off-the-shelf Time-of-Flight (ToF) depth cameras to generate high-resolution and rich-in-texture maps for applications in dark scenarios. The design of Mozart is based on our key observation that the phase components of ToF measurements can be manipulated to expose texture information. Through in-depth analysis of the physical reflection model, we show that the textures can be exposed and enhanced using highly compute-efficient phase manipulation functions. By exploiting the physics texture models, we propose an autoencoder-based unsupervised learning approach that can automatically learn efficient representations from phase components to generate high-resolution maps. We implemented Mozart on several Android smartphone models1, and an edge testbed with standalone ToF camera platforms for various applications in the dark. The results show that Mozart can work in real time and delivers significant improvement over existing sensing technologies. Therefore, Mozart offers a low-cost, high-performance sensing technology for next-generation applications in the dark.
Xiaomin Ouyang, Li Pan 0004, Wenrui Lu, Guoliang Xing, Xiaoming Liu 0002
MobiSys4
2023 Construction of an aspect-level sentiment analysis model for online medical reviews
Yuehua Zhao, Linyi Zhang, Chenxi Zeng, Wenrui Lu, Yidan Chen
Inf. Process. Manag.4
2022 HiToF: a ToF camera system for capturing high-resolution textures
abstract
We present a demonstration of an enhanced Time-of-Flight (ToF) depth system named HiToF, which can expose high-resolution textures from captured depth maps. By design, a ToF camera can easily capture the depth maps of a scene while largely omitting the corresponding texture information, which is often critical for the performance of many depth applications. HiToF is developed to address this issue by generating enhanced depth maps with high-resolution textures. The key idea is to manipulate the phase components used in the measurement of time-of-flight for the received IR light. In this demo, we showcase our implementation using off-the-shelf ToF cameras and engage audience with an interactive experience in various scenarios, which illustrates the system's effectiveness in improving the performance of ToF cameras in depth applications.
Xiaomin Ouyang, Li Pan 0004, Wenrui Lu, Xiaoming Liu 0002, Guoliang Xing
MobiCom4