VLDB 2026 Research / reviewers in the wild / expert
Neiwen Ling
dblp:228/5904
· DBLP profile ↗
18ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0003-2072-1502ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 17 · 6 first-author · 16 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TimelyLLM: Time-sensitive LLM Serving System for Physical-I/O Limited AgentsabstractLarge Language Models (LLMs) are increasingly integrated into Physical-I/O limited agents, such as robots and voice assistants, which execute outputs sequentially. However, existing LLM serving systems typically employ a throughput-oriented batching mechanism, ignoring the large gap between LLM generation speed and the constrained physical I/O rates of agents, thus wasting execution slack and worsening resource contention. Besides, they treat all tokens equally and cannot anticipate the execution implications of different content, preventing scheduling aligned with agent-side behavior. To address it, we propose a new system named TimelyLLM that coordinates LLM generation with the physical behavior of agents. TimelyLLM introduces a novel segmented generation and scheduling mechanism, strategically leveraging the time gap between agent plan generation and execution to reduce contention and improve response latency under multi-agent workloads. We implement TimelyLLM on top of a widely-used LLM serving framework. We also build a dataset collection system to construct serving workloads from real-world robots, including drones, robot arms, and quadruped robots. Our evaluation demonstrates that TimelyLLM improves the time utility up to 1.52×, and reduces the overall waiting time by 84%. Neiwen Ling, Anurag Khandelwal, Lin Zhong 0001 |
MobiSys | 1 |
| 2026 | Time-Sensitive Multi-DNN Inference on CPU-GPU Edge PlatformsabstractIn recent years, Deep Neural Networks (DNNs) have been increasingly adopted in a wide range of time-critical applications running on edge platforms equipped with heterogeneous multiprocessors. Given the limited resources available on these platforms, efficiently utilizing both CPU and GPU resources for time-sensitive DNN inference is crucial. However, this cross-processor inference paradigm poses significant challenges due to inherent performance imbalances between different processors. In this paper, we introduce BlastNet, a system that leverages duo-blocks—a novel model inference abstraction designed to enable highly efficient cross-processor, time-sensitive DNN inference. Each duo-block features a dual model structure, facilitating fine-grained, alternate inference across different processors. Duo-blocks are optimized during design and dynamically scheduled at runtime to maximize the resource utilization of CPU and GPU. To address memory constraints on edge devices, we also propose a duo-block selection algorithm that selectively constructs duo-blocks based on performance gains. BlastNet is implemented on an indoor autonomous driving platform and three popular edge platforms. Extensive evaluations demonstrate that BlastNet reduces the deadline missing rate by$35.07\,\%$with only a mere$1.63 \%$loss in model accuracy. Neiwen Ling, Wenrui Lu, Xuan Huang 0001, Nan Guan, Zhenyu Yan 0002, Guoliang Xing |
IEEE Trans. Mob. Comput. | 1 |
| 2026 | An Efficient Edge-Cloud Collaboration System With Foundational Models for Open-Set IoT ApplicationsabstractArtificial intelligence (AI) models have been widely deployed on edge devices, enabling various IoT applications. However, lightweight on-device AI models on resource-limited edge devices hinder their adaptability to dynamic environments and tasks. Despite the superior generalization capabilities of recently developed Foundation Models (FMs), utilizing their extensive knowledge on the resource-constrained edge platforms remains unexplored. In this work, we introduce DeepEdgeFM, an edge-cloud collaborative system with FMs that enables open-set learning, simultaneously achieving generalizability and efficiency for IoT applications. DeepEdgeFM employs a spatiotemporalaware semantic customization approach that leverages spatial, temporal, and domain-specific knowledge from FMs to continuously customize edge models using unlabeled sensor data in emerging IoT environments. Meanwhile, DeepEdgeFM utilizes a dynamic model switching strategy to selectively query the knowledge of FMs based on sensor-data uncertainty and real-time network fluctuations. We implement DeepEdgeFM on five FMs and multi-modal large language models (MLLMs), covering four types of sensor data modalities. We evaluate DeepEdgeFM on two edge platforms, five public datasets, and two self-collected datasets covering both indoor and outdoor real-world environments. The results show that DeepEdgeFM outperforms state-ofthe- art baselines, achieving up to an 18.6% accuracy gain and a 38.6 Bufang Yang, Wenrui Lu, Lixing He, Neiwen Ling, Zhenyu Yan 0002, Guoliang Xing, Xian Shuai, Xiaozhe Ren, Xin Jiang 0002 |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Time-sensitive AI System for Physical AgentsabstractThe integration of Artificial Intelligence (AI) into physical agents, such as drones, robots, and autonomous vehicles, demands not only intelligence but also timely perception and decision-making. This work presents a time-sensitive AI system architecture that addresses the critical timing challenges inherent in such intelligent physical agents. We propose system-level designs that enable deadline-aware Deep Neural Network (DNN) execution on edge/embedded platforms, along with time-sensitive Large Language Model (LLM) serving on edge servers, thereby providing end-to-end support for time-critical physical intelligence. Neiwen Ling |
MobiSys | 1 |
| 2025 | TypeFly: Low-Latency Drone Planning With Large Language ModelsabstractRecent advancements in robot planning using large language models (LLMs) have demonstrated significant potential, primarily due to LLMs' capabilities to understand natural language commands and generate executable plans in various languages. However, in time-sensitive and interactive applications involving mobile robots, particularly drones, the sequential token generation process inherent to LLMs introduces substantial latency, i.e., response time, during the control plan generation. In this paper, we present a system called ChatFly that tackles this latency problem using a combination of a novel programming language called MiniSpec and its runtime to reduce both the response time and generation time for the robot plan. That is, instead of asking an LLM to write a program (robotic plan) in the popular but verbose Python, ChatFly gets it to do it in MiniSpec specially designed for token efficiency and stream interpreting. Using a set of challenging drone tasks, we show that design choices made by ChatFly can reduce the average response time to 74% compared to existing works and provide a more consistent user experience, enabling responsive and intelligent LLM-based drone control. Xiaojing Yu, Neiwen Ling, Lin Zhong 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | Soar: Design and Deployment of A Smart Roadside Infrastructure System for Autonomous DrivingabstractRecently, smart roadside infrastructure (SRI) has demonstrated the potential of achieving fully autonomous driving systems. To explore the potential of infrastructure-assisted autonomous driving, this paper presents the design and deployment of Soar, the first end-to-end SRI system specifically designed to support autonomous driving systems. Soar consists of both software and hardware components carefully designed to overcome various system and physical challenges. Soar can leverage the existing operational infrastructure like street lampposts for a lower barrier of adoption. Soar adopts a new communication architecture that comprises a bi-directional multi-hop I2I network and a downlink I2V broadcast service, which are designed based on off-the-shelf 802.11ac interfaces in an integrated manner. Soar also features a hierarchical DL task management framework to achieve desirable load balancing among nodes and enable them to collaborate efficiently to run multiple data-intensive autonomous driving applications. We deployed a total of 18 Soar nodes on existing lampposts on campus, which have been operational for over two years. Our real-world evaluation shows that Soar can support a diverse set of autonomous driving applications and achieve desirable real-time performance and high communication reliability. Our findings and experiences in this work offer key insights into the development and deployment of next-generation smart roadside infrastructure and autonomous driving systems. Shuyao Shi, Neiwen Ling, Zhehao Jiang, Xuan Huang 0001, Xiaoguang Zhao, Bufang Yang, Chen Bian, Jingfei Xia, Zhenyu Yan 0002, Raymond W. Yeung, Guoliang Xing |
MobiCom | 2 |
| 2024 | Timely Fusion of Surround Radar/Lidar for Object Detection in Autonomous Driving SystemsabstractFusing Radar and Lidar sensor data can fully utilize their complementary advantages and provide more accurate reconstruction of the surrounding for autonomous driving systems. Surround Radar/Lidar can provide 360° view sampling with the minimal cost, which are promising sensing hardware solutions for autonomous driving systems. However, due to the intrinsic physical constraints, the rotating speed of surround Radar, and thus the frequency to generate Radar data frames, is much lower than surround Lidar. Existing Radar/Lidar fusion methods have to work at the low frequency of surround Radar, which cannot meet the high responsiveness requirement of autonomous driving systems. This paper develops techniques to fuse surround Radar/Lidar with working frequency only limited by the faster surround Lidar instead of the slower surround Radar, based on widely-used object detection model called MVDNet. The basic idea of our approach is simple: we let MVDNet work with temporally unaligned data from Radar/Lidar, so that fusion can take place at any time when a new Lidar data frame arrives, instead of waiting for the slow Radar data frame. However, directly applying MVDNet to temporally unaligned Radar/Lidar data greatly degrades its object detection accuracy. The key information revealed in this paper is that we can achieve high output frequency with little accuracy loss by enhancing the training procedure to explore the temporal redundancy in MVDNet so that it can tolerate the temporal unalignment of input data. We explore several different ways of training enhancement and compare them quantitatively with experiments. Tao Hu 0018, Neiwen Ling, Guoliang Xing, Chun Jason Xue, Nan Guan |
RTCSA | 3 |
| 2024 | Poster Abstract: Tasking Heterogeneous Sensor Systems with LLMsabstractDespite the extensive use of sensors enabling intelligent applications, the complementary potential of co-existing sensor systems is often not fully utilized, limiting more advanced applications. This paper introduces a novel solution using Large Language Models (LLMs) to coordinate sensor systems for handling complex user queries. It defines a sensor language for sensor systems, including vocabulary set and grammar rules, analogous to natural language components, enabling LLMs to translate user intentions into sensor coordination plans. Preliminary results show that our approach significantly outperforms the existing solution at plan generation, execution and response generation stages. Kaiwei Liu 0001, Bufang Yang, Lilin Xu, Yunqi Guo, Neiwen Ling, Guoliang Xing, Xian Shuai, Xiaozhe Ren, Xin Jiang 0002, Zhenyu Yan 0002 |
SenSys | 5 |
| 2023 | CoEdge: A Cooperative Edge System for Distributed Real-Time Deep Learning TasksabstractRecent years have witnessed the emergence of a new class of cooperative edge systems in which a large number of edge nodes can collaborate through local peer-to-peer connectivity. In this paper, we propose CoEdge, a novel cooperative edge system that can support concurrent data/compute-intensive deep learning (DL) models for distributed real-time applications such as city-scale traffic monitoring and autonomous driving. First, CoEdge includes a hierarchical DL task scheduling framework that dispatches DL tasks to edge nodes based on their computational profiles, communication overhead, and real-time requirements. Second, CoEdge can dramatically increase the execution efficiency of DL models by batching sensor data and aggregating the inferences of the same model. Finally, we propose a new edge containerization approach that enables an edge node to execute concurrent DL tasks by partitioning the CPU and GPU workloads into different containers. We extensively evaluate CoEdge on a self-deployed smart lamppost testbed on a university campus. Our results show that CoEdge can achieve up to reduction on deadline missing rate compared to baselines. Zhehao Jiang, Neiwen Ling, Xuan Huang 0001, Shuyao Shi, Chenhao Wu 0006, Xiaoguang Zhao, Zhenyu Yan 0002, Guoliang Xing |
IPSN | 2 |
| 2023 | Harmony: Heterogeneous Multi-Modal Federated Learning through Disentangled Model TrainingabstractMulti-modal sensing systems are increasingly prevalent in real-world applications such as health monitoring and autonomous driving. Most multi-modal learning approaches need to access users' raw data, which poses significant concerns to users' privacy. Federated learning (FL) provides a privacy-aware distributed learning framework. However, current FL approaches have not addressed the unique challenges of heterogeneous multi-modal FL systems, such as modality heterogeneity and significantly longer training delay. In this paper, we propose Harmony, a new system for heterogeneous multi-modal federated learning. Harmony disentangles the multi-modal network training in a novel two-stage framework, namely modality-wise federated learning and federated fusion learning. By integrating a novel balance-aware resource allocation mechanism in modality-wise FL and exploiting modality biases in federated fusion learning, Harmony improves the model accuracy under non-i.i.d. data distributions and speeds up system convergence. We implemented Harmony on a real-world multi-modal sensor testbed deployed in the homes of 16 elderly subjects for Alzheimer's Disease monitoring. Our evaluation on the testbed and three large-scale public datasets of different applications show that, Harmony outperforms by up to 46.35% accuracy over state-of-the-art baselines and saves up to 30% training delay. Xiaomin Ouyang, Heming Fu, Sitong Cheng, Li Pan 0004, Neiwen Ling, Guoliang Xing, Jianwei Huang 0001 |
MobiSys | 6 |
| 2023 | EdgeFM: Leveraging Foundation Model for Open-set Learning on the EdgeabstractDeep Learning (DL) models have been widely deployed on IoT devices with the help of advancements in DL algorithms and chips. However, the limited resources of edge devices make these on-device DL models hard to be generalizable to diverse environments and tasks. Although the recently emerged foundation models (FMs) show impressive generalization power, how to effectively leverage the rich knowledge of FMs on resource-limited edge devices is still not explored. In this paper, we propose EdgeFM, a novel edge-cloud cooperative system with open-set recognition capability. EdgeFM selectively uploads unlabeled data to query the FM on the cloud and customizes the specific knowledge and architectures for edge models. Meanwhile, EdgeFM conducts dynamic model switching at run-time taking into account both data uncertainty and dynamic network variations, which ensures the accuracy always close to the original FM. We implement EdgeFM using two FMs on two edge platforms. We evaluate EdgeFM on three public datasets and two self-collected datasets. Results show that EdgeFM can reduce the end-to-end latency up to 3.2x and achieve 34.3% accuracy increase compared with the baseline. Bufang Yang, Lixing He, Neiwen Ling, Zhenyu Yan 0002, Guoliang Xing, Xian Shuai, Xiaozhe Ren, Xin Jiang 0002 |
SenSys | 3 |
| 2023 | Miriam: Exploiting Elastic Kernels for Real-time Multi-DNN Inference on Edge GPUabstractMany applications such as autonomous driving and augmented reality, require the concurrent running of multiple deep neural networks (DNN) that poses different levels of real-time performance requirements. However, coordinating multiple DNN tasks with varying levels of criticality on edge GPUs remains an area of limited study. Unlike server-level GPUs, edge GPUs are resource-limited and lack hardware-level resource management mechanisms for avoiding resource contention. Therefore, we propose Miriam, a contention-aware task coordination framework for multi-DNN inference on edge GPU. Miriam consolidates two main components, an elastic-kernel generator, and a runtime dynamic kernel coordinator, to support mixed critical DNN inference. To evaluate Miriam, we build a new DNN inference benchmark based on CUDA with diverse representative DNN workloads. Experiments on two edge GPU platforms show that Miriam can increase system throughput by 92% while only incurring less than 10% latency overhead for critical tasks, compared to state of art baselines. Neiwen Ling, Nan Guan, Guoliang Xing |
SenSys | 2 |
| 2023 | Poster Abstract: Unifying On-device Tensor Program Optimization through Large Foundation ModelabstractWe present TensorBind, a novel approach aimed at unifying different hardware architectures for compilation optimization. Our proposed framework establishes an embedding space to seamlessly bind diverse hardware platforms together. By leveraging this unified representation, TensorBind enables efficient tensor program optimization techniques across a wide range of hardware platforms. We provide experimental results demonstrating the essentiality and adaptability of TensorBind in translating tensor program optimization records across multiple hardware architectures, thus revolutionizing compilation optimization strategies and facilitating the development of high-performance compilation systems over heterogeneous devices. Neiwen Ling, Kaiwei Liu 0001, Nan Guan, Guoliang Xing |
SenSys | 2 |
| 2022 | An Indoor Smart Traffic Dataset and Data Collection System: DatasetabstractSmart traffic is an emerging research area gaining more attention due to a class of emerging applications such as autonomous driving. Most smart traffic scenarios are outdoors, which are hard to collect traffic data and build demanding sensing systems. In this work, an indoor smart traffic testbed with an F1TENTH autonomous driving vehicle is built, allowing the collection of traffic datasets under different scenarios and performing various smart traffic tasks. This novel data collection system and collected dataset can help research teams build various smart traffic systems and evaluate indoor smart traffic datasets. The collected traffic light dataset is publicly available at the link1. Neiwen Ling, Nan Guan, Heming Fu, Guoliang Xing |
SenSys | 1 |
| 2022 | BlastNet: Exploiting Duo-Blocks for Cross-Processor Real-Time DNN InferenceabstractIn recent years, Deep Neural Network (DNN) has been increasingly adopted by a wide range of time-critical applications running on edge platforms with heterogeneous multiprocessors. To meet the stringent timing requirements of these applications, heterogeneous CPU and GPU resources must be efficiently utilized for the inference of multiple DNN models. Such a cross-processor real-time DNN inference paradigm poses major challenges due to the inherent performance imbalance among different processors and the lack of real-time support for cross-processor inference from existing deep learning frameworks. In this work, we propose a new system named BlastNet that exploits duo-block - a new model inference abstraction to support highly efficient cross-processor real-time DNN inference. Each duo-block has a dual model structure, enabling efficient fine-grained inference alternatively across different processors. BlastNet employs a novel block-level Neural Architecture Search (NAS) technique to generate duo-blocks, which accounts for computing characteristics and communication overhead. The duo-blocks are optimized at design time and then dynamically scheduled to achieve high resource utilization of heterogeneous CPU and GPU at runtime. BlastNet is implemented on an indoor autonomous driving platform and three popular edge platforms. Extensive results show that BlastNet achieves 35.07 % less deadline missing rate with a mere 1.63% of model accuracy loss. Neiwen Ling, Xuan Huang 0001, Nan Guan, Zhenyu Yan 0002, Guoliang Xing |
SenSys | 1 |
| 2022 | Aaron: Compile-Time Kernel Adaptation for Multi-DNN Inference Acceleration on Edge GPUabstractAI applications powered by deep learning are increasingly running on edge devices. Meanwhile, many real-world IoT applications demand multiple real-time tasks to run on the same device, for example, to achieve both object tracking and image segmentation simultaneously on an augmented reality glass. However, the current solutions can not yet support such multi-tenant real-time DNN inference on edge devices. Techniques such as on-device model compression trade inference accuracy for speed, while traditional DNN compilers mainly focus on single-tenant DNN model optimization. To fill this gap, we propose Aaron, which leverages DNN compiling techniques to accelerate multi-DNN inference on edge GPU based on compile-time kernel adaptation with no accuracy loss. Aaron integrates both DNN graph and kernel optimization to maximize on-device parallelism and minimize contention brought by concurrent inference. Neiwen Ling, Nan Guan, Guoliang Xing |
SenSys | 2 |
| 2021 | RT-mDL: Supporting Real-Time Mixed Deep Learning Tasks on Edge PlatformsabstractRecent years have witnessed an emerging class of real-time applications, e.g., autonomous driving, in which resource-constrained edge platforms need to execute a set of real-time mixed Deep Learning (DL) tasks concurrently. Such an application paradigm poses major challenges due to the huge compute workload of deep neural network models, diverse performance requirements of different tasks, and the lack of real-time support from existing DL frameworks. In this paper, we present RT-mDL, a novel framework to support mixed real-time DL tasks on edge platform with heterogeneous CPU and GPU resource. RT-mDL aims to optimize the mixed DL task execution to meet their diverse real-time/accuracy requirements by exploiting unique compute characteristics of DL tasks. RT-mDL employs a novel storage-bounded model scaling method to generate a series of model variants, and systematically optimizes the DL task execution by joint model variants selection and task priority assignment. To improve the CPU/GPU utilization of mixed DL tasks, RT-mDL also includes a new priority-based scheduler which employs a GPU packing mechanism and executes the CPU/GPU tasks independently. Our implementation on an F1/10 autonomous driving testbed shows that, RT-mDL can enable multiple concurrent DL tasks to achieve satisfactory real-time performance in traffic light detection and sign recognition. Moreover, compared to state-of-the-art baselines, RT-mDL can reduce deadline missing rate by 40.12% while only sacrificing 1.7% model accuracy. Neiwen Ling, Kai Wang 0018, Guoliang Xing, Daqi Xie |
SenSys | 1 |
| 2018 | ECRT: An Edge Computing System for Real-Time Image-based Object TrackingabstractReal-time image-based object tracking from live video is of great importance for several smart city applications like surveillance, intelligent traffic management and autonomous driving. Although recent deep learning systems can achieve satisfactory tracking performance, they incur significant compute overhead, which prevents them from wide adoption on resource-constrained IoT platforms. In this demonstration, we present an Edge Computing system for Real-time object Tracking (ECRT) for resource-constrained devices. The key feature of our system is that it intelligently partitions compute-intensive tasks such as inferencing a convolutional neural network(CNN) into two parts, which are executed locally on an IoT device and/or on the edge server. Moreover, ECRT can minimize the power consumption of IoT devices while taking into consideration the dynamic network environment and user requirement on end to end delay. Zhehao Jiang, Neiwen Ling, Xian Shuai, Guoliang Xing |
SenSys | 3 |