Borui Li 0001

dblp:135/9401-1 · DBLP profile ↗
← Back
31ranked-venue papers
12as first author
27since 2021 · last 2026
0000-0001-5262-2483ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 17 · 6 first-author · 15 since 2021Systems, architecture and hardware · 7 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DigimonGPT: An Evolvable Agent with Hierarchical Human-like Memory for Video Question Answering
abstract
Video question answering (VideoQA), whose goal is to produce answers through the integration of linguistic and visual understanding, has emerged as a significant research focus. Although Large Multimodal Models (LMMs) and autonomous agent methods have achieved notable advances in VideoQA, excessive computational overhead and restricted multimodal interaction capabilities limit their ability to facilitate the continuous evolution of the VideoQA system. To address the challenge, we introduce DigimonGPT, an evolvable VideoQA agent inspired by cognitive psychology. Specifically, DigimonGPT integrates a multimodal memory mechanism to achieve the continuous evolution of VideoQA systems. An intra-video declarative memory contains fundamental features of the video and semantic contexts extracted from historical QA pairs. Another inter-task procedural memory encodes task-solving experience for further question answering. Additionally, we introduce a hierarchical memory replay mechanism for VideoQA that selects appropriate memories by their relevance and question complexity. Extensive experiments demonstrate that DigimonGPT's accuracy averagely outperforms 13.71% on NExT-QA datasets and 9.89% on Intent-QA datasets over LMM and autonomous agents.
Borui Li 0001, Xingcai Zhang, Tianen Liu
AAAI1
2026 Synthesizing mmWave range-doppler data from videos for privacy-preserving human activity recognition
abstract
Abstract Millimeter-wave radar has shown significant potential in privacy-preserving human activity recognition. However, the lack of diverse radar datasets across various scenarios poses a challenge to the robustness and generalization of deep learning models. To address this limitation, existing works mainly focus on synthesizing micro-doppler data from video, range-doppler data, which provides an extra dimension, has been overlooked due to challenges caused by signal offsets. In this paper, we propose a comprehensive approach for synthesizing range-doppler data from videos by leveraging computer vision techniques and principles of camera imaging. Furthermore, we implement a map enhancement and classification model to facilitate human activity recognition. Our approach is validated on a custom dataset, where the proposed range-doppler synthesis method and classification model achieve an accuracy of 97.3% for activity recognition tasks. This performance is comparable to that of vision-based HAR methods, demonstrating the effectiveness of our proposed scheme in achieving privacy-preserving human activity recognition.
Xuehan Zhang, Shuai Wang 0008, Zhiyuan Cui, Borui Li 0001, Xiaolei Zhou 0001, Zhao-Dong Xu, Shuai Wang 0021
CCF Trans. Pervasive Comput. Interact.4
2026 Real-Batch: Real-Time Adaptive Batch Processing for Accurate Object Detection in Autonomous Driving
abstract
Video object detection stands as a pivotal element within the burgeoning landscape of autonomous driving systems. The exigency to fulfill stringent real-time requisites, while upholding both precision and efficiency in detection, underscores its significance. Although extant methodologies enhance either accuracy or efficiency through the exploitation of spatio-temporal inter-dependencies within the video context, their propensity to conduct detection on discrete frames begets superfluous computations and curbed real-time efficacy. This paper introduces a pioneering approach, called Real-Batch, tailored explicitly to redress this quandary. Real-Batch ingeniously processes batches of video frames uniformly, effectually winnowing out repetitive object detection occurrences. Our methodology is rigorously evaluated on the real-word datasets, scrutinizing four key metrics: accuracy, efficiency, informational value, and adherence to timing constraints. The comprehensive findings substantiate that Real-Batch yields an unparalleled maximal surge in accuracy and efficiency, increasing of 4.2%-13.2% and 24.7%-43.9%, respectively, offering promising advancements for autonomous driving systems.
Tianen Liu, Shuai Wang 0008, Borui Li 0001, Zheng Dong 0002, Guang Wang 0001, Wei Gong 0001, Tian He 0001
IEEE Trans. Mob. Comput.3
2026 InfSquad: SLO-Aware Serverless Machine Learning Inference With Wasm-Assisted Hybrid Functions
abstract
Though serverless computing offers transformative benefits for deploying machine learning (ML) services, it faces challenges in meeting strict real-time service-level objectives (SLOs) of ML inference while maintaining resource efficiency. Fortunately, a new binary instruction format, WebAssembly (or Wasm), offers a promising solution for serverless ML inferences thanks to its short startup time and efficient execution compared with traditional container-based solutions. Therefore, we introduce InfSquad, a serverless ML inference framework designed to balance SLO-aware execution with resource efficiency. The key design of InfSquad is a Wasm-assisted hybrid serverless function runtime to harness the complementary strengths of Wasm and traditional containerized function runtime. Based on the hybrid runtime, InfSquad advocates an SLO-aware runtime scheduling approach that delivers efficient serverless inference. InfSquad leverages a proactive runtime recycling mechanism to increase resource efficiency further. Experiments on real-world applications show that, compared with state-of-the-art serverless systems, InfSquad achieves 19.8% 85.1% SLO violation reduction with 32.5% 74.8% fewer resources.
Borui Li 0001, Hai Wang 0019, Wei Xi 0003, Shuai Wang 0008, Tian He 0001
IEEE Trans. Serv. Comput.2
2025 MobiLoRA: Accelerating LoRA-based LLM Inference on Mobile Devices via Context-aware KV Cache Optimization
abstract
Deploying large language models (LLMs) with low-rank adaptation (LoRA) on mobile devices is promising due to their capability to complete diverse domain-specific tasks while ensuring privacy and accessibility. In this paper, we introduce MobiLoRA to accelerate LoRA-based LLM inference on mobile devices. MobiLoRA focuses on optimizing the key-value (KV) caches due to the limited computing and memory resources of mobile devices. The key insight of MobiLoRA lies in the utilization of two contexts for on-device LoRA serving: semantic-level contexts, such as prompts with shared prefixes, and system-level contexts, such as the application status (e.g., foreground or killed) of LLM requests. Specifically, for semantic-level contexts, MobiLoRA proposes similarity-aware delta encoding, which leverages token-wise similarity in KV caches across LoRA adapters for efficient storage and reuse. Furthermore, MobiLoRA advocates context-aware KV cache management to optimize cache retention and eviction considering the system-level contexts. We fully implement MobiLoRA and compare it with state-of-the-art LLM serving frameworks using real-world mobile device traces. Results show that MobiLoRA accelerates LoRA-based LLM inference by 57.6% on mobile devices.
Borui Li 0001, Haoran Ma 0006, Ligeng Chen
ACL (1)1
2025 TensorShield: Safeguarding On-Device Inference by Shielding Critical DNN Tensors with TEE
abstract
To safeguard user data privacy, on-device inference has emerged as a prominent paradigm on mobile and Internet of Things (IoT) devices. This paradigm involves deploying a model provided by a third party on local devices to perform inference tasks. However, it exposes the private model to two primary security threats: model stealing (MS) and membership inference attacks (MIA). To mitigate these risks, existing wisdom deploys models within Trusted Execution Environments (TEEs), which is a secure isolated execution space. Nonetheless, the constrained secure memory capacity in TEEs makes it challenging to achieve full model security with low inference latency.
Tong Sun 0006, Hailong Lin, Borui Li 0001, Yixiao Teng, Yi Gao 0001, Wei Dong 0001
CCS4
2025 InfScaler: Enabling Efficient ML Inference Serving on Multi-Accelerator Edge Devices via Asymmetric Auto-Scaling
abstract
Nowadays, there is a growing trend to deploy machine learning (ML) models on edge devices. To cope with the increasing resource requirements of current ML models, multi-accelerator edge devices that integrate CPU, GPU, NPU, or TPU in a single SoC gain popularity. However, we observe that existing ML inference serving frameworks are poor in utilizing the unique hardware architecture of these edge devices. In this paper, we present InFSCALER, an efficient ML inference serving framework tailored for multi-accelerator edge devices. InfSCALER discovers the architectural bottleneck of ML models and designs a bottleneck-aware asymmetric auto-scaling technique to facilitate efficient resource allocation for ML models on the edge. Furthermore, InfScaler capitalizes on the hardware’s unified memory feature inherent to edge devices, ensuring efficient data sharing between the asymmetrically scaled model partitions. Our experimental results show that InfScaler achieves up to $\mathbf{1 2 6. 5 9 \%}$ throughput improvement and $\mathbf{2 7. 3 2 \%}$ resource reduction while satisfying the latency requirements compared with the state-of-the-art inference serving approaches.
Borui Li 0001, Tiange Xia, Shuai Wang 0021, Shuai Wang 0008
DAC1
2025 EdgeSched: Adaptive User-Space Scheduling for Serverless Functions in Edge Computing
abstract
Serverless edge computing is attracting increasing attention due to its management-free deployment and on-demand resource provision characteristics. However, the limited resources on edge servers make efficient task scheduling crucial. Specifically, the diverse and dynamic requests in edge scenarios and the intertwined networking and computing process in serverless edge make the existing one-for-all CPU scheduling strategy fall flat. To address the CPU scheduling problem above, we introduce EdgeSched, an adaptive user-space scheduling framework for serverless functions in edge computing. EdgeSched proposes an enclave-based hybrid scheduling technique that separates network I/O tasks from diverse computational tasks efficiently to different enclaves and applies the best-fit scheduling policy on each enclave. Furthermore, EdgeSched leverages deep reinforcement learning to adaptively schedule CPU resources into different enclaves and adapts scheduling different algorithms in the user space. Through experiments on real-world edge devices and workloads, we observe a more than$2 \times$performance improvement compared to state-of-the-art methods.
Chengqing Zhao, Borui Li 0001, Shuai Wang 0021, Shuai Wang 0008
ICPADS2
2025 InfiniCL: Elastic Continual Learning for Resource-Constrained Edge Devices
abstract
On-device continual learning (CL) enables lifelong and privacy-preserving learning for various edge intelligent applications. Increasing the number of model parameters as new learning tasks emerge is effective in ensuring learning quality but inefficient in memory cost, especially for resource-constrained devices. In this paper, we introduce InfiniCL, the first ondevice CL system that dynamically balances memory cost and learning quality. A key idea behind InfiniCL is elastic continual learning: selectively freezing layers in the expanding model and periodically distilling the model, preventing unbounded memory growth while preserving learning quality for new tasks. This novel CL paradigm opens a new challenging problem: how to decide the memory allocation of the model and data to achieve better learning Quality of Service (QoS) under the limited memory budget? To alleviate this challenge, we further propose a Bayesian Optimization-driven algorithm to jointly optimize layer freezing selection and data-model memory allocation. Evaluations show that InfiniCL outperforms state-of-the-art methods on diverse memory constraints, achieving 5.34-7.15% and 2.72-9.36% higher accuracy on CIFAR-100 and ImageNet-100, respectively.
Chenyu Lu, Mengyang Liu, Fang Dong 0001, Borui Li 0001, Ruiting Zhou, Shiyao Ji
IWQoS4
2025 nnWeb: Towards efficient WebGPU-based DNN inference via automatic collaborative offloading
Bing Dong, Muhan Yuan, Borui Li 0001, Zhao-Dong Xu, Shuai Wang 0008
Comput. Networks5
2025 WaWoT: Towards Flexible and Efficient Web of Things Services via WebAssembly on Resource-Constrained IoT Devices
abstract
Web of Things (WoT) is an emerging concept to connect IoT devices to the web using standard interfaces. This provides interoperability between different IoT platforms and enables seamless integration with web and cloud services. However, running sophisticated web services directly on resource-constrained IoT devices is challenging due to limitations in memory, computation, and energy. This paper proposesWaWoT, aWasm-based framework for flexible and efficientWebofThings services.WaWoTallows flexible WoT service development using annotations and automatic partitioning. It also enables dynamic service migration using WebAssembly modules to adapt placement between IoT devices and web clients. We also introduce an ahead-of-time compiler optimized for low memory usage through techniques like streamed compilation and trimming. For energy efficiency, we use optimizations like bulk instruction writing and direct I/O accessing. Safety is ensured through compile-time and run-time analyses to guarantee sandboxed execution. Evaluations demonstrateWaWoTexhibits better flexibility than existing WoT development approaches. Furthermore,WaWoTcan also reduce RAM usage by 84.9x and energy consumption by 1.9x-4.9x over existing WebAssembly runtimes. Overall, it enables efficient, safe, and flexible WoT services on constrained IoT devices.
Borui Li 0001, Hongchang Fan, Yi Gao 0001, Wei Dong 0001
IEEE Trans. Computers1
2025 Jarvis: Zero-code Prototyping of IoT Applications with Composable Hardware-Software Abstractions
abstract
The Internet of Things (IoT) has become an integral part of daily life, enabling seamless interaction between humans and the physical world. However, prototyping IoT applications remains an arduous task, requiring expertise in both hardware and software development. Current low-code and zero-code development approaches fail to address the tight coupling between hardware and software as well as the performance of the generated application, limiting their applicability. We introduce Jarvis, a zero-code prototyping framework for IoT applications that leverages composable hardware-software abstractions. By abstracting physical constraints and introducing a parameter-free self-reflection generation mechanism, Jarvis enables large language models to understand and address the coupling constraints in IoT development with minimized token cost and hallucination. Jarvis also considers the cost and energy efficiency of the resulting IoT prototype, optimizing the performance of the generated solutions. Evaluations demonstrate that Jarvis outperforms state-of-the-art methods, reducing hardware costs by 23.1%–75.5%, saving power consumption by 18.7%–95.2%, and lowering token usage for prototyping by 32.1%–86.2%.
Borui Li 0001, Shuai Wang 0008, Tian He 0001
ACM Trans. Internet Things2
2025 FluidEdge: Expediting Serverless Machine Learning Inference via Bottleneck-Aware Auto-Scaling on Edge SoCs
abstract
Mobile applications based on machine learning (ML) are increasingly relying on offloading to the edge devices for low-latency, resource-efficient computation. Applying serverless computing for these ML applications on the edge offers a promising solution for handling dynamic workloads while meeting user-specified latency service-level objectives (SLOs). However, existing serverless frameworks, with their coarse-grained data parallelism and rigid model partitioning, are inadequate for ML inference on widely adopted edge System-on-Chip (SoC) devices. This paper presents FluidEdge, an edge-native serverless inference framework. FluidEdge identifies bottleneck operators in ML models and addresses them through a novel fine-grained intra-function latency-sensitive auto-scaling approach that dynamically scales inference bottlenecks during online serving. Additionally, it employs inter-function scaling to further prevent latency SLO violations and leverages the unified memory of edge SoCs for efficient data sharing during inference. Experimental results demonstrate that FluidEdge achieves a 37.4% latency improvement and 67.3%-87.6% SLO violation reduction compared to best-performed state-of-the-art serverless inference frameworks.
Borui Li 0001, Tiange Xia, Shuai Wang 0021, Chenhong Cao, Zheng Dong 0002, Shuai Wang 0008
IEEE Trans. Mob. Comput.1
2024 Privacy-preserving Human Activity Recognition via Video-based Range-Doppler Synthesis
abstract
As an important branch of IoT applications, Human activity recognition (HAR) is widely used in daily life, particularly through vision-based methods. However, vision-based HAR has serious privacy issues. How to better and low-cost protect the privacy of users who have already installed the relevant devices is a problem that needs to be solved. To address this challenge, we can solve it by transforming video to privacy-preserving mmWave data. Existing studies have primarily focused on synthesizing micro-Doppler data from video, but there is a lack of methods for synthesizing range-Doppler data. Thus, we present a comprehensive method for synthesizing range-Doppler data from videos and subsequently utilize this synthetic data for HAR. Experimentally, we deploy our range-Doppler synthesis method and classification model on a custom dataset. Experimental results indicate that the model trained with synthetic data achieves accuracy on the custom dataset by 95.7%, which is comparable to the accuracy of vision-based HAR works, and demonstrate that the scheme proposed in this paper achieves privacy-preserving HAR.
Zhiyuan Cui, Luoyu Mei, Siyuan Pei, Borui Li 0001, Xiaolei Zhou 0001
CSCWD4
2024 dTEE: A Declarative Approach to Secure IoT Applications Using TrustZone
abstract
Internet of Things (IoT) applications have recently been widely used in safety-critical scenarios. To prevent sensitive information leaks, IoT device vendors provide hardware-assisted protections, called Trusted Execution Environments (TEEs), like ARM Trust-Zone. Programming a TEE-based application requires separate code for two components, significantly slowing down the development process. Existing solutions tackle this issue by automatic code partition while not successfully applying it in two complicated scenarios: adding trusted logic and interactions with secure peripherals.We propose dTEE, a declarative approach to secure IoT applications based on TrustZone. dTEE proposes a rapid approach that enables developers to declare tiered-sensitive variables and functions of existing applications. Besides, dTEE automatically transforms device drivers into trusted ones. We evaluate dTEE on four real-world IoT applications and seven micro-benchmarks. Results show that dTEE achieves high expressiveness for supporting 50% more applications than existing approaches and reduces 90% of the lines of code against handcrafted development.
Tong Sun 0006, Borui Li 0001, Yixiao Teng, Yi Gao 0001, Wei Dong 0001
IPSN2
2024 SimEnc: A High-Performance Similarity-Preserving Encryption Approach for Deduplication of Encrypted Docker Images
Tong Sun 0006, Borui Li 0001, Jiamei Lv, Yi Gao 0001, Wei Dong 0001
USENIX ATC3
2024 BLEdge: Edge-centric Programming for BLE Applications with Multi-connection Optimization
abstract
Recent years have witnessed the rapid growth of IoT (Internet of Things). Bluetooth Low Energy (BLE) is one of the most popular wireless protocols to implement IoT applications because of its energy efficiency and low-cost properties. However, the development of BLE applications is time-consuming and exhausting. Users are required to write programs for both sides of a BLE connection using complicated low-level APIs. Moreover, it needs much expertise for developers to set appropriate parameters in accordance to different application requirements, especially when there exist multiple concurrent BLE connections. To address these problems, we propose BLEdge , an edge-centric programming approach for BLE applications with multi-connection optimization. First, we propose a wireless bus abstraction for BLE programming. With this, users can write BLE applications in an edge-centric way, as if the BLE-connected peripherals are physically attached to the edge node. Second, we advocate an optimization approach for BLE connection parameters. This optimization approach considers the time slot collision problem under a multi-connection scenario. We conduct extensive experiments with the nRF52840DK platform. Experiment results show that BLEdge can reduce 62.50% to 90.55% LOC (Lines of Code) when developing BLE applications. Furthermore, our parameter optimization approach can reduce up to 42.23% energy consumption.
Borui Li 0001, Jiamei Lv, Wei Dong 0001
ACM Trans. Sens. Networks2
2023 The First Measurement Study of Target Wake Time Mechanism in 802.11ax on COTS Devices
abstract
Nowadays, the 802.11ax protocol is widely adopted in the local area networks. Compared with the former version of 802.11 standards, 802.11ax brings a new power-saving mechanism named TWT (Target Wake Time). In this paper, we conduct a comprehensive measurement study on how TWT mechanism performs on the commercial-off-the-shelf devices in our daily life. To me more specific, we focus on the impact on power consumption, network performance and stability of TWT mechanism on the smartphones. Counterintuitively, we find that TWT mechanism is not as power-saving as we thought in the experimental setup we built, and it also poses nagative impacts on the network stability. We further perform a root-cause analysis to find the rationale behind the measurement results and we also summarized the implications and suggestions to the stakeholders of 802.11ax.
Chengqing Zhao, Borui Li 0001, Shuai Wang 0008, Tian He 0001
ICC2
2023 WebInf: Accelerating WebGPU-based In-browser DNN Inference via Adaptive Model Partitioning
abstract
Artificial intelligence (AI) model inference performance in browsers is constrained, and transmitting data to the server consumes substantial transfer time by cloud computing. In this paper, we investigate the status quo of cloud and browser processing and explore model computation partitioning methods. Our study is rooted in WebGPU and employs the Tensorflow.js framework, encompassing seven AI models spanning computer vision, natural language processing, and automatic speech recognition domains. Leveraging the characteristics of neural network layers, we find a significant performance boost through a method that partitions AI models at layer granularity. We design a system called WebInf to partition AI models at layer granularity between the browser and server for faster inferencing-based adaptive model partitioning. WebInf supports diverse hardware, wireless networks, neural network structures, servers, and adaptive partitioning models for optimal inference performance. We evaluate WebInf on two laptops and servers, demonstrating that WebInf yields inference time improvements of 30% and 52%, respectively, when compared to separate inference execution in servers and browsers. The improvements can even peak at 33% and 69% respectively.
Bing Dong, Tianen Liu, Borui Li 0001, Xiaolei Zhou 0001, Shuai Wang 0008, Zhao-Dong Xu
ICPADS3
2023 RT-BLE: Real-time Multi-Connection Scheduling for Bluetooth Low Energy
Jiamei Lv, Borui Li 0001, Wei Dong 0001
INFOCOM3
2023 LinkLab 2.0: A Multi-tenant Programmable IoT Testbed for Experimentation with Edge-Cloud Integration
Wei Dong 0001, Borui Li 0001, Kaijie Gong, Wenzhao Zhang, Yi Gao 0001
NSDI2
2022 Bringing webassembly to resource-constrained iot devices for seamless device-cloud integration
abstract
Recent years have witnessed the progressive integration between IoT (Internet of Things) devices and the cloud server, which promotes the efficiency and interoperability of IoT applications. WebAssembly, known for its performance and portability, is considered a promising technology to bridge the heterogeneity between devices and the server. Nevertheless, resource-constrained devices, which are commonly deployed in the wild, have difficulty participating in this device-cloud integration because they can hardly run WebAssembly efficiently.
Borui Li 0001, Hongchang Fan, Yi Gao 0001, Wei Dong 0001
MobiSys1
2022 Edge-Centric Programming for IoT Applications With Automatic Code Partitioning
abstract
IoT application development usually involves separate programming at the device side and server side. While separate programming style is sufficient for many simple applications, it is not suitable for many complex applications that involve complex interactions and intensive data processing. We proposeEdgeProg, an edge-centric programming approach to simplify IoT application programming, motivated by the increasing popularity of edge computing. With EdgeProg, users could write application logic in a centralized manner with an augmented If-This-Then-That (IFTTT) syntax and virtual sensor mechanism. The program can be processed at the edge server, which can automatically generate the actual application code and intelligently partition the code into device code and server code, for achieving the optimal latency. EdgeProg employs dynamic linking and loading to deploy the device code on a variety of IoT devices, which do not run any application-specific codes at the start. Results show that EdgeProg achieves an average reduction of 20.96%, 27.8% and 79.41% in terms of execution latency, energy consumption, and lines of code compared with state-of-the-art approaches.
Borui Li 0001, Wei Dong 0001
IEEE Trans. Computers1
2021 WiProg: A WebAssembly-based Approach to Integrated IoT Programming
abstract
Programming a complete IoT application usually requires separated programming for device, edge and/or cloud sides, which slows down the development process and makes the project hardly portable. Existing solutions tackle this problem by proposing a single coherent language while leaving two issues unsolved: efficient migration among the three sides and the platform dependency of the binaries. We propose WiProg, an integrated approach to IoT application programming based on WebAssembly. WiProg proposes an edge-centric programming approach that enables developers to write the IoT application as if it runs on the edge. This is achieved by the peripheral-accessing SDKs and annotations specifying the computation placement. WiProg automatically processes the program to insert auxiliary code and then compile it to WebAssembly. At runtime, WiProg leverages dynamic code offloading with compact memory snapshotting to achieve efficient execution. WiProg also provides interfaces for the customization of offloading policies. Results on real-world applications and computation benchmarks show that WiProg achieves an average reduction by 18.7%~54.3% and 20.1%~57.6% in terms of energy consumption and execution time.
Borui Li 0001, Wei Dong 0001, Yi Gao 0001
INFOCOM1
2021 ThingSpire OS: a WebAssembly-based IoT operating system for cloud-edge integration
abstract
We advocate ThingSpire OS, a new IoT operating system based on WebAssembly for cloud-edge integration. By design, WebAssembly is considered as the first-class citizen in ThingSpire OS to achieve coherent execution among IoT device, edge and cloud. Furthermore, ThingSpire OS enables efficient execution of WebAssembly on resource-constrained devices by implementing a WebAssembly runtime based on Ahead-of-Time (AoT) compilation with a small footprint, achieves seamless inter-module communication wherever the modules locate, and leverages several optimizations such as lightweight preemptible invocation for memory isolation and control-flow integrity. We implement a prototype of ThingSpire OS and conduct preliminary evaluations on its inter-module communication performance.
Borui Li 0001, Hongchang Fan, Yi Gao 0001, Wei Dong 0001
MobiSys1
2021 Automatic Generation of IoT Device Platforms With AutoLink
abstract
With the development of the Internet-of-Things (IoT) industry, developers are no longer content with just prototyping a valid system but eager to create a mature IoT system that explores low power consumption or high extensibility instead. In this article, we present AutoLink, an automatic generation system of IoT device platforms. Users may write AutoLink metaprogram with an expressive syntax to specify their diverse requirements (e.g., battery lifetime, interface extensibility, execution time, and cost) of the generated IoT device platform. Taking the metaprogram as an input, AutoLink automatically transforms it into corresponding optimization problems and generates the optimal hardware configuration that meets user requirements best. Toward this, AutoLink also offers a cross-platform, duty cycle-aware power model and a time model to estimate the lifetime and execution period of an IoT system. We implement AutoLink and evaluate its performance using real-world IoT applications. Results show that AutoLink generates the optimal hardware configuration that meets diverse user requirements. Moreover, AutoLink achieves superior power estimation accuracy of IoT device platforms compared with the state-of-the-art approach.
Borui Li 0001, Wei Dong 0001
IEEE Internet Things J.1
2021 Queec: QoE-aware Edge Computing for IoT Devices under Dynamic Workloads
abstract
Many IoT applications have the requirements of conducting complex IoT events processing (e.g., speech recognition) that are hardly supported by low-end IoT devices due to limited resources. Most existing approaches enable complex IoT event processing on low-end IoT devices by statically allocating tasks to the edge or the cloud. In this article, we present Queec, a QoE-aware edge computing system for complex IoT event processing under dynamic workloads. With Queec, the complex IoT event processing tasks that are relatively computation-intensive for low-end IoT devices can be transparently offloaded to nearby edge nodes at runtime. We formulate the problem of scheduling multi-user tasks to multiple edge nodes as an optimization problem, which minimizes the overall offloading latency of all tasks while avoiding the overloading problem. We implement Queec on low-end IoT devices, edge nodes, and the cloud. We conduct extensive evaluations, and the results show that Queec reduces 56.98% of the offloading latency on average compared with the state-of-the-art under dynamic workloads, while incurring acceptable overhead.
Borui Li 0001, Wei Dong 0001, Gaoyang Guan, Tao Gu 0001, Jiajun Bu, Yi Gao 0001
ACM Trans. Sens. Networks1
2020 EdgeProg: Edge-centric Programming for IoT Applications
abstract
IoT application development usually involves separate programming at the device side and server side. While separate programming style is sufficient for many simple applications, it is not suitable for many complex applications that involve complex interactions and intensive data processing. We propose EdgeProg, an edge-centric programming approach to simplify IoT application programming, motivated by the increasing popularity of edge computing. With EdgeProg, users could write application logic in a centralized manner with an augmented If-This-Then-That (IFTTT) syntax and virtual sensor mechanism. The program can be processed at the edge server, which can automatically generate the actual application code and intelligently partition the code into device code and server code, for achieving the optimal latency. EdgeProg employs dynamic linking and loading to deploy the device code on a variety of IoT devices, which do not run any application-specific codes at the start. Results show that EdgeProg achieves an average reduction of 20.96% and 79.41% in terms of execution latency and lines of code, compared with state-of-the-art approaches.
Borui Li 0001, Wei Dong 0001
ICDCS1
2020 TinyLink 2.0: integrating device, cloud, and client development for IoT applications
abstract
The recent years have witnessed the rapid growth of IoT (Internet of Things) applications. A typical IoT application usually consists of three essential parts: the device side, the cloud side, and the client side. The development of a complete IoT application is very difficult for non-expert developers because it involves drastically different technologies and complex interactions between different sides. Unlike traditional IoT development platforms which use separate approaches for these three sides, we present TinyLink 2.0, an integrated IoT development approach with a single coherent language. It achieves high expressiveness for diverse IoT applications by an enhanced IFTTT rule design and a virtual sensor mechanism which helps developers express application logic with machine learning. Moreover, TinyLink 2.0 optimizes the IoT application performance by using both static and dynamic optimizers, especially for resource-constrained IoT devices. We implement TinyLink 2.0 and evaluate it with eight case studies, a user study, and a detailed evaluation of the proposed programming language as well as the performance optimizers. Results show that TinyLink 2.0 can speed up IoT development significantly compared with existing approaches from both industry and academia, while still achieving high expressiveness.
Gaoyang Guan, Borui Li 0001, Yi Gao 0001, Yuxuan Zhang 0003, Jiajun Bu, Wei Dong 0001
MobiCom2
2020 TinyLink: A Holistic System for Rapid Development of IoT Applications
abstract
Rapid development is essential for IoT (Internet of Things) application developers to obtain first-mover advantages and reduce the development cost. In this article, we present TinyLink, a holistic system for rapid development of IoT applications. The key idea of TinyLink is to use a top-down approach for designing both the hardware and the software of IoT applications. Developers write the application code in a C-like language to specify the key logic of their applications, without dealing with the details of the specific hardware components. Taking the application code as input, TinyLink automatically generates the hardware configuration as well as the binary program executable on the target hardware platform. TinyLink provides unified APIs for applications to interact with the underlying hardware components. We implement TinyLink and evaluate its performance using real-world IoT applications. Results show that (1) TinyLink achieves rapid development of IoT applications, reducing 52.58% of lines of code on average compared with traditional approaches; (2) TinyLink searches a much larger design space and thus can generate a superior solution for the hardware configuration, compared with the state-of-the-art approach; (3) TinyLink incurs acceptable overhead in terms of execution time and program memory.
Wei Dong 0001, Borui Li 0001, Gaoyang Guan, Zhihao Cheng, Yi Gao 0001
ACM Trans. Sens. Networks2
2019 Demo: Integrated Development of IoT Applications with OneLink
Gaoyang Guan, Yuxuan Zhang 0003, Borui Li 0001, Wei Dong 0001, Yi Gao 0001, Jiajun Bu
EWSN3