Weisong Shi

dblp:s/WeisongShi · DBLP profile ↗
← Back
204ranked-venue papers
16as first author
74since 2021 · last 2026
0000-0001-5864-4675ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 81 · 5 first-author · 28 since 2021Computer networks · 47 · 4 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 30 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 11 · 7 since 2021Software engineering, systems software and programming languages · 11 · 2 first-author · 2 since 2021Security and privacy · 10 · 5 since 2021Human-computer interaction and ubiquitous computing · 10Databases, data management, data science and information retrieval · 9 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 DORA: Dataflow-Instruction Orchestration Architecture for DNN Acceleration
abstract
Modern DNN workloads are increasingly diverse in operation types, tensor shapes, and execution dependencies, making it difficult for customized accelerators to sustain high hardware efficiency across models. We propose DORA, an instruction-based overlay architecture as well as a compilation framework that explicitly describes dataflow via a proposed ISA, enabling fine-grained control of data movement, computation, and synchronization at the layer level.
Xingzhen Chen, Zhuoping Yang, Jinming Zhuang, Shixin Ji, Sarah Schultz, Zheng Dong 0002, Weisong Shi, Peipei Zhou 0001
FCCM7
2026 DORA: Dataflow-Instruction Orchestration Architecture for DNN Acceleration
abstract
As deep neural networks develop significantly more diverse and complex, achieving high performance and efficiency on complicated DNN models faces pressing challenges. Modern DNN workloads are increasingly diverse in operation types, tensor shapes, and execution dependencies, making it difficult to sustain high hardware efficiency across models. In addition, a generic accelerator often incurs substantial overhead when executing diverse workloads.
Xingzhen Chen, Zhuoping Yang, Jinming Zhuang, Shixin Ji, Sarah Schultz, Zheng Dong 0002, Weisong Shi, Peipei Zhou 0001
ACM Great Lakes Symposium on VLSI7
2026 Efficient Deployment of Lightweight LLMs on Edge Devices: Kernel-Level Profiling and Cross-Platform Insights
Xinyao Liu, Xingzhou Zhang, Weisong Shi
ICDCS3
2026 A Graph-Based Metric on Dataset Diversity and Redundancy Evaluation
Yuankai He, Ilya Safro, Weisong Shi
IV4
2026 A Faster and More Reliable Middleware for Autonomous Driving Systems
Yuankai He, Weisong Shi
IV2
2026 Physical Intelligence on the Edge: A Vision for the Decade Ahead
Weisong Shi, Zheng Dong 0002, Peipei Zhou 0001
J. Comput. Sci. Technol.1
2026 Introduction to the Special Issue on Autonomous Driving
Weisong Shi, Zheng Dong 0002, Lily, Johannes Betz
ACM Trans. Internet Things1
2026 Real-time Guarantee of Autonomous Driving Systems: State of the Art and Challenges
abstract
Autonomous Driving (AD) has garnered significant attention in recent years across multiple domains. Despite notable advancements in algorithms and hardware, large-scale deployment of autonomous vehicles remains constrained by the lack of adequate real-time guarantees. This article comprehensively investigates real-time assurance challenges in AD systems from theoretical and practical perspectives. First, we introduce foundational concepts of real-time systems and analyze common modeling approaches, including the multi-rate DAG and processing chain DAG models. We then delve into the task scheduling and communication mechanisms of three representative middleware systems–ROS2, Cyber, and ERDOS–and the seL4 operating system. Our analysis reveals their underlying design philosophies and optimization strategies for real-time performance. The findings highlight the critical difficulty of meeting stringent real-time requirements in autonomous systems. By offering insights into current limitations and opportunities for improvement, this article aims to establish a deeper understanding of real-time assurance issues and encourage greater focus on system-level guarantees within the AD community.
Tianze Wu, Weisong Shi
ACM Trans. Internet Things2
2025 EMATO: Energy-Model-Aware Trajectory Optimization for Autonomous Driving
abstract
Autonomous driving currently lacks robust evidence of energy efficiency when using energy-model-agnostic trajectory planning. To address this, we explore how differential energy models can be effectively utilized under varying driving conditions to enhance energy efficiency. Furthermore, we propose an online nonlinear programming approach that optimizes polynomial trajectories generated by the Frenet polynomial method while incorporating traffic trajectory data and road slope predictions. Through case studies, quantitative analyses, and ablation studies conducted on both sedan and truck models, we demonstrate the effectiveness of the proposed method. [Video]
Zhaofeng Tian, Lichen Xia, Weisong Shi
ICRA3
2025 Investigating Security Threats in Multi-Tenant ROS 2 Systems
abstract
Robot Operating System (ROS) has been widely used to develop robotic applications. The first generation of ROS generally lacks security features, and ROS 2 is introduced with security support. However, security concerns still exist for running ROS in practical multi-tenant environments. In this paper, we conduct an in-depth investigation into the security of ROS 2. We focus on vulnerabilities in ROS nodes and topics and intend to explore methods to break the isolation and security mechanisms systematically. We devise a set of strategies that can be exploited by attackers to escalate privilege or cause information leakage in a multi-tenant environment. These attacks can bypass existing isolation and security mechanisms, including ROS 2's native security module. To validate our findings, we employ simulations across various real-world scenarios to demonstrate how attackers could exploit these vulnerabilities to bypass existing security mechanisms. Finally, we present several defense practices to mitigate these identified threats.
Lichen Xia, Weisong Shi
ICRA3
2025 Evaluating the Impact of Network Latency on the Teleoperation of Autonomous Vehicles
abstract
Teleoperation provides a critical fallback for autonomous vehicles (AVs), enabling remote operators or automated systems to assume control when autonomous functionality is insufficient. However, network latency remains a key challenge that can degrade control accuracy and compromise safety. This study systematically evaluates the impact of network latency on teleoperated driving through system-level experiments using a physical robotic vehicle. Controlled delays ranging from 0 ms to 240 ms were introduced via a custom latency injection module, isolating latency effects from human factors. Path-following accuracy and speed regulation were assessed using a pure-pursuit controller and LiDAR-based localization. Results show that increasing latency leads to progressive degradation in both lateral and longitudinal control, with performance declining sharply beyond 160 ms. These findings provide quantitative benchmarks for latency tolerance in real-time teleoperation and offer practical guidance for designing resilient communication and control architectures for connected and autonomous vehicles.
Mustafa Alsolami, Arpan Bhattacharjee, Weisong Shi
SEC3
2025 Drone-Based Standardized Environmental and Usage Assessment of Parks and Trails
abstract
We introduce a drone-centric workflow that complements fixed trail cameras to deliver objective, standardized, and scalable assessments of park and trail infrastructure. Traditional human audits, such as PARA and the Boston Block Walk, are labor intensive and subjective and lack the corridor scale perspective needed for comprehensive maintenance planning. Our system addresses this gap by using (i) a multirotor UAV equipped with synchronized thermal and RGB cameras to document surface conditions over vast areas and quantify defect dimensions, and (ii) low-cost fixed cameras to continuously monitor user counts, activity types, and intensity (MET) values. Drone imagery is calibrated via homography to convert pixel measurements to real-world units, and a thermal RGB data fusion pipeline detects cracks, moisture-softened patches, and other defects. A case study on a community tennis court demonstrated a mean absolute percentage error (MAPE) of 11.6% between drone-estimated and tape-measured crack lengths. Along Delaware's Jack A. Markell and James F. Hall trails, camera-based usage analytics revealed patterns consistent with intercept surveys: ≈ 20% of the users were cyclists and ≈ 98% engaged in recreational activities, while estimating speed-derived METs for walkers, runners, and cyclists. Compared to an 11-mile manual audit that required 26 raters, our approach reduced costs by approximately $4,000 and produced reusable digital evidence. We conclude that drones are the necessary backbone for corridor-scale, time-bounded condition assessments, while fixed cameras remain essential for long-term use monitoring; together, they enable auditable, repeatable, and cost-effective management of urban greenspace.
Arpan Bhattacharjee, Matthew Saponaro, Weisong Shi
SEC3
2025 Open-Vocabulary Object Detection with Driving-Aware Multi-Scale Feature Fusion for Autonomous Driving
abstract
Open-vocabulary object detection (OVD) is crucial for handling dynamic real-world driving scenarios. Inspired by YOLO-World, we propose OpenVocab-Auto, an open-vocabulary object detection framework with driving-aware multi-scale feature fusion for autonomous driving scenarios. Our system introduces three key innovations: (1) a context-adaptive prompt engine that significantly reduces computational overhead compared to global prompt strategies, (2) hierarchical vision-language alignment for improved small object detection, and (3) real-time optimization achieving 27 FPS on NVIDIA Jetson AGX Orin through TensorRT acceleration. On RTX 3080 (FP16 full model), the framework achieves 0.923 F1 for parking space detection and 0.847 F1 for zero-shot obstacle recognition. On Jetson Orin (TensorRT INT8 model), the corresponding scores are 0.811 and 0.333, respectively, under the same evaluation protocol.
Yongtao Yao, H. Peter Hofstee, Weisong Shi
SEC4
2025 iFLOW: An Intelligent and Scalable Multi-Model Federated Learning Framework on the Wheels
abstract
The high mobility characteristics of connected vehicles present noteworthy difficulties in the domain of federated learning. Based on our understanding, current federated learning strategies do not tackle the challenge of continuously training multiple models for vehicles in constant motion, which are subject to variable network conditions and changing environments. In response to this challenge, we have created and implemented iFLOW, a versatile and intelligent multi-model federated learning infrastructure specifically designed for highly mobile-connected vehicles. iFLOW addresses these challenges by integrating four key aspects: (1) a strategically devised model allocation algorithm that dynamically selects vehicle computing units for distinct model training tasks, optimizing for both resource efficiency and performance; (2) a dynamic client vehicle joining mechanism that ensures smooth participation of vehicles, even in the face of signal loss or weak connectivity, mitigating disruptions in the training process; (3) integration of a large language model (Llama3.3 70B) as an intelligent arbiter for decision-making within the framework, enhancing adaptability and robustness; and (4) real-world deployment and testing on distributed vehicular devices to validate the approach. The experimental evaluation demonstrates that iFLOW allows multiple models to train asynchronously and outperform centralized training. These results affirm the effectiveness of iFLOW in practical, real-world scenarios involving highly mobile vehicular networks.
Qiren Wang, Yongtao Yao, Nejib Ammar, Weisong Shi
IEEE Trans. Intell. Transp. Syst.4
2024 FSPDE: A Full Stack Plausibly Deniable Encryption System for Mobile Devices
abstract
In today's digital landscape, the ubiquity of mobile devices underscores the urgent need for stringent security protocols in both data transmission and storage. Plausibly deniable encryption (PDE) stands out as a pivotal solution, particularly in jurisdictions marked by rigorous regulations or increased vulnerabilities of personal data. However, the existing PDE systems for mobile platforms have evident limitations. These include vulnerabilities to multi-snapshot attacks over RAM and flash memory, an undue dependence on non-secure operating systems, traceable PDE entry point, and a conspicuous PDE application prone to reverse engineering.
Jinghui Liao, Niusen Chen, Lichen Xia, Bo Chen 0028, Weisong Shi
CODASPY5
2024 BFTRAND: Low-Latency Random Number Provider for BFT Smart Contracts
abstract
Random numbers play a crucial role in decen-tralized applications (dApps) like decentralized finance (DeFi) and non-fungible tokens (NFTs). However, their generation faces challenges due to blolckchain's deterministic and decentralized nature, risking smart contract security and ecosystem stability. Prior solutions, including Oracles, employing commit-execute schemes, suffer from higher transaction fees, extended processing times, and increased on-chain storage, compromising efficiency. This paper proposes a novel random number provider (RNP) protocol for smart contracts, eliminating dependencies on traditional commit-execute approaches. Furthermore, we systematically identify potential random number-related attacks on smart contracts, particularly Post-reveal Undo Attacks (PUAs), where attackers may reverse contract operations when randomness is unfavorable, and discuss the security requirements. Our protocol addresses these attacks by (1) incorporating distributed random beacons (D RBs) with consensus processes, bridging the semantic gap between DRB and consensus, and (2) thoroughly analyzing and classifying four types of PUA and offering robust mitigations, alongside presenting a security proof. Our experiments show the protocol significantly enhances response times and security for random number queries in smart contracts, slashing request fees by at least 89 % and reducing on-chain data by 76.4% versus current methods. This work advances the integration of DRB protocols and consensus mechanisms, securing and optimizing random number applications in dApps, thus fostering the creation of more dependable, robust systems.
Jinghui Liao, Borui Gong, Wenhai Sun, Fengwei Zhang, Zhenyu Ning, Man Ho Au, Weisong Shi
DSN7
2024 Quantitative Analysis of Storage Requirement for Autonomous Vehicles
abstract
This study addresses the critical aspect of data storage requirements for Autonomous Vehicles (AVs). With AVs generating substantial amounts of data daily, understanding these requirements is vital for AV storage systems design, enhancing vehicle safety, efficiency, and operational integrity. Through a comprehensive analysis of onboard sensor and CAN bus data, alongside a novel mathematical model, this research offers insights into the storage needs, assesses system durability, and proposes a tailored storage solution and system architecture. The findings aim to guide the development of future AV storage systems, emphasizing the importance of data-driven decision-making in AV technology advancements.
Yuankai He, Weisong Shi
HotStorage4
2024 AyE-Edge: Automated Deployment Space Search Empowering Accuracy yet Efficient Real-Time Object Detection on the Edge
Chao Wu 0006, Yifan Gong 0004, Liangkai Liu, Mengquan Li, Yushu Wu, Xuan Shen, Geng Yuan, Weisong Shi, Yanzhi Wang 0001
ICCAD9
2024 TA-ASF: Attention-Sensitive Token Sampling and Fusing for Visual Transformer Models on the Edge
abstract
Vision Transformers ($V$iTs) have made significant progress in achieving performance comparable to traditional convolutional neural networks in computer vision tasks. However, high computational complexity restricts their application to resource-constrained edge devices. Previous methods for pruning redundant tokens have shown that it is possible to balance performance and computational cost by reducing the number of tokens. Unfortunately, simply removing redundant tokens often leads to the loss of crucial information. To address this issue, we propose a novel token compression scheme called TA-ASF. This scheme considers both the global role of low-importance tokens and the redundancy among similar tokens. TA-ASF employs novel approaches for token sampling and fusion, which are directly applicable to$V$iTs without introducing additional trainable parameters. A comprehensive evaluation against several edge devices demonstrates our method effectively reduces model complexity while preserving Top-1 accuracy. Experimental results show that on the ImageNet dataset, the proposed method reduces FLOPs by 37% and increases throughput by 1.48 times on the DeiT-S model, with only a 0.1% decrease in accuracy. Specifically, on the DeiT-B model, the proposed method decreases FLOPs by 35% and increases throughput by 1.52 times while maintaining the same accuracy.
Junquan Chen, Xingzhou Zhang, Wei Zhou 0011, Weisong Shi
SEC4
2024 An Efficient Data Transmission Framework for Connected Vehicles
abstract
Connected vehicles (CVs) face significant challenges in continuous big data transmission, resulting in high transmission bandwidth costs and impacting real-time decision-making. To address this, we propose two dynamic, driving-aware compression mechanisms based on reinforcement learning and temporal compressive sensing to intelligently compress video data. These mechanisms adapt to driving conditions, reducing bandwidth while preserving sufficient information for accurate applications such as object detection and ensuring high-quality reconstruction when needed. We also implement a Vehicle-EdgeServer-Cloud (VEC) closed-loop framework that integrates these mechanisms. Specifically, a lightweight vehicle model performs real-time detection on compressed data (measurements), while the EdgeServer receives measurements and reconstructs scenes if needed. The measurements, reconstructed video, and analysis results are then sent to the cloud for vehicle model updates. Unlike conventional methods, our framework seamlessly adapts across vehicles, Edge-Servers, and the cloud, supporting efficient data transmission and dynamic model updates. Extensive evaluations were conducted on our designed roadside unit platform and robotic vehicle, both equipped with industry-grade sensors and computing units. The results demonstrate an 18x reduction in bandwidth at 320KB/s while maintaining high detection accuracy and reconstruction quality compared to non-adaptive measurements, highlighting the framework's promising real-world applications for CVs.
Yongtao Yao, Junzhou Chen 0002, Sidi Lu, Weisong Shi
SEC5
2024 Tentacles: A Middleware with Multi-Network Communication Reliability for Vehicle-Infrastructure Cooperative Autonomous Driving
abstract
Vehicle-Infrastructure Cooperative Autonomous Driving (CAD) is a new paradigm of autonomous driving, which relies on the cooperation between intelligent roads and autonomous vehicles. This paradigm has been shown to be safer and more efficient compared to the on-vehicle-only autonomous driving paradigm. Our real-world deployment data indicate that the effectiveness of Vehicle-Infrastructure CAD is constrained by the reliability and performance of commercial communication networks. This paper targets this exact problem and proposes Tentacles, a middleware to achieve high communication reliability between intelligent roads and autonomous vehicles, in the context of Vehicle-Infrastructure CAD. Specifically, Tentacles dynamically matches Vehicle-Infrastructure CAD applications and the underlying communication technologies based on varying communication performance and quality needs. Evaluation results confirm that Tentacles reduces deadline violations by more than 88%, significantly improving the reliability of Vehicle-Infrastructure CAD systems.
Tianze Wu, Sa Wang, Yungang Bao, Weisong Shi
VTC Fall4
2024 WiLDAR: WiFi Signal-Based Lightweight Deep Learning Model for Human Activity Recognition
abstract
In recent years, the WiFi channel state information (CSI) has been increasingly used for human activity recognition (HAR) during activities of daily living, because of nonintrusiveness and privacy preserving properties. However, most previous works require complex processing of CSI signals, and the large number of classification network parameters significantly increases the recognition time and deployment costs. Accordingly, a WiFi signal-based lightweight deep learning (WiLDAR) network is developed in this study to ensure systematic operation on edge computing devices. We combine the random convolution kernel with deep separable convolution and residual structure, so that WiLDAR can easily extract CSI signal features without filtering and denoising. The parameter number and training time of WiLDAR are, thus, much less than those of previous neural networks. In addition, a tiny HAR system using only Raspberry Pi and router is implemented. Experiments verify that WiLDAR can achieve real-time HAR on Internet of Things devices, which makes HAR deployment more convenient. We test WiLDAR on three different fine-grained action data sets to achieve 99%, 93.5%, and 97.5% recognition accuracy, respectively. The demonstrated learning capability of WiLDAR makes it an excellent option for the remote HAR.
Fuxiang Deng, Emil Jovanov, Houbing Song, Weisong Shi, Yuan Zhang 0007, Wenyao Xu
IEEE Internet Things J.4
2024 E3-UAV: An Edge-Based Energy-Efficient Object Detection System for Unmanned Aerial Vehicles
abstract
Motivated by the advances in deep learning techniques, the application of unmanned aerial vehicle (UAV)-based object detection has proliferated across a range of fields, including vehicle counting, fire detection, and city monitoring. While most existing research studies only a subset of the challenges inherent to UAV-based object detection, there are few studies that balance various aspects to design a practical system for energy consumption reduction. In response, we present the E3-UAV, an edge-based energy-efficient object detection system for UAVs. The system is designed to dynamically support various UAV devices, edge devices, and detection algorithms, with the aim of minimizing energy consumption by deciding the most energy-efficient flight parameters (including flight altitude, flight speed, detection algorithm, and sampling rate) required to fulfill the detection requirements of the task. We first present an effective evaluation metric for actual tasks and construct a transparent energy consumption model based on hundreds of actual flight data to formalize the relationship between energy consumption and flight parameters. Then, we present a lightweight energy-efficient priority decision algorithm based on a large quantity of actual flight data to assist the system in deciding flight parameters. Finally, we evaluate the performance of the system, and our experimental results demonstrate that it can significantly decrease energy consumption in real-world scenarios. Additionally, we provide four insights that can assist researchers and engineers in their efforts to study UAV-based object detection further.
Jiashun Suo, Xingzhou Zhang, Weisong Shi, Wei Zhou 0011
IEEE Internet Things J.3
2024 VPI: Vehicle Programming Interface for Vehicle Computing
Baofu Wu, Ren Zhong, Jian Wan 0001, Ji-Lin Zhang, Weisong Shi
J. Comput. Sci. Technol.6
2024 Joint Service Request Scheduling and Container Retention in Serverless Edge Computing for Vehicle-Infrastructure Collaboration
abstract
Lightweight and layered structure containers in serverless edge computing (SEC) provide flexible service configurations and computing for vehicles with diverse service requests in the Vehicle-Infrastructure Collaboration (VIC) environment. Despite progress in service request scheduling for the VIC system, the effect of layer sharing between different service images on request scheduling has not been fully explored. Additionally, the cold-start latency of service containers in SEC can significantly degrade the responsiveness of vehicle services, and container retention is proposed to minimize its impact and improve overall system performance. However, the existing research neglects the complex coupling relationship between request scheduling and container retention decisions, while focusing on the single decision optimization problem. Consequently, minimizing system costs by single decision optimization may not achieve the effect of joint decision optimization. To bridge this gap, we study the joint service request scheduling and container retention problem based on layer sharing and container caching. First, we model the joint decision problem with specific constraints and aim to minimize the long-term system cost while considering vehicle mobility. Second, an online co-decision scheme called Onco is proposed to solve the problem, which incorporates request scheduling and container retention for multiple vehicle services. Finally, both synthetic and real trace-driven simulation experiments have been conducted to evaluate the performance of Onco. The experimental results show that Onco outperforms state-of-the-art baselines in terms of system cost reduction and response time improvement.
Shihong Hu, Zhihao Qu, Bin Tang 0002, Guanghui Li 0001, Weisong Shi
IEEE Trans. Mob. Comput.6
2023 Learning Pruned Structure and Weights Simultaneously from Scratch: an Attention based Approach
abstract
As a deep learning model typically contains millions of trainable weights, there has been a growing demand for a more efficient network structure with reduced storage space and improved run-time efficiency. Pruning is one of the most popular network compression techniques. In this paper, we propose a novel unstructured pruning pipeline, Attention-based Simultaneous sparse structure and Weight Learning (ASWL). In ASWL, an efficient algorithm is proposed to calculate the pruning ratios layer-wisely from attentions, and both weights for the dense network and the sparse network are tracked so that the pruned structure is simultaneously learned from randomly initialized weights. Our experiments on MNIST, Cifar10, and ImageNet show that ASWL achieves superior pruning results in terms of accuracy, pruning ratio and operating efficiency when compared with state-of-the-art network pruning methods.
Qisheng He, Weisong Shi, Ming Dong 0001
IEEE Big Data2
2023 An Open Approach to Energy-Efficient Autonomous Mobile Robots
abstract
Autonomous mobile robots (AMRs) have the capability to execute a wide range of tasks with minimal human intervention. However, one of the major limitations of AMRs is their limited battery life, which often results in interruptions to their task execution and the need to reach the nearest charging station. Optimizing energy consumption in AMRs has become a critical challenge in their deployment. Through empirical studies on real AMRs, we have identified a lack of coordination between computation and control as a major source of energy inefficiency. In this paper, we propose a comprehensive energy prediction model that provides real-time energy consumption for each component of the AMR. Additionally, we propose three path models to address the obstacle avoidance problem for AMRs. To evaluate the performance of our energy prediction and path models, we have developed a customized AMR called Donkey, which has the capability for fine-grained (millisecond-level) end-to-end power profiling. Our energy prediction model demonstrated an accuracy of over 90% in our evaluations. Finally, we applied our energy prediction model to obstacle avoidance and guided energy-efficient path selection, resulting in up to a 44.8% reduction in energy consumption compared to the baseline.
Liangkai Liu, Ren Zhong, Aaron Willcock, Nathan Fisher, Weisong Shi
ICRA5
2023 Poster: Edge-Assisted Over-the-Air Software Updates
abstract
The exploration of software Over-the-Air (OTA) updates for automotive applications is currently very limited. Our work introduces an edge-assisted framework for automotive OTA updates that carefully accounts for various factors, including different software models in vehicles, communication distances, and cluster sizes. We present valuable insights using key evaluation metrics like update speed, data transmission efficiency, and success rate, accompanied by a thorough scalability analysis. Our research involves three distinct vehicle software models: ResNet-18 (46.8 MB), ResNet-50 (102.5 MB), and Faster R-CNN (175.2 MB). These models are used to evaluate update performance across eight distance categories ranging from 0 to 21 meters with a 3-meter interval. We also utilize diverse computing platforms to assess the success rate and conduct a comprehensive scalability analysis. This innovative approach significantly advances our understanding and practical implementation of OTA updates in the automotive field.
Arpan Bhattacharjee, Hamza Mahmood, Sidi Lu, Nejib Ammar, Akila Ganlath, Weisong Shi
SEC6
2023 Poster: Enhancing Autonomous Vehicles Safety Through Edge-Based Anomaly Detection
abstract
With the growth of vehicular computing capacity, there is an increasing demand for real-time data processing. However, data is sometimes not optimal for purposes such as storage or training. To address this issue, we propose a solution to enhance vehicle safety by generating abnormal image data. We then leverage machine learning algorithms to detect and classify these anomalies while vehicles are in operation. An edge-based anomaly detection approach will be applied to prevent accidents and enhance the safety of connected vehicles.
Qiren Wang, Ruijie Feng, Weisong Shi
SEC3
2023 DICE: Dynamic In-Situ Control for Edge-Based Applications
abstract
This paper focuses on addressing computational constraints and energy limitations prevalent in edge-based applications through an innovative approach, dynamic in-situ control for edge-based applications (DICE). DICE capitalizes on the burgeoning trend in vehicle sensor technologies, such as camera, Radar, and LiDAR, which are becoming increasingly powerful and capable of performing pre-processing computations. DICE introduces a concept of "downstream offloading", which distinguishes it from traditional offloading approaches that typically offload computational tasks from edge devices to more powerful Edge Servers. In contrast, DICE offloads part of the computational tasks from the Edge Server to the sensor itself, thereby optimizing data processing at the source and reducing the volume of data transmission required. This approach not only addresses the latency bottleneck frequently encountered in energy-intensive neural networks but also enhances the efficiency of data processing by selectively filtering out non-critical frames based on event-triggering mechanisms. DICE leverages the unique strengths of portable devices such as smartwatches and smartphones, even with their inherent computational and power limitations. The framework consists of an adaptive control layer for dynamic task allocation and an application layer designed to deploy quantized models on System on Chips (SoCs) like TinyML, thereby improving the efficiency of AI-driven applications while conservatively utilizing energy. This system proposes a sustainable, energy-efficient pathway for future edge-based applications.
Yongtao Yao, Liping Julia Zhu, Weisong Shi
SEC3
2023 Toward Resilient Network Slicing for Satellite-Terrestrial Edge Computing IoT
abstract
Satellite–terrestrial edge computing networks (STECNs) emerged as a global solution to support multiple Internet of Things (IoT) applications in 6G networks. The enabling technologies to slice STECNs, such as software-defined networking (SDN), satellite edge computing (EC), and network function virtualization (NFV) are key to realizing this vision. In this article, we survey and analyze network slicing (NS) solutions for STECNs. We discuss slice management and orchestration for different STECNs integration architectures, satellite EC, mmWave/THz, and artificial intelligence solutions to make NS adaptive. In addition, we identify challenges and open issues to slice STECNs. In particular, resilient NS is crucial for essential and critical services. Network failures are unavoidable in large networks and can cause significant disruptions in NS, compromising many services. To this end, we present a resilient NS design to cope with failures and guarantee service continuity which is agnostic to the integration architecture and inherently multidomain. Further, we present strategies to achieve resilient networking and slicing in STECNs, including planning and provisioning of redundant network resources, design rules for service level agreement decomposition, and cross-domain solutions to detect and mitigate failures. Finally, promising future research directions are highlighted. This article provides valuable guidelines for slicing STECNs and will benefit key sectors, such as smart healthcare, e-commerce, Industrial IoT, education, and among others.
Haitham H. Esmat, Beatriz Lorenzo, Weisong Shi
IEEE Internet Things J.3
2023 A Novel Architecture Combining Oracle With Decentralized Learning for IIoT
abstract
The rapid development of digital technology is reshaping the architecture of the Industrial Internet of Things (IIoT). The traditional architecture cannot process vast amounts of data exchanges and provide entities with trust. The future IIoT is expected to be a decentralized architecture in which blockchain and digital twin-driven IIoT can enable trusted data exchanges. However, this architecture cannot obtain huge amounts of external real-time data and isolated data. Moreover, it cannot handle complex industrial computing tasks. Therefore, we combine oracle with decentralized learning to propose a novel IIoT-oriented digital twin architecture. We also propose an effective decentralized collaboration mechanism to support external data and resources exchanges. Moreover, we propose a novel computing collaboration mechanism to expand the learning capabilities of the industrial ecology. Experiments show that our proposed paradigm has less processing time, a more stable process, and better learning ability compared to other paradigms.
Yijing Lin, Zhipeng Gao 0001, Weisong Shi, Qian Wang 0015, Huangqi Li, Miaomiao Wang 0003, Yang Yang 0006, Lanlan Rui
IEEE Internet Things J.3
2023 CA-DTS: A Distributed and Collaborative Task Scheduling Algorithm for Edge Computing Enabled Intelligent Road Network
Shihong Hu, Quyuan Luo, Guanghui Li 0001, Weisong Shi
J. Comput. Sci. Technol.4
2023 Erratum to: CA-DTS: A Distributed and Collaborative Task Scheduling Algorithm for Edge Computing Enabled Intelligent Road Network
Shihong Hu, Quyuan Luo, Guanghui Li 0001, Weisong Shi
J. Comput. Sci. Technol.4
2023 Deep Reinforcement Learning Based Computation Offloading and Trajectory Planning for Multi-UAV Cooperative Target Search
abstract
Unmanned aerial vehicles (UAVs) are widely used for surveillance and monitoring to complete target search tasks. However, the short battery life and moderate computational capability hinder UAVs to process computation-intensive tasks. The emerging edge computing technologies can alleviate this problem by offloading tasks to the ground edge servers. How to evaluate the search process so as to make optimal offloading decisions and make optimal flying trajectories represent fundamental research challenges. In this paper, we propose to utilize the concept of uncertainty to evaluate the search process, which reflects the reliability of the target search results. Thereafter, we propose a deep reinforcement learning (DRL) technique to jointly make optimal computation offloading decisions and flying orientation choices for multi-UAV cooperative target search. Specifically, we first formulate an uncertainty minimization problem based on the established system model. By introducing a reward function, we prove that the uncertainty minimization problem is equivalent to a reward maximization problem, which is further analyzed by a Markov decision process (MDP). To obtain the optimal task offloading decisions and flying orientation choices, a deep Q-network (DQN) based DRL architecture with a separated Q-network is then proposed. Finally, extensive simulations validate the effectiveness of the proposed techniques, and comprehensive discussions on how different parameters affect the search performance are given.
Quyuan Luo, Tom H. Luan, Weisong Shi, Pingzhi Fan
IEEE J. Sel. Areas Commun.3
2023 VCD-FL: Verifiable, Collusion-Resistant, and Dynamic Federated Learning
abstract
Federated learning (FL) is essentially a distributed machine learning paradigm that enables the joint training of a global model by aggregating gradients from participating clients without exchanging raw data. However, a malicious aggregation server may deliberately return designed results without any operation to save computation overhead, or even launch privacy inference attacks using crafted gradients. There are only a few schemes focusing on verifiable FL, and yet they cannot achieve collusion-resistant verification. In this paper, we propose the first Verifiable, Collusion-resistant, and Dynamic FL (VCD-FL) to tackle this issue. Specifically, we first optimize Lagrange interpolation by gradient grouping and compression for achieving efficient verifiability of FL. To protect clients’ data privacy against collusion attacks, we propose a lightweight commitment scheme using irreversible gradient transformation. By integrating the proposed efficient verification mechanism with the novel commitment scheme, our VCD-FL can detect whether or not the aggregation server is involved in collusion attacks. Moreover, considering that clients might go offline due to some reason such as network anomaly and client crash, we adopt the secret sharing technique to eliminate the effect of federation dynamics on FL. To the best of our knowledge, this is the first work to achieve collusion-resistant verification and collusion attack detection with supporting the correctness, privacy, and dynamics. Finally, we theoretically prove the effectiveness of our VCD-FL, make comprehensive comparisons, and conduct a series of experiments on MNIST dataset with MLP and CNN models. The theoretical proof and experimental analysis demonstrate that our VCD-FL is computationally efficient, robust against collusion attacks, and able to support the dynamics of FL.
Sheng Gao 0002, Jingjie Luo, Jianming Zhu 0002, Xuewen Dong, Weisong Shi
IEEE Trans. Inf. Forensics Secur.5
2023 WatchDog: Real-time Vehicle Tracking on Geo-distributed Edge Nodes
abstract
Vehicle tracking, a core application to smart city video analytics, is becoming more widely deployed than ever before thanks to the increasing number of traffic cameras and recent advances in computer vision and machine-learning. Due to the constraints of bandwidth, latency, and privacy concerns, tracking tasks are more preferable to run on edge devices sitting close to the cameras. However, edge devices are provisioned with a fixed amount of computing budget, making them incompetent to adapt to time-varying and imbalanced tracking workloads caused by traffic dynamics. In coping with this challenge, we propose WatchDog, a real-time vehicle tracking system that fully utilizes edge nodes across the road network. WatchDog leverages computer vision tasks with different resource-accuracy tradeoffs, and decomposes and schedules tracking tasks judiciously across edge devices based on the current workload to maximize the number of tasks while ensuring a provable response time-bound at each edge device. Extensive evaluations have been conducted using real-world city-wide vehicle trajectory datasets, achieving exceptional tracking performance with a real-time guarantee.
Zheng Dong 0002, Yan Lu 0006, Guangmo Tong, Yuanchao Shu, Shuai Wang 0008, Weisong Shi
ACM Trans. Internet Things6
2023 Reinforcement Learning for Adaptive Video Compressive Sensing
abstract
We apply reinforcement learning to video compressive sensing to adapt the compression ratio. Specifically, video snapshot compressive imaging (SCI), which captures high-speed video using a low-speed camera is considered in this work, in which multiple ( B ) video frames can be reconstructed from a snapshot measurement. One research gap in previous studies is how to adapt B in the video SCI system for different scenes. In this article, we fill this gap utilizing reinforcement learning (RL). An RL model, as well as various convolutional neural networks for reconstruction, are learned to achieve adaptive sensing of video SCI systems. Furthermore, the performance of an object detection network using directly the video SCI measurements without reconstruction is also used to perform RL-based adaptive video compressive sensing. Our proposed adaptive SCI method can thus be implemented in low cost and real time. Our work takes the technology one step further towards real applications of video SCI.
Sidi Lu, Xin Yuan 0002, Aggelos K. Katsaggelos, Weisong Shi
ACM Trans. Intell. Syst. Technol.4
2023 Fuel Rate Prediction for Heavy-Duty Trucks
abstract
Fuel cost contributes significantly to the high operation cost of heavy-duty trucks. Developing fuel rate prediction models is the cornerstone of fuel consumption optimization approaches for heavy-duty trucks. However, limited by accurate features directly related to the truck’s fuel consumption, state-of-the-art models show poor performance and are rarely deployed in practice. In this paper, we use the truck’s engine management system (EMS) and Instant Fuel Meter (IFM) to collect a three-month dataset during the period of December 2019 to June 2020. Seven prediction models, including linear regression, polynomial regression, MLP, CNN, LSTM, CNN-LSTM, and AutoML, are investigated and evaluated to predict real-time fuel rate. The evaluation results show that the EMS and IFM dataset help to improve the coefficient of determination of traditional linear/polynomial models from 0.87 to 0.96, while learning-based approach AutoML improves the coefficient of determination to attain 0.99. Besides, we explore the actual deployment of fuel rate prediction with transfer learning and path planning for autonomous driving.
Liangkai Liu, Wei Li 0111, Dawei Wang 0006, Ruigang Yang, Weisong Shi
IEEE Trans. Intell. Transp. Syst.6
2023 CEC: A Containerized Edge Computing Framework for Dynamic Resource Provisioning
abstract
Container has been widely used in application development and management systems. However, there are two major challenges faced in the real deployment at edge servers. The varying workload of service requests and the startup delay of containers force a flexible resource provisioning scheme in containerized edge computing. To this end, we proposeCEC, a containerized edge computing framework for dynamic resource provisioning, and especially for the smart connected community where exists multiple intelligent applications.CECintegrates workload prediction and resource pre-provisioning to enable low latency of user service requests and high utilization of edge resources. First, we present an online periodic request prediction algorithm. Then, we designed a control-based resource pre-provisioning algorithm based on the predicted request distribution, which is a self-adaptive controller to tune the resource for containers. We evaluate the performance ofCECby simulation and system experiments. The simulation experiment shows that the prediction accuracy of the proposed algorithm is higher than other two prediction algorithms. The testbed experiments demonstrate that the control-based resource pre-provisioning algorithm has low service latency and high resource utilization compared with baselines.
Shihong Hu, Weisong Shi, Guanghui Li 0001
IEEE Trans. Mob. Comput.2
2023 A Proactive On-Demand Content Placement Strategy in Edge Intelligent Gateways
abstract
Bandwidth-intensive applications transmit large-scale video data in the network. It causes backhaul bottlenecks and affects user experience. Deploying edge cache on an access point (AP) is a popular method to bring content files closer to end-users, but it faces significant challenges, especially in efficiently predicting and satisfying different users’ future content requests with limited cache capacity. In this article, we propose an intelligent gateway assisted edge cache deployment strategy (GACD), which jointly considers traffic usage patterns in multiple APs and the impact of new content on the cache performance. In GACD, The cache content placement problem is formulated as a many-to-one bidirectional matching problem with a dynamic quota allocation, aiming to improve cache resource utilization and minimize the average delivery latency. To address this problem, we design a heterogeneous information networks based prediction algorithm to predict end-users’ potential preference of new content files. Then, we adapt the seasonal autoregressive integrated moving average model for traffic usage prediction, and propose a many-to-one matching algorithm to achieve dynamic matching quota adjustment and efficient cache content placement. We conduct extensive real-world trace-based experiments to validate the performance of GACD. Compared with six alternative cache strategies, GACD improves the hit rate by 23.9% on average, reduces the average content delivery delay by 19.02%, and increases the accuracy by 31.02% on average.
Hui Sun 0002, Kewei Sha, Shaoyuan Huang, Xiaofei Wang 0001, Weisong Shi
IEEE Trans. Parallel Distributed Syst.6
2023 LARS: A Latency-Aware and Real-Time Scheduling Framework for Edge-Enabled Internet of Vehicles
abstract
With the development of Internet of Things and mobile computing, the explosive proliferation of latency-sensitive applications raises high computation demands for mobile devices. To this end, offloading computation of applications to edge-enabled Internet of Vehicles (IoV) has emerged as an effective solution. However, most of the existing studies on this issue assume that IoV can be easily formed in the practical environment, and neglect the dependency relationship between tasks of the offloading application. In this article, we first give several observations based on the analysis results of the real traffic dataset to verify the feasibility of aggregating vehicular resources in the real world. Then, we design a Latency-aware Real-time Scheduling Framework for the edge-enabled IoV, named LARS, in which mobile users can offload applications to LARS, and the offloading tasks can be scheduled to the appropriate vehicular resources in real-time. First, we propose a clustering-based algorithm to generateHerds, which treats connected vehicles as edge computation resources to provide cooperative computing services. Second, considering the dependency relationship between tasks in the job, we present a greedy-based task scheduling algorithm for offloading jobs, the objective of which is to minimize the total latency of the job as well as maximize the resource utilization ofHerds. The simulation experiment based on the real traffic dataset shows thatHerdsgenerated by the proposed clustering-based algorithm can maintain a stable period to provide computing service, and the experiments on testbed include two case studies demonstrate that the superiority of the proposed scheme compared to baselines, in terms of latency and resource utilization.
Shihong Hu, Guanghui Li 0001, Weisong Shi
IEEE Trans. Serv. Comput.3
2023 Resource Optimization of MAB-Based Reputation Management for Data Trading in Vehicular Edge Computing
abstract
Vehicles are hesitant to upload data to edge servers in vehicle edge computing (VEC) as many vehicle data collected and perceived by various on-board sensors contain sensitive and personal information and lack economic incentive. Instead of free access to shared data, encrypted data trading will alleviate security and privacy concerns and provide an incentive for vehicle owners to share their data. The edge server needs to pay the price in data trading, and reputation management is a great method to help it trade with reliable and available vehicles. In this paper, we propose a multi-armed bandit (MAB)-based reputation management scheme, so the edge servers can select the high reputation vehicles for data trading, which can ensure the credibility and reliability of the data. The encryption scheme is applied to achieve the required transmission security level and defend the rights and interests of the edge server. On the other hand, implementing security measures will consume the computation and communication resources of the vehicles. We formulate an optimization problem that maximizes the revenue of vehicles in data trading under the constraints of time delay, energy consumption, and security level. Simulation results demonstrate that the proposed scheme is effective and efficient for vehicle reputation management, data trading selection, and resource allocation.
Huizi Xiao, Lin Cai 0001, Jie Feng 0004, Qingqi Pei, Weisong Shi
IEEE Trans. Wirel. Commun.5
2023 Joint Optimization of Security Strength and Resource Allocation for Computation Offloading in Vehicular Edge Computing
abstract
Vehicular Edge Computing (VEC) is a promising new paradigm that has attracted much attention in recent years, which can enhance the storage and computing capabilities of vehicular networks to provide users with low latency and high-quality services. Due to the open access and unreliable wireless channels, some appropriate security measures should be implemented in the VEC to ensure information security. However, the operation of the security mechanism dominates supererogatory computing resources, thus affecting the performance of VEC systems. The scarcity of computation and energy resources of the vehicles conflicts with the requirement of tasks for time delay and information security. In this paper, taking the driving velocity and position of the vehicles, the number of lanes, the model and density of the attackers, and security strength into consideration, we formulate a max-min optimization problem to jointly optimize offloading decision, transmit power, task computation frequency, encryption computation frequency, edge computation frequency, and block length to obtain optimal secure information capacity and local computation delay. The formulated optimization problem is a mixed integer nonlinear programming (MINLP), which is intractable. We apply the generalized benders decomposition (GBD)-based method to solve it. The simulation results show that our proposed algorithms have convergence and effectiveness and achieve fairness among vehicles on the road.
Huizi Xiao, Jun Zhao 0007, Jie Feng 0004, Lei Liu 0031, Qingqi Pei, Weisong Shi
IEEE Trans. Wirel. Commun.6
2022 Speedster: An Efficient Multi-party State Channel via Enclaves
abstract
State channel network is the most popular layer-2 solution to the issues of scalability, high transaction fees, and low transaction throughput of public Blockchain networks. However, the existing works have limitations that curb the wide adoption of the technology, such as the expensive creation and closure of channels, strict synchronization between the main chain and off-chain channels, frozen deposits, and inability to execute multi-party smart contracts. In this work, we present Speedster, an account-based state-channel system that aims to address the above issues. To this end, Speedster leverages the latest development of secure hardware to create dispute-free certified channels that can be operated efficiently off the Blockchain. Speedster is peer-to-peer decentralized and provides better privacy protection than prior channel projects. It supports fast native multi-party contract execution, which is previously unavailable in TEE-enabled channel networks. Compared to the Lightning Network, Speedster improves the throughput by about 10,000X and generates 97%$ less on-chain data with a comparable network scale.
Jinghui Liao, Fengwei Zhang, Wenhai Sun, Weisong Shi
AsiaCCS4
2022 BlueScale: a scalable memory architecture for predictable real-time computing on highly integrated SoCs
abstract
In real-time embedded computing, time-predictability and performance are required simultaneously by memory transactions. However, with increasingly more elements being integrated into hardware, memory interconnects become a critical stumbling block to satisfying timing correctness, due to lack of hardware and scheduling scalability. In this paper, we propose a new hierarchically distributed memory interconnect, BlueScale, managing memory transactions using identical Scale Elements, which ensures hardware scalability. The Scale Element introduces two nested priority queues, achieving iterative compositional scheduling for memory transactions, guaranteeing transaction tasks' scheduling schedulability. Associated with the new architecture, a theoretical model is established to improve BlueScale's real-time performance.
Zhe Jiang 0004, Kecheng Yang 0001, Neil C. Audsley, Nathan Fisher, Weisong Shi, Zheng Dong 0002
DAC5
2022 To Turn or Not To Turn, SafeCross is the Answer
abstract
Blind area has plagued drivers’ safety ever since the dawn of automobiles. Thanks to the fast-growing vision-based perception technologies, autonomous driving systems can monitor the driving circumstance through a 360-degree view, and hence most blind areas can be avoided. However, in the left turn scenario at an intersection, the opposite road may be blocked by another vehicle parking at the same intersection (see Fig. 1), and in this case, the blind area cannot be observed by the onboard perception module of the autonomous vehicle. A potential fatal collision may occur if the autonomous vehicle turns left while a vehicle is running through the blind area. In this paper, we propose Safecross, a framework that oversees an intersection and delivers blind area warnings to the left-turn vehicles at the intersection if running vehicles are detected in the blind area. In order to provide accurate and reliable real-time warnings in all possible weather conditions, the architecture of Safecross has four major components: video pre-processing (VP) module, video classification (VC) module, few-shot learning (FL) module, and model switching (MS) module. Especially, the VP and VC modules will train a basic model to identify the blind area when a blocking vehicle appears at the intersection. Since the range of the blind area varies in different weather conditions, the FL and MS modules can adapt the basic model to the new condition in real-time to make the blind area identification more accurate. Intuitively, if the blind area is identified timely and accurately, the left-turn throughput of the intersection can be maximized. We have conducted extensive experiments to evaluate our proposed framework. The experiments are performed on a total of 2855 video segments with a time span of 180 days, including sunny, rainy, and snowy weather conditions. Experimental results show how Safecross can guarantee the vehicle’s safety while increasing the left-turn traffic throughput by 50%.
Baofu Wu, Yuankai He, Zheng Dong 0002, Jian Wan 0001, Weisong Shi
ICDCS6
2022 Poster: Autonomous mobile robot for indoor enhanced living
abstract
The world population is aging at a rapid pace. It is projected that by 2050, the total number of older individuals (65 and older) will more than double their current number to rise to 1.57 billion [1]. Further, there is a growing shortage of caregivers and registered nurses to fill the emotional, and physical needs of the elderly population [2]. As a result, the world is projected to face the challenge of providing a good quality of life for its older population. Several nationwide projects are in exploratory stages to integrate technology into the long-term care regime [3]. Some of the notable ones are Smart-BEAR1, and PHArA-ON2, and SHINESeniors3. The underlying goal is to care for the elderly remotely so that elderly can age independently in their own homes while ensuring that they receive prompt help when needed.
Weisong Shi
SEC2
2022 Poster: Towards Efficient Multilayer Collaboration for CAV Applications
abstract
Connected and autonomous vehicles (CAV s) are facing increasing amounts of data and more complex data analysis, which creates challenges for them to make reliable decisions in real-time. To enable time-sensitive CAV applications, we design and implement a vehicle-edge-cloud framework that integrates compressed imaging (CI) and edge computing into CAV systems. Specifically, a lightweight model is used on the vehicle to perform real-time detection based on optical domain compressed data (called measurements). The edge is responsible for receiving the measurements and performing video reconstruction to support (more accurate) analysis based on the reconstructed video with a trigger. At the same time, the measurements, reconstructed videos, and analysis results are sent to the cloud to continuously update the vehicle model. In addition, we apply reinforcement learning to adapt the compression rate in different driving scenarios. The proposed framework is fully evaluated using our designed roadside platform and outdoor delivery vehicles.
Sidi Lu, Weisong Shi
SEC2
2022 Towards Edge-enabled Distributed Computing Framework for Heterogeneous Android-based Devices
abstract
In this paper, we propose an Android-based distributed computing framework for accelerating DNN inference on Android edge devices. We experimentally demonstrate that the proposed distributed framework can reduce CPU utilization by 24 % (making the the CPU utilization close to that of idle status), reduce power consumption by 59.8 % to 71.8 %, without leading to high-bandwidth througput. The proposed framework can be applied to various Android devices to enable cooperation among edge devices in a distributed computing manner, accelerate DNN inference, and enrich the functionality of Android devices to enhance user experience.
Yongtao Yao, Weisong Shi
SEC4
2022 Prophet: Realizing a Predictable Real-time Perception Pipeline for Autonomous Vehicles
abstract
We have witnessed the broad adoption of Deep Neu-ral Networks (DNNs) in autonomous vehicles (AV). As a safety-critical system, deadline-based scheduling is used to guarantee the predictability of the AV system. However, non-negligible time variations exist for most DNN models in an AV system, even when the whole system is just running one model. The fact that multiple DNNs are running on the same platform makes the time variations issue even more severe. However, none of the existing works have thoroughly studied the root cause of the time variation issue. In the first part of the paper, we conducted a comprehensive empirical study. We found that the inference time variations for a single DNN model are mainly caused by the DNN's multi-stage/multi-branch structure, which has a dynamic number of proposals or raw points. In addition, we found that the uncoordinated contention and cooperation are the roots of the time variations for multi-tenant DNNs inference. Second, based on these insights, we proposed the Prophet system that addresses the time variations in the AV perception system in two steps. The first step is to predict the time variations based on the intermediate results like proposals and raw points. The second step is coordinating the multi-tenant DNNs to ensure the execution progress is close to each other. From the evaluation results on the KITTI dataset, the time prediction of a single model all achieve higher than 91% accuracy for Faster R-CNN, LaneNet, and PINet. Besides, the perception fusion delay is bounded to 150ms, and the fusion drop ratio is reduced from 5.4% to less than 1 percent.
Liangkai Liu, Zheng Dong 0002, Yanzhi Wang 0001, Weisong Shi
RTSS4
2022 A Cross-layer Plausibly Deniable Encryption System for Mobile Devices
Niusen Chen, Bo Chen 0028, Weisong Shi
SecureComm3
2022 SC-UDA: Style and Content Gaps aware Unsupervised Domain Adaptation for Object Detection
abstract
Current state-of-the-art object detectors can have significant performance drop when deployed in the wild due to domain gaps with training data. Unsupervised Domain Adaptation (UDA) is a promising approach to adapt detectors for new domains/environments without any expensive label cost. Previous mainstream UDA works for object detection usually focused on image-level and/or feature-level adaptation by using adversarial learning methods. In this work, we show that such adversarial-based methods can only reduce domain style gap, but cannot address the domain content gap that is also important for object detectors. To overcome this limitation, we propose the SC-UDA framework to concurrently reduce both gaps: We propose fine-grained domain style transfer to reduce the style gaps with finer image details preserved for detecting small objects; Then we leverage the pseudo label-based self-training to reduce content gaps; To address pseudo label error accumulation during self-training, novel optimizations are proposed, including uncertainty-based pseudo labeling and imbalanced mini-batch sampling strategy. Experiment results show that our approach consistently outperforms prior state-of-the-art methods (up to 8.6%, 2.7% and 2.5% mAP on three UDA benchmarks).
Fuxun Yu, Di Wang 0003, Yinpeng Chen, Nikolaos Karianakis, Pei Yu, Dimitrios Lymberopoulos, Sidi Lu, Weisong Shi, Xiang Chen 0010
WACV9
2022 EdgeWare: toward extensible and flexible middleware for connected vehicle services
Sidi Lu, Yongtao Yao, Zhifeng Yu, Weisong Shi
CCF Trans. High Perform. Comput.6
2022 Autonomous Driving Security: State of the Art and Challenges
abstract
The autonomous driving industry has mushroomed over the past decade. Although autonomous driving has undoubtedly become one of the most promising technologies of this century, its development faces multiple challenges, of which security is the major concern. In this article, we present a thorough analysis of autonomous driving security. First, the attack surface of autonomous driving is presented. After an analysis of the operation of autonomous driving in terms of key components and technologies, the security of autonomous driving is elaborated in four dimensions: 1) sensors; 2) operating system; 3) control system; and 4) vehicle-to-everything (V2X) communication. Sensor security is examined from five components, which are mainly responsible for self-positioning and environmental perception. The analysis of operating system security, the second dimension, is concentrated on the robot operating system. Concerning the control system security, the controller area network is approached mainly from vulnerabilities and protection measures. The fourth dimension, V2X communication security, is probed from four categories of attacks: 1) authenticity/identification; 2) availability; 3) data integrity; and 4) confidentiality with corresponding solutions. Moreover, the drawbacks of existing methods adopted in the four dimensions are also provided. Finally, a conceptual multilayer defense framework is proposed to secure the information flow from external communication to the physical autonomous vehicle.
Cong Gao 0002, Weisong Shi, Zhongmin Wang 0001, Yanping Chen 0006
IEEE Internet Things J.3
2022 Blockchain-Enabled Efficient Dynamic Cross-Domain Deduplication in Edge Computing
abstract
As the rapid proliferation of Internet of Things (IoT) and edge computing, large amounts of data are needed to be stored and transmitted in the online storage system. Data deduplication can be adopted to improve communication efficiency and minimize storage space. However, in edge computing, data deduplication brings security and functionality requirements that are still unsatisfied. Most existing schemes are vulnerable to brute-force attacks and single-point attacks. Moreover, they impose a heavy burden on resource-constrained edge nodes and do not support cross-domain deduplication. Blockchain is a promising technology because the programmable smart contract can be utilized to perform cross-domain deduplication and guarantee the traceability of data. In this article, an efficient dynamic cross-domain deduplication scheme in blockchain-enabled edge computing is proposed to solve the above problems. Specifically, the smart contract is employed to assist cross-domain deduplication, which also can reduce the storage pressure of edge nodes. Meanwhile, a hash proof system-based oblivious pseudorandom function is created to reduce the time cost of key generation and achieve the security requirements of resistance to brute-force attacks and single-point attacks. The technology of accumulators is adopted to achieve Proofs of Ownership (PoO), which can prevent duplicate-faking attacks. The security analysis demonstrates that the proposed scheme has a higher security level. The performance evaluation shows that the proposed scheme significantly reduces computation cost and communication overhead, compared with other existing schemes. The smart contract is implemented in the Ethereum test network (i.e., Rinkeby), which shows acceptable gas cost even the functions are called frequently.
Yang Ming 0001, Chenhao Wang 0005, Hang Liu 0008, Yi Zhao 0011, Jie Feng 0004, Ning Zhang 0007, Weisong Shi
IEEE Internet Things J.7
2022 FlexEdge: Dynamic Task Scheduling for a UAV-Based On-Demand Mobile Edge Server
abstract
With the large number of cameras deployed in smart industrial parks and smart campuses, edge devices and location-fixed edge servers are deployed near to these cameras and help transmit video streams to data center for video analytics; however, location-fixed edge servers are difficult to adapt to computation-intensive and delay-sensitive video analytics tasks in hot spot, and it is also challenging to execute tasks in natural disasters in which the infrastructure is damaged. Moreover, task migration methods are used to balance the load of edge servers caused by irregular movement of detected objects, but it results in extra data transmission overhead. Therefore, unmanned aerial vehicles (UAVs) with computing and communication resources are widely used to optimize mobile edge video analysis; however, existing solutions formulate the UAV-based lowest latency and energy consumption by jointly optimizing the task allocation strategy and UAV location to be a multiobjective optimization problem, based on which the Pareto optimum solution set, including task allocation strategies and UAV locations, can find multiple solutions but not a unique solution. It makes the solution difficult to be applied in video analytics with the UAV hover location decision-making scheme and task allocation strategy. In this article, we propose a flexible cloud-edge collaborative scheduling strategy based on a UAV namedFlexEdge. We first normalize values of execution time and energy consumption, and then convert the multiobjective optimization problem into a single-objective optimization problem by using the weighted sum of the two metrics as the optimization objective. We also proved the task allocation strategy based on execution time, energy consumption, and the UAV hover location decision-making scheme as an NP-hard problem. We propose a flexible and lightweight genetic algorithm (FGA) based on a polysomy-strengthening elitist genetic algorithm in FlexEdge to address the NP-hard problem. FlexEdge not only achieves optimal task allocation and UAV location to minimize the weighted sum of execution time and energy consumption but also provides computing resources and reliable network connection to reduce task offloading overload, which is validated by comprehensive performance evaluation.
Hui Sun 0002, Bo Zhang 0111, Xiuye Zhang, Kewei Sha, Weisong Shi
IEEE Internet Things J.6
2022 Authentication Security Level and Resource Optimization of Computation Offloading in Edge Computing Systems
abstract
Edge computing brings computation and storage resources to the edge of the mobile network to meet strict delay and high demanding applications. However, edge network environments are more vulnerable to malicious attacks. Reliable communication in networks usually relies heavily on verifying the authentication of content and identity. Nevertheless, enhancing the authentication security level means occupying more computation resources, time, and energy. There is a tradeoff between security level improvement and resource optimization in computation offloading. In this article, we take maximizing the authentication security level of the Merkle tree signature as segmental of the optimization objective and consider the different hash algorithms deployed on the edge servers to make the offloading decision. Specifically, to weight the time delay and authentication security level simultaneously, we formulate a minimum optimization problem to jointly optimize the offloading decision, packet transmitting rate, edge computation frequency, and data blocks number. Simulation results show that our proposed algorithms have well convergence and effectiveness and provide a tradeoff between time delay and authentication security level.
Huizi Xiao, Qingqi Pei, Xifei Song, Weisong Shi
IEEE Internet Things J.4
2022 Characterizing Co-Located Workloads in Alibaba Cloud Datacenters
abstract
Workload characteristics are vital for both data center operation and job scheduling in co-located data centers, where online services and batch jobs are deployed on the same production cluster. In this article, a comprehensive analysis is conducted on Alibaba's cluster-trace-v2018 of a production cluster of 4034 machines. The findings and insights are the following: (1) The workload on the production cluster poses a daily cyclical fluctuation, in terms of CPU and disk I/O utilization, and the memory system has become the performance bottleneck of a co-located cluster. (2) Batch jobs including their tasks and derived instances can be approximated as Zipf distribution. However, for all batch jobs with directed acyclic graph dependency, they suffer from co-location with online services since the online services are highly prioritized. (3) The resource usages of containers have similar cyclical fluctuation consistent with the whole cluster, while their memory usages remain approximately constant. (4) The number of batch jobs co-located with online services is dependent on the mispredictions per kilo instructions of online services. In order to guarantee the QoS of online services, when the MPKI of online services rises, the number of batch jobs to be co-located on the same machine should decrease.
Congfeng Jiang, Yitao Qiu, Weisong Shi, Zhefeng Ge, Shenglei Chen, Christophe Cérin, Zujie Ren, Guoyao Xu, Jiangbin Lin
IEEE Trans. Cloud Comput.3
2022 Vehicle Selection and Resource Optimization for Federated Learning in Vehicular Edge Computing
abstract
As a distributed deep learning paradigm, federated learning (FL) provides a powerful tool for the accurate and efficient processing of on-board data in vehicular edge computing (VEC). However, FL involves the training and transmission of model parameters, which consumes the vehicles’ precious energy resources and takes up much time. It is a departure from many applications with severe real-time requirements in VEC. And the capabilities and data quality of each vehicle are distinct that will affect the performance of training the model. Therefore, it is crucial to select the appropriate vehicles to participate in learning tasks and optimize resource allocation under learning time and energy consumption constraints. In this paper, taking the vehicle position and velocity into consideration, we formulate a min-max optimization problem to jointly optimize the on-board computation capability, transmission power, and local model accuracy to achieve the minimum cost in the worst case of FL. Specifically, we propose a greedy algorithm to select vehicles with higher image quality dynamically, and it keeps the system’s overall cost to a minimum in FL. The formulated optimization problem is a nonlinear programming problem, so we decompose it into two subproblems. For the resource allocation problem, we use the Lagrangian dual problem and the subgradient projection method to approximate the optimal value iteratively. For the local model accuracy problem, we develop an adaptive harmony algorithm for heuristic search. The simulation results show that our proposed algorithms have well convergence and effectiveness and achieve a tradeoff between cost and fairness.
Huizi Xiao, Jun Zhao 0007, Qingqi Pei, Jie Feng 0004, Lei Liu 0031, Weisong Shi
IEEE Trans. Intell. Transp. Syst.6
2022 Serving at the Edge: An Edge Computing Service Architecture Based on ICN
abstract
Different from cloud computing, edge computing moves computing away from the centralized data center and closer to the end-user. Therefore, with the large-scale deployment of edge services, it becomes a new challenge of how to dynamically select the appropriate edge server for computing requesters based on the edge server and network status. In the TCP/IP architecture, edge computing applications rely on centralized proxy servers to select an appropriate edge server, which leads to additional network overhead and increases service response latency. Due to its powerful forwarding plane, Information-Centric Networking (ICN) has the potential to provide more efficient networking support for edge computing than TCP/IP. However, traditional ICN only addresses named data and cannot well support the handle of dynamic content. In this article, we propose an edge computing service architecture based on ICN, which contains the edge computing service session model, service request forwarding strategies, and service dynamic deployment mechanism. The proposed service session model can not only keep the overhead low but also push the results to the computing requester immediately once the computing is completed. However, the service request forwarding strategies can forward computing requests to an appropriate edge server in a distributed manner. Compared with the TCP/IP-based proxy solution, our forwarding strategy can avoid unnecessary network transmissions, thereby reducing the service completion time. Moreover, the service dynamic deployment mechanism decides whether to deploy an edge service on an edge server based on service popularity, so that edge services can be dynamically deployed to hotspot, further reducing the service completion time.
Zhenyu Fan, Wang Yang 0002, Fan Wu 0014, Weisong Shi
ACM Trans. Internet Techn.5
2022 NLUBroker: A QoE-driven Broker System for Natural Language Understanding Services
abstract
Cloud-based Natural Language Understanding (NLU) services are becoming more popular with the development of artificial intelligence. More applications are integrated with cloud-based NLU services to enhance the way people communicate with machines. However, with NLU services provided by different companies powered by unrevealed AI technology, how to choose the best one is a problem for developers. Existing tools that can provide guidance to developers and make recommendations based on their needs are severely limited. This article comprehensively evaluates multiple state-of-the-art NLU services, and the results indicate that there is no absolute winner for different usage requirements. Motivated by this observation, we provide several insights and propose NLUBroker , a Quality of Experience-driven (QoE-driven) broker system, to select the proper service according to the environment. NLUBroker senses the client and service status and leverages a solution to the multi-armed bandit problem to conduct online learning, aiming to achieve maximum expected QoE. The performance of NLUBroker is evaluated in both simulation and real-world environments, and the evaluation results demonstrate that NLUBroker is an efficient solution for selecting NLU services. It is adaptive to changes in the environment, outperforms three baseline methods we evaluated and improves overall QoE up to 1.5× for the evaluated state-of-the-art NLU services.
Lanyu Xu, Arun Iyengar, Weisong Shi
ACM Trans. Internet Techn.3
2022 Minimizing the Delay and Cost of Computation Offloading for Vehicular Edge Computing
abstract
The development of autonomous driving poses significant demands on computing resource, which is challenging to resource-constrained vehicles. To alleviate the issue, Vehicular edge computing (VEC) has been developed to offload real-time computation tasks from vehicles. However, with multiple vehicles contending for the communication and computation resources at the same time for different applications, how to efficiently schedule the edge resources toward maximal system welfare represents a fundamental issue in VEC. This article aims to provide a detailed analysis on the delay and cost of computation offloading for VEC and minimize the delay and cost from the perspective of multi-objective optimization. Specifically, we first establish an offloading framework with communication and computation for VEC, where computation tasks with different requirements for computation capability are considered. To pursue a comprehensive performance improvement during computation offloading, we then formulate a multi-objective optimization problem to minimize both the delay and cost by jointly considering the offloading decision, allocation of communication and computation resources. By applying the game theoretic analysis, we propose a particle swarm optimization based computation offloading (PSOCO) algorithm to obtain the Pareto-optimal solutions to the multi-objective optimization problem. Extensive simulation results verify that our proposed PSOCO outperforms counterparts. Based on the results, we also present a comprehensive analysis and discussion on the relationship between delay and cost among the Pareto-optimal solutions.
Quyuan Luo, Changle Li, Tom H. Luan, Weisong Shi
IEEE Trans. Serv. Comput.4
2021 The Case for Adaptive Deep Neural Networks in Edge Computing
abstract
Deep Neural Networks (DNNs) are an application class that benefit from being distributed across the edge and cloud. A DNN is partitioned such that specific layers of the DNN are deployed onto the edge and the cloud to meet performance and privacy objectives. However, there is limited understanding of: whether and how evolving operational conditions (increased CPU and memory utilization at the edge or reduced data transfer rates between the edge and cloud) affect the performance of already deployed DNNs, and whether a new partition configuration is required to maximize performance. A DNN that adapts to changing operational conditions is referred to as an ‘adaptive DNN’. This paper investigates whether there is a case for adaptive DNNs by considering four questions: (i) Are DNNs sensitive to operational conditions? (ii) How sensitive are DNNs to operational conditions? (iii) Do individual or a combination of operational conditions equally affect DNNs? (iv) Is DNN partitioning sensitive to hardware architectures? The exploration is carried out in the context of 8 pre-trained DNN models and the results presented are from analyzing nearly 8 million data points. The results highlight that network conditions affect DNN performance more than CPU or memory related operational conditions. Repartitioning is noted to provide a performance gain in a number of cases, but a specific trend is not noted in relation to the underlying hardware architecture. Nonetheless, the need for adaptive DNNs is confirmed.
Francis McNamee, Schahram Dustdar, Peter Kilpatrick, Weisong Shi, Ivor T. A. Spence, Blesson Varghese
CLOUD4
2021 ChatCache: A Hierarchical Semantic Redundancy Cache System for Conversational Services at Edge
abstract
The spatial-temporal locality has been observed in various scenarios for conversational services with either voice or text requests. Given the current cloud-based processing mechanism, integrating such a service with caching is a promising way to improve responsiveness, reduce in-network transmission, and avoid computational redundancy. Goes beyond precise redundancy and fuzzy redundancy, semantic redundancy adapts to the diversity in command expression, and is considered as a practical solution for conversational services. In this paper, we introduce a hierarchical cache design inspired by semantic redundancy for conversational services. We propose a scalable edge system ChatCache to incorporate the hierarchical cache design and serve single or multiple users. We discussed the cache efficiency with different similarity match policies, and evaluate the responsiveness and scalability of ChatCache on heterogeneous edge platforms. On most of the evaluated platforms, ChatCache reduces user-perceived latency by more than 91.7% for voice requests, more than 81.6% for text requests. The throughput of ChatCache reaches 42.6 throughput tps for voice requests, and 64.4 tps for text requests, which is comparable with mainstream cloud cognitive services. The promising evaluation results show the capability of ChatCache in reducing the user-perceived latency and computation redundancy with high response accuracy for conversational services.
Lanyu Xu, Arun Iyengar, Weisong Shi
CLOUD3
2021 Privacy-Preserving Neural Network Inference Framework via Homomorphic Encryption and SGX
abstract
Edge computing is a promising paradigm that pushes computing, storage, and energy to the networks' edge. It utilizes the data nearby the users to provide real-time, energy-efficient, and reliable services. Neural network inference in edge computing is a powerful tool for various applications. However, edge server will collect more personal sensitive information of users inevitably. It is the most basic requirement for users to ensure their security and privacy while obtaining accurate inference results. Homomorphic encryption (HE) technology is confidential computing that directly performs mathematical computing on encrypted data. But it only can carry out limited addition and multiplication operation with very low efficiency. Intel software guard extension (SGX) can provide a trusted isolation space in the CPU to ensure the confidentiality and integrity of code and data executed. But several defects are hard to overcome due to hardware design limitations when applying SGX in inference services. This paper proposes a hybrid framework utilizing SGX to accelerate the HE-based convolutional neural network (CNN) inference, eliminating the approximation operations in HE to improve inference accuracy in theory. Besides, SGX is also taken as a built-in trusted third party to distribute keys, thereby improving our framework's scalability and flexibility. We have quantified the various CNN operations in the respective cases of HE and SGX to provide the foresight practice. Taking the connected and autonomous vehicles as a case study in edge computing, we implemented this hybrid framework in CNN to verify its feasibility and advantage.
Huizi Xiao, Qingyang Zhang 0001, Qingqi Pei, Weisong Shi
ICDCS4
2021 TrustZone Enhanced Plausibly Deniable Encryption System for Mobile Devices
Jinghui Liao, Bo Chen 0028, Weisong Shi
SEC3
2021 A trusted and collaborative framework for deep learning in IoT
Qingyang Zhang 0001, Hong Zhong 0001, Weisong Shi, Lu Liu 0001
Comput. Networks3
2021 CCPrune: Collaborative channel pruning for learning compact convolutional networks
Yanming Chen 0002, Yiwen Zhang 0001, Weisong Shi
Neurocomputing4
2021 Computing Systems for Autonomous Driving: State of the Art and Challenges
abstract
The recent proliferation of computing technologies (e.g., sensors, computer vision, machine learning, and hardware acceleration) and the broad deployment of communication mechanisms (e.g., dedicated short-range communication, cellular vehicle-to-everything, 5G) have pushed the horizon of autonomous driving, which automates the decision and control of vehicles by leveraging the perception results based on multiple sensors. The key to the success of these autonomous systems is making a reliable decision in real-time fashion. However, accidents and fatalities caused by early deployed autonomous vehicles arise from time to time. The real traffic environment is too complicated for current autonomous driving computing systems to understand and handle. In this article, we present state-of-the-art computing systems for autonomous driving, including seven performance metrics and nine key technologies, followed by 12 challenges to realize autonomous driving. We hope this article will gain attention from both the computing and automotive communities and inspire more research in this direction.
Liangkai Liu, Sidi Lu, Ren Zhong, Baofu Wu, Yongtao Yao, Qingyang Zhang 0001, Weisong Shi
IEEE Internet Things J.7
2021 CLONE: Collaborative Learning on the Edges
abstract
The proliferation of edge computing technologies has boosted the development of new applications for a plethora of edge devices. However, many applications face privacy issues and bandwidth limitations. To solve these limitations, we propose a collaborative learning framework on the edges, named CLONE, which is steered by the real-world data sets collected from a large electric vehicle (EV) company and a grocery store of a shopping mall, respectively. We categorize two application scenarios for CLONE, i.e., CLONE in the training stage (CLONE_training) and CLONE in the inference stage (CLONE_inference). As to CLONE_training, we choose the failure prediction of EV battery and associated components as the first use case. While as for CLONE_inference, customer tracking in a grocery store is selected as another case study. In this work, the goal of the CLONE is to support real-time training and inference for connected vehicles and marketing intelligence services. Our experimental results on the EV data show that CLONE is able to reduce model training time without sacrificing algorithm performance. Furthermore, the experimental results on the video data from the grocery store reveal that CLONE is a useful approach to solve the multitarget multicamera tracking problem in a collaborative fashion.
Sidi Lu, Yongtao Yao, Weisong Shi
IEEE Internet Things J.3
2021 AC4AV: A Flexible and Dynamic Access Control Framework for Connected and Autonomous Vehicles
abstract
Sensing data plays a pivotal role in connected and autonomous vehicles (CAVs), enabling CAV to perceive surroundings. For example, malicious applications might tamper this life-critical data, resulting in erroneous driving decisions and threatening the safety of passengers. Access control, one of the promising solutions to protect data from unauthorized access, is urgently needed for vehicle sensing data. However, due to the intrinsic complexity of vehicle sensing data, including historical and real time, and access patterns of different data sources, there is currently no suitable access control framework that can systematically solve this problem; current frameworks only focus on one aspect. In this article, we propose a novel and flexible access control framework,AC4AV, which aims to support various access control models, and provide APIs for dynamically adjusting access control models and developing customized access control models, thus supporting access control research on CAV for the community. In addition, we propose a data abstraction method to clearly identify data, applications, and access operations in CAV, and therefore is easily able to configure the permits of each data and application in access control policies. We have implemented a prototype to demonstrate our architecture on NATS for real-time data and NGINX for historical data, and three access control models as built-in models. We measured the performance of ourAC4AVwhile applying these access control models to real-time and historical data. The experimental results show that the framework has little impact on real-time data access within a tolerable range.
Qingyang Zhang 0001, Hong Zhong 0001, Jie Cui 0004, Lingmei Ren, Weisong Shi
IEEE Internet Things J.5
2021 Study of interconnect errors, network congestion, and applications characteristics for throttle prediction on a large scale HPC system
Saurabh Gupta 0002, Tirthak Patel, Michael Wilder, Weisong Shi, Song Fu, Christian Engelmann, Devesh Tiwari
J. Parallel Distributed Comput.5
2021 Self-Learning Based Computation Offloading for Internet of Vehicles: Model and Algorithm
abstract
With the fast development of Internet of Vehicles (IoV), various types of computation-intensive vehicular applications pose significant challenges to resource-constrained vehicles. The emerging Vehicular Edge Computing (VEC) and Edge Intelligence (EI) can alleviate this situation by offloading the computation tasks of vehicles to the roadside edge servers. However, with many vehicles contending for the communication and computation resources at the same time, how to quickly and efficiently make an optimal computation offloading decision for individual vehicles represents a fundamental research issue. In this paper, we propose a self-learning based distributed computation offloading scheme for IoV. Note that without any centralized controller, a fully distributed algorithm is necessary. The proposed scheme is devised based on a game-theoretic model. Specifically, through establishing an offloading framework with communication and computation for IoV, the computation offloading problem is first formulated as a distributed offloading decision-making game, in which each vehicle as a player makes its best response decision to minimize its joint cost (including latency and offloading cost). The existence of Nash Equilibrium can be proved. We then propose a self-learning based distributed computation offloading (DISCO) algorithm to reach the Nash Equilibrium, where a mutually satisfactory solution among vehicles is obtained and no vehicle is willing to change its decision. Using extensive simulations, we verify that DISCO can outperform the counterparts and achieve at least an order-of-magnitude improvement on time overhead and 88% performance gain on message overhead, only at up to 12% performance loss on joint cost over the centralized scheme.
Quyuan Luo, Changle Li, Tom H. Luan, Weisong Shi, Weigang Wu
IEEE Trans. Wirel. Commun.4
2020 Making Disk Failure Predictions SMARTer!
Sidi Lu, Tirthak Patel, Yongtao Yao, Devesh Tiwari, Weisong Shi
FAST6
2020 Edge Compression: An Integrated Framework for Compressive Imaging Processing on CAVs
abstract
Machine vision is the key to the successful deployment of many Advanced Driver Assistant System (ADAS) / Automated Driving System (ADS) functions, which require accurate high-resolution video processing in a real-time manner. Conventional approaches are either to reduce the frame rate or reduce the related frame size of the conventional camera videos, which lead to undesired consequences such as losing informative high-speed information and/or small objects in the video frames.Unlike conventional cameras, Compressive Imaging (CI) cameras are the promising implications of Compressive Sensing, which is an emerging field with the revelation that the optical domain compressed signal (a small number of linear projections of the original video image data) contains sufficient high-speed information for reconstruction and processing. Yet, CI cameras usually need complicated algorithms to retrieve the desired signal, leading to the corresponding high energy consumption. In this paper, we take a step further to the real applications of CI cameras in connected and autonomous vehicles (CAVs), with the primary goal of accelerating accurate video analysis and decreasing energy consumption. We propose a novel Vehicle Edge Server-Cloud closed-loop framework called Edge Compression for CI processing on CAVs. Our comprehensive experiments with four public datasets demonstrate that the detection accuracy of the compressed video images (named measurements) generated by the CI camera is close to the accuracy on reconstructed videos and comparable to the true value, which paves the way of applying CI in CAVs. Finally, six important observations with supporting evidence and analysis are presented to provide practical implications for researchers and domain experts. The code to reproduce our results is available at https://www.thecarlab.oryoutcomes/software.
Sidi Lu, Xin Yuan 0002, Weisong Shi
SEC3
2020 A Cross-Layer Optimization Framework for Distributed Computing in IoT Networks
abstract
In Internet-of-Thing (IoT) networks, enormous low-power IoT devices execute latency-sensitive yet computation intensive machine learning tasks. However, the energy is usually scarce for IoT devices, especially for some without battery and relying on solar power or other renewables forms. In this paper, we introduce a cross-layer optimization framework for distributed computing among low-power IoT devices. Specifically, a programming layer design for distributed IoT networks is presented by addressing the problems of application partition, task scheduling, and communication overhead mitigation. Furthermore, the associated federated learning and local differential privacy schemes are developed in the communication layer to enable distributed machine learning with privacy preservation. In addition, we illustrate a three-dimensional network architecture with various network components to facilitate efficient and reliable information exchange among IoT devices. Moreover, a model quantization design for IoT devices is illustrated to reduce the cost of information exchange. Finally, a parallel and scalable neuromorphic computing system for IoT devices is established to achieve energy-efficient distributed computing platforms in the hardware layer. Based on the introduced cross-layer optimization framework, IoT devices can execute their machine learning tasks in an energy-efficient way while guaranteeing data privacy and reducing communication costs.
Bodong Shang, Shiya Liu, Sidi Lu, Yang Yi 0002, Weisong Shi, Lingjia Liu 0001
SEC5
2020 EdgeMask: An Edge-based Privacy Preserving Service for Video Data Sharing
abstract
Preserving privacy in image and video data captured from public environments is essential for any research group that leverages, publishes, or shares such data. Although there are several research efforts attempting to resolve the privacy issues, they had quality and efficiency limitations. In this work, we proposed EdgeMask as a privacy preserving service that leverages edge computing and deep learning models to propose a real-time object segmentation approach and analyze the input data using parallel computing and speed up the object removal. Our experimental results indicate that EdgeMask reduces the computational time considerably.
Samira Taghavi, Weisong Shi
SEC2
2020 CHA: A Caching Framework for Home-based Voice Assistant Systems
abstract
Voice assistant systems are becoming immersive in our daily lives nowadays. However, current voice assistant systems rely on the cloud for command understanding and fulfillment, resulting in unstable performance and unnecessary frequent network transmission. In this paper, we introduce CHA, an edge-based caching framework for voice assistant systems, and especially for smart homes where resource-restricted edge devices can be deployed. Located between the voice assistant device and the cloud, CHA introduces a layered architecture with modular design in each layer. By introducing an understanding module and adaptive learning, CHA understands the user's intent with high accuracy. By maintaining a cache, CHA reduces the interaction with the cloud and provides fast and stable responses in a smart home. Targeting on resource-constrained edge devices, CHA uses joint classification and model pruning on a pre-trained language model to achieve performance and system efficiency. We compare CHA to the status quo solution of voice assistant systems and show that CHA benefits voice assistant systems. We evaluate CHA on three edge devices that differ in hardware configuration and demonstrate its ability to meet the latency and accuracy demands with efficient resource utilization. Our evaluation shows that compared to the current solution for voice assistant systems, CHA can provide at least 70% speedup in responses for frequently asked voice commands with less than 13% CPU consumption, and less than 9% memory consumption when running on a Raspberry Pi.
Lanyu Xu, Arun Iyengar, Weisong Shi
SEC3
2020 A deep neural network compression algorithm based on knowledge transfer for edge devices
Yanming Chen 0002, Chao Li 0028, Luqi Gong, Yiwen Zhang 0001, Weisong Shi
Comput. Commun.6
2020 Energy aware edge computing: A survey
Congfeng Jiang, Tiantian Fan, Honghao Gao, Weisong Shi, Liangkai Liu, Christophe Cérin, Jian Wan 0001
Comput. Commun.4
2020 π-Hub: Large-scale video learning, storage, and retrieval on heterogeneous hardware platforms
Jie Tang 0003, Shaoshan Liu, Jie Cao 0003, Bolin Ding, Jean-Luc Gaudiot, Weisong Shi
Future Gener. Comput. Syst.7
2020 EdgeABC: An architecture for task offloading and resource allocation in the Internet of Things
Kaile Xiao, Zhipeng Gao 0001, Weisong Shi, Xuesong Qiu 0001, Yang Yang 0006, Lanlan Rui
Future Gener. Comput. Syst.3
2020 EdgeVCD: Intelligent Algorithm-Inspired Content Distribution in Vehicular Edge Computing Network
abstract
Vehicular edge computing (VEC), which integrates mobile-edge computing (MEC) into vehicular networks, can provide more capability for executing resource-hungry applications and lower latency for connected vehicles. Distributing the result content to connected vehicles is vital for them to take proper actions based on computing results. However, the increasing number of connected vehicles and the limited communication resources make the content distribution a challenge. Besides, the diversity of connected vehicles and contents makes it more challenging for content distribution. To address this issue, in this article, we propose EdgeVCD, an intelligent algorithm-inspired content distribution scheme. Specifically, we first propose a dual-importance (DI) evaluation approach to reflect the relationship between the Priority of Vehicles (PoV) and the Priority of Contents (PoC). To make use of the limited communication resources, we then formulate an optimization problem to maximize the system utility for content distribution. To solve the complex optimization problem effectively, we first divide the road into small segments. Then, we propose a fuzzy-logic-based method to select the most proper content replica vehicle (CRV) for aiding content distribution and redefine the number of content request vehicles in each segment. Thereafter, the optimization problem is transformed into a nonlinear integer programming problem. Inspired by the artificial immune system, we propose an immune clone-based algorithm to solve it, which has a fast convergence to an optimal solution. Extensive simulations validate the effectiveness of our proposed EdgeVCD in terms of system utility, average utility, and convergence.
Quyuan Luo, Changle Li, Tom H. Luan, Weisong Shi
IEEE Internet Things J.4
2020 Collaborative Data Scheduling for Vehicular Edge Computing via Deep Reinforcement Learning
abstract
With the development of autonomous driving, the surging demand for data communications as well as computation offloading from connected and automated vehicles can be expected in the foreseeable future. With the limited capacity of both communication and computing, how to efficiently schedule the usage of resources in the network toward best utilization represents a fundamental research issue. In this article, we address the issue by jointly considering the communication and computation resources for data scheduling. Specifically, we investigate on the vehicular edge computing (VEC) in which edge computing-enabled roadside unit (RSU) is deployed along the road to provide data bandwidth and computation offloading to vehicles. In addition, vehicles can collaborate among each other with data relays and collaborative computing via vehicle-to-vehicle (V2V) communications. A unified framework with communication, computation, caching, and collaborative computing is then formulated, and a collaborative data scheduling scheme to minimize the system-wide data processing cost with ensured delay constraints of applications is developed. To derive the optimal strategy for data scheduling, we further model the data scheduling as a deep reinforcement learning problem which is solved by an enhanced deep $Q$ -network (DQN) algorithm with a separate target $Q$ -network. Using extensive simulations, we validate the effectiveness of the proposal.
Quyuan Luo, Changle Li, Tom H. Luan, Weisong Shi
IEEE Internet Things J.4
2020 VU: Edge Computing-Enabled Video Usefulness Detection and its Application in Large-Scale Video Surveillance Systems
abstract
In the era of smart and connected communities, video surveillance systems, which typically involve tens to thousands of cameras, have increasingly become prominent components for public safety. In current practice, when a failure occurs in a video surveillance system, the operation and maintenance teams usually spend a substantial amount of time locating and identifying the failure; hence, the fast online response cannot be guaranteed in a large-scale video surveillance system. Meanwhile, the video data that contains potential failures consumes bandwidth that could be used for useful video data. The useless video will waste the scarce bandwidth in the network and storage usage in the cloud. The emergence of edge computing is highly promising in video preprocessing with an edge camera. A video surveillance system is a killer application for edge computing. In this article, we propose an edge computing-enabled video usefulness (i.e., VU) model for large-scale video surveillance systems. We also explore its application, e.g., early failure detection and bandwidth improvement. According to the usefulness of the video data, the VU model can locate a failure and send it to end-users on the fly. In this article, our goals are threefold: 1) proposing a comprehensive VU model. To the best of our knowledge, this is the first work to explore the feasibility of the VU model and to determine VU values in a real application; 2) reducing the mean time to detection (i.e., MTTD) efficiently via edge computing-enabled fast online failure detection approaches; and 3) relieving the network bandwidth for large-scale video surveillance systems. Our experimental results demonstrate the approaches in VU model accurately detect failures that were collected from a video surveillance system with approximately 4000 cameras. The MTTD is substantially shortened by the fast online detection approaches. The video data with the worst VU values is mostly discarded to lessen overload of the network.
Hui Sun 0002, Weisong Shi
IEEE Internet Things J.2
2020 DAER: A Resource Preallocation Algorithm of Edge Computing Server by Using Blockchain in Intelligent Driving
abstract
The introduction of edge computing (EC) in intelligent driving allows the vehicle to offload tasks to the EC server closer to the vehicle side, creating a new paradigm for task offloading and resource allocation. The movement of the vehicle, the time sensitivity of the processing data, and the resource allocation of the EC server have become bottlenecks of the rapid development of intelligent driving. In this article, we jointly considered the problems of the network economy and resource allocation. In order to eliminate dependence on third parties, we propose a resource transaction architecture based on the blockchain. Moreover, we propose the dynamic allocation algorithm of edge resources (DAERs) based on the double auction mechanism to maximize the satisfaction of users and service providers of edge computing (SPs), where the DAER algorithm is implemented in the form of smart contracts in the blockchain architecture. In particular, we propose the state search algorithm that can improve the prediction accuracy of the staged destination of the vehicle to help allocate resources reasonably. Through simulation experiments, we verify the superior performance of the DAER algorithm in terms of resource utilization rate and the satisfaction of both parties participating in the auction.
Kaile Xiao, Weisong Shi, Zhipeng Gao 0001, Congcong Yao, Xuesong Qiu 0001
IEEE Internet Things J.2
2020 A Case for Adaptive Resource Management in Alibaba Datacenter Using Neural Networks
Sa Wang, Yan-Hai Zhu, Shan-Pei Chen, Tianze Wu, Wen-Jie Li, Xusheng Zhan, Haiyang Ding, Weisong Shi, Yungang Bao
J. Comput. Sci. Technol.8
2019 An Empirical Study of Quad-Level Cell (QLC) NAND Flash SSDs for Big Data Applications
abstract
As the SSD technology develops, quad-level cell (QLC) NAND based SSD is gradually being introduced to the market. And as such, we evaluate the QLC technologys impact on the landscape of modern datacenters. Since a large number of applications and workloads in the modern datacenters have far more read requests than they write requests, QLC SSD provides a promising solution. For example, real-time analytics and big data, machine and deep learning, and read-intensive AI applications are all read hungry perfectly suited for the QLC. Its favorable performance (especially in reads), high capacity and density greatly help modern datacenters to provide more efficient services to their customers. At the same time, the low cost of QLC SSD also helps to lower the cost of operation for datacenters. Additionally, we explore the state-of-art QLC SSD from the system architecture point of view to shows its key advancements from previous technologies. By conducting a comprehensive performance evaluation of QLC SSD, we are able to compare it with other types of SSD and analyze factors that impact its performance.
Shuwen Liang, Zhi Qiao 0001, Sihai Tang, Jacob Hochstetler, Song Fu, Weisong Shi, Hsing-bung Chen
IEEE BigData6
2019 OpenEI: An Open Framework for Edge Intelligence
abstract
In the last five years, edge computing has attracted tremendous attention from industry and academia due to its promise to reduce latency, save bandwidth, improve availability, and protect data privacy to keep data secure. At the same time, we have witnessed the proliferation of AI algorithms and models which accelerate the successful deployment of intelligence mainly in cloud services. These two trends, combined together, have created a new horizon: Edge Intelligence (EI). The development of EI requires much attention from both the computer systems research community and the AI community to meet these demands. However, existing computing techniques used in the cloud are not applicable to edge computing directly due to the diversity of computing sources and the distribution of data sources. We envision that there missing a framework that can be rapidly deployed on edge and enable edge AI capabilities. To address this challenge, in this paper we first present the definition and a systematic review of EI. Then, we introduce an Open Framework for Edge Intelligence (OpenEI), which is a lightweight software platform to equip edges with intelligent processing and data sharing capability. We analyze four fundamental EI techniques which are used to build OpenEI and identify several open problems based on potential research directions. Finally, four typical application scenarios enabled by OpenEI are presented.
Xingzhou Zhang, Yifan Wang 0005, Sidi Lu, Liangkai Liu, Lanyu Xu, Weisong Shi
ICDCS6
2019 MobileEdge: Enhancing On-Board Vehicle Computing Units Using Mobile Edges for CAVs
abstract
As the rapid growth of connected and autonomous vehicles (CAVs) and 5G intensifies, more third-party applications are increasingly being deployed on CAVs. They not only improve user experience but also provide more helpful services, for example, enhancing public safety by recognizing criminals in real-time videos. Current CAVs prefer to process collected data on the vehicle to avoid long transmission latency and extra network cost. However, due to the limitations of the on-board vehicle computing unit (VCU) and increasing use of computing-intensive in-vehicle applications, the burden of on-board VCU has sharply increased, which may affect driving safety. In particular, for existing vehicles on the road, adding more computing devices is a challenge if not impossible due to cost concerns. Inspired by edge computing, we propose a novel platform, MobileEdge, to enhance the computing capability of the unchangeable on-board VCU, which leverages mobile devices as edge nodes, e.g., the passengers' smartphones, by offloading computing tasks to them for collaboratively computing. Moreover, MobileEdge provides the dynamic management of mobile devices, monitoring device status and interfaces for customizable task offloading strategies and eventually achieves optimal task scheduling. We build a prototype to demonstrate the designed platform and evaluate three task offloading strategies which were implemented based on the developed interfaces. The results show that MobileEdge significantly reduces the application response latency. Compared with the baseline which does not employ task offloading, the response latency is almost near real-time when more computing resources are available. In addition, the proposed shortest response latency strategy outperforms the best overall task scheduling among the three strategies.
Qingyang Zhang 0001, Youhuizi Li, Hong Zhong 0001, Weisong Shi
ICPADS5
2019 Near-Data Processing-Enabled and Time-Aware Compaction Optimization for LSM-tree-based Key-Value Stores
abstract
With the growing volume of storage systems, the traditional relational databases cannot reach the high performance required by big-data applications. As high-throughput alternatives to relational databases, LSM-tree-based key-value stores (KV stores in short) are confronted with degraded write performance during compaction under update-intensive workloads. To address this issue, we design and implement a time-aware compaction optimization framework for KV stores called TStore. TStore explores the near-data processing (i.e., NDP) model. It dynamically partitions compaction tasks into both host and NDP-enabled device to minimize the total time of compaction. The partitioned compaction tasks are conducted by the host and the device in parallel. The NDP-based devices exhibit low-latency, high-performance and high-bandwidth capability, thus facilitating key-value stores. TStore can not only accomplish compaction for KV stores, but also improve overall performance by removing bottleneck in compaction. Results show that the TStore with an NDP framework can achieve 3.8x and 1.9x performance improvement over LevelDB and Co-KV under the db_bench workload. In addition, the TStore-enabled KV store outperforms LevelDB and Co-KV by a factor of 3.6x and 1.9x in throughput and 72.0% and 48.9% in latency, respectively, under realistic workloads generated by YCSB.
Hui Sun 0002, Jianzhong Huang 0001, Song Fu, Zhi Qiao 0001, Weisong Shi
ICPP6
2019 CalmWPC: A buffer management to calm down write performance cliff for NAND flash-based storage systems
Hui Sun 0002, Jianzhong Huang 0001, Xiao Qin 0001, Weisong Shi
Future Gener. Comput. Syst.5
2019 Collaborative Compaction Optimization System using Near-Data Processing for LSM-tree-based Key-Value Stores
Hui Sun 0002, Jianzhong Huang 0001, Weisong Shi
J. Parallel Distributed Comput.4
2019 Edge Computing for Autonomous Driving: Opportunities and Challenges
abstract
Safety is the most important requirement for autonomous vehicles; hence, the ultimate challenge of designing an edge computing ecosystem for autonomous vehicles is to deliver enough computing power, redundancy, and security so as to guarantee the safety of autonomous vehicles. Specifically, autonomous driving systems are extremely complex; they tightly integrate many technologies, including sensing, localization, perception, decision making, as well as the smooth interactions with cloud platforms for high-definition (HD) map generation and data storage. These complexities impose numerous challenges for the design of autonomous driving edge computing systems. First, edge computing systems for autonomous driving need to process an enormous amount of data in real time, and often the incoming data from different sensors are highly heterogeneous. Since autonomous driving edge computing systems are mobile, they often have very strict energy consumption restrictions. Thus, it is imperative to deliver sufficient computing power with reasonable energy consumption, to guarantee the safety of autonomous vehicles, even at high speed. Second, in addition to the edge system design, vehicle-to-everything (V2X) provides redundancy for autonomous driving workloads and alleviates stringent performance and energy constraints on the edge side. With V2X, more research is required to define how vehicles cooperate with each other and the infrastructure. Last, safety cannot be guaranteed when security is compromised. Thus, protecting autonomous driving edge computing systems against attacks at different layers of the sensing and computing stack is of paramount concern. In this paper, we review state-of-the-art approaches in these areas as well as explore potential solutions to address these challenges.
Shaoshan Liu, Liangkai Liu, Jie Tang 0003, Bo Yu 0014, Yifan Wang 0005, Weisong Shi
Proc. IEEE6
2019 Edge Computing [Scanning the Issue]
abstract
In recent years, with the proliferation of the Internet of Things (IoT) and the wide penetration of wireless networks, the number of edge devices and the data generated from the edge have been growing rapidly. According to International Data Corporation (IDC) prediction[20], global data will reach 180 zettabytes (ZB), and 70% of the data generated by IoT will be processed on the edge of the network by 2025. IDC also forecasts that more than 150 billion devices will be connected worldwide by 2025. In this case, the centralized processing mode based on cloud computing is not efficient enough to handle the data generated by the edge. The centralized processing model uploads all data to the cloud data center through the network and leverages its supercomputing power to solve the computing and storage problems, which enables the cloud services to create economic benefits. However, in the context of IoT, traditional cloud computing has several shortcomings.
Weisong Shi, George Pallis 0001
Proc. IEEE1
2019 Guest Editor's Introduction: Special Section on Fog/Edge Computing and Services
abstract
The papers in this special section focus on fog computing and services. The emerging Internet of Things (IoT) and rich cloud services have helped create the need for fog computing (also known as edge computing), in which data processing occurs in part at the network edge or anywhere along the cloud-to-endpoint continuum that can best meet user requirements, rather than completely in a relatively small number of massive clouds. Fog computing could address latency concerns, devices’ limited processing and storage capabilities and battery life, network bandwidth constraints and costs, and many security and privacy concerns that arise from the emerging IoT.
Weisong Shi, Tao Zhang 0005, Qun Li 0001
IEEE Trans. Serv. Comput.1
2018 Reliability Characterization of Solid State Drives in a Scalable Production Datacenter
abstract
In recent years, NAND flash-based solid state drives (SSD) have been widely used in datacenters due to their better performance compared with the traditional hard disk drives. However, little is known about the reliability characteristics of SSDs in production systems. Existing works study the statistical distributions of SSD failures in the field. However, they do not go deep into SSD drives and investigate the unique error types and health dynamics that distinguish SSDs from hard disk drives. In this paper, we explore the SSD-specific SMART (Self-Monitoring, Analysis, and Reporting Technology) attributes to conduct an in-depth analysis of SSD reliability in a production environment. Data is collected from a scalable production system having several physical locations. Our dataset contains over a million records with more than twenty attributes. We leverage machine learning technologies, specifically data clustering and correlation analysis methods, to discover groups of SSDs which have different health status and relations among SSD-specific SMART attributes. Our results show that 1) Media wear affects the reliability of SSDs more than any other factors, and 2) SSDs transit from one health group to another which infers the reliability degradation of those drives. To the best of our knowledge, this is the first study that investigates SSD-specific SMART data to characterize SSD reliability in a production environment.
Shuwen Liang, Zhi Qiao 0001, Jacob Hochstetler, Song Fu, Weisong Shi, Devesh Tiwari, Hsing-bung Chen, Bradley W. Settlemyer, David Richard Montoya
IEEE BigData6
2018 Teaching Autonomous Driving Using a Modular and Integrated Approach
abstract
Introduction: Teaching autonomous driving is a challenging task. Indeed, most existing autonomous driving teaching activities focus on a few of the technologies involved. This not only fails to provide a comprehensive coverage, but also sets a high entry barrier for students with different backgrounds. Objective: The primary objective of this study is to present a modular, integrated approach towards teaching autonomous driving. Methods: We organize the technologies used in autonomous driving into modules. This is described in the textbook we have developed as well as a series of multimedia online lectures designed to provide technical overview for each module. Once the students have understood these modules, the experimental platforms for integration we have developed allow the students to fully understand how the modules interact with each other. Results: To verify this teaching approach, we present three case studies: an introductory class on autonomous driving for students with only a basic technology background; a new session in an existing embedded systems class to demonstrate how embedded system technologies can be applied towards autonomous driving; and an industry professional training session to quickly bring up experienced engineers to work in autonomous driving. The results show that students can maintain a high interest level and make great progress by starting with familiar concepts before moving onto other modules. Conclusions: Autonomous driving is not one single technology, but rather a complex system integrating many technologies. Our modular and integrated approach is an effective method in teaching autonomous driving.
Jie Tang 0003, Shaoshan Liu, Songwen Pei, Stéphane Zuckerman, Chen Liu 0001, Weisong Shi, Jean-Luc Gaudiot
COMPSAC (1)6
2018 Understanding and Analyzing Interconnect Errors and Network Congestion on a Large Scale HPC System
abstract
Today's High Performance Computing (HPC) systems are capable of delivering performance in the order of petaflops due to the fast computing devices, network interconnect, and back-end storage systems. In particular, interconnect resilience and congestion resolution methods have a major impact on the overall interconnect and application performance. This is especially true for scientific applications running multiple processes on different compute nodes as they rely on fast network messages to communicate and synchronize frequently. Unfortunately, the HPC community lacks state-of-practice experience reports that detail how different interconnect errors and congestion events occur on large-scale HPC systems. Therefore, in this paper, we process and analyze interconnect data of the Titan supercomputer to develop a thorough understanding of interconnects faults, errors and congestion events. We also study the interaction between interconnect, errors, network congestion and application characteristics.
Saurabh Gupta 0002, Tirthak Patel, Michael Wilder, Weisong Shi, Song Fu, Christian Engelmann, Devesh Tiwari
DSN5
2018 OpenVDAP: An Open Vehicular Data Analytics Platform for CAVs
abstract
In this paper, we envision the future connected and autonomous vehicles (CAVs) as a sophisticated computer on wheels, with substantial on-board sensors as data sources and a variety of services running on top to support autonomous driving or other functions. In general, these services are computationally expensive, especially for the machine learning based applications (e.g., CNN-based object detection). Nevertheless, the on-board computation unit possess limited compute resources, raising a huge challenge to deploy these computation-intensive services on the vehicle. On the contrary, the cloud-based architecture conceptually with unconstrained resources suffers from unexpected extended latency that attributes to the large-scale Internet data transmission; thus, adversely affecting the services' real-time performance, quality of services and user experiences. To address this dilemma, inspired by the promising edge computing paradigm, we propose to build an Open Vehicular Data Analytics Platform (OpenVDAP) for CAVs, which is a full-stack edge based platform including an on-board computing/communication unit, an isolation-supported and security & privacy-preserved vehicle operation system, an edge-aware application library, as well as an optimal workload of?oading and scheduling strategy, allowing CAVs to dynamically detect each service's status, computation overhead and the optimal of?oading destination so that each service could be finished within an acceptable latency and limited bandwidth consumption. Most importantly, contrast to the proprietary platform, OpenVDAP is an open-source platform that offers free APIs and real-?eld vehicle data to the researchers and developers in the community, allowing them to deploy and evaluate applications on the real environment.
Qingyang Zhang 0001, Yifan Wang 0005, Xingzhou Zhang, Liangkai Liu, Xiaopei Wu, Weisong Shi, Hong Zhong 0001
ICDCS6
2018 EASE: Energy Efficiency and Proportionality Aware Virtual Machine Scheduling
abstract
Servers have different energy efficiency and energy proportionality (EP) due to their hardware configuration (i.e., CPU generation and memory installation) and workload. However, current virtual machine (VM) scheduling in virtualized environments will saturate servers without considering their energy efficiency and EP differences. This article will discuss EASE, the energy efficiency and proportionality aware VM scheduling approach. EASE first executes customized computing intensive, memory intensive, and hybrid benchmarks to calculate a server's energy efficiency and EP. Then it schedules VMs to servers to keep them working at their peak energy efficiency point (or optimal working range). This step improves the overall energy efficiency of the cluster and the data center. For performance guarantee, EASE migrates VMs from servers under highly contending conditions. The experimental results on real clusters show that power consumption can be saved 37.07% ~ 49.98% in the homogeneous cluster. The average completion time of the computing intensive VMs increases only 0.31 % ~ 8.49%. In the heterogeneous nodes, the power consumption of the computing intensive VMs can be reduced by 44.22 %. The job completion time can be saved by 53.80%.
Congfeng Jiang, Yumei Wang, Dongyang Ou, Yeliang Qiu, Youhuizi Li, Jian Wan 0001, Weisong Shi, Christophe Cérin
SBAC-PAD8
2018 Characterizing the Effectiveness of Query Optimizer in Spark
abstract
In the big data community, Spark has been widely used for processing interactive queries. Spark employs a query optimizer, called Catalyst, to provides a set of optimization rules and supports Cost-Based Optimization (CBO). In this paper, we investigated the effectiveness of the optimization rules and cost-based optimization in Catalyst. We conducted comprehensive validation experiments by varying the data volume and cluster scale, and found that the execution time of most TPC-H queries were reduced slightly even when query optimizations are applied. We derived some interesting observations on Catalyst, which can help the community better understand and improve the query optimizer of Spark in future.
Zujie Ren, Na Yun, Weisong Shi, Youhuizi Li, Jian Wan 0001, Lihua Yu, Xinxin Fan
SERVICES3
2018 On security challenges and open issues in Internet of Things
Kewei Sha, Wei Wei 0043, T. Andrew Yang, Zhiwei Wang 0003, Weisong Shi
Future Gener. Comput. Syst.5
2018 Distributed Collaborative Execution on the Edges and Its Application to AMBER Alerts
abstract
In the Internet of Everything era, billions of geographically distributed things will connect to the Internet and generate hundreds of zettabytes of data per year. Pushing that data to the cloud requires tremendous network bandwidth cost and latency. This is too onerous for some latency-sensitive applications, such as vehicle tracking using city-wide cameras. One application currently limited by such obstacles is the America's Missing Broadcast Emergency Response (AMBER) Alert system-but edge computing could transform this system's capabilities. Edge computing is a new computing paradigm that greatly diminishes data transmission and response latency by processing data at the proximity of data sources. However, most vision-based analytics are compute-intensive, and an edge device might be overwhelmed given tens of frames each second for real-time analysis. Also, the system needs a customized and flexible interface to implement efficient tracking strategies. To meet these needs, here we extend a big data processing framework, called Firework, to support collaboration between multiple edge devices and customizable task-scheduling strategies. Based on this extended version of Firework, we implement the AMBER alert assistant (A3), which efficiently tracks and locates a vehicle by analyzing city cameras' data in real time. We also propose two kinds of customized task-scheduling algorithms for vehicle tracking in A3. Comprehensive evaluation results show that A3 achieves real-time video analytics by collaborating among multiple edge devices; and the proposed location-direction-related diffusion strategy effectively controls the searching area for vehicle tracking by smartly selecting candidate cameras.
Qingyang Zhang 0001, Quan Zhang 0001, Weisong Shi, Hong Zhong 0001
IEEE Internet Things J.3
2018 Firework: Data Processing and Sharing for Hybrid Cloud-Edge Analytics
abstract
Now we are entering the era of the Internet of Everything (IoE) and billions of sensors and actuators are connected to the network. As one of the most sophisticated IoE applications, real-time video analytics is promising to significantly improve public safety, business intelligence, and healthcare & life science, among others. However, cloud-centric video analytics requires that all video data must be preloaded to a centralized cluster or the cloud, which suffers from high response latency and high cost of data transmission, given the scale of zettabytes of video data generated by IoE devices. Moreover, video data is rarely shared among multiple stakeholders due to various concerns, which restricts the practical deployment of video analytics that takes advantages of many data sources to make smart decisions. Furthermore, there is no efficient programming interface for developers and users to easily program and deploy IoE applications across geographically distributed computation resources. In this paper, we present a new computing framework,Firework, which facilitates distributed data processing and sharing for IoE applications via a virtual shared data view and service composition. We designed an easy-to-use programming interface forFireworkto allow developers to program onFirework. This paper describes the system design, implementation, and programming interface of Firework. The experimental results of a video analytics application demonstrate thatFireworkreduces up to 19.52 percent of response latency and at least 72.77 percent of network bandwidth cost, compared to a cloud-centric solution.
Quan Zhang 0001, Qingyang Zhang 0001, Weisong Shi, Hong Zhong 0001
IEEE Trans. Parallel Distributed Syst.3
2017 EdgeOS_H: A Home Operating System for Internet of Everything
abstract
The proliferation of Internet of Everything (IoE) and the success of rich Cloud services have pushed the horizon of a new computing paradigm, Edge Computing, which calls for processing the data at the edge of the network. Smart home as a typical IoE application is being widely adapted into people's lives. Edge Computing has the potential to empower the smart home, but it needs more contribution from the community before it truly benefits our lives. In this paper, we present the vision of EdgeOSH, a home operating system for Internet of Everything. We further discuss functional challenges, namely programming interface, self-management, data management, security & privacy, and naming, as well as non-functional challenges, such as user experience, system cost, delay, and the lack of availability of open testbed. Within each challenge we also discuss the potential directions that are worth further investigation.
Jie Cao 0003, Lanyu Xu, Raef Abdallah, Weisong Shi
ICDCS4
2017 Energy Proportional Servers: Where Are We in 2016?
abstract
The huge energy consumption in data centers produces not only high electricity bill but also tremendous carbon footprints. Although today's servers and data centers of leading internet companies are more energy efficient than ever before, the fluctuations in external workload and internal resource utilization calls for energy proportional computing. Insight into server energy proportionality can help improve workload placement while also reducing energy consumption. In this paper, we investigate all 477 valid published results of SPECpower_ssj benchmark from 2007 to 2016Q3 and reorganize them by hardware availability year for more accurate analysis on production servers. Through comprehensive analysis we find that: (1) The specious stagnation of energy proportionality in recent years is mainly caused by the adoption of processors of specific microarchitecture and is not the indicative trend of energy proportionality improvement. (2) Microarchitecture evolution has more influence on energy efficiency improvement than energy proportionality. (3) Today's servers' peak energy efficiencies are shifting from 100% resource utilization to 80% or 70% utilization and server energy proportionality improves with such shifting. We then conduct extensive experiments on 4 rack servers to investigate the energy efficiency variations under different hardware configurations, including memory per core installation and processor frequency scaling. Our experiments show that hardware configuration has significant impact on server's energy efficiency. Our findings presented in this paper provide useful insights and guidance to system designers, as well as data center operators for energy proportionality aware workload placement and energy savings.
Congfeng Jiang, Yumei Wang, Dongyang Ou, Weisong Shi
ICDCS5
2017 LAVEA: Latency-Aware Video Analytics on Edge Computing Platform
abstract
We present LAVEA, a system built for edge computing, which offloads computation tasks between clients and edge nodes, collaborates nearby edge nodes, to provide low-latency video analytics at places closer to the users. We have utilized an edge-first design to minimize the response time, and compared various task placement schemes tailed for inter-edge collaboration. Our results reveal that the client-edge configuration has task speedup against local or client-cloud configurations.
Shanhe Yi, Zijiang Hao, Qingyang Zhang 0001, Quan Zhang 0001, Weisong Shi, Qun Li 0001
ICDCS5
2017 Quantifying the Isolation Characteristics in Container Environments
Yusen Wu 0001, Zujie Ren, Weisong Shi, Jian Wan 0001
NPC4
2017 Mitigating Data Sparsity Using Similarity Reinforcement-Enhanced Collaborative Filtering
abstract
The data sparsity problem has attracted significant attention in collaborative filtering-based recommender systems. To alleviate data sparsity, several previous efforts employed hybrid approaches that incorporate auxiliary data sources into recommendation techniques, like content, context, or social relationships. However, due to privacy and security concerns, it is generally difficult to collect such auxiliary information. In this article, we focus on the pure collaborative filtering methods without relying on any auxiliary data source. We propose an improved memory-based collaborative filtering approach enhanced by a novel similarity reinforcement mechanism. It can discover potential similarity relationships between users or items by making better use of known but limited user-item interactions, thus to extract plentiful historical rating information from similar neighbors to make more reliable and accurate rating predictions. This approach integrates user similarity reinforcement and item similarity reinforcement into a comprehensive framework and lets them enhance each other. Comprehensive experiments conducted on several public datasets demonstrate that, in the face of data sparsity, our approach achieves a significant improvement in prediction accuracy when compared with the state-of-the-art memory-based and model-based collaborative filtering algorithms.
Weisong Shi, Hong Li 0004
ACM Trans. Internet Techn.2
2017 Realistic and Scalable Benchmarking Cloud File Systems: Practices and Lessons from AliCloud
abstract
The past decade has witnessed the rapid boom of cloud computing. Many public cloud infrastructures have been implemented and serve millions of tenants. Cloud file systems, which take charge of petabyte-scale data storage, play a crucial role in the performance of cloud infrastructures. Typical cloud file systems, including GFS, HDFS and Ceph, have attracted notable research efforts for performance evaluation and optimization. However, due to the heterogeneity and complexity of I/O workload characteristics in cloud environments, it is still challenging to conduct an accurate and efficient performance evaluation. To address this problem, we collected a two-week I/O workload trace from a 2,500-node production cluster in AliCloud, which is one of the largest cloud providers in Asia. Using the AliCloud trace, we characterized the I/O workload and data distribution, and compared two cloud services in multiple perspectives, including the request arrival pattern, request size, data population and so on. A list of observations and implications were derived and applied to help design a cloud file system benchmarking suite, called Porcupine. Porcupine aims to deploy a scalable and efficient performance evaluation on cloud file systems using realistic I/O workloads. We conducted a group of validation experiments, which demonstrated that Porcupine can achieve high accuracy and scalability. This paper provides our experiences and lessons in generating I/O workloads and deploying performance tests on cloud file systems, which we believe will be insightful to the cloud computing community in general.
Zujie Ren, Weisong Shi, Jian Wan 0001, Jiangbin Lin
IEEE Trans. Parallel Distributed Syst.2
2016 iGen: A Realistic Request Generator for Cloud File Systems Benchmarking
abstract
Benchmarking is a traditional approach for system performance evaluation and optimization. Over the past decades, a variety of file systems, e.g., GFS, HDFS and Ceph, have been designed and implemented, serving as the key components in cloud infrastructures. With the mature of those cloud file systems, the demands for performance evaluation and comparison are also rising. However, due to the complexity and heterogeneity of I/O workloads in cloud infrastructures, it is still challenging to generate realistic I/O workloads. System developers often use traditional file system benchmarks and make inaccurate assumptions on workload generation, yielding to misleading results. To address this problem, we investigate the characteristics of I/O requests in a production cloud infrastructure at Alibaba Cloud Computing, which is one of the biggest cloud providers in Asia. We proposed a flexible framework iGen to mimic I/O request arrivals. One of the salient features of the iGen is that the request arrival process is modeled by three statistics properties, request arrival rate, inter-arrival time distribution, and request periodicity. According to these properties, the iGen can determine the sequence of requests and the inter-arrival time between two subsequent requests. We use the iGen to emulate a real workload that collected from Alibaba cloud platform. Experimental results show that high accuracy and flexibility of the iGen.
Zujie Ren, Weisong Shi, Jiangbin Lin
CLOUD3
2016 H2O: A hybrid and hierarchical outlier detection method for large scale data protection
abstract
Data protection is the process of backing up data in case of a data loss event. It is one of the most critical routine activities for every organization. Detecting abnormal backup jobs is important to prevent data protection failures and ensure the service quality. Given the large scale backup endpoints and the variety of backup jobs, from a backup-as-a-service provider viewpoint, we need a scalable and flexible outlier detection method that can model a huge number of objects and well capture their diverse patterns. In this paper, we introduce H2O, a novel hybrid and hierarchical method to detect outliers from millions of backup jobs for large scale data protection. Our method automatically selects an ensemble of outlier detection models for each multivariate time series composed by the backup metrics collected for each backup endpoint by learning their exhibited characteristics. Interactions among multiple variables are considered to better detect true outliers and reduce false positives. In particular, a new seasonal-trend decomposition based outlier detection method is developed, considering the interactions among variables in the form of common trends, which is robust to the presence of outliers in the training data. The model selection process is hierarchical, following a global to local fashion. The final outlier is determined through an ensemble learning by multiple models. Built on top of Apache Spark, H2O has been deployed to detect outliers in a large and complex data protection environment with more than 600,000 backup endpoints and 3 million daily backup jobs. To the best of our knowledge, this is the first work that selects and constructs large scale outlier detection models for multivariate time series on Big Data platforms.
Quan Zhang 0001, Ramani Routray, Weisong Shi
IEEE BigData4
2016 Increasing large-scale data center capacity by statistical power control
abstract
Given the high cost of large-scale data centers, an important design goal is to fully utilize available power resources to maximize the computing capacity. In this paper we present Ampere, a novel power management system for data centers to increase the computing capacity by over-provisioning the number of servers. Instead of doing power capping that degrades the performance of running jobs, we use a statistical control approach to implement dynamic power management by indirectly affecting the workload scheduling, which can enormously reduce the risk of power violations. Instead of being a part of the already over-complicated scheduler, Ampere only interacts with the scheduler with two basic APIs. Instead of power control on the rack level, we impose power constraint on the row level, which leads to more room for over provisioning.
Guosai Wang, Shuhao Wang, Weisong Shi, Yinghang Zhu, Dianming Hu, Longbo Huang, Xin Jin 0008, Wei Xu 0005
EuroSys4
2016 Edge Computing: Vision and Challenges
abstract
The proliferation of Internet of Things (IoT) and the success of rich cloud services have pushed the horizon of a new computing paradigm, edge computing, which calls for processing the data at the edge of the network. Edge computing has the potential to address the concerns of response time requirement, battery life constraint, bandwidth cost saving, as well as data safety and privacy. In this paper, we introduce the definition of edge computing, followed by several case studies, ranging from cloud offloading to smart home and city, as well as collaborative edge to materialize the concept of edge computing. Finally, we present several challenges and opportunities in the field of edge computing, and hope this paper will gain attention from the community and inspire more research in this direction.
Weisong Shi, Jie Cao 0003, Quan Zhang 0001, Youhuizi Li, Lanyu Xu
IEEE Internet Things J.1
2016 Application configuration selection for energy-efficient execution on multicore systems
Shinan Wang, Weisong Shi, Devesh Tiwari
J. Parallel Distributed Comput.3
2016 eCope: Workload-Aware Elastic Customization for Power Efficiency of High-End Servers
abstract
Hardware components, especially CPU and memory, have made a lot of progress in terms of energy efficiency in the last decade. However, it is still far from the ideal energy proportional. Motivated by the recent observations that the energy efficiency of hardware components varies to a great extent depending on the workload characteristics, we propose eCope, workload-aware elastic customization for power efficiency of high-end servers, to reduce power consumption by workload aware and hardware customization for servers in datacenters. Our unique contribution is that eCope platform can take advantage of any configurable hardware that fits our assumption to improve the energy proportionality for various kinds of services without knowing the details of the target service. We illustrate three case studies to show how can we apply our idea to typical real-world back-end services (file system, database services and web-based services).
Shinan Wang, Weisong Shi, Yanfeng He
IEEE Trans. Cloud Comput.3
2015 MyPalmVein: A palm vein-based low-cost mobile identification system for wide age range
abstract
Traditional methods of access control such as password token or identification card are being replaced by biometrics recognition technology in many fields because of their limitation in reliability and usability. Vein pattern identification is outstanding compared to other biometrics methods such as fingerprint or iris due to its dependability and ease of use. At the same time, current biometrics recognition systems and applications have some disadvantages in usability, cost, and supported user age range. In this paper, we propose MyPalmVein, which is a low-cost identification system based on vein pattern recognition. We also introduce a mobile version of MyPalmVein system. The evaluation of MyPalmVeins performance on a public database of 250 subjects as well as field tests enrolled around 300 subjects from 2 months to 65+ years old shows that MyPalmVein is very accurate on wide age range user.
Jie Cao 0003, Mingyang Xu, Weisong Shi, Zhifeng Yu, Abdulbaset Salim, Paul Kilgore
HealthCom3
2015 ALARM: A novel fall detection algorithm based on personalized threshold
abstract
Since threshold-based fall detection has been widely studied by many research groups, accuracy is still a main limitation affected by personal factors. To this end, a personalized threshold extraction approach being adapted for the fall detection usage for different individual is proposed to increase the fall detection accuracy. Moreover, we also implement a fall detection algorithm called ALARM to verify the feasibility of the proposed threshold extraction approach. Results of comprehensive evaluation show it has high accuracy of 96.76% for fall detection, while the sensitivity and the specificity are 92.01% and 99.13%, respectively, based on the data collected from 8 volunteers.
Lingmei Ren, Weisong Shi, Zhifeng Yu, Jie Cao 0003
HealthCom2
2015 Security, trust, and resilience of distributed networks and systems
abstract
The rapid growth of distributed and networking technologies has made our information system more vulnerable to attack threats, malicious behaviors, and unpredictable failures. As the emergence of botnets and advanced persistent threat attacks, the traditional defense technology cannot cope well with the new large-scale and obfuscated malwares. In distributed and virtualized environments, the trust risk of applications has been increased considerably, which made it vital to propose new trust access control technologies. In addition, the complicated system usually comprises a large number of components that are susceptible to unpredictable failures. We need new designs of resilient infrastructure and dependable services. The papers in this special issue focus on the security, trust, and resilience management for distributed and networking computing paradigms, such as wireless sensor network, P2P network, ad hoc networks, virtualized network, and software-defined network. The contributions of these papers are outlined next. To locate the real source of the Internet attacks, existing work is easy to be evaded by attackers and difficult to justify the stepping stones. Sheng Wen et al. introduce the consistent causality probability to detect the stepping stones. They formulate the ranges of abnormal causality probabilities according to the different network conditions and further implement self-adaptive methods to capture stepping stones. To extract signatures for malwares, most existing string-based signatures extracting methods have the problem of inaccuracy and time consuming. Sun Hao et al. present a system for automatically extracting signatures from large-scale malwares, named AutoMal. The system can extract both byte signatures and hashed signatures from the malware network flows with high accuracy. Multi-interface multi-channel can reduce the channel interference and improve the network capacity for multi-hop wireless ad hoc networks. Tong Zhao et al. design a dynamic channel assignment algorithm that can dynamically switch the channels to the less busy ones by monitoring the channel usages. Moreover, the algorithm is designed in a fully distributed way with low overhead For the task allocation in wireless sensor networks (WSNs), traditional solutions for high-performance computing cannot be directly used in WSNs because of limitations of resource availability and shared communication medium. Wenzhong Guo et al. design a discrete particle swarm optimization to generate a structure of the parallel coalitions, and then introduce the game theory and redesigned fitness function to find the Nash equilibrium point for the purpose of improving the effectiveness of scheduling and the reliability of the network. Clustering approach has been considered one of the most effective measures for wireless sensor networks. Xiao-Hui Kuang et al. propose a novel energy-efficient clustering approach based on convergence degree chain, which is named ECACD. ECACD can improve the stability of topology, reduce the energy consumption, and decrease the communication cost. The distributed hash table (DHT) technology is widely used, which needs to take into account the real-time response and dynamic network maintenance for distributed communication systems. Kai Shuang et al. propose a hierarchical DHT lookup service named Comb, which is organized as a two-layered architecture; workload is distributed evenly among nodes, and most queries can be routed in no more than two hops. Few access control models have been proposed for security issues of multi-domain and virtualized network management. Yang Luo et al. enhance the classic role-based access control model through two concepts: domain and virtual machine. They define the virtualized role based access control (VRBAC) model in which authorized users can migrate or copy virtual machines from one domain to another without causing a conflict. In software-defined networks, network operating systems (NOSes) are required to share or exchange reachability and topological information. Pingping Lin et al. proposes a west–east bridge mechanism for distributed heterogeneous NOSes to cooperate in enterprise/data center/intra-autonomous system networks. Monitoring Border Gateway Protocol (BGP) is an effective way to improve the security of inter-domain routing. Ning Hu et al. present a cooperative BGP monitoring method called the cooperative information sharing model (CoISM). CoISM can provide a more comprehensive information view by introducing information diffuse reflection based on initiative inquiry and making use of the relativity of monitoring information. Securing mobile devices such as smart phones is inherently difficult. René Mayrhofer et al. review recent research results, systematically analyze the technical issues of securing mobile device platforms against different threats, and suggest potential approaches to create human-verifiable secure communication with components or services within partially untrusted devices. We would like to thank the editor-in-chief, Professor Hsiao-Hwa Chen, and co-editor-in-chief, Professor Hamid R. Sharif, for providing us the opportunity to host this special issue. We thank Prof. Guojun Wang for his great help in the organization of the special issue. We also thank all the authors who contributed their papers. Last but not least, we appreciate the work of many reviewers for this special issue.
Jinshu Su, Xiaofeng Wang 0002, Weisong Shi, Indrakshi Ray
Secur. Commun. Networks3
2015 Energy-Aware Scheduling of MapReduce Jobs for Big Data Applications
abstract
The majority of large-scale data intensive applications executed by data centers are based on MapReduce or its open-source implementation, Hadoop. Such applications are executed on large clusters requiring large amounts of energy, making the energy costs a considerable fraction of the data center's overall costs. Therefore minimizing the energy consumption when executing each MapReduce job is a critical concern for data centers. In this paper, we propose a framework for improving the energy efficiency of MapReduce applications, while satisfying the service level agreement (SLA). We first model the problem of energy-aware scheduling of a single MapReduce job as an Integer Program. We then propose two heuristic algorithms, called energy-aware MapReduce scheduling algorithms (EMRSA-I and EMRSA-II), that find the assignments of map and reduce tasks to the machine slots in orderto minimize the energy consumed when executing the application. We perform extensive experiments on a Hadoop cluster to determine the energy consumption and execution time for several workloads from the HiBench benchmark suite including TeraSort, PageRank, and K-means clustering, and then use this data in an extensive simulation study to analyze the performance of the proposed algorithms. The results show that EMRSA-I and EMRSA-II are able to find near optimal job schedules consuming approximately 40 percent less energy on average than the schedules obtained by a common practice scheduler that minimizes the makespan.
Lena Mashayekhy, Mahyar Nejad, Daniel Grosu, Quan Zhang 0001, Weisong Shi
IEEE Trans. Parallel Distributed Syst.5
2015 LTPPM: a location and trajectory privacy protection mechanism in participatory sensing
abstract
The ubiquity of mobile devices has facilitated the prevalence of participatory sensing, whereby ordinary citizens use their private mobile devices to collect regional information and to share with participators. However, such applications may endanger the users' privacy by revealing their locations and trajectories information. Most of existing solutions, which hide a user's location information with a coarse region, are under k-anonymity model. Yet, they may not be applicable in some participatory sensing applications that require precise location information. The goals are seemingly contradictory: to protect a user's location privacy while simultaneously providing precise location information for a high quality of service. In this paper, we propose a method to meet both goals. Through selecting a certain number of a user's partners, it can protect the user's location privacy while providing precise location information. The user's trajectory privacy can be protected by constructing several trajectories that are similar to the user's trajectory in an interval time T. Finally, we utilize a new metric, called slope ratio, to evaluate the partners' selection algorithm that we proposed. Then, we measure the privacy level that the location and trajectory privacy protection mechanism LTPPM can achieve. The analysis and simulation results show that LTPPM can protect the user's location and trajectory privacy effectively and also provide a high quality of service in participatory sensing. Copyright © 2012 John Wiley & Sons, Ltd.
Sheng Gao 0002, Jianfeng Ma 0001, Weisong Shi, Guoxing Zhan
Wirel. Commun. Mob. Comput.3
2014 A framework for component selection in collaborative sensing application development
abstract
Wireless sensor network-based technologies and applications have attracted a lot of attention in the past two decades because of their huge potential to change people’s way of life. These applications usually need close collaboration among multiple sensors, gateways, services and end users. When dev
Jie Cao 0003, Lingmei Ren, Weisong Shi, Zhifeng Yu
CollaborateCom3
2014 BatteryExtender: an adaptive user-guided tool for power management of mobile devices
abstract
The battery life of mobile devices is one of their most important resources. Much of the literature focuses on accurately profiling the power consumption of device components or enabling application developers to develop energy-efficient applications through fine-grained power profiling. However, there is a lack of tools to enable users to extend battery life on demand. What can users do if they need their device to last for a specific duration in order to perform a specific task? To this extent, we developed BatteryExtender, a user-guided power management tool that enables the reconfiguration of the device's resources based on the workload requirement, similar to the principle of creating virtual machines in the cloud. It predicts the battery life savings based on the new configuration, in addition to predicting the impact of running applications on the battery life. Through our experimental analysis, BatteryExtender decreased the energy consumption between 10.03% and 20.21%, and in rare cases by up to 72.83%. The accuracy rate ranged between 92.37% and 99.72%.
Grace Metri, Weisong Shi, Monica Brockmeyer, Abhishek Agrawal
UbiComp2
2014 Privacy preserving shortest path routing with an application to navigation
Yong Xi, Loren Schwiebert, Weisong Shi
Pervasive Mob. Comput.3
2014 A Design and Analysis Framework for Thermal-Resilient Hard Real-Time Systems
abstract
We address the challenge of designing predictable real-time systems in an unpredictable thermal environment where environmental temperature may dynamically change (e.g., implantable medical devices). Towards this challenge, we propose a control-theoretic design methodology that permits a system designer to specify a set of hard real-time performance modes under which the system may operate. The system automatically adjusts the real-time performance mode based on the external thermal stress. We show (via analysis, simulations, and a hardware testbed implementation) that our control design framework is stable and control performance is equivalent to previous real-time thermal approaches, even under dynamic temperature changes. A crucial and novel advantage of our framework over previous real-time control is the ability to guarantee hard deadlines even under transitions between modes. Furthermore, our system design permits the calculation of a new metric called thermal resiliency that characterizes the maximum external thermal stress that any hard real-time performance mode can withstand. Thus, our design framework and analysis may be classified as a thermal stress analysis for real-time systems.
Pradeep M. Hettiarachchi, Nathan Fisher, Masud Ahmed, Le Yi Wang, Shinan Wang, Weisong Shi
ACM Trans. Embed. Comput. Syst.6
2014 Workload Analysis, Implications, and Optimization on a Production Hadoop Cluster: A Case Study on Taobao
abstract
Understanding the characteristics of MapReduce workloads in a Hadoop cluster is the key to making optimal configuration decisions and improving the system efficiency and throughput. However, workload analysis on a Hadoop cluster, particularly in a large-scale e-commerce production environment, has not been well studied yet. In this paper, we performed a comprehensive workload analysis using the trace collected from a 2000-node Hadoop cluster at Taobao, which is the biggest online e-commerce enterprise in Asia, ranked 10th in the world as reported by Alexa. The results of the workload analysis are representative and generally consistent with the data warehouses for e-commerce web sites, which can help researchers and engineers understand the workload characteristics of Hadoop in their production environments. Based on the observations and implications derived from the trace, we designed a workload generator Ankus, to expedite the performance evaluation and debugging of new mechanisms. Ankus supports synthesizing an e-commerce style MapReduce workload at a low cost. Furthermore, we proposed and implemented a job scheduling algorithm, Fair4S , which is designed to be biased towards small jobs. Small jobs account for the majority of the workload, and most of them require instant and interactive responses, which is an important phenomenon at production Hadoop systems. The inefficiency of Hadoop fair scheduler for handling small jobs motivates us to design the Fair4S, which introduces pool weights and extends job priorities to guarantee the rapid responses for small jobs. Experimental evaluation verified that the Fair4S accelerates the average waiting times of small jobs by a factor of 7 compared with the fair scheduler.
Zujie Ren, Jian Wan 0001, Weisong Shi, Xianghua Xu
IEEE Trans. Serv. Comput.3
2013 Advanced topics on cloud computing
Yanming Shen, Keqiu Li, Weisong Shi
J. Comput. Syst. Sci.3
2013 Preface
Erik R. Altman, Weisong Shi
J. Comput. Sci. Technol.2
2013 Cost-Aware Cooperative Resource Provisioning for Heterogeneous Workloads in Data Centers
abstract
Recent cost analysis shows that the server cost still dominates the total cost of high-scale data centers or cloud systems. In this paper, we argue for a new twist on the classical resource provisioning problem: heterogeneous workloads are a fact of life in large-scale data centers, and current resource provisioning solutions do not act upon this heterogeneity. Our contributions are threefold: first, we propose a cooperative resource provisioning solution, and take advantage of differences of heterogeneous workloads so as to decrease their peak resources consumption under competitive conditions; second, for four typical heterogeneous workloads: parallel batch jobs, web servers, search engines, and MapReduce jobs, we build an agile system PhoenixCloud that enables cooperative resource provisioning; and third, we perform a comprehensive evaluation for both real and synthetic workload traces. Our experiments show that our solution could save the server cost aggressively with respect to the noncooperative solutions that are widely used in state-of-the-practice hosting data centers or cloud systems: for example, EC2, which leverages the statistical multiplexing technique, or RightScale, which roughly implements the elastic resource provisioning technique proposed in related state-of-the-art work.
Jianfeng Zhan, Lei Wang 0004, Weisong Shi, Chuliang Weng, Xiutao Zang
IEEE Trans. Computers4
2013 TrPF: A Trajectory Privacy-Preserving Framework for Participatory Sensing
abstract
The ubiquity of the various cheap embedded sensors on mobile devices, for example cameras, microphones, accelerometers, and so on, is enabling the emergence of participatory sensing applications. While participatory sensing can benefit the individuals and communities greatly, the collection and analysis of the participators' location and trajectory data may jeopardize their privacy. However, the existing proposals mostly focus on participators' location privacy, and few are done on participators' trajectory privacy. The effective analysis on trajectories that contain spatial-temporal history information will reveal participators' whereabouts and the relevant personal privacy. In this paper, we propose a trajectory privacy-preserving framework, named TrPF, for participatory sensing. Based on the framework, we improve the theoretical mix-zones model with considering the time factor from the perspective of graph theory. Finally, we analyze the threat models with different background knowledge and evaluate the effectiveness of our proposal on the basis of information entropy, and then compare the performance of our proposal with previous trajectory privacy protections. The analysis and simulation results prove that our proposal can protect participators' trajectories privacy effectively with lower information loss and costs than what is afforded by the other proposals.
Sheng Gao 0002, Jianfeng Ma 0001, Weisong Shi, Guoxing Zhan, Cong Sun 0001
IEEE Trans. Inf. Forensics Secur.3
2013 LOBOT: Low-Cost, Self-Contained Localization of Small-Sized Ground Robotic Vehicles
abstract
It is often important to obtain the real-time location of a small-sized ground robotic vehicle when it performs autonomous tasks either indoors or outdoors. We propose and implement LOBOT, a low-cost, self-contained localization system for small-sized ground robotic vehicles. LOBOT provides accurate real-time, 3D positions in both indoor and outdoor environments. Unlike other localization schemes, LOBOT does not require external reference facilities, expensive hardware, careful tuning or strict calibration, and is capable of operating under various indoor and outdoor environments. LOBOT identifies the local relative movement through a set of integrated inexpensive sensors and well corrects the localization drift by infrequent GPS-augmentation. Our empirical experiments in various temporal and spatial scales show that LOBOT keeps the positioning error well under an accepted threshold.
Guoxing Zhan, Weisong Shi
IEEE Trans. Parallel Distributed Syst.2
2013 A Two-Tiered On-Demand Resource Allocation Mechanism for VM-Based Data Centers
abstract
In a shared virtual computing environment, dynamic load changes as well as different quality requirements of applications in their lifetime give rise to dynamic and various capacity demands, which results in lower resource utilization and application quality using the existing static resource allocation. Furthermore, the total required capacities of all the hosted applications in current enterprise data centers, for example, Google, may surpass the capacities of the platform. In this paper, we argue that the existing techniques by turning on or off servers with the help of virtual machine (VM) migration is not enough. Instead, finding an optimized dynamic resource allocation method to solve the problem of on-demand resource provision for VMs is the key to improve the efficiency of data centers. However, the existing dynamic resource allocation methods only focus on either the local optimization within a server or central global optimization, limiting the efficiency of data centers. We propose a two-tiered on-demand resource allocation mechanism consisting of the local and global resource allocation with feedback to provide on-demand capacities to the concurrent applications. We model the on-demand resource allocation using optimization theory. Based on the proposed dynamic resource allocation mechanism and model, we propose a set of on-demand resource allocation algorithms. Our algorithms preferentially ensure performance of critical applications named by the data center manager when resource competition arises according to the time-varying capacity demands and the quality of applications. Using Rainbow, a Xen-based prototype we implemented, we evaluate the VM-based shared platform as well as the two-tiered on-demand resource allocation mechanism and algorithms. The experimental results show that Rainbow without dynamic resource allocation (Rainbow-NDA) provides 26 to 324 percent improvements in the application performance, as well as 26 percent higher average CPU utilization than traditional service computing framework, in which applications use exclusive servers. The two-tiered on-demand resource allocation further improves performance by 9 to 16 percent for those critical applications, 75 percent of the maximum performance improvement, introducing up to 5 percent performance degradations to others, with 1 to 5 percent improvements in the resource utilization in comparison with Rainbow-NDA.
Yuzhong Sun, Weisong Shi
IEEE Trans. Serv. Comput.3
2012 Experimental Analysis of Application Specific Energy Efficiency of Data Centers with Heterogeneous Servers
abstract
Energy efficiency is an important issue for data centers given the amount of energy they consume yearly. However, there is still a gap of understanding of how exactly the application type and the heterogeneity of servers and their configuration impact the energy efficiency of data centers. To this end, we introduce the notion of Application Specific Energy Efficiency (ASEE) in order to rank energy efficiency of heterogeneous servers based on the hosted applications. We conducted extensive sets of experiments using three benchmarks: TPC-W, BS Seeker, and Matrix Stress mark. We observed that each server has different ASEE value based on the type of application running, the size of the virtual machine, the application load, and the scalability factor. In some cases, we witnessed 70% of ASEE improvement by changing the virtual machine size within the same node while keeping an identical load. In different cases, we witnessed up to 86% of ASEE improvement by running the same application with the same load within the same size of virtual machine but on different nodes. Our observation has many implications which include but are not limited to improving virtual machine scheduling based on the ASEE rank of the node. Another implication stresses on the importance of accurate prediction of application load and selecting the appropriate virtual machine size in order to improve the ASEE.
Grace Metri, Soumyasudharsan Srinivasaraghavan, Weisong Shi, Monica Brockmeyer
IEEE CLOUD3
2012 Effective and efficient?: bilingual sentiment lexicon extraction using collocation alignment
abstract
Bilingual sentiment lexicon is fundamental resource for cross-language sentiment analysis but its compilation remains a major bottleneck in computational linguistics. Traditional word alignment algorithm faces with the status of large alignment space, which may introduce redundant computations as well as alignment errors. In this paper, we use collocation alignment to extract bilingual sentiment lexicon overcoming the drawbacks of word alignment. The idea of collocation alignment is inspired by the strong cohesion between feature words and opinion words in sentiment corpus. Experimental results show that our approach not only decreases the computing time dramatically but also improves the precision of extracted bilingual word pairs due to the smaller alignment space.
Zheng Lin 0001, Songbo Tan, Xueqi Cheng 0001, Xueke Xu, Weisong Shi
CIKM5
2012 Dual-JT: Toward the high availability of JobTracker in Hadoop
abstract
MapReduce is a state-of-the-art computation paradigm that is becoming widely used for processing large-scale datasets. Hadoop is an open-source implementation of MapReduce and follows a masterCslave architecture. This architecture makes Hadoop suffer from a single point of failure in the JobTracker. In this paper, we design a solution to resolve the single point of failure of the Job Tracker and then enhance its availability. In this solution, a standby Job Tracker is introduced to act as a hot backup node of the active Job Tracker. The standby Job Tracker synchronizes the job execution process with the active Job Tracker by collecting and parsing the job log. If the active Job Tracker fails, the standby Job Tracker can take over quickly. This solution is implemented in Hadoop 0.20.x. Extensive experiments illustrate that this solution effectively enhances the availability of Job Tracker. A big production cluster in a large e-Commerce company has adopted this solution, which avoids interrupting job submission and execution when the Job Tracker fails or restarts.
Jian Wan 0001, Minggang Liu, Xixiang Hu, Zujie Ren, Weisong Shi
CloudCom6
2012 Topic 8: Distributed Systems and Algorithms
Andrzej M. Goscinski, Marios Mavronicolas, Weisong Shi, Yong Meng Teo
Euro-Par3
2012 EaSync: A Transparent File Synchronization Service across Multiple Machines
Huajian Mao, Nong Xiao 0001, Weisong Shi, Yutong Lu
NPC5
2012 The Design and Analysis of Thermal-Resilient Hard-Real-Time Systems
abstract
We address the challenge of designing predictable real-time systems in an unpredictable thermal environment where environmental temperature may dynamically change (e.g., implantable medical devices). Towards this challenge, we propose a control-theoretic design methodology which permits a system designer to specify a set of hard-real-time performance modes under which the system may operate. The system automatically adjusts the real-time performance mode based on the external thermal stress. We show (via analysis, simulations, and a hardware testbed implementation) that our control-design framework is stable and control performance is equivalent to previous real-time thermal approaches, even under dynamic temperature changes. A crucial and novel advantage of our framework over previous real-time control is the ability to guarantee hard deadlines even under transitions between modes. Furthermore, our system design permits the calculation of a new metric called thermal resiliency which characterizes the maximum external thermal stress that any hard-real-time performance mode can withstand. Thus, our design framework and analysis may be classified as a thermal stress analysis for real-time systems.
Pradeep M. Hettiarachchi, Nathan Fisher, Masud Ahmed, Le Yi Wang, Shinan Wang, Weisong Shi
IEEE Real-Time and Embedded Technology and Applications Symposium6
2012 Wukong: A cloud-oriented file service for mobile Internet devices
Huajian Mao, Nong Xiao 0001, Weisong Shi, Yutong Lu
J. Parallel Distributed Comput.3
2012 ACM/Springer Mobile Networks and Applications (MONET) Special Issue on "Collaborative Computing: Networking, Applications and Worksharing"
Weisong Shi, James B. D. Joshi, Tao Zhang 0005, Eun K. Park, Juan Quemada
Mob. Networks Appl.1
2012 Design and Implementation of TARF: A Trust-Aware Routing Framework for WSNs
abstract
The multihop routing in wireless sensor networks (WSNs) offers little protection against identity deception through replaying routing information. An adversary can exploit this defect to launch various harmful or even devastating attacks against the routing protocols, including sinkhole attacks, wormhole attacks, and Sybil attacks. The situation is further aggravated by mobile and harsh network conditions. Traditional cryptographic techniques or efforts at developing trust-aware routing protocols do not effectively address this severe problem. To secure the WSNs against adversaries misdirecting the multihop routing, we have designed and implemented TARF, a robust trust-aware routing framework for dynamic WSNs. Without tight time synchronization or known geographic information, TARF provides trustworthy and energy-efficient route. Most importantly, TARF proves effective against those harmful attacks developed out of identity deception; the resilience of TARF is verified through extensive evaluation with both simulation and empirical experiments on large-scale WSNs under various scenarios including mobile and RF-shielding network conditions. Further, we have implemented a low-overhead TARF module in TinyOS; as demonstrated, this implementation can be incorporated into existing routing protocols with the least effort. Based on TARF, we also demonstrated a proof-of-concept mobile target detection application that functions well against an antidetection mechanism.
Guoxing Zhan, Weisong Shi, Hongmei Deng 0001
IEEE Trans. Dependable Secur. Comput.2
2012 Providing hierarchical lookup service for P2P-VoD systems
abstract
Supporting random jump in P2P-VoD systems requires efficient lookup for the “best” suppliers, where “best” means the suppliers should meet two requirements: content match and network quality match . Most studies use a DHT-based method to provide content lookup; however, these methods are neither able to meet the network quality requirements nor suitable for VoD streaming due to the large overhead. In this paper, we propose Mediacoop, a novel hierarchical lookup scheme combining both content and quality match to provide random jumps for P2P-VoD systems. It exploits the play position to efficiently locate the candidate suppliers with required data (content match), and performs refined lookup within the candidates to meet quality match. Theoretical analysis and simulation results show that Mediacoop is able to achieve lower jump latency and control overhead than the typical DHT-based method. Moreover, we implement Mediacoop in a BitTorrent-like P2P-VoD system called CoolFish and make optimizations for such “total cache” applications. The implementation and evaluation in CoolFish show that Mediacoop is able to improve user experiences, especially the jump latency, which verifies the practicability of our design.
Tieying Zhang, Xueqi Cheng 0001, Jianming Lv, Zhenhua Li 0001, Weisong Shi
ACM Trans. Multim. Comput. Commun. Appl.5
2012 In Cloud, Can Scientific Communities Benefit from the Economies of Scale?
abstract
The basic idea behind cloud computing is that resource providers offer elastic resources to end users. In this paper, we intend to answer one key question to the success of cloud computing: in cloud, can small-to-medium scale scientific communities benefit from the economies of scale? Our research contributions are threefold: first, we propose an innovative public cloud usage model for small-to-medium scale scientific communities to utilize elastic resources on a public cloud site while maintaining their flexible system controls, i.e., create, activate, suspend, resume, deactivate, and destroy their high-level management entities-service management layers without knowing the details of management. Second, we design and implement an innovative system-DawningCloud, at the core of which are lightweight service management layers running on top of a common management service framework. The common management service framework of DawningCloud not only facilitates building lightweight service management layers for heterogeneous workloads, but also makes their management tasks simple. Third, we evaluate the systems comprehensively using both emulation and real experiments. We found that for four traces of two typical scientific workloads: High-Throughput Computing (HTC) and Many-Task Computing (MTC), DawningCloud saves the resource consumption maximally by 59.5 and 72.6 percent for HTC and MTC service providers, respectively, and saves the total resource consumption maximally by 54 percent for the resource provider with respect to the previous two public cloud solutions. To this end, we conclude that small-to-medium scale scientific communities indeed can benefit from the economies of scale of public clouds with the support of the enabling system.
Lei Wang 0004, Jianfeng Zhan, Weisong Shi
IEEE Trans. Parallel Distributed Syst.3
2011 Optimizing Web Browser on Many-Core Architectures
abstract
As more and more Web applications emerging on sever end today, the Web browser on client end has become a host of a variety of applications other than just rendering static Web pages. This leads to more and more performance requirements of a Web browser, for which user experience is very important. This situation may become more urgency when on handheld devices. Some efforts like redesign a new Web browser have been made to overcome this problem. In this paper, we address this issue by optimizing the main processes of the Web browser on a state-of-the-art 64-core architecture, Godson-T, which was developed at Chinese Academy of Sciences, as multi-/many-core architecture to be the mainstream processor in the upcoming years. We start a new core to process a new tab when facing up to intensive URL requests, and we use scratch-pad memory (SPM) of each core as a local buffer to store the HTML source data to be processed to reduce off-chip memory access and exploit more data locality, otherwise, we use DTA to transfer HTML data for backup. Experiments conducted on the cycle-accurate simulator show that, starting each tab process by a new core could obtain 5.7% to 50% speedup with different number of cores used to process corresponding URL requests, with on-chip scratchpad memory of each core used to store the HTML data, more speedup could be achieved when number of cores increase. Also, as Data Transfer Agent (DTA) used to transfer the HTML data, the backup of HTML data can get 2X to 5X speedups according to different data amount.
Lingjun Fan, Weisong Shi, Shibin Tang, Dongrui Fan
PDCAT2
2011 Automatic performance debugging of SPMD-style parallel programs
Xu Liu 0001, Jianfeng Zhan, Kunlin Zhan, Weisong Shi, Dan Meng 0002, Lei Wang 0004
J. Parallel Distributed Comput.4
2011 SensorTrust: A resilient trust model for wireless sensing systems
Guoxing Zhan, Weisong Shi, Hongmei Deng 0001
Pervasive Mob. Comput.2
2010 SPARTAN: A framework for Smart Phone Assisted Real-Time Health care Network design
abstract
Leveraging body area sensor network (BASN) for health care is a very promising application domain for wireless sensor networks. In a typical BASN health care application, usually, bio-sensors and environmental-sensors connect to a local Preprocessing Unit (PU) first, e.g., a smartphone or a laptop,
Shinan Wang, Weisong Shi, Bengt B. Arnetz, Clairy Wiholm
CollaborateCom2
2010 TARF: A Trust-Aware Routing Framework for Wireless Sensor Networks
Guoxing Zhan, Weisong Shi, Hongmei Deng 0001
EWSN2
2010 Incremental Sensor Node Deployment for Low Cost and Highly Available WSNs
abstract
We attack the sensor network deployment problem. We define the deployment problem as the problem of deciding how many sensor nodes should be deployed in the sensor field over how many phases during its lifetime. We target the optimal deployment strategy that meets user-defined availability requirement with minimum total cost taking into consideration node failures and changing field trip to sensor node cost ratio. We model WSN availability and total cost as functions of the deployment plan, then, we formalize the deployment problem as a 2D optimization problem. Our modeling enables us to explore cost-benefit tradeoffs, we believe, this is a solid step toward bringing cost as an explicit dimension in the design space of WSN protocols. We compare the performance of the optimized solution (denoted as pro-active) to more ad-hoc solutions: on-demand and at-front. The former strategy schedules future deployments only on demand. The latter strategy deploys all nodes at front with no later field trips. Using extensive simulations, we show that proactive outperforms at-front and on-demand in terms of total cost per availability unit in all application scenarios. For example, using pro-active costs $7 compared to $40 and $280 per total uptime in case of on-demand and at-front, respectively.
Safwan Al-Omari, Weisong Shi
MSN2
2010 Differentiated Replication Strategy in Data Centers
Tung Nguyen 0002, Anthony Cutway, Weisong Shi
NPC3
2010 Queue Waiting Time Aware Dynamic Workflow Scheduling in Multicluster Environments
Zhifeng Yu, Weisong Shi
J. Comput. Sci. Technol.2
2010 A reputation-driven scheduler for autonomic and sustainable resource sharing in Grid computing
Zhengqiang Liang, Weisong Shi
J. Parallel Distributed Comput.2
2010 TRECON: A Trust-Based Economic Framework for Efficient Internet Routing
abstract
The fragility and the poor resilience of the Internet are manifested by the severe impact of network activities and the slow recovery after an earthquake damaged undersea cables and disrupted telephone and Internet access in East Asia in December 2006. Except the inefficiency of routing protocols, lack of efficient network monitoring mechanisms and lack of economic incentives to encourage service providers (SPs) to act cooperatively and promptly are other important reasons. In this paper, we build a trust-based economic framework called TRECON to address these open problems in Internet routing. The novelty of TRECON is combining an adaptive personalized trust model with an economic approach to provide independent trust-based routing among SPs. TRECON provides flexible policy support based on the trust-based economic mechanism so that autonomous organizations with varied interests and optimization criteria can be smoothly integrated together to achieve better adaptiveness and self-management. Through introducing the economic model, TRECON explores a new way to solve the economic problems and incentives issues in the collaboration among SPs. To show the flexibility of routing policies support, we propose four typical routing policies under the TRECON framework. We evaluate our approach by comparing these four trust-derived routing policies with the classical global shortest path routing approach. We find that the policy based on trustworthiness performs much better than all other policies under different network topologies in terms of delay, success delivery rate, and economic effects.
Zhengqiang Liang, Weisong Shi
IEEE Trans. Syst. Man Cybern. Part A2
2009 Utility analysis for Internet-oriented server consolidation in VM-based data centers
abstract
Server consolidation based on virtualization technology will simplify system administration, reduce the cost of power and physical infrastructure, and improve utilization in today's Internet-service-oriented enterprise data centers. How much power and how many servers for the underlying physical infrastructure are saved via server consolidation in VM-based data centers is of great interest to administrators and designers of those data centers. Various workload consolidations differ in saving power and physical servers for the infrastructure. The impacts caused by virtualization to those concurrent services are fluctuating considerably which may have a great effect on server consolidation. This paper proposes a utility analytic model for Internet-oriented server consolidation in VM-based data centers, modelling the interaction between server arrival requests with several QoS requirements, and capability flowing amongst concurrent services, based on the queuing theory. According to features of those services' workloads, this model can provide the upper bound of consolidated physical servers needed to guarantee QoS with the same loss probability of requests as in dedicated servers. At the same time, it can also evaluate the server consolidation in terms of power and utility of physical servers. Finally, we verify the model via a case study comprised of one e-book database service and one e-commerce Web service, simulated respectively by TPC-W and SPECweb2005 benchmarks. Our experiments show that the model is simple but accurate enough. The VM-based server consolidation saves up to 50% physical infrastructure, up to 53% power, and improves 1.7 times in CPU resource utilization, without any degradation of concurrent services' performance, running on Rainbow - our virtual computing platform.
Yuzhong Sun, Weisong Shi
CLUSTER4
2009 DiSK: A distributed shared disk cache for HPC environments
abstract
Data movement within high performance environments can be a large bottleneck to the overall performance of programs. With the addition of continuous storage and usage of older data, the back end storage is becoming a larger problem than the improving network and computational nodes. This has led us
Brandon Szeliga, Tung Nguyen 0002, Weisong Shi
CollaborateCom3
2009 Post abstract: Role-based deceptive detection and filtering in WSNs
Shinan Wang, Kewei Sha, Weisong Shi
IPSN3
2009 SensorTrust: a resilient trust model for WSNs
abstract
We present SensorTrust, a trust model to evaluate the trustworthiness of nodes in hierarchical wireless sensor networks, focusing on data integrity.
Guoxing Zhan, Weisong Shi, Hongmei Deng 0001
SenSys2
2009 A survey on dynamic Web content generation and delivery techniques
Jayashree Ravi, Zhifeng Yu, Weisong Shi
J. Netw. Comput. Appl.3
2008 Data Quality and Failures Characterization of Sensing Data in Environmental Applications
Kewei Sha, Guoxing Zhan, Safwan Al-Omari, Tim Calappi, Weisong Shi, Carol J. Miller
CollaborateCom5
2008 Toward low cost and highly reliable sensor networks deployment
abstract
In this paper, we attack the sensor network deployment problem. We define the deployment problem as the problem of deciding how many sensor nodes should be deployed in the sensor field and the number of deployment phases. We model WSN availability and the total cost as functions of the deployment strategy and use our modeling in seeking an optimal deployment strategy that meets user-defined availability requirement with minimum total cost.
Safwan Al-Omari, Weisong Shi
CoNEXT2
2008 Probabilistic Adaptive Anonymous Authentication in Vehicular Networks
Yong Xi, Kewei Sha, Weisong Shi, Loren Schwiebert, Tao Zhang 0005
J. Comput. Sci. Technol.3
2008 Consistency-driven data quality management of networked sensor systems
Kewei Sha, Weisong Shi
J. Parallel Distributed Comput.2
2008 Analysis of ratings on trust inference in open environments
Zhengqiang Liang, Weisong Shi
Perform. Evaluation2
2008 Mobile anonymity of dynamic groups in vehicular networks
abstract
Abstract Vehicular networks are an emerging network to improve safety, efficiency, and convenience of the existing transportation system. In order to preserve privacy in vehicular networks, it is desirable to keep users anonymous. In this paper, we argue that existing approaches, for example, ring singature and anonymous authentication, are susceptible to intersection attack in vehicular networks, resulting in reduced anonymity, even no anonymity. To quantify the achievable anonymity, we first define a graph‐based model called mobile anonymity for a general dynamic environment. We show that under this model, anonymity can be achieved on the system level as long as the constructed anonymity graph is strongly connected. We then extend this model into vehicular networks by applying it to ring signature and anonymous authentication in vehicular networks. We propose two strategies, the random (RND) strategy and the latest‐preferred (LPR) strategy, and evaluate them via simulation using the C3 simulator. The results show that it is promising to apply RND strategy in vehicular networks. Copyright © 2008 John Wiley & Sons, Ltd.
Yong Xi, Weisong Shi, Loren Schwiebert
Secur. Commun. Networks2
2008 Adaptive Secure Access to Remote Services in Mobile Environments
abstract
Since the inception of service-oriented computing paradigm, we have witnessed a plethora of services deployed across a broad spectrum of applications, ranging from conventional RPC-based services to SOAP-based Web services. Likewise, the proliferation of mobile devices has enabled the remote "on the move" access of these services from anywhere at any time. Secure access to these services is challenging especially in a mobile computing environment with heterogeneous modalities. Conventional static access control mechanisms are not able to accommodate complex secure access requirements. In this paper, we propose an adaptive secure access mechanism to address this problem. Our mechanism consists of two components: an adaptive access control module and an adaptive function invocation module. It not only adapts access control policies to diverse requirements, but also introduces function invocation adaptation during access, which is the missing part of existing access control models. We have successfully applied the proposed adaptive secure access mechanism to a computer-assisted surgery application called UbiCAS. Performance evaluation shows that with limited overhead, our technique enforces secure access to the services provided by the UbiCAS system in a flexible way.
Hanping Lufei, Weisong Shi, Vipin Chaudhary
IEEE Trans. Serv. Comput.2
2007 FlexFetch: A History-Aware Scheme for I/O Energy Saving in Mobile Computing
abstract
Extension of battery lifetime has always been a major issue for mobile computing. While more and more data are involved in mobile computing, energy consumption caused by I/O operations becomes increasingly large. In a pervasive computing environment, the requested data can be stored both on the local disk of a mobile computer by using the hoarding technique, and on the remote server, where data are accessible via wireless communication. Based on the current operational states of local disk (active or standby), the amount of data to be requested (small or large), and currently available wireless bandwidth (strong or weak reception), data access source can be adaptively selected to achieve maximum energy reduction. To this end, we propose a profile-based I/O management scheme, FlexFetch, that is aware of access history and adaptive to current access environment. Our simulation experiments driven by real-life traces demonstrate that the scheme can significantly reduce energy consumption in a mobile computer compared with existing representative schemes.
Feng Chen 0005, Song Jiang 0001, Weisong Shi, Weikuan Yu
ICPP3
2007 An Adaptive Rescheduling Strategy for Grid Workflow Applications
abstract
Scheduling is the key to the performance of grid workflow applications. Various strategies are proposed, including static scheduling strategies which map jobs to resources before execution time, or dynamic alternatives which schedule individual job only when it is ready to execute. While sizable work supports the claim that the static scheduling performs better for workflow applications than the dynamic one, it is questioned how a static schedule works effectively in a grid environment which changes constantly. This paper proposes a novel adaptive rescheduling concept, which allows the workflow planner works collaboratively with the run time executor and reschedule in a proactive way had the grid environment changes significantly. An HEFT-based adaptive rescheduling algorithm is presented, evaluated and compared with traditional static and dynamic strategies respectively. The experiment results show that the proposed strategy not only outperforms the dynamic one but also improves over the traditional static one. Furthermore we observed that it performs more efficiently with data intensive application of higher degree of parallelism.
Zhifeng Yu, Weisong Shi
IPDPS2
2007 Enforcing Privacy Using Symmetric Random Key-Set in Vehicular Networks
abstract
Vehicular networks have attracted extensive attentions in recent years for their promises in improving safety and enabling other value-added services. Security and privacy are two integrated issues in the deployment of vehicular networks. Privacy-preserving authentication is a key technique in addressing these two issues. We propose a random keyset based authentication protocol that preserves user privacy under the zero-trust policy, in which no central authority is trusted with the user privacy. We show that the protocol can efficiently authenticate users without compromising their privacy with theoretical analysis. Malicious user identification and key revocation are also described
Yong Xi, Kewei Sha, Weisong Shi, Loren Schwiebert, Tao Zhang 0005
ISADS3
2007 Availability Modeling and Analysis of Autonomous In-Door WSNs
abstract
Availability analysis and modeling in autonomous and remotely administered systems that are composed of cheap and failure-prone components is vital to redundancy management, which includes the prediction of the required number of components and the way these components are scheduled ON and OFF. Targeting the application of wireless sensor networks for the monitoring of elder people living in their apartments, we use techniques from reliability theory to model the WSN as a kappa-out-of m system with independent components. In addition to predicting the required redundancy to meet the desired availability behavior early in the planning phase, we show that scheduling these nodes ON and OFF later on in the operational phase does indeed improve availability over the entire system lifetime. To validate our model, we design a simulator using nesC/TOSSIM. Our analytical and experimental results show that using node scheduling almost doubles the expected WSN total uptime.
Safwan Al-Omari, Weisong Shi
MASS2
2007 Energy-aware QoS for application sessions across multiple protocol domains in mobile computing
Hanping Lufei, Weisong Shi
Comput. Networks2
2006 Redundancy-Aware Topology Management in Wireless Sensor Networks
abstract
Extending the lifetime of wireless sensor networks remains the most challenging and demanding requirement that impedes large-scale deployments. Studies show that considerable energy saving can be achieved only by putting a node's radio into full sleep mode. In this paper we present RAT, which is a redundancy-aware topology management protocol. RAT selects a minimum set of active nodes that are good enough to maintain connectivity, and allows others to sleep and save energy. RAT is designed and implemented with underlying wireless channel irregularity in mind. Scalability and low overhead are the other primary design goals of RAT as well. We implement RAT in the context of Score, which is a cross-layer framework that provides RAT with the neighbor set and allows RAT to coordinate its SLEEP and ACTIVE state changes with the routing layer smoothly. Using TinyOS and PowerTOSSIM, we implement RAT on top of Score. Comparing with the all-active scenario, RAT simulation results show a total energy consumption decrease of 67% in a one-to-many routing scenario and up to 87% in a many-to-one routing scenario
Safwan Al-Omari, Weisong Shi
CollaborateCom2
2006 TRECON: A Framework for Enforcing Trusted ISP Peering
abstract
We envision that neglecting economic factors and trustworthiness evaluation of ISPs is one of the obstacles to developing next generation Internet (NGI). In this paper, we take the initial step to build a general framework called TRECON, which combines an adaptive personalized trust model (alphaPET) with an economic-based approach and provides independent routing among ISPs. With TRECON, autonomous organizations (e.g., ISPs) with varied interests and optimization criteria are smoothly integrated together to achieve better scalability, isolation and self-management. The evaluation results show that the trust-based strategy (TRU) performs much better than the global shortest path routing (SPA) approach in terms of delay, reliability and economic incentives.
Zhengqiang Liang, Weisong Shi
ICCCN2
2006 Shubac: a searchable P2P network utilizing dynamic paths for client/server anonymity
abstract
A general approach to achieve anonymity on P2P networks is to construct an indirect path between client and server for each data transfer. The indirection, together with randomness in the selection of intermediate nodes, provides a guarantee of anonymity to some extent. It, however, comes at the cost of a large communication overhead. In this paper, we present Shubac, a searchable, anonymous peer to peer (P2P) overlay network. It implements a flexible dynamic path approach that shrinks paths in size to reduce overhead and delays and meanwhile reconfigures paths dynamically throughout a communication to maintain a high level of privacy. This dynamic path approach enables Shubac to make a good tradeoff between anonymity and efficiency.
Aharon S. Brodie, Cheng-Zhong Xu 0001, Weisong Shi
IPDPS3
2006 Preserving source location privacy in monitoring-based wireless sensor networks
abstract
While a wireless sensor network is deployed to monitor certain events and pinpoint their locations, the location information is intended only for legitimate users. However, an eavesdropper can monitor the traffic and deduce the approximate location of monitored objects in certain situations. We first describe a successful attack against the flooding-based phantom routing, proposed in the seminal work by Celal Ozturk, Yanyong Zhang, and Wade Trappe. Then, we propose GROW (Greedy Random Walk), a two-way random walk, i.e., from both source and sink, to reduce the chance an eavesdropper can collect the location information. We improve the delivery rate by using local broadcasting and greedy forwarding. Privacy protection is verified under a backtracking attack model. The message delivery time is a little longer than that of the broadcasting-based approach, but it is still acceptable if we consider the enhanced privacy preserving capability of this new approach. At the same time, the energy consumption is less than half the energy consumption of flooding-base phantom routing, which is preferred in a low duty cycle, environmental monitoring sensor network
Yong Xi, Loren Schwiebert, Weisong Shi
IPDPS3
2006 An Application-Aware Event-Oriented MAC Protocol in Multimodality Wireless Sensor Networks
Junzhao Du, Weisong Shi
MSN2
2006 Score: a sensor core framework for cross-layer design
abstract
We present Score, a sensor core framework for cross-layer design in wireless sensor networks. Network components running in the context of Score have the ability to collaborate without the need for pair-wise interfaces. This collaboration promotes protocol optimization in the resource constrained wireless sensor networks, a technique widely known as cross-layer design. We also demonstrate the advantage of Score through three example network components.
Safwan Al-Omari, Junzhao Du, Weisong Shi
QSHINE3
2006 e-QoS: energy-aware QoS for application sessions across multiple protocol domains in mobile computing
abstract
In this paper we propose a novel energy-aware QoS model, e-QoS, for application sessions that might across multiple protocol domains. The model provides the QoS guarantee by dynamically selecting and adapting application protocols. To the best of our knowledge, our model is the first attempt to address QoS adaptation at the application session level by proposing a new QoS metric called session lifetime. To show the effectiveness of the proposed scheme, we have implemented a case study: instant messenger applications between two PocketPCs. Experiment shows that the session lifetime has been successfully extended to the value negotiated by two PocketPCs with very diverse battery capacities.
Hanping Lufei, Weisong Shi
QSHINE2
2006 Fractal: A mobile code-based framework for dynamic application protocol adaptation
Hanping Lufei, Weisong Shi
J. Parallel Distributed Comput.2
2006 Special issue: Security in grid and distributed systems
Weisong Shi, Cheng-Zhong Xu 0001, Xiaobo Zhou 0002
J. Parallel Distributed Comput.1
2006 Performance evaluation of peer-to-peer Web caching systems
Weisong Shi, Yonggen Mao
J. Syst. Softw.1
2005 Performance evaluation of rating aggregation algorithms in reputation systems
abstract
Ratings (also known as recommendations, referrals, and feedbacks) provide an efficient and effective way to build trust relationship amongst peers in open environments. The key to the success of ratings is the rating aggregation algorithm. Several rating aggregation algorithms have been proposed, however, all of them are evaluated in an ad-hoc fashion so that it is difficult to compare the effects of these schemes. In this paper, we argue that what is missing is to evaluate different aggregation schemes in the same context. We first classify all state-of-the-art aggregating algorithms into five categories, and then comprehensively evaluate them in the context of a general decentralized trust inference model with respect to their resistance to different factors, such as dynamic behavior of peers and raters, dishonest ratings, and so on. The simulation results show that complicated algorithms are not always a good choice if we take the implementation cost and resistance to bad raters into consideration.
Zhengqiang Liang, Weisong Shi
CollaborateCom2
2005 Asymmetry-Aware Link Quality Services in Wireless Sensor Networks
Junzhao Du, Weisong Shi, Kewei Sha
EUC2
2005 Secure Application-Aware Service Differentiation in Public Area Wireless Networks
Weisong Shi, Sharun Santhosh, Hanping Lufei
J. Comput. Sci. Technol.1
2005 Accelerating Dynamic Web Content Delivery Using Keyword-based Fragment Detection
Daniel Brodie, Amrish Gupta, Weisong Shi
J. Web Eng.3
2005 Enforcing Cooperative Resource Sharing in Untrusted P2P Computing Environments
Zhengqiang Liang, Weisong Shi
Mob. Networks Appl.2
2004 On the Effects of Bandwidth Reduction Techniques in Distributed Applications
Hanping Lufei, Weisong Shi, Lucia Zamorano
EUC2
2004 Peer-to-Peer Web Caching: Hype or Reality?
Yonggen Mao, Zhaoming Zhu, Weisong Shi
ICPADS3
2004 Cegor: An Adaptive Distributed File System for Heterogeneous Network Environments
Weisong Shi, Sharun Santhosh, Hanping Lufei
ICPADS1
2004 Application-Aware Service Differentiation in PAWNs
abstract
We have witnessed the increasing demand for pervasive Internet access from public area wireless networks (PAWNs). The diverse service requirements from end users necessitate an efficient service differentiation mechanism, which should satisfy two goals: end-user fairness and maximizing the utilization of wireless link. However, we found that the existing best-effort based service model is not enough to satisfy either goal. We have proposed an application-aware service differentiation mechanism which takes both application semantics and user requirements into consideration. The results show that our proposed method outperforms two other bandwidth allocation approaches, best effort and static allocation, in terms of both client fairness and wireless link bandwidth utilization, especially in heavy load environments.
Hanping Lufei, Sivakumar Sellamuthu, Sharun Santhosh, Weisong Shi
ICPP4
2004 Accelerating Dynamic Web Content Delivery Using Keyword-Based Fragment Detection
Daniel Brodie, Amrish Gupta, Weisong Shi
ICWE3
2004 Workload Characterization of Uncacheable HTTP Content
Zhaoming Zhu, Yonggen Mao, Weisong Shi
ICWE3
2004 Revisiting the lifetime of wireless sensor networks
abstract
Prolonging the lifetime of wireless sensor networks (WSN) is one of the most important goals in the sensor network research. A lot of work has been done to achieve this goal; however, current definition of lifetime is either superficial or impractical. In this paper, we take the first step to modeling the lifetime of a wireless sensor network by considering the relationship between the whole sensor network and individual sensors, as well as the importance of different sensors based on their positions. We envision that the proposed lifetime model can be used to evaluate energy-efficient protocols and algorithms, which is validated by simulation results.
Kewei Sha, Weisong Shi
SenSys2
2003 Modeling object characteristics of dynamic Web content
Weisong Shi, Eli Collins, Vijay Karamcheti
J. Parallel Distributed Comput.1
2002 Modeling object characteristics of dynamic Web content
abstract
Although requests for dynamic and personalized content have become a significant part of Internet traffic, there is an absence of (1) good models describing characteristics of dynamic web content, and (2) synthetic content generators, which can be used for simulation-based study of new techniques for serving such content. This paper addresses both of these shortcomings. Its primary contribution is a set of models that capture the characteristics of dynamic content at the sub-document level in terms of independent parameters such as the distributions of object sizes and their freshness times, and derived parameters such as content reusability across time and linked documents. A secondary contribution is a publicly available Java-based dynamic content emulator, DYCE, which uses these models to generate edge side include (ESI) based dynamic content in response to requests for both whole documents and separate objects.
Weisong Shi, Eli Collins, Vijay Karamcheti
GLOBECOM1
2000 Using Confidence Interval to Summarize the Evaluating Results of DSM Systems
Weisong Shi
J. Comput. Sci. Technol.1
1999 Write Detection in Home-Based Software DSMs
Weiwu Hu, Weisong Shi
Euro-Par2
1999 Affinity-Based Self Scheduing for Software Shared Memory Systems
Weisong Shi
HiPC1
1999 Adaptive Write Detection in Home-based Software DSMs
abstract
Write detection is essential in multiple-writer protocols to identify writes to shared pages so that these writes can be correctly propagated. Software DSMs that implement multiple-writer protocol normally employ the virtual memory page fault to detect writes to shared pages. It write-protects shared pages at the beginning of an interval to detect writes of the interval. This paper proposes a new write detection scheme in a home-based software DSM called JIAJIA. It automatically recognizes single write to a shared page by its home host and assumes the page will continue to be written by the home host in the future until the page is written by remote hosts. During the period the page is assumed to be singly written by its home host, no write detection of this home page is required and page faults caused by home host write detection call are avoided. Evaluation with some well-known DSM benchmarks reveals that the new write detection can reduce page faults dramatically and improve performance significantly.
Weiwu Hu, Weisong Shi
HPDC2
1999 Dynamic Task Migration in Home-based Software DSM Systems
abstract
Dynamic task migration is an effective strategy to maximize the performance and resource utilization in metacomputing environments. Traditionally, however, a "task" means only the corresponding code of computation, i.e., the data related to this computation is usually neglected. As such, when one task is migrated from processor A to processor B, the data required by this task remains on processor A. Thus, the processor B has to perform remote communication when it executes this task, eliminating the advantage of task migration, or even further degrading the performance. Hence, the definition of the traditional "task" should be revisited. We define a task as follows: Task=Computation subtask+Data subtask. Computation subtask is the program code to be executed, while the data subtask is the operations to access the related data located in memory. In fact, as the speed gap between processors and memory becomes larger and larger, the importance of data subtask becomes more obvious than before. Therefore, we argue that both subtasks should be migrated to a new processor during task migration. Based on this observation, we propose a dynamic loop-level task migration scheme. This scheme is implemented within the context of the JIAJIA software DSM system (W. Ha et al., 1999). The evaluation results show that the task migration scheme improves the performance of our benchmark applications by 36% to 50% compared with static task allocation schemes. As a result, the new scheme performs an average of 30% better than other computation-only migration schemes.
Weisong Shi, Weiwu Hu, M. Rasit Eskicioglu
HPDC1
1999 Where does the time go in software DSMs? - Experiences with JIAJIA
Weisong Shi, Weiwu Hu
J. Comput. Sci. Technol.1
1998 A framework of memory consistency models
Weiwu Hu, Weisong Shi
J. Comput. Sci. Technol.2
1998 A lock-based cache coherence protocol for scope consistency
Weiwu Hu, Weisong Shi
J. Comput. Sci. Technol.2