VLDB 2026 Research / reviewers in the wild / expert
Ragunathan Rajkumar
dblp:16/36 · also Ragunathan Raj Rajkumar, Raj Rajkumar
· DBLP profile ↗
152ranked-venue papers
10as first author
14since 2021 · last 2026
0000-0002-0983-7259ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 50 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 40 · 6 first-author · 1 since 2021Artificial intelligence and machine learning · 15 · 8 since 2021Computer networks · 12 · 1 first-authorHuman-computer interaction and ubiquitous computing · 4 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 3Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Security and privacy · 2Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
39 papers |
Embedded and real-time systems · 63% Distributed systems · 10% Energy-efficient computing · 10% | |
| Artificial intelligence
4 papers |
Autonomous driving · 42% 3D vision · 29% Graph learning · 15% | |
| Computer networks
12 papers |
Internet of things and sensor networks · 56% Vehicular, aerial and satellite networks · 18% Network optimization and economics · 16% | |
| Software engineering, system software, and programming languages
4 papers |
Software testing · 53% Operating systems · 47% |
Topics — the 30 heaviest of 103, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Embedded and real-time systems
real-time scheduling |
1.6 | 23 | 2018 | CycleTandem: Energy-Saving Scheduling for Real-Time Systems with Hardware Accelerators · RTSS 2018 vMPCP: A Synchronization Framework for Multi-core Virtual Machines · RTSS 2014 Segment-Fixed Priority Scheduling for Self-Suspending Real-Time Tasks · RTSS 2013 |
Embedded and real-time systems › real-time scheduling
schedulability analysis |
0.6 | 7 | 2014 | vMPCP: A Synchronization Framework for Multi-core Virtual Machines · RTSS 2014 Scheduling Parallel Real-Time Tasks on Multi-core Processors · RTSS 2010 On the Scheduling of Mixed-Criticality Real-Time Task Sets · RTSS 2009 |
Internet of things and sensor networks
wireless sensor network |
0.5 | 6 | 2014 | Poster abstract: a harmony of sensors: achieving determinism in multi-application sensor networks · IPSN 2014 Low-power clock synchronization using electromagnetic energy radiating from AC power lines · SenSys 2009 FireFly Mosaic: A Vision-Enabled Wireless Sensor Networking System · RTSS 2007 |
Computer vision › 3D vision
3d object detection |
0.4 | 1 | 2020 | Point-GNN: Graph Neural Network for 3D Object Detection in a Point Cloud · CVPR 2020 |
Machine learning › Graph learning
graph neural network |
0.4 | 1 | 2020 | Point-GNN: Graph Neural Network for 3D Object Detection in a Point Cloud · CVPR 2020 |
Computer vision › 3D vision › 3d object detection › point cloud object detection
LiDAR-based 3D object detection |
0.4 | 1 | 2020 | Point-GNN: Graph Neural Network for 3D Object Detection in a Point Cloud · CVPR 2020 |
Embedded and real-time systems
cyber-physical system platforms |
0.4 | 2 | 2018 | Tools and Methodologies for Autonomous Driving Systems · Proc. IEEE 2018 Low-power clock synchronization using electromagnetic energy radiating from AC power lines · SenSys 2009 |
Robotics › Autonomous driving
perception |
0.4 | 3 | 2020 | A multi-sensor fusion system for moving object detection and tracking in urban driving environments · ICRA 2014 Point-GNN: Graph Neural Network for 3D Object Detection in a Point Cloud · CVPR 2020 Self-Driving Vehicles: The Challenges and Opportunities Ahead · SenSys 2015 |
Parallel and multicore computing › task scheduling
accelerator scheduling |
0.3 | 1 | 2018 | CycleTandem: Energy-Saving Scheduling for Real-Time Systems with Hardware Accelerators · RTSS 2018 |
Distributed systems
clock synchronization |
0.3 | 2 | 2016 | Timeline: An Operating System Abstraction for Time-Aware Applications · RTSS 2016 Monitoring Timing Constraints in Distributed Real-Time Systems · RTSS 1992 |
Embedded and real-time systems
cyber-physical systems |
0.3 | 2 | 2012 | A Cyber-Physical Future · Proc. IEEE 2012 Cyber-physical systems: the next computing revolution · DAC 2010 |
GPUs and heterogeneous computing
GPU computing |
0.2 | 2 | 2018 | RGEM: A Responsive GPGPU Execution Model for Runtime Engines · RTSS 2011 CycleTandem: Energy-Saving Scheduling for Real-Time Systems with Hardware Accelerators · RTSS 2018 |
Distributed systems
resource sharing |
0.2 | 2 | 2014 | vMPCP: A Synchronization Framework for Multi-core Virtual Machines · RTSS 2014 Resource Sharing in Reservation-Based Systems · RTSS 2001 |
Embedded and real-time systems › real-time scheduling
multiprocessor scheduling |
0.2 | 2 | 2010 | Scheduling Parallel Real-Time Tasks on Multi-core Processors · RTSS 2010 Coordinated Task Scheduling, Allocation and Synchronization on Multiprocessors · RTSS 2009 |
Robotics › Autonomous driving
perception and planning |
0.2 | 1 | 2023 | Keynote: Rising to the Challenge of Autonomous Vehicles · PERCOM 2023 |
Computer vision › Video understanding and tracking › object tracking
moving object detection and tracking |
0.2 | 1 | 2014 | A multi-sensor fusion system for moving object detection and tracking in urban driving environments · ICRA 2014 |
Robotics › Robot navigation and mapping
sensor fusion |
0.2 | 1 | 2014 | A multi-sensor fusion system for moving object detection and tracking in urban driving environments · ICRA 2014 |
Embedded and real-time systems
real-time virtualization |
0.2 | 1 | 2014 | vMPCP: A Synchronization Framework for Multi-core Virtual Machines · RTSS 2014 |
Embedded and real-time systems
real-time operating systems |
0.2 | 4 | 2007 | FireFly Mosaic: A Vision-Enabled Wireless Sensor Networking System · RTSS 2007 Nano-RK: An Energy-Aware Resource-Centric RTOS for Sensor Networks · RTSS 2005 Resource Sharing in Reservation-Based Systems · RTSS 2001 |
Distributed systems
fault tolerance |
0.2 | 4 | 2012 | SAFER: System-level Architecture for Failure Evasion in Real-time Applications · RTSS 2012 High availability in the real-time publisher/subscriber inter-process communication model · RTSS 1996 Monitoring Timing Constraints in Distributed Real-Time Systems · RTSS 1992 |
Embedded and real-time systems
distributed real-time systems |
0.2 | 3 | 2012 | SAFER: System-level Architecture for Failure Evasion in Real-time Applications · RTSS 2012 High availability in the real-time publisher/subscriber inter-process communication model · RTSS 1996 Monitoring Timing Constraints in Distributed Real-Time Systems · RTSS 1992 |
Embedded and real-time systems › real-time scheduling
resource reservation |
0.2 | 4 | 2005 | Multi-Granularity Resource Reservations · RTSS 2005 Nano-RK: An Energy-Aware Resource-Centric RTOS for Sensor Networks · RTSS 2005 Resource Sharing in Reservation-Based Systems · RTSS 2001 |
Embedded and real-time systems › real-time scheduling › schedulability analysis
self-suspending tasks |
0.2 | 1 | 2013 | Segment-Fixed Priority Scheduling for Self-Suspending Real-Time Tasks · RTSS 2013 |
Network optimization and economics › resource allocation
qos-aware resource allocation |
0.1 | 1 | 2012 | QoS-Based Resource Allocation for Next-Generation Spacecraft Networks · RTSS 2012 |
Network optimization and economics
resource allocation |
0.1 | 1 | 2012 | QoS-Based Resource Allocation for Next-Generation Spacecraft Networks · RTSS 2012 |
Vehicular, aerial and satellite networks
satellite networks |
0.1 | 1 | 2012 | QoS-Based Resource Allocation for Next-Generation Spacecraft Networks · RTSS 2012 |
GPUs and heterogeneous computing
GPU scheduling |
0.1 | 1 | 2011 | TimeGraph: GPU Scheduling for Real-Time Multi-Tasking Environments · USENIX ATC 2011 |
Embedded and real-time systems › real-time scheduling › parallel real-time task scheduling
real-time GPU scheduling |
0.1 | 1 | 2011 | TimeGraph: GPU Scheduling for Real-Time Multi-Tasking Environments · USENIX ATC 2011 |
Internet of things and sensor networks
time synchronization |
0.1 | 2 | 2009 | Low-power clock synchronization using electromagnetic energy radiating from AC power lines · SenSys 2009 FireFly Mosaic: A Vision-Enabled Wireless Sensor Networking System · RTSS 2007 |
Internet of things and sensor networks › neighbor discovery
asynchronous neighbor discovery |
0.1 | 1 | 2010 | U-connect: a low-latency energy-efficient asynchronous neighbor discovery protocol · IPSN 2010 |
Methods — techniques the papers use, named apart from their topics
model-based design · 1.1reference architecture · 1.0clock synchronization protocols · 0.5graph neural network · 0.4box merging · 0.4auto-registration · 0.4static frequency scaling · 0.3partitioning · 0.3dynamic graph decomposition · 0.3Q-RAM extension · 0.3clock synchronization protocol · 0.2sensor infrastructure · 0.2synchronization protocol · 0.2sensor fusion · 0.2schedulability analysis · 0.2kalman filtering · 0.2utilization bound · 0.2response time analysis · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On V2X Communications for Autonomous Navigation of Work Zones and Road Incidents
Gregory T. Su, Nishad Sahu, Shounak Sural, Ragunathan Rajkumar |
IV | 5 |
| 2026 | WorkZone3D: A Multimodal Dataset for 3D Work Zone Perception in Autonomous DrivingabstractWork zones are essential to maintain, repair and upgrade our roadways. However, they introduce complex, dynamic and challenging environments for autonomous vehicles to navigate safely. To help address this challenge, we introduce the first publicly available, large-scale, multimodal work zone dataset collected with an autonomous vehicle consisting of multiple synchronized lidars and high-resolution cameras. Our WorkZone3D dataset covers various work zone elements such as cones, barrels, and channelizers, and provides 3D annotation boxes for these objects. We also propose a auto-annotation pipeline that can produce high-quality 3D labels, assimilating data across frames, even for rare classes which do not have pre-trained 3D object detection models to start with. We evaluate unimodal and multimodal models on our dataset, showing the critical role of sensor fusion in accurate 3D localization of such small objects at a distance, often having very few lidar points on them. Our results demonstrate the usefulness of WorkZone3D for generalization to real-world scenarios. Our code and dataset are available at https://github.com/ssuralcmu/WorkZone3D.git. Shounak Sural, Nishad Sahu, Ragunathan Rajkumar |
WACV | 3 |
| 2025 | Towards the Safe Operation of Autonomous Vehicles in Work ZonesabstractAutonomous vehicles (AVs) promise significant advances in transportation safety and efficiency. However, navigating roadway work zones, which can be rather complex and dynamic, remains a significant challenge. This paper presents the results of a large study that addresses the challenges, requirements, solutions and practical experiences of AVs driving safely through work zones. We begin by proposing a taxonomy of work zone scenarios and analyzing their structures and attributes. We next discuss the perception, routing, behavioral and path-planning requirements for AVs to safely navigate these scenarios. We then offer methods to meet these requirements and investigate the impact of range and AV speed on perception confidence levels for work zone detection. We evaluate our solutions in a co-simulation environment, on a closed test-track and on public roadways across more than 20 work zone scenarios specified by the Pennsylvania Department of Transportation (PennDoT). Video demonstrations illustrate the feasibility of safe and reliable navigation of AVs in a wide variety of work zones. Nishad Sahu, Gregory T. Su, Shounak Sural, Ragunathan Rajkumar |
IV | 5 |
| 2025 | Enhanced Safety Messages (ESM): A Practical Alternative to V2X Basic Safety MessagesabstractA very important use case of V2X communications is the enhancement of roadway safety by utilizing the transmission of road information among vehicles. The SAE Basic Safety Message (BSM) is the most common standard used to transmit road event information and their locations based on global latitudinal and longitudinal coordinates of transmitting vehicles. In practice, however, global coordinate estimations are inherently limited by the accuracy of Global Navigation Satellite Systems (GNSS) such as GPS. GNSS signals can also be unavailable in urban canyons and tunnels, be spoofed to force incorrect localization, or be restricted to low accuracy due to the sparsity of available groundbased corrections (such as RTK base stations) in rural and remote areas. In this paper, we introduce Enhanced Safety Messages (ESMs), a backward-compatible BSM replacement that adopts the well-established foundations of information redundancy in safety-critical systems to avoid catastrophic failure by providing both absolute and relative coordinate frames to robustly describe vehicle locations. This position information redundancy in two different coordinate systems, one from external sources and another from local sensing, effectively addresses the BSM drawbacks of relying entirely on GNSS signals. Specifically, ESM includes LaneContext and MapContext to address two of the most common driving environments of open spaces and urban roadways. Location communication over ESM is enhanced by additionally specifying the connected vehicle's driving lane, its offset from the center of the lane, and its longitudinal position along the road segment. Our experimental evaluation confirms that ESM's inclusion of relative positioning rectifies the core BSM weaknesses in relying only on global GNSS coordinates. Gregory T. Su, Ragunathan Rajkumar |
VTC2025-Spring | 2 |
| 2024 | ContextualFusion: Context-Based Multi-Sensor Fusion for 3D Object Detection in Adverse Operating ConditionsabstractThe fusion of multimodal sensor data streams such as camera images and lidar point clouds plays an important role in the operation of autonomous vehicles (AVs). Robust perception across a range of adverse weather and lighting conditions is specifically required for AVs to be deployed widely. While multi-sensor fusion networks have been previously developed for perception in sunny and clear weather conditions, these methods show a significant degradation in performance under night-time and poor weather conditions. In this paper, we propose a simple yet effective technique called ContextualFusion to incorporate the domain knowledge about cameras and lidars behaving differently across lighting and weather variations into 3D object detection models. Specifically, we design a Gated Convolutional Fusion (GatedConv) approach for the fusion of sensor streams based on the operational context. To aid in our evaluation, we use the open-source simulator CARLA to create a multimodal adverse-condition dataset called AdverseOp3D to address the shortcomings of existing datasets being biased towards daytime and good-weather conditions. Our ContextualFusion approach yields an mAP improvement of 6.2% over state-of-the-art methods on our context-balanced synthetic dataset. Finally, our method enhances state-of-the-art 3D objection performance at night on the real-world NuScenes dataset with a significant mAP improvement of 11.7%. Shounak Sural, Nishad Sahu, Ragunathan Rajkumar |
IV | 3 |
| 2024 | MPMP: A Protocol to Transmit Long Messages for V2X ApplicationsabstractAutonomous vehicles (AVs) can operate more safely with vehicular communication technology. A Vehicle-to-Everything (V2X) communication application domain of particular interest is work zones, which may require significant amounts of information to fully describe their shapes and contents. However, existing V2X protocols are limited in how much data can be transmitted in a single packet. In this paper, we propose and evaluate the Multi-Packet Memo Protocol (MPMP), a V2X communication protocol for broadcasting long messages, called memos. MPMP can transmit up to 1 MB per memo and correctly assemble packets received in an out-of-order sequence. Multiple receivers can simultaneously begin receiving packets at different points in the sequence. Latency and reliability measurements show that long messages can be transmitted reliably with acceptable delays for safe CAV operations in work zones. Gregory T. Su, Ragunathan Rajkumar |
VTC Fall | 2 |
| 2023 | Keynote: Rising to the Challenge of Autonomous VehiclesabstractAutonomous vehicles (AVs) have garnered immense interest and investments for more than a decade and a half. Nevertheless, large-scale AV deployments do not seem viable in the near future. In this talk, the speaker will address questions like “What went wrong?”, “Is AI the answer?”, “Can (and how do) we course-correct?” and “Which future contributions will matter?”. Finally, challenges that must be addressed by the research and engineering communities will be discussed. Ragunathan Rajkumar |
PERCOM | 1 |
| 2022 | A-DRIVE: Autonomous Deadlock Detection and Recovery at Road Intersections for Connected and Automated VehiclesabstractConnected and Automated Vehicles (CAVs) are highly expected to improve traffic throughput and safety at road intersections, single-track lanes, and construction zones. However, multiple CAVs can block each other and create a mutual deadlock around these road segments (i) when vehicle systems have a failure, such as a communication failure, control failure, or localization failure and/or (ii) when vehicles use a long shared road segment. In this paper, we present an Autonomous Deadlock Detection and Recovery Protocol at Intersections for Automated Vehicles named A-DRIVE that is a decentralized and time-sensitive technique to improve traffic throughput and shorten worst-case recovery time. To enable the deadlock recovery with automated vehicles and with human-driven vehicles, A-DRIVE includes two components: V2V communication-based A-DRIVE and Local perception-based A-DRIVE. V2V communication-based A-DRIVE is designed for homogeneous traffic environments in which all the vehicles are connected and automated. Local perception-based A-DRIVE is for mixed traffic, where CAVs, non-connected automated vehicles, and human-driven vehicles co-exist and cooperate with one another. Since these two components are not exclusive, CAVs inclusively and seamlessly use them in practice. Finally, our simulation results show that A-DRIVE improves traffic throughput compared to a baseline protocol. Shunsuke Aoki 0001, Ragunathan Rajkumar |
IV | 2 |
| 2022 | FT-DeepNets: Fault-Tolerant Convolutional Neural Networks with Kernel-based DuplicationabstractDeep neural network (deepnet) applications play a crucial role in safety-critical systems such as autonomous vehicles (AVs). An AV must drive safely towards its destination, avoiding obstacles, and respond quickly when the vehicle must stop. Any transient errors in software calculations or hardware memory in these deepnet applications can potentially lead to dramatically incorrect results. Therefore, assessing and mitigating any transient errors and providing robust results are important for safety-critical systems. Previous research on this subject focused on detecting errors and then recovering from the errors by re-running the network. Other approaches were based on the extent of full network duplication such as the ensemble learning-based approach to boost system fault-tolerance by leveraging each model’s advantages. However, it is hard to detect errors in a deep neural network, and the computational overhead of full redundancy can be substantial.We first study the impact of the error types and locations in deepnets. We next focus on selecting which part should be duplicated using multiple ranking methods to measure the order of importance among neurons. We find that the duplication overhead for computation and memory is a trade-off between algorithmic performance and robustness. To achieve higher robustness with less system overhead, we present two error protection mechanisms that only duplicate parts of the network from critical neurons. Finally, we substantiate the practical feasibility of our approach and evaluate the improvement in the accuracy of a deepnet in the presence of errors. We demonstrate these results using a case study with real-world applications on an Nvidia GeForce RTX 2070Ti GPU and an Nvidia Xavier embedded platform used by automotive OEMs. Iljoo Baek, Wei Chen 0124, Soheil Samii, Ragunathan Rajkumar |
WACV | 5 |
| 2022 | Cyber Traffic Light: Safe Cooperation for Autonomous Vehicles at Dynamic IntersectionsabstractAutonomous driving systems are becoming increasingly feasible and highly expected to be the heart of intelligent transportation systems. To deploy the autonomous driving vehicles on public roads, one of the practical challenges might be safe cooperation and collaboration among multiple vehicles, in particular when conflicts arise on shared road segments, such as road intersections, merge points, construction zones, single-track lanes, and center turn lane. In the current traffic systems, human drivers navigate these regions using a combination of traffic rules, social norms, courtesy, hand signals, and common sense. In this paper, we identify and classify such Dynamic Intersections that might lead to vehicle accidents and/or deadlocks and that might appear almost anytime and anywhere on public roads. In addition, we present a cooperative dynamic intersection protocol that uses on-board perception systems and vehicular communications for peer-to-peer negotiation. Under our protocol, autonomous driving vehicles can create a vehicular communication-based traffic manager named Cyber Traffic Light when congestion arises. Cyber Traffic Light works as a self-organizing, self-planning, and self-optimizing traffic manager, and it allocates the green period for vehicles coming from the multiple directions. Finally, we showed that our decentralized protocol has much higher traffic throughput, compared to two simple protocols while guaranteeing road safety. Shunsuke Aoki 0001, Ragunathan Rajkumar |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Extended VINS-Mono: A Systematic Approach for Absolute and Relative Vehicle Localization in Large-Scale Outdoor EnvironmentsabstractWe present a systematic approach called Extended VINS-Mono to utilize VINS-Mono, a state-of-the-art monocular visual-inertial relative localization method, targeting practical vehicle localization in large-scale outdoor road environments. Our proposed fusion approach associates multiple independent localization methods and provides multiple (projected) state estimates in a desired coordinate system to satisfy different accuracy, rate and latency requirements. We extend VINS-Mono with absolute localization methods like GNSS and relative localization methods like Kalman-filter-based INS to provide global state estimation for navigation/routing and local state estimation for planning/control. Additionally, Extended VINS-Mono addresses two significant drawbacks in VINS-Mono for use in large-scale outdoor road environments. First, motion on an almost planar road surface will make scale unobservable in VINS-Mono. Secondly, moving objects in dynamic scenarios will degrade accuracy. We handle the scale estimation problem of VINS-Mono by extending its (re-)initialization process with speed readings and introducing a speed factor for use with graph optimization. A dynamic feature-point filter method with masks from DNN-based object detection handles dynamic environments and re-collects feature points on stationary objects like parked cars. Better global accuracy is obtained with Extended VINS-Mono, compared to VINS-Mono, in a 25 km-trip journey through highways, tunnels, urban areas and suburban areas in Pittsburgh. Thus, Extended VINS-Mono can be used for reliable and accurate absolute localization in dynamic road environments. We also evaluate the accuracy, localization rate and latency of multiple (projected) state estimates in the global coordinate system from multiple localization methods. Our fusion method is therefore able to satisfy different localization requirements of various tasks on an intelligent vehicle. Mengwen He, Ragunathan Rajkumar |
IROS | 2 |
| 2021 | MultiCruise: Eco-Lane Selection Strategy with Eco-Cruise Control for Connected and Automated VehiclesabstractConnected and Automated Vehicles (CAVs) have real-time information from the surrounding environment by using local on-board sensors, V2X (Vehicle-to-Everything) communications, pre-loaded vehicle-specific lookup tables, and map database. CAVs are capable of improving energy efficiency by incorporating these information. In particular, Eco-Cruise and Eco-Lane Selection on highways and/or motorways have immense potential to save energy, because there are generally fewer traffic controllers and the vehicles keep moving in general. In this paper, we present a cooperative and energy-efficient lane-selection strategy named MultiCruise, where each CAV selects one among multiple candidate lanes that allows the most energy-efficient travel. MultiCruise incorporates an Eco-Cruise component to select the most energy-efficient lane. The Eco-Cruise component calculates the driving parameters and prospective energy consumption of the ego vehicle for each candidate lane, and the Eco-Lane Selection component uses these values. As a result, MultiCruise can account for multiple data sources, such as the road curvature and the surrounding vehicles' velocities and accelerations. The eco-autonomous driving strategy, MultiCruise, is tested, designed and verified by using a co-simulation test platform that includes autonomous driving software and realistic road networks to study the performance under realistic driving conditions. Our experimental evaluations show that our eco-autonomous Mul-tiCruise saves up to 8.5% fuel consumption. Shunsuke Aoki 0001, Lung En Jan, Junfeng Zhao 0006, Anand Bhat, Chen-Fang Chang, Ragunathan Rajkumar |
IV | 6 |
| 2021 | Using Thermal Vision for Extended VINS-Mono to Localize Vehicles in Large-Scale Outdoor Road EnvironmentsabstractA monocular VIO (Visual-Inertial Odometry) system provides a compact, low-cost, and easily-deployed configuration for relative localization. However, using thermal vision for VIO is much less studied than using a visible-spectrum camera. A thermal-vision camera works under all lighting conditions and has been used to detect pedestrians, cars and animals at nighttime to provide ADAS (Advanced Driver Assistance System) functions. Common problems in directly using thermal images in conventional VIO methods are: (1) lower signal-to-noise ratio and fewer reliable feature points in the texture-less thermal images, (2) periodic recalibration hampers thermal image capturing and feature-point tracking, and (3) the “jello” effect of the rolling shutter readout architecture is sensitive to aggressive vehicle maneuvers. Extended VINS-Mono, proposed in our previous work [1], aims at providing relative and absolute localization of a vehicle in large-scale outdoor road environments by introducing (1) absolute localization methods to enable VINS-Mono to output local and global state estimates simultaneously, (2) vehicle speed readings for fast (re-)initialization and reliable scale estimates, and (3) DNN-based object detection methods to remove nonstationary objects from the visible scene. In this paper, we show that Extended VINS-Mono can use thermal images to provide relative and absolute localization even when light conditions are very poor. We conducted several experiments on a 25 Km-trip journey through highways, tunnels, urban areas and suburban areas in Pittsburgh, USA during daytime and nighttime to evaluate the performance of Extended Thermal VINS-Mono, including (re-)initialization, accuracy, rate, and latency. Our evaluation confirms that using thermal vision for localization satisfies localization requirements in large-scale outdoor road environments when the visible-spectrum camera performs poorly. Mengwen He, Ragunathan Rajkumar |
IV | 2 |
| 2021 | Opportunities and Challenges for Flagman Recognition in Autonomous VehiclesabstractAutonomous vehicles promise significant advances in transportation safety, efficiency and comfort. However, achieving the goal of full autonomy is impeded by the need to address several operational challenges encountered in practice. Gesture recognition of flagmen on roads is one such set of challenges. An autonomous vehicle needs to make safe decisions and facilitate forward progress in the presence of road construction workers and flagmen. However, human gestures under diverse environmental conditions are very varied and represent significant complexity. In this work, we present (i) a taxonomy of challenges for organizing traffic gestures, (ii) a sizeable flagman gesture dataset, and (iii) extensive experiments on practical algorithms for gesture recognition. We categorize traffic gestures according to their semantics, flagman appearances and the environmental context. We then collect a dataset covering a range of common flagman gestures with and without props such as signs and flags. Finally, we develop a recognition algorithm using different feature representations of the human pose and perform extensive ablation experiments on each component. Weijing Shi, Ragunathan Rajkumar, Eran Kishon |
IV | 2 |
| 2020 | Point-GNN: Graph Neural Network for 3D Object Detection in a Point CloudabstractIn this paper, we propose a graph neural network to detect objects from a LiDAR point cloud. Towards this end, we encode the point cloud efficiently in a fixed radius near-neighbors graph. We design a graph neural network, named Point-GNN, to predict the category and shape of the object that each vertex in the graph belongs to. In Point-GNN, we propose an auto-registration mechanism to reduce translation variance, and also design a box merging and scoring operation to combine detections from multiple vertices accurately. Our experiments on the KITTI benchmark show the proposed approach achieves leading accuracy using the point cloud alone and can even surpass fusion-based algorithms. Our results demonstrate the potential of using the graph neural network as a new approach for 3D object detection. The code is available at https://github.com/WeijingShi/Point-GNN. Weijing Shi, Ragunathan Rajkumar |
CVPR | 2 |
| 2020 | Co-simulation Platform for Developing InfoRich Energy-Efficient Connected and Automated VehiclesabstractWith advances in sensing, computing and communication technologies, Connected and Automated Vehicles (CAVs) are becoming feasible. The advent of CAVs presents new opportunities to improve the energy efficiency of individual vehicles. However, testing and verifying energy-efficient autonomous driving systems are difficult due to safety considerations and repeatability. In this paper, we present a co-simulation platform to develop and test novel vehicle eco-autonomous driving technologies named InfoRich, which incorporates the information from on-board sensors, V2X communications, and map database. The co-simulation platform includes eco-autonomous driving software, vehicle dynamics and powertrain (VD&PT) model, and a traffic environment simulator. Also, we utilize synthetic drive cycles derived from real-world driving data to test the strategies under realistic driving scenarios. To build road networks from the real-world driving data, we develop an Automated Parser and Calculator for Map/Scenario named AutoPASCAL. Overall, the simulation platform provides a realistic vehicle model, powertrain model, sensor model, traffic model, and road-network model to enable the evaluation of the energy efficiency of eco-autonomous driving. Shunsuke Aoki 0001, Lung En Jan, Junfeng Zhao 0006, Anand Bhat, Ragunathan Rajkumar, Chen-Fang Chang |
IV | 5 |
| 2020 | CARSS: Client-Aware Resource Sharing and Scheduling for Heterogeneous ApplicationsabstractModern hardware accelerators such as GP-GPUs and DSPs are commonly being used in real-time settings such as high-performance multimedia systems and autonomous vehicles. In fact, the throughput of a wide variety of computationally demanding tasks from 3D graphics and rendering to image processing and deep learning can benefit from such specialized hardware. Such heterogeneity can affect the performance of applications running simultaneously on the same accelerator. Prior studies on resource sharing and scheduling on hardware accelerators have not attempted to account for this context. In this work, we provide a portable tagging-based cooperative scheduler and resource monitor for use by heterogeneous applications sharing a single hardware accelerator in a soft real-time environment. We also offer practical insight into how various types of applications use the hardware accelerators differently. We substantiate the feasibility of our approach and evaluate the improvement of various scheduling policies over a proprietary scheduler in several case-studies with real-world applications on 2 NVIDIA platforms: a GeForce GTX 1070 GPU and an Xavier embedded platform1. Although we focus on GPUs in this paper, our underlying observations and framework can also be used for sharing execution on other types of hardware accelerators.1The video demo has been uploaded to https://youtu.be/pziS1btsr9c Iljoo Baek, Matthew Harding, Akshit Kanda, Kyung Ryeol Choi, Soheil Samii, Ragunathan Rajkumar |
RTAS | 6 |
| 2020 | Student Session: Practical Insights on Acceleration for 3D Lidar Data Processingabstract3D Lidar has become a widely used sensor technology in autonomous vehicles by providing accurate distance information. However, lidar pointcloud processing often involves sophisticated algorithms, and takes a lot of computational power. Many prior approaches relied on a GPU-based parallel programming model, such as CUDA, to accelerate these computations. However, little attention has been given to comparing different methods for selecting the most-suited programming and parallelization approaches for a given computing system. We present our findings and insights identified by implementing various parallel approaches considering both CPUs and GPUs. We also demonstrate significant acceleration results using a real-world perception algorithm developed to detect road boundaries. Finally, we compare the pros and cons of each method in terms of system architecture, programming model, and resource utilization to yield a better understanding of choosing the best parallelization approach for a given optimization objective. Iljoo Baek, Kamal Fuseini, Ragunathan Rajkumar |
RTCSA | 3 |
| 2020 | Error Vulnerabilities and Fault Recovery in Deep-Learning Frameworks for Hardware AcceleratorsabstractHardware accelerators such as GP-GPUs, Tensor Cores, and Deep-Learning Accelerators (DLA) are increasingly being used in real-time settings such as autonomous vehicles (AVs). In such deployments, any software errors and process failures in hardware systems can lead to critical faults in AVs. Therefore, assessing and mitigating hardware accelerator faults are critical requirements for safety-critical systems. Past work on this subject focused on simulated and injected software and hardware faults to understand and analyze the behavior of the software stack and the entire system. However, programming errors and process failures caused when using software frameworks must also be considered. In this paper, we present experiments which show that widely used deep-learning frameworks are vulnerable to programming mistakes and errors. We first focus on memory-related programming errors caused by applications using deep-learning frameworks that facilitate high-performance inferencing. We next find that a reset to recover from any fault imposes significant time penalties in reloading a pre-trained deep neural network model. To reduce these fault recovery times, we propose fault recovery mechanisms that checkpoint and resume the network based on the inference stage when an error is detected. Finally, we substantiate the practical feasibility of our approach and evaluate the improvement in recovery times11A demo video clip demonstrating our recovery algorithm has been uploaded to Youtube: https://www.youtube.com/watch?v=xwUYdJdA5oM.. We use a case-study with real-world applications on an Nvidia GeForce GTX 1070 GPU and an Nvidia Xavier embedded platform, which is commonly used by multiple automotive OEMs. Iljoo Baek, Sourav Panda, Nandha Kishore Srinivasan, Soheil Samii, Ragunathan Rajkumar |
RTCSA | 6 |
| 2020 | Fault-Tolerance Support for Adaptive AUTOSAR Platforms using SOME/IPabstractModern automobiles with driving-assist features are inherently safety-critical. Strict safety requirements and the introduction of self-driving capabilities have increased the demands on the computing and communication systems within automobiles. The AUTomotive Open System ARchitecture (AUTOSAR) Adaptive platform aims to meet these industry requirements by supporting high-performance computing devices and high-bandwidth communication technologies. The AUTOSAR Adaptive platform leverages Scalable service-Oriented MiddlewarE over IP (SOME/IP), an automotive middleware solution, that supports the exchange of control messages across various devices of different sizes and operating systems. Typically, in order to guarantee the safe execution of software, automobiles employ redundancy for crucial software tasks to tolerate permanent crash faults. The Adaptive AUTOSAR standard does not specify any fault-tolerance requirements. In this paper, we highlight some gaps in the current AUTOSAR Adaptive Platfrom standard (version 18.10) and provide suggestions to address them. We present our framework to support fault-tolerant execution using different replication strategies for the AUTOSAR Adaptive Platform. We analyze the fault detection and recovery-time bounds of our solution for applications using SOME/IP. We validate our model experimentally and present our evaluation results. Anand Bhat, Soheil Samii, Ragunathan Rajkumar |
RTCSA | 3 |
| 2019 | Thin-Plate Spline-based Adaptive 3D Surround ViewabstractA “Bird's Eye View” (or Surround View) is a popular feature in modern cars and is particularly useful for parking and unparking purposes. This 3D reconstruction of the view uses multiple fisheye cameras and has traditionally been performed by a computer vision-based approach on a fixed mesh. This method is computationally expensive and does not always give ideal results for different types of surroundings. This paper discusses the design and implementation of a Thin-Plate Spline (TPS) algorithm for creating a 3D surround view of the vehicle with configurable vantage points and can be computed efficiently. Furthermore, it can choose from different meshes to adaptively generate a match with different surround-view environments on the fly. By surveying different environments and testing the implementation in a real car, we are able to demonstrate the practical feasibility of our solution11The video demos of our algorithm have been uploaded to Youtube: https://www.youtube.com/watch?v=-boYsUSA52c, https://www.youtube.com/watch?v=NED435Uj12E. https://www.youtube.com/watch?v=2sZTCIi4X2M&t=2s. Our approach is competitive with the-state-of-the-art 3D reconstruction algorithms in computer vision while having a considerably lower runtime complexity. Iljoo Baek, Akshit Kanda, Tzu Chieh Tai, Anchan Saxena, Ragunathan Rajkumar |
IV | 5 |
| 2019 | Fractional GPUs: Software-Based Compute and Memory Bandwidth Reservation for GPUsabstractGPUs are increasingly being used in real-time systems, such as autonomous vehicles, due to the vast performance benefits that they offer. As more and more applications use GPUs, more than one application may need to run on the same GPU in parallel. However, real-time systems also require predictable performance from each individual applications which GPUs do not fully support in a multi-tasking environment. Nvidia recently added a new feature in their latest GPU architecture that allows limited resource provisioning. This feature is provided in the form of a closed-source kernel module called the Multi-Process Service (MPS). However, MPS only provides the capability to partition the compute resources of GPU and does not provide any mechanism to avoid inter-application conflicts within the shared memory hierarchy. In our experiments, we find that compute resource partitioning alone is not sufficient for performance isolation. In the worst case, due to interference from co-running GPU tasks, read/write transactions can observe a slowdown of more than 10x. In this paper, we present Fractional GPUs (FGPUs), a software-only mechanism to partition both compute and memory resources of a GPU to allow parallel execution of GPU workloads with performance isolation. As many details of GPU memory hierarchy are not publicly available, we first reverse-engineer the information through various micro-benchmarks. We find that the GPU memory hierarchy is different from that of the CPU, making it well-suited for page coloring. Based on our findings, we were able to partition both the L2 cache and DRAM for multiple Nvidia GPUs. Furthermore, we show that a better strategy exists for partitioning compute resources than the one used by MPS. An FGPU combines both this strategy and memory coloring to provide superior isolation. We compare our FGPU implementation with Nvidia MPS. Compared to MPS, FGPU reduces the average variation in application runtime, in a multi-tenancy environment, from 135% to 9%. To allow multiple applications to use FGPUs seamlessly, we ported Caffe, a popular framework used for machine learning, to use our FGPU API. Saksham Jain, Iljoo Baek, Shige Wang, Ragunathan Rajkumar |
RTAS | 4 |
| 2019 | V2V-based Synchronous Intersection Protocols for Mixed Traffic of Human-Driven and Self-Driving VehiclesabstractSelf-driving vehicles are expected to be at the core of future transportation systems. Over time, these vehicles might enhance traffic efficiency and safety, especially at road intersections. However, there will likely be a long transition period before human-driven vehicles will be completely replaced by automated vehicles. Intersection safety and efficiency might be barely improved if automated vehicles just follow the traffic light signals. In this paper, we present a decentralized intersection protocol named the Distributed Synchronous Intersection Protocol (DSIP) that is for mixed traffic environments, where human-driven and self-driving vehicles cooperate with one another in order to avoid vehicle collisions and possible deadlocks while improving traffic efficiency. In DSIP, all automated vehicles use both Vehicle-to-Vehicle (V2V) communications and traffic lights to traverse the intersection safely. On the other hand, human-driven vehicles simply follow the traffic lights just like they do today. Under the protocol, the automated vehicles synchronize when there are no human-driven vehicles around the intersection. In addition, Cooperative Perception is used to detect the presence of human-driven vehicles, where all automated and connected vehicles sense and share the presence of the human-driven vehicles at the intersection. Our simulation results show that DSIP increases the traffic throughput of the intersections compared to common signalized intersections and other V2V-based intersection protocols. Shunsuke Aoki 0001, Ragunathan Rajkumar |
RTCSA | 2 |
| 2019 | Practical task allocation for software fault-tolerance and its implementation in embedded automotive systems
Anand Bhat, Soheil Samii, Ragunathan Rajkumar |
Real Time Syst. | 3 |
| 2019 | Many suspensions, many problems: a review of self-suspending tasks in real-time systemsabstractIn general computing systems, a job (process/task) may suspend itself whilst it is waiting for some activity to complete, e.g., an accelerator to return data. In real-time systems, such self-suspension can cause substantial performance/schedulability degradation. This observation, first made in 1988, has led to the investigation of the impact of self-suspension on timing predictability, and many relevant results have been published since. Unfortunately, as it has recently come to light, a number of the existing results are flawed. To provide a correct platform on which future research can be built, this paper reviews the state of the art in the design and analysis of scheduling algorithms and schedulability tests for self-suspending tasks in real-time systems. We provide (1) a systematic description of how self-suspending tasks can be handled in both soft and hard real-time systems; (2) an explanation of the existing misconceptions and their potential remedies; (3) an assessment of the influence of such flawed analyses on partitioned multiprocessor fixed-priority scheduling when tasks synchronize access to shared resources; and (4) a discussion of the computational complexity of analyses for different self-suspension task models. Jian-Jia Chen, Geoffrey Nelissen, Wen-Hung Kevin Huang, Maolin Yang 0004, Björn B. Brandenburg, Konstantinos Bletsas 0001, Cong Liu 0005, Pascal Richard, Frédéric Ridouard, Neil C. Audsley, Ragunathan Rajkumar, Dionisio de Niz, Georg von der Brüggen |
Real Time Syst. | 11 |
| 2019 | CSIP: A Synchronous Protocol for Automated Vehicles at Road IntersectionsabstractIntersection management is one of the main challenging issues in road safety because intersections are a leading cause of traffic congestion and accidents. In fact, more than 44% of all reported crashes in the U.S. occur around intersection areas, which, in turn, has led to 8,500 fatalities and approximately 1 million injuries every year. With vehicles expected to become self-driving, the question is whether high throughput can be obtained through intersections while keeping them safe. A spatio-temporal intersection protocol named the Ballroom Intersection Protocol (BRIP) [8] was recently proposed in the literature to address this situation. Under this protocol, automated and connected vehicles arrive at and go through an intersection in a cooperative fashion with no vehicle needing to stop, while maximizing the intersection throughput. Though no vehicles run into one another under ideal environments with BRIP, vehicle accidents can occur when the self-driving vehicles have location errors and/or control system failure. In this article, we present a safe and practical intersection protocol named the Configurable Synchronous Intersection Protocol (CSIP). CSIP is a more general and resilient version of BRIP. CSIP utilizes a certain inter-vehicular distance to meet safety requirements in the presence of GPS inaccuracies and control failure. The inter-vehicular distances under CSIP are much more acceptable and comfortable to human passengers due to longer inter-vehicular distances that do not cause fear. With CSIP, the inter-vehicular distances can also be changed at each intersection to account for different traffic volumes, GPS accuracy levels, and geographical layout of intersections. Our simulation results show that CSIP never leads to traffic accidents even when the system has typical location errors, and that CSIP increases the traffic throughput of the intersections compared to common signalized intersections. Shunsuke Aoki 0001, Ragunathan Rajkumar |
ACM Trans. Cyber Phys. Syst. | 2 |
| 2018 | Recovery Time Considerations in Real-Time Systems Employing Software Fault ToleranceabstractSafety-critical real-time systems like modern automobiles with advanced driving-assist features must employ redundancy for crucial software tasks to tolerate permanent crash faults. This redundancy can be achieved by using techniques like active replication or the primary-backup approach. In such systems, the recovery time which is the amount of time it takes for a redundant task to take over execution on the failure of a primary task becomes a very important design parameter. The recovery time for a given task depends on various factors like task allocation, primary and redundant task priorities, system load and the scheduling policy. Each task can also have a different recovery time requirement (RTR). For example, in automobiles with automated driving features, safety-critical tasks like perception and steering control have strict RTRs, whereas such requirements are more relaxed in the case of tasks like heating control and mission planning. In this paper, we analyze the recovery time for software tasks in a real-time system employing Rate-Monotonic Scheduling (RMS). We derive bounds on the recovery times for different redundant task options and propose techniques to determine the redundant-task type for a task to satisfy its RTR. We also address the fault-tolerant task allocation problem, with the additional constraint of satisfying the RTR of each task in the system. Given that the problem of assigning tasks to processors is a well-known NP-hard bin-packing problem we propose computationally-efficient heuristics to find a feasible allocation of tasks and their redundant copies. We also apply the simulated annealing method to the fault-tolerant task allocation problem with RTR constraints and compare against our heuristics. Anand Bhat, Soheil Samii, Ragunathan Rajkumar |
ECRTS | 3 |
| 2018 | Will Distributed Computing Revolutionize Peace? The Emergence of Battlefield IoTabstractAn upcoming frontier for distributed computing might literally save lives in future military operations. In civilian scenarios, significant efficiencies were gained from interconnecting devices into networked services and applications that automate much of everyday life from smart homes to intelligent transportation. The ecosystem of such applications and services is collectively called the Internet of Things (IoT). Can similar benefits be gained in a military context by developing an IoT for the battlefield? This paper describes unique challenges in such a context as well as potential risks, mitigation strategies, and benefits. Tarek F. Abdelzaher, Nora Ayanian, Tamer Basar, Suhas N. Diggavi, Jana Diesner, Deepak Ganesan, Ramesh Govindan, Susmit Jha, Tancrède Lepoint, Benjamin M. Marlin, Klara Nahrstedt, David M. Nicol, Ragunathan Rajkumar, Stephen Russell 0001, Sanjit A. Seshia, Fei Sha, Prashant J. Shenoy, Mani Srivastava 0001, Gaurav S. Sukhatme, Ananthram Swami, Paulo Tabuada, Don Towsley, Nitin H. Vaidya, Venugopal V. Veeravalli |
ICDCS | 13 |
| 2018 | Real-time Detection, Tracking, and Classification of Moving and Stationary Objects using Multiple Fisheye ImagesabstractThe ability to detect pedestrians and other moving objects is crucial for an autonomous vehicle. This must be done in real-time with minimum system overhead. This paper discusses the implementationof a surround view system to identify moving as well as static objects that are close to the ego vehicle. The algorithm works on 4 views captured by fisheye cameras which are merged into a single frame. The moving object detection and tracking solution uses minimal system overhead to isolate regions of interest (ROIs) containing moving objects. These ROIs are then analyzed using a deep neural network (DNN) to categorize the moving object. With deployment and testing on a real car in urban environments, we have demonstrated the practical feasibility of the solution.11The video demos of our algorithm have been uploaded to Youtube: https://youtu.be/vpoCfC724iA, https://youtu.be/2X4aqH2bMBs Iljoo Baek, Albert Davies, Geng Yan, Ragunathan Rajkumar |
Intelligent Vehicles Symposium | 4 |
| 2018 | QuartzV: Bringing Quality of Time to Virtual MachinesabstractCyber-physical systems are increasingly interconnected and distributed. Examples range from factory-scale industrial robotics to regional-scale smart grids. Therefore, to enable dynamic coordination at scale among geo-distributed physical endpoints, the intelligence behind these systems will often be hosted in the cloud. However, most CPS applications are inherently safety-critical, and require low-latency responses. Hence, a hierarchy of edge cloudlets and the cloud can be used to offload computationally and data-intensive workloads. While low latency is key, a shared sense of time with the added notion of Quality of Time (QoT) is useful for fault detection, and enables fault-tolerant coordinated action in distributed CPS. Given that most public clouds and cloudlets provide multi-tenancy using virtualized units of computing, we aim to introduce the notion of QoT to virtual machines. The use of virtual machines entails the use of a hypervisor, which adds additional timing uncertainty due to relatively higher jitter in clock-read and timer-interrupt latencies. Hence, the use of virtualization presents a challenge in terms of observing and guaranteeing the QoT delivered to an application. To meet these challenges, we present the QuartzV extension to the QoT Stack for Linux, to make virtual machines QoT-aware. We utilize the open-source QEMU-KVM hypervisor, and illustrate the para-virtual design choices that are key for delivering near-native levels of timing performance in virtual machines. We also demonstrate the utility of QuartzV by using it in the development of an industrial-automation application. Experimental evaluations also show the efficacy of QuartzV with respect to the native and fully-virtualized cases. Sandeep D'Souza, Ragunathan Rajkumar |
RTAS | 2 |
| 2018 | Analytical Enhancements and Practical Insights for MPCP with Self-SuspensionsabstractHardware accelerators such as GP-GPUs and DSPs are being increasingly used in computationally-intensive real-time and multimedia systems. System efficiency is often increased when CPU tasks suspend while using these devices. In this paper, we extend the existing Multiprocessor Priority Ceiling Protocol (MPCP) schedulability analysis in this particular context. We present three methods to improve the traditional MPCP analysis that reduces pessimism in analyzing blocking times. Two of these methods, the request-driven and the job-driven approaches, are motivated by prior work and are adapted to MPCP. The third combines these two approaches in a novel way to consistently outperform either on its own. We note that our underlying observations are general, and that such methods can also be used for analyzing other real-time synchronization protocols. Experimental results indicate that our analytical improvements result in a significantly higher schedulability compared to the traditional recursion-based analysis, even when self-suspensions are not considered. Our approach is also competitive with and often outperforms the linear-programming-based FMLP+ analysis, while having a considerably lower runtime complexity. We further substantiate the practical feasibility of suspension-based MPCP and examine its benefits over the busy-waiting approach by presenting a case-study on an NVIDIA TX2 embedded platform using real-world vision applications. Pratyush Patel, Iljoo Baek, Hyoseung Kim 0001, Ragunathan Rajkumar |
RTAS | 4 |
| 2018 | CycleTandem: Energy-Saving Scheduling for Real-Time Systems with Hardware AcceleratorsabstractCyber-physical systems such as autonomous vehicles need to process and analyze multiple simultaneous streams of sensor data in real-time. Therefore, these systems require powerful multi-core platforms with hardware accelerators such as GP-GPUs. These accelerators generally consume significant amounts of power. Therefore, power management is required to ensure that task deadlines are met while staying within the energy and thermal constraints of the system. In these systems, most tasks execute using a combination of CPU and accelerator resources. Hence, the power of the CPU and the accelerator needs to be managed in tandem. To reduce energy consumption, commercially-available accelerators such as GP-GPUs and DSPs expose interfaces to scale their operating voltage and frequency. Hence, we propose the CycleTandem static frequency-scaling technique to co-optimize the operating frequencies of both the CPU and the hardware accelerator. Based on practical considerations of real-world platforms, we consider various energy-management scenarios where the accelerator or CPU frequencies may or may not be adjustable, and propose the CycleSolo family of algorithms for such contexts. Furthermore, we also study partitioning techniques to reduce the operating frequency when multi-core processors are used in conjunction with hardware accelerators. Experimental evaluations indicate that our proposed techniques can yield significant energy savings. We also present a case-study on the NVIDIA TX2 embedded platform to illustrate the energy savings delivered by our proposed techniques. Sandeep D'Souza, Ragunathan Rajkumar |
RTSS | 2 |
| 2018 | A server-based approach for predictable GPU access with improved analysis
Hyoseung Kim 0001, Pratyush Patel, Shige Wang, Ragunathan Rajkumar |
J. Syst. Archit. | 4 |
| 2018 | Tools and Methodologies for Autonomous Driving SystemsabstractDue to the advent of active safety features and automated driving capabilities, the scope and complexity of embedded computing systems within automobiles continue to increase. Moreover, with the introduction of communication technologies like dedicated short range communications (DSRC), autonomous vehicles can also communicate with each other, pedestrians and the infrastructure. As the number of applications and levels of autonomy within such connected and autonomous vehicles (CAVs) increase the inherent safety-critical nature of CAVs imposes strict requirements in terms of testing and verification of system correctness. Hence, appropriate tools and methodologies are to manage these increasing complexity and testing requirements. In this paper, we present a standard reference architecture for CAVs and the tools and methodologies we use to model, design, develop and test systems to realize CAV applications. Anand Bhat, Shunsuke Aoki 0001, Ragunathan Rajkumar |
Proc. IEEE | 3 |
| 2018 | Schedulability Analysis of Tasks with Corunner-Dependent Execution TimesabstractConsider fixed-priority preemptive partitioned scheduling of constrained-deadline sporadic tasks on a multiprocessor. A task generates a sequence of jobs and each job has a deadline that must be met. Assume tasks have Corunner-dependent execution times; i.e., the execution time of a job J depends on the set of jobs that happen to execute (on other processors) at instants when J executes. We present a model that describes Corunner-dependent execution times. For this model, we show that exact schedulability testing is co-NP-hard in the strong sense. Facing this complexity, we present a sufficient schedulability test, which has pseudo-polynomial-time complexity if the number of processors is fixed. We ran experiments with synthetic software benchmarks on a quad-core Intel multicore processor with the Linux/RK operating system and found that for each task, its maximum measured response time was bounded by the upper bound computed by our theory. Björn Andersson, Hyoseung Kim 0001, Dionisio de Niz, Mark Klein 0003, Ragunathan Rajkumar, John P. Lehoczky |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2018 | Predictable Shared Cache Management for Multi-Core Real-Time VirtualizationabstractReal-time virtualization has gained much attention for the consolidation of multiple real-time systems onto a single hardware platform while ensuring timing predictability. However, a shared last-level cache (LLC) on modern multi-core platforms can easily hamper the timing predictability of real-time virtualization due to the resulting temporal interference among consolidated workloads. Since such interference caused by the LLC is highly variable and may have not even existed in legacy systems to be consolidated, it poses a significant challenge for real-time virtualization. In this article, we propose a predictable shared cache management framework for multi-core real-time virtualization. Our framework introduces two hypervisor-level techniques, vLLC and vColoring, that enable the cache allocation of individual tasks running in a virtual machine (VM), which is not achievable by the current state of the art. Our framework also provides a cache management scheme that determines cache allocation to tasks, designs VMs in a cache-aware manner, and minimizes the aggregated utilization of VMs to be consolidated. As a proof of concept, we implemented vLLC and vColoring in the KVM hypervisor running on x86 and ARM multi-core platforms. Experimental results with three different guest OSs (i.e., Linux/RK, vanilla Linux, and MS Windows Embedded) show that our techniques can effectively control the cache allocation of tasks in VMs. Our cache management scheme yields a significant utilization benefit compared to other approaches while satisfying timing constraints. Hyoseung Kim 0001, Ragunathan Rajkumar |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2017 | Mixed-criticality processing pipelinesabstractWhile a number of schemes exist for mixed-criticality scheduling in a single processor setting, no solution exists to cover the industry need for end-to-end scheduling across multiple processors in a pipeline. In this paper, we present an end-to-end zero-slack rate-monotonic scheme (ZSRM) based on real-time pipelines, called the ZSRM pipeline scheduler, that addresses this need. Under ZSRM, each task is associated with a parameter called zero-slack instant, and whenever a higher-criticality job has not finished at its zero-slack instant relative to its arrival time, all jobs of lower criticality are suspended to meet the deadline of the higher-criticality job. We develop a new schedulability test and algorithm for computing the zero-slack instants of tasks scheduled across a pipeline. Dionisio de Niz, Björn Andersson, Hyoseung Kim 0001, Mark Klein 0003, Linh T. X. Phan, Ragunathan Rajkumar |
DATE | 6 |
| 2017 | Thermal Implications of Energy-Saving SchedulersabstractIn many real-time systems, continuous operation can raise processor temperature, potentially leading to system failure, bodily harm to users, or a reduction in the functional lifetime of a system. Static power dominates the total power consumption, and is also directly proportional to the operating temperature. This reduces the effectiveness of frequency scaling and necessitates the use of sleep states. In this work, we explore the relationship between energy savings and system temperature in the context of fixed-priority energy-saving schedulers, which utilize a processor’s deep-sleep state to save energy. We derive insights from a well-known thermal model, and are able to identify proactive design choices which are independent of system constants and can be used to reduce processor temperature. Our observations indicate that, while energy savings are key to lower temperatures, not all energy-efficient solutions yield low temperatures. Based on these insights, we propose the SysSleep and ThermoSleep algorithms, which enable a thermally-effective sleep schedule. We also derive a lower bound on the optimal temperature achievable by energy-saving schedulers. Additionally, we discuss partitioning and task phasing techniques for multi-core processors, which require all cores to synchronously transition into deep sleep, as well as those which support independent deep-sleep transitions. We observe that, while energy optimization is straightforward in some cases, the dependence of temperature on partitioning and task phasing makes temperature minimization non-trivial. Evaluations show that compared to the existing purely energy-efficient design methodology, our proposed techniques yield lower temperatures along with significant energy savings. Sandeep D'Souza, Ragunathan Rajkumar |
ECRTS | 2 |
| 2017 | Practical Task Allocation for Software Fault-Tolerance and Its Implementation in Embedded Automotive SystemsabstractDue to the advent of active safety features and automated driving capabilities, the complexity of embedded computing systems within automobiles continues to increase. Such advanced driver assistance systems (ADAS) are inherently safety-critical and must tolerate failures in any subsystem. However, fault-tolerance in safety-critical systems has been traditionally supported by hardware replication, which is prohibitively expensive in terms of cost, weight, and size for the automotive market. Recent work has studied the use of software-based fault-tolerance techniques that utilize task-level hot and cold standbys to tolerate fail-stop processor and task failures. The benefit of using standbys is maximal when a task and any of its standbys obey the placement constraint of not being co-located on the same processor. We propose a new heuristic based on a "tiered" placement constraint, and show that our heuristic produces a better task assignment that saves at least one processor up to 40% of the time relative to the best known heuristic to date. We then introduce a task allocation algorithm that, for the first time to our knowledge, leverages the run-time attributes of cold standbys. Our empirical study finds that our heuristic uses no more than one additional processor in most cases relative to an optimal allocation that we construct for evaluation purposes using a creative technique. We have designed and implemented our software fault-tolerance framework in AUTOSAR, an automotive industry standard. We use this implementation to provide an experimental evaluation of our task-level fault-tolerance features. Finally, we present an analysis of the worst-case behavior of our task recovery features. Anand Bhat, Soheil Samii, Ragunathan Rajkumar |
RTAS | 3 |
| 2017 | A configurable synchronous intersection protocol for self-driving vehiclesabstractRoad intersections represent a leading cause of traffic congestion and accidents. In fact, more than 44% of all reported crashes in the U.S. occur within intersection areas, which, in turn, lead to 8,500 fatalities and approximately 1 million injuries every year. The expected advent of self-driving vehicles raises the question of how automated and connected vehicles can be used to improve throughput at intersections while keeping them safe. A spatio-temporal intersection protocol named the Ballroom Intersection Protocol (BRIP) [5] was recently proposed in the literature to address this situation. Under this protocol, automated vehicles arrive at and go through an intersection in a cooperative fashion with no vehicle needing to stop, while maximizing the intersection throughput. Though no vehicles run into one another under ideal environments with BRIP, vehicle accidents can occur when the self-driving vehicles have location errors. In this paper, we present a safe and practical intersection protocol named the Configurable Synchronous Intersection Protocol (CSIP) that is a more general and resilient version of BRIP. CSIP utilizes a certain inter-vehicle distance to meet safety requirements against GPS inaccuracy and control failure. In addition, the inter-vehicle distances under CSIP are much more acceptable to human passengers due to longer inter-vehicle distances that do not cause fear. With CSIP, the inter-vehicle distances can also be changed at each intersection to account for different traffic volumes and GPS inaccuracies. Our simulation results show that CSIP never leads to traffic accidents even when the system has typical location errors, and that CSIP increases the traffic throughput of the intersections compared to common signalized intersections. Shunsuke Aoki 0001, Ragunathan Rajkumar |
RTCSA | 2 |
| 2017 | A server-based approach for predictable GPU access controlabstractWe propose a server-based approach to manage a general-purpose graphics processing unit (GPU) in a predictable and efficient manner. Our proposed approach introduces a GPU server task that is dedicated to handling GPU requests from other tasks on their behalf. The GPU server ensures bounded time to access the GPU, and allows other tasks to suspend during their GPU computation to save CPU cycles. By doing so, we address the two major limitations of the existing real-time synchronization-based GPU management approach: busy waiting and long priority inversion. We implemented a prototype of the server-based approach on a real embedded platform. This case study demonstrates the practicality and effectiveness of the server-based approach. Experimental results indicate that the server-based approach yields significant improvements in task schedulability over the existing synchronization-based approach in most practical settings. Although we focus on a GPU in this paper, the server-based approach can also be used for other types of computational accelerators. Hyoseung Kim 0001, Pratyush Patel, Shige Wang, Ragunathan Rajkumar |
RTCSA | 4 |
| 2016 | Sleep Scheduling for Energy-Savings in Multi-core ProcessorsabstractAs transistor geometries get smaller, static leakage power dominates the power consumption in modern processors. This phenomenon diminishes the ability of frequency scaling-based techniques to save energy. Modern processors also provide sleep states which minimize leakage power by gating portions of the processor and/or the system clock. This paper presents partitioned fixed-priority scheduling solutions for utilizing these sleep states to efficiently schedule periodic real-time tasks, and maximize energy savings on multi-core processors. The techniques presented rely on an Enhanced Version of Energy-Saving Rate-Harmonized Scheduling (ES-RHS), and our newly proposed Energy-Saving Rate-Monotonic Scheduling (ES-RMS) policy to maximize the time the processor spends in the lowest-power deep sleep state. In some modern multi-core processors, all cores need to transition synchronously into deep sleep. For this class of processors, we present a partitioning technique called Max-SyncSleep which utilizes a priori task information, to maximize the synchronous deep sleep duration across all processing cores. The performance of Max-SyncSleep is compared to the classical Worst-Fit Decreasing load balancing heuristic. We also illustrate the benefits of using ES-RMS over ES-RHS for this class of processors. For processors which allow cores to individually transition into deep sleep, we prove that, while utilizing ES-RHS on each core, any feasible partition can optimally utilize all of the processor's idle durations to put it into deep sleep. Experimental evaluations indicate that the proposed techniques can provide significant energy savings. Sandeep D'Souza, Anand Bhat, Ragunathan Rajkumar |
ECRTS | 3 |
| 2016 | Real-time cache management for multi-core virtualizationabstractReal-time virtualization techniques have been investigated with the primary goal of consolidating multiple real-time systems onto a single hardware platform while ensuring timing predictability. However, a shared last-level cache (LLC) on recent multi-core platforms can easily hamper timing predictability due to the resulting temporal interference among consolidated workloads. Since such interference caused by the LLC is highly variable and may have not even existed in legacy systems to be consolidated, it poses a significant challenge for real-time virtualization. In this paper, we propose a real-time cache management framework for multi-core virtualization. Our framework introduces two hypervisor-level techniques, vLLC and vColoring, that enable the cache allocation of individual tasks running in a virtual machine (VM), which is not achievable by the current state of the art. Our framework also provides a cache management scheme that determines cache allocation to tasks, designs VMs in a cache-aware manner, and minimizes the aggregated utilization of VMs to be consolidated. As a proof of concept, we implemented vLLC and vColoring in the KVM hypervisor running on x86 and ARM multi-core platforms. Experimental results with three different guest OSs, namely Linux/RK, vanilla Linux and MS Windows Embedded, show that our techniques can effectively control the cache allocation of tasks in VMs. Our cache management scheme yields a significant utilization benefit compared to other approaches. Hyoseung Kim 0001, Ragunathan Rajkumar |
EMSOFT | 2 |
| 2016 | Timeline: An Operating System Abstraction for Time-Aware ApplicationsabstractHaving a shared and accurate sense of time is critical to distributed Cyber-Physical Systems (CPS) and the Internet of Things (IoT). Thanks to decades of research in clock technologies and synchronization protocols, it is now possible to measure and synchronize time across distributed systems with unprecedented accuracy. However, applications have not benefited to the same extent due to limitations of the system services that help manage time, and hardware-OS and OS-application interfaces through which timing information flows to the application. Due to the importance of time awareness in a broad range of emerging applications, running on commodity platforms and operating systems, it is imperative to rethink how time is handled across the system stack. We advocate the adoption of a holistic notion of Quality of Time (QoT) that captures metrics such as resolution, accuracy, and stability. Building on this notion we propose an architecture in which the local perception of time is a controllable operating system primitive with observable uncertainty, and where time synchronization balances applications' timing demands with system resources such as energy and bandwidth. Our architecture features an expressive application programming interface that is centered around the abstraction of a timeline - a virtual temporal coordinate frame that is defined by an application to provide its components with a shared sense of time, with a desired accuracy and resolution. The timeline abstraction enables developers to easily write applications whose activities are choreographed across time and space. Leveraging open source hardware and software components, we have implemented an initial Linux realization of the proposed timeline-driven QoT stack on a standard embedded computing platform. Results from its evaluation are also presented. Fatima M. Anwar 0001, Sandeep D'Souza, Andrew Colquhoun Symington, Adwait Dongare, Ragunathan Rajkumar, Anthony Rowe 0001, Mani Srivastava 0001 |
RTSS | 5 |
| 2016 | Bounding and reducing memory interference in COTS-based multi-core systems
Hyoseung Kim 0001, Dionisio de Niz, Björn Andersson, Mark Klein 0003, Onur Mutlu, Ragunathan Rajkumar |
Real Time Syst. | 6 |
| 2015 | Ballroom Intersection Protocol: Synchronous Autonomous Driving at IntersectionsabstractRoad intersections are considered to be serious bottlenecks in urban transportation. More than 44% of all reported crashes in U.S. Occur within intersection areas, which in turn lead to 8,500 fatalities and approximately 1 million injuries every year. Furthermore, because traffic traveling in one direction is generally stopped at busy intersections to allow traffic to flow in another direction, an intersection creates traffic congestion and frustration. The impact of road intersections on traffic delays leads to enormous waste of human and natural resources. According to the 2011 Urban Mobility Report, the delay endured by the average commuter was 34 hours, which costs in aggregate more than $100 billion each year in the U.S. With the advances in Cyber-Physical Systems (CPS), autonomous driving as a part of Intelligent Transportation Systems (ITS) is likely to be at the heart of urban transportation in the future. Autonomous vehicles have been demonstrated successfully at the DARPA Urban Challenge. General Motors' Electrical-Networked Vehicle, CMU's autonomous vehicle and Google's car are just a few other recently unveiled examples. Therefore, it is critical to address safety and throughput concerns as one of the main challenges for autonomous driving through intersections. In this paper, we propose a spatio-temporal technique called the Ballroom Intersection Protocol (BRIP) to manage the safe and efficient passage of autonomous vehicles through intersections. To achieve high throughput at intersections, BRIP aims to maximize the utilization of the capacity of the intersection area by increasing parallelism. By enforcing a synchronized arrival of autonomous vehicles at intersections, BRIP allows vehicles approaching from all directions to simultaneously and continuously cross without stopping behind or inside the intersection area. Our simulation results show that we are able to avoid collisions and increase the throughput of the intersections by up to 96.24% compared to common signalized intersections. Under BRIP, the optimal intersection capacity utilization of 100% is achievable in certain cases. Seyed (Reza) Azimi, Gaurav Bhatia, Ragunathan Rajkumar, Priyantha Mudalige |
RTCSA | 3 |
| 2015 | Responsive and Enforced Interrupt Handling for Real-Time System VirtualizationabstractThe increasing performance of modern processors makes virtualization a viable solution for consolidating real-time systems into a single hardware platform. Although real-time task scheduling in a virtual machine can benefit from hierarchical scheduling, unbounded interrupt handling time and vulnerability to interrupt storms make practitioners hesitant to virtualize interrupt-driven real-time applications. In this paper, we propose vINT, an interrupt handling scheme designed for real-time system virtualization. vINT provides a pseudo-VCPU abstraction dedicated for interrupt handling, which overcomes the limits imposed by the timing parameters of virtual CPUs in an analyzable way. vINT also accounts for and enforces interrupt handling and resulting execution flows within a guest virtual machine. vINT does not require any change to the guest OS code, so it can be used for virtualizing proprietary, closed-source OSs. We analyze interrupt handling time as well as VCPU and task schedulability, with and without vINT. Our experimental results indicate that vINT achieves timely interrupt handling while providing as good task schedulability as when it is not used. Our case study based on a prototype implementation on the KVM hyper visor shows that vINT yields significant benefits in reducing interrupt handling time and in protecting real-time tasks against interrupt storms permeating into the virtual machine. Hyoseung Kim 0001, Shige Wang, Ragunathan Rajkumar |
RTCSA | 3 |
| 2015 | Self-Driving Vehicles: The Challenges and Opportunities AheadabstractSelf-driving vehicles seem to have become quite the rage in popular culture over the past 3 years or so. Jumpstarted by the DARPA Grand Challenges, the promise of self-driving vehicles does have the potential to revolutionize modern transportation. This talk will provide some insights on many basic questions that, however, still remain unanswered. What are the technological barriers? What can or cannot be sensed? Can vehicles recognize and comprehend as good as (or better than) humans? What role does connectivity play (if any)? Will the technology be affordable for the masses? How do issues like liability, insurance, regulations and societal acceptance impact deployment? The talk will be based on road experiences interspersed with some speculation. Ragunathan Rajkumar |
SenSys | 1 |
| 2014 | Detection and tracking of boundary of unmarked roads
Young-Woo Seo, Ragunathan Rajkumar |
FUSION | 2 |
| 2014 | A multi-sensor fusion system for moving object detection and tracking in urban driving environmentsabstractA self-driving car, to be deployed in real-world driving environments, must be capable of reliably detecting and effectively tracking of nearby moving objects. This paper presents our new, moving object detection and tracking system that extends and improves our earlier system used for the 2007 DARPA Urban Challenge. We revised our earlier motion and observation models for active sensors (i.e., radars and LIDARs) and introduced a vision sensor. In the new system, the vision module detects pedestrians, bicyclists, and vehicles to generate corresponding vision targets. Our system utilizes this visual recognition information to improve a tracking model selection, data association, and movement classification of our earlier system. Through the test using the data log of actual driving, we demonstrate the improvement and performance gain of our new tracking system. Hyunggi Cho, Young-Woo Seo, B. V. K. Vijaya Kumar, Ragunathan Rajkumar |
ICRA | 4 |
| 2014 | Poster abstract: a harmony of sensors: achieving determinism in multi-application sensor networks
Vikram Gupta, Nuno Pereira 0001, Eduardo Tovar, Ragunathan Rajkumar |
IPSN | 4 |
| 2014 | Utilizing instantaneous driving direction for enhancing lane-marking detectionabstractOur earlier lane-marking detection method identified lane-markings appearing on an input image based on the intensity contrast between lane-markings pixels and their neighboring pixels. This detection results in outputs with nearly-zero false negatives, but with many false positives. To filter out these false positives in a principled way, we utilize the driving direction of a roadway. We do this because longitudinal lane-markings delineate the driving direction of a road and the orientations of any true, longitudinal lane-markings appearing on input images should be aligned with this direction. To approximate the driving direction of a road, we detect the vanishing point on a horizon line and draw a line to link the image coordinates of the detected vanishing point to those of the center of the image bottom. We then filter out any lane-marking blobs if their orientations are not aligned with that of the approximated driving direction. Through testing with streets and inter-city highway images, the proposed method demonstrates its effectiveness. Young-Woo Seo, Ragunathan Rajkumar |
Intelligent Vehicles Symposium | 2 |
| 2014 | Bounding memory interference delay in COTS-based multi-core systemsabstractIn commercial-off-the-shelf (COTS) multi-core systems, a task running on one core can be delayed by other tasks running simultaneously on other cores due to interference in the shared DRAM main memory. Such memory interference delay can be large and highly variable, thereby posing a significant challenge for the design of predictable real-time systems. In this paper, we present techniques to provide a tight upper bound on the worst-case memory interference in a COTS-based multi-core system. We explicitly model the major resources in the DRAM system, including banks, buses and the memory controller. By considering their timing characteristics, we analyze the worst-case memory interference delay imposed on a task by other tasks running in parallel. To the best of our knowledge, this is the first work bounding the request re-ordering effect of COTS memory controllers. Our work also enables the quantification of the extent by which memory interference can be reduced by partitioning DRAM banks. We evaluate our approach on a commodity multi-core platform running Linux/RK. Experimental results show that our approach provides an upper bound very close to our measured worst-case interference. Hyoseung Kim 0001, Dionisio de Niz, Björn Andersson, Mark Klein 0003, Onur Mutlu, Ragunathan Rajkumar |
RTAS | 6 |
| 2014 | Energy-efficient allocation of real-time applications onto Heterogeneous ProcessorsabstractSelf-powered vehicles that interact with the physical world, such as spacecraft, require computing platforms with predictable timing behavior and a low energy demand. Energy consumption can be reduced by choosing energy-efficient designs for both hardware and software components of the platform. We leverage the state-of-the-art in energy-efficient hardware design by adopting Heterogeneous Multi-core Processors with support for Dynamic Voltage and Frequency Scaling and Dynamic Power Management. We address the problem of allocating real-time software components onto heterogeneous cores such that total energy is minimized. Our approach is to start from an analytically justified target load distribution and find a task assignment heuristic that approximates it. Our analysis shows that neither balancing the load nor assigning all load to the “cheapest” core is the best load distribution strategy, unless the cores are extremely alike or extremely different. The optimal load distribution is then formulated as a solution to a convex optimization problem. A heuristic that approximates this load distribution and an alternative method that leverages the solution explicitly are proposed as viable task assignment methods. The proposed methods are compared to state-of-the-art on simulated problem instances and in a case study of a soft-real-time application on an off-the-shelf ARM big.LITTLE heterogeneous processor. Alexei Colin, Arvind Kandhalu, Ragunathan Rajkumar |
RTCSA | 3 |
| 2014 | Network-Harmonized Scheduling for multi-application sensor networksabstractSupport for multiple concurrent applications is an important enabler for promoting the use of sensor networks as an infrastructure technology, where multiple users can deploy their applications independently. In such a scenario, different applications on a node may transmit packets at distinct periods, causing the node to change from sleep to active state more often, which negatively impacts the energy consumption of the whole network. In this paper, we propose to batch the transmissions together by defining a harmonizing period to align the transmissions from multiple applications at periodic boundaries. This harmonizing period is then leveraged to design a protocol that coordinates the transmissions across nodes and provides real-time guarantees in a multi-hop network. This protocol, which we call Network- Harmonized Scheduling (NHS), takes advantage of the periodicity introduced to assign offsets to nodes at different hop-levels such that collisions are always avoided, and deterministic behavior is enforced. NHS is a light-weight and distributed protocol that does not require any global state-keeping mechanism. We implemented NHS on the Contiki operating system and show how it can achieve a duty-cycle comparable to an ideal TDMA approach. Vikram Gupta, Nuno Pereira 0001, Shashank Gaur, Eduardo Tovar, Ragunathan Rajkumar |
RTCSA | 5 |
| 2014 | vMPCP: A Synchronization Framework for Multi-core Virtual MachinesabstractThe virtualization of real-time systems has received much attention for its many benefits, such as the consolidation of individually developed real-time applications while maintaining their implementations. However, the current state of the art still lacks properties required for resource sharing among real-time application tasks in a multi-core virtualization environment. In this paper, we propose vMPCP, a synchronization framework for the virtualization of multi-core real-time systems. Vmpcp exposes the executions of critical sections of tasks in a guest virtual machine to the hyper visor. Using this approach, vMPCP reduces and bounds blocking time on accessing resources shared within and across virtual CPUs (VCPUs) assigned on different physical CPU cores. Vmpcp supports periodic server and deferrable server policies for the VCPU budget replenish policy, with an optional budget overrun to reduce blocking times. We provide the VCPU and task schedulability analyses under vMPCP, with different VCPU budget supply policies, with and without overrun. Experimental results indicate that, under vMPCP, deferrable server outperforms periodic server when overrun is used, with as much as 80% more task sets being schedulable. The case study using our hyper visor implementation shows that vMPCP yields significant benefits compared to a virtualization-unaware multi-core synchronization protocol, with 29% shorter response time on average. Hyoseung Kim 0001, Shige Wang, Ragunathan Rajkumar |
RTSS | 3 |
| 2014 | Predicting dynamic computational workload of a self-driving carabstractThis study aims at developing a method that predicts the CPU usage patterns of software tasks running on a self-driving car. To ensure safety of such dynamic systems, the worst-case-based CPU utilization analysis has been used; however, the nature of dynamically changing driving contexts requires more flexible approach for an efficient computing resource management. To better understand the dynamic CPU usage patterns, this paper presents an effort of designing a feature vector to represent the information of driving environments and of predicting, using regression methods, the selected tasks' CPU usage patterns given specific driving contexts. Experiments with real-world vehicle data show a promising result and validate the usefulness of the proposed method. Young-Woo Seo, Junsung Kim 0001, Ragunathan Rajkumar |
SMC | 3 |
| 2014 | Memory reservation and shared page management for real-time systems
Hyoseung Kim 0001, Ragunathan Rajkumar |
J. Syst. Archit. | 2 |
| 2014 | Utility-Based Resource Overbooking for Cyber-Physical SystemsabstractTraditional hard real-time scheduling algorithms require the use of the worst-case execution times to guarantee that deadlines will be met. Unfortunately, many algorithms with parameters derived from sensing the physical world suffer large variations in execution time, leading to pessimistic overall utilization, such as visual recognition tasks. In this article, we present ZS-QRAM, a scheduling approach that enables the use of flexible execution times and application-derived utility to tasks in order to maximize total system utility. In particular, we provide a detailed description of the algorithm, the formal proofs for its temporal protection, and a detailed, evaluation. Our evaluation uses the Utility Degradation Resilience (UDR) showing that ZS-QRAM is able to obtain 4× as much UDR as ZSRM, a previous overbooking approach, and almost 2× as much UDR as Rate-Monotonic with Period Transformation (RM/TP). We then evaluate a Linux kernel module implementation of our scheduler on an Unmanned Air Vehicle (UAV) platform. We show that, by using our approach, we are able to keep the tasks that render the most utility by degrading lower-utility ones even in the presence of highly dynamic execution times. Dionisio de Niz, Lutz Wrage, Anthony Rowe 0001, Ragunathan Rajkumar |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2013 | A Coordinated Approach for Practical OS-Level Cache Management in Multi-core Real-Time SystemsabstractMany modern multi-core processors sport a large shared cache with the primary goal of enhancing the statistic performance of computing workloads. However, due to resulting cache interference among tasks, the uncontrolled use of such a shared cache can significantly hamper the predictability and analyzability of multi-core real-time systems. Software cache partitioning has been considered as an attractive approach to address this issue because it does not require any hardware support beyond that available on many modern processors. However, the state-of-the-art software cache partitioning techniques face two challenges: (1) the memory co-partitioning problem, which results in page swapping or waste of memory, and (2) the availability of a limited number of cache partitions, which causes degraded performance. These are major impediments to the practical adoption of software cache partitioning. In this paper, we propose a practical OS-level cache management scheme for multi-core real-time systems. Our scheme provides predictable cache performance, addresses the aforementioned problems of existing software cache partitioning, and efficiently allocates cache partitions to schedule a given task set. We have implemented and evaluated our scheme in Linux/RK running on the Intel Core i7 quad-core processor. Experimental results indicate that, compared to the traditional approaches, our scheme is up to 39% more memory space efficient and consumes up to 25% less cache partitions while maintaining cache predictability. Our scheme also yields a significant utilization benefit that increases with the number of tasks. Hyoseung Kim 0001, Arvind Kandhalu, Ragunathan Rajkumar |
ECRTS | 3 |
| 2013 | Towards a viable autonomous driving research platformabstractWe present an autonomous driving research vehicle with minimal appearance modifications that is capable of a wide range of autonomous and intelligent behaviors, including smooth and comfortable trajectory generation and following; lane keeping and lane changing; intersection handling with or without V2I and V2V; and pedestrian, bicyclist, and workzone detection. Safety and reliability features include a fault-tolerant computing system; smooth and intuitive autonomous-manual switching; and the ability to fully disengage and power down the drive-by-wire and computing system upon E-stop. The vehicle has been tested extensively on both a closed test field and public roads. Junqing Wei, Jarrod M. Snider, Junsung Kim 0001, John M. Dolan, Ragunathan Rajkumar, Bakhtiar Litkouhi |
Intelligent Vehicles Symposium | 5 |
| 2013 | Utility-based resource overbooking for Cyber-Physical SystemsabstractThe tight coupling among computation, sensing and control found in Cyber-Physical Systems (CPS) often requires information processing to be completed within strict timing deadlines. Traditional hard real-time scheduling algorithms require the use of the worst-case execution times to guarantee that deadlines will be met. Unfortunately, many algorithms with parameters derived from sensing the physical world suffer from large variations in execution time, which leads to pessimistic overall utilization. For example, object tracking in a computer vision system is highly dependent on the number and size of the objects within the camera's field of view. In this paper, we present the formal description of ZS-QRAM [8], a scheduling approach that allows system designers to flexibly assign execution times and application-derived utility to tasks in order to maximize total system utility even in the presence of highly variable processing estimates. In particular, we provide a detailed description of the algorithm, the formal proofs for its temporal protection and a detail evaluation. Our evaluation uses the Utility Degradation Resilience (UDR) metric presented in [8]. Our results show that ZS-QRAM is able to obtain four times as much UDR as ZSRM, a previous overbooking approach, and almost twice as much UDR as Rate-Monotonic with Period Transformation (RM/TP) even when the latter does not provide temporal protection. Dionisio de Niz, Lutz Wrage, Anthony Rowe 0001, Ragunathan Rajkumar |
RTCSA | 4 |
| 2013 | Segment-Fixed Priority Scheduling for Self-Suspending Real-Time TasksabstractRecent trends in System-on-a-Chip show that an increasing number of special-purpose processors are being added to improve the efficiency of common operations. Unfortunately, the use of these processors may introduce suspension delays incurred by communication, synchronization and external I/O operations. When these processors are used in real-time systems, conventional schedulability analyses incorporate these delays in the worst-case execution/response time, hence significantly reducing the schedulable utilization. In this paper, we provide schedulability analyses and propose segment-fixed priority scheduling for self-suspending tasks. We model the tasks as segments of execution separated by suspensions. We start from providing response-time analyses for self-suspending tasks under Rate Monotonic Scheduling (RMS). While RMS is shown to not be optimal, it can be used effectively in some special cases that we have identified. We then derive a utilization bound for the cases as a function of the ratio of the suspension duration to the period of the tasks. For general cases, we develop a segment-fixed priority scheduling scheme. Our scheme assigns individual segments different priorities and phase offsets that are used for phase enforcement to control the unexpected self-suspending nature. With the exact schedulability analysis designed for our scheme, our experiments show that the proposed scheme provides up to 40 times more schedulable utilization than RMS. Junsung Kim 0001, Björn Andersson, Dionisio de Niz, Ragunathan Rajkumar |
RTSS | 4 |
| 2012 | pCOMPATS: Period-Compatible Task Allocation and Splitting on Multi-core ProcessorsabstractExtensive research is underway to build chips with potentially hundreds of cores. In this paper, we consider the problem of scheduling periodic real-time tasks on multi-core processors. We develop a task partitioning algorithm called Period-Compatible-Allocation and Task-Splitting (pCOMPATS) for fixed-priority scheduling of preemptive hard real-time tasks where the utilization of each of the tasks is less than 50%. pCOMPATS clusters compatible tasks together with task splitting to improve the achievable utilization. We show that as the number of cores increases, the least upper bound on schedulable utilization achieved using pCOMPATS approaches 100% per core. To the best of our knowledge, this is the first result that shows that the utilization bound improves as the number of processing cores in the system increases. We refer to tasks having utilization greater than or equal to 50% as heavy tasks and provide a task-partitioning algorithm called pCOMPATS-HT for allocating such tasks. We show that the upper bound on schedulable utilization when tasks are scheduled using pCOMPATS-HT is at most 72%. We also evaluate the performance of pCOMPATS and other well-known partitioning techniques, and show that using pCOMPATS provides much better schedulable utilization in the average case. We characterize the overhead of pCOMPATS using measurements on an Intel Core i7 processor running Linux/RK. The overheads are seen to be low on the platform, making pCOMPATS to be practical. Our results are especially useful in the context of future many-core processors with dozens to hundreds of cores per processor. Arvind Kandhalu, Karthik Lakshmanan, Junsung Kim 0001, Ragunathan Rajkumar |
IEEE Real-Time and Embedded Technology and Applications Symposium | 4 |
| 2012 | ExSched: An External CPU Scheduler Framework for Real-Time SystemsabstractScheduling theory and algorithms have been well studied in the real-time systems literature. Many useful approaches and solutions have appeared in different problem domains. While their theoretical effectiveness has been extensively discussed, the community is now facing implementation challenges that show the impact of the algorithms in practice. In this paper, we propose a scheduler framework, called ExSched, which enables different schedulers to be developed for different operating system (OS) platforms without any modifications to the OS itself, using a unified interface. The framework will easily keep up with changes in the kernel since it is only dependent on a few kernel primitives. The usefulness of this framework is that scheduling policies can be implemented as external plug-ins. They can simply use the ExSched interface instead of platform-dependent functions, since platform details are abstracted by ExSched. The advantage for industry is that they would more easily keep up with new kernel versions since ExSched does not require patches. The advantage for academia is that we could focus on the development of schedulers instead of tedious and time-consuming installations of patched kernels. Our prototype implementation of ExSched supports Linux and Vx Works and it comes with example schedulers which include hierarchical and multi-core schedulers in addition to traditional fixed-priority scheduling (FPS) and earliest deadline first (EDF) algorithms. Mikael Asberg, Thomas Nolte, Shinpei Kato, Ragunathan Rajkumar |
RTCSA | 4 |
| 2012 | Shared-Page Management for Improving the Temporal Isolation of Memory Reservations in Resource KernelsabstractMemory reservation provides real-time applications with guaranteed memory access to a specified amount of physical memory. However, previous work on memory reservation primarily focused on private pages, and did not pay attention to shared pages, which are widely used in current operating systems. With previous schemes, a real-time application may experience unexpected timing delays from other applications through shared pages that are shared by another process, even though the application has enough free pages in its reservation. In this paper, we describe problems with shared pages in real-time applications, and propose a shared-page management mechanism to enhance the temporal isolation of memory reservations in resource kernels that use resource reservation. The proposed mechanism consists of two techniques, Shared-Page Conservation (SPC) and Shared-Page Eviction Lock (SPEL), each of which prevents timing penalties caused by the seemingly arbitrary eviction of shared pages. The mechanism can manage shared data for inter-process communication and shared libraries, as well as pages shared by the kernel's copy-on-write technique and file caches. We have implemented and evaluated our schemes on the Linux/RK platform, but it can be applied to other operating systems with paged virtual memory. Hyoseung Kim 0001, Ragunathan Rajkumar |
RTCSA | 2 |
| 2012 | QoS-Based Resource Allocation for Next-Generation Spacecraft NetworksabstractCurrent spacecraft systems generally have monolithic structures, but a "fractionated" architecture is being considered for next generation spacecrafts. A fractionated spacecraft system is a cluster of independent modules that communicate wirelessly to maintain cluster flight formations and realize the functions usually performed by a monolithic satellite. The envisioned benefits of the fractionated approach include enhanced responsiveness, greater flexibility, robustness and co-existence of multiple missions from different sources with varying degree of trust. The fractionated architecture, however, introduces significant new challenges from the perspective of resource allocation and management. The mobile nature of the clusters and the modules within the cluster implies that the network topology is highly time-varying. A cluster with multiple missions can require messages to be transmitted across the network with varying degrees of QoS requirements such as timeliness and data delivery reliability. The system must determine the appropriate and timely resource allocation for these missions. In this paper, we address these resource allocation challenges by introducing an abstraction of dynamic graphs, and extending the QoS based Resource Allocation Model (Q-RAM) to operate on these dynamic graphs. We develop a mechanism to decompose a dynamic graph into multiple static sub-graphs using which the resource allocation problem is partitioned into multiple sub-problems within each of these static sub-graphs. We have experimentally evaluated our solution by building a simulation framework called SatSim, which can handle a variety of satellite configurations and mobility models. The proposed solution is shown to achieve a near-optimal solution for the resource allocation problem in time-varying networks, while reducing time complexity significantly. Arvind Kandhalu, Ragunathan Rajkumar |
RTSS | 2 |
| 2012 | SAFER: System-level Architecture for Failure Evasion in Real-time ApplicationsabstractRecent trends towards increasing complexity in distributed embedded real-time systems pose challenges in designing and implementing a reliable system such as a self-driving car. The conventional way of improving reliability is to use redundant hardware to replicate the whole (sub)system. Although hardware replication has been widely deployed in hard real-time systems such as avionics, space shuttles and nuclear power plants, it is significantly less attractive to many applications because the amount of necessary hardware multiplies as the size of the system increases. The growing needs of flexible system design are also not consistent with hardware replication techniques. To address the needs of dependability through redundancy operating in real-time, we propose a layer called SAFER(System-level Architecture for Failure Evasion in Real-time applications) to incorporate configurable task-level fault-tolerance features to tolerate fail-stop processor and task failures for distributed embedded real-time systems. To detect such failures, SAFER monitors the health status and state information of each task and broadcasts the information. When a failure is detected using either time-based failure detection or event-based failure detection, SAFER reconfigures the system to retain the functionality of the whole system. We provide a formal analysis of the worst-case timing behaviors of SAFER features. We also describe the modeling of a system equipped with SAFER to analyze timing characteristics through a model-based design tool called SysWeaver. SAFER has been implemented on Ubuntu 10.04 LTS and deployed on Boss, an award-winning autonomous vehicle developed at Carnegie Mellon University. We show various measurements using simulation scenarios used during the 2007 DARPA Urban Challenge. Finally, we present a case study of failure recovery by SAFER when node failures are injected. Junsung Kim 0001, Gaurav Bhatia, Ragunathan Rajkumar, Markus Jochim |
RTSS | 3 |
| 2012 | A Cyber-Physical FutureabstractCyber-physical systems (CPSs) couple the cyber aspects of computing and communications tightly with the dynamics and physics of physical systems operating in the world around us. This emerging multidisciplinary frontier will enable revolutionary changes in the way humans live. Just like the Internet has transformed how national economies are intertwined, how humans interact, and how commerce is conducted, CPSs will transform how humans interact with and control the physical environment to the greater benefit of society. Major economic sectors that will see dramatic advances will include transportation, energy, buildings, healthcare, manufacturing, physical infrastructure, agriculture, and defense among others. The CPS community will comprise computer scientists, engineers of all stripes as well as many biologists, chemists, and physicians. The arts will also leverage the innovations of CPSs and inspire even more novel inventions in the scientific and engineering realms that combine the world of computing with the physical world. In short, the cyber-physical world of the future will be both very different and very welcome. Ragunathan Rajkumar |
Proc. IEEE | 1 |
| 2012 | Overload provisioning in mixed-criticality cyber-physical systemsabstractCyber-physical systems are an emerging class of applications that require tightly coupled interaction between the computational and physical worlds. These systems are typically realized using sensor/actuator interfaces connected with processing backbones. Safety is a primary concern in cyber-physical systems since the actuators directly influence the physical world. However, unexpected or unusual conditions in the physical world can manifest themselves as increased workload demands being offered to the computational infrastructure of a cyber-physical system. Guaranteeing system safety under overload conditions is therefore a prime concern in developing and deploying cyber-physical systems. In this work, we study this problem in the context of a radar surveillance system, where tasks have different levels of criticality or influence on system safety . In the face of overloads, we observe that the desirable property in such systems is that the more critical tasks continue to meet their timing requirements. We capture this mixed-criticality overload requirement using a formal overload-tolerance metric called ductility . Using this overload-tolerance metric, we first develop our solution in the context of uniprocessor systems, where we show that Zero-Slack scheduling (ZS) algorithms can be used to improve the overload behavior in mixed-criticality cyber-physical systems compared to existing fixed-priority scheduling algorithms like Rate-Monotonic Scheduling (RMS) and Criticality-As-Priority-Assignment (CAPA). Leveraging these results, we then develop a criticality-aware task allocation algorithm called Compress-on-Overload Packing (COP) for dealing with multiprocessor cyber-physical systems. Evaluation results show that COP achieves up to five times better ductility than traditional load balancing bin-packing algorithms like Worst-Fit Decreasing (WFD). Finally, we apply ZS and COP to the radar surveillance system to demonstrate the resulting improvement in system overload behavior. Our implementation of the Zero-Slack scheduler is available as a part of the Linux/RK project, which provides resource kernel extensions for Linux. Karthik Lakshmanan, Dionisio de Niz, Ragunathan Rajkumar, Gabriel A. Moreno |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2012 | Guest Editorial Special Section on Cyber-Physical Systems and Cooperating ObjectsabstractThe four papers in this special section present examples of recent advances in the state-of-the-art of cyber-physical systems and cooperating objects. Chenyang Lu 0001, Ragunathan Rajkumar, Eduardo Tovar |
IEEE Trans. Ind. Informatics | 2 |
| 2011 | Resource Sharing in GPU-Accelerated Windowing SystemsabstractRecent windowing systems allow graphics applications to directly access the graphics processing unit (GPU) for fast rendering. However, application tasks that render frames on the GPU contend heavily with the windowing server that also accesses the GPU to blit the rendered frames to the screen. This resource-sharing nature of direct rendering introduces core challenges of priority inversion and temporal isolation in multi-tasking environments. In this paper, we identify and address resource-sharing problems raised in GPU-accelerated windowing systems. Specifically, we propose two protocols that enable application tasks to efficiently share the GPU resource in the X Window System. The Priority Inheritance with X server (PIX) protocol eliminates priority inversion caused in accessing the GPU, and the Reserve Inheritance with X server (RIX) protocol addresses the same problem for resource-reservation systems. Our design and implementation of these protocols highlight the fact that neither the X server nor user applications need modifications to use our solutions. Our evaluation demonstrates that multiple GPU-accelerated graphics applications running concurrently in the X Window System can be correctly prioritized and isolated by the PIX and the RIX protocols. Shinpei Kato, Karthik Lakshmanan, Yutaka Ishikawa, Ragunathan Rajkumar |
IEEE Real-Time and Embedded Technology and Applications Symposium | 4 |
| 2011 | Mixed-Criticality Task Synchronization in Zero-Slack SchedulingabstractRecent years have seen an increasing interest in the scheduling of mixed-criticality real-time systems. These systems are composed of groups of tasks with different levels of criticality deployed over the same processor(s). Such systems must be able to accommodate additional execution-time requirements that may occasionally be needed. When overload conditions develop, critical tasks must still meet their timing constraints at the expense of less critical tasks. Zero-slack scheduling algorithms are promising candidates for such systems. These algorithms guarantee that all tasks meet their deadlines when no overload occurs, and that criticality ordering is satisfied under overloads. Unfortunately, when mutually exclusive resources are shared across tasks, these guarantees are voided. Furthermore, the dual-execution modes of tasks in mixed-criticality systems violate the assumptions of traditional real-time synchronization protocols like PCP and hence the latter cannot be used directly. In this paper, we develop extensions to real-time synchronization protocols (Priority Inheritance and Priority Ceiling Protocol) that coordinate the mode changes of the zero-slack scheduler. We analyze the properties of these new protocols and the blocking terms they introduce. We maintain the deadlock avoidance property of our PCP extension, called the Priority and Criticality Ceiling Protocol (PCCP), and limit the blocking to only one critical section for each of the zero-slack scheduling execution modes. We also develop techniques to accommodate the blocking terms arising from synchronization, in calculating the zero-slack instants used by the scheduler. Finally, we conduct an experimental evaluation of PCCP. Our evaluation shows that PCCP is able to take advantage of the capacity of zero-slack schedulers to reclaim unused over-provisioning of resources that are only used in critical execution modes. This allows PCCP to accommodate larger blocking terms. Karthik Lakshmanan, Dionisio de Niz, Ragunathan Rajkumar |
IEEE Real-Time and Embedded Technology and Applications Symposium | 3 |
| 2011 | SAGA: Tracking and Visualization of Building EnergyabstractIn this paper, we present SAGA, a system for building energy management that provides robust multi-hop wireless sensing, actuation, device management, plotting of historical data and a server backend API to support remote access. Information collected from sensor nodes is stored locally for quick retrieval even if outside network connectivity is lost or unavailable. When network connectivity is available, data is pushed using the extensible Message Passing Protocol (XMPP). This enables external server-side archiving of historical events and additional processing as well as secure bi-directional communication from gateways behind firewalls or with dynamic IP address like those found in broadband connected homes. SAGA provides a web interface that allows devices to be easily configured with aliases and grouped together based on sensor type. Individual and groups of sensors can be plotted to show relative comparisons of different sensor values. For example, selecting to plot a group of energy metering devices shows the overall distribution of energy consumption per-device. A local multi-resolution storage system provides optimized access to recent high-resolution data along with quick retrieval of pre-computed long-term averages. SAGA has been deployed in three homes around the Pittsburgh area, collecting data for more than two years. Maxim Buevich, Anthony Rowe 0001, Ragunathan Rajkumar |
RTCSA (2) | 3 |
| 2011 | A Framework for Programming Sensor Networks with Scheduling and Resource-Sharing OptimizationsabstractSeveral projects in the recent past have aimed at promoting Wireless Sensor Networks as an infrastructure technology, where several independent users can submit applications that execute concurrently across the network. Concurrent multiple applications cause significant energy-usage overhead on sensor nodes, that cannot be eliminated by traditional schemes optimized for single-application scenarios. In this paper, we outline two main optimization techniques for reducing power consumption across applications. First, we describe a compiler based approach that identifies redundant sensing requests across applications and eliminates those. Second, we cluster the radio transmissions together by concatenating packets from independent applications based on Rate-Harmonized Scheduling. Vikram Gupta, Eduardo Tovar, Karthik Lakshmanan, Ragunathan Rajkumar |
RTCSA (2) | 4 |
| 2011 | Energy-Aware Partitioned Fixed-Priority Scheduling for Chip Multi-processorsabstractEnergy management is becoming an increasingly important problem in application domains ranging from embedded devices to data centers. In many such systems, multi-core processors are projected as a promising technology to achieve improved performance with a lower power envelope. Managing the application power consumption under timing constraints poses significant challenges in these emerging platforms. In this paper, we study the energy-efficient scheduling of periodic real time tasks with implicit deadlines on chip multi-core processors (CMPs). We specifically consider processors with a single voltage and clock frequency domain, such as the state-of-the-art embedded multi-core NVIDIA Tegra 2processor and enterprise-class processors such as Intel'sItanium 2, i5, i7 and IBM's Power 6 and Power 7series. The major contributions of this work are (i)we prove that Worst-Fit-Decreasing (WFD) task partitioning when Rate-Monotonic Scheduling (RMS) is used has an approximation ratio of 1.71 for the problem of minimizing the schedulable operating frequency with partitioned fixed-priority scheduling, (ii) we illustrate the major shortcoming of WFD with RMS resulting from not considering task periods during allocation, and(iii) we propose a Single-clock domain multi-processor Frequency Assignment Algorithm (SFAA) that determines a globally energy-efficient frequency while including task period relationships. Our evaluation results show that SFAA provides significant energy gains when compared to WFD. In fact SFAA is shown to save up to 55% more power compared to WFD for an octa-core processor. Arvind Kandhalu, Junsung Kim 0001, Karthik Lakshmanan, Ragunathan Rajkumar |
RTCSA (1) | 4 |
| 2011 | Making WSN TDMA Practical: Stealing Slots Up and Down the TreeabstractTime Division Multiple Access (TDMA) communication protocols in wireless sensor networks provide collision-free communication that increases energy-efficiency while maintaining deterministic packet latencies. The TDMA-ASAP[1] protocol proposed stealing neighbor's slots when networks are running at low-duty cycles to reduce the potentially large latencies of the TDMA cycle size. In this paper, we further reduce the end-to-end latencies by intelligently spreading slots across the TDMA cycle and by enhancing the stealing opportunities in the schedule, stealing slots scheduled for downstream (control) and upstream (data) messages. We also provide a practical time synchronization algorithm that operates within the TDMA schedule. In order to evaluate our new schemes, we carried out both simulation studies to show scalability and a test bed implementation of Fire Fly wireless sensor nodes to show feasibility. Our schemes provide higher peak throughput (nearly 2x) as compared to a common low-power-listen contention-based (LPL-CSMA) protocol and improves the average packet latency by as much as 5x as compared to existing TDMA protocols without slot-stealing and up to 2x as compared to TDMA-ASAP. John Yackovich, Daniel Mossé, Anthony Rowe 0001, Ragunathan Rajkumar |
RTCSA (1) | 4 |
| 2011 | RGEM: A Responsive GPGPU Execution Model for Runtime EnginesabstractGeneral-purpose computing on graphics processing units, also known as GPGPU, is a burgeoning technique to enhance the computation of parallel programs. Applying this technique to real-time applications, however, requires additional support for timeliness of execution. In particular, the non-preemptive nature of GPGPU, associated with copying data to/from the device memory and launching code onto the device, needs to be managed in a timely manner. In this paper, we present a responsive GPGPU execution model (RGEM), which is a user-space runtime solution to protect the response times of high-priority GPGPU tasks from competing workload. RGEM splits a memory-copy transaction into multiple chunks so that preemption points appear at chunk boundaries. It also ensures that only the highest-priority GPGPU task launches code onto the device at any given time, to avoid performance interference caused by concurrent launches. A prototype implementation of an RGEM-based CUDA runtime engine is provided to evaluate the real-world impact of RGEM. Our experiments demonstrate that the response times of high-priority GPGPU tasks can be protected under RGEM, whereas their response times increase in an unbounded fashion without RGEM support, as the data sizes of competing workload increase. Shinpei Kato, Karthik Lakshmanan, Mihir Kelkar, Yutaka Ishikawa, Ragunathan Rajkumar |
RTSS | 6 |
| 2011 | Nano-CF: A coordination framework for macro-programming in Wireless Sensor NetworksabstractWireless Sensor Networks (WSN) are being used for a number of applications involving infrastructure monitoring, building energy monitoring and industrial sensing. The difficulty of programming individual sensor nodes and the associated overhead have encouraged researchers to design macro-programming systems which can help program the network as a whole or as a combination of subnets. Most of the current macro-programming schemes do not support multiple users seamlessly deploying diverse applications on the same shared sensor network. As WSNs are becoming more common, it is important to provide such support, since it enables higher-level optimizations such as code reuse, energy savings, and traffic reduction. In this paper, we propose a macro-programming framework called Nano-CF, which, in addition to supporting in-network programming, allows multiple applications written by different programmers to be executed simultaneously on a sensor networking infrastructure. This framework enables the use of a common sensing infrastructure for a number of applications without the users being concerned about the applications already deployed on the network. The framework also supports timing constraints and resource reservations using the Nano-RK operating system. Nano-CF is efficient at improving WSN performance by (a) combining multiple user programs, (b) aggregating packets for data delivery, and (c) satisfying timing and energy specifications using Rate-Harmonized Scheduling. Using representative applications, we demonstrate that Nano-CF achieves 90% reduction in Source Lines-of-Code (SLoC) and 50% energy savings from aggregated data delivery. Vikram Gupta, Junsung Kim 0001, Aditi Pandya, Karthik Lakshmanan, Ragunathan Rajkumar, Eduardo Tovar |
SECON | 5 |
| 2011 | Efficient Elastic Resource Management for Dynamic Embedded SystemsabstractDynamic resource management is generating growing interest as a way to simplify systems deployment, react to changing operational conditions and improve efficiency in using system resources. Buttazzo et al proposed the elastic task model (ETM) that dynamically moves the tasks instantiation period around a nominal value, increasing or decreasing the bandwidth that each task uses from the resource. However such computation requires a number of iteration steps. This paper identifies the QoS-based resource distribution problem and proposes a generic framework that encompasses the ETM. An algorithm is presented for the resource distribution that computes the results in minimum time, outperforming the original ETM scheme in terms of computational complexity. Ricardo Marau, Karthik Lakshmanan, Paulo Pedreiras, Luís Almeida 0001, Ragunathan Rajkumar |
TrustCom | 5 |
| 2011 | TimeGraph: GPU Scheduling for Real-Time Multi-Tasking Environments
Shinpei Kato, Karthik Lakshmanan, Ragunathan Rajkumar, Yutaka Ishikawa |
USENIX ATC | 3 |
| 2011 | CPU scheduling and memory management for interactive real-time applications
Shinpei Kato, Yutaka Ishikawa, Ragunathan Rajkumar |
Real Time Syst. | 3 |
| 2010 | Cyber-physical systems: the next computing revolutionabstractCyber-physical systems (CPS) are physical and engineered systems whose operations are monitored, coordinated, controlled and integrated by a computing and communication core. Just as the internet transformed how humans interact with one another, cyber-physical systems will transform how we interact with the physical world around us. Many grand challenges await in the economically vital domains of transportation, health-care, manufacturing, agriculture, energy, defense, aerospace and buildings. The design, construction and verification of cyber-physical systems pose a multitude of technical challenges that must be addressed by a cross-disciplinary community of researchers and educators. Ragunathan Rajkumar, Insup Lee 0001, Lui Sha, John A. Stankovic |
DAC | 1 |
| 2010 | Integrated end-to-end timing analysis of networked AUTOSAR-compliant systemsabstractAs Electronic Control Units (ECUs) and embedded software functions within an automobile keep increasing in number, the scale and complexity of automotive embedded systems is growing at a very rapid pace. Hence, the automotive industry has been developing the Automotive Open System Architecture (AUTOSAR) to harness the reusability of common interfaces to communication buses, real-time operating systems and services. These common interfaces foster ease of adoption, interoperability, maintainability, predictability, and analyzability. However, realizing such standards also requires strong support from end-to-end design tool chains. In this paper, we describe some key analytical components that together characterize the end-to-end timing properties of hierarchical bus structures composed of FlexRay, CANbus and LINbus. Our analysis shows that the practical constraints imposed by standards such as AUTOSAR can lead to higher levels of schedulable resource utilization. This reduces both the overall component count and cost, while facilitating easy enhancements. Our analytical results show (a) how a schedulable utilization of 100% can be obtained for time-triggered FlexRay static segments under AUTOSAR compliance, (b) average-case schedulable utilization of 87% for the event-triggered CAN bus, and (c) similarities between LINbus and FlexRay analyses. We generalize the analytical results from different bus technologies, by exploiting their common underlying structure to enable an integrated end-to-end timing analysis of hierarchical heterogeneous networks. These together yield an end-to-end framework to analyze heterogeneously networked AUTOSAR-compliant automotive systems. Karthik Lakshmanan, Gaurav Bhatia, Ragunathan Rajkumar |
DATE | 3 |
| 2010 | AIRS: Supporting Interactive Real-Time Applications on Multicore PlatformsabstractModern real-time systems increasingly operate with multiple interactive applications. While these systems often require reliable quality of service (QoS) for the applications, even under heavy workloads, many existing CPU schedulers are not very capable of satisfying such requirements. In this paper, we design and implement an Advanced Interactive and Real-time Scheduler, called AIRS. AIRS is aimed at supporting systems that run multiple interactive real-time applications, particularly on multicore platforms. It provides a new CPU reservation mechanism to enhance the QoS of the overall system. The reservation algorithm is based on the prior Constant Bandwidth Server (CBS) algorithm, but is more flexible and efficient, when multiple applications reserve CPU bandwidth. It also provides a new multicore scheduler to improve the absolute CPU bandwidth available for the applications to perform well. The scheduling algorithm is subject to the prior Earliest Deadline First with Window-constraint Migration (EDF-WM) algorithm, but is extended to work with the new CPU reservation mechanism. Experimental evaluation shows that AIRS delivers higher quality to simultaneous playback of multiple movies than the existing real-time scheduler. It also demonstrates that AIRS offers hard timing guarantees for randomly-generated task sets with heavy workloads. Shinpei Kato, Ragunathan Rajkumar, Yutaka Ishikawa |
ECRTS | 2 |
| 2010 | Utilization-based schedulability analysis for switched Ethernet aiming dynamic QoS managementabstractEthernet switches are typically found in many large-scale distributed real-time systems providing low-end transactions as well as bulk backbone routing to real-time applications. The FTT-SE protocol (Flexible Time-Triggered communication over Switched Ethernet) is a recent proposal to bypass the limitations of conventional switches in terms of real-time behavior while catering for growing requirements on dynamic reconfigurability and adaptability. For this end, this paper develops linear time-complexity and memory-efficient on-line admission control tests based on utilization bounds for Rate-Monotonic and EDF scheduling on Ethernet switches using FTT-SE, which are suited for dynamic Quality of Service (QoS) management. Our analysis also has broader applicability in general periodic task sets with bounded release delays. For FTT-SE with 100 Mbps links and 1500 bytes of maximum packet size, our sufficient schedulability condition achieves an utilization bound of 61% for RMS and 88% for EDF. Simulation results on randomly generated task sets demonstrate that such bounds are within 18% and 5% utilization of the ideal tests for RMS and EDF, respectively. Ricardo Marau, Luís Almeida 0001, Paulo Pedreiras, Karthik Lakshmanan, Ragunathan Rajkumar |
ETFA | 5 |
| 2010 | Resource Allocation in Distributed Mixed-Criticality Cyber-Physical SystemsabstractLarge-scale distributed cyber-physical systems will have many sensors/actuators (each with local micro-controllers), and a distributed communication/computing backbone with multiple processors. Many cyber-physical applications will be safety critical and in many cases unexpected workload spikes are likely to occur due to unpredictable changes in the physical environment. In the face of such overload scenarios, the desirable property in such systems is that the most critical applications continue to meet their deadlines. In this paper, we capture this mixed-criticality property by developing a formal overload-resilience metric called ductility. The generality of ductility enables it to evaluate any scheduling algorithm from the perspective of mixed-criticality cyber-physical systems. In distributed cyber-physical systems, this ductility is the result of both the task-to-processor packing (a.k.a bin packing) and the uniprocessor scheduling algorithms used. In this paper, we present a ductility-maximization packing algorithm to complement our previous work on mixed-criticality uniprocessor scheduling. Our packing algorithm, known as Compress-on-Overload Packing (COP) is a criticality-aware greedy bin-packing algorithm that maximizes the tolerance of high-criticality tasks to overloads. We compare the ductility of COP against the Worst-Fit Decreasing (WFD) bin-packing heuristic used traditionally for load balancing in distributed systems, and show that the performance of COP dominates WFD in the average case and can reach close to five times better ductility when resources are limited. Finally, we illustrate the practical use of COP in distributed cyber-physical systems using a radar surveillance application, and provide an overview of the entire process from assigning task criticality levels to evaluating its performance Karthik Lakshmanan, Dionisio de Niz, Ragunathan Rajkumar, Gabriel A. Moreno |
ICDCS | 3 |
| 2010 | U-connect: a low-latency energy-efficient asynchronous neighbor discovery protocolabstractMobile sensor nodes can be used for a wide variety of applications such as social networks and location tracking. An important requirement for all such applications is that the mobile nodes need to actively discover their neighbors with minimal energy and latency. Nodes in mobile networks are not necessarily synchronized with each other, making the neighbor discovery problem all the more challenging. In this paper, we propose a neighbor discovery protocol called U-Connect, which achieves neighbor discovery at minimal and predictable energy costs while allowing nodes to pick dissimilar duty-cycles. We provide a theoretical formulation of this asynchronous neighbor discovery problem, and evaluate it using the power-latency product metric. We analytically establish that U-Connect is an 1.5-approximation algorithm for the symmetric asynchronous neighbor discovery problem, whereas existing protocols like Quorum and Disco are 2-approximation algorithms. We evaluate the performance of U-Connect and compare the performance of U-Connect with that of existing neighbor discovery protocols. We have implemented U-Connect on our custom portable FireFly Badge hardware platform. A key aspect of our implementation is that it uses a slot duration of only 250μs, and achieves orders of magnitude lower latency for a given duty cycle compared to existing schemes for wireless sensor networks. We provide experimental results from our implementation on a network of around 20 sensor nodes. Finally, we also describe a Friend-Finder application that uses the neighbor discovery service provided by U-Connect. Arvind Kandhalu, Karthik Lakshmanan, Ragunathan Rajkumar |
IPSN | 3 |
| 2010 | Scheduling Self-Suspending Real-Time Tasks with Rate-Monotonic PrioritiesabstractRecent results have shown that the feasibility problem of scheduling periodic tasks with self-suspensions is NP-hard in the strong sense. We observe that a variation of the problem statement that includes sporadic tasks instead of periodic tasks results in a simple characterization of the critical scheduling instant. This in turn leads to an exact characterization of the critical instant for self-suspending tasks with respect to the interference (preemption) from higher-priority sporadic tasks. Using this characterization, we provide pseudo-polynomial response-time tests for analyzing the schedulability of such self-suspending tasks. Self-suspending tasks can also result in more worst-case interference to lower-priority tasks than their equivalent non-suspending counterparts with zero suspension intervals. Hence, we develop a dynamic slack enforcement scheme, which guarantees that the worst-case interference caused by suspending sporadic tasks is no more than the worst-case interference arising from equivalent non-suspending sporadic tasks without suspension intervals. The worst-case response time of self-suspending sporadic tasks themselves is also shown to be unaffected by dynamic slack enforcement, thereby making it optimal. In order to reduce the runtime complexity of slack enforcement, a static slack enforcement scheme is also developed. Empirical analysis of these schemes and the previously studied period enforcement algorithm shows that static slack enforcement achieves within 3% of the breakdown utilization of dynamic slack enforcement, while period enforcement achieves within 14% of dynamic slack enforcement. System designers can take advantage of these different execution control policies depending on their taskset utilizations and implementation constraints. Karthik Lakshmanan, Ragunathan Rajkumar |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |
| 2010 | Scheduling Parallel Real-Time Tasks on Multi-core ProcessorsabstractMassively multi-core processors are rapidly gaining market share with major chip vendors offering an ever increasing number of cores per processor. From a programming perspective, the sequential programming model does not scale very well for such multi-core systems. Parallel programming models such as OpenMP present promising solutions for more effectively using multiple processor cores. In this paper, we study the problem of scheduling periodic real-time tasks on multiprocessors under the fork join structure used in OpenMP. We illustrate the theoretical best-case and worst-case periodic fork-join task sets from a processor utilization perspective. Based on our observations of these task sets, we provide a partitioned preemptive fixed-priority scheduling algorithm for periodic fork-join tasks. The proposed multiprocessor scheduling algorithm is shown to have a resource augmentation bound of 3.42, which implies that any task set that is feasible on m unit speed processors can be scheduled by the proposed algorithm on m processors that are 3:42 times faster. Karthik Lakshmanan, Shinpei Kato, Ragunathan Rajkumar |
RTSS | 3 |
| 2010 | Rate-Harmonized Scheduling and Its Applicability to Energy ManagementabstractThis paper presents a family ofRate-Harmonized Schedulersthat can be used in reservation-based operating systems to naturally cluster task execution and lump processor idle durations. While traditional approaches to energy management have focused on reducingdynamic switching powerthrough Dynamic Voltage and Frequency Scaling (DVFS), processor technology trends predict a future in whichstatic leakage powerwill begin to dominate. To this end, most modern processors provide built-in support for sleep modes withlow leakage power. However, substantial time is required to switch in/out of such sleep modes due to mechanical oscillator stabilization delays. Significant opportunities for energy saving are potentially missed due to idle gaps between executing tasks that are shorter than the time required to enter the sleep mode. Armed with apriori workload information,reservation-basedoperating systems can potentially eliminate such wasted idle durations using Rate-Harmonized Scheduling. AnEnergy-Saving Rate-Harmonized Schedulerguarantees thateveryidle duration can be used to switch into sleep mode. This paper also provides extensions to Rate-Harmonized Scheduling to support multicore processors. Empirical evaluation results are provided from an implementation in the nano-RK operating system for wireless sensor networks. Energy-Saving Rate-Harmonized Scheduling saves 16.8% energy compared to conventional Rate-Monotonic Scheduling for the task set used inSensor Andrewproject. At low utilization levels, Energy-Saving Rate-Harmonized Scheduling can save up to 39% energy on randomly generated task sets. Anthony Rowe 0001, Karthik Lakshmanan, Haifeng Zhu 0001, Ragunathan Rajkumar |
IEEE Trans. Ind. Informatics | 4 |
| 2009 | Partitioned Fixed-Priority Preemptive Scheduling for Multi-core ProcessorsabstractEnergy and thermal considerations are increasingly driving system designers to adopt multi-core processors. In this paper, we consider the problem of scheduling periodic real-time tasks on multi-core processors using fixed-priority preemptive scheduling. Specifically, we focus on the partitioned (static binding) approach, which statically allocates tasks to processing cores. The well-established 50% bound for partitioned multiprocessor scheduling [10] can be overcome by task-splitting (TS) [19], which allows a task to be split across more than one core. We prove that a utilization bound of 60% per core can be achieved by the partitioned deadline-monotonic scheduling (PDMS) class of algorithms on implicit-deadline task sets, when the highest-priority task on each processing core is allowed to be split (HPTS). Given the widespread usage of fixed-priority scheduling in commercial real-time and non real-time operating systems (e.g. VxWorks, Linux), establishing such utilization bounds is both relevant and useful. We also show that a specific instance of PDMS_ HPTS, where tasks are allocated in the decreasing order of size, called PDMS_HPTS_ DS, has a utilization bound of 65% on implicit deadline task-sets. The PDMS_ HPTS_ DS algorithm also achieves a utilization bound of 69% on lightweight implicit-deadline task-sets where no single task utilization exceeds 41.4%. The average-case behavior of PDMS_ HPTS_ DS is studied using randomly generated task-sets, and it is seen to have an average schedulable utilization of 88%. We also characterize the overhead of task-splitting using measurements on an Intel Core 2 Duo processor. Karthik Lakshmanan, Ragunathan Rajkumar, John P. Lehoczky |
ECRTS | 2 |
| 2009 | Demo abstract: The Sensor Andrew infrastructure for large-scale campus-wide sensing and actuation
Anthony Rowe 0001, Mario Berges, Gaurav Bhatia, Ethan Goldman, Ragunathan Rajkumar, Lucio Soibelman |
IPSN | 5 |
| 2009 | Real-Time Video Surveillance over IEEE 802.11 Mesh NetworksabstractIn recent years, there has been an increase in video surveillance systems in public and private environments due to a heightened sense of security. The next generation of surveillance systems will be able to annotate video and locally coordinate the tracking of objects while multiplexing hundreds of video streams in real-time. In this paper, we present OmniEye, a wireless distributed real-time surveillance system composed of wireless smart cameras. OmniEye is comprised of custom-designed smart camera nodes called DSPcams that communicate using an IEEE 802.11 mesh network. These cameras provide wide-area coverage and local processing with the ability to direct a sparse number of high-resolution pan, tilt and zoom (PTZ) cameras that can home onto targets of interest. Each DSPcam performs local processing to help classify events and pro-actively draw an operator's attention when necessary. In video-streaming applications, maintaining high network utilization is required in order to maximize image quality as well as the number of cameras. Our experiments show that by using the standard 802.11 DCF MAC protocol for communication, the system does not scale beyond 5-6 cameras while each camera is streaming at 1 Mbps. Also, we see high levels of jitter in video transmissions. This performance degrades further for multi-hop scenarios due to the presence of hidden nodes. In order to improve the system's scalability and reliability, we propose a Time-Synchronized Application- level MAC protocol (TSAM) capable of operating on top of existing 802.11 protocols using commodity off-the-shelf hardware. Through analysis and experimental validation, we show how TSAM is able to improve throughput and provide bounded delay. Unlike traditional CSMA-based systems, TSAM gracefully degrades in a fair manner so that existing streams can still deliver data. Arvind Kandhalu, Anthony Rowe 0001, Ragunathan Rajkumar, Chingchun Huang, Chao-Chun Yeh |
IEEE Real-Time and Embedded Technology and Applications Symposium | 3 |
| 2009 | Coordinated Task Scheduling, Allocation and Synchronization on MultiprocessorsabstractChip-multiprocessors represent a dominant new shift in the field of processor design. Better utilization of such technology in the real-time context requires coordinated approaches to task allocation, scheduling, and synchronization. In this paper, we characterize various scheduling penalties arising from multiprocessor task synchronization, including (i) blocking delays on global critical sections, (ii) back-to-back execution due to jitter from blocking, and (iii) multiple priority inversions due to remote resource sharing. We analyze the impact of these scheduling penalties under different execution control policies (ECPs) which compensate for the scheduling penalties incurred by tasks due to remote blocking. Subsequently, we develop a synchronization-aware task allocation algorithm for explicitly accommodating these global task synchronization penalties. The key idea of our algorithm is to bundle tasks that access a common shared resource and co-locate them, thereby transforming global resource sharing into local sharing. This approach reduces the above-mentioned penalties associated with remote task synchronization. Experimental results indicate that such a coordinated approach to scheduling, allocation, and synchronization yields significant benefits (as much as 50% savings in terms of required number of processing cores). An implementation of this approach is available as a part of our RT-MAP library, which uses the pthreads implementation of Linux-2.6.22. Karthik Lakshmanan, Dionisio de Niz, Ragunathan Rajkumar |
RTSS | 3 |
| 2009 | On the Scheduling of Mixed-Criticality Real-Time Task SetsabstractThe functional consolidation induced by the cost reduction trends in embedded systems can force tasks of different criticality (e.g. ABS Brakes with DVD) to share a processor and interfere with each other. These systems are known as mixed criticality systems. While traditional temporal isolation techniques prevent all inter-task interference, they waste utilization because they need to reserve for the absolute worst-case execution time (WCET) for all tasks. In many mixed-criticality systems the WCET is not only rare, but at times difficult to calculate, such as the time to localize all possible objects in an obstacle avoidance algorithm. In this situation it is more appropriate to allow the execution time to grow by stealing cycles from lower-criticality tasks. Even more crucial is the fact that temporal isolation techniques can stop a high-criticality task (that was overrunning its nomimal WCET) to allow a low-criticality task to run, making the former miss its deadline. We identify this as the criticality inversion problem. In this paper, we characterize the criticality inversion problem and present a new scheduling scheme called zero-slack scheduling that implements an alternative protection scheme we refer to as asymmetric protection. This protection only prevents interference from lower-criticality to higher-criticality tasks and improves the schedulable utilization. We use an offline algorithm with two parts: a zero-slack calculation algorithm, and a slack analysis algorithm. The zero-slack calculation algorithm minimizes the utilization needed by a task set by reducing the time low-criticality tasks are preempted by high-criticality ones. This algorithm can be used with priority-based preemptive schedulers (e.g. RMS, EDF). The slack analysis algorithm is specific for each priority-based preemptive scheduler and we develop and evaluated the one for RMS. We prove that this algorithm provides the same level of protection against criticality inversion as the best known priority assignment for this purpose, criticality as priority assignment (CAPA). We also prove that zero-slack RM provides the same level of schedulable utilization as RMS when all tasks have equal criticality levels. Finally, we present our implementation of the runtime enforcement mechanisms in Linux/RK to demonstrate its practicality. Dionisio de Niz, Karthik Lakshmanan, Ragunathan Rajkumar |
RTSS | 3 |
| 2009 | Low-power clock synchronization using electromagnetic energy radiating from AC power linesabstractClock synchronization is highly desirable in many sensor networking applications. It enables event ordering, coordinated actuation, energy-efficient communication and duty cycling. This paper presents a novel low-power hardware module for achieving global clock synchronization by tuning to the magnetic field radiating from existing AC power lines. This signal can be used as a global clock source for battery-operated sensor nodes to eliminate drift between nodes over time even when they are not passing messages. With this scheme, each receiver is frequency-locked with each other, but there is typically a phase-offset between them. Since these phase offsets tend to be constant, a higher-level compensation protocol can be used to globally synchronize a sensor network. We present the design of an LC tank receiver circuit tuned to the AC 60Hz signal which we call a Syntonistor. The Syntonistor incorporates a low-power microcontroller that filters the signal induced from AC power lines generating a pulse-per-second output for easy interfacing with sensor nodes. The hardware consumes less than 58μW which is 2--3 times lower than the idle state of most sensor networking MAC protocols. Next, we evaluate a software clock-recovery technique running on the local microcontroller that minimizes timing jitter and provides robustness to noise. Finally, we provide a protocol that sets a global notion of time by accounting for phase-offsets. We evaluate the synchronization accuracy and energy performance as compared to in-band message passing schemes. The use of out-of-band signals for clock synchronization has the useful property of decoupling the synchronization scheme from any particular MAC protocol. Our experiments show that over a 11 day period, eight nodes distributed across the floor of the CIC building on Carnegie Mellon's campus remained synchronized on an average to less than 1ms without exchanging any radio messages beyond the initialization phase. Anthony Rowe 0001, Vikram Gupta, Ragunathan Rajkumar |
SenSys | 3 |
| 2008 | Distributed Resource Kernels: OS Support for End-To-End Resource IsolationabstractThe notion of resource reservation for obtaining real-time scheduling guarantees and enforcement of resource usage has gained strong support in recent years. However, much work on resource reservation has primarily focused on single-processor systems. In this paper, we propose the distributed resource kernel frame wo rk to deploy distributed real-time applications with end-to-end timing constraints, and to efficiently enforce and monitor their usage. Modern distributed real-time systems host multiple applications, where each application can span two or more processors. Timing bugs in one distributed application can affect the timing properties of other applications in the system. Our framework introduces the abstraction of a distributed resource container as an isolated virtual operating environment for a distributed real-time application. We have implemented this framework by extending our open-source single- node Linux/RK platform (R. Rajkumar et al., 1998). A deployment and monitoring tool called dMon is also provided. We evaluate the framework's ability to provide timing guarantees by stress-testing the system using the Distributed Hartstone benchmarks. An audio processing pipeline is then used to illustrate the temporal isolation support provided by the Distributed RK framework. The distributed container abstraction can also be extended in the future to support security and fault-tolerance attributes. Karthik Lakshmanan, Ragunathan Rajkumar |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |
| 2008 | Coexistence of Real-Time and Interactive & Batch Tasks in DVS SystemsabstractInteractive and batch tasks typically have aperiodic random demands and arrival patterns. Generally, interactive tasks are assigned high priority for high responsiveness. Batch tasks with less timing criticality are scheduled in background. Unfortunately, most real-time DVS algorithms focus only on the real-time task workload and timing constraints in determining the operating power-optimized clock frequency. This approach can often leave insufficient cycles for servicing interactive and batch tasks and lead to unacceptable tardiness in conventional applications. We present a power-management framework which ensures that conventional applications will obtain acceptable response times and workload throughput without breaking the temporal constraints of real-time tasks that use resource reservation. We propose two solutions: Background-Preservingand Background-On-Demand algorithms. The first scheme is straightforward and increases the clock frequencies of all tasks to accommodate a future non-real-time workload. The second scheme assigns two modes of frequencies to each task, normal mode and turbo mode. The turbo mode is triggered by the presence of a pending non-real-time task in the system. We also provide the integrated versions of both schemes with our dynamic slack reclamation DVS scheme, called the Progressive algorithm. The integrated versions exploit the slack time from underused reserves for saving more power without performance degradation in all applications. Saowanee Saewong, Ragunathan Rajkumar |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |
| 2008 | Rate-Harmonized Scheduling for Saving EnergyabstractEnergy consumption continues to be a major concern in multiple application domains including power-hungry data centers, portable and wearable devices, mobile communication devices and wireless sensor networks. While energy-constrained, many such applications must meet timing and QoS constraints for sensing, actuation or multimedia data processing. Many modern power-aware processors and microcontrollers have built-in support for active, idle and sleep operating modes. In sleep mode, substantially more energy savings can be obtained but it requires a significant amount of time to switch into and out of that mode. Hence, a significant amount of energy is lost due to idle gaps between executing tasks that are shorter than the required time for the processor to enter the sleep mode. We present a technique called rate-harmonized scheduling that naturally clusters task execution such that processor idle times are lumped together. We next introduce the energy-saving rate-harmonized scheduler which guarantees that every idle duration on the processor can be used to put the processor into sleep mode. This property can be used to even eliminate the idle power mode in processors but nevertheless it is predictable, analyzable, and saves more energy. We finally evaluate the practical benefits of rate-harmonized scheduling implemented in the nano-RK real-time operating system [1] for wireless sensor networks. Anthony Rowe 0001, Karthik Lakshmanan, Haifeng Zhu 0001, Ragunathan Rajkumar |
RTSS | 4 |
| 2008 | RT-Link: A global time-synchronized link protocol for sensor networks
Anthony Rowe 0001, Rahul Mangharam, Ragunathan Rajkumar |
Ad Hoc Networks | 3 |
| 2008 | MEERA: Cross-Layer Methodology for Energy Efficient Resource Allocation in Wireless NetworksabstractIn many portable devices, wireless network interfaces consume upwards of 30% of scarce system energy. Reducing the transceiver's power consumption to extend the system lifetime has therefore become a design goal. Our work is targeted at this goal and is based on the following two observations. First, conventional energy management approaches have focused independently on minimizing the fixed energy cost (by shutdown) and on scalable energy costs (by leveraging, for example, the modulation, code-rate and transmission power). These two energy management approaches present a tradeoff. For example, lower modulation rates and transmission power minimize the variable energy component, but this shortens the sleep duration thereby increasing fixed energy consumption. Second, in order to meet the quality of service (QoS) timeliness requirements for multiple users, we need to determine to what extent each system in the network may sleep and scale. Therefore, we propose a two-phase methodology that resolves the sleep-scaling tradeoff across the physical, communications and link layers at design time and schedules nodes at runtime with near optimal energy-efficient configurations in the solution space. As a result, we are able to achieve very low run-time overheads. Our methodology is applied to a case study on delivering a guaranteed QoS for multiple users with MPEG-4 video over a slow-fading channel. By exploiting runtime controllable parameters of actual RF components and a modified 802.11 medium access controller, system lifetime is increased by a factor of 3-to-10 in comparison with conventional techniques. Sofie Pollin, Rahul Mangharam, Bruno Bougard, Liesbet Van der Perre, Ingrid Moerman, Ragunathan Rajkumar, Francky Catthoor |
IEEE Trans. Wirel. Commun. | 6 |
| 2007 | FireFly Mosaic: A Vision-Enabled Wireless Sensor Networking SystemabstractWith the advent of CMOS cameras, it is now possible to make compact, cheap and low-power image sensors capable of on-board image processing. These embedded vision sensors provide a rich new sensing modality enabling new classes of wireless sensor networking applications. In order to build these applications, system designers need to overcome challenges associated with limited bandwidth, limited power, group coordination and fusing of multiple camera views with various other sensory inputs. Real-time properties must be upheld if multiple vision sensors are to process data, communicate with each other and make a group decision before the measured environmental feature changes. In this paper, we present FireFly Mosaic, a wireless sensor network image processing framework with operating system, networking and image processing primitives that assist in the development of distributed vision-sensing tasks. Each FireFly Mosaic wireless camera consists of a FireFly (Rowe et al., 2006) node coupled with a CMUcam3 (Rowe et al., 2007) embedded vision processor. The FireFly nodes run the nano-RK (Eswaran et al., 2005) real-time operating system and communicate using the RT-link (Rowe et al., 2006) collision-free TDMA link protocol. Using FireFly Mosaic, we demonstrate an assisted living application capable of fusing multiple cameras with overlapping views to discover and monitor daily activities in a home. Using this application, we show how an integrated platform with support for time synchronization, a collision-free TDMA link layer, an underlying RTOS and an interface to an embedded vision sensor provides a stable framework for distributed real-time vision processing. To the best of our knowledge, this is the first wireless sensor networking system to integrate multiple coordinating cameras performing local processing. Anthony Rowe 0001, Dhiraj Goel, Ragunathan Rajkumar |
RTSS | 3 |
| 2007 | Using micro-climate sensing to enhance RF localization in assisted living environmentsabstractIn this paper, we propose micro-climate sensing as an effective means of enhancing conventional RF-based localization. Our system targets people-tracking applications in dynamic indoor environments, such as nursing homes, hospitals and office spaces that require simple deployment and where conventional RF-based tracking using signal strengths alone is very likely to suffer from time-varying signal attenuation and inevitable changes in the environment over time such as new furniture arrangements, people traffic, changing obstacle patterns etc. To help mitigate these effects, we use time-synchronized windows of sensor samples to dynamically associate a mobile node with its nearest beacon nodes. Comparisons and matches are always relative to the ambient attributes at the time of localization, and hence our technique automatically evolves with environmental changes. We consider this property of localization techniques to be a significant contribution and a necessary requirement for any long- lived localization system. In assisted-living environments, sensor networks likely already have basic sensors to collect contextual information about users and to monitor the environment. We propose using light, humidity, temperature and audio data samples over a short window of time to model the micro-climate of a beacon node. Using microclimate matching in conjunction with RF signal strength decreases the worst-case localization error significantly by a factor of more than 3 (from 25 m to 8 m) while making the system more resilient to environment changes. Microclimate data helps ensure at least room level location tracking even in buildings like hospitals with many rooms in close proximity. Anthony Rowe 0001, Zane Starr, Ragunathan Rajkumar |
SMC | 3 |
| 2007 | FireFly: a cross-layer platform for real-time embedded wireless networks
Rahul Mangharam, Anthony Rowe 0001, Ragunathan Rajkumar |
Real Time Syst. | 3 |
| 2007 | MEERA: cross-layer methodology for energy efficient resource allocation in wireless networksabstractIn many portable devices, wireless network interfaces consume upwards of 30% of scarce system energy. Reducing the transceiver's power consumption to extend the system lifetime has therefore become a design goal. Our work is targeted at this goal and is based on the following two observations. First, conventional energy management approaches have focused independently on minimizing the fixed energy cost (by shutdown) and on scalable energy costs (by leveraging, for example, the modulation, code-rate and transmission power). These two energy management approaches present a tradeoff. For example, lower modulation rates and transmission power minimize the variable energy component, but this shortens the sleep duration thereby increasing fixed energy consumption. Second, in order to meet the quality of service (QoS) timeliness requirements for multiple users, we need to determine to what extent each system in the network may sleep and scale. Therefore, we propose a two-phase methodology that resolves the sleep-scaling tradeoff across the physical, communications and link layers at design time and schedules nodes at runtime with near optimal energy-efficient configurations in the solution space. As a result, we are able to achieve very low run-time overheads. Our methodology is applied to a case study on delivering a guaranteed QoS for multiple users with MPEG-4 video over a slow-fading channel. By exploiting runtime controllable parameters of actual RF components and a modified 802.11 medium access controller, system lifetime is increased by a factor of 3-to-10 in comparison with conventional techniques Sofie Pollin, Rahul Mangharam, Bruno Bougard, Liesbet Van der Perre, Ingrid Moerman, Ragunathan Rajkumar, Francky Catthoor |
IEEE Trans. Wirel. Commun. | 6 |
| 2006 | MAX: A Maximal Transmission Concurrency MAC for Wireless Networks with Regular StructureabstractMulti-hop wireless networks facilitate applications in metropolitan area broadband, home multimedia, surveillance and industrial control networks. Many of these applications require high end-to- end throughput and/or bounded delay. Random access link-layer protocols such as carrier sense multiple access (CSMA) which are widely used in single-hop networks perform poorly in the multi-hop regime and provide no end-to-end QoS guarantees. The primary causes for their poor performance are uncoordinated interference and unfairness in exclusive access of the shared wireless medium. Furthermore, random access schemes do not leverage spatial reuse effectively and require routes to be link- aware. In this paper, we propose and study MAX, a time-division- multiplexed resource allocation framework for multi-hop networks with regular topologies. MAX tiling delivers optimal end-to-end throughput across arbitrarily large regularly structured networks while providing bounded delay. It outperforms CSMA-based random access protocols by a factor of 5 to 8. The MAX approach also supports network services including flexible uplink and downlink bandwidth management, deterministic route admission control, and optimal gateway placement. MAX has been implemented on IEEE 802.15.3 embedded nodes and a test-bed of 50 nodes has been deployed both indoors and outdoors. Rahul Mangharam, Ragunathan Rajkumar |
BROADNETS | 2 |
| 2006 | GrooveNet: A Hybrid Simulator for Vehicle-to-Vehicle NetworksabstractVehicular networks are being developed for efficient broadcast of safety alerts, real-time traffic congestion probing and for distribution of on-road multimedia content. In order to investigate vehicular networking protocols and evaluate the effects of incremental deployment it is essential to have a topology-aware simulation and test-bed infrastructure. While several traffic simulators have been developed under the intelligent transport system initiative, their primary motivation has been to model and forecast vehicle traffic flow and congestion from a queuing perspective. GrooveNet is a hybrid simulator which enables communication between simulated vehicles, real vehicles and between real and simulated vehicles. By modeling inter-vehicular communication within a real street map-based topography it facilitates protocol design and also in-vehicle deployment. GrooveNet's modular architecture incorporates mobility, trip and message broadcast models over a variety of link and physical layer communication models. It is easy to run simulations of thousands of vehicles in any US city and to add new models for networking, security, applications and vehicle interaction. GrooveNet supports multiple network interfaces, GPS and events triggered from the vehicle's on-board computer. Through simulation, we are able to study the message latency, and coverage under various traffic conditions. On-road tests over 400 miles lend insight to required market penetration Rahul Mangharam, Daniel S. Weller, Ragunathan Rajkumar, Priyantha Mudalige, Fan Bai 0002 |
MobiQuitous | 3 |
| 2006 | Voice over Sensor NetworksabstractWireless sensor networks have traditionally focused on low duty-cycle applications where sensor data are reported periodically in the order of seconds or even longer. This is due to typically slow changes in physical variables, the need to keep node costs low and the goal of extending battery lifetime. However, there is a growing need to support real-time streaming of audio and/or low-rate video even in wireless sensor networks for use in emergency situations and short-term intruder detection. In this paper, we present FireFly, a time-synchronized sensor network platform for real-time data streaming across multiple hops. FireFly is composed of several integrated layers including specialized low-cost hardware, a sensor network operating system, a real-time link layer and network scheduling which together provide efficient support for applications with timing constraints. In order to achieve high end-to-end throughput, bounded latency and predictable lifetime, we employ hardware-based time synchronization. Multiple tasks including audio sampling, networking and sensor reading are scheduled using the nano-RK RTOS. We have implemented RT-Link, a TDMA-based link layer protocol for message exchange on well-defined time slots and pipelining along multiple hops. We use this platform to support 2-way audio streaming concurrently with sensing tasks. For interactive voice, we investigate TDMA-based slot scheduling with balanced bi-directional latency while meeting audio timeliness requirements. Finally, we describe our experimental deployment of 42 nodes in a coal mine, and present measurements of the end-to-end throughput, jitter, packet loss and voice quality Rahul Mangharam, Anthony Rowe 0001, Ragunathan Rajkumar, Ryohei Suzuki |
RTSS | 3 |
| 2006 | RT-Link: A Time-Synchronized Link Protocol for Energy- Constrained Multi-hop Wireless NetworksabstractWe propose RT-link, a time-synchronized link protocol for real-time wireless communication in industrial control, surveillance and inventory tracking. RT-link provides predictable lifetime for battery-operated embedded nodes, bounded end-to-end delay across multiple hops, and collision-free operation. We investigate the use of hardware-based time-synchronization for infrastructure nodes by using an AM carrier-current radio for indoors and atomic clock receivers for outdoors. Mobile nodes are synchronized via in-band software synchronization within the same framework. We identify three key observations in the design and deployment of RT-link: (a) hardware-based global-time synchronization is a robust and scalable option to in-band software-based techniques, (b) achieving global time-synchronization is both economical and convenient for indoor and outdoor deployments, (c) RT-link achieves a practical lifetime of over 2 years. Through analysis and simulation, we show that RT-link outperforms energy-efficient link protocols such as B-MAC in terms of node lifetime and end-to-end latency. The protocol supports flexible services such as on-demand end-to-end rate control and logical topology control. We implemented RT-link on the CMU FireFly sensor platform and have integrated it within the nano-RK real-time sensor OS. A 42-node network with sub-20 mus synchronization accuracy has been deployed for 3 weeks in the NIOSH Mining Research Laboratory and within two 5-story campus buildings Anthony Rowe 0001, Rahul Mangharam, Ragunathan Rajkumar |
SECON | 3 |
| 2006 | Integrated QoS-aware resource management and scheduling with multi-resource constraints
Ragunathan Rajkumar, Jeffery P. Hansen, John P. Lehoczky |
Real Time Syst. | 2 |
| 2005 | Energy-Aware Memory Firewalling for QoS-Sensitive ApplicationsabstractThis paper presents operating system abstractions for managing physical memory and paging that can be used to improve both timing predictability and the run-time performance of soft real-time tasks. First, we propose a memory reservation scheme which allows any application to reserve a portion of the total system memory pages for its exclusive use. If the application's memory needs exceed its memory reservation, its pages are swapped within its own reservation, thereby containing the performance effects of its memory access profile to its reservation. Swap space for the application is also reserved. A memory reservation can be shared by multiple threads/applications, and reservations can be used hierarchically, with children using only a portion of their parent's reservation. Next, we propose a novel methodology to determine reservation sizes for an embedded task-set that optimizes the overall performance of the system. We also show how an application can leverage customized predictable page-replacement policies to minimize performance penalties from avoidable page faults ("capacity misses"). Anand Eswaran, Ragunathan Rajkumar |
ECRTS | 2 |
| 2005 | Is there an optimum dynamic load balancing scheme?abstractSeveral dynamic load balancing schemes have been proposed in the literature to solve the hot spot problem in cellular networks. In O. K. Tonguz and E. Yanmaz (2003), a unified theoretical framework for the performance analysis of dynamic load balancing schemes in cellular networks have been developed, and closed-form expressions for the call blocking probability of several existing dynamic load balancing schemes have been derived. However, to the best of our knowledge, an analysis showing which one of these dynamic load balancing schemes is optimum, in terms of the achievable call blocking probability, has not been reported before. In this paper, we explore, for the first time, if a global optimum dynamic load balancing scheme exists, and we examine if and under which conditions the existing dynamic load balancing schemes can achieve optimality. Evsen Yanmaz, Ozan K. Tonguz, Ragunathan Rajkumar |
GLOBECOM | 3 |
| 2005 | Optimal fixed and scalable energy management for wireless networksabstractIn many devices, wireless network interfaces consume upwards of 30% of scarce portable system energy. Extending the system lifetime by minimizing communication power consumption has therefore become a priority. Conventional energy management techniques focus independently on minimizing the fixed energy consumption of the transceiver circuit or on scalable transmission control. Fixed energy consumption is reduced by maximizing the transceiver shutdown interval. In contrast, variable transmission rate, coding and power can be leveraged to minimize energy costs. These two energy management approaches present a tradeoff in minimizing the overall system energy. For example, variable energy costs are minimized by transmitting at a lower modulation rate and transmission power, but this also shortens the sleep duration thereby increasing fixed energy consumption. We present a methodology for energy-efficient resource allocation across the physical layer, communications layer and link layer. Our methodology is aimed at providing QoS for multiple users with bursty MPEG-4 video over a time-varying channel. We evaluate our scheme by exploiting control knobs of actual RF components over a modified IEEE 802.11 MAC. Our results indicate that the system lifetime is increased by a factor of 2 to 5 compared to the gains of conventional techniques. Rahul Mangharam, Ragunathan Rajkumar, Sofie Pollin, Francky Catthoor, Bruno Bougard, Liesbet Van der Perre, Ingrid Moerman |
INFOCOM | 2 |
| 2005 | Scalable QoS-Based Resource Allocation in Hierarchical Networked EnvironmentabstractIn this paper, we study the problem of allocating end-to-end bandwidth to each of multiple traffic flows in a large-scale network. We adopt the QoS-based resource allocation model (Q-RAM) (K-S. Lui et al., 2000), whereby each flow derives an utility based on the amount of its allocated bandwidth. Our goal therefore is to maximize the total utility derived across all network flows. The NP-hard nature of the resource allocation problem is compounded by the need to select an appropriate path between each source-destination pair. We propose a hierarchical decomposition scheme that allows the resource allocation problem to be solved in a decentralized and scalable fashion. The hierarchy we use is based on a (natural) partitioning of the network into subnets, with resource allocation decisions made on a subnet-by-subnet basis. A novel distributed transaction scheme is used to ensure that resource allocations are consistent across all the subnets traversed by each flow. We provide both analytical and experimental evidence to show that our scheme is very scalable and yet does not sacrifice the quality of the allocations. Ragunathan Rajkumar, Jeffery P. Hansen, John P. Lehoczky |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |
| 2005 | Diff-EDF: A Simple Mechanism for Differentiated EDF ServiceabstractMany existing and emerging network applications such as voice-over-IP, videoconferencing and online gaming have end-to-end timing requirements. Despite the real-time demands of these applications, they are usually deployed on best-effort networks such as the Internet. This results in unpredictable and often unsatisfactory performance. In this paper we propose a simple and novel task (or packet) scheduling algorithm Diff-EDF (differentiated earliest deadline first) which can meet the real-time needs of these applications while continuing to provide best effort service to nonreal time traffic. In our system we consider each flow as having stochastic traffic characteristics, a stochastic deadline and a maximum allowable miss rate. The Diff-EDF service meets the flow miss rate requirements through the combination of an admission control test and a scheduling algorithm similar to EDF (earliest deadline first). However, unlike standard EDF scheduling each flow receives a deadline bias based on the flow's miss rate requirement. Applying this bias allows the miss rate to be controlled on a flow-by-flow basis. Both the admission control test and the bias selection algorithms can be computed as a linear function of the flow traffic parameters and the logarithms of the miss rate requirements resulting in an efficient implementation. In this paper, we presented the proposed system structure, protocols, algorithms, analysis and experiments. Experiments with randomly generated and real-life data closely match values predicted by the theory. Haifeng Zhu 0001, John P. Lehoczky, Jeffery P. Hansen, Ragunathan Rajkumar |
IEEE Real-Time and Embedded Technology and Applications Symposium | 4 |
| 2005 | Nano-RK: An Energy-Aware Resource-Centric RTOS for Sensor NetworksabstractMany sensor networking applications such as surveillance and environmental monitoring are time-sensitive in nature. To support such applications, we design and implement Nano-RK, a reservation-based real-time operating system (RTOS) with multi-hop networking support for use in wireless sensor networks. We support fixed-priority preemptive multitasking for guaranteeing that task deadlines are met, along with support for CPU and network bandwidth reservations. Tasks can specify their resource demands and the operating system provides timely, guaranteed and controlled access to CPU cycles and network packets in resource-constrained embedded sensor environments. We also introduce the concept of virtual energy reservations that allows the OS to enforce energy budgets associated with a sensing task by controlling resource accesses. A lightweight wireless networking stack supports packet forwarding, routing and TDMA-based network scheduling. Nano-RK has been implemented on the Atmel ATMEGA128 processor with the Chipcon CC2420 802.15.4 transceiver chip. Our results show that a light-weight embedded resource kernel with rich functionality and timing support is practical and constitutes a simple and alternative paradigm for supporting distributed sensing tasks. Anand Eswaran, Anthony Rowe 0001, Ragunathan Rajkumar |
RTSS | 3 |
| 2005 | Multi-Granularity Resource ReservationsabstractResource reservation has been recently supported by many real-time operating systems to provide applications with guaranteed and timely access to system resources. Typically, reservations are based on the worst-case requirements, and therefore can inflate resource demands unnecessarily. Many multimedia applications such as MPEG video streams (1) have high worst-case to average-case demand ratio and (2) can tolerate some deadline misses. To support such applications, we propose a "multi-granularity" reservation model. Instead of the classical {C, T, D} model of resource reservation, the multi-granular reserve specification is given by {{C,T,D},...,{C/sup x/, /spl epsi//sup x/T/sub i/},...,{C/sup y/, /spl epsi//sup y/T/sub i/}} which represents a guarantee of the highest-granularity reserve for C units of resource during every successive periodic interval of T only as long as the resource usage by each of its low-granularity reserves (e.g., C/sup x/ units of resource in every recurring time of /spl epsi//sup x/T/sub i/, /spl epsi//sup x/ /spl isin/ Z/sup +/) is maintained. This multi-granular reservation approach delivers higher system utilization than the pessimistic strategy of worst-case reservation and better temporal isolation than other stochastic and heuristic guarantees in the literature. We perform a detailed schedulability analysis of this model using deadline-monotonic scheduling and derive an appropriate admission control test. We also present detailed analyses and simulation results comparing our reservation scheme for MPEG-4 streams with average-case resource reservation, constant bandwidth server (CBS), and (m, k)-firm guarantee. Saowanee Saewong, Ragunathan Rajkumar |
RTSS | 2 |
| 2005 | Undergraduate embedded system education at Carnegie MellonabstractEmbedded systems encompass a wide range of applications, technologies, and disciplines, necessitating a broad approach to education. We describe embedded system coursework during the first 4 years of university education (the U.S. undergraduate level). Embedded application curriculum areas include: small and single-microcontroller applications, control systems, distributed embedded control, system-on-chip, networking, embedded PCs, critical systems, robotics, computer peripherals, wireless data systems, signal processing, and command and control. Additional cross-cutting skills that are important to embedded system designers include: security, dependability, energy-aware computing, software/systems engineering, real-time computing, and human--computer interaction. We describe lessons learned from teaching courses in many of these areas, as well as general skills taught and approaches used, including a heavy emphasis on course projects to teach system skills. Philip Koopman, Howie Choset, Rajeev Gandhi, Bruce H. Krogh, Diana Marculescu, Priya Narasimhan, JoAnn M. Paul, Ragunathan Rajkumar, Daniel P. Siewiorek, Asim Smailagic, Peter Steenkiste, Donald E. Thomas |
ACM Trans. Embed. Comput. Syst. | 8 |
| 2004 | Resource Management of Highly Configurable TasksabstractSummary form only given. We present an extension to our QoS optimization algorithm, Q-RAM, that can improve optimization time by several orders of magnitude when managing highly configurable tasks. A highly configurable task is one with a large number of QoS dimensions and/or a large number of quality levels on those dimensions. For example, an application that has ten QoS dimensions with ten quality levels each will have 10/sup 10/ setpoints, or ways in which it can be configured. While the existing Q-RAM algorithm has been shown to be a very effective resource management tool, it must still explicitly perform computations on all of the setpoints for each task. For tasks with 10/sup 10/ setpoints or more, this is clearly impractical. The key idea presented here is a new approximation algorithm for the concave majorant step in Q-RAM. By using this algorithm in a filtering step, the best performing subset of the setpoints can be quickly found without explicitly examining all of the setpoints. The idea is validated using a phased array radar system as an example application. Jeffery P. Hansen, Ragunathan Rajkumar, John P. Lehoczky |
IPDPS | 3 |
| 2004 | Design Trade-Offs for Networks with Soft End-to-End Timing ConstraintsabstractAs broadband capabilities on wired, wireless and mobile phone networks proliferate, real-time multimedia traffic is expected to consume higher portions of the available bandwidth. Usage models in such networks range from low-cost VoIP to high-cost, high-quality, high-resolution video streams along with many intermediate data streams. Video flows, however, have stochastic processing requirements, and they may result in very low utilization levels if traditional real-time scheduling techniques are used. This calls for new analytical methods. In this paper, we illustrate a set of new techniques for reasoning about the lateness of such stochastic flows, and apply it to engineer a wired multihop real-time network. We make use of 4 dominant parameters to characterize the design space: end-to-end deadline, acceptable lateness, the number of hops and the network workload. Given any three of these parameters, the remaining parameter can be computed. For instance, if a stochastic flow passes through 10 nodes, its end-to-end deadline is 150 ms, and no more than 0.1% of the packets can be late, what is the maximum allowable system workload that would satisfy the lateness requirements? This maximum allowable workload can then be used to develop for the admission control policy to guarantee that the specified timeliness requirements are met. This paper provides insights into real-time network effects and rules of thumb to properly engineer a high utilization real-time EDF (earliest deadline first) network. Simulation and experimental results validate our methods. Our techniques are further applicable to heterogeneous network cases where some nodes on the flow path are bottlenecks or have cross traffic from other flows. Haifeng Zhu 0001, John P. Lehoczky, Jeffery P. Hansen, Ragunathan Rajkumar |
IEEE Real-Time and Embedded Technology and Applications Symposium | 4 |
| 2004 | Integrated Resource Management and Scheduling with Multi-Resource ConstraintsabstractDynamic real-time systems such as phased-array radars must manage multiple resources, satisfy energy constraints and make frequent on-line scheduling decisions. These systems are hard to manage because task and system requirements change rapidly (e.g. in radar systems, the targets/tasks in the sky are moving continuously) and must satisfy a multitude of constraints. Their highly dynamic nature and stringent time constraints lead to complex cross-layer interactions in these systems. Therefore, the design of such systems has long been a conservative and/or unpredictable mixture of pre-computed schedules, pessimistic resource allocations, cautious energy usage and operator intuition. In this paper, we present an integrated approach that simultaneously maximizes overall system utility, performs task scheduling and satisfies multi-resource constraints. Using a phased-array radar system, we show that our approach can reconfigure settings of 100 tracks at every 0.7 sec in real-time, and performs within 0.1% of the achievable optimal solution. Jeffery P. Hansen, Ragunathan Rajkumar, John P. Lehoczky |
RTSS | 3 |
| 2004 | Real-Time Operating Systems
John A. Stankovic, Ragunathan Rajkumar |
Real Time Syst. | 2 |
| 2003 | Scalable Resource Allocation for Multi-Processor QoS OptimizationabstractWe present scalable QoS optimization algorithms for allocating resources to tasks in a multi-processor environment. Given a set of tasks, each of which is capable of running at one of several different QoS levels, the algorithms can select a QoS operating point, the number of replicas for fault-tolerance and the processors on which to run the replicas so as to maximize overall system QoS. The algorithms are extensions of Q-RAM (QoS-based Resource Allocation Model) [5] and fix two deficiencies with the basic algorithm. The first is that the existing algorithm is weak in making resource trade-off decisions such as to which processor to map a task. The second was that it was not scalable to very large numbers of resources such as in a large multi-processor system. In this paper we present two new algorithms which significantly enhance the ability of Q-RAM to make resource tradeoff decisions. We also introduce a hierarchical decomposition scheme which enables QoS optimization to be performed on problems with thousands of resources and thousands of tasks. Ragunathan Rajkumar, Jeffery P. Hansen, John P. Lehoczky |
ICDCS | 2 |
| 2003 | Time weaver: a software-through-models framework for embedded real-time systemsabstractEmbedded real-time systems are deployed in a wide range of application domains including transportation systems, automated manufacturing, process control, defense, aerospace, and telecommunications. These systems must satisfy not only logical functional requirements but also para-functional properties such as timeliness, Quality of Service (QoS) and reliability. The cross-cutting behaviors imposed by these para-functional properties and dependencies on operational characteristics (e.g. hardware, OS and middleware platforms used) have traditionally led to hard-to-code, hard-to-understand and hard-to-change software. The net result is that productivity improvements in embedded software development have been miniscule compared to improvements in computing and network technologies. We propose a software-through-models framework for the concurrent construction of behavioral models and executable code. Our framework, which can lead to high degrees of cost-effective reuse of embedded software components, decomposes inter-component relationships with an abstraction named coupler. This decomposition enables the separation of para-functional aspects into multiple semantic dimensions (e.g. timing, event flow, concurrency, fault-tolerance, deployment) that can be modified independent of one another. The impact of changes in one dimension on the realization of other dimensions is automatically projected and managed. Platform dependencies are also captured separately, enabling a code-generation subsystem to re-use the same components across a wide range of heterogeneous platforms and applications. System components can be recursively composed or decomposed. An analyzable software structure is enforced such that the end-to-end timing behavior of the resulting system can be verified. A visual tool called Time Weaver supports the framework and has been used to model avionics systems, automotive systems and signal processing systems. The software can be downloaded from[4]. Dionisio de Niz, Ragunathan Rajkumar |
LCTES | 2 |
| 2002 | Analysis of Hierar hical Fixed-Priority SchedulingabstractReservation-based operating systems provide applications with guaranteed and timely access to system resources. One of their chief benefits is temporal isolation, which prevents the timing mis-behavior of one task from interfering with other tasks. Such a benefit is appealing enough that many systems [2, 8] desire to recursively apply this reservation model to each of their components. This recursive application provides flexible load isolation among applications, users and other high-level resource management entities such as aggregated flows for network bandwith. The hierarchical reservation study can be applied to hierarchical schedulers [5, 6], that support heterogenous scheduling algorithms. We propose and analyze a hierarchical reservation model in the context of fixed-priority scheduling, rate-monotonic and deadline-monotonic, as used in systems such as the Resource Kernel [11]. Detailed schedulability analyses under both deferrable-server and sporadic-server replenishment schemes, including exact completion time tests under hierarchical deadline-monotonic schedulers, are presented. We also derive the least upper scheduling bound for hierarchicalrate-monotonic schedulers. Finally, we describe how to apply multi-reserve PCP [4 ], an extension of the Priority Ceiling Protocol for reservation-based systems, to allow tasks to share non-preemptable resources across the hierarchy. Saowanee Saewong, Ragunathan Rajkumar, John P. Lehoczky, Mark Klein 0003 |
ECRTS | 2 |
| 2002 | Critical power slope: understanding the runtime effects of frequency scalingabstractEnergy efficiency is becoming an increasingly important feature for both mobile and high-performance server systems. Most processors designed today include power management features that provide processor operating points which can be used in power management algorithms. However, existing power management algorithms implicitly assume that lower performance points are more energy efficient than higher performance points. Our empirical observations indicate that for many systems, this assumption is not valid.We introduce a new concept called critical power slope to explain and capture the power-performance characteristics of systems with power management features. We evaluate three systems - a clock throttled Pentium laptop, a frequency scaled PowerPC platform, and a voltage scaled system to demonstrate the benefits of our approach. Our evaluation is based on empirical measurements of the first two systems, and publicly available data for the third. Using critical power slope, we explain why on the Pentium-based system, it is energy efficient to run only at the highest frequency, while on the PowerPC-based system, it is energy efficient to run at the lowest frequency point. We confirm our results by measuring the behavior of a web serving benchmark. Furthermore, we extend the critical power slope concept to understand the benefits of voltage scaling when combined with frequency scaling. We show that in some cases, it may be energy efficient not to reduce voltage below a certain point. Akihiko Miyoshi, Charles Lefurgy, Eric Van Hensbergen, Ramakrishnan Rajamony, Ragunathan Rajkumar |
ICS | 5 |
| 2002 | Optimal Partitioning for Quantized EDF SchedulingabstractThe quantized earliest deadline first (Q-EDF) scheduling policy assumes a limited number of priority levels available to express task (or packet) deadlines. This situation arises in computer and communication systems in which priority levels must be expressed using a limited number of bits. The range of possible deadlines is partitioned into bins, and a task whose deadline lies within the range of a bin is assigned a quantized deadline equal to the left endpoint of the bin. We assume soft real-time tasks arrive according to a renewal process, have random service requirements, and have random deadlines drawn from a probability distribution. Using real-time queueing theory, a general method for computing the long run fraction of tasks that miss their actual deadlines (or their quantized deadlines) is presented. Results are given for uniform deadline distributions, and the optimal bin partitioning is determined for this distribution. The theoretical results show excellent agreement with simulation results. Finally, assuming only the min, mean, and max task deadlines are specified, we determine the worst-case deadline distribution given any specified priority bin choice, then optimize the priority bin choice to minimize the worst-case lateness relative to EDF. For this Q-EDF partition, we determine the number of bins needed to achieve a given performance relative to EDF. We show that for practical cases with soft deadlines, 3 bits are sufficient to achieve nearly the same performance with Q-EDF as with pure EDF. Haifeng Zhu 0001, Jeffery P. Hansen, John P. Lehoczky, Ragunathan Rajkumar |
RTSS | 4 |
| 2001 | Optimization of Quality of Service in Dynamic SystemsabstractResource allocation policies which signi cantly improve Quality of Service (QoS) while minimizing QoS instability are presented. In a dynamic system in which applications are continuously entering and departing the system, QoS instability occurs when the system over-commits resources and can not meet the resource demands of new admission requests. The system is forced to either reject the request or degrade one or more existing tasks. In this paper we introduce three admission control policies and compare their QoS performance/stability tradeo s. We show that by maintaining a small resource reserve modeled as a competing application in the QoS optimization framework, we can achieve QoS levels that are over 90% of the theoretical maximum while reducing instability by one to two orders of magnitude. Jeffery P. Hansen, John P. Lehoczky, Ragunathan Rajkumar |
IPDPS | 3 |
| 2001 | Resource Sharing in Reservation-Based SystemsabstractThe resource-sharing problem in priority-driven realtime systems has been studied at length, with the result that some effective and practical solutions are available for both fixed-priority and dynamic-priority systems. In recent years, real-time operating systems have begun to support the resource reservation paradigm, providing a "temporal isolation" abstraction. However the problem of sharing logical resources across reserved applications has not been extensively studied. In this paper we consider both the theoretical and practical implications of such resource-sharing in reservation-based systems. Moreover we provide some experimental results from the implementation of our proposed schemes in Linux/RK, a "resource kernel" that supports reservations. Dionisio de Niz, Luca Abeni, Saowanee Saewong, Ragunathan Rajkumar |
RTSS | 4 |
| 2000 | Constructing Real-time Group Communication Middleware Using the Resource KernelabstractGroup communication is a widely studied paradigm which is often used in building real-time and fault-tolerant distributed systems. RTCAST is a real-time group communication protocol which has been designed to work with commercial, non-real-time, off-the-shelf hardware and operating systems, such as Solaris, Linux and Windows NT. RTCAST makes probabilistic real-time guarantees based on assumptions about the performance of the underlying system. Unfortunately, the high variability of the access to system resources that these operating systems provide may limit the predictability of the real-time guarantees provided by RTCAST. By taking advantage of a service that provides resource scheduling and reservation in these operating systems, both the hardness and timing granularity of RTCAST's real-time services can be greatly improved. This paper describes an implementation of RTCAST which makes use of the Resource Kernel to provide highly predictable, real-time communication guarantees. Scott Iekel-Johnson, Farnam Jahanian, Akihiko Miyoshi, Dionisio de Niz, Ragunathan Rajkumar |
RTSS | 5 |
| 2000 | Operating system support for the management of hard real-time disk traffic
Anastasio Molano, Ángel Viña, Ragunathan Rajkumar |
J. Syst. Archit. | 3 |
| 2000 | Guest Editor's Introduction: 1997 IEEE Real-Time Technologies and Applications Symposiumabstract—————————— ✦ —————————— 1I NTRODUCTION EAL-TIME systems are computing and communication systems that must deliver their services in timely fashion. Typically, these systems monitor, control, and interact with their physical environment. Classical real-time systems include feedback control systems, radar signal processing systems, avionics and space-based systems. The advent of human-friendly data types in digital computing and communications in the form of voice, audio, and video traffic introduced real-time computing to the mainstream. Two recent trends have accelerated this trend in recent years. First, the ongoing convergence of data communications, telephony, and entertainment demands that the tight timing constraints required by telephony systems be married with the flexibility and heterogeneous traffic of data communications. Second, pervasive (or ubiquitous) computing with sensors and actuators embedded everywhere in the physical environment we live in must also operate on a timely basis. The wide spectrum of timing constraints that real-time systems need to satisfy often requires that resources must be explicitly managed and allocated. Those who are still skeptical of the need to do explicit resource management in a “plentiful resource environment” should run one or more of the following practical tests: Ragunathan Rajkumar |
IEEE Trans. Computers | 1 |
| 1999 | A Scalable Solution to the Multi-Resource QoS ProblemabstractThe problem of maximizing system utility by allocating a single finite resource to satisfy discrete Quality of Service (QoS) requirements of multiple applications along multiple QoS dimensions was studied previously. In this paper we consider the more complex problem of apportioning multiple finite resources to satisfy the QoS needs of multiple applications along multiple QoS dimensions. In other words, each application, such as video-conferencing, needs multiple resources to satisfy its QoS requirements. We evaluate and compare three strategies to solve this provably NP-hard problem. We show that dynamic programming and mixed integer programming compute optimal solutions to this problem but exhibit very long running times. We then adapt the mixed integer programming problem to yield near-optimal results with smaller running times. Finally, we present an approximation algorithm based on a local search technique that is less than 5% away from the optimal solution but which is more than two orders of magnitude faster. Perhaps more significantly, the local search technique turns out to be very scalable and robust as the number of resources required by each application increases. Chen Lee, John P. Lehoczky, Daniel P. Siewiorek, Ragunathan Rajkumar, Jeffery P. Hansen |
RTSS | 4 |
| 1999 | Cooperative Scheduling of Multiple ResourcesabstractObtaining simultaneous and timely access to multiple resources is known to be an NP-complete problem. Complete resource decoupling is, therefore, often used for managing end-to-end delays in distributed real-time system where each processor is scheduled independent of the others. This decoupling approach unfortunately fails when multiple resources must be managed within a single node. Resources such as disk bandwidth and network bandwidth are available on a single node but must be managed by their host processor by means of device drivers, filesystem or protocol services. The host processor acting as a controlling resource, therefore, must play multiple roles. One, it is used by applications on that node. Two, it is used to control and manage other (time-shared) controlled resources including disk bandwidth and network bandwidth. These two roles, unfortunately can often be at odds with one another. In this paper we investigate the problem of co-scheduling controlling and controlled resources. We propose the use of a Cooperative Scheduling Server (CS S), which is a dedicated server that manages one specific controlled resource (like disk bandwidth, network bandwidth, inter-process communication, etc.) while using a controlling resource (like the processor). Two core ideas underlie our approach. First, a single (aperiodic) server is created on a controlling resource (such as a CPU) to handle all local requests for a controlled resource (such as disk bandwidth). This implies that conjuctive admission control must be carried out on both the controlling and controlled resources. Secondly, timing constraints at the application level are partitioned into multiple stages, each of which will be guaranteed to complete on a particular resource. RTFS is a real-time filesystem that provides disk bandwidth guarantees under light CPU loads. With a cooperative scheduling server (FS-CSS) for this disk-based filesystem, disk bandwidth guarantees can be obtained under both heavy CPU and disk workloads. We describe the design and implementation of FS-CSS for providing disk bandwidth guarantees. We conclude with a detailed performance evaluation of FS-CSS. Saowanee Saewong, Ragunathan Rajkumar |
RTSS | 2 |
| 1998 | Dynamic disk bandwidth management and metadata pre-fetching in a real-time file systemabstractThe authors focus on two practical considerations that arise in the design of a real-time file system. Firstly, disk bandwidth management should be dynamic, which in turn would allow a QoS manager to dynamically reallocate disk bandwidth to running applications based on their changing needs. Secondly, real-time access to file system data structures should be deterministic, in order to avoid unexpected latencies when accessing files from disk. These issues have implications to the design of the file system and to its schedulability analysis. They address both these problems and present an implementation in RTFS (Real-Time Filesystem Server), a real-time file system supporting disk bandwidth reservation running on top of the Real-Time Mach microkernel. Finally, quantitative comparisons of actual achieved file system bandwidth and response times are used to validate the approach. Anastasio Molano, Ragunathan Rajkumar, Kanaka Juvva |
ECRTS | 2 |
| 1998 | Practical Solutions for QoS-Based Resource AllocationabstractThe QoS based Resource Allocation Model (Q-RAM) proposed by R. Rajkumar et al. (1998) presented an analytical approach for satisfying multiple quality of service dimensions in a resource constrained environment. Using this model, available system resources can be apportioned across multiple applications such that the net utility that accrues to the end users of those applications is maximized. We present several practical solutions to allocation problems that were beyond the limited scope of Q-RAM. We show that the Q-RAM problem of finding the optimal resource allocation to satisfy multiple QoS dimensions is NP hard. We then present a polynomial solution for this resource allocation problem which yields a solution within a provably fixed and short distance from the optimal allocation. Secondly, Q-RAM dealt mainly with the problem of apportioning a single resource to satisfy multiple QoS dimensions. We study the converse problem of apportioning multiple resources to satisfy a single QoS dimension. In practice, this problem becomes complicated, since a single QoS dimension perceived by the user can be satisfied using different combinations of available resources. We show that this problem can be formulated as a mixed integer programming problem that can be solved efficiently to yield an optimal resource allocation. We also present the run times of these optimizations to illustrate how these solutions can be applied in practice. A good understanding of these solutions will yield insights into the general problem of apportioning multiple resources to satisfy simultaneously multiple QoS dimensions of multiple concurrent applications. Ragunathan Rajkumar, Chen Lee, John P. Lehoczky, Daniel P. Siewiorek |
RTSS | 1 |
| 1997 | Real-time filesystems - Guaranteeing timing constraints for disk accesses in RT-MachabstractTraditional real-time systems have largely avoided the use of disks due to their relative slow speeds and their unpredictability. However, many real-time applications including multimedia systems and real-time database applications benefit significantly from the use of disks to store and access real-time data. We investigate the problem of obtaining guaranteed timely access to files on a disk in a real-time system. Our study focuses on several aspects of this problem of providing a real-time filesystem. First, we consider the use of two real-time disk scheduling algorithms: earliest deadline scheduling and just-in-time scheduling, a variation of aperiodic servers for the disk. The latter algorithm is designed to improve disk throughput that can be hurt when a real-time scheduling algorithm such as EDF is applied directly. Admission control policies with practically acceptable properties of performance and usability are provided. Next, we design and implement a real-time filesystem on the RT-Mach microkernel-based system running a real-time shell. The new interface we develop is based on RT-Mach's resource reservation paradigm and provides guaranteed and timely access for multiple concurrent applications requiring disk bandwidth with different timing and volume requirements. Finally, we perform a detailed performance evaluation of the real-time filesystem including its raw performance. We show the following positive but rather surprising result: our real-time scheduling filesystem not only provides guaranteed and timely access but also does so at relatively high levels of throughput. Traditional disk scheduling algorithms offer completely unacceptable file access latencies for real-time applications and do so only at slightly higher throughput. Anastasio Molano, Kanaka Juvva, Ragunathan Rajkumar |
RTSS | 3 |
| 1997 | A resource allocation model for QoS managementabstractQuality of service (QoS) has been receiving wide attention in many research communities including networking, multimedia systems, real-time systems and distributed systems. In large distributed systems such as those used in defense systems, on-demand service and inter-networked systems, applications contending for system resources must satisfy timing, reliability and security constraints as well as application-specific quality requirements. Allocating sufficient resources to different applications in order to satisfy various requirements is a fundamental problem in these situations. A basic yet flexible model for performance-driven resource allocations can therefore be useful in making appropriate tradeoffs. We present an analytical model for QoS management in systems which must satisfy application needs along multiple dimensions such as timeliness, reliable delivery schemes, cryptographic security and data quality. We refer to this model as Q-RAM (QoS-based Resource Allocation Model). The model assumes a system with multiple concurrent applications, each of which can operate at different levels of quality based on the system resources available to it. The goal of the model is to be able to allocate resources to the various applications such that the overall system utility is maximized under the constraint that each application can meet its minimum needs. We identify resource profiles of applications which allow such decisions to be made efficiently and in real-time. We also identify application utility functions along different dimensions which are composable to form unique application requirement profiles. We use a video-conferencing system to illustrate the model. Ragunathan Rajkumar, Chen Lee, John P. Lehoczky, Daniel P. Siewiorek |
RTSS | 1 |
| 1996 | High availability in the real-time publisher/subscriber inter-process communication modelabstractThe real time publisher/subscriber (RT P/S) communications model has been proposed as a flexible and powerful interprocess communication model for distributed real time systems (R. Rajkumar et al., 1995). It can also be used as the underlying framework for supporting building blocks such as extensible cells and replaceable software units far building evolvable distributed real time systems. However, for the model to be adopted in practice, it must tolerate processor failures and allow repaired processors to rejoin the system on a dynamic basis. Such processor failures and rejoins must also not hurt the efficiency of the steady state publication/subscription of messages by repetitive real time processes. We present extensions to the RT P/S model to support these capabilities. The solution is structured in two layers. First, an efficient processor membership protocol layer, based on F. Cristian's (1988; 1991) periodic broadcast membership protocol, detects processor failures and rejoins. It provides strong semantics and exhibits a finite delay in detecting processor failures. Secondly, idempotence properties, weak interleaving needs and the benign impact of node failures within the RT P/S information structure enable us to transfer consistent state to newly joining daemons and to manage changes to the information elegantly. The changes are orthogonal to the communication programming interface and also maintain very efficient and analyzable steady state real time execution paths. These protocols have been successfully built in the context of both feedback control and multimedia dissemination applications. Ragunathan Rajkumar, Michael Gagliardi |
RTSS | 1 |
| 1995 | Future Distributed Embedded and Real-Time Applications Will Be Adaptive: Meanings, Challenges and Research Paradigms (Panel)abstractSummary form only given, as follows. Static models are not appropriate for next-generation distributed real-time applications that are likely to be adaptive in nature, (for example, to provide a high degree of fault tolerance). During the last few years, the real-time systems community has started to counter this criticism by extending traditional work to cover newer application domains, The central problem remains, however, that the concept of adaptivity is often domain-specific and sometimes ill-defined in the context of bringing distributed real-time systems concept into better focus. Accordingly, be it resolved that future distributed embedded and real-time applications will be adaptive and that meanings, challenges and research paradigms await discovery. The charge to the panel is to defend (or to dismiss as fluff) the above resolution. Aloysius K. Mok, Constance L. Heitmeyer, Kevin Jeffay, Michael B. Jones, C. Douglass Locke, Ragunathan Rajkumar |
ICDCS | 6 |
| 1994 | Generalized rate-monotonic scheduling theory: a framework for developing real-time systemsabstractReal-time computing systems are used to control telecommunication systems, defense systems, avionics, and modern factories. Generalized rate-monotonic scheduling theory, is a recent development that has had large impact on the development of real-time systems and open standards. In this paper we provide an up-to-date and self-contained review of generalized rate-monotonic scheduling theory. We show how this theory can be applied in practical system development, where special attention must be given to facilitate concurrent development by geographically distributed programming teams and the reuse of existing hardware and software components.> Lui Sha, Ragunathan Rajkumar, Shirish S. Sathaye |
Proc. IEEE | 2 |
| 1994 | Runtime Monitoring of Timing Constraints in Distributed Real-Time Systems
Farnam Jahanian, Ragunathan Rajkumar, Sitaram C. V. Raju |
Real Time Syst. | 2 |
| 1993 | Processor Group Membership Protocols: Specification, Design, and ImplementationabstractThe specification, design and implementation of a set of protocols to solve the processor group membership problem in distributed systems are presented. These group membership protocols were developed as part of a toolkit for building distributed/parallel applications on a cluster of workstations. The group membership service forms the lowest layer in the toolkit, and is the glue which unifies all other layers. The membership service supports three distinct protocols: weak, strong, and hybrid. These protocols differ significantly in the level of consistency and the number of messages exchanged in reaching agreement. The modular implementation of these protocols and the optimization techniques used to enhance their performance are described.> Farnam Jahanian, Sameh A. Fakhouri, Ragunathan Rajkumar |
SRDS | 3 |
| 1992 | Monitoring Timing Constraints in Distributed Real-Time SystemsabstractA run-time environment for monitoring distributed real-time systems is described. In particular, the authors focus on the problem of detecting violations of timing assertions in an environment in which the real-time tasks run on multiple processors, and timing constraints can be either interprocessor or intraprocessor constraints. Constraint violations are detected at the earliest possible time by deriving and checking intermediate constraints. If the violations must be detected as early as possible, then the problem of minimizing the number of messages to be exchanged between the processors becomes intractable. The authors characterize a subclass of timing constraints that occur commonly in distributed real-time systems and whose message requirements can be minimized. They also take into account the drift among the various processor clocks when detecting a violation of a timing assertion. Finally, an implementation of a distributed run-time monitor is described.> Sitaram C. V. Raju, Ragunathan Rajkumar, Farnam Jahanian |
RTSS | 2 |
| 1991 | A Real-Time Locking ProtocolabstractThe authors examine a priority driven two-phase lock protocol called the read/write priority ceiling protocol. It is shown that this protocol leads to freedom from mutual deadlock. In addition, a high-priority transactions can be blocked by lower priority transactions for at most the duration of a single embedded transaction. These properties can be used by schedulability analysis to guarantee that a set of periodic transactions using this protocol can always meet its deadlines. Finally, the performance of this protocol is examined for randomly arriving transactions using simulation studies.> Lui Sha, Ragunathan Rajkumar, Sang Hyuk Son, Chun-Hyon Chang |
IEEE Trans. Computers | 2 |
| 1990 | Real-Time Scheduling Support in Futurebus+abstractA simple but efficient architecture for building multiprocessors is to connect several processors to a common backplane bus. The backplane acts as a shared resource in this architecture and contention for its use by different bus modules must be resolved. In a real-time system, this backplane must also provide scheduling support such that the timing behavior of the resulting system is analyzable. In addition, the support primitives for real-time scheduling on a backplane bus must also be constrained by the economic considerations associated with a bus standard that is intended to support both time sharing and real-time applications. The authors review the design considerations to support real-time systems in the IEEE Futurebus+ backplane specification and describe how this backplane can be used to satisfy timing constraints in priority-driven real-time systems.> Lui Sha, John P. Lehoczky, Ragunathan Rajkumar |
RTSS | 3 |
| 1990 | Priority Inheritance Protocols: An Approach to Real-Time SynchronizationabstractAn investigation is conducted of two protocols belonging to the priority inheritance protocols class; the two are called the basic priority inheritance protocol and the priority ceiling protocol. Both protocols solve the uncontrolled priority inversion problem. The priority ceiling protocol solves this uncontrolled priority inversion problem particularly well; it reduces the worst-case task-blocking time to at most the duration of execution of a single critical section of a lower-priority task. This protocol also prevents the formation of deadlocks. Sufficient conditions under which a set of periodic tasks using this protocol may be scheduled is derived.> Lui Sha, Ragunathan Rajkumar, John P. Lehoczky |
IEEE Trans. Computers | 2 |
| 1989 | Mode Change Protocols for Priority-Driven Preemptive Scheduling
Lui Sha, Ragunathan Rajkumar, John P. Lehoczky, Krithi Ramamritham |
Real Time Syst. | 2 |
| 1988 | Real-Time Synchronization Protocols for MultiprocessorsabstractThe authors investigate the synchronization problem in the context of priority-driven preemptive scheduling on shared-memory multiprocessors. Unfortunately, a direct application of synchronization mechanisms such as the Ada rendezvous, semaphores, or monitors can lead to uncontrolled priority inversion: a high job being blocked by a lower priority job for an indefinite period of time. A task allocation scheme based on the generalized protocol is outlined.> Ragunathan Rajkumar, Lui Sha, John P. Lehoczky |
RTSS | 1 |
| 1987 | On Countering the Effects of Cycle-Stealing in a Hard Real-Time Environment
Ragunathan Rajkumar, Lui Sha, John P. Lehoczky |
RTSS | 1 |
| 1986 | Solutions for Some Practical Problems in Prioritized Preemptive Scheduling
Lui Sha, John P. Lehoczky, Ragunathan Rajkumar |
RTSS | 3 |