Bryan Donyanavard

dblp:122/3449 · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
13since 2021 · last 2025
0000-0002-6990-2577ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 3 first-author · 12 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Generating and Predicting Output Perturbations in Image Segmenters
abstract
Image segmentation applications are a core component of safety-critical autonomous software pipelines. Sensor data input noise can lead to segmentation output corruption that threatens safety in both DNN- and transformer-based segmenters. Previous work has proposed methods for generating malicious noise to cause DNN- and transformer-based object detection and classification output corruption. We perform the same task for image segmentation applications using genetic algorithms for optimization. We then propose a novel method to predict whether an input image will yield a corrupted segmentation output due to noise. We evaluate the optimal noise generation and corruption prediction on state-of-the-art image segmenters YOLOv8 and DETR. We observe that we can (a) cause segmentation output corruption with noise that is undetectable to the human eye and unrelated to the corrupted region of the image; and (b) predict output corruption due to image noise with over 96% accuracy.
Matthew Bozoukov, Nguyen Anh Vu Doan, Bryan Donyanavard
DATE3
2025 SCHED: Safe CPU Scheduling Framework with Reinforcement Learning and Decision Trees for Autonomous Vehicles
abstract
Autonomous vehicles (AVs) require consistently low-latency computations; however, operating system (OS) CPU scheduler can lead to high tail latency that threatens timely decision-making and safety. General-purpose OS schedulers prioritize fairness and throughput over individual task deadlines, posing challenges for complex AV workloads with mixed-criticality tasks. Despite extensive research on CPU scheduling, reinforcement learning (RL) has not been explored at the OS level due to feasibility concerns. This paper presents SCHED, the first RL-based CPU scheduling framework optimized for OS integration. SCHED learns a scheduling policy via reinforcement learning and deploys it as a lightweight, decision-tree-based scheduler. We demonstrate that SCHED remains sufficiently fast for OS-level integration compared to existing OS schedulers. Furthermore, we validate SCHED on a realistic autonomous vehicle pipeline, demonstrating its practical potential in maintaining low latency and ensuring safer, more responsive AV operations.
Dongjoo Seo, Changhoon Sung 0001, Ping-Xiang Chen, Bryan Donyanavard, Nikil Dutt
VTC2025-Spring4
2025 Runtime Adaptivity for Efficient Neural Network Inference on Autonomous Systems
abstract
Neural network pruning and dynamic training have emerged as key techniques for optimizing deep learning models to meet the constraints of resource-limited systems. However, achieving both efficiency and adaptability without compromising safety or performance remains a significant challenge in real-time autonomous applications. We present Back to the Future and USA-Nets , two complementary approaches that address this challenge. Back to the Future combines pruning with dynamic routing to enable latency gains and dynamic reconfiguration at runtime, allowing a pruned model to seamlessly revert to the full model when unsafe or anomalous behavior is detected. USA-Nets extend this concept by enabling runtime adaptability through dynamically trained networks that can adjust their width without requiring additional annotated data or excessive storage overhead. Together, these methods deliver significant performance improvements while maintaining safety and flexibility, as evidenced by experimental results demonstrating that Back to the Future achieves a 32× faster reversion time compared to loading the full model, and USA-Nets achieve up to 85% latency reduction with minimal accuracy degradation. These innovations pave the way for efficient, adaptable, and safe deployment of deep learning models in diverse real-time and resource-constrained environments, with future work focusing on advanced pruning techniques and runtime optimizations.
Danny Abraham, Biswadip Maity, Bryan Donyanavard, Nikil Dutt
ACM Trans. Embed. Comput. Syst.3
2025 OASIS: Optimized Adaptive System for Intelligent SLAM
abstract
Visual Simultaneous Localization and Mapping (VSLAM) is essential for mobile autonomous systems operating in complex dynamic environments. VSLAM algorithms are computationally intensive and must execute in real-time on resource-constrained embedded devices. Variations in environmental complexity can lead to longer frame processing times, causing dropped frames, lost localization information, and degraded accuracy. To address these challenges, we introduce OASIS, a novel adaptive approximation method that dynamically reduces input frame areas based on realtime visual importance. Unlike traditional optimizations that require adjusting internal SLAM parameters, OASIS selectively minimizes computation by adaptively filtering less critical image regions, significantly reducing computational load. Evaluations on the EuRoC MAV dataset demonstrate that our approach balances accuracy and system predictability, achieving up to a 71.8% reduction in worst-case pose estimation errors. OASIS offers a significant advancement in reliable, predictable, and energy-efficient SLAM tailored for mobile autonomous robotic applications.
Alles Rebel, Nikil Dutt, Bryan Donyanavard
ACM Trans. Embed. Comput. Syst.3
2024 Back to the Future: Reversible Runtime Neural Network Pruning for Safe Autonomous Systems
abstract
Neural network pruning has emerged as a technique to reduce the size of networks at the cost of accuracy to enable deployment in resource-constrained systems. However, low-accuracy pruned models may compromise the safety of realtime autonomous systems when encountering unpredictable scenarios, e.g., due to anomalous or emergent behavior. We propose Back to the Future: a novel approach that combines pruning with dynamic routing to achieve both latency gains and dynamic reconfiguration to meet desired accuracy at runtime. Our approach enables the pruned model to quickly revert to the full model when unsafe behavior is detected, enhancing safety and reliability. Experimental results demonstrate that our swapping approach is 32× faster than loading the original model from disk, providing seamless reversion to the accurate version of the model, demonstrating its applicability for safe autonomous systems design.
Danny Abraham, Biswadip Maity, Bryan Donyanavard, Nikil Dutt
DATE3
2023 Tutorial: MARS: A Framework for Runtime Monitoring, Modeling, and Management of Realtime Systems
Bryan Donyanavard, Nikil Dutt, Biswadip Maity, Parth Malani, Tiago Rogério Mück
CODES+ISSS1
2023 Lightning Talk: The New Era of Computational Cognitive Intelligence
abstract
The triple whammy of variability in platforms (e.g., process variability), applications (e.g., dynamic use cases in autonomous systems), and the environment (e.g., context) renders ineffective the classical computational/algorithmic/numerical computing paradigm in dealing with the inherent runtime dynamism and uncertainty faced by emerging systems. We posit that this requires a fundamental change from classical "static" computing to a new era that deploys a computational cognitive intelligence (CCI) paradigm that is able to learn and evolve at runtime. The CCI paradigm empowers systems to be adaptable and evolvable by exploiting biologically-inspired cognitive intelligence principles.
Nikil Dutt, Bryan Donyanavard
DAC2
2023 Information Processing Factory 2.0 - Self-awareness for Autonomous Collaborative Systems
abstract
This paper summarizes the talks of a special session on the IPF 2.0 project, a collaborative German-US research project that leverages self-awareness principles for the self-management of distributed systems of autonomous multiprocessor systems-on-chip (MPSoCs).
Nora Sperling, Alex Bendrick, Dominik Stöhrmann, Rolf Ernst, Bryan Donyanavard, Florian Maurer 0003, Oliver Lenke, Anmol Surhonne, Andreas Herkersdorf, Walaa Amer, Caio Batista de Melo, Ping-Xiang Chen, Quang Anh Hoang, Rachid Karami, Biswadip Maity, Paul Nikolian, Mariam Rakka, Dongjoo Seo, Saehanseul Yi, Minjun Seo, Nikil Dutt, Fadi J. Kurdahi
DATE5
2022 ProSwap: Period-aware Proactive Swapping to Maximize Embedded Application Performance
abstract
Linux prevents errors due to physical memory limits by swapping out active application memory from main memory to secondary storage. Swapping degrades application performance due to swap-in/out latency overhead. To mitigate the swapping overhead in periodic applications, we present ProSwap: a period-aware proactive and adaptive swapping policy for em-bedded systems. ProSwap exploits application periodic behavior to proactively swap-out rarely-used physical memory pages, creating more space for active processes. A flexible memory reclamation time-window enables adaptation to memory limitations that vary between applications. We demonstrate ProSwap's efficacy for an autonomous vehicle application scenario executing multi-application pipelines, and show that our policy achieves up to 1.26×performance gain via proactive swapping.
Dongjoo Seo, Biswadip Maity, Ping-Xiang Chen, Dukyoung Yun, Bryan Donyanavard, Nikil Dutt
NAS5
2022 Online Learning for Orchestration of Inference in Multi-user End-edge-cloud Networks
abstract
Deep-learning-based intelligent services have become prevalent in cyber-physical applications, including smart cities and health-care. Deploying deep-learning-based intelligence near the end-user enhances privacy protection, responsiveness, and reliability. Resource-constrained end-devices must be carefully managed to meet the latency and energy requirements of computationally intensive deep learning services. Collaborative end-edge-cloud computing for deep learning provides a range of performance and efficiency that can address application requirements through computation offloading. The decision to offload computation is a communication-computation co-optimization problem that varies with both system parameters (e.g., network condition) and workload characteristics (e.g., inputs). However, deep learning model optimization provides another source of tradeoff between latency and model accuracy. An end-to-end decision-making solution that considers such computation-communication problem is required to synergistically find the optimal offloading policy and model for deep learning services. To this end, we propose a reinforcement-learning-based computation offloading solution that learns optimal offloading policy considering deep learning model selection techniques to minimize response time while providing sufficient accuracy. We demonstrate the effectiveness of our solution for edge devices in an end-edge-cloud system and evaluate with a real-setup implementation using multiple AWS and ARM core configurations. Our solution provides 35% speedup in the average response time compared to the state-of-the-art with less than 0.9% accuracy reduction, demonstrating the promise of our online learning framework for orchestrating DL inference in end-edge-cloud systems.
Sina Shahhosseini, Dongjoo Seo, Anil Kanduri, Sung-Soo Lim, Bryan Donyanavard, Amir-Mohammad Rahmani, Nikil Dutt
ACM Trans. Embed. Comput. Syst.6
2021 Cross-layer Configuration Optimization for Localization on Resource-constrained Devices
abstract
Mobile devices are increasingly expected to sup-port high-performance cyber-physical applications in small form factors, e.g., drones and rovers. However, the gap between hardware limitations of these devices and application requirements is still prohibitive – conflicting goals such as robust, accurate, and efficient execution must be managed carefully to achieve acceptable operation. In this paper, we explore the tradeoff between performance and efficiency in such cyber-physical systems, specifically with respect to localization (a core task for any mobile autonomous device). We perform a design space exploration (DSE) given a number of configurable parameters for both localization algorithm and platform layers. Given the configuration space, we formulate a cross-layer multi-objective optimization problem to explore the tradeoff between localization accuracy and power consumption. We then propose a predictive model for robust execution that can be used to determine desirable configurations at runtime in the face of environmental changes.
Sandra Hernández, José Araújo, Patric Jensfelt, Ananya Muddukrishna, Bryan Donyanavard
IROS6
2021 SEAMS: Self-Optimizing Runtime Manager for Approximate Memory Hierarchies
abstract
Memory approximation techniques are commonly limited in scope, targeting individual levels of the memory hierarchy. Existing approximation techniques for a full memory hierarchy determine optimal configurations at design-time provided a goal and application. Such policies are rigid: they cannot adapt to unknown workloads and must be redesigned for different memory configurations and technologies. We propose SEAMS: the first self-optimizing runtime manager for coordinating configurable approximation knobs across all levels of the memory hierarchy. SEAMS continuously updates and optimizes its approximation management policy throughout runtime for diverse workloads. SEAMS optimizes the approximate memory configuration to minimize energy consumption without compromising the quality threshold specified by application developers. SEAMS can (1) learn a policy at runtime to manage variable application quality of service ( QoS ) constraints, (2) automatically optimize for a target metric within those constraints, and (3) coordinate runtime decisions for interdependent knobs and subsystems. We demonstrate SEAMS’ ability to efficiently provide functions (1)–(3) on a RISC-V Linux platform with approximate memory segments in the on-chip cache and main memory. We demonstrate SEAMS’ ability to save up to 37% energy in the memory subsystem without any design-time overhead. We show SEAMS’ ability to reduce QoS violations by 75% with < 5% additional energy.
Biswadip Maity, Bryan Donyanavard, Anmol Surhonne, Amir-Mohammad Rahmani, Andreas Herkersdorf, Nikil Dutt
ACM Trans. Embed. Comput. Syst.2
2021 Chauffeur: Benchmark Suite for Design and End-to-End Analysis of Self-Driving Vehicles on Embedded Systems
abstract
Self-driving systems execute an ensemble of different self-driving workloads on embedded systems in an end-to-end manner, subject to functional and performance requirements. To enable exploration, optimization, and end-to-end evaluation on different embedded platforms, system designers critically need a benchmark suite that enables flexible and seamless configuration of self-driving scenarios, which realistically reflects real-world self-driving workloads’ unique characteristics. Existing CPU and GPU embedded benchmark suites typically (1) consider isolated applications, (2) are not sensor-driven, and (3) are unable to support emerging self-driving applications that simultaneously utilize CPUs and GPUs with stringent timing requirements. On the other hand, full-system self-driving simulators (e.g., AUTOWARE, APOLLO) focus on functional simulation, but lack the ability to evaluate the self-driving software stack on various embedded platforms. To address design needs, we present Chauffeur, the first open-source end-to-end benchmark suite for self-driving vehicles with configurable representative workloads. Chauffeur is easy to configure and run, enabling researchers to evaluate different platform configurations and explore alternative instantiations of the self-driving software pipeline. Chauffeur runs on diverse emerging platforms and exploits heterogeneous onboard resources. Our initial characterization of Chauffeur on different embedded platforms – NVIDIA Jetson TX2 and Drive PX2 – enables comparative evaluation of these GPU platforms in executing an end-to-end self-driving computational pipeline to assess the end-to-end response times on these emerging embedded platforms while also creating opportunities to create application gangs for better response times. Chauffeur enables researchers to benchmark representative self-driving workloads and flexibly compose them for different self-driving scenarios to explore end-to-end tradeoffs between design constraints, power budget, real-time performance requirements, and accuracy of applications.
Biswadip Maity, Saehanseul Yi, Dongjoo Seo, Leming Cheng, Sung-Soo Lim, Jongchan Kim 0001, Bryan Donyanavard, Nikil Dutt
ACM Trans. Embed. Comput. Syst.7
2020 Emergent Control of MPSoC Operation by a Hierarchical Supervisor / Reinforcement Learning Approach
abstract
MPSoCs increasingly depend on adaptive resource management strategies at runtime for efficient utilization of resources when executing complex application workloads. In particular, conflicting demands for adequate computation performance and power-/energy-efficiency constraints make desired application goals hard to achieve. We present a hierarchical, cross-layer hardware/software resource manager capable of adapting to changing workloads and system dynamics with zero initial knowledge. The manager uses rule-based reinforcement learning classifier tables (LCTs) with an archive-based backup policy as leaf controllers. The LCTs directly manipulate and enforce MPSoC building block operation parameters in order to explore and optimize potentially conflicting system requirements (e.g., meeting a performance target while staying within the power constraint). A supervisor translates system requirements and application goals into per-LCT objective functions (e.g., core instructions-per-second (IPS). Thus, the supervisor manages the possibly emergent behavior of the low-level LCT controllers in response to 1) switching between operation strategies (e.g., maximize performance vs. minimize power; and 2) changing application requirements. This hierarchical manager leverages the dual benefits of a software supervisor (enabling flexibility), together with hardware learners (allowing quick and efficient optimization). Experiments on an FPGA prototype confirmed the ability of our approach to identify optimized MPSoC operation parameters at runtime while strictly obeying given power constraints.
Florian Maurer 0003, Bryan Donyanavard, Amir-Mohammad Rahmani, Nikil Dutt, Andreas Herkersdorf
DATE2
2019 SOSA: Self-Optimizing Learning with Self-Adaptive Control for Hierarchical System-on-Chip Management
abstract
Resource management strategies for many-core systems dictate the sharing of resources among applications such as power, processing cores, and memory bandwidth in order to achieve system goals. System goals require consideration of both system constraints (e.g., power envelope) and user demands (e.g., response time, energy-efficiency). Existing approaches use heuristics, control theory, and machine learning for resource management. They all depend on static system models, requiring a priori knowledge of system dynamics, and are therefore too rigid to adapt to emerging workloads or changing system dynamics.
Bryan Donyanavard, Tiago Rogério Mück, Amir-Mohammad Rahmani, Nikil Dutt, Armin Sadighi, Florian Maurer 0003, Andreas Herkersdorf
MICRO1
2018 SPECTR: Formal Supervisory Control and Coordination for Many-core Systems Resource Management
abstract
Resource management strategies for many-core systems need to enable sharing of resources such as power, processing cores, and memory bandwidth while coordinating the priority and significance of system- and application-level objectives at runtime in a scalable and robust manner. State-of-the-art approaches use heuristics or machine learning for resource management, but unfortunately lack formalism in providing robustness against unexpected corner cases. While recent efforts deploy classical control-theoretic approaches with some guarantees and formalism, they lack scalability and autonomy to meet changing runtime goals. We present SPECTR, a new resource management approach for many-core systems that leverages formal supervisory control theory (SCT) to combine the strengths of classical control theory with state-of-the-art heuristic approaches to efficiently meet changing runtime goals. SPECTR is a scalable and robust control architecture and a systematic design flow for hierarchical control of many-core systems. SPECTR leverages SCT techniques such as gain scheduling to allow autonomy for individual controllers. It facilitates automatic synthesis of the high-level supervisory controller and its property verification. We implement SPECTR on an Exynos platform containing ARM»s big.LITTLE-based heterogeneous multi-processor (HMP) and demonstrate that SPECTR»s use of SCT is key to managing multiple interacting resources (e.g., chip power and processing cores) in the presence of competing objectives (e.g., satisfying QoS vs. power capping). The principles of SPECTR are easily applicable to any resource type and objective as long as the management problem can be modeled using dynamical systems theory (e.g., difference equations), discrete-event dynamic systems, or fuzzy dynamics.
Amir-Mohammad Rahmani, Bryan Donyanavard, Tiago Rogério Mück, Kasra Moazzemi, Axel Jantsch, Onur Mutlu, Nikil Dutt
ASPLOS2
2018 Gain scheduled control for nonlinear power management in CMPs
abstract
Dynamic voltage and frequency scaling (DVFS) is a well-established technique for power management of thermal-or energy-sensitive chip multiprocessors (CMPs). In this context, linear control theoretic solutions have been successfully implemented to control the voltage-frequency knobs. However, modern CMPs with a large range of operating frequencies and multiple voltage levels display nonlinear behavior in the relationship between frequency and power. State-of-the-art linear controllers therefore under-optimize DVFS operation. We propose a Gain Scheduled Controller (GSC) for nonlinear runtime power management of CMPs that simplifies the controller implementation of systems with varying dynamic properties by utilizing an adaptive control theoretic approach in conjunction with static linear controllers. Our design improves the accuracy of the controller over a static linear controller with minimal overhead. We implement our approach on an Exynos platform containing ARM's big.LITTLE-based heterogeneous multi-processor (HMP) and demonstrate that the system's response to changes in target power is improved by 2x while operating up to 12% more efficiently for tracking accuracy.
Bryan Donyanavard, Amir-Mohammad Rahmani, Tiago Rogério Mück, Kasra Moazemmi, Nikil Dutt
DATE1
2018 Design methodologies for enabling self-awareness in autonomous systems
abstract
This paper deals with challenges and possible solutions for incorporating self-awareness principles in EDA design flows for autonomous systems. We present a holistic approach that enables self-awareness across the software/hardware stack, from systems-on-chip to systems-of-systems (autonomous car) contexts. We use the Information Processing Factory (IPF) metaphor as an exemplar to show how self-awareness can be achieved across multiple abstraction levels, and discuss new research challenges. The IPF approach represents a paradigm shift in platform design by envisioning the move towards a consequent platform-centric design in which the combination of self-organizing learning and formal reactive methods guarantee the applicability of such cyber-physical systems in safety-critical and high-availability applications.
Armin Sadighi, Bryan Donyanavard, Thawra Kadeed, Kasra Moazzemi, Tiago Rogério Mück, Ahmed Nassar 0001, Amir-Mohammad Rahmani, Thomas Wild, Nikil Dutt, Rolf Ernst, Andreas Herkersdorf, Fadi J. Kurdahi
DATE2
2018 ShaVe-ICE: Sharing Distributed Virtualized SPMs in Many-Core Embedded Systems
abstract
Traditional approaches for managing software-programmable memories (SPMs) do not support sharing of distributed on-chip memory resources and, consequently, miss the opportunity to better utilize those memory resources. Managing on-chip memory resources in many-core embedded systems with distributed SPMs requires runtime support to share memory resources between various threads with different memory demands running concurrently. Runtime SPM managers cannot rely on prior knowledge about the dynamically changing mix of threads that will execute and therefore should be designed in a way that enables SPM allocations for any unpredictable mix of threads contending for on-chip memory space. This article proposes ShaVe-ICE , an operating-system-level solution, along with hardware support, to virtualize and ultimately share SPM resources across a many-core embedded system to reduce the average memory latency. We present a number of simple allocation policies to improve performance and energy. Experimental results show that sharing SPMs could reduce the average execution time of the workload up to 19.5% and reduce the dynamic energy consumed in the memory subsystem up to 14%.
Majid Namaki-Shoushtari, Bryan Donyanavard, Luis Angel D. Bathen, Nikil Dutt
ACM Trans. Embed. Comput. Syst.2
2017 PoIiCym: rapid prototyping of resource management policies for HMPs
abstract
Heterogeneous Multiprocessors (HMPs) are becoming pervasive in current modern embedded platforms (e.g. mobile devices). These platforms often provide better power-performance tradeoffs than their homogeneous predecessors; however, novel and intelligent resource management policies are required to manage the added complexity of heterogeneous platforms and exploit their power-performance benefits. In this paper we propose PoliCym, a framework for the prototyping, validating, and deploying resource management policies for heterogeneous platforms. PoliCym provides two main benefits to resource management policy developers and to the research community: 1) a trace-based offline simulator allows policies to be quickly prototyped, debugged, and validated on top of arbitrary platform configurations; and 2) a light-weight sensing-actuation interface allows the same policies to be efficiently deployed on top of Linux-based systems without the need for implementation changes or additional development cycles. We evaluate our light-weight interface in terms of overhead and validate the PoliCym offline simulator for an ARM big.LITTLE based HMP platform running Linux.
Tiago Rogério Mück, Bryan Donyanavard, Nikil Dutt
RSP2
2016 SPMPool: Runtime SPM Management for Memory-Intensive Applications in Embedded Many-Cores
Hossein Tajik, Bryan Donyanavard, Nikil Dutt, Janmartin Jahn, Jörg Henkel
ACM Trans. Embed. Comput. Syst.2