EDBT 2026 Demo / reviewers in the wild / expert
Joshua Bakita
dblp:221/5097
· DBLP profile ↗
16ranked-venue papers
7as first author
13since 2021 · last 2026
0009-0007-2856-5774ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GDDRHammer: Greatly Disturbing DRAM Rows - Cross-Component Rowhammer Attacks From Modern GPUs
Yichang Hu, Noah Brown, Joshua Bakita, Tianlong Chen 0001, Daniel Genkin, Andrew Kwong |
SP | 4 |
| 2025 | Hardware Compute Partitioning on NVIDIA GPUs for Composable Systems
Joshua Bakita, James H. Anderson |
ECRTS | 1 |
| 2025 | Concurrent FFT Execution on GPUs in Real-TimeabstractFourier transforms are vital for a broad range of signal-processing applications. Accelerating FFTs with GPUs offers an orders-of-magnitude improvement vs. CPU-only FFT computation. However, two problems arise when executing FFT tasks with other GPU work. First, concurrent GPU use introduces unpredictability in the form of lengthy response times. Second, it is unclear how to best parameterize and schedule FFT tasks to meet the throughput and timeliness constraints of real-time signal processing. This work investigates how FFT and other GPU-using tasks can concurrently access a GPU while maintaining bounded response-time guarantees without sacrificing throughput. In our experiments, the techniques proposed by this work result in an up to 17% improvement in worst-case FFT response times. Syed W. Ali, Joseph Goh, Joshua Bakita, Samarjit Chakraborty, James H. Anderson |
PDP | 3 |
| 2025 | Work in Progress: Increasing Schedulability via on-GPU SchedulingabstractGPUs are increasingly needed to run a variety of tasks in embedded systems, from object recognition to conver-sational chat. Some of these tasks are safety-critical, real-time tasks, where completing each by its deadline is essential for system safety. To meet the practical constraints of real-world systems, these tasks much also be run efficiently. Unfortunately, current techniques to schedule GPU-using tasks onto a single GPU while respecting deadlines impart high overheads, leading to inefficiency and substantial capacity loss during formal analysis. We address this problem by moving GPU scheduling from the CPU to the GPU. Our approach limits overheads, increasing the proportion of CPU tasks which can meet their deadlines by as much as 12.1% while increasing available GPU capacity. Joshua Bakita, James H. Anderson |
RTAS | 1 |
| 2025 | Work-in-Progress: Enabling Transparent Priority Scheduling for NVIDIA GPUsabstractModern embedded systems increasingly rely on GPUs for safety-critical tasks such as object detection in autonomous vehicles and data visualization in medical diagnostics. These applications require running high-priority workloads alongside lower-priority tasks on shared GPU hardware due to cost, power, or space constraints. However, NVIDIA GPUs employ round-robin scheduling that treats all tasks equally, causing unpredictable delays for critical workloads. Existing GPU priority scheduling solutions require application modifications, introduce high overhead, or depend on specialized embedded hardware. We present a GPU-level static priority scheduler built on the nvsched framework that requires no application changes and runs on commodity GPUs. Our scheduler maintains tasks in a priority-ordered queue and uses a parameter communication system to enable dynamic priority updates from userspace. Experiments with MobileNetV3 inference under GPU contention demonstrate 32-41% lower maximum latency for high-priority tasks when using our scheduler compared to NVIDIA's default scheduler. Our approach provides practical priority control for GPU workloads while maintaining low scheduling overhead. Noah Weaver, Joshua Bakita |
RTSS | 2 |
| 2025 | The advantage of the GPU as a real-time AI accelerator
Joshua Bakita, James H. Anderson |
Real Time Syst. | 1 |
| 2024 | Demystifying NVIDIA GPU Internals to Enable Reliable GPU ManagementabstractAs GPU-dependent artificial intelligence and ma-chine learning workloads increasingly come to embedded, safety-critical systems-such as self-driving cars-real-time predictabil-ity for GPU-using tasks becomes essential. This paper identifies flaws in three different real-time GPU management approaches that are largely the result of incomplete information about NVIDIA GPU internals. Details concerning this missing information are elucidated via experiments. Based on this information, key rules of GPU scheduling are identified and shown necessary for safe GPU management. Joshua Bakita, James H. Anderson |
RTAS | 1 |
| 2023 | Hardware Compute Partitioning on NVIDIA GPUsabstractEmbedded and autonomous systems are increasingly integrating AI/ML features, often enabled by a hardware accelerator such as a GPU. As these workloads become increasingly demanding, but size, weight, power, and cost constraints remain unyielding, ways to increase GPU capacity are an urgent need. In this work, we provide a means by which to spatially partition the computing units of NVIDIA GPUs transparently, allowing of tidled capacity to be reclaimed via safe and efficient GPU sharing. Our approach works on any NVIDIA GPU since 2013, and can be applied via our easy-to-use, user-space library titled libsmctrl. We back the design of our system with deep investigations into the hardware scheduling pipeline of NVIDIA GPUs. We provide guidelines for the use of our system, and demonstrate it via an object detection case study using YOLOv2. Joshua Bakita, James H. Anderson |
RTAS | 1 |
| 2022 | Minimizing DAG Utilization by Exploiting SMTabstractParallel workloads are commonly modeled as directed acyclic graphs (DAGs). While DAG scheduling is an important tool, it is plagued by capacity loss; it is not uncommon to see half of a platform go unused. Here this loss is attacked from a new direction: reducing per-DAG utilization prior to assigning computing cores to a DAG. Specifically, simultaneous multithreading (SMT) is used to schedule individual nodes of a DAG task in parallel on the same physical computing core. An optimization program is given that applies SMT to a DAG in a way that minimizes total utilization without compromising correctness. Results for both individual DAGs and systems of DAGs are evaluated using both a large-scale study of synthetic DAGs and a case study. Optimal use of the program can reduce DAG utilization and required core counts by over 40% in the best cases and by 25% in nearly half of cases. Runtime requirements for the optimization program are considered, and a tunable parameter is provided to make tradeoffs between runtime and optimality, allowing even DAGs with 500 nodes to benefit. Sims Osborne, Joshua Bakita, Tyler Yandrofski, James H. Anderson |
RTAS | 2 |
| 2022 | Enabling GPU Memory Oversubscription via Transparent Paging to an NVMe SSDabstractSafety-critical embedded systems are experiencing increasing computational and memory demands as edge-computing and autonomous systems gain adoption. Main memory (DRAM) is often scarce, and existing mechanisms to support DRAM oversubscription, such as demand paging or compile-time transformations, either imply serious CPU capacity loss, or put unacceptable constraints on program structure. This work proposes an alternative: paging GPU rather than CPU memory buffers directly to permanent storage to enable efficient and predictable memory oversubscription. This paper focuses on why GPU paging is useful and how it can be efficiently implemented. Specifically, a GPU paging implementation is proposed as an extension to NVIDIA's embedded Linux GPU drivers. In experiments reported herein, this implementation was seen to be three times faster end-to-end than demand paging, with 81% lower overheads. It also achieved speeds above the fastest prexisting Linux userspace I/O APIs with low DRAM and bus interference to CPU tasks—at most a 17% slowdown. Joshua Bakita, James H. Anderson |
RTSS | 1 |
| 2021 | Simultaneous Multithreading in Mixed-Criticality Real-Time SystemsabstractSimultaneous multithreading (SMT) enables enhanced computing capacity by allowing multiple tasks to execute concurrently on the same computing core. Despite its benefits, its use has been largely eschewed in work on real-time systems due to concerns that tasks running on the same core may adversely interfere with each other. In this paper, the safety of using SMT in a mixed-criticality multicore context is considered in detail. To this end, a prior open-source framework called MC2(mixedcriticality on multicore), which provides features for mitigating cache and memory interference, was re-implemented to support SMT on an SMT-capable multicore platform. The creation of this new, configurable MC2variant entailed producing the first operating-system implementations of several recently proposed real-time SMT schedulers and tying them together within a mixed-criticality context. These schedulers introduce new spatialisolation challenges, which required introducing isolation at both the L2 and L3 cache levels. The efficacy of the resulting MC2variant is demonstrated via three experimental efforts. The first involved obtaining execution data using a wide range of benchmark suites, including TACLeBench, DIS, SD-VBS, and synthetic microbenchmarks. The second involved conducting a large-scale overhead-aware schedulability study, parameterized by the collected benchmark data, to elucidate schedulability tradeoffs. The third involved experiments involving case-study task systems. In the schedulability study, the use of SMT proved capable of increasing platform capacity by an average factor of 1.32. In the case-study experiments, deadline misses of highly critical tasks were never observed. Joshua Bakita, Shareef Ahmed, Sims Osborne, F. Donelson Smith, James H. Anderson |
RTAS | 1 |
| 2021 | TimeWall: Enabling Time Partitioning for Real-Time Multicore+Accelerator PlatformsabstractAcross a range of safety-critical domains, an evolution is underway to endow embedded systems with "thinking" capabilities by using artificial-intelligence (AI) techniques. This evolution is being fueled by the availability of high-performance embedded hardware, typically multicore machines augmented with accelerators. Unfortunately, existing software certification processes rely on time partitioning to isolate system components, and this sense of isolation can be broken by accelerator usage. To address this issue, this paper presents TimeWall, a time-partitioning framework for multicore+accelerator platforms. When applied alongside existing methods for alleviating spatial interference, TimeWall can help enable component-wise certification on multicore+accelerator platforms. The challenges in realizing a TimeWall implementation are discussed in detail in this paper. Additionally, the temporal isolation TimeWall affords is examined experimentally, including via a case study of a computer-vision perception application, on a real platform. Tanya Amert, Zelin Tong, Sergey Voronov, Joshua Bakita, F. Donelson Smith, James H. Anderson |
RTSS | 4 |
| 2021 | Statically optimal dynamic soft real-time semi-partitioned scheduling
Clara Hobbs, Zelin Tong, Joshua Bakita, James H. Anderson |
Real Time Syst. | 3 |
| 2019 | Simultaneous Multithreading Applied to Real TimeabstractExisting models used in real-time scheduling are inadequate to take advantage of simultaneous multithreading (SMT), which has been shown to improve performance in many areas of computing, but has seen little application to real-time systems. The SMART task model, which allows for combining SMT and real time by accounting for the variable task execution costs caused by SMT, is introduced, along with methods and conditions for scheduling SMT tasks under global earliest-deadline-first scheduling. The benefits of using SMT are demonstrated through a large-scale schedulability study in which we show that task systems with utilizations 30% larger than what would be schedulable without SMT can be correctly scheduled. Sims Osborne, Joshua Bakita, James H. Anderson |
ECRTS | 2 |
| 2019 | Re-Thinking CNN Frameworks for Time-Sensitive Autonomous-Driving Applications: Addressing an Industrial ChallengeabstractVision-based perception systems are crucial for profitable autonomous-driving vehicle products. High accuracy in such perception systems is being enabled by rapidly evolving convolution neural networks (CNNs). To achieve a better understanding of its surrounding environment, a vehicle must be provided with full coverage via multiple cameras. However, when processing multiple video streams, existing CNN frameworks often fail to provide enough inference performance, particularly on embedded hardware constrained by size, weight, and power limits. This paper presents the results of an industrial case study that was conducted to re-think the design of CNN software to better utilize available hardware resources. In this study, techniques such as parallelism, pipelining, and the merging of per-camera images into a single composite image were considered in the context of a Drive PX2 embedded hardware platform. The study identifies a combination of techniques that can be applied to increase throughput (number of simultaneous camera streams) without significantly increasing per-frame latency (camera to CNN output) or reducing per-stream accuracy. Ming Yang 0036, Shige Wang, Joshua Bakita, Thanh Vu 0001, F. Donelson Smith, James H. Anderson, Jan-Michael Frahm |
RTAS | 3 |
| 2018 | Avoiding Pitfalls when Using NVIDIA GPUs for Real-Time Tasks in Autonomous SystemsabstractNVIDIA's CUDA API has enabled GPUs to be used as computing accelerators across a wide range of applications. This has resulted in performance gains in many application domains, but the underlying GPU hardware and software are subject to many non-obvious pitfalls. This is particularly problematic for safety-critical systems, where worst-case behaviors must be taken into account. While such behaviors were not a key concern for earlier CUDA users, the usage of GPUs in autonomous vehicles has taken CUDA programs out of the sole domain of computer-vision and machine-learning experts and into safety-critical processing pipelines. Certification is necessary in this new domain, which is problematic because GPU software may have been developed without any regard for worst-case behaviors. Pitfalls when using CUDA in real-time autonomous systems can result from the lack of specifics in official documentation, and developers of GPU software not being aware of the implications of their design choices with regards to real-time requirements. This paper focuses on the particular challenges facing the real-time community when utilizing CUDA-enabled GPUs for autonomous applications, and best practices for applying real-time safety-critical principles. Ming Yang 0036, Nathan Otterness, Tanya Amert, Joshua Bakita, James H. Anderson, F. Donelson Smith |
ECRTS | 4 |