EDBT 2026 Demo / reviewers in the wild / expert
Tanya Amert
dblp:208/0838
· DBLP profile ↗
16ranked-venue papers
8as first author
8since 2021 · last 2026
0000-0002-1943-4242ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 3 since 2021Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Critical-Section Granularity for Multi-Resource Systems with Nested Critical SectionsabstractTypical models of resource-sharing real-time tasks provide the worst-case task execution time and the duration of each critical section during which a resource is accessed. Such models abstract away the detailed behavior of a task, which may make several individual accesses within a single critical section, incurring access overhead once for the entire critical section. Considering a more fine-grained model with individual access and non-access segments at the forefront gives more control in the system design process; choosing how accesses are grouped into critical sections enables balancing trade-offs between overhead and blocking based on other system parameters, including task deadlines. This paper presents an optimal approach for multi-resource systems that allow any given task to use up to two resources, including support for nested critical sections. This is achieved by extending analysis to multiple resources and constructing a Quadratically-Constrained Integer Program to determine critical sections. This approach is compared to heuristics on the basis of schedulability and its runtime is explored. Further extension to support more resources per task is discussed. Catherine E. Nemitz, Tanya Amert, Jonad Pulaj |
ECRTS | 2 |
| 2023 | Work-in-Progress: Impacts of Critical-Section Granularity When Accessing Shared ResourcesabstractThe prevalence of computer-vision applications in autonomous vehicles necessitates the use of graphics processing units (GPUs), which must be shared among tasks due to size, weight, power, and cost constraints. Such sharing is typically managed via locking protocols. However, experiments reveal a trade-off in the granularity of such sharing; GPU operations' execution times are reduced if multiple GPU accesses are grouped in a single lock request (i.e., form a single critical section) rather than assuming only one access per critical section. This grouping exposes a broader trade-off between the extra lock overhead (and blocking) incurred by an individual task, and the analytical blocking experienced by all other tasks in the system. This trade-off is expressed herein via an extended resource model and explored using example task systems, demonstrating the impact of different access-grouping heuristics. Tanya Amert, Catherine E. Nemitz |
RTSS | 1 |
| 2021 | Timing-Predictable Vision Processing for Autonomous SystemsabstractVision processing for autonomous systems today involves implementing machine learning algorithms and vision processing libraries on embedded platforms consisting of CPUs, GPUs and FPGAs. Because many of these use closed-source proprietary components, it is very difficult to perform any timing analysis on them. Even measuring or tracing their timing behavior is challenging, although it is the first step towards reasoning about the impact of different algorithmic and implementation choices on the end-to-end timing of the vision processing pipeline. In this paper we discuss some recent progress in developing tracing, measurement and analysis infrastructure for determining the timing behavior of vision processing pipelines implemented on state-of-the-art FPGA and GPU platforms. Tanya Amert, Michael Balszun, Martin Geier 0001, F. Donelson Smith, James H. Anderson, Samarjit Chakraborty |
DATE | 1 |
| 2021 | CUPiDRT: Detecting Improper GPU Usage in Real-Time ApplicationsabstractComputer-vision applications typically rely on graphics processing units (GPUs) to accelerate computations. However, prior work has shown that care must be taken when using GPUs in real-time systems subject to strict timing constraints; without such care, GPU use can easily lead to unexpected delays not only on the GPU device but also on the host CPU. In this paper, a software library is presented that can detect the improper use of GPUs for safety-critical computer-vision applications. This library was used to analyze several GPU-using sample applications available as part of OpenCV, a popular computer-vision library, revealing the presence of issues in all ten applications considered. Additionally, a case study is presented, detailing the response-time improvements to one of the applications when such issues are corrected. Tanya Amert, James H. Anderson |
ISORC | 1 |
| 2021 | TimeWall: Enabling Time Partitioning for Real-Time Multicore+Accelerator PlatformsabstractAcross a range of safety-critical domains, an evolution is underway to endow embedded systems with "thinking" capabilities by using artificial-intelligence (AI) techniques. This evolution is being fueled by the availability of high-performance embedded hardware, typically multicore machines augmented with accelerators. Unfortunately, existing software certification processes rely on time partitioning to isolate system components, and this sense of isolation can be broken by accelerator usage. To address this issue, this paper presents TimeWall, a time-partitioning framework for multicore+accelerator platforms. When applied alongside existing methods for alleviating spatial interference, TimeWall can help enable component-wise certification on multicore+accelerator platforms. The challenges in realizing a TimeWall implementation are discussed in detail in this paper. Additionally, the temporal isolation TimeWall affords is examined experimentally, including via a case study of a computer-vision perception application, on a real platform. Tanya Amert, Zelin Tong, Sergey Voronov, Joshua Bakita, F. Donelson Smith, James H. Anderson |
RTSS | 1 |
| 2021 | AI Meets Real-Time: Addressing Real-World Complexities in Graph Response-Time AnalysisabstractArtificial-intelligence algorithms are enabling ever more sophisticated autonomous features in safety-critical application domains. These algorithms can be quite complex—consisting of many tasks interconnected in processing graphs—and often must execute on complex heterogeneous hardware—typically multicore machines augmented with one or more hardware accelerators. To further complicate matters, these processing graphs often must be supported in contexts where a large system is broken into smaller components. With this confluence of factors, existing response-time analysis for processing graphs is not applicable. In this paper, such analysis is extended to address these complexities in systems where components are isolated via time partitioning. Additionally, graph restructuring methods are presented that enable response-time bounds to be reduced. Sergey Voronov, Tanya Amert, James H. Anderson |
RTSS | 3 |
| 2021 | The price of schedulability in cyclic workloads: The history-vs.-response-time-vs.-accuracy trade-off
Tanya Amert, Ming Yang 0036, Sergey Voronov, Saujas Nandi, Thanh Vu 0001, James H. Anderson, F. Donelson Smith |
J. Syst. Archit. | 1 |
| 2021 | Concurrency groups: a new way to look at real-time multiprocessor lock nesting
Catherine E. Nemitz, Tanya Amert, Manish Goyal 0002, James H. Anderson |
Real Time Syst. | 2 |
| 2020 | The Price of Schedulability in Multi-Object Tracking: The History-vs.-Accuracy Trade-OffabstractAutonomous vehicles often employ computer-vision (CV) algorithms that track the movements of pedestrians and other vehicles to maintain safe distances from them. These algorithms are usually expressed as real-time processing graphs that have cycles due to back edges that provide history information. If immediate back history is required, then such a cycle must execute sequentially. Due to this requirement, any graph that contains a cycle with utilization exceeding 1.0 is categorically unschedulable, i.e., bounded graph response times cannot be guaranteed. Unfortunately, such cycles can occur in practice, particularly if conservative execution-time assumptions are made, as befits a safety-critical system. This dilemma can be obviated by allowing older back history, which enables parallelism in cycle execution at the expense of possibly affecting the accuracy of tracking. However, the efficacy of this solution hinges on the resulting history-vs.-accuracy trade-off that it exposes. In this paper, this trade-off is explored in depth through an experimental study conducted using the open-source CARLA autonomous-driving simulator. Somewhat surprisingly, easing away from always requiring immediate back history proved to have only a marginal impact on accuracy in this study. Tanya Amert, Ming Yang 0036, Saujas Nandi, Thanh Vu 0001, James H. Anderson, F. Donelson Smith |
ISORC | 1 |
| 2019 | OpenVX and Real-Time Certification: The Troublesome HistoryabstractMany computer-vision (CV) applications used in autonomous vehicles rely on historical results, which introduce cycles in processing graphs. However, existing response-time analysis breaks down in the presence of cycles, either by failing completely or by drastically sacrificing parallelism or CV accuracy. To address this situation, this paper presents a new graph-based task model, based on the recently ratified OpenVX standard, that includes historical requirements and their induced cycles as first-class concepts. Using this model, response-time bounds for graphs that may contain cycles are derived. These bounds expose a tradeoff between responsiveness and CV accuracy that hinges on the extent of allowed parallelism. This tradeoff is illustrated via a CV case study involving pedestrian tracking. In this case study, the methods proposed in this paper enabled significant improvements in both analytical and observed response times, with acceptable CV accuracy, compared to prior methods. Tanya Amert, Sergey Voronov, James H. Anderson |
RTSS | 1 |
| 2019 | Real-time multiprocessor locks with nesting: optimizing the common case
Catherine E. Nemitz, Tanya Amert, James H. Anderson |
Real Time Syst. | 2 |
| 2018 | Using Lock Servers to Scale Real-Time Locking Protocols: Chasing Ever-Increasing Core CountsabstractDuring the past decade, parallelism-related issues have been at the forefront of real-time systems research due to the advent of multicore technologies. In the coming years, such issues will loom ever larger due to increasing core counts. Having more cores means a greater potential exists for platform capacity loss when the available parallelism cannot be fully exploited. In this paper, such capacity loss is considered in the context of real-time locking protocols. In this context, lock nesting becomes a key concern as it can result in transitive blocking chains that force tasks to execute sequentially unnecessarily. Such chains can be quite long on a larger machine. Contention-sensitive real-time locking protocols have been proposed as a means of "breaking" transitive blocking chains, but such protocols tend to have high overhead due to more complicated lock/unlock logic. To ease such overhead, the usage of lock servers is considered herein. In particular, four specific lock-server paradigms are proposed and many nuances concerning their deployment are explored. Experiments are presented that show that, by executing cache hot, lock servers can enable reductions in lock/unlock overhead of up to 86%. Such reductions make contention-sensitive protocols a viable approach in practice. Catherine E. Nemitz, Tanya Amert, James H. Anderson |
ECRTS | 2 |
| 2018 | Avoiding Pitfalls when Using NVIDIA GPUs for Real-Time Tasks in Autonomous SystemsabstractNVIDIA's CUDA API has enabled GPUs to be used as computing accelerators across a wide range of applications. This has resulted in performance gains in many application domains, but the underlying GPU hardware and software are subject to many non-obvious pitfalls. This is particularly problematic for safety-critical systems, where worst-case behaviors must be taken into account. While such behaviors were not a key concern for earlier CUDA users, the usage of GPUs in autonomous vehicles has taken CUDA programs out of the sole domain of computer-vision and machine-learning experts and into safety-critical processing pipelines. Certification is necessary in this new domain, which is problematic because GPU software may have been developed without any regard for worst-case behaviors. Pitfalls when using CUDA in real-time autonomous systems can result from the lack of specifics in official documentation, and developers of GPU software not being aware of the implications of their design choices with regards to real-time requirements. This paper focuses on the particular challenges facing the real-time community when utilizing CUDA-enabled GPUs for autonomous applications, and best practices for applying real-time safety-critical principles. Ming Yang 0036, Nathan Otterness, Tanya Amert, Joshua Bakita, James H. Anderson, F. Donelson Smith |
ECRTS | 3 |
| 2018 | Making OpenVX Really "Real Time"abstractOpenVX is a recently ratified standard that was expressly proposed to facilitate the design of computer-vision (CV) applications used in real-time embedded systems. Despite its real-time focus, OpenVX presents several challenges when validating real-time constraints. Many of these challenges are rooted in the fact that OpenVX only implicitly defines any notion of a schedulable entity. Under OpenVX, CV applications are specified in the form of processing graphs that are inherently considered to execute monolithically end-to-end. This monolithic execution hinders parallelism and can lead to significant processing-capacity loss. Prior work partially addressed this problem by treating graph nodes as schedulable entities, but under OpenVX, these nodes represent rather coarse-grained CV functions, so the available parallelism that can be obtained in this way is quite limited. In this paper, a much more fine-grained approach for scheduling OpenVX graphs is proposed. This approach was designed to enable additional parallelism and to eliminate schedulability-related processing-capacity loss that arises when programs execute on both CPUs and graphics processing units (GPUs). Response-time analysis for this new approach is presented and its efficacy is evaluated via a case study involving an actual CV application. Ming Yang 0036, Tanya Amert, Kecheng Yang 0001, Nathan Otterness, James H. Anderson, F. Donelson Smith, Shige Wang |
RTSS | 2 |
| 2018 | Physics-Inspired Garment Recovery from a Single-View ImageabstractMost recent garment capturing techniques rely on acquiring multiple views of clothing, which may not always be readily available, especially in the case of pre-existing photographs from the web. As an alternative, we propose a method that is able to compute a 3D model of a human body and its outfit from a single photograph with little human interaction. Our algorithm is not only able to capture the global shape and overall geometry of the clothing, it can also extract the physical properties (i.e., material parameters needed for simulation) of cloth. Unlike previous methods using full 3D information (i.e., depth, multi-view images, or sampled 3D geometry), our approach achieves garment recovery from a single-view image by using physical, statistical, and geometric priors and a combination of parameter estimation, semantic parsing, shape/pose recovery, and physics-based cloth simulation. We demonstrate the effectiveness of our algorithm by re-purposing the reconstructed garments for virtual try-on and garment transfer applications and for cloth animation on digital characters. Zherong Pan, Tanya Amert, Ke Wang 0021, Licheng Yu, Tamara L. Berg, Ming C. Lin |
ACM Trans. Graph. | 3 |
| 2017 | GPU Scheduling on the NVIDIA TX2: Hidden Details RevealedabstractThe push towards fielding autonomous-driving capabilities in vehicles is happening at breakneck speed. Semi-autonomous features are becoming increasingly common, and fully autonomous vehicles are optimistically forecast to be widely available in just a few years. Today, graphics processing units (GPUs) are seen as a key technology in this push towards greater autonomy. However, realizing full autonomy in mass-production vehicles will necessitate the use of stringent certification processes. Currently available GPUs pose challenges in this regard, as they tend to be closed-source “black boxes” that have features that are not publicly disclosed. For certification to be tenable, such features must be documented. This paper reports on such a documentation effort. This effort was directed at the NVIDIA TX2, which is one of the most prominent GPU-enabled platforms marketed today for autonomous systems. In this paper, important aspects of the TX2's GPU scheduler are revealed as discerned through experimental testing and validation. Tanya Amert, Nathan Otterness, Ming Yang 0036, James H. Anderson, F. Donelson Smith |
RTSS | 1 |