EDBT 2026 Demo / reviewers in the wild / expert
Aditya Ukarande
dblp:273/4363
· DBLP profile ↗
4ranked-venue papers
2as first author
3since 2021 · last 2025
0009-0009-2401-4902ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RIMMS: Runtime Integrated Memory Management System for Heterogeneous ComputingabstractEfficient memory management in heterogeneous systems is increasingly challenging due to diverse compute architectures (e.g., CPU, GPU, and FPGA) and dynamic task mappings not known at compile time. Existing approaches often require programmers to manage data placement and transfers explicitly, or assume static mappings that limit portability and scalability. This article introduces RIMMS (Runtime Integrated Memory Management System), a lightweight, runtime-managed, hardware-agnostic memory abstraction layer that decouples application development from low-level memory operations. RIMMS transparently tracks data locations, manages consistency, and supports efficient memory allocation across heterogeneous compute elements without requiring platform-specific tuning or code modifications. We integrate RIMMS into a baseline runtime and evaluate with complete radar signal processing applications across CPU+GPU and CPU+FPGA platforms. RIMMS delivers up to 2.43× speedup on GPU-based and 1.82× on FPGA-based systems over the baseline. Compared to IRIS, a recent heterogeneous runtime system, RIMMS achieves up to 3.08X speedup and matches the performance of native CUDA implementations while significantly reducing programming complexity. Despite operating at a higher abstraction level, RIMMS incurs only 1–2 cycles of overhead per memory management call, making it a low-cost solution. These results demonstrate RIMMS’s ability to deliver high performance and enhanced programmer productivity in dynamic, real-world heterogeneous environments. Serhan Gener, Aditya Ukarande, Shilpa Mysore Srinivasa Murthy, Md Sahil Hassan, Joshua Mack, Chaitali Chakrabarti, Ümit Y. Ogras, Ali Akoglu |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2024 | PACT: Accurate Power Analysis and Carbon Emission Tracking for SustainabilityabstractThe energy consumption of artificial intelligence (AI) workloads has exceeded 1 GWh/day, highlighting the urgent need to analyze their energy use and carbon emissions beyond just focusing on accuracy and performance. Current energy and carbon tracking tools focus on GPU and CPU, approximating other system components. This paper proposes PACT, an accurate system-wide power analysis and carbon emission tracking methodology that accounts for all system components. It first measures the total power consumed by the hardware resources while running a specific workload. Then, it uses performance counters that capture the dynamic workload characteristics to model the power consumption and then carbon emission. Quantitative evaluations show PACT's superior accuracy, with an average error of less than 1% in estimating energy consumption and carbon emissions, compared to existing tools, with average errors ranging from 20.2% to 37.9% for energy consumption and from 19.6% to 38.3% for carbon emissions. Aditya Ukarande, Toygun Basaklar, Mingcong Cao, Ümit Y. Ogras |
ISLPED | 1 |
| 2022 | Locality-Aware CTA Scheduling for Gaming ApplicationsabstractThe compute work rasterizer or the GigaThread Engine of a modern NVIDIA GPU focuses on maximizing compute work occupancy across all streaming multiprocessors in a GPU while retaining design simplicity. In this article, we identify the operational aspects of the GigaThread Engine that help it meet those goals but also lead to less-than-ideal cache locality for texture accesses in 2D compute shaders, which are an important optimization target for gaming applications. We develop three software techniques, namely LargeCTAs , Swizzle , and Agents , to show that it is possible to effectively exploit the texture data working set overlap intrinsic to 2D compute shaders. We evaluate these techniques on gaming applications across two generations of NVIDIA GPUs, RTX 2080 and RTX 3080, and find that they are effective on both GPUs. We find that the bandwidth savings from all our software techniques on RTX 2080 is much higher than the bandwidth savings on baseline execution from inter-generational cache capacity increase going from RTX 2080 to RTX 3080. Our best-performing technique, Agents , records up to a 4.7% average full-frame speedup by reducing bandwidth demand of targeted shaders at the L1-L2 and L2-DRAM interfaces by 23% and 32%, respectively, on the latest generation RTX 3080. These results acutely highlight the sensitivity of cache locality to compute work rasterization order and the importance of locality-aware cooperative thread array scheduling for gaming applications. Aditya Ukarande, Suryakant Patidar, Ram Rangan |
ACM Trans. Archit. Code Optim. | 1 |
| 2020 | Zeroploit: Exploiting Zero Valued Operands in Interactive Gaming ApplicationsabstractIn this article, we first characterize register operand value locality in shader programs of modern gaming applications and observe that there is a high likelihood of one of the register operands of several multiply, logical-and, and similar operations being zero, dynamically. We provide intuition, examples, and a quantitative characterization for how zeros originate dynamically in these programs. Next, we show that this dynamic behavior can be gainfully exploited with a profile-guided code optimization called Zeroploit that transforms targeted code regions into a zero-(value-)specialized fast path and a default slow path. The fast path benefits from zero-specialization in two ways, namely: (a) the backward slice of the other operand of a given multiply or logical-and can be skipped dynamically, provided the only use of that other operand is in the given instruction, and (b) the forward slice of instructions originating at the given instruction can be zero-specialized, potentially triggering further backward slice specializations from operations of that forward slice as well. Such specialization helps the fast path avoid redundant dynamic computations as well as memory fetches, while the fast-slow versioning transform helps preserve functional correctness. With an offline value profiler and manually optimized shader programs, we demonstrate that Zeroploit is able to achieve an average speedup of 35.8% for targeted shader programs, amounting to an average frame-rate speedup of 2.8% across a collection of modern gaming applications on an NVIDIA® GeForce RTX™ 2080 GPU. Ram Rangan, Mark Stephenson, Aditya Ukarande, Shyam Murthy, Virat Agarwal, Marc Blackstein |
ACM Trans. Archit. Code Optim. | 3 |