EDBT 2026 Demo / reviewers in the wild / expert
Omais Shafi
dblp:276/9907
· DBLP profile ↗
5ranked-venue papers
5as first author
4since 2021 · last 2024
0000-0002-0054-5161ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Repercussions of Using DNN Compilers on Edge GPUs for Real Time and Safety Critical Systems: A Quantitative AuditabstractRapid advancements in edge devices have led to a large deployment of deep neural network (DNN) based workloads. To utilize the resources at the edge effectively, many DNN compilers are proposed that efficiently map the high level DNN models developed in frameworks like PyTorch, Tensorflow, Caffe, and so on into minimum deployable lightweight execution engines. For real time applications like ADAS, these compiler optimized engines should give precise, reproducible, and predictable inferences, both in-terms of runtime and output consistency. This article is the first effort in empirically auditing state-of-the-art DNN compilers viz TensorRT, AutoTVM, and AutoScheduler. We characterize the NN compilers based on their performance predictability w.r.t inference latency, output reproducibility, hardware utilization, and so on and based on that provide various recommendations. Our methodology and findings can potentially help the application developers, in making informed decision about the choice of DNN compiler, in a real time safety critical setting. Omais Shafi, Mohammad Khalid Pandit, Amarjeet Saini, Gayathri Ananthanarayanan, Rijurekha Sen |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2022 | DynCNN: Application Dynamism and Ambient Temperature Aware Neural Network Scheduler in Edge Devices for Traffic ControlabstractRoad traffic congestion increases vehicular emissions and air pollution. Traffic rule violation causes road accidents. Both pollution and accidents take tremendous social and economic toll worldwide, and more so in developing countries where the skewed vehicle to road infrastructure ratio amplifies the problems. Automating traffic intersection management to detect and penalize traffic rule violations and reduce traffic congestion, is the focus of this paper, using state-of-the-art Convolutional Neural Network (CNN) on traffic camera feeds. There are however non-trivial challenges in handling the chaotic, non-laned traffic scenes in developing countries. Maintaining high throughput is one of the challenges, as broadband connectivity to remote GPU servers is absent in developing countries, and embedded GPU platforms on roads need to be low cost due to budget constraints. Additionally, ambient temperatures in developing country cities can go to 45-50 degree Celsius in summer, where continuous embedded processing can lead to lower lifetimes of the embedded platforms. In this paper, we present DynCNN, an application dynamism and ambient temperature aware controller for Neural Network concurrency. DynCNN effectively uses processor heterogeneity to control the number of threads and frequencies on the accelerator to manage application utility under strict thermal and power thresholds. We evaluate the efficiency of DynCNN on three different commercially available embedded GPUs (Jetson TX2TM, Xavier NXTM and Xavier AGXTM) using a real traffic intersection’s 40 days’ dataset. Experimental results show that in comparison to all existing state-of-the art- GPU governors for two different CPU settings, DynCNN reduces the average temperature and power by ~12°C and 68.82% respectively for one CPU setting (Baseline1) and similarly, it improves the performance by around 31.2% compared to the other CPU setting (Baseline2). Omais Shafi, Sachin Chauhan, Gayathri Ananthanarayanan, Rijurekha Sen |
COMPASS | 1 |
| 2021 | CuckoOnsai: An Efficient Memory Authentication Using Amalgam of Cuckoo Filters and Integrity TreesabstractThe off-chip main memory data can be extracted or tampered by an adversary having physical access to a device. In modern secure designs, such tampering or data attacks can be prevented by storing the integrity tree built on top of encryption counters. However, such approaches have significant performance and on-chip storage overheads. This paper proposes CuckoOnsai, which uses the combination of a novel on-chip per core Cuckoo filter and an off-chip small integrity tree. Compared to the recent competing scheme, CuckoOnsai improves performance by 13.1% and reduces the space overheads and $\mathrm{ED}^{2}$ by 50% and 26.1%, respectively. Omais Shafi, Ismi Abidi |
DAC | 1 |
| 2021 | FreqCounter: Efficient Cacheability of Encryption and Integrity Tree Counters in Secure Processors
Omais Shafi, Janibul Bashir |
J. Syst. Archit. | 1 |
| 2020 | SecSched: Flexible Scheduling in Secure ProcessorsabstractTrusted execution environments (TEEs) are an integral part of modern processors because security has become a very important concern. However, many such environments are bedeviled by the high cost of context switches, particularly when there is a switch from secure mode to non-secure mode owing primarily to cache pollution and TLB-flushing overheads. State-of-the-art implementations create a secure shared memory channel between a thread running in secure mode and a thread running in non-secure mode, which invokes system calls on its behalf. We argue that this is inefficient, and it is possible to reduce the overheads significantly by efficiently storing the context of secure threads and intelligent scheduling. In this paper, we propose a new scheduling algorithm SecSched that uses Cuckoo filters to capture the context of a thread. We schedule threads with similar contexts on the same core to leverage the effects of the locality. Our algorithm requires minimal hardware enhancements that are limited to maintaining a Cuckoo filter per core and a thread with the addition of few performance counters per thread to keep track of the miss counts. We show that with these minimal changes we can increase the performance of a suite of OS-intensive workloads by 27.6% with a minimal area overhead (around 0.04%). Omais Shafi, Janibul Bashir |
PACT | 1 |