Ila Gokarn

dblp:286/2394 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
10since 2021 · last 2025
0009-0001-6977-945XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 RA-MOSAIC: Resource Adaptive Edge AI Optimization over Spatially Multiplexed Video Streams
abstract
Sustaining real-time, high-fidelity AI-based vision perception on edge devices is challenging due to both the high computational overhead of increasingly “deeper” Deep Neural Networks (DNNs) and the increasing resolution/quality of camera sensors. Such high-throughput vision perception is even more challenging in multi-tenancy systems, where video streams from multiple such high-quality cameras need to share the same GPU resource on a single edge device. Criticality-aware canvas-based processing is a promising paradigm that decomposes multiple concurrent video streams into Regions of Interest (RoI) and spatially channels the limited computational resources to selected RoI with higher “resolution,” thereby moderating the tradeoff between computational load, task fidelity, and processing throughput. RA-MOSAIC (Resource Adaptive MOSAIC) employs such canvas-based processing, while further tuning the incoming video streams and available resources on-demand to allow the system to adapt to dynamic changes in workload (often arising from variations in the number or size of relevant objects observed by individual cameras). RA-MOSAIC utilizes two distinct and synergistic concepts. First, at the camera sensor, a bandwidth-adaptive and lightweight Bandwidth-Adaptive Camera Transmission (BACT) method applies differential downsampling to create mixed-resolution individual frames that preferentially preserve resolution for critical RoIs, before being transmitted to the edge node. Second, at the edge, BACT video streams received from multiple cameras are decomposed into multi-scale RoI tiles and spatially packed using a novel workload-adaptive bin-packing strategy into a single “canvas frame.” Notably, the canvas frame itself is dynamically sized such that the edge device can opportunistically provide higher processing throughput for selected high-priority tiles during periods of lower aggregate workloads. To demonstrate RA-MOSAIC’s gains in processing throughput and perception fidelity, we evaluate RA-MOSAIC on a single NVIDIA Jetson TX2 edge device for two benchmark tasks—Drone-based Pedestrian Detection and Automatic License Plate Recognition. In a bandwidth-constrained wireless environment, RA-MOSAIC employs a batch size of 1 to pack up to 6 concurrent video streams on a dynamically sized canvas frame to provide (i) 14.3% gain in object detection accuracy and (ii) 11.11% gain in throughput on average (up to 20 FPS per camera, cumulatively 120 FPS), over our previous work MOSAIC, a naive canvas-based baseline. Compared to prior state-of-the-art baselines such as batched inference over extracted RoI, RA-MOSAIC provides a very significant, 29.6% gain in accuracy for a comparable throughput. Similarly, RA-MOSAIC dramatically outperforms bandwidth-adaptive baselines, such as First Come First Serve (FCFS) ( \(\leq 1\%\) accuracy gain but 5.6× or 566.67% throughput gain) and uniform grid packing (17% accuracy improvement and 5% throughput gain).
Ila Gokarn, Yigong Hu, Tarek F. Abdelzaher, Archan Misra
ACM Trans. Multim. Comput. Commun. Appl.1
2024 JIGSAW: Edge-based Streaming Perception over Spatially Overlapped Multi-Camera Deployments
abstract
We present JIGSAW, a novel system that performs edge-based streaming perception over multiple video streams, while additionally factoring in the redundancy offered by the spatial overlap often exhibited in urban, multi-camera deployments. To assure high streaming throughput, JIGSAW extracts and spatially multiplexes multiple regions-of-interest from different camera frames into a smaller canvas frame. Moreover, to ensure that perception stays abreast of evolving object kinematics, JIGSAW includes a utility-based weighted scheduler to preferentially prioritize and even skip object-specific tiles extracted from an incoming stream of camera frames. Using the CityflowV2 traffic surveillance dataset, we show that JIGSAW can simultaneously process 25 cameras on a single Jetson TX2 with a 66.6% increase in accuracy and a simultaneous 18x (1800%) gain in cumulative throughput (475 FPS), far outperforming competitive baselines.
Ila Gokarn, Yigong Hu, Tarek F. Abdelzaher, Archan Misra
ICME1
2024 Criticality Aware Canvas-based Visual Perception at the Edge
abstract
Efficient and effective machine perception remains a formidable challenge in sustaining high fidelity and high throughput of perception tasks on affordable edge devices. This is especially due to the continuing increase in resolution of sensor streams (e.g., video input streams generated by 4K/8K cameras and neuromorphic event cameras that produce ≥ 10 MEvents/second) and computational complexity of Deep Neural Network (DNN) models, which overwhelms edge platforms, adversely impacting machine perception efficiency. Given the insufficiency of the available computation resources, a question then arises on whether selected regions/components of the perception task can be prioritized (and executed preferentially) to achieve highest task fidelity while adhering to the resource budget. This extended abstract explores the paradigm of Canvas-based Processing and criticality-awareness in the context of multi-sensor machine perception pipelines on resource-constrained platforms, in guiding perception pipelines and systems on "what" to pay attention to in the sensing field and "when", to maximize overall perception fidelity under computational constraints and moderate the processing throughput-vs-accuracy trade-off.
Ila Gokarn
MobiSys1
2024 Poster: Profiling Event Vision Processing on Edge Devices
abstract
As RGB camera resolutions and frame-rates improve, their increased energy requirements make it challenging to deploy fast, efficient, and low-power applications on edge devices. Newer classes of sensors, such as the biologically inspired neuromorphic event-based camera, capture only changes in light intensity per-pixel to achieve operational superiority in sensing latency (O(μs)), energy consumption (O(mW)), high dynamic range (140dB), and task accuracy such as in object tracking, over traditional RGB camera streams. However, highly dynamic scenes can yield an event rate of up to 12MEvents/second, the processing of which could overwhelm resource-constrained edge devices. Efficient processing of high volumes of event data is crucial for ultra-fast machine vision on edge devices. In this poster, we present a profiler that processes simulated event streams from RGB videos into 6 variants of framed representations for DNN inference on an NVIDIA Jetson Orin AGX, a representative edge device. The profiler evaluates the trade-offs between the volume of events evaluated, the quality of the processed event representation, and processing time to present the design choices available to an edge-scale event camera-based application observing the same RGB scenes. We believe that this analysis opens up the exploration of novel system designs for real-time low-power event vision on edge devices.
Ila Gokarn, Archan Misra
MobiSys1
2024 EyeGraph: Modularity-aware Spatio Temporal Graph Clustering for Continuous Event-based Eye Tracking
abstract
Continuous tracking of eye movement dynamics plays a significant role in developing a broad spectrum of human-centered applications, such as cognitive skills (visual attention and working memory) modeling, human-machine interaction, biometric user authentication, and foveated rendering. Recently neuromorphic cameras have garnered significant interest in the eye-tracking research community, owing to their sub-microsecond latency in capturing intensity changes resulting from eye movements. Nevertheless, the existing approaches for event-based eye tracking suffer from several limitations: dependence on RGB frames, label sparsity, and training on datasets collected in controlled lab environments that do not adequately reflect real-world scenarios. To address these limitations, in this paper, we propose a dynamic graph-based approach that uses a neuromorphic event stream captured by Dynamic Vision Sensors (DVS) for high-fidelity tracking of pupillary movement. More specifically, first, we present EyeGraph, a large-scale multi-modal near-eye tracking dataset collected using a wearable event camera attached to a head-mounted device from 40 participants -- the dataset was curated while mimicking in-the-wild settings, accounting for varying mobility and ambient lighting conditions. Subsequently, to address the issue of label sparsity, we adopt an unsupervised topology-aware approach as a benchmark. To be specific, (a) we first construct a dynamic graph using Gaussian Mixture Models (GMM), resulting in a uniform and detailed representation of eye morphology features, facilitating accurate modeling of pupil and iris. Then (b) apply a novel topologically guided modularity-aware graph clustering approach to precisely track the movement of the pupil and address the label sparsity in event-based eye tracking. We show that our unsupervised approach has comparable performance against the supervised approaches while consistently outperforming the conventional clustering approaches.
Nuwan Sriyantha Bandara, Thivya Kandappu, Argha Sen, Ila Gokarn, Archan Misra
NeurIPS4
2024 Algorithms for Canvas-Based Attention Scheduling with Resizing
abstract
Canvas-based attention scheduling was recently pro-posed to improve the efficiency of real-time machine perception systems. This framework introduces a notion of focus locales, referring to those areas where the attention of the inference system should “allocate its attention”. Data from these locales (e.g., parts of the input video frames containing objects of interest) are packed together into a smaller canvas frame which is processed by the downstream machine learning algorithm. Compared with processing the entire input data frame, this practice saves resources while maintaining inference quality. Previous work was limited to a simplified solution where the focus locales are quantized to a small set of allowed sizes for the ease of packing into the canvas in a best-effort manner. In this paper, we remove this limiting constraint thus obviating quantization, and derive the first spatiotemporal schedulability bound for objects of arbitrary sizes in a canvas-based attention scheduling framework. We further allow object resizing and design a set of scheduling algorithms to adapt to varying workloads dynamically. Experiments on a representative AI-powered embedded platform with a real-world video dataset demonstrate the improvements in performance and inform the design and capacity planning of modern real-time machine perception pipelines.
Yigong Hu, Ila Gokarn, Shengzhong Liu, Archan Misra, Tarek F. Abdelzaher
RTAS2
2023 Underprovisioned GPUs: On Sufficient Capacity for Real-Time Mission-Critical Perception
abstract
Recent work suggests that computing resources, such as GPUs in real-time edge-based perception systems, need not have sufficient capacity to keep up with the input frame rates of all input devices (e.g., cameras) at their full-frame resolution. Rather, they can be under-provisioned because only parts of any given frame need to be inspected (i.e., paid attention to). This paper derives an attention allocation policy, called canvas-based attention scheduling that decides which parts of each frame of each device to inspect, and a corresponding schedulability condition that relates the spatiotemporal properties of surrounding objects to the ability of the edge-based perception subsystem to keep up with the state of the environment in real-time. It provides a quantitative estimate of adequate computing capacity for the expected perception workload. We implement a canvas-based attention scheduler for an object detection application and perform an empirical comparative study based on actual GPU hardware and surveillance videos. Results show that canvas-based attention scheduling keeps up with the environment while using a much smaller GPU capacity, compared with prior approaches.
Yigong Hu, Ila Gokarn, Shengzhong Liu, Archan Misra, Tarek F. Abdelzaher
ICCCN2
2023 MOSAIC: Spatially-Multiplexed Edge AI Optimization over Multiple Concurrent Video Sensing Streams
abstract
Sustaining high fidelity and high throughput of perception tasks over vision sensor streams on edge devices remains a formidable challenge, especially given the continuing increase in image sizes (e.g., generated by 4K cameras) and complexity of DNN models. One promising approach involves criticality-aware processing, where the computation is directed selectively to "critical" portions of individual image frames. We introduce MOSAIC, a novel system for such criticality-aware concurrent processing of multiple vision sensing streams that provides a multiplicative increase in the achievable throughput with negligible loss in perception fidelity. MOSAIC determines critical regions from images received from multiple vision sensors and spatially bin-packs these regions using a novel multi-scale Mosaic Across Scales (MoS) tiling strategy into a single `canvas frame', sized such that the edge device can retain sufficiently high processing throughput. Experimental studies using benchmark datasets for two tasks, Automatic License Plate Recognition and Drone-based Pedestrian Detection, shows that MOSAIC, executing on a Jetson TX2 edge device, can provide dramatic gains in the throughput vs. fidelity tradeoff. For instance, for drone-based pedestrian detection, for a batch size of 4, MOSAIC can pack input frames from 6 cameras to achieve (a) 4.75X (475%) higher throughput (23 FPS per camera, cumulatively 138FPS) with ≤ 1% accuracy loss, compared to a First Come First Serve (FCFS) processing paradigm.
Ila Gokarn, Hemanth Reddy Sabbella, Yigong Hu, Tarek F. Abdelzaher, Archan Misra
MMSys1
2023 Work-in-Progress: Algorithms for Canvas-Based Attention Scheduling with Resizing
abstract
In real-time machine inference literature, canvas-based attention scheduling was recently introduced as an effective scheduling algorithm for real-time perception pipelines. In this framework, a notion of focus locales is maintained, referring to those locales on which the perception subsystem “focuses its attention”. Data from these locales (e.g., parts of input video frames corresponding to objects of interest) are packed into smaller bins called canvas frames that are then processed by the AI pipeline. The practice saves resources compared to processing the entirety of the original full frames. While prior work on canvas-based scheduling derived a schedulability bound, their bound applies only if focus locales are quantized into a small set of allowable container sizes for ease of packing into the canvas. In this work, we explore the possibility of removing this limiting assumption thus obviating quantization for a new bound, and generalizing the scheduling policy to allow for object resizing. Experiments on a representative AI-powered embedded platform with a real-world video dataset demonstrate improvements in efficiency in the presence and empirically validate the new bound. The result informs the design and capacity planning of modern real-time machine perception pipelines.
Yigong Hu, Ila Gokarn, Shengzhong Liu, Archan Misra, Tarek F. Abdelzaher
RTSS2
2021 VibranSee: Enabling Simultaneous Visible Light Communication and Sensing
abstract
Driven by the ubiquitous proliferation of low-cost LED luminaires, visible light communication (VLC) has been established as a high-speed communications technology based on the high-frequency modulation of an optical source. In parallel, Visible Light Sensing (VLS) has recently demonstrated how vision-based at-a-distance sensing of mechanical vibrations (e.g., of factory equipment) can be performed using high frequency optical strobing. However, to date, exemplars of VLC and VLS have been explored in isolation, without consideration of their mutual dependencies. In this work, we explore whether and how high-throughput VLC and high-coverage VLS can be simultaneously supported. We first demonstrate the existence of a fundamental VLC-vs.-VLS tradeoff, driven by the duty cycle of the strobing light source: a larger duty cycle results in higher VLC throughput but reduced VLS coverage, and vice versa. To overcome this limitation, we evaluate two approaches: (a) time-multiplexed VLC and VLS on a single strobe, and (b) harmonic multi-strobing, where multiple light sources are strobed synchronously to effectively create low-duty cycle harmonics of the base strobe frequency. Finally, we present VibranSee, an approach that improves harmonic multi-strobing by adaptively tuning both (a) the strobe duty cycle and (b) the number of strobing harmonics used. Using both analytical studies and prototype-based experiments, we show VibranSee's benefits: it simultaneously achieves VLC data goodput that is ideally only 18.6% lower (and 23.9% lower for an actual working prototype) than the maximum communication rate and infers over 96.6% (100% for the prototype) of possible vibration frequencies.
Ila Gokarn, Archan Misra
SECON1