VLDB 2026 Research / reviewers in the wild / expert
Shishira R. Maiya
dblp:230/4408
· DBLP profile ↗
7ranked-venue papers
4as first author
6since 2021 · last 2024
0000-0002-5346-9510ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Explaining the Implicit Neural Canvas: Connecting Pixels to Neurons by Tracing Their ContributionsabstractThe many variations of Implicit Neural Representations (INRs), where a neural network is trained as a continuous representation of a signal, have tremendous practical utility for downstream tasks including novel view synthesis, video compression, and image super-resolution. Unfortunately, the inner workings of these networks are seriously under-studied. Our work, eXplaining the Implicit Neural Canvas (XINC), is a unified framework for explaining properties of INRs by examining the strength of each neuron's contribution to each output pixel. We call the aggregate of these contribution maps the Implicit Neural Canvas and we use this concept to demonstrate that the INRs we study learn to “see” the frames they represent in surprising ways. For ex-ample, INRs tend to have highly distributed representations. While lacking high-level object semantics, they have a sig-nificant bias for color and edges, and are almost entirely space-agnostic. We arrive at our conclusions by examining how objects are represented across time in video INRs, using clustering to visualize similar neurons across layers and architectures, and show that this is dominated by motion. These insights demonstrate the general usefulness of our analysis framework. Namitha Padmanabhan, Matthew Gwilliam, Pulkit Kumar, Shishira R. Maiya, Max Ehrlich, Abhinav Shrivastava |
CVPR | 4 |
| 2024 | Latent-INR: A Flexible Framework for Implicit Representations of Videos with Discriminative Semantics
Shishira R. Maiya, Matthew Gwilliam, Max Ehrlich, Abhinav Shrivastava |
ECCV (15) | 1 |
| 2024 | LEIA: Latent View-Invariant Embeddings for Implicit 3D Articulation
Archana Swaminathan, Kamal Gupta 0002, Shishira R. Maiya, Vatsal Agarwal, Abhinav Shrivastava |
ECCV (16) | 4 |
| 2023 | Unifying the Harmonic Analysis of Adversarial Attacks and Robustness
Shishira R. Maiya, Max Ehrlich, Vatsal Agarwal, Ser-Nam Lim, Tom Goldstein, Abhinav Shrivastava |
BMVC | 1 |
| 2023 | NIRVANA: Neural Implicit Representations of Videos with Adaptive Networks and Autoregressive Patch-Wise ModelingabstractImplicit Neural Representations (INR) have recently shown to be powerful tool for high-quality video compression. However, existing works are are limiting as they do not exploit the temporal redundancy in videos, leading to a long encoding time. Additionally, these methods have fixed architectures which do not scale to longer videos or higher resolutions. To address these issues, we propose NIRVANA, which treats videos as groups of frames and fits separate networks to each group performing patch-wise prediction. The video representation is modeled autoregressively, with networks fit on a current group initialized using weights from the previous group's model. To enhance efficiency, we quantize the parameters during training, requiring no post-hoc pruning or quantization. When compared with previous works on the benchmark UVG dataset, NIRVANA improves encoding quality from 37.36 to 37.70 (in terms of PSNR) and the encoding speed by 12x, while maintaining the same compression rate. In contrast to prior video INR works which struggle with larger resolution and longer videos, we show that our algorithm scales naturally due to its patch-wise and autoregressive design. Moreover, our method achieves variable bitrate compression by adapting to videos with varying inter-frame motion. NIRVANA also achieves 6x decoding speed scaling well with more GPUs, making it practical for various deployment scenarios.11The project site can be found here. Shishira R. Maiya, Sharath Girish, Max Ehrlich, Hanyu Wang 0002, Kwot Sin Lee, Patrick Poirson, Pengxiang Wu, Abhinav Shrivastava |
CVPR | 1 |
| 2021 | The Lottery Ticket Hypothesis for Object RecognitionabstractRecognition tasks, such as object recognition and key-point estimation, have seen widespread adoption in recent years. Most state-of-the-art methods for these tasks use deep networks that are computationally expensive and have huge memory footprints. This makes it exceedingly difficult to deploy these systems on low power embedded devices. Hence, the importance of decreasing the storage requirements and the amount of computation in such models is paramount. The recently proposed Lottery Ticket Hypothesis (LTH) states that deep neural networks trained on large datasets contain smaller subnetworks that achieve on par performance as the dense networks. In this work, we perform the first empirical study investigating LTH for model pruning in the context of object detection, instance segmentation, and keypoint estimation. Our studies reveal that lottery tickets obtained from Imagenet pretraining do not transfer well to the downstream tasks. We provide guidance on how to find lottery tickets with up to 80% overall sparsity on different sub-tasks without incurring any drop in the performance. Finally, we analyse the behavior of trained tickets with respect to various task attributes such as object size, frequency, and difficulty of detection. Sharath Girish, Shishira R. Maiya, Kamal Gupta 0002, Hao Chen 0066, Larry Davis 0001, Abhinav Shrivastava |
CVPR | 2 |
| 2020 | Rethinking Retinal Landmark Localization as Pose Estimation: Naïve Single Stacked Network for Optic Disk and Fovea DetectionabstractAutomatic detection of optic disk and fovea, the two fundamental biological landmarks of the retinal system, is crucial to track the disease progression in a diabetic patient. Recent advances in this direction were mostly limited to applying CNN based networks to aggressively extract visual geometric features. In a departure from that practice, we put forward the notion of treating the landmark detection problem in human eye scans as a pose estimation problem owing to the anatomical geometrical relationship between optic disk and fovea. In this regard, we present Naive Single Stacked Hourglass (NSSH) network which learns the spatial orientation and pixel intensity contrast between optic disk and fovea to accurately pinpoint their locations. NSSH network significantly reduces the mean squared loss, thus outperforming all previously known techniques and establishing a state of the art in both optic disk and fovea localization tasks. Shishira R. Maiya, Puneet Mathur |
ICASSP | 1 |