EDBT 2026 Demo / reviewers in the wild / expert
Sai Saketh Rambhatla
dblp:180/2805
· DBLP profile ↗
13ranked-venue papers
4as first author
9since 2021 · last 2025
0009-0002-2135-8311ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action RecognitionabstractVideo understanding requires effective modeling of both motion and appearance information, particularly for few-shot action recognition. While recent advances in point tracking have been shown to improve few-shot action recognition, two fundamental challenges persist: selecting informative points to track and effectively modeling their motion patterns. We present Trokens, a novel approach that transforms trajectory points into semantic-aware relational tokens for action recognition. First, we introduce a semantic-aware sampling strategy to adaptively distribute tracking points based on object scale and semantic relevance. Second, we develop a motion modeling framework that captures both intra-trajectory dynamics through the Histogram of Oriented Displacements (HoD) and inter-trajectory relationships to model complex action patterns. Our approach effectively combines these trajectory tokens with semantic features to enhance appearance features with motion information, achieving state-of-the-art performance across six diverse few-shot action recognition benchmarks: Something-Something-V2 (both full and small splits), Kinetics, UCF101, HMDB51, and FineGym. For project page see https://trokens-iccv25.github.io Pulkit Kumar, Shuaiyi Huang, Matthew Walmer, Sai Saketh Rambhatla, Abhinav Shrivastava |
ICCV | 4 |
| 2024 | InstanceDiffusion: Instance-Level Control for Image GenerationabstractText-to-image diffusion models produce high quality images but do not offer control over individual instances in the image. We introduce InstanceDiffusion that adds precise instance-level control to text-to-image diffusion models. InstanceDiffusion supports free-form language conditions per instance and allows flexible ways to specify instance locations such as simple single points, scribbles, bounding boxes or intricate instance segmentation masks, and combinations thereof We propose three major changes to text-to-image models that enable precise instance-level control. Our UniFusion block enables instance-level conditions for text-to-image models, the ScaleU block improves image fidelity, and our Multi-instance Sampler improves generations for multiple instances. InstanceDiffusion significantly surpasses specialized state-of-the-art models for each location condition. Notably, on the COCO dataset, we out-perform previous state-of-the-art by 20.4%$AP_{50}^{box}$for box inputs, and 25.4% IoU for mask inputs. Xudong Wang 0007, Trevor Darrell, Sai Saketh Rambhatla, Rohit Girdhar, Ishan Misra |
CVPR | 3 |
| 2024 | Factorizing Text-to-Video Generation by Explicit Image Conditioning
Rohit Girdhar, Mannat Singh, Quentin Duval, Samaneh Azadi, Sai Saketh Rambhatla, Akbar Shah, Xi Yin 0001, Devi Parikh, Ishan Misra |
ECCV (62) | 6 |
| 2024 | Trajectory-Aligned Space-Time Tokens for Few-Shot Action Recognition
Pulkit Kumar, Namitha Padmanabhan, Luke Luo, Sai Saketh Rambhatla, Abhinav Shrivastava |
ECCV (36) | 4 |
| 2023 | MOST: Multiple Object localization with Self-supervised Transformers for object discoveryabstractWe tackle the challenging task of unsupervised object localization in this work. Recently, transformers trained with self-supervised learning have been shown to exhibit object localization properties without being trained for this task. In this work, we present Multiple Object localization with Self-supervised Transformers (MOST) that uses features of transformers trained using self-supervised learning to localize multiple objects in real world images. MOST analyzes the similarity maps of the features using box counting; a fractal analysis tool to identify tokens lying on foreground patches. The identified tokens are then clustered together, and tokens of each cluster are used to generate bounding boxes on foreground regions. Unlike recent state-of-the-art object localization methods, MOST can localize multiple objects per image and outperforms SOTA algorithms on several object localization and discovery benchmarks on PASCAL-VOC 07, 12 and COCO20k datasets. Additionally, we show that MOST can be used for self-supervised pretraining of object detectors, and yields consistent improvements on fully, semi-supervised object detection and unsupervised region proposal generation.Our project is publicly available at rssaketh.github.io/most. Sai Saketh Rambhatla, Ishan Misra, Rama Chellappa, Abhinav Shrivastava |
ICCV | 1 |
| 2023 | SparseDet: Improving Sparsely Annotated Object Detection with Pseudo-positive MiningabstractTraining with sparse annotations is known to reduce the performance of object detectors. Previous methods have focused on proxies for missing ground truth annotations in the form of pseudo-labels for unlabeled boxes. We observe that existing methods suffer at higher levels of sparsity in the data due to noisy pseudo-labels. To prevent this, we propose an end-to-end system that learns to separate the proposals into labeled and unlabeled regions using Pseudo-positive mining. While the labeled regions are processed as usual, self-supervised learning is used to process the unlabeled regions thereby preventing the negative effects of noisy pseudo-labels. This novel approach has multiple advantages such as improved robustness to higher sparsity when compared to existing methods. We conduct exhaustive experiments on five splits on the PASCAL-VOC and COCO datasets achieving state-of-the-art performance. We also unify various splits used across literature for this task and present a standardized benchmark. On average, we improve by 2.6, 3.9 and 9.6 mAP over previous state-of-the-art methods on three splits of increasing sparsity on COCO. Our project is publicly available at cs.umd.edu/~sakshams/SparseDet. Saksham Suri, Sai Saketh Rambhatla, Rama Chellappa, Abhinav Shrivastava |
ICCV | 2 |
| 2022 | An Empirical Analysis of Boosting Deep NetworksabstractBoosting is a method for finding a highly accurate classifier by linearly combining many “weak” classifiers, each of which may be only moderately accurate. Thus, boosting is a method for learning an ensemble of classifiers. While boosting has been shown to be very effective for decision trees, its impact on neural networks has not been extensively studied. Using standard object recognition datasets, we verify experimentally the well-known result that a boosted ensemble of decision trees usually generalizes much better on testing data than a single decision tree with the same number of parameters. In contrast, using the same datasets and boosting algorithms, our experiments show the opposite to be true when using neural networks (both convolutional neural networks (CNNs) and multilayer perceptrons (MLPs)). We find that a single neural network usually generalizes better than a boosted ensemble of smaller neural networks with the same total number of parameters. While this is an experimental investigation, more theoretical research is warranted to understand the role of boosting in deep learning-based classifiers. Sai Saketh Rambhatla, Michael J. Jones 0001, Rama Chellappa |
IJCNN | 1 |
| 2021 | Towards Discovery and Attribution of Open-world GAN Generated ImagesabstractWith the recent progress in Generative Adversarial Networks (GANs), it is imperative for media and visual forensics to develop detectors which can identify and attribute images to the model generating them. Existing works have shown to attribute images to their corresponding GAN sources with high accuracy. However, these works are limited to a closed set scenario, failing to generalize to GANs unseen during train time and are therefore, not scalable with a steady influx of new GANs. We present an iterative algorithm for discovering images generated from previously unseen GANs by exploiting the fact that all GANs leave distinct fingerprints on their generated images. Our algorithm consists of multiple components including network training, out-of-distribution detection, clustering, merge and refine steps. Through extensive experiments, we show that our algorithm discovers unseen GANs with high accuracy and also generalizes to GANs trained on unseen real datasets. We additionally apply our algorithm to attribution and discovery of GANs in an online fashion as well as to the more standard task of real/fake detection. Our experiments demonstrate the effectiveness of our approach to discover new GANs and can be used in an open-world setup. Sharath Girish, Saksham Suri, Sai Saketh Rambhatla, Abhinav Shrivastava |
ICCV | 3 |
| 2021 | The Pursuit of Knowledge: Discovering and Localizing Novel Categories using Dual MemoryabstractWe tackle object category discovery, which is the problem of discovering and localizing novel objects in a large unlabeled dataset. While existing methods show results on datasets with less cluttered scenes and fewer object in-stances per image, we present our results on the challenging COCO dataset. Moreover, we argue that, rather than discovering new categories from scratch, discovery algorithms can benefit from identifying what is already known and focusing their attention on the unknown. We propose a method that exploits prior knowledge about certain object types to discover new categories by leveraging two memory modules, namely Working and Semantic memory. We show the performance of our detector on the COCO minival dataset to demonstrate its in-the-wild capabilities. Sai Saketh Rambhatla, Rama Chellappa, Abhinav Shrivastava |
ICCV | 1 |
| 2020 | Detecting Human-Object Interactions via Functional GeneralizationabstractWe present an approach for detecting human-object interactions (HOIs) in images, based on the idea that humans interact with functionally similar objects in a similar manner. The proposed model is simple and efficiently uses the data, visual features of the human, relative spatial orientation of the human and the object, and the knowledge that functionally similar objects take part in similar interactions with humans. We provide extensive experimental validation for our approach and demonstrate state-of-the-art results for HOI detection. On the HICO-Det dataset our method achieves a gain of over 2.5% absolute points in mean average precision (mAP) over state-of-the-art. We also show that our approach leads to significant performance gains for zero-shot HOI detection in the seen object setting. We further demonstrate that using a generic object detector, our model can generalize to interactions involving previously unseen objects. Ankan Bansal, Sai Saketh Rambhatla, Abhinav Shrivastava, Rama Chellappa |
AAAI | 2 |
| 2019 | Body Part Alignment and Temporal Attention Pooling for Video-Based Person Re-Identification
Sai Saketh Rambhatla, Michael J. Jones 0001 |
BMVC | 1 |
| 2019 | A Dual-Path Model With Adaptive Attention for Vehicle Re-IdentificationabstractIn recent years, attention models have been extensively used for person and vehicle re-identification. Most re-identification methods are designed to focus attention on key-point locations. However, depending on the orientation, the contribution of each key-point varies. In this paper, we present a novel dual-path adaptive attention model for vehicle re-identification (AAVER). The global appearance path captures macroscopic vehicle features while the orientation conditioned part appearance path learns to capture localized discriminative features by focusing attention on the most informative key-points. Through extensive experimentation, we show that the proposed AAVER method is able to accurately re-identify vehicles in unconstrained scenarios, yielding state of the art results on the challenging dataset VeRi-776. As a byproduct, the proposed system is also able to accurately predict vehicle key-points and shows an improvement of more than 7% over state of the art. The code for key-point estimation model is available at https://github.com/Pirazh/Vehicle_Key_ Point_Orientation_Estimation. Pirazh Khorramshahi, Amit Kumar 0013, Neehar Peri, Sai Saketh Rambhatla, Jun-Cheng Chen, Rama Chellappa |
ICCV | 4 |
| 2016 | Camera based estimation of respiration rate by analyzing shape and size variation of structured lightabstractRespiration rate is a key parameter that is monitored in intensive care units. The current solutions for respiration rate require that a sensor is placed in contact with the subject to derive it. However, this will cause discomfort to the subject and may damage the fragile skin if the subject were a neonate. There are several other applications where the respiration signal is used in clinical settings. The respiration signal is used to correct for organ motion during diagnosis scans, e.g., computed tomography. It is expected that the respiration signal results in clear peaks and troughs during the breathing cycle of the subject so that these corrections can be made. We propose a contactless system and method by using a camera to derive a faithful respiration signal. We project a circular dot of light onto the chest and abdomen region of the subject and monitor the shape and size changes of it using a camera. In an oblique projection of the structured light, both the shape and size of the dot will vary with breathing. We segment the dot and derive the respiration signal from it. Numerical results presented show that we are able to obtain the respiration rate accurately. Vishnu Vardhan Makkapati, Sai Saketh Rambhatla |
ICASSP | 2 |