VLDB 2026 Research / reviewers in the wild / expert
Sirisha Rambhatla
dblp:123/4808
· DBLP profile ↗
16ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-9389-727XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Boundary-aware semantic segmentation for ice hockey rink registration
Amir Nazemi, Stephie Liu, Sirisha Rambhatla, Yuhao Chen 0001, David A. Clausi |
Comput. Vis. Image Underst. | 4 |
| 2025 | LOCATEdit: Graph Laplacian Optimized Cross Attention for Localized Text-Guided Image Editing
Achint Soni, Meet Soni, Sirisha Rambhatla |
ICCV | 3 |
| 2025 | SubTrack++ : Gradient Subspace Tracking for Scalable LLM TrainingabstractTraining large language models (LLMs) is highly resource-intensive due to their massive number of parameters and the overhead of optimizer states. While recent work has aimed to reduce memory consumption, such efforts often entail trade-offs among memory efficiency, training time, and model performance. Yet, true democratization of LLMs requires simultaneous progress across all three dimensions. To this end, we propose SubTrack++ that leverages Grassmannian gradient subspace tracking combined with projection-aware optimizers, enabling Adam’s internal statistics to adapt to subspace changes. Additionally, employing recovery scaling, a technique that restores information lost through low-rank projections, further enhances model performance. Our method demonstrates SOTA convergence by exploiting Grassmannian geometry, **reducing training wall-time by up to 65%** compared to the best performing baseline, LDAdam, while preserving the reduced memory footprint. Code is at https://github.com/criticalml-uw/SubTrack. Sahar Rajabi, Nayeema Nonta, Sirisha Rambhatla |
NeurIPS | 3 |
| 2025 | Bringing Equity to Classification: Domain Generalization for Domain-Linked ClassesabstractDomain generalization (DG) focuses on transferring domain-invariant knowledge from multiple source (training) domains to an a priori unseen target domain(s). This task implicitly requires that classes of interest are expressed in multiple sources (domain-shared) to break spurious domain-class correlations. However, real-world data scarcity challenges may often result in classes present in only a specific domain (domain-linked), which we show leads to extremely poor generalization. In this work, we introduce the domain-linked DG task to the community and develop a methodology to learn useful domain-invariant representations from domain-shared classes for domain-linked ones. Specifically, we propose FOND, a Fairness-inspired and cONtrastive learning objective for Domain-linked DG. Rigorous and reproducible experimental results communicate that FOND accomplishes state-of-the-art improvements for domain-linked classes, given a sufficient number of domain-shared classes and with minimal performance trade-offs. Complementary to these contributions, we theoretically analyze this task and provide practical insights for domain-linked class generalizability. Kimathi Kaai, Saad Hossain, Sirisha Rambhatla |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Why do Variational Autoencoders Really Promote Disentanglement?abstractDespite not being designed for this purpose, the use of variational autoencoders (VAEs) has proven remarkably effective for disentangled representation learning (DRL). Recent research attributes this success to certain characteristics of the loss function that prevent latent space rotation, or hypothesize about the orthogonality properties of the decoder by drawing parallels with principal component analysis (PCA). This hypothesis, however, has only been tested experimentally for linear VAEs, and the theoretical justification still remains an open problem. Moreover, since real-world VAEs are often inherently non-linear due to the use of neural architectures, understanding DRL capabilities of real-world VAEs remains a critical task. Our work takes a step towards understanding disentanglement in real-world VAEs to theoretically establish how the orthogonality properties of the decoder promotes disentanglement in practical applications. Complementary to our theoretical contributions, our experimental results corroborate our analysis. Code is available at https://github.com/criticalml-uw/Disentanglement-in-VAE. Pratik Bhowal, Achint Soni, Sirisha Rambhatla |
ICML | 3 |
| 2024 | Seeing Beyond the Crop: Using Language Priors for Out-of-Bounding Box Keypoint PredictionabstractAccurate estimation of human pose and the pose of interacting objects, like a hockey stick, is crucial for action recognition and performance analysis, particularly in sports. Existing methods capture the object along with the human in the bounding boxes, assuming all keypoints are visible within the bounding box. This necessitates larger bounding boxes to capture the object, introducing unnecessary visual features and hindering performance in real-world cluttered environments. We propose a simple image and text-based multimodal solution TokenCLIPose that addresses this limitation. Our approach focuses solely on human keypoints within the bounding box, treating objects as unseen. TokenCLIPose leverages the rich semantic representations endowed by language for inducing keypoint-specific context, even for occluded keypoints. We evaluate the performance of TokenCLIPose on a real-world Ice-Hockey dataset, and demonstrate its generalizability through zero-shot transfer to a smaller Lacrosse dataset. Additionally, we showcase its flexibility on CrowdPose, a popular occlusion benchmark with keypoints within the bounding box. Our method significantly improves over state-of-the-art approaches on all three datasets, with gains of 4.36\%, 2.35\%, and 3.8\%, respectively. Bavesh Balaji, Jerrin Bright, Yuhao Chen 0001, Sirisha Rambhatla, John S. Zelek, David A. Clausi |
NeurIPS | 4 |
| 2023 | Domain-Guided Spatio-Temporal Self-Attention for Egocentric 3D Pose EstimationabstractVision-based ego-centric 3D human pose estimation (ego-HPE) is essential to support critical applications of xR-technologies. However, severe self-occlusions and strong distortion introduced by the fish-eye view from the head mounted camera, make ego-HPE extremely challenging. To address these challenges, we propose a domain-guided spatio-temporal transformer model that leverages information specific to ego-views. Powered by this domain-guided transformer, we build Egocentric Spatio-Temporal Self-Attention Network (Ego-STAN), which uses 2D image representations and spatio-temporal attention to address both distortions and self-occlusions in ego-HPE. Additionally, we introduce a spatial concept called feature map tokens (FMT) which endows Ego-STAN with the ability to draw complex spatio-temporal information encoded in ego-centric videos. Our quantitative evaluation on the contemporary xR-EgoPose dataset, achieves a 38.2% improvement on the highest error joints against the SOTA ego-HPE model, while accomplishing a 22% decrease in the number of parameters. Finally, we also demonstrate the generalization capabilities of our model to real-world HPE tasks beyond ego-views achieving 7.7% improvement on 2D human pose estimation with the Human3.6M dataset. Our code is also made available at: https://github.com/jmpark0808/Ego-STAN Jinman Park, Kimathi Kaai, Saad Hossain, Norikatsu Sumi, Sirisha Rambhatla, Paul W. Fieguth |
KDD | 5 |
| 2022 | I-SEA: Importance Sampling and Expected Alignment-Based Deep Distance Metric Learning for Time Series Analysis and EmbeddingabstractLearning effective embeddings for potentially irregularly sampled time-series, evolving at different time scales, is fundamental for machine learning tasks such as classification and clustering. Task-dependent embeddings rely on similarities between data samples to learn effective geometries. However, many popular time-series similarity measures are not valid distance metrics, and as a result they do not reliably capture the intricate relationships between the multi-variate time-series data samples for learning effective embeddings. One of the primary ways to formulate an accurate distance metric is by forming distance estimates via Monte-Carlo-based expectation evaluations. However, the high-dimensionality of the underlying distribution, and the inability to sample from it, pose significant challenges. To this end, we develop an Importance Sampling based distance metric -- I-SEA -- which enjoys the properties of a metric while consistently achieving superior performance for machine learning tasks such as classification and representation learning. I-SEA leverages Importance Sampling and Non-parametric Density Estimation to adaptively estimate distances, enabling implicit estimation from the underlying high-dimensional distribution, resulting in improved accuracy and reduced variance. We theoretically establish the properties of I-SEA and demonstrate its capabilities via experimental evaluations on real-world healthcare datasets. Sirisha Rambhatla, Zhengping Che, Yan Liu 0002 |
AAAI | 1 |
| 2021 | DL4Burn: Burn Surgical Candidacy Prediction using Multimodal Deep Learning
Sirisha Rambhatla, Samantha Huang, Loc Trinh, Mengfei Zhang, Boyuan Long, Mingtao Dong, Vyom Unadkat, Haig Yenikomshian, Justin Gillenwater, Yan Liu 0002 |
AMIA | 1 |
| 2021 | Physics-aware Spatiotemporal Modules with Auxiliary Tasks for Meta-LearningabstractModeling the dynamics of real-world physical systems is critical for spatiotemporal prediction tasks, but challenging when data is limited. The scarcity of real-world data and the difficulty in reproducing the data distribution hinder directly applying meta-learning techniques. Although the knowledge of governing partial differential equations (PDE) of the data can be helpful for the fast adaptation to few observations, it is mostly infeasible to exactly find the equation for observations in real-world physical systems. In this work, we propose a framework, physics-aware meta-learning with auxiliary tasks, whose spatial modules incorporate PDE-independent knowledge and temporal modules utilize the generalized features from the spatial modules to be adapted to the limited data, respectively. The framework is inspired by a local conservation law expressed mathematically as a continuity equation and does not require the exact form of governing equation to model the spatiotemporal observations. The proposed method mitigates the need for a large number of real-world tasks for meta-learning by leveraging spatial information in simulated data to meta-initialize the spatial modules. We apply the proposed framework to both synthetic and real-world spatiotemporal prediction tasks and demonstrate its superior performance with limited observations. Sungyong Seo, Chuizheng Meng, Sirisha Rambhatla, Yan Liu 0002 |
IJCAI | 3 |
| 2021 | Cross-Node Federated Graph Neural Network for Spatio-Temporal Data ModelingabstractVast amount of data generated from networks of sensors, wearables, and the Internet of Things (IoT) devices underscores the need for advanced modeling techniques that leverage the spatio-temporal structure of decentralized data due to the need for edge computation and licensing (data access) issues. While federated learning (FL) has emerged as a framework for model training without requiring direct data sharing and exchange, effectively modeling the complex spatio-temporal dependencies to improve forecasting capabilities still remains an open problem. On the other hand, state-of-the-art spatio-temporal forecasting models assume unfettered access to the data, neglecting constraints on data sharing. To bridge this gap, we propose a federated spatio-temporal model -- Cross-Node Federated Graph Neural Network (CNFGNN) -- which explicitly encodes the underlying graph structure using graph neural network (GNN)-based architecture under the constraint of cross-node federated learning, which requires that data in a network of nodes is generated locally on each node and remains decentralized. CNFGNN operates by disentangling the temporal dynamics modeling on devices and spatial dynamics on the server, utilizing alternating optimization to reduce the communication cost, facilitating computations on the edge devices. Experiments on the traffic flow forecasting task show that CNFGNN achieves the best forecasting performance in both transductive and inductive learning settings with no extra computation cost on edge devices, while incurring modest communication cost. Chuizheng Meng, Sirisha Rambhatla, Yan Liu 0002 |
KDD | 2 |
| 2021 | Interpretable and Trustworthy Deepfake Detection via Dynamic PrototypesabstractIn this paper we propose a novel human-centered approach for detecting forgery in face images, using dynamic prototypes as a form of visual explanations. Currently, most state-of-the-art deepfake detections are based on black-box models that process videos frame-by-frame for inference, and few closely examine their temporal inconsistencies. However, the existence of such temporal artifacts within deepfake videos is key in detecting and explaining deepfakes to a supervising human. To this end, we propose Dynamic Prototype Network (DPNet) - an interpretable and effective solution that utilizes dynamic representations (i.e., prototypes) to explain deepfake temporal artifacts. Extensive experimental results show that DPNet achieves competitive predictive performance, even on unseen testing datasets such as Google's DeepFakeDetection, DeeperForensics, and Celeb-DF, while providing easy referential explanations of deepfake dynamics. On top of DPNet's prototypical framework, we further formulate temporal logic specifications based on these dynamics to check our model's compliance to desired temporal behaviors, hence providing trustworthiness for such critical detection systems. Loc Trinh, Michael Tsang, Sirisha Rambhatla, Yan Liu 0002 |
WACV | 3 |
| 2020 | Provable Online CP/PARAFAC Decomposition of a Structured Tensor via Dictionary LearningabstractWe consider the problem of factorizing a structured 3-way tensor into its constituent Canonical Polyadic (CP) factors. This decomposition, which can be viewed as a generalization of singular value decomposition (SVD) for tensors, reveals how the tensor dimensions (features) interact with each other. However, since the factors are a priori unknown, the corresponding optimization problems are inherently non-convex. The existing guaranteed algorithms which handle this non-convexity incur an irreducible error (bias), and only apply to cases where all factors have the same structure. To this end, we develop a provable algorithm for online structured tensor factorization, wherein one of the factors obeys some incoherence conditions, and the others are sparse. Specifically we show that, under some relatively mild conditions on initialization, rank, and sparsity, our algorithm recovers the factors exactly (up to scaling and permutation) at a linear rate. Complementary to our theoretical results, our synthetic and real-world data evaluations showcase superior performance compared to related techniques. Sirisha Rambhatla, Xingguo Li, Jarvis D. Haupt |
NeurIPS | 1 |
| 2020 | How does This Interaction Affect Me? Interpretable Attribution for Feature InteractionsabstractMachine learning transparency calls for interpretable explanations of how inputs relate to predictions. Feature attribution is a way to analyze the impact of features on predictions. Feature interactions are the contextual dependence between features that jointly impact predictions. There are a number of methods that extract feature interactions in prediction models; however, the methods that assign attributions to interactions are either uninterpretable, model-specific, or non-axiomatic. We propose an interaction attribution and detection framework called Archipelago which addresses these problems and is also scalable in real-world settings. Our experiments on standard annotation labels indicate our approach provides significantly more interpretable explanations than comparable methods, which is important for analyzing the impact of interactions on predictions. We also provide accompanying visualizations of our approach that give new insights into deep neural networks. Michael Tsang, Sirisha Rambhatla, Yan Liu 0002 |
NeurIPS | 2 |
| 2019 | NOODL: Provable Online Dictionary Learning and Sparse Coding
Sirisha Rambhatla, Xingguo Li, Jarvis D. Haupt |
ICLR (Poster) | 1 |
| 2018 | Robust PCA via Dictionary Based Outlier PursuitabstractIn this paper, we examine the problem of locating vector outliers from a large number of inliers, with a particular focus on the case where the outliers are represented in a known basis or dictionary. Using a convex demixing formulation, we provide provable guarantees for exact recovery of the space spanned by the inliers and the supports of the outlier columns, even when the rank of inliers is high and the number of outliers is a constant proportion of total observations. Comprehensive numerical experiments on both synthetic and hyper-spectral imaging real datasets demonstrate the efficiency of our proposed method. Xingguo Li, Jineng Ren, Sirisha Rambhatla, Jarvis D. Haupt |
ICASSP | 3 |