Gautam Singh

dblp:35/2642 · DBLP profile ↗
← Back
20ranked-venue papers
10as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 9 first-author · 9 since 2021Computer networks · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Dreamweaver: Learning Compositional World Models from Pixels
abstract
Humans have an innate ability to decompose their perceptions of the world into objects and their attributes, such as colors, shapes, and movement patterns. This cognitive process enables us to imagine novel futures by recombining familiar concepts. However, replicating this ability in artificial intelligence systems has proven challenging, particularly when it comes to modeling videos into compositional concepts and generating unseen, recomposed futures without relying on auxiliary data, such as text, masks, or bounding boxes. In this paper, we propose __Dreamweaver__, a neural architecture designed to discover hierarchical and compositional representations from raw videos and generate compositional future simulations. Our approach leverages a novel Recurrent Block-Slot Unit (RBSU) to decompose videos into their constituent objects and attributes. In addition, Dreamweaver uses a multi-future-frame prediction objective to capture disentangled representations for dynamic concepts more effectively as well as static concepts. In experiments, we demonstrate our model outperforms current state-of-the-art baselines for world modeling when evaluated under the DCI framework across multiple datasets. Furthermore, we show how the modularized concept representations of our model enable compositional imagination, allowing the generation of novel videos by recombining attributes from previously seen objects. [cun-bjy.github.io/dreamweaver-website](https://cun-bjy.github.io/dreamweaver-website/)
Junyeob Baek, Yi-Fu Wu, Gautam Singh, Sungjin Ahn
ICLR3
2025 Cross-Slice Online Federated Learning Framework for Anomaly Detection in 5G Networks
abstract
Network slicing in fifth generation (5G) networks has become a key enabler for tailored resource allocation to meet diverse use cases. Operators can instantiate isolated 5G slices to meet various service-level agreements (SLAs) such as quality-of-service (QoS) profiles and latency constraints. Despite being isolated, slices introduce new attack surfaces in the 5G core. Traditional anomaly detection (AD) methods are supervised, centralized, and do not scale out across slices. Machine Learning (ML) models for AD require frequent retraining to account for drifts in traffic patterns. Centralized methods increase the risk of leaking user data, as network flow records containing Personally Identifiable Information (PII) leave the slice. To address these gaps, we propose an online, unsupervised federated learning (FL) framework that enables slices to continuously train a local ML model on unlabeled network flow data from core network functions (NFs) in their own slice and aggregate parameter updates in a privacy-preserving manner to form a robust global model. It integrates with the Network Data Analytics Function (NWDAF), which collects data from NFs and provides FL capabilities. This cross-slice knowledge transfer accelerates AD while preventing leakage of PII by keeping data within slices. With the help of FL, our proposed framework detects previously unseen anomalous traffic patterns in a particular slice in real time, achieving a macro F1 score of 93.83% on a test dataset containing new anomalous traffic flows. The results confirm that cross-slice online FL keeps AD robust, responsive, and privacy-compliant in sliced 5G core networks. To the best of our knowledge, this is the first FL-based AD framework that detects both new and recurring anomalous traffic patterns in heterogeneous 5G slices using unlabeled data in real time.
Arjit Gupta, Gautam Singh, A. Antony Franklin
MSWiM2
2024 Parallelized Spatiotemporal Slot Binding for Videos
abstract
While modern best practices advocate for scalable architectures that support long-range interactions, object-centric models are yet to fully embrace these architectures. In particular, existing object-centric models for handling sequential inputs, due to their reliance on RNN-based implementation, show poor stability and capacity and are slow to train on long sequences. We introduce Parallelizable Spatiotemporal Binder or PSB, the first temporally-parallelizable slot learning architecture for sequential inputs. Unlike conventional RNN-based approaches, PSB produces object-centric representations, known as slots, for all time-steps in parallel. This is achieved by refining the initial slots across all time-steps through a fixed number of layers equipped with causal attention. By capitalizing on the parallelism induced by our architecture, the proposed model exhibits a significant boost in efficiency. In experiments, we test PSB extensively as an encoder within an auto-encoding framework paired with a wide variety of decoder options. Compared to the state-of-the-art, our architecture demonstrates stable training on longer sequences, achieves parallelization that results in a 60% increase in training speed, and yields performance that is on par with or better on unsupervised 2D and 3D object-centric scene decomposition and understanding.
Gautam Singh, Yue Wang 0041, Jiawei Yang 0002, Boris Ivanovic, Sungjin Ahn, Marco Pavone 0001, Tong Che
ICML1
2024 Slot State Space Models
abstract
Recent State Space Models (SSMs) such as S4, S5, and Mamba have shown remarkable computational benefits in long-range temporal dependency modeling. However, in many sequence modeling problems, the underlying process is inherently modular and it is of interest to have inductive biases that mimic this modular structure. In this paper, we introduce SlotSSMs, a novel framework for incorporating independent mechanisms into SSMs to preserve or encourage separation of information. Unlike conventional SSMs that maintain a monolithic state vector, SlotSSMs maintains the state as a collection of multiple vectors called slots. Crucially, the state transitions are performed independently per slot with sparse interactions across slots implemented via the bottleneck of self-attention. In experiments, we evaluate our model in object-centric learning, 3D visual reasoning, and long-context video understanding tasks, which involve modeling multiple objects and their long-range temporal dependencies. We find that our proposed design offers substantial performance gains over existing sequence modeling methods. Project page is available at \url{https://slotssms.github.io/}
Jindong Jiang, Fei Deng 0001, Gautam Singh, Minseung Lee, Sungjin Ahn
NeurIPS3
2023 Neural Systematic Binder
Gautam Singh, Yeongbin Kim, Sungjin Ahn
ICLR1
2023 Object-Centric Slot Diffusion
abstract
The recent success of transformer-based image generative models in object-centric learning highlights the importance of powerful image generators for handling complex scenes. However, despite the high expressiveness of diffusion models in image generation, their integration into object-centric learning remains largely unexplored in this domain. In this paper, we explore the feasibility and potential of integrating diffusion models into object-centric learning and investigate the pros and cons of this approach. We introduce Latent Slot Diffusion (LSD), a novel model that serves dual purposes: it is the first object-centric learning model to replace conventional slot decoders with a latent diffusion model conditioned on object slots, and it is also the first unsupervised compositional conditional diffusion model that operates without the need for supervised annotations like text. Through experiments on various object-centric tasks, including the first application of the FFHQ dataset in this field, we demonstrate that LSD significantly outperforms state-of-the-art transformer-based decoders, particularly in more complex scenes, and exhibits superior unsupervised compositional generation quality. In addition, we conduct a preliminary investigation into the integration of pre-trained diffusion models in LSD and demonstrate its effectiveness in real-world image segmentation and generation. Project page is available at https://latentslotdiffusion.github.io
Jindong Jiang, Fei Deng 0001, Gautam Singh, Sungjin Ahn
NeurIPS3
2023 Imagine the Unseen World: A Benchmark for Systematic Generalization in Visual World Models
abstract
Systematic compositionality, or the ability to adapt to novel situations by creating a mental model of the world using reusable pieces of knowledge, remains a significant challenge in machine learning. While there has been considerable progress in the language domain, efforts towards systematic visual imagination, or envisioning the dynamical implications of a visual observation, are in their infancy. We introduce the Systematic Visual Imagination Benchmark (SVIB), the first benchmark designed to address this problem head-on. SVIB offers a novel framework for a minimal world modeling problem, where models are evaluated based on their ability to generate one-step image-to-image transformations under a latent world dynamics. The framework provides benefits such as the possibility to jointly optimize for systematic perception and imagination, a range of difficulty levels, and the ability to control the fraction of possible factor combinations used during training. We provide a comprehensive evaluation of various baseline models on SVIB, offering insight into the current state-of-the-art in systematic visual imagination. We hope that this benchmark will help advance visual systematic compositionality.
Yeongbin Kim, Gautam Singh, Junyeong Park, Caglar Gulcehre, Sungjin Ahn
NeurIPS2
2022 Illiterate DALL-E Learns to Compose
Gautam Singh, Fei Deng 0001, Sungjin Ahn
ICLR1
2022 Simple Unsupervised Object-Centric Learning for Complex and Naturalistic Videos
abstract
Unsupervised object-centric learning aims to represent the modular, compositional, and causal structure of a scene as a set of object representations and thereby promises to resolve many critical limitations of traditional single-vector representations such as poor systematic generalization. Although there have been many remarkable advances in recent years, one of the most critical problems in this direction has been that previous methods work only with simple and synthetic scenes but not with complex and naturalistic images or videos. In this paper, we propose STEVE, an unsupervised model for object-centric learning in videos. Our proposed model makes a significant advancement by demonstrating its effectiveness on various complex and naturalistic videos unprecedented in this line of research. Interestingly, this is achieved by neither adding complexity to the model architecture nor introducing a new objective or weak supervision. Rather, it is achieved by a surprisingly simple architecture that uses a transformer-based image decoder conditioned on slots and the learning objective is simply to reconstruct the observation. Our experiment results on various complex and naturalistic videos show significant improvements compared to the previous state-of-the-art.
Gautam Singh, Yi-Fu Wu, Sungjin Ahn
NeurIPS1
2021 Structured World Belief for Reinforcement Learning in POMDP
abstract
Object-centric world models provide structured representation of the scene and can be an important backbone in reinforcement learning and planning. However, existing approaches suffer in partially-observable environments due to the lack of belief states. In this paper, we propose Structured World Belief, a model for learning and inference of object-centric belief states. Inferred by Sequential Monte Carlo (SMC), our belief states provide multiple object-centric scene hypotheses. To synergize the benefits of SMC particles with object representations, we also propose a new object-centric dynamics model that considers the inductive bias of object permanence. This enables tracking of object states even when they are invisible for a long time. To further facilitate object tracking in this regime, we allow our model to attend flexibly to any spatial location in the image which was restricted in previous models. In experiments, we show that object-centric belief provides a more accurate and robust performance for filtering and generation. Furthermore, we show the efficacy of structured world belief in improving the performance of reinforcement learning, planning and supervised reasoning.
Gautam Singh, Skand Vishwanath Peri, Jung-Hyun Kim 0007, Sungjin Ahn
ICML1
2020 SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition
Zhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun, Gautam Singh, Fei Deng 0001, Jindong Jiang, Sungjin Ahn
ICLR5
2020 Robustifying Sequential Neural Processes
abstract
When tasks change over time, meta-transfer learning seeks to improve the efficiency of learning a new task via both meta-learning and transfer-learning. While the standard attention has been effective in a variety of settings, we question its effectiveness in improving meta-transfer learning since the tasks being learned are dynamic and the amount of context can be substantially smaller. In this paper, using a recently proposed meta-transfer learning model, Sequential Neural Processes (SNP), we first empirically show that it suffers from a similar underfitting problem observed in the functions inferred by Neural Processes. However, we further demonstrate that unlike the meta-learning setting, the standard attention mechanisms are not effective in meta-transfer setting. To resolve, we propose a new attention mechanism, Recurrent Memory Reconstruction (RMR), and demonstrate that providing an imaginary context that is recurrently updated and reconstructed with interaction is crucial in achieving effective attention for meta-transfer learning. Furthermore, incorporating RMR into SNP, we propose Attentive Sequential Neural Processes-RMR (ASNP-RMR) and demonstrate in various tasks that ASNP-RMR significantly outperforms the baselines.
Jaesik Yoon, Gautam Singh, Sungjin Ahn
ICML2
2019 Sequential Neural Processes
abstract
Neural Processes combine the strengths of neural networks and Gaussian processes to achieve both flexible learning and fast prediction in stochastic processes. However, a large class of problems comprise underlying temporal dependency structures in a sequence of stochastic processes that Neural Processes (NP) do not explicitly consider. In this paper, we propose Sequential Neural Processes (SNP) which incorporates a temporal state-transition model of stochastic processes and thus extends its modeling capabilities to dynamic stochastic processes. In applying SNP to dynamic 3D scene modeling, we introduce the Temporal Generative Query Networks. To our knowledge, this is the first 4D model that can deal with the temporal dynamics of 3D scenes. In experiments, we evaluate the proposed methods in dynamic (non-stationary) regression and 4D scene inference and rendering.
Gautam Singh, Jaesik Yoon, Youngsung Son, Sungjin Ahn
NeurIPS1
2018 Semantic Parsing for Technical Support Questions
abstract
Technical support problems are very complex. In contrast to regular web queries (that contain few keywords) or factoid questions (which are a few sentences), these problems usually include attributes like a detailed description of what is failing (symptom), steps taken in an effort to remediate the failure (activity), and sometimes a specific request or ask (intent). Automating support is the task of automatically providing answers to these problems given a corpus of solution documents. Traditional approaches to this task rely on information retrieval and are keyword based; looking for keyword overlap between the question and solution documents and ignoring these attributes. We present an approach for semantic parsing of technical questions that uses grammatical structure to extract these attributes as a baseline, and a CRF based model that can improve performance considerably in the presence of annotated data for training. We also demonstrate that combined with reasoning, these attributes help outperform retrieval baselines.
Abhirut Gupta, Anupama Ray, Gargi Dasgupta, Gautam Singh, Pooja Aggarwal, Prateeti Mohapatra
COLING4
2017 Neev: A cognitive support agent for content improvement in hardware tickets
abstract
IT service providers differentiate themselves through offering after-sales support for hardware and software products. Thus, businesses, including large corporations, have intricate work-flows for servicing such support requests while reducing man-hours needed. These work-flows generally operate through a ticketing system for resolving customer issues. A lot of man-hours are spent in searching old tickets for correct problem and resolution for such issues. Support requests pertaining to enterprise hardware are more challenging than desktop support for end-user products. Enterprise hardware requires deeper diagnosis involving several systems and expertise of multiple agents. In this work we propose a cognitive agent, Neev, which helps in mitigating the problem in a three-fold fashion (1) retrieving a summary of relevant ticket text (2) Tagging the relevant parts as a part-of-the-problem or a part-of-the-solution (3) Focusing on the precise problem and solution. We evaluate the performance of our system using a rank-based metric where a ticket extraction is successful if the problem or solution occur in the top-n suggestions. We report the results for varying top-n values for both problem and solution on varying severity of the tickets. We find that the accuracy for problem extraction in top-1 is 62% and it reaches 86% and 94% for top-3 and top-5 cases, respectively. Furthermore, the accuracy for solution extraction reaches 62% and 88% for top-3 and top-8 cases, respectively.
Nishtha Madaan, Gautam Singh, Arun Kumar 0002, Gargi Dasgupta
IM2
2015 A subtractive clustering scheme for text-independent online writer identification
abstract
This paper proposes a text-independent writer identification framework for online handwritten text. The method utilizes an unsupervised learning scheme termed `subtractive clustering' to discover the unique writing styles of a given author. Subtractive clustering has been adopted in the literature for the problems of image segmentation and speaker identification. To the best of our knowledge, its applicability in the domain of writer identification is yet to be explored. Unlike traditional clustering techniques such as k-means and fuzzy c-means, the subtractive clustering algorithm does not rely on the initial choice of seed points. Instead, it locates the high density regions in the feature space, and this make this scheme an interesting exploration to capture the writing styles of an author (referred to as `prototypes'). The discovered prototypes from the clustering algorithm are subsequently employed to score the authorship of an unknown handwritten text. In addition, inspired from the t f-idf approach used in document retrieval, we propose a modified scoring scheme for identifying the writer. The efficacy of the algorithms are evaluated on the paragraphs from the IAM-Online Handwritten Database.
Gautam Singh, Suresh Sundaram 0001
ICDAR1
2014 Introspective semantic segmentation
abstract
Traditional approaches for semantic segmentation work in a supervised setting assuming a fixed number of semantic categories and require sufficiently large training sets. The performance of various approaches is often reported in terms of average per pixel class accuracy and global accuracy of the final labeling. When applying the learned models in the practical settings on large amounts of unlabeled data, possibly containing previously unseen categories, it is important to properly quantify their performance by measuring a classifier's introspective capability. We quantify the confidence of the region classifiers in the context of a non-parametric k-nearest neighbor (k-NN) framework for semantic segmentation by using the so called strangeness measure. The proposed measure is evaluated by introducing confidence based image ranking and showing its feasibility on a dataset containing a large number of previously unseen categories.
Gautam Singh, Jana Kosecka
WACV1
2013 Nonparametric Scene Parsing with Adaptive Feature Relevance and Semantic Context
abstract
This paper presents a nonparametric approach to semantic parsing using small patches and simple gradient, color and location features. We learn the relevance of individual feature channels at test time using a locally adaptive distance metric. To further improve the accuracy of the nonparametric approach, we examine the importance of the retrieval set used to compute the nearest neighbours using a novel semantic descriptor to retrieve better candidates. The approach is validated by experiments on several datasets used for semantic parsing demonstrating the superiority of the method compared to the state of art approaches.
Gautam Singh, Jana Kosecka
CVPR1
2013 Localization in Urban Environments Using a Panoramic Gist Descriptor
abstract
Vision-based topological localization and mapping for autonomous robotic systems have received increased research interest in recent years. The need to map larger environments requires models at different levels of abstraction and additional abilities to deal with large amounts of data efficiently. Most successful approaches for appearance-based localization and mapping with large datasets typically represent locations using local image features. We study the feasibility of performing these tasks in urban environments using global descriptors instead and taking advantage of the increasingly common panoramic datasets. This paper describes how to represent a panorama using the global gist descriptor, while maintaining desirable invariance properties for location recognition and loop detection. We propose different gist similarity measures and algorithms for appearance-based localization and an online loop-closure detection method, where the probability of loop closure is determined in a Bayesian filtering framework using the proposed image representation. The extensive experimental validation in this paper shows that their performance in urban environments is comparable with local-feature-based approaches when using wide field-of-view images.
Ana Cristina Murillo, Gautam Singh, Jana Kosecka, Josechu J. Guerrero
IEEE Trans. Robotics2
2012 Acquiring semantics induced topology in urban environments
abstract
Methods for acquisition and maintenance of an environment model are central to a broad class of mobility and navigation problems. Towards this end, various metric, topological or hybrid models have been proposed. Due to recent advances in sensing and recognition, acquisition of semantic models of the environments have gained increased interest in the community. In this work, we will demonstrate a capability of using weak semantic models of the environment to induce different topological models, capturing the spatial semantics of the environment at different levels. In the first stage of the model acquisition, we propose to compute semantic layout of the street scenes imagery by recognizing and segmenting buildings, roads, sky, cars and trees. Given such semantic layout, we propose an informative feature characterizing the layout and train a classifier to recognize street intersections in challenging urban inner city scenes. We also show how the evidence of different semantic concepts can induce useful topological representation of the environment, which can aid navigation and localization tasks. To demonstrate the approach, we carry out experiments on a challenging dataset of omnidirectional inner city street views and report the performance of both semantic segmentation and intersection classification.
Gautam Singh, Jana Kosecka
ICRA1