Maneesh Kumar Singh 0001

dblp:263/9205-1 · also Maneesh Singh 0001 · DBLP profile ↗
← Back
42ranked-venue papers
6as first author
15since 2021 · last 2025
0000-0002-7414-1813ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2Theory of computation · 1
YearPublicationVenuePosition
2025 Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)
abstract
Transformer-based language models, though not explicitly trained to mimic brain recordings, have demonstrated surprising alignment with brain activity. Progress in these models—through increased size, instruction-tuning, and multimodality—has led to better representational alignment with neural data. Recently, a new class of instruction-tuned multimodal LLMs (MLLMs) have emerged, showing remarkable zero-shot capabilities in open-ended multimodal vision tasks. However, it is unknown whether MLLMs, when prompted with natural instructions, lead to better brain alignment and effectively capture instruction-specific representations. To address this, we first investigate the brain alignment, i.e., measuring the degree of predictivity of neural visual activity using text output response embeddings from MLLMs as participants engage in watching natural scenes. Experiments with 10 different instructions (like image captioning, visual question answering, etc.) show that MLLMs exhibit significantly better brain alignment than vision-only models and perform comparably to non-instruction-tuned multimodal models like CLIP. We also find that while these MLLMs are effective at generating high-quality responses suitable to the task-specific instructions, not all instructions are relevant for brain alignment. Further, by varying instructions, we make the MLLMs encode instruction-specific visual concepts related to the input image. This analysis shows that MLLMs effectively capture count-related and recognition-related concepts, demonstrating strong alignment with brain activity. Notably, the majority of the explained variance of the brain encoding models is shared between MLLM embeddings of image captioning and other instructions. These results indicate that enhancing MLLMs' ability to capture more task-specific information could allow for better differentiation between various types of instructions, and hence improve their precision in predicting brain responses.
Subba Reddy Oota, Akshett Rai Jindal, Ishani Mondal, Khushbu Pahwa, Satya Sai Srinath Namburi, Manish Shrivastava 0001, Maneesh Kumar Singh 0001, Raju S. Bapi, Manish Gupta 0001
ICLR7
2025 Multi-modal brain encoding models for multi-modal stimuli
abstract
Despite participants engaging in unimodal stimuli, such as watching images or silent videos, recent work has demonstrated that multi-modal Transformer models can predict visual brain activity impressively well, even with incongruent modality representations. This raises the question of how accurately these multi-modal models can predict brain activity when participants are engaged in multi-modal stimuli. As these models grow increasingly popular, their use in studying neural activity provides insights into how our brains respond to such multi-modal naturalistic stimuli, i.e., where it separates and integrates information across modalities through a hierarchy of early sensory regions to higher cognition (language regions). We investigate this question by using multiple unimodal and two types of multi-modal models—cross-modal and jointly pretrained—to determine which type of models is more relevant to fMRI brain activity when participants are engaged in watching movies (videos with audio). We observe that both types of multi-modal models show improved alignment in several language and visual regions. This study also helps in identifying which brain regions process unimodal versus multi-modal information. We further investigate the contribution of each modality to multi-modal alignment by carefully removing unimodal features one by one from multi-modal representations, and find that there is additional information beyond the unimodal embeddings that is processed in the visual and language regions. Based on this investigation, we find that while for cross-modal models, their brain alignment is partially attributed to the video modality; for jointly pretrained models, it is partially attributed to both the video and audio modalities. These findings serve as strong motivation for the neuro-science community to investigate the interpretability of these models for deepening our understanding of multi-modal information processing in brain.
Subba Reddy Oota, Khushbu Pahwa, Mounika Marreddy, Maneesh Kumar Singh 0001, Manish Gupta 0001, Raju S. Bapi
ICLR4
2025 Sovereign & Shared: Frugally Scalable Multilingual-Multimodal AI for Bharat
abstract
The movement for Sovereign AI is accelerating. Meeting its promise requires vertically integrated AI stacks -spanning data, models, and reasoning systems- that remain sovereign while adhering to shared scientific principles around which global research communities can coalesce. This talk presents BharatGen as a sovereign-yet-shared effort to make AI work for all: creation of datasets, benchmarks, and models that natively support Indian languages, dialects, and code mixing across text, speech, and vision; data pipelines grounded in local realities; and frugal methods that reduce cost and lower barriers. We outline our journey to date across language infrastructure, efficient training and distillation, and early sector pilots. The R&D deep dive will draw from some of our recent work on cross-lingual knowledge distillation for low-resource languages, tokenization/phonetic design for code-mix robustness, or trustworthy document AI with visual grounding focusing on robustness under dialect/code-mix shift, and latency/cost trade-offs. We hope to inspire other Sovereign-AI efforts, especially in the low-resource ecosystems of the Global South and close by inviting international collaborations toward principled research to build people-serving AI.
Maneesh Kumar Singh 0001
ACM Multimedia1
2023 Unveiling The Mask of Position-Information Pattern Through the Mist of Image Features
abstract
Recent studies have shown that paddings in convolutional neural networks encode absolute position information which can negatively affect the model performance for certain tasks. However, existing metrics for quantifying the strength of positional information remain unreliable and frequently lead to erroneous results. To address this issue, we propose novel metrics for measuring and visualizing the encoded positional information. We formally define the encoded information as Position-information Pattern from Padding (PPP) and conduct a series of experiments to study its properties as well as its formation. The proposed metrics measure the presence of positional information more reliably than the existing metrics based on PosENet and tests in F-Conv. We also demonstrate that for any extant (and proposed) padding schemes, PPP is primarily a learning artifact and is less dependent on the characteristics of the underlying padding schemes.
Chieh Hubert Lin, Hung-Yu Tseng, Hsin-Ying Lee 0001, Maneesh Kumar Singh 0001, Ming-Hsuan Yang 0001
ICML4
2022 AutoSDF: Shape Priors for 3D Completion, Reconstruction and Generation
abstract
Powerful priors allow us to perform inference with in-sufficient information. In this paper, we propose an au-toregressive prior for 3D shapes to solve multimodal 3D tasks such as shape completion, reconstruction, and gener-ation. We model the distribution over 3D shapes as a non-sequential autoregressive distribution over a discretized, low-dimensional, symbolic grid-like latent representation of 3D shapes. This enables us to represent distributions over 3D shapes conditioned on information from an arbitrary set of spatially anchored query locations and thus perform shape completion in such arbitrary settings (e.g. generating a complete chair given only a view of the back leg). We also show that the learned autoregressive prior can be leveraged for conditional tasks such as single-view reconstruction and language-based generation. This is achieved by learning task-specific ‘naive’ conditionals which can be approxi-mated by light-weight models trained on minimal paired data. We validate the effectiveness of the proposed method using both quantitative and qualitative evaluation and show that the proposed method outperforms the specialized state-of-the-art methods trained for individual tasks. The project page with code and video visualizations can be found at https://yccyenchicheng.github.io/AutoSDF/.
Paritosh Mittal, Yen-Chi Cheng, Maneesh Kumar Singh 0001, Shubham Tulsiani
CVPR3
2022 Hierarchical Semantic Regularization of Latent Spaces in StyleGANs
Tejan Karmali, Rishubh Parihar, Susmit Agrawal, Harsh Rangwani, Varun Jampani, Maneesh Kumar Singh 0001, Venkatesh Babu Radhakrishnan
ECCV (15)6
2022 DocInfer: Document-level Natural Language Inference using Optimal Evidence Selection
abstract
We present DocInfer -a novel, end-to-end Document-level Natural Language Inference model that builds a hierarchical document graph enriched through inter-sentence relations (topical, entity-based, concept-based), performs paragraph pruning using the novel SubGraph Pooling layer, followed by optimal evidence selection based on REINFORCE algorithm to identify the most important context sentences for a given hypothesis.Our evidence selection mechanism allows it to transcend the input length limitation of modern BERT-like Transformer models while presenting the entire evidence together for inferential reasoning.We show this is an important property needed to reason on large documents where the evidence may be fragmented and located arbitrarily far from each other.Extensive experiments on popular corpora -DocNLI, ContractNLI, and ConTRoL datasets, and our new proposed dataset called CaseHoldNLI on the task of legal judicial reasoning, demonstrate significant performance gains of 8-12% over SOTA methods.Our ablation studies validate the impact of our model.Performance improvement of ∼ 3 -6% on annotation-scarce downstream tasks of fact verification, multiple-choice QA, and contract clause retrieval demonstrates the usefulness of DocInfer beyond primary NLI tasks.
Puneet Mathur, Gautam Kunapuli, Riyaz A. Bhat, Manish Shrivastava 0001, Dinesh Manocha, Maneesh Kumar Singh 0001
EMNLP6
2022 Acoustic Representation Learning on Breathing and Speech Signals for COVID-19 Detection
Debottam Dutta, Debarpan Bhattacharya, Sriram Ganapathy, Amir Hossein Poorjam, Deepak Mittal, Maneesh Kumar Singh 0001
INTERSPEECH6
2022 SPLICEOUT: A Simple and Efficient Audio Augmentation Method
abstract
Time masking has become a de facto augmentation technique for speech and audio tasks, including automatic speech recognition (ASR) and audio classification, most notably as a part of SpecAugment.In this work, we propose SPLICEOUT, a simple modification to time masking which makes it computationally more efficient.SPLICEOUT performs comparably to (and sometimes outperforms) SpecAugment on a wide variety of speech and audio tasks, including ASR for seven different languages using varying amounts of training data, as well as on speech translation, sound and music classification, thus establishing itself as a broadly applicable audio augmentation method.SPLICEOUT also provides additional gains when used in conjunction with other augmentation techniques.Apart from the fully-supervised setting, we also demonstrate that SPLICEOUT can complement unsupervised representation learning with performance gains in the semi-supervised and self-supervised settings.
Arjit Jain, Pranay Reddy Samala, Deepak Mittal, Preethi Jyothi, Maneesh Kumar Singh 0001
INTERSPEECH5
2022 VQ-Flows: Vector quantized local normalizing flows
abstract
Normalizing flows provide an elegant approach to generative modeling that allows for efficient sampling and exact density evaluation of unknown data distributions. However, current techniques have significant limitations in their expressivity when the data distribution is supported on a low-dimensional manifold or has a non-trivial topology. We introduce a novel statistical framework for learning a mixture of local normalizing flows as “chart maps” over the data manifold. Our framework augments the expressivity of recent approaches while preserving the signature property of normalizing flows, that they admit exact density evaluation. We learn a suitable atlas of charts for the data manifold via a vector quantized auto-encoder (VQ-AE) and the distributions over them using a conditional flow. We validate experimentally that our probabilistic framework enables existing approaches to better model data distributions over complex manifolds.
Sahil Sidheekh, Chris B. Dock, Tushar Jain, Radu V. Balan, Maneesh Kumar Singh 0001
UAI5
2022 Is My Model Using The Right Evidence? Systematic Probes for Examining Evidence-Based Tabular Reasoning
abstract
Abstract Neural models command state-of-the-art performance across NLP tasks, including ones involving “reasoning”. Models claiming to reason about the evidence presented to them should attend to the correct parts of the input while avoiding spurious patterns therein, be self-consistent in their predictions across inputs, and be immune to biases derived from their pre-training in a nuanced, context- sensitive fashion. Do the prevalent *BERT- family of models do so? In this paper, we study this question using the problem of reasoning on tabular data. Tabular inputs are especially well-suited for the study—they admit systematic probes targeting the properties listed above. Our experiments demonstrate that a RoBERTa-based model, representative of the current state-of-the-art, fails at reasoning on the following counts: it (a) ignores relevant parts of the evidence, (b) is over- sensitive to annotation artifacts, and (c) relies on the knowledge encoded in the pre-trained language model rather than the evidence presented in its tabular inputs. Finally, through inoculation experiments, we show that fine- tuning the model on perturbed data does not help it overcome the above challenges.
Vivek Gupta 0001, Riyaz A. Bhat, Atreya Ghosal, Manish Shrivastava 0001, Maneesh Kumar Singh 0001, Vivek Srikumar
Trans. Assoc. Comput. Linguistics5
2021 Learning to Stylize Novel Views
abstract
We tackle a 3D scene stylization problem — generating stylized images of a scene from arbitrary novel views given a set of images of the same scene and a reference image of the desired style as inputs. Direct solution of combining novel view synthesis and stylization approaches lead to results that are blurry or not consistent across different views. We propose a point cloud-based method for consistent 3D scene stylization. First, we construct the point cloud by back-projecting the image features to the 3D space. Second, we develop point cloud aggregation modules to gather the style information of the 3D scene, and then modulate the features in the point cloud with a linear transformation matrix. Finally, we project the transformed features to 2D space to obtain the novel views. Experimental results on two diverse datasets of real-world scenes validate that our method generates consistent stylized novel view synthesis results against other alternative approaches.
Hsin-Ping Huang, Hung-Yu Tseng, Saurabh Saini, Maneesh Kumar Singh 0001, Ming-Hsuan Yang 0001
ICCV4
2021 Deep Implicit Surface Point Prediction Networks
abstract
Deep neural representations of 3D shapes as implicit functions have been shown to produce high fidelity models surpassing the resolution-memory trade-off faced by the explicit representations using meshes and point clouds. However, most such approaches focus on representing closed shapes. Unsigned distance function (UDF) based approaches have been proposed recently as a promising alternative to represent both open and closed shapes. However, since the gradients of UDFs vanish on the surface, it is challenging to estimate local (differential) geometric properties like the normals and tangent planes which are needed for many downstream applications in vision and graphics. There are additional challenges in computing these properties efficiently with a low-memory footprint. This paper presents a novel approach that models such surfaces using a new class of implicit representations called the closest surface-point (CSP) representation. We show that CSP allows us to represent complex surfaces of any topology (open or closed) with high fidelity. It also allows for accurate and efficient computation of local geometric properties. We further demonstrate that it leads to efficient implementation of downstream algorithms like sphere-tracing for rendering the 3D surface as well as to create explicit mesh-based representations. Extensive experimental evaluation on the ShapeNet dataset validate the above contributions with results surpassing the state-of-the-art. Code and data are available at https://sites.google.com/view/cspnet.
Rahul Venkatesh, Tejan Karmali, Sarthak Sharma, Aurobrata Ghosh, Venkatesh Babu Radhakrishnan, László A. Jeni, Maneesh Kumar Singh 0001
ICCV7
2021 Perturb, Predict & Paraphrase: Semi-Supervised Learning using Noisy Student for Image Captioning
abstract
Recent semi-supervised learning (SSL) methods are predominantly focused on multi-class classification tasks. Classification tasks allow for easy mixing of class labels during augmentation which does not trivially extend to structured outputs such as word sequences that appear in tasks like image captioning. Noisy Student Training is a recent SSL paradigm proposed for image classification that is an extension of self-training and teacher-student learning. In this work, we provide an in-depth analysis of the noisy student SSL framework for the task of image captioning and derive state-of-the-art results. The original algorithm relies on computationally expensive data augmentation steps that involve perturbing the raw images and computing features for each perturbed image. We show that, even in the absence of raw image augmentation, the use of simple model and feature perturbations to the input images for the student model are beneficial to SSL training. We also show how a paraphrase generator could be effectively used for label augmentation to improve the quality of pseudo labels and significantly improve performance. Our final results in the limited labeled data setting (1% of the MS-COCO labeled data) outperform previous state-of-the-art approaches by 2.5 on BLEU4 and 11.5 on CIDEr scores.
Arjit Jain, Pranay Reddy Samala, Preethi Jyothi, Deepak Mittal, Maneesh Kumar Singh 0001
IJCAI5
2021 Investigating Feature Selection and Explainability for COVID-19 Diagnostics from Cough Sounds
abstract
In this paper, we propose an approach to automatically classify COVID-19 and non-COVID-19 cough samples based on the combination of both feature engineering and deep learning models. In the feature engineering approach, we develop a support vector machine classifier over high dimensional (6373D) space of acoustic features. In the deep learning-based approach, on the other hand, we apply a convolutional neural network trained on the log-mel spectrograms. These two methodologically diverse models are then combined by fusing the probability scores of the models. The proposed system, which ranked 9th on the 2021 Diagnosing COVID-19 using Acoustics (Di- COVA) challenge leaderboard, obtained an area under the receiver operating characteristic curve (AUC) of 0:81 on the blind test data set, which is a 10:9 absolute improvement compared to the baseline. Moreover, we analyze the explainability of the deep learning-based model when detecting COVID-19 from cough signals. Copyright © 2021 ISCA.
Flávio Ávila, Amir Hossein Poorjam, Deepak Mittal, Charles Dognin, Ananya Muguli, Srikanth Raj Chetupalli, Sriram Ganapathy, Maneesh Kumar Singh 0001
Interspeech9
2020 Accelerating Column Generation via Flexible Dual Optimal Inequalities with Application to Entity Resolution
abstract
In this paper, we introduce a new optimization approach to Entity Resolution. Traditional approaches tackle entity resolution with hierarchical clustering, which does not benefit from a formal optimization formulation. In contrast, we model entity resolution as correlation-clustering, which we treat as a weighted set-packing problem and write as an integer linear program (ILP). In this case, sources in the input data correspond to elements and entities in output data correspond to sets/clusters. We tackle optimization of weighted set packing by relaxing integrality in our ILP formulation. The set of potential sets/clusters can not be explicitly enumerated, thus motivating optimization via column generation. In addition to the novel formulation, we also introduce new dual optimal inequalities (DOI), that we call flexible dual optimal inequalities, which tightly lower-bound dual variables during optimization and accelerate column generation. We apply our formulation to entity resolution (also called de-duplication of records), and achieve state-of-the-art accuracy on two popular benchmark datasets. Our F-DOI can be extended to other weighted set-packing problems.
Vishnu Suresh Lokhande, Maneesh Kumar Singh 0001, Julian Yarkony
AAAI3
2020 ProAlignNet: Unsupervised Learning for Progressively Aligning Noisy Contours
V. S. R. Veeravasarapu, Abhishek Goel, Deepak Mittal, Maneesh Kumar Singh 0001
CVPR4
2020 Infoprint: Information Theoretic Digital Image Forensics
abstract
Tampered images pose a serious predicament since digitized media is a ubiquitous part of our lives. These are facilitated by the availability of image editing software and recent advances in deep Generative Adversarial Networks (GANs). We propose an innovative method to formulate the problem of 10-calizing manipulated regions in fake images as a deep representation learning problem using the Information Bottleneck (IB) principle. We devise a convolutional neural net-based architecture, InfoPrint (IP), that uses variational inference to approximate the IB formulation. Testing on three standard datasets, we demonstrate that InfoPrint outperforms the state-of-the-art by 3% points or more. Additionally, we demonstrate that it has the ability to to detect alterations made by inpainting GANs.
Aurobrata Ghosh, Steve Cruz, Subbu Veeravasarapu, Maneesh Kumar Singh 0001, Terrance E. Boult
ICIP5
2020 Multiscale System for Alzheimer's Dementia Recognition Through Spontaneous Speech
Erik Edwards, Charles Dognin, Bajibabu Bollepalli, Maneesh Kumar Singh 0001
INTERSPEECH4
2020 Progressive Domain Adaptation for Object Detection
abstract
Recent deep learning methods for object detection rely on a large amount of bounding box annotations. Collecting these annotations is laborious and costly, yet supervised models do not generalize well when testing on images from a different distribution. Domain adaptation provides a solution by adapting existing labels to the target testing data. However, a large gap between domains could make adaptation a challenging task, which leads to unstable training processes and sub-optimal results. In this paper, we propose to bridge the domain gap with an intermediate domain and progressively solve easier adaptation subtasks. This intermediate domain is constructed by translating the source images to mimic the ones in the target domain. To tackle the domain-shift problem, we adopt adversarial learning to align distributions at the feature level. In addition, a weighted task loss is applied to deal with unbalanced image quality in the intermediate domain. Experimental results show that our method performs favorably against the state-of-the-art method in terms of the performance on the target domain.
Han-Kai Hsu, Chun-Han Yao, Yi-Hsuan Tsai, Wei-Chih Hung, Hung-Yu Tseng, Maneesh Kumar Singh 0001, Ming-Hsuan Yang 0001
WACV6
2020 DRIT++: Diverse Image-to-Image Translation via Disentangled Representations
Hsin-Ying Lee 0001, Hung-Yu Tseng, Qi Mao 0002, Jia-Bin Huang 0001, Yu-Ding Lu, Maneesh Kumar Singh 0001, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.6
2020 On Lipschitz Bounds of General Convolutional Neural Networks
abstract
Many convolutional neural networks (CNN's) have a feed-forward structure. In this paper, we model a general framework for analyzing the Lipschitz bounds of CNN's and propose a linear program that estimates these bounds. Several CNN's, including the scattering networks, the AlexNet and the GoogleNet, are studied numerically. In these practical numerical examples, estimations of local Lipschitz bounds are compared to these theoretical bounds. Based on the Lipschitz bounds, we next establish concentration inequalities for the output distribution with respect to a stationary random input signal. The Lipschitz bound is further used to perform nonlinear discriminant analysis that measures the separation between features of different classes.
Dongmian Zou, Radu V. Balan, Maneesh Kumar Singh 0001
IEEE Trans. Inf. Theory3
2019 Sampling Bias in Deep Active Classification: An Empirical Study
abstract
Ameya Prabhu, Charles Dognin, Maneesh Singh. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Ameya Prabhu, Charles Dognin, Maneesh Kumar Singh 0001
EMNLP/IJCNLP (1)3
2019 Compositional Video Prediction
abstract
We present an approach for pixel-level future prediction given an input image of a scene. We observe that a scene is comprised of distinct entities that undergo motion and present an approach that operationalizes this insight. We implicitly predict future states of independent entities while reasoning about their interactions, and compose future video frames using these predicted states. We overcome the inherent multi-modality of the task using a global trajectory-level latent random variable, and show that this allows us to sample diverse and plausible futures. We empirically validate our approach against alternate representations and ways of incorporating multi-modality. We examine two datasets, one comprising of stacked objects that may fall, and the other containing videos of humans performing activities in a gym, and show that our approach allows realistic stochastic video prediction across these diverse settings. See project website (https://judyye.github.io/CVP/) for video predictions.
Yufei Ye 0001, Maneesh Kumar Singh 0001, Abhinav Gupta 0001, Shubham Tulsiani
ICCV2
2018 Disentangling Factors of Variation with Cycle-Consistent Variational Auto-encoders
Ananya Harsh Jha, Saket Anand, Maneesh Kumar Singh 0001, V. S. R. Veeravasarapu
ECCV (3)3
2018 Diverse Image-to-Image Translation via Disentangled Representations
Hsin-Ying Lee 0001, Hung-Yu Tseng, Jia-Bin Huang 0001, Maneesh Kumar Singh 0001, Ming-Hsuan Yang 0001
ECCV (1)4
2018 Adversarial Learning of Raw Speech Features for Domain Invariant Speech Recognition
abstract
Recent advances in neural network based acoustic modelling have shown significant improvements in automatic speech recognition (ASR) performance. In order for acoustic models to be able to handle large acoustic variability, large amounts of labeled data is necessary, which are often expensive to obtain. This paper explores the application of adversarial training to learn features from raw speech that are invariant to acoustic variability. This acoustic variability is referred to as a domain shift in this paper. The experimental study presented in this paper leverages the architecture of Domain Adversarial Neural Networks (DANNs) [1] which uses data from two different domains. The DANN is a Y-shaped network that consists of a multi-layer CNN feature extractor module that is common to a label (senone) classifier and a so-called domain classifier. The utility of DANNs is evaluated on multiple datasets with domain shifts caused due to differences in gender and speaker accents. Promising empirical results indicate the strength of adversarial training for unsupervised domain adaptation in ASR, thereby emphasizing the ability of DANNs to learn domain invariant features from raw speech.
Aditay Tripathi, Aanchan Mohan, Saket Anand, Maneesh Kumar Singh 0001
ICASSP4
2018 Disconnected Manifold Learning for Generative Adversarial Networks
abstract
Natural images may lie on a union of disjoint manifolds rather than one globally connected manifold, and this can cause several difficulties for the training of common Generative Adversarial Networks (GANs). In this work, we first show that single generator GANs are unable to correctly model a distribution supported on a disconnected manifold, and investigate how sample quality, mode dropping and local convergence are affected by this. Next, we show how using a collection of generators can address this problem, providing new insights into the success of such multi-generator GANs. Finally, we explain the serious issues caused by considering a fixed prior over the collection of generators and propose a novel approach for learning the prior and inferring the necessary number of generators without any supervision. Our proposed modifications can be applied on top of any other GAN model to enable learning of distributions supported on disconnected manifolds. We conduct several experiments to illustrate the aforementioned shortcoming of GANs, its consequences in practice, and the effectiveness of our proposed modifications in alleviating these issues.
Mahyar Khayatkhoei, Maneesh Kumar Singh 0001, Ahmed M. Elgammal
NeurIPS2
2017 Unsupervised Representation Learning by Sorting Sequences
abstract
We present an unsupervised representation learning approach using videos without semantic labels. We leverage the temporal coherence as a supervisory signal by formulating representation learning as a sequence sorting task. We take temporally shuffled frames (i.e., in non-chronological order) as inputs and train a convolutional neural network to sort the shuffled sequences. Similar to comparison-based sorting algorithms, we propose to extract features from all frame pairs and aggregate them to predict the correct order. As sorting shuffled image sequence requires an understanding of the statistical temporal structure of images, training with such a proxy task allows us to learn rich and generalizable visual representation. We validate the effectiveness of the learned representation using our method as pre-training on high-level recognition problems. The experimental results show that our method compares favorably against state-of-the-art methods on action recognition, image classification, and object detection tasks.
Hsin-Ying Lee 0001, Jia-Bin Huang 0001, Maneesh Kumar Singh 0001, Ming-Hsuan Yang 0001
ICCV3
2012 Discrete texture traces: Topological representation of geometric context
abstract
Modeling representations of image patches that are quasi-invariant to spatial deformations is an important problem in computer vision. In this paper, we propose a novel concept, the texture trace, that allows sparse patch representations which are quasi-invariant to smooth deformations and robust against occlusions. We first propose a continuous domain model, the profile trace, which is a function only of the topological properties of an image and is by construction invariant to any homeomorphic transformation of the domain. We analyze its theoretical properties and then derive a discrete-domain approximation, the Discrete Texture Trace (DTT). DTTs are designed to be computationally practical and shown by a set of controlled experiments to be quasi-invariant to smooth spatial deformations as well as common image perturbations. We then show how DTTs can be naturally adapted to the incremental tracking problem, yielding highly precise results on par with the state of the art on challenging real data without using heavy machine learning tools. Indeed, we show that with even just using one image at the start of a sequence (i.e. no incremental updating), our method already outperforms four of six state of the art methods of the recent literature on challenging sequences.
Jan Ernst, Maneesh Kumar Singh 0001, Visvanathan Ramesh
CVPR2
2012 Fast radial symmetry detection under affine transformations
abstract
The fast radial symmetry (FRS) transform has been very popular for detecting interest points based on local radial symmetry1. Although FRS delivers good performance at a relatively low computational cost and is very well suited for a variety of real-time computer vision applications, it is not invariant to perspective distortions. Moreover, even perfectly (radially) symmetric visual patterns in the real world are perceived by us after a perspective projection. In this paper, we propose a systematic extension to the FRS transform to make it invariant to (bounded) cases of perspective projection - we call this transform the generalized FRS or GFRS transform. We show that GFRS inherits the basic characteristics of FRS and retains its computational efficiency. We demonstrate the wide applicability of GFRS by applying it to a variety of natural images to detect radially symmetric patterns that have undergone significant perspective distortions. Subsequently, we build a nucleus detector based on the GFRS transform and apply it to the important problem of digital histopathology. We demonstrate superior performance over state-of-the-art nuclei detection algorithms, validated using ROC curves.
Jie Ni, Maneesh Kumar Singh 0001, Claus Bahlmann
CVPR2
2011 Predicate Logic Based Image Grammars for Complex Pattern Recognition
Vinay D. Shet, Maneesh Kumar Singh 0001, Claus Bahlmann, Visvanathan Ramesh, Jan Neumann, Larry Davis 0001
Int. J. Comput. Vis.2
2010 Illumination compensation based change detection using order consistency
abstract
We present a change detection method resistant to global and local illumination variations for use in visual surveillance scenarios. Approaches designed thus far for robustness to illumination change are generally based either on color normalization, texture (e.g. edges, rank order statistics, etc.), or illumination compensation. Normalization based methods sacrifice discriminability while texture based methods cannot operate on texture-less regions. Both types of method can produce large missing regions in the distance image which in turn pose problems for higher-level processing tasks that may be shape or region-based and require accurate foreground masks (e.g. person detection and tracking, crowd segmentation, etc.). Texture based methods have an additional problem in that they produce false alarms due to textures induced by local illumination effects (e.g. cast shadows). In this paper we propose a compensation based approach for change detection. Prior work on compensation has largely taken an empirical approach, and has not dealt with the important problem of rejecting outliers when they dominate the scene. In contrast, our generative approach and systematic handling of outliers enables us to achieve robustness to illumination change while eliminating the problems mentioned above. Furthermore, the computational complexity of our method is low enough for real-time performance. Results comparing images taken under strongly different illumination conditions, demonstrate the power and generality of the proposed method.
Vasu Parameswaran, Maneesh Kumar Singh 0001, Visvanathan Ramesh
CVPR2
2010 In a Blink of an Eye and a Switch of a Transistor: Cortically Coupled Computer Vision
abstract
Our society's information technology advancements have resulted in the increasingly problematic issue of information overload—i.e., we have more access to information than we can possibly process. This is nowhere more apparent than in the volume of imagery and video that we can access on a daily basis—for the general public, availability of YouTube video and Google Images, or for the image analysis professional tasked with searching security video or satellite reconnaissance. Which images to look at and how to ensure we see the images that are of most interest to us, begs the question of whether there are smart ways to triage this volume of imagery. Over the past decade, computer vision research has focused on the issue of ranking and indexing imagery. However, computer vision is limited in its ability to identify interesting imagery, particularly as “interesting” might be defined by an individual. In this paper we describe our efforts in developing brain–computer interfaces (BCIs) which synergistically integrate computer vision and human vision so as to construct a system for image triage. Our approach exploits machine learning for real-time decoding of brain signals which are recorded noninvasively via electroencephalography (EEG). The signals we decode are specific for events related to imagery attracting a user's attention. We describe two architectures we have developed for this type of cortically coupled computer vision and discuss potential applications and challenges for the future.
Paul Sajda, Eric Pohlmeyer, Jun Wang 0006, Lucas C. Parra, Christoforos Christoforou, Jacek Dmochowski, Barbara Hanna, Claus Bahlmann, Maneesh Kumar Singh 0001, Shih-Fu Chang
Proc. IEEE9
2009 State-of-the-art on spatio-temporal information-based video retrieval
Sameer Singh 0002, Maneesh Kumar Singh 0001, Y. S. Zhu
Pattern Recognit.3
2008 Order consistent change detection via fast statistical significance testing
abstract
Robustness to illumination variations is a key requirement for the problem of change detection which in turn is a fundamental building block for many visual surveillance applications. The use of ordinal measures is a powerful way of filtering out illumination dependency in representing appearance, and several such measures have been proposed in the past for change detection. By design, these measures are invariant to unknown monotonic transformations that may be caused due to global illumination changes or automatic camera gain. However, previous work has left theoretical and practical gaps that limit their full potential from being realized. For instance, random noise has not been given a principled treatment. In this paper, we formulate the change detection problem in terms of order consistency and show that in the presence of noise with known statistical properties, significance tests for order consistency yield much better results than the state of the art. Since ordinal measures require a reordering of patches, they are usually expensive in practice (O(n*log n) at best). We improve upon this by connecting the problem to monotonic regression, and applying a fast algorithm from the corresponding literature. We also show that good trade offs between speed and accuracy can be made by quantization to achieve accurate and very fast matching algorithms in practice. We demonstrate superior performance on statistical simulations as well as real image sequences.
Maneesh Kumar Singh 0001, Vasu Parameswaran, Visvanathan Ramesh
CVPR1
2006 Prior-Constrained Scale-Space Mean Shift
abstract
This paper proposes a new variational bound optimization framework for incorporating spatial prior information to the mean shift-based data-driven mode analysis, offering flexible control of the mean shift convergence. Two forms of Gaussian spatial priors are considered. Attractive prior pulls the convergence toward a desired location. Repulsive prior pushes away from such a location. Using a generic variational optimization formulation via construction of quadratic lower and upper bounds, we show that the priorconstrained mean shift step can be interpreted as an information fusion of the data and prior terms in the sense of the best linear unbiased estimator. This approach is used to propose a mode parsing algorithm using the inhibitionof-return principle. The proposed algorithm is used for a semi-automatic 3D segmentation of lung nodules in CT data for evaluating its effectiveness. Our experiments demonstrate that the proposed solution can successfully segment challenging wall-attached cases. 1
Kazunori Okada, Maneesh Kumar Singh 0001, Visvanathan Ramesh
BMVC2
2005 Robust Pulmonary Nodule Segmentation in CT: Improving Performance for Juxtapleural Cases
Kazunori Okada, Visvanathan Ramesh, Arun Krishnan, Maneesh Kumar Singh 0001, Umut Akdemir
MICCAI (2)4
2004 A Robust Probabilistic Estimation Framework for Parametric Image Models
Maneesh Kumar Singh 0001, Himanshu Arora, Narendra Ahuja
ECCV (1)1
2003 Regression based Bandwidth Selection for Segmentation using Parzen Windows
abstract
We consider the problem of segmentation of images that can be modelled as piecewise continuous signals having unknown, nonstationary statistics. We propose a solution to this problem which first uses a regression framework to estimate the image PDF, and then mean-shift to find the modes of this PDF. The segmentation follows from mode identification wherein pixel clusters or image segments are identified with unique modes of the multimodal PDF. Each pixel is mapped to a mode using a convergent, iterative process. The effectiveness of the approach depends upon the accuracy of the (implicit) estimate of the underlying multimodal density function and thus on the bandwidth parameters used for its estimate using Parzen windows. Automatic selection of bandwidth parameters is a desired feature of the algorithm. We show that the proposed regression-based model admits a realistic framework to automatically choose bandwidth parameters which minimizes a global error criterion. We validate the theory presented with results on real images.
Maneesh Kumar Singh 0001, Narendra Ahuja
ICCV1
2002 Mean-Shift Segmentation with Wavelet-based Bandwidth Selection
abstract
Recently, various non-linear techniques for segmentation have been proposed based on non-parametric density estimation. These approaches model image data as clusters of pixels in the combined range-domain space, using kernel based techniques to represent the underlying, multi-modal Probability Density Function (PDF). In Mean-shift based segmentation, pixel clusters or image segments are identified with unique modes of the multi-modal PDF by mapping each pixel to a mode using a convergent, iterative process. The advantages of such approaches include flexible modeling of the image and noise processes and consequent robustness in segmentation. An important issue is the automatic selection of scale parameters a problem far from satisfactorily addressed. In this paper, we propose a regression-based model which admits a realistic framework to choose scale parameters. Results on real images are presented.
Maneesh Kumar Singh 0001, Narendra Ahuja
WACV1
1999 Segmentation Based Denoising Using Multiple Compaction Domains
abstract
In this paper, we propose a novel segmentation based denoising algorithm. Segmentation yields intrinsically homogeneous and extrinsically heterogeneous regions. A denoising algorithm that uses Multiple Compaction Domains (MCD) is then applied on each of the resulting segments. Such a scheme retains important perceptual information in the segment boundaries while the denoising algorithm operates only on homogeneous segments. Further, the MCD algorithm is demonstrably superior to the classical denoising algorithms using transform domain thresholding. Our algorithm yields better perceptual quality and superior PSNR as compared to MATLAB's adaptive Wiener filter.
Maneesh Kumar Singh 0001, Prakash Ishwar, Krishna Ratakonda, Narendra Ahuja
ICIP (1)1