Pradeep Natarajan

dblp:95/5978 · DBLP profile ↗
← Back
39ranked-venue papers
9as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 24 · 8 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SAIGE-GPU: accelerating genome- and phenome-wide association studies using GPUs
abstract
MOTIVATION: Genome-wide association studies (GWAS) at biobank scale are computationally intensive, especially for admixed populations requiring robust statistical models. SAIGE is a widely used method for generalized linear mixed-model GWAS but is limited by its CPU-based implementation, making phenome-wide association studies impractical for many research groups. RESULTS: We developed SAIGE-GPU, a GPU-accelerated version of SAIGE that replaces CPU-intensive matrix operations with GPU-optimized kernels. The core innovation is distributing genetic relationship matrix calculations across GPUs and communication layers. Applied to 2068 phenotypes from 635 969 participants in the Million Veteran Program, including diverse and admixed populations, SAIGE-GPU achieved a 5-fold speedup in mixed model fitting on supercomputing infrastructure and cloud platforms. We further optimized the variant association testing step through multi-core and multi-trait parallelization. Deployed on Google Cloud Platform and Azure, the method provided substantial cost and time savings. AVAILABILITY AND IMPLEMENTATION: Source code and binaries are available for download at https://github.com/saigegit/SAIGE/tree/SAIGE-GPU-1.3.3. A code snapshot is archived at Zenodo for reproducibility (DOI: [10.5281/zenodo.17642591]). SAIGE-GPU is available in a containerized format for use across HPC and cloud environments and is implemented in R/C++ and runs on Linux systems.
Alex Rodriguez, Youngdae Kim, Tarak Nath Nandi, Karl Keat, Rachit Kumar, Mitchell Conery, Rohan Bhukar, Molei Liu, John Hessington, Ketan Maheshwari, VA Million Veteran Program, Edmon Begoli, Georgia Tourassi, Pradeep Natarajan, Benjamin F. Voight, John Michael Gaziano, Scott M. Damrauer, Katherine P. Liao, Jennifer E. Huffman, Anurag Verma, Ravi K. Madduri
Bioinform.14
2025 Robust pleiotropy-decomposed polygenic scores identify distinct contributions to elevated coronary artery disease polygenic risk
abstract
BACKGROUND: Polygenic risk score (PRS) have proved to offer robust risk prediction for coronary artery disease (CAD). However, the global CAD PRS summarizes the joint effects of all the markers in the genome, masking potential genetic heterogeneity that may be important for disease interpretation and targeted interventions. METHODS: Using summary-level data, we identified 43 significant CAD-related traits based on genetic correlations, and further classified them into eight pleiotropy clusters based on their biological functions. We then partitioned the genome into 2,353 near-independent regions. Variants in each region were assigned to the trait most genetically similar to CAD, and then were labeled with the corresponding pleiotropy cluster. We grouped variants without labels into a ninth, non-specific cluster. The Pleiotropy Decomposed (PD) PRSs for each of the nine clusters were calculated using variants assigned to each cluster for 407,903 samples of European ancestry from the UK Biobank (UKBB). RESULTS: We decomposed the CAD PRS into nine PD-PRSs and further stratified individuals with high CAD-PRS into nine subgroups. Each PD-PRS accounted for a higher proportion of the global CAD-PRS within its corresponding subgroup than in the remaining subjects with high CAD-PRS (e.g., 25.2% (0.07) vs. 10.06% (0.07) for lipids-PD-PRS). Additionally, these subgroups showed distinct clinical features. For example, in the lipids-related subgroup, lipoprotein(a) and LDL-cholesterol levels were 67.5% and 18.3% higher, respectively, compared to the remaining high-risk individuals. Furthermore, significant interactions were observed between blood pressure and BP PD-PRS, and between current smoking and respiratory system PD-PRS. CONCLUSION: Our findings suggest that PD-PRSs may reveal substantial genetic and phenotypic heterogeneity among individuals with high CAD-PRS. The unique PD-PRS compositions of each individual can highlight the relative importance of different pleiotropic regions.
Jiaqi Hu 0004, Yixuan Ye, Yunfeng Ruan, Pradeep Natarajan, Hongyu Zhao 0003
PLoS Comput. Biol.5
2024 HVCLIP: High-Dimensional Vector in CLIP for Unsupervised Domain Adaptation
Noranart Vesdapunt, Kah Kuen Fu, Pradeep Natarajan
ECCV (66)5
2023 User-Controllable Arbitrary Style Transfer via Entropy Regularization
abstract
Ensuring the overall end-user experience is a challenging task in arbitrary style transfer (AST) due to the subjective nature of style transfer quality. A good practice is to provide users many instead of one AST result. However, existing approaches require to run multiple AST models or inference a diversified AST (DAST) solution multiple times, and thus they are either slow in speed or limited in diversity. In this paper, we propose a novel solution ensuring both efficiency and diversity for generating multiple user-controllable AST results by systematically modulating AST behavior at run-time. We begin with reformulating three prominent AST methods into a unified assign-and-mix problem and discover that the entropies of their assignment matrices exhibit a large variance. We then solve the unified problem in an optimal transport framework using the Sinkhorn-Knopp algorithm with a user input ε to control the said entropy and thus modulate stylization. Empirical results demonstrate the superiority of the proposed solution, with speed and stylization quality comparable to or better than existing AST and significantly more diverse than previous DAST works. Code is available at https://github.com/cplusx/eps-Assign-and-Mix.
Jiaxin Cheng, Yue Wu 0001, Ayush Jaiswal, Xu Zhang 0022, Pradeep Natarajan, Premkumar Natarajan
AAAI5
2023 FashionNTM: Multi-turn Fashion Image Retrieval via Cascaded Memory
abstract
Multi-turn textual feedback-based fashion image retrieval focuses on a real-world setting, where users can iteratively provide information to refine retrieval results until they find an item that fits all their requirements. In this work, we present a novel memory-based method, called FashionNTM, for such a multi-turn system. Our framework incorporates a new Cascaded Memory Neural Turing Machine (CM-NTM) approach for implicit state management, thereby learning to integrate information across all past turns to retrieve new images, for a given turn. Unlike vanilla Neural Turing Machine (NTM), our CM-NTM operates on multiple inputs, which interact with their respective memories via individual read and write heads, to learn complex relationships. Extensive evaluation results show that our proposed method outperforms the previous state-of-the-art algorithm by 50.5%, on Multi-turn FashionIQ [60] – the only existing multi-turn fashion dataset currently, in addition to having a relative improvement of 12.6% on Multi-turn Shoes – an extension of the singleturn Shoes dataset [5] that we created in this work. Further analysis of the model in a real-world interactive setting demonstrates two important capabilities of our model – memory retention across turns, and agnosticity to turn order for non-contradictory feedback. Finally, user study results show that images retrieved by FashionNTM were favored by 83.1% over other multi-turn models.
Anwesan Pal, Sahil Wadhwa, Ayush Jaiswal, Rakesh Chada, Pradeep Natarajan, Henrik I. Christensen
ICCV7
2023 Performance and Failure Cause Estimation for Machine Learning Systems in the Wild
Xiruo Liu, Furqan Khan, Yue Niu 0001, Pradeep Natarajan, Rinat Khaziev, Amir Salman Avestimehr, Prateek Singhal
ICVS4
2022 Enhancing Fairness in Face Detection in Computer Vision Systems by Demographic Bias Mitigation
abstract
Fairness has become an important agenda in computer vision and artificial intelligence. Recent studies have shown that many computer vision models and datasets exhibit demographic biases and proposed mitigation strategies. These works attempt to address accuracy disparity, spurious correlations, or unbalanced representations in datasets in tasks such as face recognition, verification and expression and attribute classification. These tasks, however, all require face detection as the first preprocessing step, and surprisingly, there has been little effort in identifying or mitigating biases in face detection. Biased face detectors themselves pose a threat against fair and ethical AI systems, and their biases may be further passed on to subsequent downstream tasks such as face recognition in a computer vision pipeline. This paper therefore investigates the problem of biases in face detection, focusing on accuracy disparity of detectors between demographic groups including gender, age group, and skin tone. We collect perceived demographic attributes on a popular face detection benchmark dataset, WIDER FACE, report skewed demographic distributions, and compare detection performance between groups. In order to mitigate the biases, we apply three mitigation methods that have been introduced in the recent literature and also propose two novel methods. Experimental results show that these methods are effective in reducing demographic biases. We also discuss how the effectiveness varies by demographic attributes, detection easiness, and multiple detectors, which will shed light on this new topic of addressing face detection bias.
Jianwei Feng, Prateek Singhal, Vivek Yadav, Yue Wu 0001, Pradeep Natarajan, Varsha Hedau, Jungseock Joo
AIES7
2022 FashionVLP: Vision Language Transformer for Fashion Retrieval with Feedback
abstract
Fashion image retrieval based on a query pair of reference image and natural language feedback is a challenging task that requires models to assess fashion related information from visual and textual modalities simultaneously. We propose a new vision-language transformer based model, FashionVLP, that brings the prior knowledge contained in large image-text corpora to the domain of fashion image retrieval, and combines visual information from multiple levels of context to effectively capture fashion-related information. While queries are encoded through the transformer layers, our asymmetric design adopts a novel attention-based approach for fusing target image features without involving text or transformer layers in the process. Extensive results show that FashionVLP achieves the state-of-the-art performance on benchmark datasets, with a large 23% relative improvement on the challenging FashionIQ dataset, which contains complex natural language feedback.
Sonam Goenka, Zhaoheng Zheng, Ayush Jaiswal, Rakesh Chada, Yue Wu 0001, Varsha Hedau, Pradeep Natarajan
CVPR7
2022 Asd-Transformer: Efficient Active Speaker Detection Using Self And Multimodal Transformers
abstract
Multimodal active speaker detection (ASD) methods assign a speaking/not-speaking label per individual in a video clip. ASD is critical for applications such as natural human-computer interaction, speaker diarization, and video reframing. Recent work has shown the success of transformers in multimodal settings, thus we propose a novel framework that leverages modern transformer and concatenation mechanisms to efficiently capture the interaction between audio and video modalities for ASD. We achieve mAP similar to state-of-the-art (93.0% vs 93.5%) on the AVA-ActiveSpeaker dataset. Further, our model has ~3× smaller size (15.23MB vs 49.82MB), reduced FLOPs count (11.8 vs 14.3), and lower training time (15h vs 38h). To verify our model is making predictions from the right visual cues, we computed saliency maps over input images. We found that in addition to mouth regions, the nose, cheek, and area under the eye were helpful in identifying active speakers. Our ablation study reveals that the mouth region alone achieved lower mAP (91.9% vs 93.0%) compared to full face region, supporting our hypothesis that facial expressions in addition to mouth region are useful for ASD.
Gourav Datta, Tyler Etchart, Vivek Yadav, Varsha Hedau, Pradeep Natarajan, Shih-Fu Chang
ICASSP5
2022 Alexa Teacher Model: Pretraining and Distilling Multi-Billion-Parameter Encoders for Natural Language Understanding Systems
abstract
We present results from a large-scale experiment on pretraining encoders with non-embedding parameter counts ranging from 700M to 9.3B, their subsequent distillation into smaller models ranging from 17M-170M parameters, and their application to the Natural Language Understanding (NLU) component of a virtual assistant system. Though we train using 70% spoken-form data, our teacher models perform comparably to XLM-R and mT5 when evaluated on the written-form Cross-lingual Natural Language Inference (XNLI) corpus. We perform a second stage of pretraining on our teacher models using in-domain data from our system, improving error rates by 3.86% relative for intent classification and 7.01% relative for slot filling. We find that even a 170M-parameter model distilled from our Stage 2 teacher model has 2.88% better intent classification and 7.69% better slot filling error rates when compared to the 2.3B-parameter teacher trained only on public data (Stage 1), emphasizing the importance of in-domain data for pretraining. When evaluated offline using labeled NLU data, our 17M-parameter Stage 2 distilled model outperforms both XLM-R Base (85M params) and DistillBERT (42M params) by 4.23% to 6.14%, respectively. Finally, we present results from a full virtual assistant experimentation platform, where we find that models trained using our pretraining and distillation pipeline outperform models distilled from 85M-parameter teachers by 3.74%-4.91% on an automatic measurement of full-system user dissatisfaction.
Jack FitzGerald, Shankar Ananthakrishnan, Konstantine Arkoudas, Davide Bernardi, Abhishek Bhagia, Claudio Delli Bovi, Jin Cao 0003, Rakesh Chada, Amit Chauhan, Luoxin Chen, Anurag Dwarakanath, Satyam Dwivedi, Turan Gojayev, Karthik Gopalakrishnan 0001, Thomas Gueudré, Dilek Hakkani-Tür, Wael Hamza, Jonathan J. Hüser, Kevin Martin Jose, Haidar Khan, Beiye Liu, Jianhua Lu, Alessandro Manzotti, Pradeep Natarajan, Karolina Owczarzak, Gokmen Oz, Enrico Palumbo, Charith Peris, Chandana Satya Prakash, Stephen Rawls, Andy Rosenbaum, Anjali Shenoy, Saleh Soltan, Mukund Sridhar, Lizhen Tan, Fabian Triefenbach, Pan Wei, Shuai Zheng 0004, Gökhan Tür, Premkumar Natarajan
KDD24
2022 Model Monitoring in Practice: Lessons Learned and Open Challenges
abstract
Artificial Intelligence (AI) is increasingly playing an integral role in determining our day-to-day experiences. Increasingly, the applications of AI are no longer limited to search and recommendation systems, such as web search and movie and product recommendations, but AI is also being used in decisions and processes that are critical for individuals, businesses, and society. With AI based solutions in high-stakes domains such as hiring, lending, criminal justice, healthcare, and education, the resulting personal and professional implications of AI are far-reaching. Consequently, it becomes critical to ensure that these models are making accurate predictions, are robust to shifts in the data, are not relying on spurious features, and are not unduly discriminating against minority groups. To this end, several approaches spanning various areas such as explainability, fairness, and robustness have been proposed in recent literature, and many papers and tutorials on these topics have been presented in recent computer science conferences. However, there is relatively less attention on the need for monitoring machine learning (ML) models once they are deployed and the associated research challenges.
Krishnaram Kenthapadi, Himabindu Lakkaraju, Pradeep Natarajan, Mehrnoosh Sameki
KDD3
2021 Style-Aware Normalized Loss for Improving Arbitrary Style Transfer
abstract
Neural Style Transfer (NST) has quickly evolved from single-style to infinite-style models, also known as Arbitrary Style Transfer (AST). Although appealing results have been widely reported in literature, our empirical studies on four well-known AST approaches (GoogleMagenta [14], AdaIN [19], LinearTransfer [29], and SANet [37]) show that more than 50% of the time, AST stylized images are not acceptable to human users, typically due to under- or over-stylization. We systematically study the cause of this imbalanced style transferability (IST ) and propose a simple yet effective solution to mitigate this issue. Our studies show that the IST issue is related to the conventional AST style loss, and reveal that the root cause is the equal weightage of training samples irrespective of the properties of their corresponding style images, which biases the model towards certain styles. Through investigation of the theoretical bounds of the AST style loss, we propose a new loss that largely overcomes IST . Theoretical analysis and experimental results validate the effectiveness of our loss, with over 80% relative improvement in style deception rate and 98% relatively higher preference in human evaluation.
Jiaxin Cheng, Ayush Jaiswal, Yue Wu 0001, Pradeep Natarajan, Premkumar Natarajan
CVPR4
2021 FewshotQA: A simple framework for few-shot learning of question answering tasks using pre-trained text-to-text models
abstract
The task of learning from only a few examples (called a few-shot setting) is of key importance and relevance to a real-world setting.For question answering (QA), the current state-of-the-art pre-trained models typically need fine-tuning on tens of thousands of examples to obtain good results.Their performance degrades significantly in a few-shot setting (< 100 examples).To address this, we propose a simple fine-tuning framework that leverages pre-trained text-to-text models and is directly aligned with their pre-training framework.Specifically, we construct the input as a concatenation of the question, a mask token representing the answer span and a context.Given this input, the model is fine-tuned using the same objective as that of its pre-training objective.Through experimental studies on various few-shot configurations, we show that this formulation leads to significant gains on multiple QA benchmarks (an absolute gain of 34.2 F1 points on average when there are only 16 training examples).The gains extend further when used with larger models (Eg:-72.3F1 on SQuAD using BART-large with only 32 examples) and translate well to a multilingual setting .On the multilingual TydiQA benchmark, our model outperforms the XLM-Roberta-large by an absolute margin of upto 40 F1 points and an average of 33 F1 points in a few-shot setting (<= 64 training examples).We conduct detailed ablation studies to analyze factors contributing to these gains.
Rakesh Chada, Pradeep Natarajan
EMNLP (1)2
2021 Adversarial Mask Generation for Preserving Visual Privacy
abstract
We present a privacy preserving machine learning method for images that separates task-relevant information from task-irrelevant information. Our primary hypothesis is that by revealing the minimal number of pixels required for a task we can provide the most privacy preserving guarantees. Specifically, we propose an adversarial method that masks out task-irrelevant information from an image for preserving privacy. The proposed method only uses task-specific label information and no privacy annotations such as identity of the subject, gender, race, etc., are required. We validate the proposed method on face attribute prediction on the CelebA dataset and emotion recognition on the FER+ dataset, showing that we can preserve visual privacy with little degradation in the task performance.
Ayush Jaiswal, Yue Wu 0001, Vivek Yadav, Pradeep Natarajan
FG5
2021 Class-agnostic Object Detection
abstract
Object detection models perform well at localizing and classifying objects that they are shown during training. However, due to the difficulty and cost associated with creating and annotating detection datasets, trained models detect a limited number of object types with unknown objects treated as background content. This hinders the adoption of conventional detectors in real-world applications like large-scale object matching, visual grounding, visual relation prediction, obstacle detection (where it is more important to determine the presence and location of objects than to find specific types), etc. We propose class-agnostic object detection as a new problem that focuses on detecting objects irrespective of their object-classes. Specifically, the goal is to predict bounding boxes for all objects in an image but not their object-classes. The predicted boxes can then be consumed by another system to perform application-specific classification, retrieval, etc. We propose training and eval uation protocols for benchmarking class-agnostic detectors to advance future research in this domain. Finally, we propose (1) baseline methods and (2) a new adversarial learning framework for class-agnostic detection that forces the model to exclude class-specific information from features used for predictions. Experimental results show that adversarial learning improves class-agnostic detection efficacy.
Ayush Jaiswal, Yue Wu 0001, Pradeep Natarajan, Premkumar Natarajan
WACV3
2014 Zero-Shot Event Detection Using Multi-modal Fusion of Weakly Supervised Concepts
abstract
Current state-of-the-art systems for visual content analysis require large training sets for each class of interest, and performance degrades rapidly with fewer examples. In this paper, we present a general framework for the zeroshot learning problem of performing high-level event detection with no training exemplars, using only textual descriptions. This task goes beyond the traditional zero-shot framework of adapting a given set of classes with training data to unseen classes. We leverage video and image collections with free-form text descriptions from widely available web sources to learn a large bank of concepts, in addition to using several off-the-shelf concept detectors, speech, and video text for representing videos. We utilize natural language processing technologies to generate event description features. The extracted features are then projected to a common high-dimensional space using text expansion, and similarity is computed in this space. We present extensive experimental results on the large TRECVID MED [26] corpus to demonstrate our approach. Our results show that the proposed concept detection methods significantly outperform current attribute classifiers such as Classemes [34], ObjectBank [21], and SUN attributes[28] . Further, we find that fusion, both within as well as between modalities, is crucial for optimal performance.
Shuang Wu 0003, Sravanthi Bondugula, Florian Luisier, Xiaodan Zhuang, Pradeep Natarajan
CVPR5
2014 Text Classification via iVector Based Feature Representation
abstract
In this paper, we address the problem of text classification: classifying modern machine-printed text, handwritten text and historical typewritten text from degraded noisy documents. We propose a novel text classification approach based on iVector, a newly developed concept in speaker verification. To a given text line, the iVector is a fixed-length feature vector representation, transformed from a high-dimensional super vector based on means of Gaussian mixture model (GMM), where the text dependent component is separated from a universal background model (UBM) and can be represented by a low dimensional set of factors. We classify the text lines with a discriminative classifier - support vector machine (SVM) in iVector space. A baseline approach of text classification using GMM in feature space is also presented for evaluation purpose. Experimental results on an Arabic document database show accuracy of 92.04% for text line classification using the proposed method. Furthermore, the relative word error rate (WER) of 9.6% is decreased in optical character recognition (OCR) when coupled with the proposed iVector-SVM classifier. The proposed iVector-SVM approach is language independent, thus, can be applied to other scripts as well.
Shengxin Zha, Xujun Peng, Huaigu Cao, Xiaodan Zhuang, Pradeep Natarajan, Premkumar Natarajan
Document Analysis Systems5
2014 Text detection and recognition in natural scenes and consumer videos
abstract
We propose an end-to-end system for text detection and recognition in natural scenes and consumer videos. Maximally Stable Extremal Regions which are robust to illumination and viewpoint variations are selected as text candidates. Rich shape descriptors such as Histogram of Oriented Gradients, Gabor filter, corners and geometrical features are used to represent the candidates and classified using a support vector machine. Positively labeled candidates serve as anchor regions for word formation. We then group candidate regions based on geometric and color properties to form word boundaries. To speed up the system for practical applications, we use Partial Least Squares approach for dimensionality reduction. The detected words are binarized, filtered and passed to a hidden Markov model based Optical Character Recognition (OCR) system for recognition. We show significant improvement in text detection and recognition tasks over previous approaches on a large consumer video dataset. Furthermore, the event detection system built upon the OCR output of this approach outperformed multiple other OCR-only based submissions in the recently concluded NIST TRECVID 2013 multimedia event detection evaluations.
Xujun Peng, Xiaodan Zhuang, Pradeep Natarajan, Huaigu Cao
ICASSP4
2014 Effective representations for leveraging language content in multimedia event detection
abstract
Language content in videos from speech and overlaid or inscene video text can provide high precision signals for video event detection and retrieval. However, sporadic occurrence, content that is unrelated to the events of interest, and high error rates of current speech and text recognition systems on consumer domain video make it difficult to exploit these channels. In this paper, we study different representations of language content to address these challenges. First, we utilize likelihood weighted word lattices obtained from a Hidden Markov Model (HMM) based decoding engine to encode many alternate hypotheses, rather than relying on noisy single best hypotheses. Second, we utilize an event-independent modified term frequency-inverse document frequency (TF-IDF) weighting scheme to obtain the final feature vector. We present detailed experimental results on the TRECVID MED 2013 dataset containing ~150000 videos, and show that our representation significantly outperforms alternate representations for both speech and video text.
Shuang Wu 0003, Xiaodan Zhuang, Pradeep Natarajan
ICASSP3
2014 A constrained optimization approach to combining multiple non-local means denoising estimates
Brian Tracey, Eric L. Miller 0001, Yue Wu 0001, Pradeep Natarajan, Joseph P. Noonan
Signal Process.4
2013 Graph based multimodal word clustering for video event detection
abstract
Combining diverse low-level features from multiple modalities has consistently improved performance over a range of video processing tasks, including event detection. In our work, we study graph based clustering techniques for integrating information from multiple modalities by identifying word clusters spread across the different modalities. We present different methods to identify word clusters including word similarity graph partitioning, word-video co-clustering and Latent Semantic Indexing and the impact of different metrics to quantify the co-occurrence of words. We present experimental results on a ≈45000 video dataset used in the TRECVID MED 11 evaluations. Our experiments show that multimodal features have consistent performance gains over the use of individual features. Further, word similarity graph construction using a complete graph representation consistently improves over partite graphs and early fusion based multimodal systems. Finally, we see additional performance gains by fusing multimodal features with individual features.
Aravind Namandi Vembu, Pradeep Natarajan, Shuang Wu 0003, Rohit Prasad, Premkumar Natarajan
ICASSP2
2013 Audio self organized units for high-level event detection
Xiaodan Zhuang, Shuang Wu 0003, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan
INTERSPEECH3
2013 Compact bag-of-words visual representation for effective linear classification
abstract
Bag-of-words approaches have been shown to achieve state-of-the-art performance in large-scale multimedia event detection. However, the commonly used histogram representation of bag-of-words requires large codebook sizes and expensive nonlinear kernel based classifiers for optimal performance. To address these two issues, we present a two-part generative model for compact visual representation, based on the i-vector approach recently proposed for speech and audio modeling. First, we use a Gaussian mixture model (GMM) to model the joint distribution of local descriptors. Second, we use a low-dimensional factor representation that constrains the GMM parameters to a subspace that preserves most of the information. We further extend this method to incorporate overlapping spatial regions, forming a highly compact visual representation that achieves superior performance with fast linear classifiers. We evaluate the method on a large video dataset used in the TRECVID 2011 MED evaluation. With linear classifiers, the proposed representation, with one-tenth of the storage footprint, outperforms soft quantization histograms used in the top performing TRECVID 2011 MED systems.
Xiaodan Zhuang, Shuang Wu 0003, Pradeep Natarajan
ACM Multimedia3
2013 Ridge Regression based classifiers for large scale class imbalanced datasets
abstract
Large scale, class imbalanced data classification is a challenging task that occurs frequently in several computer vision tasks such as web video retrieval. A number of algorithms have been proposed in literature that approach this problem from different perspectives (e.g. Sampling, Cost-sensitive learning, Active learning). The challenge is two fold in this task - first the data imbalance causes many classification algorithms to learn trivial classifiers that declare all test examples to be from the majority class. Second, many algorithms do not scale to large dataset sizes. We address these two issues by using two different cost-sensitive versions of Ridge Regression as our binary classifiers. We demonstrate our approach for retrieving unstructured web videos from 10 events on the benchmark TRECVID MED 12 dataset containing ≈47000 videos. We empirically show that they perform at par with state-of-the-art support vector machine based classifiers using χ2kernels while being 30 to 60 times faster.
Devansh Arpit, Shuang Wu 0003, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan
WACV3
2013 Scene image categorization and video event detection using Naive Bayes Nearest Neighbor
abstract
We present a detailed study of Naive Bayes Nearest Neighbor (NBNN) proposed by Boiman et al., with application to scene categorization and video event detection. Our study indicates that using Dense-SIFT along with dimensionality reduction using PCA enables NBNN to obtain state-of-the-art results. We demonstrate this on two tasks: (1) scene image categorization on the UIUC 8 Sports Events Image Dataset (obtaining 84.67%) and the MIT 67 Indoor Scene Image Dataset (obtaining 48.84%); and (2) detecting videos depicting certain events of interest on the challenging MED'11 video dataset with only 15 positive training videos per event. We present an extension referred to as sparse-NBNN that constrains the number of training images that can used to match with a given test image for the image-to-class distance computation. Experiments indicate that this improves upon NBNN for handling of imbalanced training data.
Shiv Vitaladevuni, Pradeep Natarajan, Shuang Wu 0003, Xiaodan Zhuang, Rohit Prasad, Premkumar Natarajan
WACV2
2013 Hierarchical multi-channel hidden semi Markov graphical models for activity recognition
Pradeep Natarajan, Ramakant Nevatia
Comput. Vis. Image Underst.1
2012 Multimodal feature fusion for robust event detection in web videos
abstract
Combining multiple low-level visual features is a proven and effective strategy for a range of computer vision tasks. However, limited attention has been paid to combining such features with information from other modalities, such as audio and videotext, for large scale analysis of web videos. In our work, we rigorously analyze and combine a large set of low-level features that capture appearance, color, motion, audio and audio-visual co-occurrence patterns in videos. We also evaluate the utility of high-level (i.e., semantic) visual information obtained from detecting scene, object, and action concepts. Further, we exploit multimodal information by analyzing available spoken and videotext content using state-of-the-art automatic speech recognition (ASR) and videotext recognition systems. We combine these diverse features using a two-step strategy employing multiple kernel learning (MKL) and late score level fusion methods. Based on the TRECVID MED 2011 evaluations for detecting 10 events in a large benchmark set of ~45000 videos, our system showed the best performance among the 19 international teams.
Pradeep Natarajan, Shuang Wu 0003, Shiv Vitaladevuni, Xiaodan Zhuang, Stavros Tsakalidis, Unsang Park, Rohit Prasad, Premkumar Natarajan
CVPR1
2012 Multi-channel Shape-Flow Kernel Descriptors for Robust Video Event Detection and Retrieval
Pradeep Natarajan, Shuang Wu 0003, Shiv Vitaladevuni, Xiaodan Zhuang, Unsang Park, Rohit Prasad, Premkumar Natarajan
ECCV (2)1
2012 Robust Event Detection From Spoken Content In Consumer Domain Videos
Stavros Tsakalidis, Xiaodan Zhuang, Roger Hsiao, Shuang Wu 0003, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan
INTERSPEECH5
2012 Compact Audio Representation for Event Detection in Consumer Media
Xiaodan Zhuang, Stavros Tsakalidis, Shuang Wu 0003, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan
INTERSPEECH4
2011 Wavelet Band-pass Filters for Matching Multiple Templates in Real-time
abstract
Many applications in image processing and computer vision require finding a particular template in an image or a video, that is, template matching. Given a template and an input, the matching algorithm finds the region of interest (ROI) that most closely matches the template in terms of some similarity measurement. According to the way similarity measurements are performed, the template matching methods can roughly be classified into two groups: 1) patch matching schemes, such as the sum of absolute difference (SAD) , the sum of squared difference (SSD) [1], or cross correlation (XCORR), where the similarity measurement directly relies on pixel information from the patch of interest; and 2) feature matching schemes, such as invariant features [4] and bags of features [5], where similarity measurement relies on features describing the template and the frame. Patch matching methods are not robust, especially when noise, skew, or errors occur. Further, they consume a large amount of time, because of expensive sliding window search for calculating the similarity score over all possible locations. Several techniques have been explored for accelerating such matching methods, including early rejections and correlation techniques [1]. However, the computation cost could still be unaffordable when the frame size is large. Typically other techniques, like frame difference, are used to reduce the search space in applications. Feature matching methods process the template and describe it with features, which are ideally invariant to rotation, skew, noise etc. However in many cases, the use of a more complicated model for similarity measurement results in higher computational cost. Further, sliding window search is also a costly stage for such methods. While there exist known algorithms for fast search of object instances in an image using branch-and-bound techniques, in our particular problem, methods of this type have two crucial limitations. First, they require a large number of training samples for each class to learn robust classifiers. Second, interest point detectors like SIFT [5] typically do not generate sufficient number of feature points, because of the small size of the provided logo, large homogenous regions and degradations. Wavelets based approaches have been extensively used in object detection and recognition. In [3], wavelet coefficients based image histogram are collected in bins and are used for classifying logos. In [6], wavelet coefficients are directly used and trained for pedestrian detection. In [8], wavelet coefficients are selected to form rotation-invariant features by using the angular-radial transform. However, matching logos within frames using [3, 6, 8] still requires expensive window searching and thus are not appropriate for real-time processing. In this paper, we propose a new matching method using the wavelet based band-pass filters (WBPFs). Instead of using direct distance measurement requiring expensive window search, the similarity is measured in the indirect way involving two stages. In the stage of offline template processing (see Figure 1), a template is automatically described by a set of three directional WBPFs, where only salient wavelet frequency components of the template are allowed to pass. In the stage of online frame processing (see Figure 2), a frame is transformed to the wavelet domain and its sub-bands are filtered with respect to the corresponding template WBFPs. Finally, the detection is made with respect to the region of the densest responses under spatial constraints [2, 4]. We show that the proposed template matching system has a very low computational cost, which is 50 times faster than the correlation based SSD [1] and 10 times faster than the orthogonal Haar transform (OHT) based SSD [7]. Further, the proposed method does not trade-off accuracy, since the use of subtemplate information makes it robust to skew and camera view change. Experimental results demonstrate our method for real-time logo detection in broadcast videos.
Yue Wu 0001, Pradeep Natarajan, Joseph P. Noonan, Rohit Prasad, Premkumar Natarajan
BMVC2
2011 Efficient Orthogonal Matching Pursuit using sparse random projections for scene and video classification
abstract
Sparse projection has been shown to be highly effective in several domains, including image denoising and scene / object classification. However, practical application to large scale problems such as video analysis requires efficient versions of sparse projection algorithms such as Orthogonal Matching Pursuit (OMP). In particular, random projection based locality sensitive hashing (LSH) has been proposed for OMP. In this paper, we propose a novel technique called Comparison Hadamard random projection (CHRP) for further improving the efficiency of LSH within OMP. CHRP combines two techniques:(1) The Fast Johnson-Lindenstrauss Transform (FJLT) which uses a randomized Hadamard transform and sparse projection matrix for LSH, and (2) Achlioptas' random projection that uses only addition and comparison operations. Our approach provides the robustness of FJLT while completely avoiding multiplications. We empirically validate CHRP's efficacy by performing a suite of experiments for image denoising, scene classification, and video categorization. Our experiments indicate that CHRP significantly speeds-up OMP with negligible loss in classification accuracy.
Shiv Vitaladevuni, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan
ICCV2
2011 Baseline Dependent Percentile Features for Offline Arabic Handwriting Recognition
abstract
Handwritten text in Arabic and other languages exhibit significant variations in the slant and baseline of characters across words and also within a single word. Since the concept of baseline does not have a precise mathematical definition, existing approaches use heuristic methods to first identify a set of baseline relevant pixels and then fit lines/curves through them. However, for statistical features like percentiles that we use in our system, we only need an approximate curve that is close to the baseline to normalize the features. Hence we propose a two stage approach to estimate the approximate baseline. First we segment the text line into a set of components, and then estimate the baseline in each component using two methods max projection and smoothed centroid line. We incorpate the computed baseline into percentile feature computation in the BBN Byblos OCR system for an Arabic offline handwriting recognition task. Our new features, result in a 1% absolute gain and 3.1% relative gain in the word error rate on a large test set with 15K handwritten Arabic words, which is statistically significant with p-value<;0.001 using the matched pair comparison test. Further, our results show that computing fine-grained baselines from small line segments is significantly better than estimating a single baseline over the entire text line.
Pradeep Natarajan, David Belanger 0001, Rohit Prasad, Matin Kamali, Krishna Subramanian 0001, Premkumar Natarajan
ICDAR1
2011 Large-scale, real-time logo recognition in broadcast videos
abstract
Robust, real-time, multi-class logo detection in high resolution broadcast videos presents several difficult challenges. For most logos we only have a few training samples, which makes training robust classifiers hard. Also, logos could potentially occur anywhere in the image, and traditional sliding window approaches for logo/object detection are computationally intensive. We present a system that addresses these issues by first identifying a small set of possible logo locations in a frame, based on temporal continuity and multi-resolution search, and then successively pruning these locations for each logo template, using a cascade of color and edge based features. We present experimental results that demonstrate our system for detecting a total of 270 different logo classes in broadcast video from 5 different languages (English, Indonesian, Malay, Simplified and Traditional Chinese).
Pradeep Natarajan, Yue Wu 0001, Shirin Saleem, Ehry MacRostie, Fred Bernardin, Rohit Prasad, Premkumar Natarajan
ICME1
2011 Unsupervised Audio Analysis for Categorizing Heterogeneous Consumer Domain Videos
Pradeep Natarajan, Stavros Tsakalidis, Vasant Manohar, Rohit Prasad, Premkumar Natarajan
INTERSPEECH1
2011 Audio-visual fusion using bayesian model combination for web video retrieval
abstract
Combining features from multiple, heterogeneous, audio visual sources can significantly improve retrieval performance in consumer domain videos. However, such videos often contain unrelated overlaid audio content, or have significant camera motion to reliably extract visual features. We present an approach, which overcomes errors in individual feature streams by combining classifiers trained on multiple, heterogeneous feature streams using Bayesian model combination (BAYCOM). We demonstrate our method, by combining low-level audio and visual features, for classification of a large 200 hour web video corpus. The combined models outperform any of the individual features by 10%. Further, BAYCOM consistently outperforms traditional early and late fusion methods.
Vasant Manohar, Stavros Tsakalidis, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan
ACM Multimedia3
2010 Learning 3D action models from a few 2D videos for view invariant action recognition
abstract
Most existing approaches for learning action models work by extracting suitable low-level features and then training appropriate classifiers. Such approaches require large amounts of training data and do not generalize well to variations in viewpoint, scale and across datasets. Some work has been done recently to learn multi-view action models from Mocap data, but obtaining such data is time consuming and requires costly infrastructure. We present a method that addresses both these issues by learning action models from just a few video training samples. We model each action as a sequence of primitive actions, represented as functions which transform the actor's state. We formulate model learning as a curve-fitting problem, and present a novel algorithm for learning human actions by lifting 2D annotations of a few keyposes to 3D and interpolating between them. Actions are inferred by sampling the models and accumulating the feature weights learned discriminatively using a latent state Perceptron algorithm. We show results comparable to state-of-art on the standard Weizmann dataset, with a much smaller train:test ratio, and also in datasets for visual gesture recognition and cluttered grocery store environments.
Pradeep Natarajan, Vivek K. Singh 0002, Ramakant Nevatia
CVPR1
2008 View and scale invariant action recognition using multiview shape-flow models
abstract
Actions in real world applications typically take place in cluttered environments with large variations in the orientation and scale of the actor. We present an approach to simultaneously track and recognize known actions that is robust to such variations, starting from a person detection in the standing pose. In our approach we first render synthetic poses from multiple viewpoints using Mocap data for known actions and represent them in a conditional random field (CRF) whose observation potentials are computed using shape similarity and the transition potentials are computed using optical flow. We enhance these basic potentials with terms to represent spatial and temporal constraints and call our enhanced model the shape, flow, duration-conditional random field (SFD-CRF). We find the best sequence of actions using Viterbi search in the SFD-CRF. We demonstrate our approach on videos from multiple viewpoints and in the presence of background clutter.
Pradeep Natarajan, Ramakant Nevatia
CVPR1
2007 Hierarchical Multi-channel Hidden Semi Markov Models
Pradeep Natarajan, Ramakant Nevatia
IJCAI1