Abhishek Kumar 0001

dblp:67/6188-1 · DBLP profile ↗
← Back
25ranked-venue papers
10as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 9 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
19 papers
Trustworthy machine learning · 28% Transfer learning and domain adaptation · 21% Learning paradigms · 10%
Databases, data mining, and information retrieval
3 papers
Data mining · 72% Information retrieval · 28%

Topics — the 30 heaviest of 51, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
1.422024
The Group Robustness is in the Details: Revisiting Finetuning under Spurious Correlations · NeurIPS 2024
Towards Last-layer Retraining for Group Robustness with Fewer Annotations · NeurIPS 2023
Machine learning › Trustworthy machine learning › fairness
group robustness
1.422024
The Group Robustness is in the Details: Revisiting Finetuning under Spurious Correlations · NeurIPS 2024
Towards Last-layer Retraining for Group Robustness with Fewer Annotations · NeurIPS 2023
Machine learning › Trustworthy machine learning › robustness
spurious correlation
1.422024
The Group Robustness is in the Details: Revisiting Finetuning under Spurious Correlations · NeurIPS 2024
Towards Last-layer Retraining for Group Robustness with Fewer Annotations · NeurIPS 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
causal graphical model
0.912025
Score-based Causal Representation Learning: Linear and General Transformations · J. Mach. Learn. Res. 2025
Machine learning › Representation and self-supervised learning
causal representation learning
0.912025
Score-based Causal Representation Learning: Linear and General Transformations · J. Mach. Learn. Res. 2025
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.822021
Bayesian Structural Adaptation for Continual Learning · ICML 2021
Variational Inference of Disentangled Latent Concepts from Unlabeled Observations · ICLR (Poster) 2018
Machine learning › Generative modeling
generative adversarial network
0.822021
Generalized Adversarially Learned Inference · AAAI 2021
Semi-supervised Learning with GANs: Manifold Invariance with Improved Inference · NIPS 2017
Machine learning › Trustworthy machine learning › robustness › model robustness evaluation
worst-group accuracy
0.812024
The Group Robustness is in the Details: Revisiting Finetuning under Spurious Correlations · NeurIPS 2024
Machine learning › Transfer learning and domain adaptation › fine-tuning
last-layer retraining
0.712023
Towards Last-layer Retraining for Group Robustness with Fewer Annotations · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation › parameter-efficient transfer learning
selective finetuning
0.712023
Towards Last-layer Retraining for Group Robustness with Fewer Annotations · NeurIPS 2023
Machine learning › Generative modeling › adversarial inference
adversarially learned inference
0.512021
Generalized Adversarially Learned Inference · AAAI 2021
Machine learning › Learning paradigms
continual learning
0.512021
Bayesian Structural Adaptation for Continual Learning · ICML 2021
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › variational learning
variational continual learning
0.512021
Bayesian Structural Adaptation for Continual Learning · ICML 2021
Machine learning › Learning paradigms
multi-task learning
0.422017
Fully-Adaptive Feature Sharing in Multi-Task Networks with Applications in Person Attribute Classification · CVPR 2017
Learning Task Grouping and Overlap in Multi-task Learning · ICML 2012
Machine learning › Transfer learning and domain adaptation
parameter-efficient transfer learning
0.412019
SpotTune: Transfer Learning Through Adaptive Fine-Tuning · CVPR 2019
Machine learning › Efficient and distributed learning › adaptive computation
adaptive inference
0.312018
BlockDrop: Dynamic Inference Paths in Residual Networks · CVPR 2018
Machine learning › Efficient and distributed learning › adaptive computation
conditional computation
0.312018
BlockDrop: Dynamic Inference Paths in Residual Networks · CVPR 2018
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.312018
Variational Inference of Disentangled Latent Concepts from Unlabeled Observations · ICLR (Poster) 2018
Machine learning › Transfer learning and domain adaptation
domain alignment
0.312018
Co-regularized Alignment for Unsupervised Domain Adaptation · NeurIPS 2018
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot image classification
0.312018
Delta-encoder: an effective sample synthesis method for few-shot object recognition · NeurIPS 2018
Machine learning › Transfer learning and domain adaptation › few-shot learning › few-shot image classification
one-shot recognition
0.312018
Delta-encoder: an effective sample synthesis method for few-shot object recognition · NeurIPS 2018
Machine learning › Deep learning architectures and training › convolutional neural network
residual network
0.312018
BlockDrop: Dynamic Inference Paths in Residual Networks · CVPR 2018
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.312018
Co-regularized Alignment for Unsupervised Domain Adaptation · NeurIPS 2018
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning
0.312017
Semi-supervised Learning with GANs: Manifold Invariance with Improved Inference · NIPS 2017
Machine learning › Deep learning architectures and training › neural network layer design
pooling
0.312017
S3Pool: Pooling with Stochastic Spatial Sampling · CVPR 2017
Machine learning › Learning paradigms
semi-supervised learning
0.312017
Semi-supervised Learning with GANs: Manifold Invariance with Improved Inference · NIPS 2017
Data mining
clustering
0.222011
Co-regularized Multi-view Spectral Clustering · NIPS 2011
A Co-training Approach for Multi-view Spectral Clustering · ICML 2011
Data mining › clustering
multi-view clustering
0.222011
Co-regularized Multi-view Spectral Clustering · NIPS 2011
A Co-training Approach for Multi-view Spectral Clustering · ICML 2011
Data mining › clustering
spectral clustering
0.222011
Co-regularized Multi-view Spectral Clustering · NIPS 2011
A Co-training Approach for Multi-view Spectral Clustering · ICML 2011
Computational geometry
convex geometry
0.212013
Fast Conical Hull Algorithms for Near-separable Non-negative Matrix Factorization · ICML (1) 2013

Methods — techniques the papers use, named apart from their topics

fine-tuning · 1.4variational inference · 1.3score function · 0.9nonparametric latent causal model · 0.9intervention · 0.9upsampling · 0.8loss upweighting · 0.8class balancing · 0.8policy network · 0.7deep feature reweighting · 0.7separable NMF · 0.2parallel implementation · 0.2extreme ray computation · 0.2quadratic constrained quadratic program · 0.1generalized eigenvalue problem · 0.1canonical correlation analysis · 0.1co-training · 0.1co-regularization · 0.1
YearPublicationVenuePosition
2025 Score-based Causal Representation Learning: Linear and General Transformations
abstract
This paper addresses intervention-based causal representation learning (CRL) under a general nonparametric latent causal model and an unknown transformation that maps the latent variables to the observed variables. Linear and general transformations are investigated. The paper addresses both the identifiability and achievability aspects. Identifiability refers to determining algorithm-agnostic conditions that ensure the recovery of the true latent causal variables and the underlying latent causal graph. Achievability refers to the algorithmic aspects and addresses designing algorithms that achieve identifiability guarantees. By drawing novel connections between score functions (i.e., the gradients of the logarithm of density functions) and CRL, this paper designs a score-based class of algorithms that ensures both identifiability and achievability. First, the paper focuses on linear transformations and shows that one stochastic hard intervention per node suffices to guarantee identifiability. It also provides partial identifiability guarantees for soft interventions, including identifiability up to mixing with parents for general causal models and perfect recovery of the latent graph for sufficiently nonlinear causal models. Secondly, it focuses on general transformations and demonstrates that two stochastic hard interventions per node are sufficient for identifiability. This is achieved by defining a differentiable loss function whose global optima ensure identifiability for general CRL. Notably, one does not need to know which pair of interventional environments has the same node intervened. Finally, the theoretical results are empirically validated via experiments on structured synthetic data and image data.
Burak Varici, Emre Acartürk, Karthikeyan Shanmugam 0001, Abhishek Kumar 0001, Ali Tajer
J. Mach. Learn. Res.4
2024 The Group Robustness is in the Details: Revisiting Finetuning under Spurious Correlations
abstract
Modern machine learning models are prone to over-reliance on spurious correlations, which can often lead to poor performance on minority groups. In this paper, we identify surprising and nuanced behavior of finetuned models on worst-group accuracy via comprehensive experiments on four well-established benchmarks across vision and language tasks. We first show that the commonly used class-balancing techniques of mini-batch upsampling and loss upweighting can induce a decrease in worst-group accuracy (WGA) with training epochs, leading to performance no better than without class-balancing. While in some scenarios, removing data to create a class-balanced subset is more effective, we show this depends on group structure and propose a mixture method which can outperform both techniques. Next, we show that scaling pretrained models is generally beneficial for worst-group accuracy, but only in conjunction with appropriate class-balancing. Finally, we identify spectral imbalance in finetuning features as a potential source of group disparities --- minority group covariance matrices incur a larger spectral norm than majority groups once conditioned on the classes. Our results show more nuanced interactions of modern finetuned models with group robustness than was previously known. Our code is available at https://github.com/tmlabonte/revisiting-finetuning.
Tyler LaBonte, John C. Hill, Vidya Muthukumar, Abhishek Kumar 0001
NeurIPS5
2023 Towards Last-layer Retraining for Group Robustness with Fewer Annotations
abstract
Empirical risk minimization (ERM) of neural networks is prone to over-reliance on spurious correlations and poor generalization on minority groups. The recent deep feature reweighting (DFR) technique achieves state-of-the-art group robustness via simple last-layer retraining, but it requires held-out group and class annotations to construct a group-balanced reweighting dataset. In this work, we examine this impractical requirement and find that last-layer retraining can be surprisingly effective with no group annotations (other than for model selection) and only a handful of class annotations. We first show that last-layer retraining can greatly improve worst-group accuracy even when the reweighting dataset has only a small proportion of worst-group data. This implies a "free lunch" where holding out a subset of training data to retrain the last layer can substantially outperform ERM on the entire dataset with no additional data, annotations, or computation for training. To further improve group robustness, we introduce a lightweight method called selective last-layer finetuning (SELF), which constructs the reweighting dataset using misclassifications or disagreements. Our experiments present the first evidence that model disagreement upsamples worst-group data, enabling SELF to nearly match DFR on four well-established benchmarks across vision and language tasks with no group annotations and less than 3% of the held-out class annotations.
Tyler LaBonte, Vidya Muthukumar, Abhishek Kumar 0001
NeurIPS3
2021 Generalized Adversarially Learned Inference
abstract
Allowing effective inference of latent vectors while training GANs can greatly increase their applicability in various downstream tasks. Recent approaches, such as ALI and BiGAN frameworks, develop methods of inference of latent variables in GANs by adversarially training an image generator along with an encoder to match two joint distributions of image and latent vector pairs. We generalize these approaches to incorporate multiple layers of feedback on reconstructions, self-supervision, and other forms of supervision based on prior or learned knowledge about the desired solutions. We achieve this by modifying the discriminator's objective to correctly identify more than two joint distributions of tuples of an arbitrary number of random variables consisting of images, latent vectors, and other variables generated through auxiliary tasks, such as reconstruction and inpainting or as outputs of suitable pre-trained models. We design a non-saturating maximization objective for the generator-encoder pair and prove that the resulting adversarial game corresponds to a global optimum that simultaneously matches all the distributions. Within our proposed framework, we introduce a novel set of techniques for providing self-supervised feedback to the model based on properties, such as patch-level correspondence and cycle consistency of reconstructions. Through comprehensive experiments, we demonstrate the efficacy, scalability, and flexibility of the proposed approach for a variety of tasks. The appendix of the paper can be found at the following link: https://drive.google.com/file/d/1i99e682CqYWMEDXlnqkqrctGLVA9viiz/view?usp=sharing
Yatin Dandi, Homanga Bharadhwaj, Abhishek Kumar 0001, Piyush Rai
AAAI3
2021 Bayesian Structural Adaptation for Continual Learning
abstract
Continual Learning is a learning paradigm where learning systems are trained on a sequence of tasks. The goal here is to perform well on the current task without suffering from a performance drop on the previous tasks. Two notable directions among the recent advances in continual learning with neural networks are (1) variational Bayes based regularization by learning priors from previous tasks, and, (2) learning the structure of deep networks to adapt to new tasks. So far, these two approaches have been largely orthogonal. We present a novel Bayesian framework based on continually learning the structure of deep neural networks, to unify these distinct yet complementary approaches. The proposed framework learns the deep structure for each task by learning which weights to be used, and supports inter-task transfer through the overlapping of different sparse subsets of weights learned by different tasks. An appealing aspect of our proposed continual learning framework is that it is applicable to both discriminative (supervised) and generative (unsupervised) settings. Experimental results on supervised and unsupervised benchmarks demonstrate that our approach performs comparably or better than recent advances in continual learning.
Abhishek Kumar 0001, Sunabha Chatterjee, Piyush Rai
ICML1
2019 SpotTune: Transfer Learning Through Adaptive Fine-Tuning
abstract
Transfer learning, which allows a source task to affect the inductive bias of the target task, is widely used in computer vision. The typical way of conducting transfer learning with deep neural networks is to fine-tune a model pretrained on the source task using data from the target task. In this paper, we propose an adaptive fine-tuning approach, called SpotTune, which finds the optimal fine-tuning strategy per instance for the target data. In SpotTune, given an image from the target task, a policy network is used to make routing decisions on whether to pass the image through the fine-tuned layers or the pre-trained layers. We conduct extensive experiments to demonstrate the effectiveness of the proposed approach. Our method outperforms the traditional fine-tuning approach on 12 out of 14 standard datasets. We also compare SpotTune with other state-of-the-art fine-tuning strategies, showing superior performance. On the Visual Decathlon datasets, our method achieves the highest score across the board without bells and whistles.
Yunhui Guo, Humphrey Shi, Abhishek Kumar 0001, Kristen Grauman, Tajana Rosing, Rogério Feris
CVPR3
2018 BlockDrop: Dynamic Inference Paths in Residual Networks
abstract
Very deep convolutional neural networks offer excellent recognition results, yet their computational expense limits their impact for many real-world applications. We introduce BlockDrop, an approach that learns to dynamically choose which layers of a deep network to execute during inference so as to best reduce total computation without degrading prediction accuracy. Exploiting the robustness of Residual Networks (ResNets) to layer dropping, our framework selects on-the-fly which residual blocks to evaluate for a given novel image. In particular, given a pretrained ResNet, we train a policy network in an associative reinforcement learning setting for the dual reward of utilizing a minimal number of blocks while preserving recognition accuracy. We conduct extensive experiments on CIFAR and ImageNet. The results provide strong quantitative and qualitative evidence that these learned policies not only accelerate inference but also encode meaningful visual information. Built upon a ResNet-101 model, our method achieves a speedup of 20% on average, going as high as 36% for some images, while maintaining the same 76.4% top-1 accuracy on ImageNet.
Zuxuan Wu, Tushar Nagarajan, Abhishek Kumar 0001, Steven Rennie, Larry Davis 0001, Kristen Grauman, Rogério Feris
CVPR3
2018 Variational Inference of Disentangled Latent Concepts from Unlabeled Observations
Abhishek Kumar 0001, Prasanna Sattigeri, Avinash Balakrishnan
ICLR (Poster)1
2018 Co-regularized Alignment for Unsupervised Domain Adaptation
abstract
Deep neural networks, trained with large amount of labeled data, can fail to generalize well when tested with examples from a target domain whose distribution differs from the training data distribution, referred as the source domain. It can be expensive or even infeasible to obtain required amount of labeled data in all possible domains. Unsupervised domain adaptation sets out to address this problem, aiming to learn a good predictive model for the target domain using labeled examples from the source domain but only unlabeled examples from the target domain. Domain alignment approaches this problem by matching the source and target feature distributions, and has been used as a key component in many state-of-the-art domain adaptation methods. However, matching the marginal feature distributions does not guarantee that the corresponding class conditional distributions will be aligned across the two domains. We propose co-regularized domain alignment for unsupervised domain adaptation, which constructs multiple diverse feature spaces and aligns source and target distributions in each of them individually, while encouraging that alignments agree with each other with regard to the class predictions on the unlabeled target examples. The proposed method is generic and can be used to improve any domain adaptation method which uses domain alignment. We instantiate it in the context of a recent state-of-the-art method and observe that it provides significant performance improvements on several domain adaptation benchmarks.
Abhishek Kumar 0001, Prasanna Sattigeri, Kahini Wadhawan, Leonid Karlinsky, Rogério Feris, William T. Freeman, Gregory W. Wornell
NeurIPS1
2018 Delta-encoder: an effective sample synthesis method for few-shot object recognition
abstract
Learning to classify new categories based on just one or a few examples is a long-standing challenge in modern computer vision. In this work, we propose a simple yet effective method for few-shot (and one-shot) object recognition. Our approach is based on a modified auto-encoder, denoted delta-encoder, that learns to synthesize new samples for an unseen category just by seeing few examples from it. The synthesized samples are then used to train a classifier. The proposed approach learns to both extract transferable intra-class deformations, or "deltas", between same-class pairs of training examples, and to apply those deltas to the few provided examples of a novel class (unseen during training) in order to efficiently synthesize samples from that new class. The proposed method improves the state-of-the-art of one-shot object-recognition and performs comparably in the few-shot case.
Eli Schwartz, Leonid Karlinsky, Joseph Shtok, Sivan Harary, Mattias Marder, Abhishek Kumar 0001, Rogério Feris, Raja Giryes, Alexander M. Bronstein
NeurIPS6
2017 Local Group Invariant Representations via Orbit Embeddings
abstract
Invariance to nuisance transformations is one of the desirable properties of effective representations. We consider transformations that form a group and propose an approach based on kernel methods to derive local group invariant representations. Locality is achieved by defining a suitable probability distribution over the group which in turn induces distributions in the input feature space. We learn a decision function over these distributions by appealing to the powerful framework of kernel methods and generate local invariant random feature maps via kernel approximations. We show uniform convergence bounds for kernel approximation and provide generalization bounds for learning with these features. We evaluate our method on three real datasets, including Rotated MNIST and CIFAR-10, and observe that it outperforms competing kernel based approaches. The proposed method also outperforms deep CNN on Rotated MNIST and performs comparably to the recently proposed group-equivariant CNN.
Anant Raj, Abhishek Kumar 0001, Youssef Mroueh, P. Thomas Fletcher, Bernhard Schölkopf
AISTATS2
2017 Fully-Adaptive Feature Sharing in Multi-Task Networks with Applications in Person Attribute Classification
abstract
Multi-task learning aims to improve generalization performance of multiple prediction tasks by appropriately sharing relevant information across them. In the context of deep neural networks, this idea is often realized by hand-designed network architectures with layers that are shared across tasks and branches that encode task-specific features. However, the space of possible multi-task deep architectures is combinatorially large and often the final architecture is arrived at by manual exploration of this space, which can be both error-prone and tedious. We propose an automatic approach for designing compact multi-task deep learning architectures. Our approach starts with a thin multi-layer network and dynamically widens it in a greedy manner during training. By doing so iteratively, it creates a tree-like deep architecture, on which similar tasks reside in the same branch until at the top layers. Evaluation on person attributes classification tasks involving facial and clothing attributes suggests that the models produced by the proposed method are fast, compact and can closely match or exceed the state-of-the-art accuracy from strong baselines by much more expensive models.
Yongxi Lu, Abhishek Kumar 0001, Shuangfei Zhai, Yu Cheng 0001, Tara Javidi, Rogério Feris
CVPR2
2017 S3Pool: Pooling with Stochastic Spatial Sampling
Shuangfei Zhai, Hui Wu 0009, Abhishek Kumar 0001, Yu Cheng 0001, Yongxi Lu, Zhongfei Zhang, Rogério Feris
CVPR3
2017 Semi-supervised Learning with GANs: Manifold Invariance with Improved Inference
abstract
Semi-supervised learning methods using Generative adversarial networks (GANs) have shown promising empirical success recently. Most of these methods use a shared discriminator/classifier which discriminates real examples from fake while also predicting the class label. Motivated by the ability of the GANs generator to capture the data manifold well, we propose to estimate the tangent space to the data manifold using GANs and employ it to inject invariances into the classifier. In the process, we propose enhancements over existing methods for learning the inverse mapping (i.e., the encoder) which greatly improves in terms of semantic similarity of the reconstructed sample with the input sample. We observe considerable empirical gains in semi-supervised learning over baselines, particularly in the cases when the number of labeled examples is low. We also provide insights into how fake examples influence the semi-supervised learning procedure.
Abhishek Kumar 0001, Prasanna Sattigeri, P. Thomas Fletcher
NIPS1
2016 Scalable Exemplar Clustering and Facility Location via Augmented Block Coordinate Descent with Column Generation
abstract
In recent years exemplar clustering has become a popular tool for applications in document and video summarization, active learning, and clustering with general similarity, where cluster centroids are required to be a subset of the data samples rather than their linear combinations. The problem is also well-known as facility location in the operations research literature. While the problem has well-developed convex relaxation with approximation and recovery guarantees, it has number of variables grows quadratically with the number of samples. Therefore, state-of-the-art methods can hardly handle more than 10^4 samples (i.e. 10^8 variables). In this work, we propose an Augmented-Lagrangian with Block Coordinate Descent (AL-BCD) algorithm that utilizes problem structure to obtain closed-form solution for each block sub-problem, and exploits low-rank representation of the dissimilarity matrix to search active columns without computing the entire matrix. Experiments show our approach to be orders of magnitude faster than existing approaches and can handle problems of up to 10^6 samples. We also demonstrate successful applications of the algorithm on world-scale facility location, document summarization and active learning.
Ian En-Hsu Yen, Dmitry Malioutov, Abhishek Kumar 0001
AISTATS3
2016 Large-scale Submodular Greedy Exemplar Selection with Structured Similarity Matrices
Dmitry Malioutov, Abhishek Kumar 0001, Ian En-Hsu Yen
UAI2
2015 Near-separable Non-negative Matrix Factorization with ℓ1 and Bregman Loss Functions
abstract
Recently, a family of tractable NMF algorithms have been proposed under the assumption that the data matrix satisfies a separability condition (Donoho & Stodden, 2003; Arora et al., 2012). Geometrically, this condition reformulates the NMF problem as that of finding the extreme rays of the conical hull of a finite set of vectors. In this paper, we develop separable NMF algorithms with ℓ1 loss and Bregman divergences, by extending the conical hull procedures proposed in our earlier work (Kumar et al., 2013). Our methods inherit all the advantages of (Kumar et al., 2013) including scalability and noise-tolerance. We show that on foreground-background separation problems in computer vision, robust near-separable NMFs match the performance of Robust PCA, considered state of the art on these problems, with an order of magnitude faster training time. We also demonstrate applications in exemplar selection settings.
Abhishek Kumar 0001, Vikas Sindhwani
SDM1
2013 Fast Conical Hull Algorithms for Near-separable Non-negative Matrix Factorization
abstract
The separability assumption (Arora et al., 2012; Donoho & Stodden, 2003) turns non-negative matrix factorization (NMF) into a tractable problem. Recently, a new class of provably-correct NMF algorithms have emerged under this assumption. In this paper, we reformulate the separable NMF problem as that of finding the extreme rays of the conical hull of a finite set of vectors. From this geometric perspective, we derive new separable NMF algorithms that are highly scalable and empirically noise robust, and have several favorable properties in relation to existing methods. A parallel implementation of our algorithm scales excellently on shared and distributed-memory machines.
Abhishek Kumar 0001, Vikas Sindhwani, Prabhanjan Kambadur
ICML (1)1
2012 Generalized Multiview Analysis: A discriminative latent space
abstract
This paper presents a general multi-view feature extraction approach that we call Generalized Multiview Analysis or GMA. GMA has all the desirable properties required for cross-view classification and retrieval: it is supervised, it allows generalization to unseen classes, it is multi-view and kernelizable, it affords an efficient eigenvalue based solution and is applicable to any domain. GMA exploits the fact that most popular supervised and unsupervised feature extraction techniques are the solution of a special form of a quadratic constrained quadratic program (QCQP), which can be solved efficiently as a generalized eigenvalue problem. GMA solves a joint, relaxed QCQP over different feature spaces to obtain a single (non)linear subspace. Intuitively, GMA is a supervised extension of Canonical Correlational Analysis (CCA), which is useful for cross-view classification and retrieval. The proposed approach is general and has the potential to replace CCA whenever classification or retrieval is the purpose and label information is available. We outperform previous approaches for textimage retrieval on Pascal and Wiki text-image data. We report state-of-the-art results for pose and lighting invariant face recognition on the MultiPIE face dataset, significantly outperforming other approaches.
Abhishek Sharma 0001, Abhishek Kumar 0001, Hal Daumé III, David Jacobs 0001
CVPR2
2012 Learning Task Grouping and Overlap in Multi-task Learning
Abhishek Kumar 0001, Hal Daumé III
ICML1
2012 A Binary Classification Framework for Two-Stage Multiple Kernel Learning
Abhishek Kumar 0001, Alexandru Niculescu-Mizil, Koray Kavukcuoglu, Hal Daumé III
ICML1
2012 Simultaneously Leveraging Output and Task Structures for Multiple-Output Regression
abstract
Multiple-output regression models require estimating multiple functions, one for each output. To improve parameter estimation in such models, methods based on structural regularization of the model parameters are usually needed. In this paper, we present a multiple-output regression model that leverages the covariance structure of the functions (i.e., how the multiple functions are related with each other) as well as the conditional covariance structure of the outputs. This is in contrast with existing methods that usually take into account only one of these structures. More importantly, unlike most of the other existing methods, none of these structures need be known a priori in our model, and are learned from the data. Several previously proposed structural regularization based multiple-output regression models turn out to be special cases of our model. Moreover, in addition to being a rich model for multiple-output regression, our model can also be used in estimating the graphical model structure of a set of variables (multivariate outputs) conditioned on another set of variables (inputs). Experimental results on both synthetic and real datasets demonstrate the effectiveness of our method.
Piyush Rai, Abhishek Kumar 0001, Hal Daumé III
NIPS2
2011 A Co-training Approach for Multi-view Spectral Clustering
Abhishek Kumar 0001, Hal Daumé III
ICML1
2011 Co-regularized Multi-view Spectral Clustering
abstract
In many clustering problems, we have access to multiple views of the data each of which could be individually used for clustering. Exploiting information from multiple views, one can hope to find a clustering that is more accurate than the ones obtained using the individual views. Since the true clustering would assign a point to the same cluster irrespective of the view, we can approach this problem by looking for clusterings that are consistent across the views, i.e., corresponding data points in each view should have same cluster membership. We propose a spectral clustering framework that achieves this goal by co-regularizing the clustering hypotheses, and propose two co-regularization schemes to accomplish this. Experimental comparisons with a number of baselines on two synthetic and three real-world datasets establish the efficacy of our proposed approaches.
Abhishek Kumar 0001, Piyush Rai, Hal Daumé III
NIPS1
2010 Co-regularization Based Semi-supervised Domain Adaptation
abstract
This paper presents a co-regularization based approach to semi-supervised domain adaptation. Our proposed approach (EA++) builds on the notion of augmented space (introduced in EASYADAPT (EA) [1]) and harnesses unlabeled data in target domain to further enable the transfer of information from source to target. This semi-supervised approach to domain adaptation is extremely simple to implement and can be applied as a pre-processing step to any supervised learner. Our theoretical analysis (in terms of Rademacher complexity) of EA and EA++ show that the hypothesis class of EA++ has lower complexity (compared to EA) and hence results in tighter generalization bounds. Experimental results on sentiment analysis tasks reinforce our theoretical findings and demonstrate the efficacy of the proposed method when compared to EA as well as a few other baseline approaches.
Hal Daumé III, Abhishek Kumar 0001, Avishek Saha
NIPS2