VLDB 2026 Research / reviewers in the wild / expert
Abhishek Kumar 0001
dblp:67/6188-1
· DBLP profile ↗
25ranked-venue papers
10as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 9 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
19 papers |
Trustworthy machine learning · 28% Transfer learning and domain adaptation · 21% Learning paradigms · 10% | |
| Databases, data mining, and information retrieval
3 papers |
Data mining · 72% Information retrieval · 28% |
Topics — the 30 heaviest of 51, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
1.4 | 2 | 2024 | The Group Robustness is in the Details: Revisiting Finetuning under Spurious Correlations · NeurIPS 2024 Towards Last-layer Retraining for Group Robustness with Fewer Annotations · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › fairness
group robustness |
1.4 | 2 | 2024 | The Group Robustness is in the Details: Revisiting Finetuning under Spurious Correlations · NeurIPS 2024 Towards Last-layer Retraining for Group Robustness with Fewer Annotations · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › robustness
spurious correlation |
1.4 | 2 | 2024 | The Group Robustness is in the Details: Revisiting Finetuning under Spurious Correlations · NeurIPS 2024 Towards Last-layer Retraining for Group Robustness with Fewer Annotations · NeurIPS 2023 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
causal graphical model |
0.9 | 1 | 2025 | Score-based Causal Representation Learning: Linear and General Transformations · J. Mach. Learn. Res. 2025 |
Machine learning › Representation and self-supervised learning
causal representation learning |
0.9 | 1 | 2025 | Score-based Causal Representation Learning: Linear and General Transformations · J. Mach. Learn. Res. 2025 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.8 | 2 | 2021 | Bayesian Structural Adaptation for Continual Learning · ICML 2021 Variational Inference of Disentangled Latent Concepts from Unlabeled Observations · ICLR (Poster) 2018 |
Machine learning › Generative modeling
generative adversarial network |
0.8 | 2 | 2021 | Generalized Adversarially Learned Inference · AAAI 2021 Semi-supervised Learning with GANs: Manifold Invariance with Improved Inference · NIPS 2017 |
Machine learning › Trustworthy machine learning › robustness › model robustness evaluation
worst-group accuracy |
0.8 | 1 | 2024 | The Group Robustness is in the Details: Revisiting Finetuning under Spurious Correlations · NeurIPS 2024 |
Machine learning › Transfer learning and domain adaptation › fine-tuning
last-layer retraining |
0.7 | 1 | 2023 | Towards Last-layer Retraining for Group Robustness with Fewer Annotations · NeurIPS 2023 |
Machine learning › Transfer learning and domain adaptation › parameter-efficient transfer learning
selective finetuning |
0.7 | 1 | 2023 | Towards Last-layer Retraining for Group Robustness with Fewer Annotations · NeurIPS 2023 |
Machine learning › Generative modeling › adversarial inference
adversarially learned inference |
0.5 | 1 | 2021 | Generalized Adversarially Learned Inference · AAAI 2021 |
Machine learning › Learning paradigms
continual learning |
0.5 | 1 | 2021 | Bayesian Structural Adaptation for Continual Learning · ICML 2021 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › variational learning
variational continual learning |
0.5 | 1 | 2021 | Bayesian Structural Adaptation for Continual Learning · ICML 2021 |
Machine learning › Learning paradigms
multi-task learning |
0.4 | 2 | 2017 | Fully-Adaptive Feature Sharing in Multi-Task Networks with Applications in Person Attribute Classification · CVPR 2017 Learning Task Grouping and Overlap in Multi-task Learning · ICML 2012 |
Machine learning › Transfer learning and domain adaptation
parameter-efficient transfer learning |
0.4 | 1 | 2019 | SpotTune: Transfer Learning Through Adaptive Fine-Tuning · CVPR 2019 |
Machine learning › Efficient and distributed learning › adaptive computation
adaptive inference |
0.3 | 1 | 2018 | BlockDrop: Dynamic Inference Paths in Residual Networks · CVPR 2018 |
Machine learning › Efficient and distributed learning › adaptive computation
conditional computation |
0.3 | 1 | 2018 | BlockDrop: Dynamic Inference Paths in Residual Networks · CVPR 2018 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.3 | 1 | 2018 | Variational Inference of Disentangled Latent Concepts from Unlabeled Observations · ICLR (Poster) 2018 |
Machine learning › Transfer learning and domain adaptation
domain alignment |
0.3 | 1 | 2018 | Co-regularized Alignment for Unsupervised Domain Adaptation · NeurIPS 2018 |
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot image classification |
0.3 | 1 | 2018 | Delta-encoder: an effective sample synthesis method for few-shot object recognition · NeurIPS 2018 |
Machine learning › Transfer learning and domain adaptation › few-shot learning › few-shot image classification
one-shot recognition |
0.3 | 1 | 2018 | Delta-encoder: an effective sample synthesis method for few-shot object recognition · NeurIPS 2018 |
Machine learning › Deep learning architectures and training › convolutional neural network
residual network |
0.3 | 1 | 2018 | BlockDrop: Dynamic Inference Paths in Residual Networks · CVPR 2018 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation |
0.3 | 1 | 2018 | Co-regularized Alignment for Unsupervised Domain Adaptation · NeurIPS 2018 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning |
0.3 | 1 | 2017 | Semi-supervised Learning with GANs: Manifold Invariance with Improved Inference · NIPS 2017 |
Machine learning › Deep learning architectures and training › neural network layer design
pooling |
0.3 | 1 | 2017 | S3Pool: Pooling with Stochastic Spatial Sampling · CVPR 2017 |
Machine learning › Learning paradigms
semi-supervised learning |
0.3 | 1 | 2017 | Semi-supervised Learning with GANs: Manifold Invariance with Improved Inference · NIPS 2017 |
Data mining
clustering |
0.2 | 2 | 2011 | Co-regularized Multi-view Spectral Clustering · NIPS 2011 A Co-training Approach for Multi-view Spectral Clustering · ICML 2011 |
Data mining › clustering
multi-view clustering |
0.2 | 2 | 2011 | Co-regularized Multi-view Spectral Clustering · NIPS 2011 A Co-training Approach for Multi-view Spectral Clustering · ICML 2011 |
Data mining › clustering
spectral clustering |
0.2 | 2 | 2011 | Co-regularized Multi-view Spectral Clustering · NIPS 2011 A Co-training Approach for Multi-view Spectral Clustering · ICML 2011 |
Computational geometry
convex geometry |
0.2 | 1 | 2013 | Fast Conical Hull Algorithms for Near-separable Non-negative Matrix Factorization · ICML (1) 2013 |
Methods — techniques the papers use, named apart from their topics
fine-tuning · 1.4variational inference · 1.3score function · 0.9nonparametric latent causal model · 0.9intervention · 0.9upsampling · 0.8loss upweighting · 0.8class balancing · 0.8policy network · 0.7deep feature reweighting · 0.7separable NMF · 0.2parallel implementation · 0.2extreme ray computation · 0.2quadratic constrained quadratic program · 0.1generalized eigenvalue problem · 0.1canonical correlation analysis · 0.1co-training · 0.1co-regularization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Score-based Causal Representation Learning: Linear and General TransformationsabstractThis paper addresses intervention-based causal representation learning (CRL) under a general nonparametric latent causal model and an unknown transformation that maps the latent variables to the observed variables. Linear and general transformations are investigated. The paper addresses both the identifiability and achievability aspects. Identifiability refers to determining algorithm-agnostic conditions that ensure the recovery of the true latent causal variables and the underlying latent causal graph. Achievability refers to the algorithmic aspects and addresses designing algorithms that achieve identifiability guarantees. By drawing novel connections between score functions (i.e., the gradients of the logarithm of density functions) and CRL, this paper designs a score-based class of algorithms that ensures both identifiability and achievability. First, the paper focuses on linear transformations and shows that one stochastic hard intervention per node suffices to guarantee identifiability. It also provides partial identifiability guarantees for soft interventions, including identifiability up to mixing with parents for general causal models and perfect recovery of the latent graph for sufficiently nonlinear causal models. Secondly, it focuses on general transformations and demonstrates that two stochastic hard interventions per node are sufficient for identifiability. This is achieved by defining a differentiable loss function whose global optima ensure identifiability for general CRL. Notably, one does not need to know which pair of interventional environments has the same node intervened. Finally, the theoretical results are empirically validated via experiments on structured synthetic data and image data. Burak Varici, Emre Acartürk, Karthikeyan Shanmugam 0001, Abhishek Kumar 0001, Ali Tajer |
J. Mach. Learn. Res. | 4 |
| 2024 | The Group Robustness is in the Details: Revisiting Finetuning under Spurious CorrelationsabstractModern machine learning models are prone to over-reliance on spurious correlations, which can often lead to poor performance on minority groups. In this paper, we identify surprising and nuanced behavior of finetuned models on worst-group accuracy via comprehensive experiments on four well-established benchmarks across vision and language tasks. We first show that the commonly used class-balancing techniques of mini-batch upsampling and loss upweighting can induce a decrease in worst-group accuracy (WGA) with training epochs, leading to performance no better than without class-balancing. While in some scenarios, removing data to create a class-balanced subset is more effective, we show this depends on group structure and propose a mixture method which can outperform both techniques. Next, we show that scaling pretrained models is generally beneficial for worst-group accuracy, but only in conjunction with appropriate class-balancing. Finally, we identify spectral imbalance in finetuning features as a potential source of group disparities --- minority group covariance matrices incur a larger spectral norm than majority groups once conditioned on the classes. Our results show more nuanced interactions of modern finetuned models with group robustness than was previously known. Our code is available at https://github.com/tmlabonte/revisiting-finetuning. Tyler LaBonte, John C. Hill, Vidya Muthukumar, Abhishek Kumar 0001 |
NeurIPS | 5 |
| 2023 | Towards Last-layer Retraining for Group Robustness with Fewer AnnotationsabstractEmpirical risk minimization (ERM) of neural networks is prone to over-reliance on spurious correlations and poor generalization on minority groups. The recent deep feature reweighting (DFR) technique achieves state-of-the-art group robustness via simple last-layer retraining, but it requires held-out group and class annotations to construct a group-balanced reweighting dataset. In this work, we examine this impractical requirement and find that last-layer retraining can be surprisingly effective with no group annotations (other than for model selection) and only a handful of class annotations. We first show that last-layer retraining can greatly improve worst-group accuracy even when the reweighting dataset has only a small proportion of worst-group data. This implies a "free lunch" where holding out a subset of training data to retrain the last layer can substantially outperform ERM on the entire dataset with no additional data, annotations, or computation for training. To further improve group robustness, we introduce a lightweight method called selective last-layer finetuning (SELF), which constructs the reweighting dataset using misclassifications or disagreements. Our experiments present the first evidence that model disagreement upsamples worst-group data, enabling SELF to nearly match DFR on four well-established benchmarks across vision and language tasks with no group annotations and less than 3% of the held-out class annotations. Tyler LaBonte, Vidya Muthukumar, Abhishek Kumar 0001 |
NeurIPS | 3 |
| 2021 | Generalized Adversarially Learned InferenceabstractAllowing effective inference of latent vectors while training GANs can greatly increase their applicability in various downstream tasks. Recent approaches, such as ALI and BiGAN frameworks, develop methods of inference of latent variables in GANs by adversarially training an image generator along with an encoder to match two joint distributions of image and latent vector pairs. We generalize these approaches to incorporate multiple layers of feedback on reconstructions, self-supervision, and other forms of supervision based on prior or learned knowledge about the desired solutions. We achieve this by modifying the discriminator's objective to correctly identify more than two joint distributions of tuples of an arbitrary number of random variables consisting of images, latent vectors, and other variables generated through auxiliary tasks, such as reconstruction and inpainting or as outputs of suitable pre-trained models. We design a non-saturating maximization objective for the generator-encoder pair and prove that the resulting adversarial game corresponds to a global optimum that simultaneously matches all the distributions. Within our proposed framework, we introduce a novel set of techniques for providing self-supervised feedback to the model based on properties, such as patch-level correspondence and cycle consistency of reconstructions. Through comprehensive experiments, we demonstrate the efficacy, scalability, and flexibility of the proposed approach for a variety of tasks. The appendix of the paper can be found at the following link: https://drive.google.com/file/d/1i99e682CqYWMEDXlnqkqrctGLVA9viiz/view?usp=sharing Yatin Dandi, Homanga Bharadhwaj, Abhishek Kumar 0001, Piyush Rai |
AAAI | 3 |
| 2021 | Bayesian Structural Adaptation for Continual LearningabstractContinual Learning is a learning paradigm where learning systems are trained on a sequence of tasks. The goal here is to perform well on the current task without suffering from a performance drop on the previous tasks. Two notable directions among the recent advances in continual learning with neural networks are (1) variational Bayes based regularization by learning priors from previous tasks, and, (2) learning the structure of deep networks to adapt to new tasks. So far, these two approaches have been largely orthogonal. We present a novel Bayesian framework based on continually learning the structure of deep neural networks, to unify these distinct yet complementary approaches. The proposed framework learns the deep structure for each task by learning which weights to be used, and supports inter-task transfer through the overlapping of different sparse subsets of weights learned by different tasks. An appealing aspect of our proposed continual learning framework is that it is applicable to both discriminative (supervised) and generative (unsupervised) settings. Experimental results on supervised and unsupervised benchmarks demonstrate that our approach performs comparably or better than recent advances in continual learning. Abhishek Kumar 0001, Sunabha Chatterjee, Piyush Rai |
ICML | 1 |
| 2019 | SpotTune: Transfer Learning Through Adaptive Fine-TuningabstractTransfer learning, which allows a source task to affect the inductive bias of the target task, is widely used in computer vision. The typical way of conducting transfer learning with deep neural networks is to fine-tune a model pretrained on the source task using data from the target task. In this paper, we propose an adaptive fine-tuning approach, called SpotTune, which finds the optimal fine-tuning strategy per instance for the target data. In SpotTune, given an image from the target task, a policy network is used to make routing decisions on whether to pass the image through the fine-tuned layers or the pre-trained layers. We conduct extensive experiments to demonstrate the effectiveness of the proposed approach. Our method outperforms the traditional fine-tuning approach on 12 out of 14 standard datasets. We also compare SpotTune with other state-of-the-art fine-tuning strategies, showing superior performance. On the Visual Decathlon datasets, our method achieves the highest score across the board without bells and whistles. Yunhui Guo, Humphrey Shi, Abhishek Kumar 0001, Kristen Grauman, Tajana Rosing, Rogério Feris |
CVPR | 3 |
| 2018 | BlockDrop: Dynamic Inference Paths in Residual NetworksabstractVery deep convolutional neural networks offer excellent recognition results, yet their computational expense limits their impact for many real-world applications. We introduce BlockDrop, an approach that learns to dynamically choose which layers of a deep network to execute during inference so as to best reduce total computation without degrading prediction accuracy. Exploiting the robustness of Residual Networks (ResNets) to layer dropping, our framework selects on-the-fly which residual blocks to evaluate for a given novel image. In particular, given a pretrained ResNet, we train a policy network in an associative reinforcement learning setting for the dual reward of utilizing a minimal number of blocks while preserving recognition accuracy. We conduct extensive experiments on CIFAR and ImageNet. The results provide strong quantitative and qualitative evidence that these learned policies not only accelerate inference but also encode meaningful visual information. Built upon a ResNet-101 model, our method achieves a speedup of 20% on average, going as high as 36% for some images, while maintaining the same 76.4% top-1 accuracy on ImageNet. Zuxuan Wu, Tushar Nagarajan, Abhishek Kumar 0001, Steven Rennie, Larry Davis 0001, Kristen Grauman, Rogério Feris |
CVPR | 3 |
| 2018 | Variational Inference of Disentangled Latent Concepts from Unlabeled Observations
Abhishek Kumar 0001, Prasanna Sattigeri, Avinash Balakrishnan |
ICLR (Poster) | 1 |
| 2018 | Co-regularized Alignment for Unsupervised Domain AdaptationabstractDeep neural networks, trained with large amount of labeled data, can fail to generalize well when tested with examples from a target domain whose distribution differs from the training data distribution, referred as the source domain. It can be expensive or even infeasible to obtain required amount of labeled data in all possible domains. Unsupervised domain adaptation sets out to address this problem, aiming to learn a good predictive model for the target domain using labeled examples from the source domain but only unlabeled examples from the target domain. Domain alignment approaches this problem by matching the source and target feature distributions, and has been used as a key component in many state-of-the-art domain adaptation methods. However, matching the marginal feature distributions does not guarantee that the corresponding class conditional distributions will be aligned across the two domains. We propose co-regularized domain alignment for unsupervised domain adaptation, which constructs multiple diverse feature spaces and aligns source and target distributions in each of them individually, while encouraging that alignments agree with each other with regard to the class predictions on the unlabeled target examples. The proposed method is generic and can be used to improve any domain adaptation method which uses domain alignment. We instantiate it in the context of a recent state-of-the-art method and observe that it provides significant performance improvements on several domain adaptation benchmarks. Abhishek Kumar 0001, Prasanna Sattigeri, Kahini Wadhawan, Leonid Karlinsky, Rogério Feris, William T. Freeman, Gregory W. Wornell |
NeurIPS | 1 |
| 2018 | Delta-encoder: an effective sample synthesis method for few-shot object recognitionabstractLearning to classify new categories based on just one or a few examples is a long-standing challenge in modern computer vision. In this work, we propose a simple yet effective method for few-shot (and one-shot) object recognition. Our approach is based on a modified auto-encoder, denoted delta-encoder, that learns to synthesize new samples for an unseen category just by seeing few examples from it. The synthesized samples are then used to train a classifier. The proposed approach learns to both extract transferable intra-class deformations, or "deltas", between same-class pairs of training examples, and to apply those deltas to the few provided examples of a novel class (unseen during training) in order to efficiently synthesize samples from that new class. The proposed method improves the state-of-the-art of one-shot object-recognition and performs comparably in the few-shot case. Eli Schwartz, Leonid Karlinsky, Joseph Shtok, Sivan Harary, Mattias Marder, Abhishek Kumar 0001, Rogério Feris, Raja Giryes, Alexander M. Bronstein |
NeurIPS | 6 |
| 2017 | Local Group Invariant Representations via Orbit EmbeddingsabstractInvariance to nuisance transformations is one of the desirable properties of effective representations. We consider transformations that form a group and propose an approach based on kernel methods to derive local group invariant representations. Locality is achieved by defining a suitable probability distribution over the group which in turn induces distributions in the input feature space. We learn a decision function over these distributions by appealing to the powerful framework of kernel methods and generate local invariant random feature maps via kernel approximations. We show uniform convergence bounds for kernel approximation and provide generalization bounds for learning with these features. We evaluate our method on three real datasets, including Rotated MNIST and CIFAR-10, and observe that it outperforms competing kernel based approaches. The proposed method also outperforms deep CNN on Rotated MNIST and performs comparably to the recently proposed group-equivariant CNN. Anant Raj, Abhishek Kumar 0001, Youssef Mroueh, P. Thomas Fletcher, Bernhard Schölkopf |
AISTATS | 2 |
| 2017 | Fully-Adaptive Feature Sharing in Multi-Task Networks with Applications in Person Attribute ClassificationabstractMulti-task learning aims to improve generalization performance of multiple prediction tasks by appropriately sharing relevant information across them. In the context of deep neural networks, this idea is often realized by hand-designed network architectures with layers that are shared across tasks and branches that encode task-specific features. However, the space of possible multi-task deep architectures is combinatorially large and often the final architecture is arrived at by manual exploration of this space, which can be both error-prone and tedious. We propose an automatic approach for designing compact multi-task deep learning architectures. Our approach starts with a thin multi-layer network and dynamically widens it in a greedy manner during training. By doing so iteratively, it creates a tree-like deep architecture, on which similar tasks reside in the same branch until at the top layers. Evaluation on person attributes classification tasks involving facial and clothing attributes suggests that the models produced by the proposed method are fast, compact and can closely match or exceed the state-of-the-art accuracy from strong baselines by much more expensive models. Yongxi Lu, Abhishek Kumar 0001, Shuangfei Zhai, Yu Cheng 0001, Tara Javidi, Rogério Feris |
CVPR | 2 |
| 2017 | S3Pool: Pooling with Stochastic Spatial Sampling
Shuangfei Zhai, Hui Wu 0009, Abhishek Kumar 0001, Yu Cheng 0001, Yongxi Lu, Zhongfei Zhang, Rogério Feris |
CVPR | 3 |
| 2017 | Semi-supervised Learning with GANs: Manifold Invariance with Improved InferenceabstractSemi-supervised learning methods using Generative adversarial networks (GANs) have shown promising empirical success recently. Most of these methods use a shared discriminator/classifier which discriminates real examples from fake while also predicting the class label. Motivated by the ability of the GANs generator to capture the data manifold well, we propose to estimate the tangent space to the data manifold using GANs and employ it to inject invariances into the classifier. In the process, we propose enhancements over existing methods for learning the inverse mapping (i.e., the encoder) which greatly improves in terms of semantic similarity of the reconstructed sample with the input sample. We observe considerable empirical gains in semi-supervised learning over baselines, particularly in the cases when the number of labeled examples is low. We also provide insights into how fake examples influence the semi-supervised learning procedure. Abhishek Kumar 0001, Prasanna Sattigeri, P. Thomas Fletcher |
NIPS | 1 |
| 2016 | Scalable Exemplar Clustering and Facility Location via Augmented Block Coordinate Descent with Column GenerationabstractIn recent years exemplar clustering has become a popular tool for applications in document and video summarization, active learning, and clustering with general similarity, where cluster centroids are required to be a subset of the data samples rather than their linear combinations. The problem is also well-known as facility location in the operations research literature. While the problem has well-developed convex relaxation with approximation and recovery guarantees, it has number of variables grows quadratically with the number of samples. Therefore, state-of-the-art methods can hardly handle more than 10^4 samples (i.e. 10^8 variables). In this work, we propose an Augmented-Lagrangian with Block Coordinate Descent (AL-BCD) algorithm that utilizes problem structure to obtain closed-form solution for each block sub-problem, and exploits low-rank representation of the dissimilarity matrix to search active columns without computing the entire matrix. Experiments show our approach to be orders of magnitude faster than existing approaches and can handle problems of up to 10^6 samples. We also demonstrate successful applications of the algorithm on world-scale facility location, document summarization and active learning. Ian En-Hsu Yen, Dmitry Malioutov, Abhishek Kumar 0001 |
AISTATS | 3 |
| 2016 | Large-scale Submodular Greedy Exemplar Selection with Structured Similarity Matrices
Dmitry Malioutov, Abhishek Kumar 0001, Ian En-Hsu Yen |
UAI | 2 |
| 2015 | Near-separable Non-negative Matrix Factorization with ℓ1 and Bregman Loss FunctionsabstractRecently, a family of tractable NMF algorithms have been proposed under the assumption that the data matrix satisfies a separability condition (Donoho & Stodden, 2003; Arora et al., 2012). Geometrically, this condition reformulates the NMF problem as that of finding the extreme rays of the conical hull of a finite set of vectors. In this paper, we develop separable NMF algorithms with ℓ1 loss and Bregman divergences, by extending the conical hull procedures proposed in our earlier work (Kumar et al., 2013). Our methods inherit all the advantages of (Kumar et al., 2013) including scalability and noise-tolerance. We show that on foreground-background separation problems in computer vision, robust near-separable NMFs match the performance of Robust PCA, considered state of the art on these problems, with an order of magnitude faster training time. We also demonstrate applications in exemplar selection settings. Abhishek Kumar 0001, Vikas Sindhwani |
SDM | 1 |
| 2013 | Fast Conical Hull Algorithms for Near-separable Non-negative Matrix FactorizationabstractThe separability assumption (Arora et al., 2012; Donoho & Stodden, 2003) turns non-negative matrix factorization (NMF) into a tractable problem. Recently, a new class of provably-correct NMF algorithms have emerged under this assumption. In this paper, we reformulate the separable NMF problem as that of finding the extreme rays of the conical hull of a finite set of vectors. From this geometric perspective, we derive new separable NMF algorithms that are highly scalable and empirically noise robust, and have several favorable properties in relation to existing methods. A parallel implementation of our algorithm scales excellently on shared and distributed-memory machines. Abhishek Kumar 0001, Vikas Sindhwani, Prabhanjan Kambadur |
ICML (1) | 1 |
| 2012 | Generalized Multiview Analysis: A discriminative latent spaceabstractThis paper presents a general multi-view feature extraction approach that we call Generalized Multiview Analysis or GMA. GMA has all the desirable properties required for cross-view classification and retrieval: it is supervised, it allows generalization to unseen classes, it is multi-view and kernelizable, it affords an efficient eigenvalue based solution and is applicable to any domain. GMA exploits the fact that most popular supervised and unsupervised feature extraction techniques are the solution of a special form of a quadratic constrained quadratic program (QCQP), which can be solved efficiently as a generalized eigenvalue problem. GMA solves a joint, relaxed QCQP over different feature spaces to obtain a single (non)linear subspace. Intuitively, GMA is a supervised extension of Canonical Correlational Analysis (CCA), which is useful for cross-view classification and retrieval. The proposed approach is general and has the potential to replace CCA whenever classification or retrieval is the purpose and label information is available. We outperform previous approaches for textimage retrieval on Pascal and Wiki text-image data. We report state-of-the-art results for pose and lighting invariant face recognition on the MultiPIE face dataset, significantly outperforming other approaches. Abhishek Sharma 0001, Abhishek Kumar 0001, Hal Daumé III, David Jacobs 0001 |
CVPR | 2 |
| 2012 | Learning Task Grouping and Overlap in Multi-task Learning
Abhishek Kumar 0001, Hal Daumé III |
ICML | 1 |
| 2012 | A Binary Classification Framework for Two-Stage Multiple Kernel Learning
Abhishek Kumar 0001, Alexandru Niculescu-Mizil, Koray Kavukcuoglu, Hal Daumé III |
ICML | 1 |
| 2012 | Simultaneously Leveraging Output and Task Structures for Multiple-Output RegressionabstractMultiple-output regression models require estimating multiple functions, one for each output. To improve parameter estimation in such models, methods based on structural regularization of the model parameters are usually needed. In this paper, we present a multiple-output regression model that leverages the covariance structure of the functions (i.e., how the multiple functions are related with each other) as well as the conditional covariance structure of the outputs. This is in contrast with existing methods that usually take into account only one of these structures. More importantly, unlike most of the other existing methods, none of these structures need be known a priori in our model, and are learned from the data. Several previously proposed structural regularization based multiple-output regression models turn out to be special cases of our model. Moreover, in addition to being a rich model for multiple-output regression, our model can also be used in estimating the graphical model structure of a set of variables (multivariate outputs) conditioned on another set of variables (inputs). Experimental results on both synthetic and real datasets demonstrate the effectiveness of our method. Piyush Rai, Abhishek Kumar 0001, Hal Daumé III |
NIPS | 2 |
| 2011 | A Co-training Approach for Multi-view Spectral Clustering
Abhishek Kumar 0001, Hal Daumé III |
ICML | 1 |
| 2011 | Co-regularized Multi-view Spectral ClusteringabstractIn many clustering problems, we have access to multiple views of the data each of which could be individually used for clustering. Exploiting information from multiple views, one can hope to find a clustering that is more accurate than the ones obtained using the individual views. Since the true clustering would assign a point to the same cluster irrespective of the view, we can approach this problem by looking for clusterings that are consistent across the views, i.e., corresponding data points in each view should have same cluster membership. We propose a spectral clustering framework that achieves this goal by co-regularizing the clustering hypotheses, and propose two co-regularization schemes to accomplish this. Experimental comparisons with a number of baselines on two synthetic and three real-world datasets establish the efficacy of our proposed approaches. Abhishek Kumar 0001, Piyush Rai, Hal Daumé III |
NIPS | 1 |
| 2010 | Co-regularization Based Semi-supervised Domain AdaptationabstractThis paper presents a co-regularization based approach to semi-supervised domain adaptation. Our proposed approach (EA++) builds on the notion of augmented space (introduced in EASYADAPT (EA) [1]) and harnesses unlabeled data in target domain to further enable the transfer of information from source to target. This semi-supervised approach to domain adaptation is extremely simple to implement and can be applied as a pre-processing step to any supervised learner. Our theoretical analysis (in terms of Rademacher complexity) of EA and EA++ show that the hypothesis class of EA++ has lower complexity (compared to EA) and hence results in tighter generalization bounds. Experimental results on sentiment analysis tasks reinforce our theoretical findings and demonstrate the efficacy of the proposed method when compared to EA as well as a few other baseline approaches. Hal Daumé III, Abhishek Kumar 0001, Avishek Saha |
NIPS | 2 |