EDBT 2026 Demo / reviewers in the wild / expert
Vidit Jain
dblp:68/5650
· DBLP profile ↗
17ranked-venue papers
10as first author
6since 2021 · last 2024
0000-0002-7911-1074ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 8 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Learning theory · 26% Efficient and distributed learning · 20% Planning, search and constraint satisfaction · 15% | |
| Databases, data mining, and information retrieval
3 papers |
Data mining · 41% Information retrieval · 31% Web and social media mining · 14% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% |
Topics — the 25 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory › classification › multiclass classification
extreme classification |
0.8 | 1 | 2024 | Enhancing Tail Performance in Extreme Classifiers by Label Variance Reduction · ICLR 2024 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.8 | 1 | 2024 | Enhancing Tail Performance in Extreme Classifiers by Label Variance Reduction · ICLR 2024 |
Data mining › predictive modeling › classification › multi-label classification
extreme classification |
0.7 | 1 | 2023 | NGAME: Negative Mining-aware Mini-batching for Extreme Classification · WSDM 2023 |
Human-AI interaction
conversational agents |
0.5 | 1 | 2021 | Exploring Semi-Supervised Learning for Predicting Listener Backchannels · CHI 2021 |
Machine learning › Learning theory
generalization |
0.2 | 1 | 2024 | Enhancing Tail Performance in Extreme Classifiers by Label Variance Reduction · ICLR 2024 |
Web and social media mining
social network analysis |
0.2 | 1 | 2015 | Tracking the Conductance of Rapidly Evolving Topic-Subgraphs · Proc. VLDB Endow. 2015 |
Computer vision › Face, body and person analysis
face detection |
0.2 | 2 | 2013 | Online domain adaptation of a pre-trained cascade of classifiers · CVPR 2011 Adapting Classification Cascades to New Domains · ICCV 2013 |
Computer vision › Image recognition and object detection › object detection
cascade classifier |
0.2 | 1 | 2013 | Adapting Classification Cascades to New Domains · ICCV 2013 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.2 | 1 | 2013 | Adapting Classification Cascades to New Domains · ICCV 2013 |
Computer vision › Image recognition and object detection
object detection |
0.2 | 1 | 2013 | Adapting Classification Cascades to New Domains · ICCV 2013 |
Computer vision › Face, body and person analysis
face recognition |
0.1 | 2 | 2007 | People-LDA: Anchoring Topics to People using Face Recognition · ICCV 2007 Unsupervised Joint Alignment of Complex Images · ICCV 2007 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
online domain adaptation |
0.1 | 1 | 2011 | Online domain adaptation of a pre-trained cascade of classifiers · CVPR 2011 |
Information retrieval › image retrieval › web image search
image re-ranking |
0.1 | 1 | 2011 | Learning to re-rank: query-dependent image re-ranking using click data · WWW 2011 |
Information retrieval
image retrieval |
0.1 | 1 | 2011 | Learning to re-rank: query-dependent image re-ranking using click data · WWW 2011 |
Information retrieval › ranking
learning to rank |
0.1 | 1 | 2011 | Learning to re-rank: query-dependent image re-ranking using click data · WWW 2011 |
Information retrieval › ranking
query-dependent ranking |
0.1 | 1 | 2011 | Learning to re-rank: query-dependent image re-ranking using click data · WWW 2011 |
Natural language and speech › Information extraction and text analysis › event extraction
event classification |
0.1 | 1 | 2008 | Selective hidden random fields: Exploiting domain-specific saliency for event classification · CVPR 2008 |
Computer vision › Segmentation and scene understanding › image segmentation
saliency-based segmentation |
0.1 | 1 | 2008 | Selective hidden random fields: Exploiting domain-specific saliency for event classification · CVPR 2008 |
Computer vision › 3D vision
alignment-based recognition |
0.1 | 1 | 2007 | Unsupervised Joint Alignment of Complex Images · ICCV 2007 |
Computer vision › Face, body and person analysis
face alignment |
0.1 | 1 | 2007 | Unsupervised Joint Alignment of Complex Images · ICCV 2007 |
Computer vision › Face, body and person analysis › face recognition
face identification |
0.1 | 1 | 2007 | People-LDA: Anchoring Topics to People using Face Recognition · ICCV 2007 |
Natural language and speech › Information extraction and text analysis
topic model |
0.1 | 1 | 2007 | People-LDA: Anchoring Topics to People using Face Recognition · ICCV 2007 |
Computer vision › 3D vision › image registration
unsupervised joint alignment |
0.1 | 1 | 2007 | Unsupervised Joint Alignment of Complex Images · ICCV 2007 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
gaussian process regression |
0.0 | 1 | 2011 | Online domain adaptation of a pre-trained cascade of classifiers · CVPR 2011 |
Computer vision › 3D vision
object pose estimation |
0.0 | 1 | 2007 | Unsupervised Joint Alignment of Complex Images · ICCV 2007 |
Methods — techniques the papers use, named apart from their topics
semi-supervised learning · 1.1user study · 1.0regularization · 0.8loss re-calibration · 0.8knowledge distillation · 0.8transformer encoder · 0.7mini-batch training · 0.7neural model · 0.6in-memory approximation · 0.2bloom filter · 0.2few-shot positive adaptation · 0.2cascade classifier adaptation · 0.2gaussian process regression · 0.1click data · 0.1classifier cascade · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | HiGen: Hierarchy-Aware Sequence Generation for Hierarchical Text ClassificationabstractHierarchical text classification (HTC) is a complex subtask under multi-label text classification, characterized by a hierarchical label taxonomy and data imbalance. The best-performing models aim to learn a static representation by combining document and hierarchical label information. However, the relevance of document sections can vary based on the hierarchy level, necessitating a dynamic document representation. To address this, we propose HiGen, a text-generation-based framework utilizing language models to encode dynamic text representations. We introduce a level-guided loss function to capture the relationship between text and label name semantics. Our approach incorporates a task-specific pretraining strategy, adapting the language model to in-domain knowledge and significantly enhancing performance for classes with limited examples. Furthermore, we present a new and valuable dataset called ENZYME, designed for HTC, which comprises articles from PubMed with the goal of predicting Enzyme Commission (EC) numbers. Through extensive experiments on the ENZYME dataset and the widely recognized WOS and NYT datasets, our methodology demonstrates superior performance, surpassing existing approaches while efficiently handling data and mitigating class imbalance. We release our code and dataset here: https://github.com/viditjain99/HiGen. Vidit Jain, Mukund Rungta, Yuchen Zhuang, Yue Yu 0001, Mu Gao, Jeffrey Skolnick, Chao Zhang 0014 |
EACL (1) | 1 |
| 2024 | Enhancing Tail Performance in Extreme Classifiers by Label Variance ReductionabstractExtreme Classification (XC) architectures, which utilize a massive One-vs-All (OvA) classifier layer at the output, have demonstrated remarkable performance on problems with large label sets. Nonetheless, these architectures falter on tail labels with few representative samples. This phenomenon has been attributed to factors such as classifier over-fitting and missing label bias, and solutions involving regularization and loss re-calibration have been developed. This paper explores the impact of label variance - a previously unexamined factor - on the tail performance in extreme classifiers. It also develops a method to systematically reduce label variance in XC by transferring the knowledge from a specialized tail-robust teacher model to the OvA classifiers. For this purpose, it proposes a principled knowledge distillation framework, LEVER, which enhances the tail performance in extreme classifiers with formal guarantees on generalization. Comprehensive experiments are conducted on a diverse set of XC datasets, demonstrating that LEVER can enhance tail performance by around 5\% and 6\% points in PSP and coverage metrics, respectively, when integrated with leading extreme classifiers. Moreover, it establishes a new state-of-the-art when added to the top-performing Renee classifier. Extensive ablations and analyses substantiate the efficacy of our design choices. Another significant contribution is the release of two new XC datasets that are different from and more challenging than the available benchmark datasets, thereby encouraging more rigorous algorithmic evaluation in the future. Code for LEVER is available at: aka.ms/lever. Anirudh Buvanesh, Rahul Chand, Jatin Prakash, Bhawna Paliwal, Mudit Dhawan, Neelabh Madan, Deepesh Hada, Vidit Jain, Sonu Mehta, Yashoteja Prabhu, Ramachandran Ramjee, Manik Varma |
ICLR | 8 |
| 2023 | NGAME: Negative Mining-aware Mini-batching for Extreme ClassificationabstractExtreme Classification (XC) seeks to tag data points with the most relevant subset of labels from an extremely large label set. Performing deep XC with dense, learnt representations for data points and labels has attracted much attention due to its superiority over earlier XC methods that used sparse, hand-crafted features. Negative mining techniques have emerged as a critical component of all deep XC methods, allowing them to scale to millions of labels. However, despite recent advances, training deep XC models with large encoder architectures such as transformers remains challenging. This paper notices that memory overheads of popular negative mining techniques often force mini-batch sizes to remain small and slow training down. In response, this paper introduces NGAME, a light-weight mini-batch creation technique that offers provably accurate in-batch negative samples. This allows training with larger mini-batches offering significantly faster convergence and higher accuracies than existing negative sampling techniques. NGAME was found to be up to 16% more accurate than state-of-the-art methods on a wide array of benchmark datasets for extreme classification, as well as 3% more accurate at retrieving search engine queries in response to a user webpage visit to show personalized ads. In live A/B tests on a popular search engine, NGAME yielded up to 23% gains in click-through-rates. Code for NGAME is available at https://github.com/Extreme-classification/ngame Kunal Dahiya, Nilesh Gupta, Deepak Saini, Akshay Soni, Kushal Dave 0001, Jian Jiao 0007, Gururaj K, Amit Singh 0003, Deepesh Hada, Vidit Jain, Bhawna Paliwal, Anshul Mittal, Sonu Mehta, Ramachandran Ramjee, Sumeet Agarwal, Purushottam Kar, Manik Varma |
WSDM | 12 |
| 2022 | Neural Models for Output-Space Invariance in Combinatorial Problems
Yatin Nandwani, Vidit Jain, Mausam, Parag Singla |
ICLR | 2 |
| 2022 | AENeT: an attention-enabled neural architecture for fake news detection using contextual features
Vidit Jain, Rohit Kumar Kaliyar, Anurag Goswami, Pratik Narang, Yashvardhan Sharma |
Neural Comput. Appl. | 1 |
| 2021 | Exploring Semi-Supervised Learning for Predicting Listener BackchannelsabstractDeveloping human-like conversational agents is a prime area in HCI research and subsumes many tasks. Predicting listener backchannels is one such actively-researched task. While many studies have used different approaches for backchannel prediction, they all have depended on manual annotations for a large dataset. This is a bottleneck impacting the scalability of development. To this end, we propose using semi-supervised techniques to automate the process of identifying backchannels, thereby easing the annotation process. To analyze our identification module’s feasibility, we compared the backchannel prediction models trained on (a) manually-annotated and (b) semi-supervised labels. Quantitative analysis revealed that the proposed semi-supervised approach could attain 95% of the former’s performance. Our user-study findings revealed that almost 60% of the participants found the backchannel responses predicted by the proposed model more natural. Finally, we also analyzed the impact of personality on the type of backchannel signals and validated our findings in the user-study. Vidit Jain, Maitree Leekha, Rajiv Ratn Shah, Jainendra Shukla |
CHI | 1 |
| 2020 | Deep Matrix Factorization on Graphs: Application to Collaborative Filtering
Aanchal Mongia, Vidit Jain, Angshul Majumdar |
ICONIP (4) | 2 |
| 2019 | Deep Latent Factor Model for Predicting Drug Target InteractionsabstractIn drug target interaction (DTI) the interactions of some (a subset) drugs on some (a subset) targets are known. The goal is to predict the interactions of all drugs on all targets. One approach is to formulate this as a matrix completion problem, where the matrix of interactions having drugs along the rows and targets along the columns is partially filled. So far standard matrix completion approaches such as nuclear norm minimization and matrix factorization have been used to address the problem. In this work, we propose a deep matrix factorization approach to improve the prediction results. Experiments have been performed on benchmark databases and comparison carried out with some state-of-the-art algorithms. Empirically our proposed deep method, outperforms all the techniques compared against. Aanchal Mongia, Vidit Jain, Emilie Chouzenoux, Angshul Majumdar |
ICASSP | 2 |
| 2015 | Tracking the Conductance of Rapidly Evolving Topic-SubgraphsabstractMonitoring the formation and evolution of communities in large online social networks such as Twitter is an important problem that has generated considerable interest in both industry and academia. Fundamentally, the problem can be cast as studying evolving sugraphs (each subgraph corresponding to a topical community) on an underlying social graph - with users as nodes and the connection between them as edges. A key metric of interest in this setting is tracking the changes to the conductance of subgraphs induced by edge activations. This metric quantifies how well or poorly connected a subgraph is to the rest of the graph relative to its internal connections. Conductance has been demonstrated to be of great use in many applications, such as identifying bursty topics, tracking the spread of rumors, and so on. However, tracking this simple metric presents a considerable scalability challenge - the underlying social network is large, the number of communities that are active at any moment is large, the rate at which these communities evolve is high, and moreover, we need to track conductance in real-time. We address these challenges in this paper. We propose an in-memory approximation called BloomGraphs to store and update these (possibly overlapping) evolving subgraphs. As the name suggests, we use Bloom filters to represent an approximation of the underlying graph. This representation is compact and computationally efficient to maintain in the presence of updates. This is especially important when we need to simultaneously maintain thousands of evolving subgraphs. BloomGraphs are used in computing and tracking conductance of these subgraphs as edge-activations arrive. BloomGraphs have several desirable properties in the context of this application, including a small memory footprint and efficient updateability. We also demonstrate mathematically that the error incurred in computing conductance is one-sided and that in the case of evolving subgraphs the change in approximate conductance has the same sign as the change in exact conductance in most cases. We validate the effectiveness of BloomGraphs through extensive experimentation on large Twitter graphs and other social networks. Sainyam Galhotra, Amitabha Bagchi, Srikanta J. Bedathur, Maya Ramanath, Vidit Jain |
Proc. VLDB Endow. | 5 |
| 2014 | Min-d-Occur: Ensuring Future Occurrences in Streaming Sets
Vidit Jain, Sainyam Galhotra |
UAI | 1 |
| 2013 | Adapting Classification Cascades to New DomainsabstractClassification cascades have been very effective for object detection. Such a cascade fails to perform well in data domains with variations in appearances that may not be captured in the training examples. This limited generalization severely restricts the domains for which they can be used effectively. A common approach to address this limitation is to train a new cascade of classifiers from scratch for each of the new domains. Building separate detectors for each of the different domains requires huge annotation and computational effort, making it not scalable to a large number of data domains. Here we present an algorithm for quickly adapting a pre-trained cascade of classifiers - using a small number of labeled positive instances from a different yet similar data domain. In our experiments with images of human babies and human-like characters from movies, we demonstrate that the adapted cascade significantly outperforms both of the original cascade and the one trained from scratch using the given training examples. Vidit Jain, Sachin Sudhakar Farfade |
ICCV | 1 |
| 2011 | Online domain adaptation of a pre-trained cascade of classifiersabstractMany classifiers are trained with massive training sets only to be applied at test time on data from a different distribution. How can we rapidly and simply adapt a classifier to a new test distribution, even when we do not have access to the original training data? We present an on-line approach for rapidly adapting a “black box” classifier to a new test data set without retraining the classifier or examining the original optimization criterion. Assuming the original classifier outputs a continuous number for which a threshold gives the class, we reclassify points near the original boundary using a Gaussian process regression scheme. We show how this general procedure can be used in the context of a classifier cascade, demonstrating performance that far exceeds state-of-the-art results in face detection on a standard data set. We also draw connections to work in semi-supervised learning, domain adaptation, and information regularization. Vidit Jain, Erik G. Learned-Miller |
CVPR | 1 |
| 2011 | Learning to re-rank: query-dependent image re-ranking using click dataabstractOur objective is to improve the performance of keyword based image search engines by re-ranking their original results. To this end, we address three limitations of existing search engines in this paper. First, there is no straight-forward, fully automated way of going from textual queries to visual features. Image search engines therefore primarily rely on static and textual features for ranking. Visual features are mainly used for secondary tasks such as finding similar images. Second, image rankers are trained on query-image pairs labeled with relevance judgments determined by human experts. Such labels are well known to be noisy due to various factors including ambiguous queries, unknown user intent and subjectivity in human judgments. This leads to learning a sub-optimal ranker. Finally, a static ranker is typically built to handle disparate user queries. The ranker is therefore unable to adapt its parameters to suit the query at hand which again leads to sub-optimal results. We demonstrate that all of these problems can be mitigated by employing a re-ranking algorithm that leverages aggregate user click data. Vidit Jain, Manik Varma |
WWW | 1 |
| 2008 | Selective hidden random fields: Exploiting domain-specific saliency for event classificationabstractClassifying an event captured in an image is useful for understanding the contents of the image. The captured event provides context to refine models for the presence and appearance of various entities, such as people and objects, in the captured scene. Such contextual processing facilitates the generation of better abstractions and annotations for the image. Consider a typical set of consumer images with sports-related content. These images are taken mostly by amateur photographers, and often at a distance. In the absence of manual annotation or other sources of information such as time and location, typical recognition tasks are formidable on these images. Identifying the sporting event in these images provides a context for further recognition and annotation tasks. We propose to use the domain-specific saliency of the appearances of the playing surfaces, and ignore the noninformative parts of the image such as crowd regions, to discriminate among different sports. To this end, we present a variation of the hidden-state conditional random field that selects a subset of the observed features suitable for classification. The inferred hidden variables in this model represent a selection criteria desirable for the problem domain. For sports-related images, this selection criteria corresponds to the segmentation of the playing surface in the image. We demonstrate the utility of this model on consumer images collected from the Internet. Vidit Jain, Amit Singhal 0001, Jiebo Luo 0001 |
CVPR | 1 |
| 2007 | Unsupervised Joint Alignment of Complex ImagesabstractMany recognition algorithms depend on careful positioning of an object into a canonical pose, so the position of features relative to a fixed coordinate system can be examined. Currently, this positioning is done either manually or by training a class-specialized learning algorithm with samples of the class that have been hand-labeled with parts or poses. In this paper, we describe a novel method to achieve this positioning using poorly aligned examples of a class with no additional labeling. Given a set of unaligned examplars of a class, such as faces, we automatically build an alignment mechanism, without any additional labeling of parts or poses in the data set. Using this alignment mechanism, new members of the class, such as faces resulting from a face detector, can be precisely aligned for the recognition process. Our alignment method improves performance on a face recognition task, both over unaligned images and over images aligned with a face alignment algorithm specifically developed for and trained on hand-labeled face images. We also demonstrate its use on an entirely different class of objects (cars), again without providing any information about parts or pose to the learning algorithm. Gary B. Huang, Vidit Jain, Erik G. Learned-Miller |
ICCV | 2 |
| 2007 | People-LDA: Anchoring Topics to People using Face RecognitionabstractTopic models have recently emerged as powerful tools for modeling topical trends in documents. Often the resulting topics are broad and generic, associating large groups of people and issues that are loosely related. In many cases, it may be desirable to influence the direction in which topic models develop. In this paper, we explore the idea of centering topics around people. In particular, given a large corpus of images featuring collections of people and associated captions, it seems natural to extract topics specifically focussed on each person. What words are most associated with George Bush? Which with Condoleezza Rice? Since people play such an important role in life, it is natural to anchor one topic to each person. In this paper, we present People-LDA, which uses the coherence efface images in news captions to guide the development of topics. In particular, we show how topics can be refined to be more closely related to a single person (like George Bush) rather than describing groups of people in a related area (like politics). To do this we introduce a new graphical model that tightly couples images and captions through a modern face recognizer. In addition to producing topics that are people specific (using images as a guiding force), the model also performs excellent soft clustering efface images, using the language model to boost performance. We present a variety of experiments comparing our method to recent developments in topic modeling and joint image-language modeling, showing that our model has lower perplexity for face identification than competing models and produces more refined topics. Vidit Jain, Erik G. Learned-Miller, Andrew McCallum |
ICCV | 1 |
| 2006 | Discriminative Training of Hyper-feature Models for Object IdentificationabstractObject identification is the task of identifying specific objects belonging to the same class such as cars. We often need to recognize an object that we have only seen a few times. In fact, we often observe only one example of a particular object before we need to recognize it again. Thus we are interested in building a system which can learn to extract distinctive markers from a single example and which can then be used to identify the object in another image as “same ” or “different”. Previous work by Ferencz et al. introduced the notion of hyper-features, which are properties of an image patch that can be used to estimate the utility of the patch in subsequent matching tasks. In this work, we show that hyper-feature based models can be more efficiently estimated using discriminative training techniques. In particular, we describe a new hyper-feature model based upon logistic regression that shows improved performance over previously published techniques. Our approach significantly outperforms Bayesian face recognition that is considered as a standard benchmark for face recognition. 1 Vidit Jain, Andras Ferencz, Erik G. Learned-Miller |
BMVC | 1 |