VLDB 2026 Research / reviewers in the wild / expert
Bhalaji Nagarajan
dblp:244/8191
· DBLP profile ↗
10ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0003-2473-2057ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conjuring Positive Pairs for Efficient Unification of Representation Learning and Image SynthesisabstractWhile representation learning and generative modeling seek to understand visual data, unifying both domains remains unexplored. Recent Unified Self-Supervised Learning (SSL) methods have started to bridge the gap between both paradigms. However, they either extract information from discriminative pretrained models or rely solely on semantic token reconstruction, which requires an external tokenizer during training — introducing a significant computational overhead. In this work, we introduce Sorcen, a novel Unified SSL framework, incorporating a synergic Contrastive-Reconstruction objective. Our novel Contrastive objective, leverages the generative capabilities of Sorcen and eliminates the need for additional image crops or augmentations during training. Sorcen "generates" contrastive positive samples, called Echoes, directly in the semantic token space using the reconstruction objective. This on-the-fly Echo generation, enables Sorcen to operate exclusively on precomputed tokens, eliminating the need for an online tokenizer during training. Sorcen significantly reduces the computational overhead by 60.8% compared to token reconstruction SoTA. Extensive experiments on ImageNet-1k demonstrate that Sorcen outperforms the previous Unified SSL SoTA by 0.4%, 1.48 FID, 1.76%, and 1.53% on linear probing, unconditional image generation, few-shot learning, and transfer learning, respectively. Additionally, Sorcen establishes as a new single-crop MIM SoTA in linear probing and achieves SoTA performance in unconditional image generation, highlighting significant improvements and breakthroughs in Unified SSL models1. Imanol G. Estepa, Jesús M. Rodríguez-de-Vera, Ignacio Sarasua, Bhalaji Nagarajan, Petia Radeva |
WACV | 4 |
| 2026 | Precision at scale: Domain-specific datasets on-demandabstract• Precision at Scale (PaS) automatically creates domain-specific datasets on-demand. • It leverages LLMs, VLMs, and generative models for data collection and curation. • PaS includes a task-agnostic framework for assessing dataset diversity. • Training on PaS datasets outperform model pretraining on large-scale general-domain datasets. • PaS datasets efficiently fine-tune SoTA VLMs in specialized domains. Recent self-supervised learning methods rely on massive general-domain datasets for robust model pretraining. However, these datasets may lack specificity required in specialized domains. Collecting large, supervised datasets to compensate for this limitation is also cumbersome. This raises a key question: Can automatically crafted domain-specific datasets serve as efficient and effective SSL pretrainers, performing comparable to—or even surpassing—much larger state-of-the-art general-domain datasets? To address this challenge, we propose Precision at Scale (PaS) , a novel modular pipeline for automatic creation of domain-specific datasets on-demand. PaS leverages Large Language Models (LLMs) and Vision-Language Models (VLMs) through three distinct phases: Concept Generation, where LLMs identify relevant domain concepts; Image Collection, utilizing VLMs and Generative models to gather appropriate images; Data Curation, ensuring quality and relevance by eliminating unrelated or redundant images. We conduct extensive experiments across three complex domains — food, insects, and birds — proving that PaS datasets compete and often surpass existing domain-specific datasets in diversity, scale, and effectiveness as pretrainers. Models pretrained on PaS datasets outperform those trained on large-scale general-domain datasets (ImageNet-1K) by up to 21 % and surpass same-scale domain-specific datasets by 6.7 % across classification tasks. Notably, despite being an order of magnitude smaller, PaS datasets outperform ImageNet-21K pretraining, with improvements of 3.3 % in fine-tuning and 9.5 % in few-shot learning, and showing superior performance on specialized dense tasks. Furthermore, by efficiently fine-tuning pretrained VLMs like CLIP and SigLIP using low-rank methods, we achieve performance gains (+4.2 % over CLIP) in specialized domains with minimal overhead, demonstrating the versatility of PaS datasets. Jesús M. Rodríguez-de-Vera, Imanol G. Estepa, Ignacio Sarasua, Bhalaji Nagarajan, Petia Radeva |
Pattern Recognit. | 4 |
| 2025 | Adaptive Vision-Language Prompt Learners for Learning with Noisy Labels
Changhui Hu 0004, Bhalaji Nagarajan, Ricardo Marques, Petia Radeva |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | Bayesian DivideMix++ for Enhanced Learning with Noisy LabelsabstractLeveraging inexpensive and human intervention-based annotating methodologies, such as crowdsourcing and web crawling, often leads to datasets with noisy labels. Noisy labels can have a detrimental impact on the performance and generalization of deep neural networks. Robust models that are able to handle and mitigate the effect of these noisy labels are thus essential. In this work, we explore the open challenges of neural network memorization and uncertainty in creating robust learning algorithms with noisy labels. To overcome them, we propose a novel framework called "Bayesian DivideMix++" with two critical components: (i) DivideMix++, to enhance the robustness against memorization and (ii) Monte-Carlo MixMatch, which focuses on improving the effectiveness towards label uncertainty. DivideMix++ improves the pipeline by integrating the warm-up and augmentation pipeline with self-supervised pre-training and dedicated different data augmentations for loss analysis and backpropagation. Monte-Carlo MixMatch leverages uncertainty measurements to mitigate the influence of uncertain samples by reducing their weight in the data augmentation MixMatch step. We validate our proposed pipeline using four datasets encompassing various synthetic and real-world noise settings. We demonstrate the effectiveness and merits of our proposed pipeline using extensive experiments. Bayesian DivideMix++ outperforms the state-of-the-art models by considerable differences in all experiments. Our findings underscore the potential of leveraging these modifications to enhance the performance and generalization of deep neural networks in practical scenarios. Bhalaji Nagarajan, Ricardo Marques, Eduardo Aguilar 0001, Petia Radeva |
Neural Networks | 1 |
| 2024 | Decoding class dynamics in learning with noisy labelsabstractThe creation of large-scale datasets annotated by humans inevitably introduces noisy labels, leading to reduced generalization in deep-learning models. Sample selection-based learning with noisy labels is a recent approach that exhibits promising upbeat performance improvements. The selection of clean samples amongst the noisy samples is an important criterion in the learning process of these models. In this work, we delve deeper into the clean-noise split decision and highlight the aspect that effective demarcation of samples would lead to better performance. We identify the Global Noise Conundrum in the existing models, where the distribution of samples is treated globally. We propose a per-class-based local distribution of samples and demonstrate the effectiveness of this approach in having a better clean-noise split. We validate our proposal on several benchmarks - both real and synthetic, and show substantial improvements over different state-of-the-art algorithms. We further propose a new metric, classiness to extend our analysis and highlight the effectiveness of the proposed method. Source code and instructions to reproduce this paper are available at https://github.com/aldakata/CCLM/ Albert Tatjer, Bhalaji Nagarajan, Ricardo Marques, Petia Radeva |
Pattern Recognit. Lett. | 2 |
| 2023 | All4One: Symbiotic Neighbour Contrastive Learning via Self-Attention and Redundancy ReductionabstractNearest neighbour-based methods have proved to be one of the most successful self-supervised learning (SSL) approaches due to their high generalization capabilities. However, their computational efficiency decreases when more than one neighbour is used. In this paper, we propose a novel contrastive SSL approach, which we call All4One, that reduces the distance between neighbour representations using "centroids" created through a self-attention mechanism. We use a Centroid Contrasting objective along with single Neighbour Contrasting and Feature Contrasting objectives. Centroids help in learning contextual information from multiple neighbours whereas the neighbour contrast enables learning representations directly from the neighbours and the feature contrast allows learning representations unique to the features. This combination enables All4One to outperform popular instance discrimination approaches by more than 1% on linear classification evaluation for popular benchmark datasets and obtains state-of-the-art (SoTA) results. Finally, we show that All4One is robust towards embedding dimensionalities and augmentations, surpassing NNCLR and Barlow Twins by more than 5% on low dimensionality and weak augmentation settings. Source code is available in https://github.com/ImaGonEs/all4one. Imanol G. Estepa, Ignacio Sarasua, Bhalaji Nagarajan, Petia Radeva |
ICCV | 3 |
| 2023 | Deep ensemble-based hard sample mining for food recognitionabstractDeep neural networks represent a compelling technique to tackle complex real-world problems, but are over-parameterized and often suffer from over- or under-confident estimates. Deep ensembles have shown better parameter estimations and often provide reliable uncertainty estimates that contribute to the robustness of the results. In this work, we propose a new metric to identify samples that are hard to classify. Our metric is defined as coincidence score for deep ensembles which measures the agreement of its individual models. The main hypothesis we rely on is that deep learning algorithms learn the low-loss samples better compared to large-loss samples. In order to compensate for this, we use controlled over-sampling on the identified ”hard” samples using proper data augmentation schemes to enable the models to learn those samples better. We validate the proposed metric using two public food datasets on different backbone architectures and show the improvements compared to the conventional deep neural network training using different performance metrics. Bhalaji Nagarajan, Marc Bolaños, Eduardo Aguilar 0001, Petia Radeva |
J. Vis. Commun. Image Represent. | 1 |
| 2022 | Hyper-Spectral Imaging for Overlapping Plastic Flakes SegmentationabstractGiven the hyper-spectral imaging unique potentials in grasping the polymer characteristics of different materials, it is commonly used in sorting procedures. In a practical plastic sorting scenario, multiple plastic flakes may overlap which depending on their characteristics, the overlap can be reflected in their spectral signature. In this work, we use hyper-spectral imaging for the segmentation of three types of plastic flakes and their possible overlapping combinations. We propose an intuitive and simple multi-label encoding approach, bitfield encoding, to account for the overlapping regions. With our experiments, we show that the bitfield encoding improves over the baseline single-label approach and we further demonstrate its potential in predicting multiple labels for overlapping classes even when the model is only trained with non-overlapping classes. Guillem Martinez, Maya Aghaei, Martin Dijkstra, Bhalaji Nagarajan, Femke Jaarsma, Jaap van de Loosdrecht, Petia Radeva, Klaas Dijkstra |
ICIP | 4 |
| 2020 | Uncertainty-Aware Data Augmentation for Food RecognitionabstractFood recognition has recently attracted attention of many researchers. However, high food ambiguity, inter-class variability and intra-class similarity define a real challenge for the Deep learning and Computer Vision algorithms. In order to improve their performance, it is necessary to better understand what the model learns and, from this, to determine the type of data that should be additionally included for being the most beneficial to the training procedure. In this paper, we propose a new data augmentation strategy that estimates and uses the epistemic uncertainty to guide the model training. The method follows an active learning framework, where the new synthetic images are generated from the hard to classify real ones present in the training data based on the epistemic uncertainty. Hence, it allows the food recognition algorithm to focus on difficult images in order to learn their discriminatives features. On the other hand, avoiding data generation from images that do not contribute to the recognition makes it faster and more efficient. We show that the proposed method allows to improve food recognition and provides a better trade-off between micro- and macro-recall measures. Eduardo Aguilar 0001, Bhalaji Nagarajan, Rupali Khatun, Marc Bolaños, Petia Radeva |
ICPR | 2 |
| 2019 | Group Emotion Recognition in Adverse Face DetectionabstractMost techniques for group emotion recognition rely on detection on faces of people and then aggregating the facial information to interpret the group emotion of a given image. However, several faces in the image may be occluded, non-frontal, or indistinguishable, i.e., too many faces in a single image (crowd). This paper focuses on such cases and investigate alternate frameworks which does not involve face detection. The developed frameworks are applied on two datasets – EmotiC and Group Affect Database 3.0 and the results are shown to be competitive with face detection (MTCNN) based approaches. Bhalaji Nagarajan, O. V. Ramana Murthy |
FG | 1 |