EDBT 2026 Demo / reviewers in the wild / expert
Rohit Kundu
dblp:294/8373
· DBLP profile ↗
14ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0001-8665-8898ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Visibility guided Self-Supervised Occlusion-Resilient Human Pose EstimationabstractOcclusion remains a significant challenge for existing human pose estimation algorithms, often resulting in inaccurate and anatomically implausible predictions. Although recent occlusion-robust methods report strong performance, they typically rely heavily on supervised learning and privileged information, such as multiview data or temporal sequences. Furthermore, these models often fail under domain changes. Domain-adaptive human pose estimation seeks to mitigate this issue; however, when occlusions are present in the target domain, a common occurrence in real-world applications, performance of these algorithms deteriorates significantly. To address these challenges, we propose VisOR, a novel Visibility guided Self-Supervised algorithm for Occlusion-Resilient Human Pose Estimation. VisOR achieves robustness to both domain shifts and occlusions by integrating contextual reasoning with iterative pseudo-label refinement. It mitigates the overfitting to noisy labels from occluded regions via a visibility-driven curriculum learning strategy, which progressively introduces the model to increasingly occluded training samples. Additionally, VisOR is regularized by a learned human pose prior that maintains anatomical plausibility throughout the adaptation process. Recognizing the scarcity of human pose datasets with realistic occlusions, we introduce BOW Blended Occlusions in-the-Wild, a rigorously constructed context-aware synthetic benchmark designed to evaluate the occlusion resilience of human pose estimation algorithms. BOW offers a diverse range of context-aware occlusions across both indoor and outdoor environments, simulating real-world conditions. Through extensive experiments, we demonstrate that VisOR outperforms current state-of-the-art methods by ∼ 7% in challenging occluded human pose estimation benchmarks and provides a baseline performance on BOW, against existing algorithms. Arindam Dutta, Sarosij Bose, Rohit Kundu, Calvin-Khang Ta, Saketh Bachu, Konstantinos Karydis, Amit K. Roy-Chowdhury |
WACV | 3 |
| 2025 | Towards Source-Free Machine UnlearningabstractAs machine learning becomes more pervasive and data privacy regulations evolve, the ability to remove private or copyrighted information from trained models is becoming an increasingly critical requirement. Existing unlearning methods often rely on the assumption of having access to the entire training dataset during the forgetting process. However, this assumption may not hold true in practical scenarios where the original training data may not be accessible, i.e., the source-free setting. To address this challenge, we focus on the source-free unlearning scenario, where an unlearning algorithm must be capable of removing specific data from a trained model without requiring access to the original training dataset. Building on recent work, we present a method that can estimate the Hessian of the unknown remaining training data, a crucial component required for efficient unlearning. Leveraging this estimation technique, our method enables efficient zero-shot unlearning while providing robust theoretical guarantees on the unlearning performance, while maintaining performance on the remaining data. Extensive experiments over a wide range of datasets verify the efficacy of our method. Sk Miraj Ahmed, Umit Yigit Basaran, Dripta S. Raychaudhuri, Arindam Dutta, Rohit Kundu, Fahim Faisal Niloy, Basak Guler, Amit K. Roy-Chowdhury |
CVPR | 5 |
| 2025 | Towards a Universal Synthetic Video Detector: From Face or Background Manipulations to Fully AI-Generated ContentabstractExisting DeepFake detection techniques primarily focus on facial manipulations, such as face-swapping or lip-syncing. However, advancements in text-to-video (T2V) and image-to-video (I2V) generative models now allow fully AI-generated synthetic content and seamless background alterations, challenging face-centric detection methods and demanding more versatile approaches.To address this, we introduce the Universal Network for Identifying Tampered and synthEtic videos (UNITE) model, which, unlike traditional detectors, captures full-frame manipulations. UNITE extends detection capabilities to scenarios without faces, non-human subjects, and complex background modifications. It leverages a transformer-based architecture that processes domain-agnostic features extracted from videos via the SigLIP-So400M foundation model. Given limited datasets encompassing both facial/background alterations and T2V/I2V content, we integrate task-irrelevant data alongside standard DeepFake datasets in training. We further mitigate the model’s tendency to over-focus on faces by incorporating an attention-diversity (AD) loss, which promotes diverse spatial attention across video frames. Combining AD loss with cross-entropy improves detection performance across varied contexts. Comparative evaluations demonstrate that UNITE outperforms state-of-the-art detectors on datasets featuring face/background manipulations and fully synthetic T2V/I2V videos, showcasing its adaptability and generalizable detection capabilities. Rohit Kundu, Vishal Mohanty, Athula Balachandran, Amit K. Roy-Chowdhury |
CVPR | 1 |
| 2024 | Prism refraction search: a novel physics-based metaheuristic algorithm
Rohit Kundu, Soumitri Chattopadhyay, Sayan Nag, Mario A. Navarro, Diego Oliva 0001 |
J. Supercomput. | 1 |
| 2023 | Ideal: Improved Dense Local Contrastive Learning For Semi-Supervised Medical Image SegmentationabstractDue to the scarcity of labeled data, Contrastive Self-Supervised Learning (SSL) frameworks have lately shown great potential in several medical image analysis tasks. However, the existing contrastive mechanisms are sub-optimal for dense pixel-level segmentation tasks due to their inability to mine local features. To this end, we extend the concept of metric learning to the segmentation task, using a dense (dis)similarity learning for pre-training a deep encoder network, and employing a semi-supervised paradigm to fine-tune for the downstream task. Specifically, we propose a simple convolutional projection head for obtaining dense pixel-level features, and a new contrastive loss to utilize these dense projections thereby improving the local representations. A bidirectional consistency regularization mechanism involving two-stream model training is devised for the downstream task. Upon comparison, our IDEAL method outperforms the SoTA methods by fair margins on cardiac MRI segmentation. Our source codes are publicly accessible at: https://github.com/Rohit-Kundu/IDEAL-ICASSP23. Hritam Basak, Soumitri Chattopadhyay, Rohit Kundu, Sayan Nag, Rammohan Mallipeddi |
ICASSP | 3 |
| 2023 | Serf: Towards better training of deep neural networks using log-Softplus ERror activation FunctionabstractActivation functions play a pivotal role in determining the training dynamics and neural network performance. The widely adopted activation function ReLU despite being simple and effective has few disadvantages including the Dying ReLU problem. In order to tackle such problems, we propose a novel activation function called Serf which is self-regularized and non-monotonic in nature. Like Mish, Serf also belongs to the Swish family of functions. Based on several experiments on computer vision (image classification and object detection) and natural language processing (machine translation, sentiment classification and multi-modal entailment) tasks with different state-of-the-art architectures, it is observed that Serf vastly outperforms ReLU (baseline) and other activation functions including both Swish and Mish, with a markedly bigger margin on deeper architectures. Ablation studies further demonstrate that Serf based architectures perform better than those of Swish and Mish in varying scenarios, validating the effectiveness and compatibility of Serf with varying depth, complexity, optimizers, learning rates, batch sizes, initializers and dropout rates. Finally, we investigate the mathematical relation between Swish and Serf, thereby showing the impact of pre-conditioner function ingrained in the first derivative of Serf which provides a regularization effect making gradients smoother and optimization faster. Sayan Nag, Mayukh Bhattacharyya, Anuraag Mukherjee, Rohit Kundu |
WACV | 4 |
| 2023 | Deep features selection through genetic algorithm for cervical pre-cancerous cell classification
Rohit Kundu, Soham Chattopadhyay |
Multim. Tools Appl. | 1 |
| 2022 | Zero-shot Disfluency Detection for Indian LanguagesabstractDisfluencies that appear in the transcriptions from automatic speech recognition systems tend to impair the performance of downstream NLP tasks. Disfluency correction models can help alleviate this problem. However, the unavailability of labeled data in low-resource languages impairs progress. We propose using a pretrained multilingual model, finetuned only on English disfluencies, for zero-shot disfluency detection in Indian languages. We present a detailed pipeline to synthetically generate disfluent text and create evaluation datasets for four Indian languages: Bengali, Hindi, Malayalam, and Marathi. Even in the zero-shot setting, we obtain F1 scores of 75 and higher on five disfluency types across all four languages. We also show the utility of synthetically generated disfluencies by evaluating on real disfluent text in Bengali, Hindi, and Marathi. Finetuning the multilingual model on additional synthetic Hindi disfluent text nearly doubles the number of exact matches and yields a 20-point boost in F1 scores when evaluated on real Hindi disfluent text, compared to training with only English disfluent text. Rohit Kundu, Preethi Jyothi, Pushpak Bhattacharyya |
COLING | 1 |
| 2022 | Doodle It Yourself: Class Incremental Learning by Drawing a Few SketchesabstractThe human visual system is remarkable in learning new visual concepts from just a few examples. This is precisely the goal behind few-shot class incremental learning (FS-CIL), where the emphasis is additionally placed on ensuring the model does not suffer from “forgetting”. In this paper, we push the boundary further for FSCIL by addressing two key questions that bottleneck its ubiquitous application (i) can the model learn from diverse modalities other than just photo (as humans do), and (ii) what if photos are not readily accessible (due to ethical and privacy constraints). Our key innovation lies in advocating the use of sketches as a new modality for class support. The product is a “Doodle It Yourself” (DIY) FSCIL framework where the users can freely sketch a few examples of a novel class for the model to learn to recognise photos of that class. For that, we present a framework that infuses (i) gradient consensus for domain invariant learning, (ii) knowledge distillation for preserving old class information, and (iii) graph attention networks for message passing between old and novel classes. We experimentally show that sketches are better class support than text in the context of FSCIL, echoing findings elsewhere in the sketching literature. Ayan Kumar Bhunia, Gajjala Viswanatha Reddy, Subhadeep Koley, Rohit Kundu, Aneeshan Sain, Tao Xiang 0002, Yi-Zhe Song |
CVPR | 4 |
| 2022 | Pneumonia detection from lung X-ray images using local search aided sine cosine algorithm based deep feature selection methodabstractPneumonia is a major cause of death among children below the age of 5 years, globally. It is especially prevalent in developing and underdeveloped nations where the risk factors for the disease such as unhygienic living conditions, high levels of pollution and overcrowding are higher. Radiological examination (usually X-ray scans) is conducted to detect pneumonia, yet it is prone to subjective variability and can lead to disagreements among different radiologists. To detect traces of pneumonia from X-ray images, a more robust method is therefore required, which can be achieved by using a computer-aided diagnosis (CAD) system. In this study, we develop a two-stage framework, using the combination of deep learning and optimization algorithms, which is both accurate and time-efficient. In its first stage, the proposed framework extracts feature using a customized deep learning model called DenseNet-201 following the concept of transfer learning to cope with the scanty available data. In the second stage, we then reduce the feature dimension using an improved sine cosine algorithm equipped with adaptive beta hill climbing-based local search algorithm. The optimized feature subset is utilized for the classification of “Pneumonia” and “Normal” X-ray images using a support vector machines classifier. Upon an evaluation on a publicly available data set, the proposed method demonstrates the highest accuracy of 98.36% and sensitivity of 98.79% with a feature reduction of 85.55% (74 features selected out of 512), using a five-fold cross-validation scheme. Extensive additional experiments on continuous benchmark functions as well as the CEC-2017 test suite further showcase the superiority and suitability of our proposed approach in application to real-valued optimization problems. The relevant codes for the proposed method can be found in https://github.com/soumitri2001/Pneumonia-Detection-Local-Search-aided-SCA. Soumitri Chattopadhyay, Rohit Kundu, Pawan Kumar Singh 0001, Seyedali Mirjalili, Ram Sarkar |
Int. J. Intell. Syst. | 2 |
| 2022 | ET-NET: an ensemble of transfer learning models for prediction of COVID-19 infection through chest CT-scan images
Rohit Kundu, Pawan Kumar Singh 0001, Massimiliano Ferrara, Ali Ahmadian, Ram Sarkar |
Multim. Tools Appl. | 1 |
| 2022 | An ensemble approach for still image-based human action recognition
Avinandan Banerjee, Sayantan Roy, Rohit Kundu, Pawan Kumar Singh 0001, Vikrant Bhateja, Ram Sarkar |
Neural Comput. Appl. | 3 |
| 2022 | MFSNet: A multi focus segmentation network for skin lesion segmentation
Hritam Basak, Rohit Kundu, Ram Sarkar |
Pattern Recognit. | 2 |
| 2021 | Segmentation of brain MRI using an altruistic Harris Hawks' Optimization algorithm
Rajarshi Bandyopadhyay, Rohit Kundu, Diego Oliva 0001, Ram Sarkar |
Knowl. Based Syst. | 2 |