VLDB 2026 Research / reviewers in the wild / expert
Pritam Sarkar
dblp:246/5024
· DBLP profile ↗
13ranked-venue papers
11as first author
10since 2021 · last 2025
0000-0003-4000-3604ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 9 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level AlignmentabstractDespite their significant advancements, Multimodal Large Language Models
(MLLMs) often generate factually inaccurate information, referred to as hallucination.
In this work, we address object hallucinations in MLLMs, where information
is generated about an object not present in the input image. We introduce Data-augmented
Phrase-level Alignment (DPA), a novel loss which can be applied to
instruction-tuned off-the-shelf MLLMs to mitigate hallucinations, while preserving
their general vision-language capabilities. To fine-tune MLLMs with DPA, we first
generate a set of 'hallucinated' and 'correct' response pairs through generative data
augmentation by selectively altering the ground-truth information of the correct
responses at a phrase level. The DPA loss is then used to train MLLMs to reduce
the likelihood of hallucinated phrases compared to the correct ones. Our thorough
evaluation on various benchmarks confirms the effectiveness of DPA in mitigating
hallucination while retaining the out-of-the-box performance of the MLLMs on
general tasks. For instance, MLLMs finetuned with DPA, which we refer to as Hallucination
Attenuated Language and Vision Assistant (HALVA), improve F1 by up
to 13.4% on hallucination visual question-answering and reduce the hallucination
rate by up to 4.2% on image description tasks. Pritam Sarkar, Sayna Ebrahimi, Ali Etemad, Ahmad Beirami, Sercan Ö. Arik, Tomas Pfister |
ICLR | 1 |
| 2025 | Self-alignment of Large Video Language Models with Refined Regularized Preference OptimizationabstractDespite recent advances in Large Video Language Models (LVLMs), they still struggle with fine-grained temporal understanding, hallucinate, and often make simple mistakes on even simple video question-answering tasks, all of which pose significant challenges to their safe and reliable deployment in real-world applications. To address these limitations, we propose a self-alignment framework that enables LVLMs to learn from their own errors. Our proposed framework first obtains a training set of preferred and non-preferred response pairs, where non-preferred responses are generated by incorporating common error patterns that often occur due to inadequate spatio-temporal understanding, spurious correlations between co-occurring concepts, and over-reliance on linguistic cues while neglecting the vision modality, among others. To facilitate self-alignment of LVLMs with the constructed preferred and non-preferred response pairs, we introduce Refined Regularized Preference Optimization (RRPO), a novel preference optimization method that utilizes sub-sequence-level refined rewards and token-wise KL regularization to address the limitations of Direct Preference Optimization (DPO). We demonstrate that RRPO achieves more precise alignment and more stable training compared to DPO. Our experiments and analysis validate the effectiveness of our approach across diverse video tasks, including video hallucination, short- and long-video understanding, and fine-grained temporal reasoning. Pritam Sarkar, Ali Etemad |
NeurIPS | 1 |
| 2024 | XKD: Cross-Modal Knowledge Distillation with Domain Alignment for Video Representation LearningabstractWe present XKD, a novel self-supervised framework to learn meaningful representations from unlabelled videos. XKD is trained with two pseudo objectives. First, masked data reconstruction is performed to learn modality-specific representations from audio and visual streams. Next, self-supervised cross-modal knowledge distillation is performed between the two modalities through a teacher-student setup to learn complementary information. We introduce a novel domain alignment strategy to tackle domain discrepancy between audio and visual modalities enabling effective cross-modal knowledge distillation. Additionally, to develop a general-purpose network capable of handling both audio and visual streams, modality-agnostic variants of XKD are introduced, which use the same pretrained backbone for different audio and visual tasks. Our proposed cross-modal knowledge distillation improves video action classification by 8% to 14% on UCF101, HMDB51, and Kinetics400. Additionally, XKD improves multimodal action classification by 5.5% on Kinetics-Sound. XKD shows state-of-the-art performance in sound classification on ESC50, achieving top-1 accuracy of 96.5%. Pritam Sarkar, Ali Etemad |
AAAI | 1 |
| 2024 | Region-Disentangled Diffusion Model for High-Fidelity PPG-to-ECG TranslationabstractThe high prevalence of cardiovascular diseases (CVDs) calls for accessible and cost-effective continuous cardiac monitoring tools. Despite Electrocardiography (ECG) being the gold standard, continuous monitoring remains a challenge, leading to the exploration of Photoplethysmography (PPG), a promising but more basic alternative available in consumer wearables. This notion has recently spurred interest in translating PPG to ECG signals. In this work, we introduce Region-Disentangled Diffusion Model (RDDM), a novel diffusion model designed to capture the complex temporal dynamics of ECG. Traditional Diffusion models like Denoising Diffusion Probabilistic Models (DDPM) face challenges in capturing such nuances due to the indiscriminate noise addition process across the entire signal. Our proposed RDDM overcomes such limitations by incorporating a novel forward process that selectively adds noise to specific regions of interest (ROI) such as QRS complex in ECG signals, and a reverse process that disentangles the denoising of ROI and non-ROI regions. Quantitative experiments demonstrate that RDDM can generate high-fidelity ECG from PPG in as few as 10 diffusion steps, making it highly effective and computationally efficient. Additionally, to rigorously validate the usefulness of the generated ECG signals, we introduce CardioBench, a comprehensive evaluation benchmark for a variety of cardiac-related tasks including heart rate and blood pressure estimation, stress classification, and the detection of atrial fibrillation and diabetes. Our thorough experiments show that RDDM achieves state-of-the-art performance on CardioBench. To the best of our knowledge, RDDM is the first diffusion model for cross-modal signal-to-signal translation in the bio-signal domain. Debaditya Shome, Pritam Sarkar, Ali Etemad |
AAAI | 2 |
| 2023 | Self-Supervised Audio-Visual Representation Learning with Relaxed Cross-Modal SynchronicityabstractWe present CrissCross, a self-supervised framework for learning audio-visual representations. A novel notion is introduced in our framework whereby in addition to learning the intra-modal and standard 'synchronous' cross-modal relations, CrissCross also learns 'asynchronous' cross-modal relationships. We perform in-depth studies showing that by relaxing the temporal synchronicity between the audio and visual modalities, the network learns strong generalized representations useful for a variety of downstream tasks. To pretrain our proposed solution, we use 3 different datasets with varying sizes, Kinetics-Sound, Kinetics400, and AudioSet. The learned representations are evaluated on a number of downstream tasks namely action recognition, sound classification, and action retrieval. Our experiments show that CrissCross either outperforms or achieves performances on par with the current state-of-the-art self-supervised methods on action recognition and action retrieval with UCF101 and HMDB51, as well as sound classification with ESC50 and DCASE. Moreover, CrissCross outperforms fully-supervised pretraining while pretrained on Kinetics-Sound. Pritam Sarkar, Ali Etemad |
AAAI | 1 |
| 2023 | AVCAffe: A Large Scale Audio-Visual Dataset of Cognitive Load and Affect for Remote WorkabstractWe introduce AVCAffe, the first Audio-Visual dataset consisting of Cognitive load and Affect attributes. We record AVCAffe by simulating remote work scenarios over a video-conferencing platform, where subjects collaborate to complete a number of cognitively engaging tasks. AVCAffe is the largest originally collected (not collected from the Internet) affective dataset in English language. We recruit 106 participants from 18 different countries of origin, spanning an age range of 18 to 57 years old, with a balanced male-female ratio. AVCAffe comprises a total of 108 hours of video, equivalent to more than 58,000 clips along with task-based self-reported ground truth labels for arousal, valence, and cognitive load attributes such as mental demand, temporal demand, effort, and a few others. We believe AVCAffe would be a challenging benchmark for the deep learning research community given the inherent difficulty of classifying affect and cognitive load in particular. Moreover, our dataset fills an existing timely gap by facilitating the creation of learning systems for better self-management of remote work meetings, and further study of hypotheses regarding the impact of remote work on cognitive load and affective states. Pritam Sarkar, Aaron Posen, Ali Etemad |
AAAI | 1 |
| 2023 | Uncovering the Hidden Dynamics of Video Self-supervised Learning under Distribution ShiftsabstractVideo self-supervised learning (VSSL) has made significant progress in recent years. However, the exact behavior and dynamics of these models under different forms of distribution shift are not yet known. In this paper, we comprehensively study the behavior of six popular self-supervised methods (v-SimCLR, v-MoCo, v-BYOL, v-SimSiam, v-DINO, v-MAE) in response to various forms of natural distribution shift, i.e., (i) context shift, (ii) viewpoint shift, (iii) actor shift, (iv) source shift, (v) generalizability to unknown classes (zero-shot), and (vi) open-set recognition. To perform this extensive study, we carefully craft a test bed consisting of 17 in-distribution and out-of-distribution benchmark pairs using available public datasets and a series of evaluation protocols to stress-test the different methods under the intended shifts. Our study uncovers a series of intriguing findings and interesting behaviors of VSSL methods. For instance, we observe that while video models generally struggle with context shifts, v-MAE and supervised learning exhibit more robustness. Moreover, our study shows that v-MAE is a strong temporal learner, whereas contrastive methods, v-SimCLR and v-MoCo, exhibit strong performances against viewpoint shifts. When studying the notion of open-set recognition, we notice a trade-off between closed-set and open-set recognition performance if the pretrained VSSL encoders are used without finetuning. We hope that our work will contribute to the development of robust video representation learning frameworks for various real-world scenarios. The project page and code are available at: https://pritamqu.github.io/OOD-VSSL. Pritam Sarkar, Ahmad Beirami, Ali Etemad |
NeurIPS | 1 |
| 2022 | Self-Supervised ECG Representation Learning for Emotion RecognitionabstractWe exploit a self-supervised deep multi-task learning framework for electrocardiogram (ECG) -based emotion recognition. The proposed solution consists of two stages of learninga)learning ECG representations andb)learning to classify emotions. ECG representations are learned by a signal transformation recognition network. The network learns high-level abstract representations from unlabeled ECG data. Six different signal transformations are applied to the ECG signals, and transformation recognition is performed as pretext tasks. Training the model on pretext tasks helps the network learn spatiotemporal representations that generalize well across different datasets and different emotion categories. We transfer the weights of the self-supervised network to an emotion recognition network, where the convolutional layers are kept frozen and the dense layers are trained with labelled ECG data. We show that the proposed solution considerably improves the performance compared to a network trained using fully-supervised learning. New state-of-the-art results are set in classification of arousal, valence, affective states, and stress for the four utilized datasets. Extensive experiments are performed, providing interesting insights into the impact of using a multi-task self-supervised structure instead of a single-task model, as well as the optimum level of difficulty required for the pretext self-supervised tasks. Pritam Sarkar, Ali Etemad |
IEEE Trans. Affect. Comput. | 1 |
| 2021 | CardioGAN: Attentive Generative Adversarial Network with Dual Discriminators for Synthesis of ECG from PPGabstractElectrocardiogram (ECG) is the electrical measurement of cardiac activity, whereas Photoplethysmogram (PPG) is the optical measurement of volumetric changes in blood circulation. While both signals are used for heart rate monitoring, from a medical perspective, ECG is more useful as it carries additional cardiac information. Despite many attempts toward incorporating ECG sensing in smartwatches or similar wearable devices for continuous and reliable cardiac monitoring, PPG sensors are the main feasible sensing solution available. In order to tackle this problem, we propose CardioGAN, an adversarial model which takes PPG as input and generates ECG as output. The proposed network utilizes an attention-based generator to learn local salient features, as well as dual discriminators to preserve the integrity of generated data in both time and frequency domains. Our experiments show that the ECG generated by CardioGAN provides more reliable heart rate measurements compared to the original input PPG, reducing the error from 9.74 beats per minute (measured from the PPG) to 2.89 (measured from the generated ECG). Pritam Sarkar, Ali Etemad |
AAAI | 1 |
| 2021 | Happy Driver: Investigating the Effect of Mood on Preferred Style of Driving in Self-Driving CarsabstractSelf-driving cars are around the corner, yet little is known about how users of self-driving cars will react to the car’s driving style, and whether the driver’s mood affects their driving style preference. This paper explores the impact of users’ mood on driving style preference in self-driving cars. An experiment was conducted online (N=182) to investigate participants’ preference for three driving styles (conservative, moderate, aggressive) under three induced moods (calm, neutral, excited). Measures of arousal, valence, and driving satisfaction were recorded. Overall, participants scored the aggressive driving style lowest, irrespective of driver mood. Participants’ mood impacted preference, where a mismatch between driving style and mood induced prior to the driving style predicted lower driver satisfaction scores. We conclude with the design recommendation that driving styles in self-driving cars should not be overly aggressive, and drivers’ mood should be taken into consideration when designing driving styles. Rachel Phinnemore, Gabriele Cimolino, Pritam Sarkar, Ali Etemad, T. C. Nicholas Graham |
HAI | 3 |
| 2020 | Self-Supervised Learning for ECG-Based Emotion RecognitionabstractWe present an electrocardiogram (ECG) -based emotion recognition system using self-supervised learning. Our proposed architecture consists of two main networks, a signal transformation recognition network and an emotion recognition network. First, unlabelled data are used to successfully train the former network to detect specific pre-determined signal transformations in the self-supervised learning step. Next, the weights of the convolutional layers of this network are transferred to the emotion recognition network, and two dense layers are trained in order to classify arousal and valence scores. We show that our self-supervised approach helps the model learn the ECG feature manifold required for emotion recognition, performing equal or better than the fully-supervised version of the model. Our proposed method outperforms the state-of-the-art in ECG-based emotion recognition with two publicly available datasets, SWELL and AMIGOS. Further analysis highlights the advantage of our self-supervised approach in requiring significantly less data to achieve acceptable results. Pritam Sarkar, Ali Etemad |
ICASSP | 1 |
| 2019 | Classification of Cognitive Load and Expertise for Adaptive Simulation using Deep Multitask LearningabstractSimulations are a pedagogical means of enabling a risk-free way for healthcare practitioners to learn, maintain, or enhance their knowledge and skills. Such simulations should provide an optimum amount of cognitive load to the learner and be tailored to their levels of expertise. However, most current simulations are a one-type-fits-all tool used to train different learners regardless of their existing skills, expertise, and ability to handle cognitive load. To address this problem, we propose an end-to-end framework for a trauma simulation that actively classifies a participant's level of cognitive load and expertise for the development of a dynamically adaptive simulation. To facilitate this solution, trauma simulations were developed for the collection of electrocardiogram (ECG) signals of both novice and expert practitioners. A multitask deep neural network was developed to utilize this data and classify high and low cognitive load, as well as expert and novice participants. A leave-one-subject-out (LOSO) validation was used to evaluate the effectiveness of our model, achieving an accuracy of 89.4% and 96.6% for classification of cognitive load and expertise, respectively. Pritam Sarkar, Kyle Ross, Aaron J. Ruberto, Dirk Rodenburg, Paul Hungler, Ali Etemad |
ACII | 1 |
| 2019 | Computer-Aided Diagnosis using Class-Weighted Deep Neural NetworkabstractComputer-aided diagnosis has become a major focal point of Artificial Intelligence. Interpreting medical images is often time-consuming and requires significant human expertise. Hence, there is an increasing demand to use machine learning techniques to correctly classify different medical images captured by mammography, CT scans, and MRI among others. This paper presents a deep learning method for computer-aided differential diagnosis of benign and malignant breast cancer tumors by avoiding potential errors caused by poor feature selection as well as class imbalances in the dataset. We design, develop and test an end-to-end convolutional neural network architecture for two different breast cancer datasets of fine needle aspiration biopsy samples, and show that our network outperforms the state of the art. Furthermore, we have introduced a loss coefficient which can be adjusted to fine-tune the performance of our network. The proposed method can be used to support oncologists in the detection of breast cancer with high confidence. Pritam Sarkar, Vandad Davoodnia, Ali Etemad |
ICMLA | 1 |