Puneet Kumar 0003

dblp:09/5254-3 · DBLP profile ↗
← Back
14ranked-venue papers
12as first author
11since 2021 · last 2025
0000-0002-4318-1353ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 6 first-author · 5 since 2021Computer networks · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Gaze-GZ: Generalized Gaze Estimation with Multi-scale Gaze Zone Prediction
abstract
Gaze estimation models often experience significant performance degradation on cross-domain tests. Existing methods enforce the model to concentrate on isolating gaze-pertinent features by filtering out irrelevant ones. This paper proposes an advanced generalized framework for gaze estimation, which employs multi-scale gaze zone prediction as an auxiliary task to improve generalization capabilities. Specifically, each facial image is assigned a discrete zone sub-label, alongside the continuous gaze direction label. In addition, we introduce the triplet loss module and the feature consistency branch to ensure that the extracted features within each zone maintain ordered embeddings and robustness to the environmental variations, respectively. The results from comprehensive experiments reveal that the proposed method outperforms all the state-of-the-art generalized gaze estimation methods. The code will be available at github.
Puneet Kumar 0003
ICASSP2
2025 Multimodal Interpretable Depression Analysis Using Visual, Physiological, Audio and Textual Data
abstract
Motivated by depression's significant impact on global health, this work proposes MultiDepNet, a novel multi-modal interpretable depression detection system integrating visual, physiological, audio, and textual data. Through ded-icated feature extraction methods (MTCNN for video, TS-CAN for physiological, ResNet-18 for audio, and RoBERTa for text modalities) and a strategic fusion of modality-specific networks including CNN-RNN, Transformer, MLP, and ResNet-18, it achieves significant advancements in depression detection. Its performance, evaluated across four benchmark datasets (AVEC 2013, AVEC 2014, DAIC, and E-DAIC), demonstrates average MAE of 5.64, RMSE of 7.15, accuracy of 74.19%, precision of 0.7373, re-call of 0.7378, and F1 of 0.7376. It also implements a Multiviz-based interpretability mechanism that computes each modality's contribution to the model's performance. The results reveal the visual modality to be the most signifi-cant, contributing 37.88% towards depression detection.
Puneet Kumar 0003, Shreshtha Misra, Zhuhong Shao, Balasubramanian Raman
WACV1
2025 Interpretable image emotion recognition: A domain adaptation approach using facial expressions
Puneet Kumar 0003, Balasubramanian Raman
Multim. Tools Appl.1
2025 VISTANet: VIsual Spoken Textual Additive Net for Interpretable Multimodal Emotion Recognition
abstract
This paper proposes a multimodal emotion recognition system, VIsual Spoken Textual Additive Net (VISTANet), to classify emotions reflected by input containing image, speech, and text into discrete classes. A new interpretability technique, K-Average Additive exPlanation (KAAP), has been developed that identifies important visual, spoken, and textual features leading to predicting a particular emotion class. The VISTANet fuses information from image, speech, and text modalities using a hybrid of intermediate and late fusion. It automatically adjusts the weights of their intermediate outputs while computing the weighted average. The KAAP technique computes the contribution of each modality and corresponding features toward predicting a particular emotion class. To mitigate the insufficiency of multimodal emotion datasets labelled with discrete emotion classes, we have constructed the IIT-R MMEmoRec dataset consisting of images, corresponding speech and text, and emotion labels (‘angry,’ ‘happy,’ ‘hate,’ and ‘sad’). The VISTANet has resulted in an overall emotion recognition accuracy of 80.11% on the IIT-R MMEmoRec dataset using visual, spoken, and textual modalities, outperforming single or dual-modality configurations. The code and data can be accessed at github.com/MIntelligence-Group/MMEmoRec.
Puneet Kumar 0003, Sarthak Malik, Balasubramanian Raman
IEEE Trans. Affect. Comput.1
2024 Interpretable multimodal emotion recognition using hybrid fusion of speech and image data
Puneet Kumar 0003, Sarthak Malik, Balasubramanian Raman
Multim. Tools Appl.1
2023 Zero-shot learning based cross-lingual sentiment analysis for sanskrit text with insufficient labeled data
Puneet Kumar 0003, Kshitij Pathania, Balasubramanian Raman
Appl. Intell.1
2023 Affective Feedback Synthesis Towards Multimodal Text and Image Data
abstract
In this article, we have defined a novel task of affective feedback synthesis that generates feedback for input text and corresponding images in a way similar to humans responding to multimodal data. A feedback synthesis system has been proposed and trained using ground-truth human comments along with image–text input. We have also constructed a large-scale dataset consisting of images, text, Twitter user comments, and the number of likes for the comments by crawling news articles through Twitter feeds. The proposed system extracts textual features using a transformer-based textual encoder. The visual features have been extracted using a Faster region-based convolutional neural networks model. The textual and visual features have been concatenated to construct multimodal features that the decoder uses to synthesize the feedback. We have compared the results of the proposed system with baseline models using quantitative and qualitative measures. The synthesized feedbacks have been analyzed using automatic and human evaluation. They have been found to be semantically similar to the ground-truth comments and relevant to the given text–image input.
Puneet Kumar 0003, Gaurav Bhatt, Omkar Ingle, Daksh Goyal, Balasubramanian Raman
ACM Trans. Multim. Comput. Commun. Appl.1
2022 A BERT based dual-channel explainable text emotion recognition system
Puneet Kumar 0003, Balasubramanian Raman
Neural Networks1
2021 Hybrid Fusion Based Approach for Multimodal Emotion Recognition with Insufficient Labeled Data
abstract
In this paper, a deep learning based fusion approach has been proposed to classify the emotions portrayed by image and corresponding text into discrete emotion classes. The proposed method first implements intermediate fusion on image and text inputs and then applies late fusion on image, text, and intermediate fusion’s output. We have also come up with a way to handle the unavailability of labeled multimodal emotional data. We have prepared a new dataset built on Balanced Twitter for Sentiment Analysis dataset (B-T4SA) dataset containing an image, text, and emotion labels, i.e., ‘happy,’ ‘sad,’ ‘hate’ and ‘anger.’ The emotion recognition accuracy of 90.20% has been achieved by the proposed method. Along with multi-class emotion recognition, we’ve also compared the sentiment classification results and found the proposed method to perform better than the benchmark approaches.
Puneet Kumar 0003, Vedanti Khokher, Yukti Gupta, Balasubramanian Raman
ICIP1
2021 Towards the Explainability of Multimodal Speech Emotion Recognition
Puneet Kumar 0003, Vishesh Kaushik, Balasubramanian Raman
Interspeech1
2021 Deep neural network hyper-parameter tuning through twofold genetic approach
Puneet Kumar 0003, Shalini Batra, Balasubramanian Raman
Soft Comput.1
2020 End-to-end Triplet Loss based Emotion Embedding System for Speech Emotion Recognition
abstract
In this paper, an end-to-end neural embedding system based on triplet loss and residual learning has been proposed for speech emotion recognition. The proposed system learns the embeddings from the emotional information of the speech utterances. The learned embeddings are used to recognize the emotions portrayed by given speech samples of various lengths. The proposed system implements Residual Neural Network architecture. It is trained using softmax pretraining and triplet loss function. The weights between the fully connected and embedding layers of the trained network are used to calculate the embedding values. The embedding representations of various emotions are mapped onto a hyperplane, and the angles among them are computed using the cosine similarity. These angles are utilized to classify a new speech sample into its appropriate emotion class. The proposed system has demonstrated 91.67% and 64.44% accuracy while recognizing emotions for RAVDESS and IEMOCAP dataset, respectively.
Puneet Kumar 0003, Sidharth Jain, Balasubramanian Raman, Partha Pratim Roy 0001, Masakazu Iwamura
ICPR1
2020 Fast Griffin Lim based waveform generation strategy for text-to-speech synthesis
Puneet Kumar 0003, Vikas Maddukuri, Nagasai Madamshettib, Kishore KG, Sahit Sai Sriram Kavurub, Balasubramanian Raman, Partha Pratim Roy 0001
Multim. Tools Appl.2
2018 MVO-Based 2-D Path Planning Scheme for Providing Quality of Service in UAV Environment
abstract
The need to develop smart unmanned aerial vehicles (UAVs) which are capable of deciding their trajectories is increasing at a rapid pace. Due to their usage in wide range of applications, such as-military, security, communications, survey mapping, disaster management, etc., the provisioning of end-to-end quality of service (QoS) is a challenging task in UAV environment. Moreover, with limited power, the efficiency of the UAVs can be enhanced if adaptive decisions with respect to their itineraries is considered dynamically. However, most of the solutions reported in the literature are not efficient with respect to QoS preservations for various applications. Motivated by this, several recently proposed meta-heuristic optimization schemes for reactive path planning of UAVs have been explored while designing a UAV path planning problem using multiverse optimizer (MVO). By carrying out the simulations over 1000 iterations, it has been demonstrated that MVO algorithm performs better in majority of the cases with average fitness function value of 0.152 and average execution time of 33.686 s.
Puneet Kumar 0003, Sahil Garg, Shalini Batra, Neeraj Kumar 0001, Ilsun You
IEEE Internet Things J.1