VLDB 2026 Research / reviewers in the wild / expert
Snehasis Mukherjee
dblp:94/5766
· DBLP profile ↗
37ranked-venue papers
10as first author
14since 2021 · last 2025
0000-0002-2196-8980ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 9 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Adaptive adam-based optimizers using second-order weight decoupling and gradient-aware weight decay for vision transformer
Boyapati Hemanth Sai, Snehasis Mukherjee, Shiv Ram Dubey |
Mach. Vis. Appl. | 2 |
| 2024 | UnSeGArmaNet: Unsupervised Image Segmentation using Graph Neural Networks with Convolutional ARMA Filters
Kovvuri Sai Gopal Reddy, Bodduluri Saran, A. Mudit Adityaja, Saurabh J. Shigwan, Nitin Kumar 0003, Snehasis Mukherjee |
BMVC | 6 |
| 2024 | Deep Model Compression based on the Training History
S. H. Shabbeer Basha, Mohammad Farazuddin, Viswanath Pulabaigari, Shiv Ram Dubey, Snehasis Mukherjee |
Neurocomputing | 5 |
| 2024 | DFW-PP: dynamic feature weighting-based popularity prediction for social media content
Viswanatha Reddy G, B. S. N. V. Chaitanya, Prathyush P, Sumanth M, Mrinalini C, Dileep Kumar P, Snehasis Mukherjee |
J. Supercomput. | 7 |
| 2023 | Auto CNN classifier based on knowledge transferred from self-supervised model
Jaydeep Kishore, Snehasis Mukherjee |
Appl. Intell. | 2 |
| 2023 | AGLC-GAN: Attention-based global-local cycle-consistent generative adversarial networks for unpaired single image dehazing
R. S. Jaisurya, Snehasis Mukherjee |
Image Vis. Comput. | 2 |
| 2022 | Attention-based Single Image Dehazing Using Improved CycleGANabstractSingle image dehazing is a popular research topic among the researchers in computer vision, machine learning, image processing, and graphics. Most of the recent methods for single image dehazing are based upon supervised learning set up. However, supervised methods require annotation of the data, which often makes the dehazing methods biased towards the manual annotation errors. Unsupervised methods are more likely to produce realistic, clear images. However, fewer efforts are found in the literature for single image dehazing in unsupervised set up. We propose an enhanced CycleGAN architecture for Unpaired single image dehazing, with an attention-based transformer architecture embedded in the generator. The proposed transformer comprises three components: 1) A Feature Attention (FA) block combining channel attention and pixel attention mechanism, 2) A Dynamic feature enhancement block for dynamically capturing the spatial structured features and 3) An adaptive mix-up module to preserve the flow of shallow features from downsampling. Experiments on the benchmark datasets show the efficacy of the proposed method. Codes for this work are available in the link: https://github.com/rsjai47/Attention-Based-CycleDehaze. R. S. Jaisurya, Snehasis Mukherjee |
IJCNN | 2 |
| 2022 | First-person Activity Recognition by Modelling Subject - Action RelevanceabstractThe efficacy of Action Recognition methods depends upon relevance of the action (Verb) with respect to the subject (Noun). Existing methods overlook the suitability of Noun-Verb combination in defining an action. In this work, we propose an algorithm called Reduced Verb Set Generator (RVSGen) to reduce the number of possible verbs related to the actions, based upon the relevance of noun-verb combination. A dual modal fusion model for egocentric activity recognition is proposed here to combine the features extracted from the RGB channels and the Optical flow vectors for an egocentric video to recognize the human activity. Unlike state-of-the-art methods, where the key objects and temporal cues are extracted simultaneously, the proposed model first extracts the spatial features i.e., object information (Noun) from the RGB channels and then relates the Noun with the suitable action (Verb) obtained from motion information (Optical Flow). The verbs are predicted by an ConvLSTM architecture with the help of a modified softmax. The notion behind the modified softmax is to estimate the probability distribution with the reduced verb set obtained from RVSGen and the feature vector obtained from the ConvLSTM. With the help of an end-to-end trained architecture, the noun and verb are predicted which are then concatenated together constituting an action. The experiments are performed on a benchmark dataset. The results show the efficacy of the proposed method compared to the state-of-the-art. The codes related to this work can be found at: https://github.com/mpLogics/EgoAR-RVSGen. Manav Prabhakar, Snehasis Mukherjee |
IJCNN | 2 |
| 2022 | Long-term Spatio-temporal Contrastive Learning framework for Skeleton Action RecognitionabstractRecent years have been witnessing significant developments in research in human action recognition based on skeleton data. The graphical representation of the human skeleton, available with the dataset, provides opportunity to apply Graph Convolutional Networks (GCN), to avail efficient analysis of deep spatial-temporal information from the joint and skeleton structure. Most of the current works in skeleton action recognition use the temporal aspect of the video in short-term sequences, ignoring the long-term information present in the evolving skeleton sequence. The proposed long-term Spatio-temporal Contrastive Learning framework for Skeleton Action Recognition uses an encoder-decoder module. The encoder collects deep global-level (long-term) information from the complete action sequence using efficient self-supervision. The proposed encoder combines knowledge from the temporal domain with high-level information of the relative joint and structure movements of the skeleton. The decoder serves two purposes: predicting the human activity and predicting skeleton structure in the future frames. The decoder primarily uses the high-level encodings from the encoder to anticipate the action. For predicting skeleton structure, we extract an even deeper correlation in the Spatio-temporal domain and merge it with the original frame of the video. We apply a contrastive framework in the frame prediction part so that similar actions have similar predicted skeleton structure. The use of the contrastive framework throughout the proposed model helps achieve exemplary performance while employing a self-supervised aspect to the model. We test our model on the NTU-RGB-D-60 dataset and achieve state-of-the-art performance. The codes related to this work are available at: https://github.com/AnshulRustoai/Long-Term-Spatio-Temporal-Framework. Anshul Rustogi, Snehasis Mukherjee |
IJCNN | 2 |
| 2022 | An information-rich sampling technique over spatio-temporal CNN for classification of human actions in videos
S. H. Shabbeer Basha, Viswanath Pulabaigari, Snehasis Mukherjee |
Multim. Tools Appl. | 3 |
| 2021 | Single image dehazing using improved cycleGAN
B. S. N. V. Chaitanya, Snehasis Mukherjee |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | Learnable spatiotemporal feature pyramid for prediction of future optical flow in videos
Laisha Wadhwa, Snehasis Mukherjee |
Mach. Vis. Appl. | 2 |
| 2021 | AutoFCL: automatically tuning fully connected layers for handling small dataset
S. H. Shabbeer Basha, Sravan Kumar Vinakota, Shiv Ram Dubey, Viswanath Pulabaigari, Snehasis Mukherjee |
Neural Comput. Appl. | 5 |
| 2021 | AutoTune: Automatically Tuning Convolutional Neural Networks for Improved Transfer Learning
S. H. Shabbeer Basha, Sravan Kumar Vinakota, Viswanath Pulabaigari, Snehasis Mukherjee, Shiv Ram Dubey |
Neural Networks | 4 |
| 2020 | Measuring photography aesthetics with deep CNNsabstractIn spite of the recent advancements of deep learning based techniques, automatic photo aesthetic assessment still remains a challenging computer vision task. Existing approaches used to focus on providing a single aesthetic score or category (“good” or “bad”) of photograph, rather than quantifying “goodness” or “badness”. The existing algorithms often ignore the importance of different attributes contributing to the artistic quality of the photograph. To obtain the human‐interpretability of aesthetic score of photo, we advocate learning the aesthetic attributes alongwith the prediction of the general aesthetic score. We propose a multi‐task deep CNN, that collectively learns aesthetic attributes alongwith a general aesthetic score for the photograph. To understand the mathematical representation of the attributes in the proposed model, a visualization technique is proposed using back propagation of gradients. These visualization of attributes correspond to the location of objects in the images in order to find out which part of an image “triggers” the classification outcome, thus providing the insights about the model's understanding of these attributes. This paper proposes an aesthetic feature vector based on the relative foreground position of the object in the image. The proposed aesthetic features outperform the state‐of‐art methods especially for Rule of Thirds attribute. Gajjala Viswanatha Reddy, Snehasis Mukherjee, Mainak Thakur |
IET Image Process. | 2 |
| 2020 | Impact of fully connected layers on performance of convolutional neural networks for image classification
S. H. Shabbeer Basha, Shiv Ram Dubey, Viswanath Pulabaigari, Snehasis Mukherjee |
Neurocomputing | 4 |
| 2020 | Single image dehazing by approximating and eliminating the additional airlight component
Kushal Borkar, Snehasis Mukherjee |
Neurocomputing | 2 |
| 2020 | LDOP: local directional order pattern for robust face retrieval
Shiv Ram Dubey, Snehasis Mukherjee |
Multim. Tools Appl. | 2 |
| 2020 | Blind image quality assessment using a combination of statistical features and CNN
Aravind Babu Jeripothula, Santosh Kumar Velamala, Sunil Kumar Banoth, Snehasis Mukherjee |
Multim. Tools Appl. | 4 |
| 2020 | Human activity recognition in RGB-D videos by dynamic images
Snehasis Mukherjee, Leburu Anvitha, T. Mohana Lahari |
Multim. Tools Appl. | 1 |
| 2020 | Local bit-plane decoded convolutional neural network features for biomedical image retrieval
Shiv Ram Dubey, Swalpa Kumar Roy, Soumendu Chakraborty, Snehasis Mukherjee, Bidyut B. Chaudhuri |
Neural Comput. Appl. | 4 |
| 2020 | diffGrad: An Optimization Method for Convolutional Neural NetworksabstractStochastic gradient descent (SGD) is one of the core techniques behind the success of deep neural networks. The gradient provides information on the direction in which a function has the steepest rate of change. The main problem with basic SGD is to change by equal-sized steps for all parameters, irrespective of the gradient behavior. Hence, an efficient way of deep network optimization is to have adaptive step sizes for each parameter. Recently, several attempts have been made to improve gradient descent methods such as AdaGrad, AdaDelta, RMSProp, and adaptive moment estimation (Adam). These methods rely on the square roots of exponential moving averages of squared past gradients. Thus, these methods do not take advantage of local change in gradients. In this article, a novel optimizer is proposed based on the difference between the present and the immediate past gradient (i.e., diffGrad). In the proposed diffGrad optimization technique, the step size is adjusted for each parameter in such a way that it should have a larger step size for faster gradient changing parameters and a lower step size for lower gradient changing parameters. The convergence analysis is done using the regret bound approach of the online learning framework. In this article, thorough analysis is made over three synthetic complex nonconvex functions. The image categorization experiments are also conducted over the CIFAR10 and CIFAR100 data sets to observe the performance of diffGrad with respect to the state-of-the-art optimizers such as SGDM, AdaGrad, AdaDelta, RMSProp, AMSGrad, and Adam. The residual unit (ResNet)-based convolutional neural network (CNN) architecture is used in the experiments. The experiments show that diffGrad outperforms other optimizers. Also, we show that diffGrad performs uniformly well for training CNN using different activation functions. The source code is made publicly available at https://github.com/shivram1987/diffGrad. Shiv Ram Dubey, Soumendu Chakraborty, Swalpa Kumar Roy, Snehasis Mukherjee, Satish Kumar Singh, Bidyut B. Chaudhuri |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | Spontaneous Facial Micro-Expression Recognition using 3D Spatiotemporal Convolutional Neural NetworksabstractFacial expression recognition in videos is an active area of research in computer vision. However, fake facial expressions are difficult to be recognized even by humans. On the other hand, facial micro-expressions generally represent the actual emotion of a person, as it is a spontaneous reaction expressed through human face. Despite of a few attempts made for recognizing micro-expressions, still the problem is far from being a solved problem, which is depicted by the poor rate of accuracy shown by the state-of-the-art methods. A few CNN based approaches are found in the literature to recognize micro-facial expressions from still images. Whereas, a spontaneous microexpression video contains multiple frames that have to be processed together to encode both spatial and temporal information. This paper proposes two 3D-CNN methods: MicroExpSTCNN and MicroExpFuseNet, for spontaneous facial micro-expression recognition by exploiting the spatiotemporal information in CNN framework. The MicroExpSTCNN considers the full spatial information, whereas the MicroExpFuseNet is based on the 3D-CNN feature fusion of the eyes and mouth regions. The experiments are performed over CAS(ME)2and SMIC microb expression databases. The proposed MicroExpSTCNN model outperforms the state-of-the-art methods. Sai Prasanna Teja Reddy, Surya Teja Karri, Shiv Ram Dubey, Snehasis Mukherjee |
IJCNN | 4 |
| 2019 | A Conditional Adversarial Network for Scene Flow EstimationabstractThe problem of Scene flow estimation in depth videos has been attracting attention of researchers of machine vision, due to its potential application in various areas of robotics. The conventional scene flow estimation methods are difficult to use in real-time applications due to their long computational overhead. We propose a conditional adversarial network SceneFlowGAN for scene flow estimation. The proposed SceneFlowGAN uses loss function at two ends: both the generator and the discriminator. The proposed network is a first attempt to estimate scene flow using generative adversarial networks, and is able to estimate both the optical flow and disparity from the input stereo images simultaneously. The proposed method is experimented on a huge RGB-D benchmark sceneflow estimation dataset. Ravi Kumar Thakur, Snehasis Mukherjee |
RO-MAN | 2 |
| 2018 | RCCNet: An Efficient Convolutional Neural Network for Histological Routine Colon Cancer Nuclei ClassificationabstractEfficient and precise classification of histological cell nuclei is of utmost importance due to its potential applications in the field of medical image analysis. It would facilitate the medical practitioners to better understand and explore various factors for cancer treatment. The classification of histological cell nuclei is a challenging task due to the cellular heterogeneity. This paper proposes an efficient Convolutional Neural Network (CNN) based architecture for classification of histological routine colon cancer nuclei named as RCCNet. The main objective of this network is to keep the CNN model as simple as possible. The proposed RCCNet model consists of 1, 512, 868 learnable parameters which are significantly less compared to the popular CNN models such as AlexNet, CIFAR-VGG, GoogLeNet, and WRN. The experiments are conducted over publicly available routine colon cancer histological dataset “CRCHistoPhenotypes”. The results of the proposed RCCNet model are compared with five state-of-the-art CNN models in terms of the accuracy, weighted average F1 score and training time. The proposed method has achieved a classification accuracy of 80.61% and 0.7887 weighted average F1 score. The proposed RCCNet is more efficient and generalized in terms of the training time and data over-fitting, respectively. S. H. Shabbeer Basha, Soumen Ghosh, Kancharagunta Kishan Babu, Shiv Ram Dubey, Viswanath Pulabaigari, Snehasis Mukherjee |
ICARCV | 6 |
| 2018 | A Multi-Face Challenging Dataset for Robust Face RecognitionabstractFace recognition in images is an active area of interest among the computer vision researchers. However, recognizing human face in an unconstrained environment, is a relatively less-explored area of research. Multiple face recognition in unconstrained environment is a challenging task, due to the variation of view-point, scale, pose, illumination and expression of the face images. Partial occlusion of faces makes the recognition task even more challenging. The contribution of this paper is two-folds: introducing a challenging multi-face dataset (i.e., IIITS_MFace Dataset) for face recognition in unconstrained environment and evaluating the performance of state-of-the-art hand-designed and deep learning based face descriptors on the dataset. The proposed IIITS_MFace dataset contains faces with challenges like pose variation, occlusion, mask, spectacle, expressions, change of illumination, etc. We experiment with several state-of-the-art face descriptors, including recent deep learning based face descriptors like VGGFace, and compare with the existing benchmark face datasets. Results of the experiments clearly show that the difficulty level of the proposed dataset is much higher compared to the benchmark datasets. Shiv Ram Dubey, Snehasis Mukherjee |
ICARCV | 2 |
| 2018 | SceneEDNet: A Deep Learning Approach for Scene Flow EstimationabstractEstimating scene flow in RGB-D videos is attracting much interest of the computer vision researchers, due to its potential applications in robotics. The state-of-the-art techniques for scene flow estimation typically rely on the knowledge of scene structure of the frame and the correspondence between frames. However, with the availability of large RGB-D data captured from depth sensors, learning representations for estimation of scene flow has become possible. This paper introduces a first effort to apply a deep learning method for direct estimation of scene flow by presenting a fully convolutional neural network with an encoder-decoder (ED) architecture. The proposed network SceneEDNet involves estimation of three dimensional motion vectors of all the scene points from sequence of stereo images. The training for direct estimation of scene flow is done using consecutive pairs of stereo images and corresponding scene flow ground truth. Ravi Kumar Thakur, Snehasis Mukherjee |
ICARCV | 2 |
| 2018 | Measuring level of cuteness of baby images: a supervised learning scheme
Pooja Makula, Snehasis Mukherjee |
Multim. Tools Appl. | 3 |
| 2018 | Human action and event recognition using a novel descriptor based on improved dense trajectories
Snehasis Mukherjee, Krit Karan Singh |
Multim. Tools Appl. | 1 |
| 2015 | Human Action Recognition Using Dominant Pose Duplet
Snehasis Mukherjee |
ICVS | 1 |
| 2015 | Human Action Recognition Using Dominant Motion Pattern
Snehasis Mukherjee, Apurbaa Mallik, Dipti Prasad Mukherjee |
ICVS | 1 |
| 2015 | A motion-based approach to detect persons in low-resolution video
Snehasis Mukherjee, Dipti Prasad Mukherjee |
Multim. Tools Appl. | 1 |
| 2014 | Recognizing interactions between human performers by 'Dominating Pose Doublet'
Snehasis Mukherjee, Sujoy Kumar Biswas, Dipti Prasad Mukherjee |
Mach. Vis. Appl. | 1 |
| 2013 | A design-of-experiment based statistical technique for detection of key-frames
Snehasis Mukherjee, Dipti Prasad Mukherjee |
Multim. Tools Appl. | 1 |
| 2011 | Recognizing interaction between human performers using 'key pose doublet'abstractIn this paper, we propose a graph theoretic approach for recognizing interactions between two human performers present in a video clip. We watch primarily the human poses of each performer and derive descriptors that capture the motion patterns of the poses. From an initial dictionary of poses (visual words), we extract key poses (or key words) by ranking the poses on the centrality measure of graph connectivity. We argue that the key poses are graph nodes which share a close semantic relationship (in terms of some suitable edge weight function) with all other pose nodes and hence are said to be the central part of the graph. We apply the same centrality measure on all possible combinations of the key poses of the two performers to select the set of 'key pose doublets' that best represent the corresponding action. The results on standard interaction recognition dataset show the robustness of our approach when compared to the present state of the art method. Snehasis Mukherjee, Sujoy Kumar Biswas, Dipti Prasad Mukherjee |
ACM Multimedia | 1 |
| 2011 | Recognizing Human Action at a Distance in Video by Key PosesabstractIn this paper, we propose a graph theoretic technique for recognizing human actions at a distance in a video by modeling the visual senses associated with poses. The proposed methodology follows a bag-of-word approach that starts with a large vocabulary of poses (visual words) and derives a refined and compact codebook of key poses using centrality measure of graph connectivity. We introduce a “meaningful” threshold on centrality measure that selects key poses for each action type. Our contribution includes a novel pose descriptor based on histogram of oriented optical flow evaluated in a hierarchical fashion on a video frame. This pose descriptor combines both pose information and motion pattern of the human performer into a multidimensional feature vector. We evaluate our methodology on four standard activity-recognition datasets demonstrating the superiority of our method over the state-of-the-art. Snehasis Mukherjee, Sujoy Kumar Biswas, Dipti Prasad Mukherjee |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2010 | Modeling Sense Disambiguation of Human Pose: Recognizing Action at a Distance by Key Poses
Snehasis Mukherjee, Sujoy Kumar Biswas, Dipti Prasad Mukherjee |
ACCV (1) | 1 |