Vinod Pankajakshan

dblp:95/8522 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 A robust JPEG quantization step estimation method for image forensics
Chothmal Kumawat, Vinod Pankajakshan
Signal Process. Image Commun.2
2026 Improving Generalization in Deepfake Detection With Feature-Space Adversarial Consistency
Vinod Pankajakshan
IEEE Signal Process. Lett.2
2025 Audio Source Identification using Delta-Delta MFCC Features
Athar Ali, Vinod Pankajakshan
AVSS2
2025 InfraNet: An Ensemble Approach for Real-time Wildlife Detection using Infrared Thermal Imaging
abstract
Human-wildlife conflict presents significant challenges to both conservation and human safety, necessitating efficient monitoring systems for timely wildlife detection. We introduce InfraNet, an infrared object detection system designed for real-time wildlife monitoring using embedded devices. Our key contributions are (1) a new annotated infrared dataset of elephants, human, and common domestic animals, curated to capture diverse environmental conditions, and (2) an ensemble methodology that combines predictions of multiple preprocessed thermal image versions using a baseline YOLOv8m model without fine-tuning. Experimental results on a set of Elephant dataset show that the proposed ensemble approach significantly increases recall rate from 0.35 to 0.62. Additionally, the ensemble model achieves real-time inference speeds on an NVIDIA Jetson Xavier NX, making it suitable for field deployment.
Dheeraj Dhillon, Vinod Pankajakshan, Parvathi M. S, Sreejith Sajeev, Joby Thomas, Byju C, Rajesh K. R
AVSS2
2025 Boosting Deepfake Video Detection Using Non-Local Attention and Sharpness-Aware Minimization
abstract
With the rapid progress and advancement in artificial intelligence, creating synthetic media and manipulating original data has become an incredibly easy task. This poses significant challenges to the authenticity and security of the digital content. Deepfakes are realistic-looking images or videos created using deep learning techniques with the intention to deceive viewers, spread misinformation, influence public opinion, and facilitate financial fraud. While several deepfake countermeasures prove to be effective, they struggle to generalize when it comes to unseen forgeries. In this paper, we present a robust deepfake detection technique that combines ResNet-50 as a feature extraction backbone with a Non-Local Attention Block (NLAB) to enhance spatial dependencies. To improve generalization, we incorporate a Sharpness-Aware Minimization (SAM) technique, which helps the model converge to flat minima, thereby increasing its robustness to small input perturbations and unseen manipulations. In our experiments, the FaceForensics++ and Celeb-df datasets are employed. When trained on the FaceForensics++, our model attains an AUC of 99.11% on the same dataset and generalizes well to Celeb-df with an AUC of 81.39%, surpassing state-of-the-art deepfake detection techniques. Extensive experiments show the effectiveness of the proposed method in improving the performance and generalization of the deepfake detectors.
Vinod Pankajakshan
AVSS2
2020 Enhancing Perceptual Loss with Adversarial Feature Matching for Super-Resolution
abstract
Single image super-resolution (SISR) is an ill-posed problem with an indeterminate number of valid solutions. Solving this problem with neural networks would require access to extensive experience, either presented as a large training set over natural images or a condensed representation from another pre-trained network. Perceptual loss functions, which belong to the latter category, have achieved breakthrough success in SISR and several other computer vision tasks. While perceptual loss plays a central role in the generation of photo-realistic images, it also produces undesired pattern artifacts in the super-resolved outputs. In this paper, we show that the root cause of these pattern artifacts can be traced back to a mismatch between the pre-training objective of perceptual loss and the super-resolution objective. To address this issue, we propose to augment the existing perceptual loss formulation with a novel content loss function that uses the latent features of a discriminator network to filter the unwanted artifacts across several levels of adversarial similarity. Further, our modification has a stabilizing effect on non-convex optimization in adversarial training. The proposed approach offers notable gains in perceptual quality based on an extensive human evaluation study and a competent reconstruction fidelity when tested on objective evaluation metrics.
Ravi Tej Akella, Shirsendu Sukanta Halder, Arunav Pratap Shandeelya, Vinod Pankajakshan
IJCNN4
2020 A robust JPEG compression detector for image forensics
Chothmal Kumawat, Vinod Pankajakshan
Signal Process. Image Commun.2
2018 Tucker tensor decomposition-based tracking and Gaussian mixture model for anomaly localisation and detection in surveillance videos
abstract
The anomaly detection and localisation (ADL) gains remarkable interest as dealing with the complex surveillance videos for detecting the abnormal behaviour is tedious. The human effort in monitoring and classifying the abnormal object is inaccurate and time‐consuming; therefore, the method is proposed using the Tucker tensor decomposition (TTD) and classification of the objects using Gaussian mixture model (GMM). Initially, the object is detected in the frames for easy recognition using simple background subtraction. The TTD decomposes the tensor as core tensor and factor matrices and the two decomposed tensors are compared using the cosine similarity measure that determines the location of the object in the frame. Finally, the features including shape and speed of the object are extracted that is used for classification using the GMM that follows the maximum posterior probability principle to detect and locate the anomaly in the video. The experimentation for anomaly detection proves that the proposed TTD and TTD‐GMM method attains a higher rate of multiple object tracking precision, accuracy, sensitivity, and specificity at 0.96375, 0.975, 1, and 1, respectively.
Avinash Ratre, Vinod Pankajakshan
IET Comput. Vis.2
2018 A JPEG blocking artifact detector for image forensics
Dinesh Bhardwaj, Vinod Pankajakshan
Signal Process. Image Commun.2
2017 A deep learning architecture for brain tumor segmentation in MRI images
abstract
With the advent of new technologies in the field of medicine, there is rising awareness of biomechanisms, and we are better able to treat ailments than we could earlier. Deep learning has helped a lot in this endeavor. This paper deals with the application of deep learning in brain tumor segmentation. Brain tumors are difficult to segment automatically given the high variability in the shapes and sizes. We propose a novel yet simple fully convolutional network (FCN) which results in competitive performance and faster runtime than state-of-the-art model. Using the database provided for the Brain Tumor Segmentation (BraTS) challenge by the Medical Image Computing and Computer Assisted Intervention (MICCAI) society, we are able to achieve dice scores of 0.83 in the whole tumor region, 0.75 in the core tumor region and 0.72 in the enhancing tumor region, while our method is about 18 times faster than the state-of-the-art.
V. Shreyas, Vinod Pankajakshan
MMSP2
2016 Image Overlay Text Detection Based on JPEG Truncation Error Analysis
abstract
This letter proposes a new algorithm for the detection and localization of overlay text in still images. The algorithm is based on the fact that the high-contrast edges in overlay text boundaries generate truncation error when subjected to JPEG compression. The regions containing high-contrast edges in a given test image are first identified using a discrete cosine transform (DCT) domain technique and a binary truncation error map is generated. The overlay text regions in the truncation error map are generally clustered and the cluster corresponding to a text line appears in a connected form. The initial detection is refined using connected component-based processing. Experimental results on a set of $300$ images taken from news videos show that the proposed method is effective in detection of overlay text with a good accuracy.
Dinesh Bhardwaj, Vinod Pankajakshan
IEEE Signal Process. Lett.2
2010 A subjective study of visibility thresholds for wavelet domain watermarking
abstract
Ensuring watermark invisibility in digital images is a challenging task. Most watermarking techniques empirically adjust a strength parameter in order to reach the best trade-off between invisibility and robustness. A target PSNR value is typically set in order to reach a defined quality level. Some watermarking techniques exploit local activity to increase the watermark strength in some specific images areas (edges, textures). In this work we study the visibility thresholds for wavelet domain multiplicative embedding using a watermarking oriented subjective experiment protocol. Thirty four observers were enrolled for a subjective experiment, and had to adjust the watermark strength in order to best define the visibility threshold for watermarking distortions. Three watermark embedding equations were tested in various wavelet sub-bands, and the optimal equation maximizing robustness was derived for every sub-band.
Florent Autrusseau, Sylvain David, Vinod Pankajakshan
ICIP3
2010 A multi-purpose objective quality metric for image watermarking
abstract
Knowing that the watermarking community use simple statistical quality metrics in order to evaluate the watermarked image quality, the authors have recently proposed a simplified objective quality metric (OQM), called “CPA”, for watermarking applications. The metric used the contrast sensitivity function, along with an adapted error pooling, and proved to perform better than state-of-the-art OQMs. In this work, we intend to improve the performance of the CPA metric. The new metric includes the most important steps of Human Visual System (HVS) based quality metric, namely spatial frequency consideration and masking effects. Besides, this work goes further than classical image quality assessment, and several objective quality metrics will be tested in a watermarking algorithm comparison scenario. We will show that the proposed metric is both able to accurately predict the observers score in a quality assessment task, and is also able to compare watermarking algorithms altogether on a perceptual quality viewpoint.
Vinod Pankajakshan, Florent Autrusseau
ICIP1
2009 Detection of motion-incoherent components in video streams
abstract
Motion coherency has recently been identified as a desirable property for watermarks embedded within video streams in order to withstand temporal frame averaging along the motion axis. Nevertheless, no tool has been proposed to easily evaluate the motion coherency of a given watermarking system. Today, this assessment relies on a computationally expensive procedure, namely, (1) embed a watermark, (2) perform temporal frame averaging, and (3) check for the presence of the watermark. In this article, a novel oracle is designed to detect whether a video stream contains any motion-incoherent component or not. Since such incoherence can be introduced by nonmotion-coherent watermarking algorithms, this tool has proven to be most valuable to distinguish watermarked from nonwatermarked content. The oracle relies on some features extracted from error frames after motion compensation. Experimental results demonstrate the efficiency of the proposed method with uncompressed and compressed video streams.
Vinod Pankajakshan, Gwenaël J. Doërr, Prabin Kumar Bora
IEEE Trans. Inf. Forensics Secur.1