Prithwijit Guha

dblp:29/4092 · DBLP profile ↗
← Back
38ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0003-2885-0026ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 4 first-author · 6 since 2021Systems, architecture and hardware · 4 · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CAFACLite: Condition Aware Face Anchor Classification for Face Detection with Lightweight Networks
Yogesh Aggarwal, Prithwijit Guha
ICPR (15)2
2026 Vrittanta-EN: A Benchmark Dataset for Event Trigger Detection and Classification Advancing Event Understanding in English Narrative Discourse
Chaitanya Kirti, Ashish Anand, Prithwijit Guha
LREC3
2026 Vrittanta-AS: Dataset Development and Benchmarking for Event Trigger Detection and Classification in Assamese
Chaitanya Kirti, Dhrubajyoti Pathak, Ashish Anand, Prithwijit Guha
LREC4
2026 Image captioning in low resource assamese language with semantic information prior and spatially encoded transformer model
Pankaj Choudhury, Sidharth Nair, Prithwijit Guha, Sukumar Nandi
Expert Syst. Appl.3
2025 Designing Customized Lightweight Backbones for the Face Detection Task
abstract
Recent works in lightweight face detectors have mostly used pretrained backbone networks (from MobileNet or ShuffleNet series) for low latency in operation and contributed toward novel loss functions and/or efficient training strategies. In contrast, only a few works have proposed customized backbones. In this context, this work proposes the design of customized lightweight backbone networks (BBLite series) for the face detection task. This proposal emphasizes on designing a series of lightweight backbones while using the standard loss functions associated with face detection in the RetinaFace framework. The BBLite series contains backbone networks with computations ranging from 0.37 GFLOPs (BBLiteV1) to 5.72 GFLOPs (BBLiteV4ax4) and parameters ranging from 0.088M (BBLiteV4c) to 2.52M (BBLiteV4ax4). The resulting lightweight face detectors (LWFD) using different backbone networks from BBLite series are observed to achieve mAP ranging from 85.78% (BBLiteV4c-LWFD, 0.152M, 0.553 GFLOPs) to 91.95% (BBLiteV4x4-LWFD, 2.59M, 4.98 GFLOPs) on the WIDER FACE validation data subset. The BBLite based lightweight face detectors provide competitive performance (in terms of mAP, number of parameters and floating point computations) against 12 state-of-art networks on the WIDER FACE validation data-subset, FDDB and MAFA datasets.
Yogesh Aggarwal, Prithwijit Guha
IJCNN2
2025 Exploring Semantic Attributes for Image Caption Synthesis in Low-Resource Assamese Language
abstract
Research on image caption generation has predominantly focused on resource-rich languages like English, leaving resource-poor languages (like Assamese and several others) largely understudied. In this context, this paper leverages both visual and semantic attribute based features for generating captions in Assamese language. Semantic attributes refer to the significant words that represent higher-level knowledge about the image content. This work contributes through the effective use of features derived from semantic words in low resource Assamese language. The second contribution is the proposal of a Visual-Semantic Self-Attention (VSSA) module for the combination of features derived from images and semantic attributes. The VSSA module enables the image captioning model to dynamically attend to relevant regions of the image as well as the important semantic attributes, thereby leading to more contextually relevant and linguistically accurate Assamese captions. Moreover, the VSSA module is incorporated into a Transformer model to leverage the stacked attention for performance improvement. The model is trained by using both cross-entropy loss optimization and reinforcement learning approach. The effectiveness of the proposed model is evaluated through both qualitative and quantitative analyses (using BLEU-n and CIDEr metrics). The proposed model shows significant performance improvement in Assamese caption synthesis compared to previous methods, achieving 93.7% CIDEr score on the COCO-Assamese Caption (COCO-AC) dataset.
Pankaj Choudhury, Prithwijit Guha, Sukumar Nandi
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2024 Visual Question Answering with Cascade of Self- and Co-Attention Blocks
Aakansha Mishra, Ashish Anand, Prithwijit Guha
ICPR (19)3
2024 Exploration of Speech and Music Information for Movie Genre Classification
abstract
Movie genre prediction from trailers is mostly attempted in a multi-modal manner. However, the characteristics of movie trailer audio indicate that this modality alone might be highly effective in genre prediction. Movie trailer audio predominantly consists of speech and music signals in isolation or overlapping conditions. This work hypothesizes that the genre labels of movie trailers might relate to the composition of their audio component. In this regard, speech-music confidence sequences for the trailer audio are used as a feature. In addition, two other features previously proposed for discriminating speech-music are also adopted in the current task. This work proposes a time and channel Attention Convolutional Neural Network (ACNN) classifier for the genre classification task. The convolutional layers in ACNN learn the spatial relationships in the input features. The time and channel attention layers learn to focus on crucial timesteps and CNN kernel outputs, respectively. The Moviescope dataset is used to perform the experiments, and two audio-based baseline methods are employed to benchmark this work. The proposed feature set with the ACNN classifier improves the genre classification performance over the baselines. Moreover, decent generalization performance is obtained for genre prediction of movies with different cultural influences (EmoGDB).
Mrinmoy Bhattacharjee, S. R. Mahadeva Prasanna, Prithwijit Guha
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Image Caption Synthesis for Low Resource Assamese Language using Bi-LSTM with Bilinear Attention
Pankaj Choudhury, Prithwijit Guha, Sukumar Nandi
PACLIC2
2023 Clean vs. Overlapped Speech-Music Detection Using Harmonic-Percussive Features and Multi-Task Learning
abstract
Detection of speech and music signals in isolated and overlapped conditions is an essential preprocessing step for many audio applications. Speech signals have wavy and continuous harmonics, while music signals exhibit horizontally linear and discontinuous harmonic patterns. Music signals also contain more percussive components than speech signals, manifested as vertical striations in the spectrograms. In case of speech music overlap, it might be challenging for automatic feature learning systems to extract class-specific horizontal and vertical striations from the combined spectrogram representation. A pre-processing step of separating the harmonic and percussive components before training might aid the classifier. Thus, this work proposes the use of harmonic-percussive source separation method to generate features for better detection of speech and music signals. Additionally, this work also explores the traditional and cascaded-information multi-task learning (MTL) frameworks to design better classifiers. MTL framework aids the training of the main task by employing simultaneous learning of several related auxiliary tasks. Results have been reported both on synthetically generated speech music overlapped signals and real recordings. Four state-of-the-art approaches are used for performance comparison. Experiments show that harmonic and percussive decomposition of spectrograms perform better as features. Moreover, the MTL-framework based classifiers further improve performances.
Mrinmoy Bhattacharjee, S. R. Mahadeva Prasanna, Prithwijit Guha
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 GAUR: Genetic Algorithm based Unlocking of Register Transfer Level Locking
abstract
Logic locking is a technique for the protection of hardware intellectual property (IP) from malicious entities like piracy, overproduction, reverse engineering, etc. The register transfer level (RTL) locking performs the locking on RTL description for protection of the IP even from the early design cycle. TAO [12] is such a locking scheme that employs locking during the high-level synthesis (HLS) process. In this paper, we evaluate the unlocking capability of the genetic algorithm (GA) by performing attacks on the RTLs locked using TAO based technique. We demonstrate the ability of GA to unlock TAO generated RTLs in seconds. Our GA based attack is faster as compared to the Satisfiability Modulo Theories (SMT) based attack [9]. The GA based method also converges well in most of the cases as shown in the experimental results.
Gagan Gayari, Chandan Karfa, Prithwijit Guha
ACM Great Lakes Symposium on VLSI3
2022 Hardware Implementation of Low Complexity High-speed Perceptron Block
abstract
Perceptron is the basic computation unit of neural network architectures. This work proposes a resource-efficient and fast hardware for perceptron and Multi-layer Perceptron (MLP) network. The inner product computation unit and activation function unit is designed using Offset Binary Coding (OBC) and Co-ordinate Rotation Digital Computer (CORDIC) respectively. The proposed hardware is implemented on Field Programmable Logic Array (FPGA) and synthesized on 65 nm Application Specific Integrated Chips (ASIC). It achieved a speed-up of at least $21 \times$ as compared to software. The total area and power consumed was $0.085662 mm^{2}$ and 2.28 mW respectively @200 MHz.
Rituparna Choudhury, Shaik Rafi Ahamed, Prithwijit Guha
ISCAS3
2022 Only overlay text: novel features for TV news broadcast video segmentation
Raghvendra Kannao, Prithwijit Guha, Bidyut B. Chaudhuri
Multim. Tools Appl.2
2022 Speech/music classification using phase-based and magnitude-based features
Mrinmoy Bhattacharjee, S. R. Mahadeva Prasanna, Prithwijit Guha
Speech Commun.3
2021 Automatic Detection of Shouted Speech Segments in Indian News Debates
Shikha Baghel, Mrinmoy Bhattacharjee, S. R. Mahadeva Prasanna, Prithwijit Guha
Interspeech4
2021 Training Accelerator for Two Means Decision Tree
abstract
Decision trees (DTs) are profusely used in machine learning (ML) applications on account of their fast execution and high interpretability. As DT training is time-consuming, in this brief, we proposed a hardware training accelerator to speedup the training process. The proposed training accelerator is implemented on the field-programmable gate array (FPGA) having a maximum operating frequency of 62 MHz. The proposed architecture uses a combination of parallel execution for training time reduction and pipelined execution to minimize resource consumption. For a given design, the proposed hardware implementation is found to be at least 14× faster than the C-based software implementation. Moreover, the proposed architecture can be easily retrained for the next set of data using a single RESET signal. This on-the-go training makes the hardware versatile for any kind of application.
Rituparna Choudhury, Shaik Rafi Ahamed, Prithwijit Guha
IEEE Trans. Very Large Scale Integr. Syst.3
2020 Siamese Fully Convolutional Tracker with Motion Correction
abstract
Visual tracking algorithms use cues like appearance, structure, motion etc. for locating an object in a video. We propose an ensemble tracker with two components. First, a Siamese tracker that learns object appearance from a static image. Second, motion information obtained from consecutive frames using a flow estimation network. The motion information is used to correct the predictions obtained by the appearance based tracking component of the ensemble. Complementary nature of the two components (appearance and motion) lead to performance improvement as observed in experiments performed on VOT2018 and VOT2019 datasets.
Mathew Francis, Prithwijit Guha
ICPR2
2020 Multi-stage Attention based Visual Question Answering
abstract
Recent developments in the field of Visual Question Answering (VQA) have witnessed promising improvements in performance through contributions in attention based networks. Most such approaches have focused on unidirectional attention that leverage over attention from textual domain (question) on visual space. These approaches mostly focused on learning high-quality attention in the visual space. In contrast, this work proposes an alternating bi-directional attention framework. First, a question to image attention helps to learn the robust visual space embedding, and second, an image to question attention helps to improve the question embedding. This attention mechanism is realized in an alternating fashion i.e. question-to-image followed by image-to-question and is repeated for maximizing performance. We believe that this process of alternating attention generation helps both the modalities and leads to better representations for the VQA task. This proposal is benchmark on TDIUC dataset and against state-of-art approaches. Our ablation analysis shows that alternate attention is the key to achieve high performance in VOA.
Aakansha Mishra, Ashish Anand, Prithwijit Guha
ICPR3
2020 CQ-VQA: Visual Question Answering on Categorized Questions
abstract
This paper proposes CQ-VQA, a novel two-level hierarchical but end-to-end model to solve the task of visual question answering (VQA). The first level of CQ-VQA, referred to as Question Categorizer (QC), classifies questions to reduce the potential answer search space. The QC uses attended and fused features of the input question and image. The second level, referred to as Answer Predictor (AP), comprises of a set of distinct classifiers corresponding to each question category. Depending on the question category predicted by QC, only one of the classifiers of AP remains active. The loss functions of QC and AP are aggregated together to make it an end-to-end model. The proposed model (CQ-VQA) is evaluated on the TDIUC dataset and is benchmarked against state-of-the-art approaches. Results indicate a competitive or better performance of CQ-VQA.
Aakansha Mishra, Ashish Anand, Prithwijit Guha
IJCNN3
2020 A system for semantic segmentation of TV news broadcast videos
Raghvendra Kannao, Prithwijit Guha
Multim. Tools Appl.2
2020 Speech/Music Classification Using Features From Spectral Peaks
abstract
Spectrograms of speech and music contain distinct striation patterns. Traditional features represent various properties of the audio signal but do not necessarily capture such patterns. This work proposes to model such spectrogram patterns using a novel Spectral Peak Tracking (SPT) approach. Two novel time-frequency features for speech vs. music classification are proposed. The proposed features are extracted in two stages. First, SPT is performed to track a preset number of highest amplitude spectral peaks in an audio interval. In the second stage, the location and amplitudes of these peak traces are used to compute the proposed feature sets. The first feature involves the computation of mean and standard deviation of peak traces. The second feature is obtained as averaged component posterior probability vectors of Gaussian mixture models learned on the peak traces. Speech vs. music classification is performed by training various binary classifiers on these proposed features. Three standard datasets are used to evaluate the efficiency of the proposed features for speech/music classification. The proposed features are benchmarked against five baseline approaches. Finally, the best-proposed feature is combined with two contemporary deep-learning based features to show that such combinations can lead to more robust speech vs. music classification systems.
Mrinmoy Bhattacharjee, S. R. Mahadeva Prasanna, Prithwijit Guha
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Segmenting with style: detecting program and story boundaries in TV news broadcast videos
Raghvendra Kannao, Prithwijit Guha
Multim. Tools Appl.2
2018 Visual Tracking with Breeding Fireflies using Brightness from Background-Foreground Information
abstract
Visual target tracking involves object localization in image sequences. This is achieved by optimizing image feature similarity based objective functions in object state space. Meta-heuristic algorithms have shown promising results in solving hard optimization problems where gradients are not available. This motivated us to use Firefly algorithms in visual object tracking. The object state is represented by its bounding box parameters and the target is modeled by its color distribution. This work has two significant contributions. First, we propose a hybrid firefly algorithm where genetic operations are performed using Real-coded Genetic Algorithm(RGA). Here, the crossover operation is modified by incorporating parent velocity information. Second, the firefly brightness is computed from both foreground and background information (as opposed to only foreground). This helps in handling scale implosion and explosion problems. The proposed approach is benchmarked on challenging sequences from VOT2014 dataset and is compared against other baseline trackers and metaheuristic algorithms.
Pranay Kate, Mathew Francis, Prithwijit Guha
ICPR3
2017 Success based locally weighted Multiple Kernel combination
Raghvendra Kannao, Prithwijit Guha
Pattern Recognit.2
2016 Story segmentation in TV news broadcast
abstract
Segmentation of TV news broadcast into semantically meaningful stories is an essential pre-requisite for a wide range of video analytics applications. In this work we have introduced a hybrid approach for news story segmentation based on conditional random fields (CRFs). The story boundary detection problem is converted into a shot classification problem by classifying video shots into either of the four categories. These are start shot, end shot and middle shots of a story or single shot story. To achieve this classification, we have introduced two new features. These are overlay text based semantic similarity and grid-wise edge orientation histogram. The first feature measures the semantic similarity between video shots by linking them through a set of web news articles. We use overlay text with their relevance as weight to link a set of articles with the video shots. The second feature captures the variations in presentation formats. The CRF model effectively combines these two features to model the news stories. Experimental results on approximately 50 hours of news videos demonstrates the efficiency of the proposed features. We were able to achieve an F1 score of 81% with our proposed features.
Raghvendra Kannao, Prithwijit Guha
ICPR2
2016 Multiple kernel learning using data envelopment analysis and feature vector selection and projection
abstract
Multiple kernel learning methods combine a set of base kernels to produce an optimal one for a certain classification or regression problem. But selecting a set of base kernels from a plethora of kernels is not automated. We provide a criteria to select efficient base kernels. Automating the selection process of efficient base kernel requires less time and effort than manually selecting them. However, learning the weights in the ratio of which the selected kernels are to be combined is still a costly process. To calculate these combination weights, we first evaluate the efficiency of a kernel on the basis of two parameters viz. the Trace and Alignment of the kernel matrix using data envelopment analysis. A base kernel can be selected if its efficiency is 100%. After selecting a set of most efficient kernels, we combine them in the proportion to their efficiencies. Also, we want to control the complexity of the model by the method of data selection. We use feature vector selection method to cast data points to a limited number of features and then apply classical algorithms to solve our classification problems.
Gitimoni Saikia, Saroj Shivagunde, Vijaya V. Saradhi, Raghvendra Kannao, Prithwijit Guha
ICPR5
2016 Reinforcement Learning via Recurrent Convolutional Neural Networks
abstract
Deep Reinforcement Learning has enabled the learning of policies for complex tasks in partially observable environments, without explicitly learning the underlying model of the tasks. While such model-free methods do achieve considerable performance, they often ignore the structure of task. We present a more natural representation of the solutions to Reinforcement Learning (RL) problems, within 3 Recurrent Convolutional Neural Network (RCNN) architectures to better exploit this inherent structure. The forward passes of each RCNN execute an efficient Value Iteration, propagate beliefs of state in partially observable environments, and choose optimal actions respectively. Applying back-propagation to these RCNNs allows the system to explicitly learn the Transition Model and Reward Function associated with the underlying MDP, serving as an elegant alternative to classical model-based RL. We evaluate the proposed algorithms in simulation, considering a robot planning problem. We demonstrate the capability of our framework to reduce the cost of re-planning, learn accurate MDP models, and finally re-plan with learned models to achieve near-optimal policies.
Tanmay Shankar, Santosha K. Dwivedy, Prithwijit Guha
ICPR3
2016 News Program Detection in TV Broadcast Videos
abstract
Television news channels broadcast different kinds of content like debates, interviews, commercials along with news presentations. Real-time detection of these news programs or their retrieval from large volumes of stored broadcast videos is a challenging problem and is a necessary first step for broadcast analytics. News program detection is even harder in Indian context where closed caption text or program markers are not provided by TV news channels (not mandated by law). We propose a two-stage approach to classify news video segments. First, broadcast video shots are classified with multiple labels based on a set of audio-visual features. Second, sequences of these shot features are modeled to detect news programs. Another contribution of this work is the construction of a dataset of 120 hours of shot categories and news programs from Indian English news channels. We have experimented with SVM, HMM and CRF based classifiers and achieved a F1 score of 99% in detecting news programs while experimenting on our dataset.
Raghvendra Kannao, Durgaprasad Dandi, Swamy Yellapu, Prithwijit Guha
ACM Multimedia4
2016 TV Commercial Detection Using Success Based Locally Weighted Kernel Combination
Raghvendra Kannao, Prithwijit Guha
MMM (1)2
2015 An occlusion reasoning scheme for monocular pedestrian tracking in dynamic scenes
abstract
This paper looks into the problem of pedestrian tracking using a monocular, potentially moving, uncalibrated camera. The pedestrians are located in each frame using a standard human detector, which are then tracked in subsequent frames. This is a challenging problem as one has to deal with complex situations like changing background, partial or full occlusion and camera motion. In order to carry out successful tracking, it is necessary to resolve associations between the detected windows in the current frame with those obtained from the previous frame. Compared to methods that use temporal windows incorporating past as well as future information, we attempt to make decision on a frame-by-frame basis. An occlusion reasoning scheme is proposed to resolve the association problem between a pair of consecutive frames by using an affinity matrix that defines the closeness between a pair of windows and then, uses a binary integer programming to obtain unique association between them. A second stage of verification based on SURF matching is used to deal with those cases where the above optimization scheme might yield wrong associations. The efficacy of the approach is demonstrated through experiments on several standard pedestrian datasets.
Sourav Garg, Swagat Kumar, Rajesh Ratnakaram, Prithwijit Guha
AVSS4
2012 Unsupervised Language Learning for Discovered Visual Concepts
Prithwijit Guha, Amitabha Mukerjee
ACCV (4)1
2012 The Video Face Book
Nipun Pande, Dhawal Kapil, Prithwijit Guha
MMM4
2011 Formulation, detection and application of occlusion states (Oc-7) in the context of multiple object tracking
abstract
Occlusion is often thought of as a challenge for visual algorithms, specially tracking. Existing literature, however, has identified a number of occlusion categories in the context of tracking in ad hoc manner. We propose a systematic approach to formulate a set of occlusion cases by considering the spatial relations among object support(s) (projections on the image plane) with the detected foreground blob(s), to show that only 7 occlusion states are possible. We designate the resulting qualitative formalism as Oc-7, and show how these occlusion states can be detected and used effectively for the task of multi-object tracking under occlusion of various types. The object support is decomposed into overlapping patches which are tracked independently on the occurrence of occlusions. As a demonstration of the application of these occlusion states, we propose a reasoning scheme for selective tracker execution and object feature updates to track multiple objects in complex environments.
Prithwijit Guha, Amitabha Mukerjee, K. S. Venkatesh
AVSS1
2011 OCS-14 : You Can Get Occluded in Fourteen Ways
Prithwijit Guha, Amitabha Mukerjee, K. S. Venkatesh
IJCAI1
2008 Back to the future: Robust foreground extraction with reversed-time background modeling
abstract
ldquoGhostsrdquo arise in traditional background subtraction when an object starts to move, causing the exposed background to be labelled as a ghost foreground. With background model updates, the ghost may disappear after some time, but removing ghosts immediately is crucial for identifying starts, and also for improving downstream tasks like tracking, object recognition and activity analysis. Here we propose a lagged background subtraction process, where the background model is computed in reverse after a lag of k frames, resulting in a ldquoreversed-timerdquo background model. We present an algorithm that handles disparities between the forward and backward foreground blobs, and show how ghosts can be reliably eliminated, potentially permitting single-frame latency in reliably detecting starts and identifying stops. We present theoretical results that the bidirectional model results in lower false positives (such as ghosts) compared to either directional model, and develop an algorithm for identifying ghosts.
Akhilesh Kumar Sinha, Prithwijit Guha, Amitabha Mukerjee
ICPR2
2006 A Multiscale Co-linearity Statistic Based Approach to Robust Background Modeling
Prithwijit Guha, Dibyendu Palai, K. S. Venkatesh, Amitabha Mukerjee
ACCV (1)1
2006 Efficient Continuous Re-grasp Planning for Moving and Deforming Planar Objects
abstract
A novel approach to real-time tracking of three-finger planar grasp points for deforming objects is proposed. The search space of possible grasping configurations is reduced in two stages - firstly, by fixing one finger at the boundary point nearest to the object centroid and secondly, through a heuristic partitioning of the object boundary where the remaining two fingers are localized. The potential grasping configurations satisfying force closure conditions are evaluated through an objective function that maximizes the grasping span while minimizing the distance between the object centroid and the intersection of the contour normals at the finger contact points. A population based stochastic search strategy is adopted for computing the optimal grasping configurations and re-localizing them as the shape undergoes drastic translations, rotations, scaling and local deformations. Experimental results of grasp point tracking are presented for deforming planar shapes extracted from both real and synthetic image sequences. The current implementation of the proposed scheme operates at 10 Hz for grasp point tracking on shapes extracted through visual feedback
Tripuresh Mishra, Prithwijit Guha, Ashish Dutta, K. S. Venkatesh
ICRA2
2006 Appearance Based Multiple Agent Tracking Under Complex Occlusions
Prithwijit Guha, Amitabha Mukerjee, K. S. Venkatesh
PRICAI1