Achyut Mani Tripathi

dblp:173/8904 · DBLP profile ↗
← Back
24ranked-venue papers
17as first author
17since 2021 · last 2026
0000-0003-0548-9688ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 15 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 WaveletMamba: Harnessing learnable multi-scale state space dynamics for time series forecasting
Sachin Maurya, Achyut Mani Tripathi, Kedar Vithal Khadeparkar
Neural Networks3
2025 Semi-Supervised Knowledge Distillation Framework towards Lightweight Large Language Model for Spoken Language Translation
abstract
Even though large language models (LLMs) have demonstrated remarkable performance across various natural language processing tasks, their application in speech-related tasks has largely remained underexplored. This work addresses this gap by incorporating acoustic features into an LLM which can be fine-tuned for downstream direct speech-to-text translation and automatic speech recognition tasks. To address the computational demands associated with fine-tuning LLMs, a novel self and semi-supervised knowledge distillation technique is proposed to implement a lightweight LLM having 50% lesser parameters. Validated on the MuST-C and Librispeech datasets, this technique achieves over 92% of the performance of the larger LLM, demonstrating both robust performance and computational efficiency.
Tonmoy Rajkhowa, Amartya Chowdhury, Achyut Mani Tripathi, Sanjeev Sharma 0001, Om Jee Pandey
ICASSP3
2025 Bandit-based Attention Mechanism in Vision Transformers
abstract
Vision Transformers (ViT) have demonstrated remarkable performance on many computer vision tasks. However, their high computational cost and quadratic complexity pose chal-lenges for deployment in resource-constrained environments. The core of Vision Transformers is the self-attention mecha-nism which aggregates information from different image re-gions, or patches. In a conventional ViT, processing involves attention to all patches, creating a substantial computational bottleneck and extended training times. We hypothesize that applying soft attention to all patches may be unnecessary and instead focusing on relevant and significant patches (hard attention) would be sufficient. To address this, we introduce a module within the Vision Transformer that allows the at-tention mechanism to selectively process only the essential patches. We propose a novel bandit-based attention mecha-nism that leverages the idea of exploration and exploitation. The extensive experimentation across various datasets il-lustrates that the proposed bandit attention-based ViT not only achieves superior performance compared to the existing state-of-the-art vision transformer models but also results in greater throughput and lower computational time in the training as well as the inference. The code is publicly avail-able at https://github.com/aquorio15/bandit_wacv
Amartya Chowdhury, Raghuram Bharadwaj Diddigi, Prabuchandran K. J., Achyut Mani Tripathi
WACV4
2025 When Visual State Space Model Meets Backdoor Attacks
abstract
The recently proposed Visual State Space Model (VMamba), operating on the principle of state space mechanisms (SSM), processes images as a sequence of patches and outperforms Vision Transformers (ViT) in several computer vision tasks. Given their substantial design differences from CNNs and ViT, it is crucial to investigate their vulnerability to backdoor attacks and the impact of various advanced backdoor attacks on their robustness. Backdoor attacks involve embedding a specific trigger into a small subset of training images, which remains dormant until activated later. While the model performs well on clean test images, an attacker can manipulate its decisions by presenting the trigger in one of the test images. This work examines state-of-the-art (SOTA) vision architectures (ResNet, ViT, MLP-mixer, VMamba), focusing on their susceptibility to backdoor attacks and the effect of different backdoor attacks on their robustness. The well-known Visual State Space Model (VMamba) is the least susceptible to backdoor attacks among these architectures. To address this, in this paper, we propose two novel QR decomposition-based backdoor attacks that are visually imperceptible and achieve a high attack success rate (ASR) against the VMamba. We also present a qualitative analysis of the proposed backdoor attacks, explaining the reasons behind their success or failure against the VMamba model. Experiments and results conducted on two popular image datasets (CIFAR-10 and ImageNet-1K) demonstrate that the proposed backdoor attacks exceed the performance of SOTA backdoor attacks and effectively fool the recently proposed VMamba model.
Sankalp Nagaonkar, Achyut Mani Tripathi
WACV2
2024 TM-PATHVQA: 90000+ Textless Multilingual Questions for Medical Visual Question Answering
abstract
In healthcare and medical diagnostics, Visual Question Answering (VQA) may emerge as a pivotal tool in scenarios where analysis of intricate medical images becomes critical for accurate diagnoses.Current text-based VQA systems limit their utility in scenarios where hands-free interaction and accessibility are crucial while performing tasks.A speech-based VQA system may provide a better means of interaction where information can be accessed while performing tasks simultaneously.To this end, this work implements a speech-based VQA system by introducing a Textless Multilingual Pathological VQA (TM-PathVQA) dataset, an expansion of the PathVQA dataset, containing spoken questions in English, German & French.This dataset comprises 98,397 multilingual spoken questions and answers based on 5,004 pathological images along with 70 hours of audio.Finally, this work benchmarks and compares TM-PathVQA systems implemented using various combinations of acoustic and visual features.
Tonmoy Rajkhowa, Amartya Chowdhury, Sankalp Nagaonkar, Achyut Mani Tripathi, S. R. Mahadeva Prasanna
INTERSPEECH4
2023 Sub-Band Contrastive Learning-Based Knowledge Distillation For Sound Classification
abstract
Knowledge distillation(KD) technique is widely known for its outstanding ability to train a compact student model under supervision of a cumbersome pre-trained teacher network. The traditional KD technique focuses only on distilling dark knowledge using logits of a teacher network while neglecting the information regarding contrastive representation. To this end, we propose a new KD loss function that enables a student network to learn informative contrastive distribution and fine grained information from spectrogram representation of a signal thus enhancing performance of a student network for sound classification task. The experiments are conducted on two benchmark sound classification datasets, viz. ESC-10 and Audio MNIST, which illustrates that the student network trained using the proposed KD loss function outperformed the competitive KD techniques.
Achyut Mani Tripathi, Aakansha Mishra
ICASSP1
2023 Towards Multi-Lingual Audio Question Answering
Swarup Ranjan Behera, Pailla Balakrishna Reddy, Achyut Mani Tripathi, Megavath Bharadwaj Rathod, Tejesh Karavadi
INTERSPEECH3
2023 Anchor-based void detouring routing protocol in three dimensional IoT networks
abstract
In recent years, several applications of Internet of Things (IoT) have been observed in various areas including environmental monitoring, healthcare systems, cognitive smart agriculture, industrial control, smart homes , intelligent transportation systems , and traffic management. For such applications, wireless sensor networks (WSNs) are generally deployed to gather the sensed data from the targeted application field. In order to transfer the sensor node data to the gateway (sink node), novel routing protocols need to be developed, leading to reduced data transmission delay, high data throughput , and improved energy efficiency across the network. In this context, geographical routing protocol has been considered as a promising approach for the path selection in WSNs. This approach is full of scalability and multi-hop routing is performed using local decisions. However, geographical routing protocols suffer from the void node problem (VNP) i.e., a region where active nodes are not available in the direction closer to the destination. Numerous protocols have been designed to get recovery from VNP in 2D networks which cannot be directly applied to 3D networks. The 3D routing includes the networks deployed in the hilly area, high buildings, airborne region, underground, underwater and so forth. On applying the 2D routing protocols on complex 3D topology, the network may face additional problems like packet looping, routing failure, ambiguity, or increased data latency due to longer path. Further, the majority of geographical routing protocols follow the boundary of void which leads to a longer path. In order to address the aforementioned challenges, this paper presents a novel anchor-based void detouring routing (AVDR) protocol where anchor node is treated as a sub-destination which provides the direct smaller path between source and gateway nodes. The proposed method bypasses the void boundaries and directly connects source to anchor, anchor to destination, or two successive anchors. Further, anchor information is distributed to the desired region to reduce the periodic anchor advertisement process. The effectiveness of the proposed method has been tested over both, real field data set and simulated testbed with OMNET++ simulator. The results obtained over real field data set claim that the proposed method takes only 29.09 ms (ms) for transferring the data on an average. However, this value is 32.37 ms, 34.32 ms, 33.61 ms, 37.20 ms, and 38.73 ms, respectively, using A3DR, EDGR, GPSR-3D, BSMH, and RPL methods. Moreover, it is also noted that the proposed method achieves an improvement of 8.2%, 7.54%, 7.66%, 8.49%, and 8.22%, in routing stretch when compared to aforementioned methods, respectively. This improvement with respect to network overhead is 30.25%, 57.45%, 51.05%, 75.56%, and 58.89% using the proposed method.
Naveen Kumar Gupta, Rama Shankar Yadav, Rajendra Kumar Nagaria, Achyut Mani Tripathi, Om Jee Pandey
Comput. Networks5
2023 Divide and Distill: New Outlooks on Knowledge Distillation for Environmental Sound Classification
abstract
Environmental sound classification (ESC) is an important research problem with a broad range of applications including audio-based surveillance, audio-visual systems, smart homes, and robotics, among others. The recently proposed vision multi-layer perceptron-mixer (MLP-mixer) has outperformed traditional deep models (CNN or ResNet) and attained new state-of-the-art performances for several computer vision applications (image/video classification and image segmentation). Following the success of MLP-mixer, in this paper, we propose a novel audio MLP-mixer (AMM) network that classifies the different types of environmental sounds. Despite the higher performance, the high computational cost (number of trainable parameters and floating point operations) prohibits deployment of the AMM model on edge for designing real-life applications. To alleviate the aforementioned issue, in this work, we present three different knowledge distillation (KD) strategies to train a compact deep network for ESC. The proposed strategies divide the input Mel-spectrogram into patches and a lightweight deep ESC model is trained in the presence of three teacher networks under the offline KD training framework. Additionally, we have designed two novel loss functions for KD that are free from a temperature parameter that need to be set manually by a user as in the case of the traditional vanilla KD technique. We conducted our experiments on three benchmark ESC datasets namely ESC-10, Urbansound8k (US8K), and DCASE-2019 Task-1(A). The obtained results demonstrate the significance of utilization of proposed methods over other existing KD methods in terms of classification accuracy.
Achyut Mani Tripathi, Om Jee Pandey
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 Defensive Bit Planes: Defense Against Adversarial Attacks
abstract
Deep learning-based models are vulnerable to adversarial examples crafted with different adversarial attack techniques. Numerous attack methods have been proposed that utilize gradient information of deep model to craft an adversarial examples. Amongst the existing defense mechanisms, adversarial training has gained considerable attention in building robust deep models that remain effective against different adversarial attacks. However, adversarial training demands high computational cost during the development of a robust deep model. In this paper, we present a simple yet effective defense mechanism against adversarial attacks. The proposed defense mechanism uses the concept of bit plane slicing for de-noising of an input image. The efficacy of the proposed defense technique has been evaluated on two benchmark image datasets, viz. MNIST and Fashion-MNIST datasets. The experiments and results show that the proposed defence technique yields comparable and competitive performance to state-of-the-art defense techniques against adversarial attacks.
Achyut Mani Tripathi, Swarup Ranjan Behera, Konark Paul
IJCNN1
2022 Adv-IFD: Adversarial Attack Datasets for An Intelligent Fault Diagnosis
abstract
Deep learning techniques have been widely applied for performing intelligent fault diagnosis (IFD) for applications such as bearing fault diagnosis, wind turbines and drilling operations. Notwithstanding the huge success of deep learning-based models for intelligent fault diagnosis, their vulnerability against adversarial attacks have been neglected to a great extent. The adversarial attacks aim to craft an adversarial sample by addition of imperceptible perturbation to its clean data sample. In this paper, we investigate the performance of different deep models against four state-of-the-art adversarial attacks. A total of four deep models have been tested against untargeted white-box adversarial attacks. Moreover, analysis of transferability of adversarial examples across different deep models is also inspected. Experiments and results reveal that deep models for IFD are highly susceptible to adversarial examples crafted with four state-of-the-art adversarial attacks. The proposed work presents an extensive insight on adversarial samples of machinery vibration signals of the CWRU dataset. Additionally, we are also releasing a first ever adversarial attack dataset for IFD, i.e. Adv-IFD. The code used in this work and adversarial attack datasets are available at: https://github.com/achyutmani/ADV-IFD
Achyut Mani Tripathi, Swarup Ranjan Behera, Konark Paul
IJCNN1
2022 Reverse Adversarial Attack To Enhance Environmental Sound Classification
abstract
Classification of environmental sounds becomes an essential unit for several applications such as robot audition, audio-visual and ambient intelligent systems. Recent years have witnessed remarkable success of adversarial attacks for misguiding deep models with a high fooling rate. Adversarial examples crafted with adversarial techniques pose serious security threats for deployment of deep models to practical scenarios. In this paper, we show that adversarial signals crafted by performing search in the opposite direction of the search direction of the perturbations could be utilized while training of deep classifiers of environmental sound classification (ESC). The proposed reverse adversarial attack (RAA) technique can be used as data augmentation tool to enhance the accuracy of deep model that performs ESC. The deep network trained with the crafted adversarial signal shows high robustness against adversarial attacks. The efficacy of the proposed method is evaluated on two benchmark ESC datasets viz. ESC-10 and DCASE-2019 Task-1A datasets. The experiments and results show that the model trained with the proposed technique shows comparable and competitive performance to state-of-the-art (SOTA) deep ESC models.
Achyut Mani Tripathi, Swarup Ranjan Behera, Konark Paul
IJCNN1
2022 Investigation of Performance of Visual Attention Mechanisms for Environmental Sound Classification: A Comparative Study
abstract
Classification of environmental sound is an equally challenging task to classification of human speech or music. In recent years, attention mechanism has gained considerable attention in learning representative and prototypical features to resolve various research challenges from domains such as computer vision, text mining and sound classification. Advancements in the computer vision can be used for improving the performance of deep models designed to perform environmental sound classification as well. In this paper, we investigated the performance of deep network when twelve SOTA visual attention mechanisms are incorporated in the training of the deep network. The performance of the deep model is evaluated on two benchmark environmental sound classification datasets, viz. ESC-10 and DCASE-2019 task-1(A) datasets.
Achyut Mani Tripathi, Swarup Ranjan Behera, Konark Paul
IJCNN1
2022 Revamped Knowledge Distillation for Sound Classification
abstract
This paper presents a novel knowledge distillation technique that inherits knowledge from multiple deep Environment Sound Classification (ESC) models trained on spectrogram features created by dividing the spectrogram into multiple subband spectrogram. The deep models trained on sub-band spectrograms prevent information loss while performing knowledge distillation from a teacher model to a student model receiving the full spectrogram as an input. The student models' performance is evaluated on two benchmark sound datasets, viz. the ESC-10 and Audio MNIST datasets. The impact of teacher models trained with different number of sub-band features and four ensemble techniques has been investigated thoroughly to enhance the final accuracy of the student model supervised by the proposed knowledge distillation framework. Experiments and results shows that the accuracy of the student model is comparable and competitive to state-of-the-art methods for sound classification. Moreover, the student model trained on the Audio MNIST dataset attains an hitherto unpublished accuracy of 98.25%, a new benchmark for the Audio MNIST dataset. Additionally, Grad- CAM visualization of the spectrogram features is generated to identify the spectrogram's relevant regions and understand why the model classifies a signal into a specific class.
Achyut Mani Tripathi, Aakansha Mishra
IJCNN1
2022 Temporal Self Attention-Based Residual Network for Environmental Sound Classification
Achyut Mani Tripathi, Konark Paul
INTERSPEECH1
2022 Data augmentation guided knowledge distillation for environmental sound classification
Achyut Mani Tripathi, Konark Paul
Neurocomputing1
2021 Environment sound classification using an attention-based residual neural network
Achyut Mani Tripathi, Aakansha Mishra
Neurocomputing1
2020 Contextual Anomaly Detection in Time Series Using Dynamic Bayesian Network
Achyut Mani Tripathi, Rashmi Dutta Baruah
ACIIDS (2)1
2020 Acoustic Event Detection Using Fuzzy Integral Ensemble and Oriented Fuzzy Local Binary Pattern Encoded CNN
abstract
In this paper, we propose a novel ensemble classifier using an Oriented Fuzzy Local Binary Pattern Encoded Convolutional Neural Network (CNN) for acoustic event detection (AED). The CNN has been widely used to perform acoustic event detection using a spectrogram image of the acoustic signals. The efficiency of the CNN depends on representation of the spectrogram images used during the training process. We propose the Oriented Fuzzy Local Binary Pattern (OFLBP) that extracts directional texture features from the spectrogram image by inspecting neighborhood pixels present at different angles from a central pixel. The proposed OFLBP technique is capable to deal with uncertainty present in the spectrogram image. The ensemble of the trained CNN is performed by a Fuzzy Integral method. The experiment and results show the proposed method outperforms to existing AED methods to classify the ESC-50 dataset.
Achyut Mani Tripathi, Rashmi Dutta Baruah
FUZZ-IEEE1
2020 Enhancing Multivariate Time Series Classification Using LSTM and Evidence Feed Forward HMM
abstract
This paper presents a hybrid classifier that combines a Long Short Term Memory (LSTM) and an Evidence Feed Forward Hidden Markov Model (EFF-HMM) to classify multivariate time series (MTS). Learning of the EFF-HMM is performed based on mistakes of the LSTM. Confusion matrix obtained after classification of the MTS by the LSTM is employed during a learning process of the EFF-HMM. The EFF-HMM efficiently models a temporal characteristic and uncertainty of the MTS. The hybrid classifier combines strengths of the LSTM and EFFHMM to enhance the accuracy of the MTS classification. The proposed method is tested on Human Activity Recognition (HAR) dataset to classify various human activities. The experiments and results show the proposed method outperforms as compared to state of the art methods.
Achyut Mani Tripathi
IJCNN1
2020 Multivariate Time Series Classification With An Attention-Based Multivariate Convolutional Neural Network
abstract
Classification of time series is an essential requirement of various applications that demand continuous monitoring of dynamical systems such as industrial process and health care monitoring. Feature extraction plays a vital role in deciding performance of the time series models. In recent years deep learning techniques have shown an excellent performance to extract highly discriminating features for the classification of the time series. A Convolutional Neural Network (CNN) is a unified framework that performs the feature learning and classification tasks simultaneously. Using the CNN to perform the classification of the multivariate time series is still a challenging task. In this paper, we propose an attention-based multivariate convolutional neural network (AT-MVCNN) that consists of the attention feature-based input tensor scheme to encode informations across the multiple time stamps. The method is capable of learning the temporal characteristics of the multivariate time series. The efficacy of the proposed method is tested on Human Activity Recognition (HAR) and Occupancy Detection datasets. The experiments and results show the proposed method outperforms the other deep learning and traditional machine learning models.
Achyut Mani Tripathi, Rashmi Dutta Baruah
IJCNN1
2019 Incremental Cauchy Non-Negative Matrix Factorization and Fuzzy Rule-based Classifier for Acoustic Source Separation
abstract
Non-negative matrix factorization (NNMF) technique has been widely applicable for dimension reduction. NNMF has shown its usability to solve numerous challenges in areas such as signal processing, image classification, and text mining. Existing matrix factorization methods need a predefined value of rank of a base matrix. Our primary intention in this article is to automatically identify the optimum value of the rank to factorize the matrix. Here we propose an Incremental Cauchy NonNegative Matrix factorization (ICNNMF) which automatically finds the optimum rank and separates the overlapping sounds received from a single channel of a source. Classification is performed using an ensemble of one class fuzzy rule-based (FRB) classifiers. We evaluated the proposed method over the real data set, and the proposed model has shown better accuracy as compared to the traditional classifier. Moreover, the proposed matrix factorization method automatically determines the optimal value of the rank and adapt to new data.
Achyut Mani Tripathi, Rashmi Dutta Baruah
FUZZ-IEEE1
2019 Anomaly Detection in Multivariate Time Series Using Fuzzy AdaBoost and Dynamic Naive Bayesian Classifier
abstract
This paper presents a novel method to detect anomaly using Fuzzy AdaBoost and Dynamic Naive Bayesian classifier. Dynamic Naive Bayesian Classifier (DNBC) is an extension of Hidden Markov Models (HMM) that is used here to model multivariate observation sequences typically generated from multiple sensors associated to monitoring processes. The Fuzzy AdaBoost (FAB) method is used for ensembling multiple DNBCs to classify an instance as an anomaly. FAB needs Footprint of Uncertainty (FOU) of error that is further used to Figure out the misclassification error and update weights of the data samples required during the boosting process. Here, we introduce an approach to initialize the intervals of FOU using the statistical assets of data that belongs to the normal class. The efficacy of the proposed method is demonstrated through a case study on a stuck pipe problem that occurs during oil well drilling process.
Achyut Mani Tripathi, Rashmi Dutta Baruah
SMC1
2017 Acoustic event classification using Cauchy Non-negative matrix factorization and fuzzy rule-based classifier
abstract
Identification of presence of target acoustic sound or event from a single channel mixture is a challenging task of automatic sound recognition system. In presence of background noise, the detection and classification of target acoustic event becomes more difficult. Various methods have been proposed that extract features from spectrogram of sound and then the extracted features are used with traditional non negative matrix factorization for separation of overlapping sound. In his paper, we propose an approach to separate and classify single channel acoustic events. The method combines Common Fate Transformation and Cauchy Non-negative Matrix Factorization for feature extraction and finally fuzzy rule-based classifier is developed for classification. The proposed method, when applied to real data, gave high true positive rate. The method also gave better results in terms of true positive rate when compared to widely used support vector machine using the same real data. Moreover, the proposed approach is fast and can be used for the efficient separation of acoustic events from overlapping sounds.
Achyut Mani Tripathi, Rashmi Dutta Baruah
FUZZ-IEEE1