Raju S. Bapi

dblp:50/1424 · also Bapi Raju Surampudi, Bapi Surampudi Raju, Raju Surampudi Bapi, S. Bapi Raju, Surampudi Bapi Raju · DBLP profile ↗
← Back
60ranked-venue papers
3as first author
25since 2021 · last 2025
0000-0003-2204-0890ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 53 · 3 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Aligning Text/Speech Representations from Multimodal Models with MEG Brain Activity During Listening
abstract
Although speech language models are expected to align well with brain language processing during speech comprehension, recent studies have found that they fail to capture brainrelevant semantics beyond low-level features.Surprisingly, text-based language models exhibit stronger alignment with brain language regions, as they better capture brain-relevant semantics.However, no prior work has examined the alignment effectiveness of text/speech representations from multimodal models.This raises several key questions: Can speech embeddings from such multimodal models capture brain-relevant semantics through cross-modal interactions?Which modality can take advantage of this synergistic multimodal understanding to improve alignment with brain language processing?Can text/speech representations from such multimodal models outperform unimodal models?To address these questions, we systematically analyze multiple multimodal models, extracting both text-and speech-based representations to assess their alignment with MEG brain recordings during naturalistic story listening.We find that text embeddings from both multimodal and unimodal models significantly outperform speech embeddings from these models.Specifically, multimodal text embeddings exhibit a peak around 200 ms, suggesting that they benefit from speech embeddings, with heightened activity during this time period.However, speech embeddings from these multimodal models still show a similar alignment compared to their unimodal counterparts, suggesting that they do not gain meaningful semantic benefits over text-based representations.These results highlight an asymmetry in cross-modal knowledge transfer, where the text modality benefits more from speech information, but not vice versa.We make the code publicly available 1 .
Padakanti Srijith, Khushbu Pahwa, Radhika Mamidi, Raju S. Bapi, Manish Gupta 0001, Subba Reddy Oota
EMNLP4
2025 Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)
abstract
Transformer-based language models, though not explicitly trained to mimic brain recordings, have demonstrated surprising alignment with brain activity. Progress in these models—through increased size, instruction-tuning, and multimodality—has led to better representational alignment with neural data. Recently, a new class of instruction-tuned multimodal LLMs (MLLMs) have emerged, showing remarkable zero-shot capabilities in open-ended multimodal vision tasks. However, it is unknown whether MLLMs, when prompted with natural instructions, lead to better brain alignment and effectively capture instruction-specific representations. To address this, we first investigate the brain alignment, i.e., measuring the degree of predictivity of neural visual activity using text output response embeddings from MLLMs as participants engage in watching natural scenes. Experiments with 10 different instructions (like image captioning, visual question answering, etc.) show that MLLMs exhibit significantly better brain alignment than vision-only models and perform comparably to non-instruction-tuned multimodal models like CLIP. We also find that while these MLLMs are effective at generating high-quality responses suitable to the task-specific instructions, not all instructions are relevant for brain alignment. Further, by varying instructions, we make the MLLMs encode instruction-specific visual concepts related to the input image. This analysis shows that MLLMs effectively capture count-related and recognition-related concepts, demonstrating strong alignment with brain activity. Notably, the majority of the explained variance of the brain encoding models is shared between MLLM embeddings of image captioning and other instructions. These results indicate that enhancing MLLMs' ability to capture more task-specific information could allow for better differentiation between various types of instructions, and hence improve their precision in predicting brain responses.
Subba Reddy Oota, Akshett Rai Jindal, Ishani Mondal, Khushbu Pahwa, Satya Sai Srinath Namburi, Manish Shrivastava 0001, Maneesh Kumar Singh 0001, Raju S. Bapi, Manish Gupta 0001
ICLR8
2025 Multi-modal brain encoding models for multi-modal stimuli
abstract
Despite participants engaging in unimodal stimuli, such as watching images or silent videos, recent work has demonstrated that multi-modal Transformer models can predict visual brain activity impressively well, even with incongruent modality representations. This raises the question of how accurately these multi-modal models can predict brain activity when participants are engaged in multi-modal stimuli. As these models grow increasingly popular, their use in studying neural activity provides insights into how our brains respond to such multi-modal naturalistic stimuli, i.e., where it separates and integrates information across modalities through a hierarchy of early sensory regions to higher cognition (language regions). We investigate this question by using multiple unimodal and two types of multi-modal models—cross-modal and jointly pretrained—to determine which type of models is more relevant to fMRI brain activity when participants are engaged in watching movies (videos with audio). We observe that both types of multi-modal models show improved alignment in several language and visual regions. This study also helps in identifying which brain regions process unimodal versus multi-modal information. We further investigate the contribution of each modality to multi-modal alignment by carefully removing unimodal features one by one from multi-modal representations, and find that there is additional information beyond the unimodal embeddings that is processed in the visual and language regions. Based on this investigation, we find that while for cross-modal models, their brain alignment is partially attributed to the video modality; for jointly pretrained models, it is partially attributed to both the video and audio modalities. These findings serve as strong motivation for the neuro-science community to investigate the interpretability of these models for deepening our understanding of multi-modal information processing in brain.
Subba Reddy Oota, Khushbu Pahwa, Mounika Marreddy, Maneesh Kumar Singh 0001, Manish Gupta 0001, Raju S. Bapi
ICLR6
2025 Ambiguous Medical Image Segmentation Using Diffusion Schrödinger Bridge
Lalith Bharadwaj Baru, Kamalaker Dadi, Tapabrata Chakraborti, Raju S. Bapi
MICCAI (4)4
2024 Enhancing Healthcare with EOG: A Novel Approach to Sleep Stage Classification
abstract
We introduce an innovative approach to automated sleep stage classification using electrooculogram (EOG) signals, addressing the discomfort and impracticality associated with electroencephalogram (EEG) data acquisition. In addition, this approach is untapped in the field, highlighting its potential for novel insights and contributions. Our proposed SE-Resnet-Transformer model effectively classifies five distinct sleep stages from raw EOG signals. Extensive validation on publicly available databases (SleepEDF-20, SleepEDF-78, and SHHS) reveals performance, with macro-F1 scores of 74.72, 70.63, and 69.26, respectively. The model excels in identifying Rapid Eye Movement (REM) sleep, a crucial aspect of sleep disorder investigations. We also provide insight into the internal mechanisms of the model using techniques such as GradCAM and t-SNE plots. Our method improves the accessibility of sleep stage classification while decreasing the need for EEG modalities. This development will have promising implications for healthcare and the incorporation of wearable technology into sleep studies, thereby advancing the field’s potential for enhanced diagnostics and patient comfort.
Suvadeep Maiti, Shivam Kumar Sharma, Raju S. Bapi
ICASSP3
2024 HyperGALE: ASD Classification via Hypergraph Gated Attention with Learnable HyperEdges
abstract
Autism Spectrum Disorder (ASD) is a neurodevelopmental condition characterized by varied social cognitive challenges and repetitive behavioral patterns. Identifying reliable brain imaging-based biomarkers for ASD has been a persistent challenge due to the spectrum’s diverse symptomatology. Existing baselines in the field have made significant strides in this direction, yet there remains room for improvement in both performance and interpretability. We propose HyperGALE, which builds upon the hypergraph by incorporating learned hyperedges and gated attention mechanisms. This approach has led to substantial improvements in the model’s ability to interpret complex brain graph data, offering deeper insights into ASD biomarker characterization. Evaluated on the extensive ABIDE II dataset, HyperGALE not only improves interpretability but also demonstrates statistically significant enhancements in key performance metrics compared to both previous baselines and the foundational hypergraph model. The advancement HyperGALE brings to ASD research highlights the potential of sophisticated graph-based techniques in neurodevelopmental studies. The source code and implementation instructions are available at GitHub.
Mehul Arora, Chirag Shantilal Jain, Lalith Bharadwaj Baru, Kamalaker Dadi, Raju S. Bapi
IJCNN5
2023 Intended Outcomes Expand In Time: Evidence from the Temporal Reproduction Task
Rohan R. Donapati, Anuj Shukla, Raju S. Bapi
CogSci3
2023 Neural Architecture of Speech
abstract
A vast literature on brain encoding has effectively harnessed deep neural network models for accurately predicting brain activations from visual or text stimuli. Unfortunately, there is not much work on brain encoding for speech stimuli. The few existing studies on brain encoding for speech stimuli transcribe speech to text and then leverage text-only models for encoding, thereby ignoring audio signals completely. However, recently several speech representation learning models have revolutionized the field of speech processing. Inspired by the recent progress on deep learning models for speech, we present a first systematic study on understanding human speech processing by probing neural speech models to predict both language and auditory brain region activations. In particular, we investigate 30 speech representation models grouped into four categories: (i) traditional feature engineering, (ii) generative, (iii) predictive, and (iv) contrastive, to study how these models encode the speech stimuli and align with human brain activity for the Moth Radio Hour fMRI (functional magnetic resonance imaging) dataset. We find that both contrastive (Wav2Vec2.0) and predictive models (HuBERT, Data2Vec) are very accurate. Specifically, Data2Vec aligns the best with both language and auditory brain regions among all investigated models. We make our code publicly available1.
Subba Reddy Oota, Khushbu Pahwa, Mounika Marreddy, Manish Gupta 0001, Raju S. Bapi
ICASSP5
2023 Emergence of Direction Selectivity and Motion Strength in Dot Motion Task Through Deep Reinforcement Learning Networks
abstract
Deep Reinforcement learning is beginning to be useful for studying neural representations in the brain because of its ability to combine decision-making and representation. Here, we use it to study a dot motion perceptual decision-making task in a high-dimensional setting where the inputs are akin to those used in psychological experiments. This end-to-end model gives a unique insight into how these networks solve the task providing a background on how the brain could solve this task. We find that the network can show properties similar to the middle temporal visual area (MT) in the brain, which code for direction and motion strength. We find the emergence of direction selectivity purely through reward-based training and graded firing coding motion strength and make a testable prediction that the MT population would also have coherence-selective neurons.
Dolton Fernandes, Pramod S. Kaushik, Raju S. Bapi
IJCNN3
2023 Dynamic functional connectivity analysis in individuals with Autism Spectrum Disorder
abstract
Autism spectrum disorder (ASD) is a neurodevelopmental disorder that predominantly occurs in children. Previous brain research in ASD has mainly studied biomarkers based on the functional connectivity characterized by the correlation of static temporal signals. However, brain connectivity is dynamic and varies extensively among brain states. The main aim of the paper is to understand the fundamental group differences between ASD patients and typically developing (TD) subjects using dynamic functional connectivity (dFNC) analysis. In this study, we investigated the dFNC between 53 independent components among 188 ASD and 195 TD subjects. We estimated dFNC using sliding window-based approaches and identified four distinct dynamic states through hard-clustering analysis. Hyper-connectivity within the cognitive control domain, between cognitive control and default mode network, has been identified among ASD subjects. Hyper-connectivity within the default mode network has been found among TD individuals. Further, we estimated the dynamic temporal properties such as fractional time spent, and mean dwell time per state and observed significant differences between ASD and TD groups. ASD subjects are found to have significantly longer dwell time in state 4 when compared to TD individuals. We also found a significantly increased occurrence of state 4 in ASD subjects and states 1 and 3 in TD subjects. While there is broad consensus in the brain network profiles between static functional connectivity (sFNC) and dFNC, the temporal profile of brain state dynamics is additionally available with dFNC analysis and may potentially contribute to disease biomarkers.
Pindi Krishna Chandra Prasad, Kamalaker Dadi, Raju S. Bapi
IJCNN3
2023 Application of Graph Theoretic measures for assessing efficacy of Stroke Rehabilitation
abstract
One of the challenges in stroke rehabilitation is to identify bio-markers that correlate with amelioratory changes in the recovery of brain function that can be verified by a clinician. For strokes related to upper extremity, clinicians use Fugl Meyer Assessment - Upper Extremity (FMA-UE) score for verification. We hypothesize that even before clinical measures of function recovery (FMA-UE) show facilitatory changes, structural and functional changes in the brain connectome indicate early changes in stroke patients. Toward establishing such early bio-markers of rehabilitation related changes in the brain plasticity, we propose to use graph theoretic measures on the structural (SC) and functional connectivity (FC) matrices. We used longitudinal multi-modal neuroimaging data acquired within 1 to 6 months of the onset of stroke and after 3 months of rehabilitation for 15 acute ischemic stroke subjects with deficit in motor function of the upper extremity. We compared structural and functional network properties of the brain between the baseline (pre-) and post-rehabilitation stages. While significant changes are observed in the nodal properties, global network properties did not reach statistical significance. However, when correlation patterns across pre-, post-rehabilitation, and healthy controls are investigated, some interesting patterns emerge that are not captured by statistical analysis as well as the clinical assessment scores. Global functional network properties of patients post-rehabilitation resemble that of a healthy control group, than that of pre-rehabilitation (Baseline). We also observed that greater the change in normalized percent change in FMA-UE scores, closer are the correlation maps of global graph metrics of FC to that of the healthy controls. We did not observe any such resemblance in correlation patterns in structural connectivity (SC) of pre- and post-rehabilitation with that of the healthy control group. Overall, the results point out the suitability of using graph metrics for characterizing early markers of functional change in the brain regions.
Upadrasta Naga Sita Sravanthi, Kamalakar Dadi, Raju S. Bapi, P. N. Sylaja, C. Kesavadas, Srijithesh PR, Rinta Paul
IJCNN3
2023 Modelling Grid Navigation Using Reinforcement Learning Linear Ballistic Accumulators
abstract
Reinforcement Learning (RL) models constitute an important subset of models used in studying many facets of human learning including Motor Sequence Learning. However, conventional action selection in RL models such as softmax-based choice rules lack biological plausibility and do not offer mechanistic explanations. Furthermore, they also do not use response time data in model fitting, which can be indicative of the difference in value of alternate choices as perceived by subjects. Evidence Accumulation Models(EAM) such as Linear Ballistic Accumulators(LBA) provide a solution to many of the above problems. In this study, we use RL algorithms integrated with an LBA model to model human behaviour in the Grid-Sailing Task. The task involves navigating a grid to reach a goal position where the participant can choose from three possible actions. We fit RLLBA models using three different RL algorithms: a model using only Model Based updates, and two models that arbitrates between Model Based and Model Free learning, where Weight Based Arbitration is used in one and Value of Information(Vol) based Arbitration is used in the other. When following the actions of human subjects, we find a significant negative correlation between the Variance in Q-values and Response Time, motivating a competition based mechanism of action selection. When the three models are fit to data we find that VoI based models provide the best fit to data. We find that such models are able to predict actions made by human participants with accuracy comparable to that of conventional RL models and also predict the response times taken by the subject within reasonable margins. We discuss the implications of these results and outline scenarios where it would be advantageous to use RLEAM especially in the domain of Motor Learning.
Gautham Venugopal, Raju S. Bapi
IJCNN2
2023 Speech Taskonomy: Which Speech Tasks are the most Predictive of fMRI Brain Activity?
abstract
International audience
Subba Reddy Oota, Veeral Agarwal, Mounika Marreddy, Manish Gupta 0001, Raju S. Bapi
INTERSPEECH5
2022 Deep Learning for Brain Encoding and Decoding
Subba Reddy Oota, Jashn Arora, Manish Gupta 0001, Raju S. Bapi, Mariya Toneva
CogSci4
2022 Relative Numerical Context Affects Temporal Processing
Anuj Shukla, Raju S. Bapi
CogSci2
2022 Multi-view and Cross-view Brain Decoding
abstract
Can we build multi-view decoders that can decode concepts from brain recordings corresponding to any view (picture, sentence, word cloud) of stimuli? Can we build a system that can use brain recordings to automatically describe what a subject is watching using keywords or sentences? How about a system that can automatically extract important keywords from sentences that a subject is reading? Previous brain decoding efforts have focused only on single view analysis and hence cannot help us build such systems. As a first step toward building such systems, inspired by Natural Language Processing literature on multi-lingual and cross-lingual modeling, we propose two novel brain decoding setups: (1) multi-view decoding (MVD) and (2) cross-view decoding (CVD). In MVD, the goal is to build an MV decoder that can take brain recordings for any view as input and predict the concept. In CVD, the goal is to train a model which takes brain recordings for one view as input and decodes a semantic vector representation of another view. Specifically, we study practically useful CVD tasks like image captioning, image tagging, keyword extraction, and sentence formation. Our extensive experiments lead to MVD models with ~0.68 average pairwise accuracy across view pairs, and also CVD models with ~0.8 average pairwise accuracy across tasks. Analysis of the contribution of different brain networks reveals exciting cognitive insights: (1) Models trained on picture or sentence view of stimuli are better MV decoders than a model trained on word cloud view. (2) Our extensive analysis across 9 broad regions, 11 language sub-regions and 16 visual sub-regions of the brain help us localize, for the first time, the parts of the brain involved in cross-view tasks like image captioning, image tagging, sentence formation and keyword extraction. We make the code publicly available.
Subba Reddy Oota, Jashn Arora, Manish Gupta 0001, Raju S. Bapi
COLING4
2022 Visio-Linguistic Brain Encoding
abstract
Brain encoding aims at reconstructing fMRI brain activity given a stimulus. There exists a plethora of neural encoding models which study brain encoding for single mode stimuli: visual (pretrained CNNs) or text (pretrained language models). Few recent papers have also obtained separate visual and text representation models and performed late-fusion using simple heuristics. However, previous work has failed to explore the co-attentive multi-modal modeling for visual and text reasoning. In this paper, we systematically explore the efficacy of image and multi-modal Transformers for brain encoding. Extensive experiments on two popular datasets, BOLD5000 and Pereira, provide the following insights. (1) We find that VisualBERT, a multi-modal Transformer, significantly outperforms previously proposed single-mode CNNs, image Transformers as well as other previously proposed multi-modal models, thereby establishing new state-of-the-art. (2) The regions such as LPTG, LMTG, LIFG, and STS which have dual functionalities for language and vision, have higher correlation with multi-modal models which reinforces the fact that these models are good at mimicing the human brain behavior. (3) The supremacy of visio-linguistic models raises the question of whether the responses elicited in the visual regions are affected implicitly by linguistic processing even when passively viewing images. Future fMRI tasks can verify this computational insight in an appropriate experimental setting. We make our code publicly available.
Subba Reddy Oota, Jashn Arora, Vijay Rowtula, Manish Gupta 0001, Raju S. Bapi
COLING5
2022 A Phenomenological Deep Oscillatory Neural Network Model to Capture the Whole Brain Dynamics in Terms of BOLD Signal
Anirban Bandyopadhyay, Dipayan Biswas, Raju S. Bapi, V. Srinivasa Chakravarthy
ICONIP (2)4
2022 Multiple Kernel Learning for Modeling Resting State EEG Connectomes using Structural Connectivity of the Brain
abstract
An active area of research in cognitive science is characterizing the relationship between brain structure and the observed functional activations. Recent graph diffusion models have had great success in mapping whole-brain, resting-state dynamics measured using functional Magnetic Resonance Imaging (fMRI) to the brain structure derived using diffusion and T1 brain imaging. Here we test the application of one such graph diffusion method called the Multiple Kernel Learning (MKL) model. MKL model, formulated as a reaction-diffusion system using Wilson-Cowan equations, combines multiple diffusion kernels at different scales to predict functional connectome (FC) arising from a fixed structural connectome (SC). Our simulation results demonstrate that the MKL model successfully mapped the relationship between SC and FC from five different Electroen-cephalogram (EEG) bands (delta, theta, alpha, beta, and gamma). We used simultaneously acquired EEG-fMRI and NODDI dataset of 17 participants. The correlation between predicted FC and ground truth FC was higher for EEG bands than for fMRI data. The prediction accuracy peaked for the alpha band, and the highest frequency band, gamma had the lowest prediction accuracy. To the best of our knowledge, this is the first such end-to-end application of multiple kernel graph diffusion framework for modeling EEG data. One of the important features of MKL model is its ability to incorporate structural connectivity features into the generative model that predicts the EEG functional connectivity.
P. L. Ammar Ahmed, Archi Yadav, Avinash Sharma 0001, Raju S. Bapi
IJCNN4
2022 Characterizing the Dynamic Reorganization in Healthy Ageing and Classification of Brain Age
abstract
During healthy ageing, the brain networks undergo various topological and functional alterations. Previous studies have shown that the dedifferentiation of the functional modules could be one of the hallmarks of large-scale brain networks and alterations through the lifespan. This modular organization and alterations may be critically linked to a variety of neurodegenerative disorders and cognitive deficits encountered during ageing. In spite of accumulating evidence based on tracking static functional connectivity (FC) and modularity in characterizing dedifferentiation associated with ageing, there is a gap in understanding the brain dynamics of modular segregation and integration through the lifespan. Using the Cam-CAN dataset (young: 18–44, mean 32 years, old: 65–88, mean 75 years), we characterize the modular reorganization using dynamic measures like flexibility, to find characteristic nodes that make up the stable core and flexible periphery in the young and old age groups. In this study, we hypothesize that the nodes that exhibit higher flexibility in the older age groups will be negatively correlated with modularity since these nodes ‘compensate’ for the functional integration while ensuring that the segregation is efficient. Our results demonstrate that the regions from the Default Mode network (DMN) show a negative correlation with modularity in the old age groups. Further, nodes from Limbic, SensoriMotor (SMN) and Salience networks show a positive correlation with modularity. These networks that are responsible for higher-order cognitive functions, e.g., decision making, attentional control, cognitive flexibility are found to make up a stable core as evidenced by their low flexibility scores. We also trained various classifiers using node flexibility scores as features for the binary (young vs old) classification task. Support Vector Machine (SVM) with Gaussian kernel trained on a reduced-dimensional feature set gave the best classification results. The features (nodes) that are found to be important for classification concur with those identified through the data-driven network measures based analysis. In summary, we anticipate that these findings can help identify the regions that are responsible for the reorganization and maintenance of a rich diversity of functional repertoire in healthy ageing.
Arpita Dash, Raju S. Bapi, Dipanjan Roy, P. K. Vinod
IJCNN2
2022 Multiple GraphHeat Networks for Structural to Functional Brain Mapping
abstract
Over the last decade, there has been growing interest in learning the mapping from structural connectivity (SC) to functional connectivity (FC) of the brain. The spontaneous brain activity fluctuations during the resting-state as captured by functional MRI (rsfMRI) contain rich non-stationary dynamics over a relatively fixed structural connectome. Among the modeling approaches, graph diffusion-based methods with single and multiple diffusion kernels approximating static or dynamic functional connectivity have shown promise in predicting the FC given the SC. However, these methods are computationally expensive, not scalable, and fail to capture the complex dynamics underlying the whole process. Recently, deep learning methods such as GraphHeat networks along with graph diffusion have been shown to handle complex relational structures while preserving global information. In this paper, we propose multiple GraphHeat networks (M-GHN), a novel approach for mapping SC-FC. M-GHN enables us to model multiple heat kernel diffusion over the brain graph for approximating the complex Reaction Diffusion phenomenon. We argue that the proposed deep learning method overcomes the scalability and computational inefficiency issues but can still learn the SC-FC mapping successfully. Training and testing were done using the rsfMRI data of 100 participants from the human connectome project (HCP), and the results establish the viability of the proposed model. On the HCP dataset of 100 participants, the M-GHN achieves a high Pearson correlation of 0.747. Furthermore, experiments demonstrate that M-GHN outperforms the existing methods in learning the complex nature of human brain function.
Subba Reddy Oota, Archi Yadav, Arpita Dash, Raju S. Bapi, Avinash Sharma 0001
IJCNN4
2022 Deep Learning Approach for Classification and Interpretation of Autism Spectrum Disorder
abstract
Autism spectrum disorder (ASD) is a neurodevelopmental disorder predominantly found in children. The current behavior-based diagnosis of ASD is arduous and requires expertise. Therefore, it is appealing to develop an accurate computer-aided tool for diagnosing ASD. Although resting-state functional magnetic resonance imaging (rsfMRI) has proven to be successful in capturing the neural organization of the brain, automated detection of ASD using rsfMRI scans is a challenging task due to heterogeneity in the dataset and limited sample size. This paper proposes a Multilayer Perceptron (MLP) based classification model with auto encoder pretraining for classifying ASD from Typically Developing (TD) using rsfMRI scans obtained from the ABIDE-1 dataset. Our model achieves new state-of-the-art performance on the ABIDE-1 dataset with a 10-fold cross-validation accuracy of 74.82%. Further, we use the Integrated Gradients (IG) and DeepLIFT techniques to identify the correlations between brain regions that contribute most to the classification task. Our analysis identifies the following regions, Left Lingual Gyrus, Right Insula Lobe, Right Cuneus, Right Middle Frontal Gyrus, Left Superior Temporal Gyrus to be associated with ASD. Interestingly, these regions in the brain are primarily responsible for social cognition, language, attention, decision making and visual processing, which are known to be altered in ASD.
Pindi Krishna Chandra Prasad, Yash Khare, Kamalaker Dadi, P. K. Vinod, Raju S. Bapi
IJCNN5
2022 mulEEG: A Multi-view Representation Learning on EEG Signals
Vamsi Kumar, Likith Reddy, Shivam Kumar Sharma, Kamalaker Dadi, Chiranjeevi Yarra, Raju S. Bapi, Srijithesh Rajendran
MICCAI (3)6
2022 Neural Language Taskonomy: Which NLP Tasks are the most Predictive of fMRI Brain Activity?
abstract
Several popular Transformer based language models have been found to be successful for text-driven brain encoding. However, existing literature leverages only pretrained text Transformer models and has not explored the efficacy of task-specific learned Transformer representations. In this work, we explore transfer learning from representations learned for ten popular natural language processing tasks (two syntactic and eight semantic) for predicting brain responses from two diverse datasets: Pereira (subjects reading sentences from paragraphs) and Narratives (subjects listening to the spoken stories). Encoding models based on task features are used to predict activity in different regions across the whole brain. Features from coreference resolution, NER, and shallow syntax parsing explain greater variance for the reading activity. On the other hand, for the listening activity, tasks such as paraphrase generation, summarization, and natural language inference show better encoding performance. Experiments across all 10 task representations provide the following cognitive insights: (i) language left hemisphere has higher predictive brain activity versus language right hemisphere, (ii) posterior medial cortex, temporo-parieto-occipital junction, dorsal frontal lobe have higher correlation versus early auditory and auditory association cortex, (iii) syntactic and semantic tasks display a good predictive performance across brain regions for reading and listening stimuli resp.
Subba Reddy Oota, Jashn Arora, Veeral Agarwal, Mounika Marreddy, Manish Gupta 0001, Raju S. Bapi
NAACL-HLT6
2021 IMLE-Net: An Interpretable Multi-level Multi-channel Model for ECG Classification
abstract
Early detection of cardiovascular diseases is crucial for effective treatment and an electrocardiogram (ECG) is pivotal for diagnosis. The accuracy of Deep Learning based methods for ECG signal classification has progressed in recent years to reach cardiologist-level performance. In clinical settings, a cardiologist makes a diagnosis based on the standard 12-channel ECG recording. Automatic analysis of ECG recordings from a multiple-channel perspective has not been given enough attention, so it is essential to analyze an ECG recording from a multiple-channel perspective. We propose a model that leverages the multiple-channel information available in the standard 12-channel ECG recordings and learns patterns at the beat, rhythm, and channel level. The experimental results show that our model achieved a macro-averaged ROC-AUC score of 0.9216, mean accuracy of 88.85% and a maximum F1 score of 0.8057 on the PTB-XL dataset. The attention visualization results from the interpretable model are compared against the cardiologist’s guidelines to validate the correctness and usability.
Likith Reddy, Vivek Talwar, Shanmukh Alle, Raju S. Bapi, U. Deva Priyakumar
SMC4
2020 Value-of-Information based Arbitration between Model-based and Model-free Control
Krishn Bera, Yash Mandilwar, Anuj Shukla, Raju S. Bapi
CogSci4
2020 Motor Chunking During Sequence Learning in Grid-Navigation Tasks
Krishn Bera, Anuj Shukla, Raju S. Bapi
CogSci3
2020 Grid-Navigation Tasks involve Skill Learning
Krishn Bera, Anuj Shukla, Raju S. Bapi
CogSci3
2020 HED-ID: An Affective Adaptation Model Explaining the Intensity-Duration Relationship of Emotion
abstract
Intensity and duration are both pertinent aspects of an emotional experience, yet how they are related is unclear. Though stronger emotions usually last longer, sometimes they abate faster than the weaker ones. We present a quantitative model of affective adaptation, the process by which emotional responses to unchanging affective stimuli weaken with time, that addresses this intensity-duration problem. The model, described by three simple linear algebraic equations, assumes that the relationship between an affective stimulus and its experiencer can be broken down into three parameters. Self-relevance and explanation level combine multiplicatively to determine emotion intensity whereas the interaction of these with explanatory ease determines its duration. The model makes predictions, consistent with available empirical data, about emotion intensity, its duration, and adaptation speed for different scenarios. It predicts when the intensity-duration correlation is positive, negative or even absent, thus offering a solution to the intensity-duration problem. The model also addresses the shortcomings of past models of affective adaptation with its enhanced predictive power and by offering a more complete explanation to empirical observations that earlier models explain inadequately or fail to explain altogether. The model has potential applications in areas such as virtual reality training, games, human-computer interactions, and robotics.
John Eric Steephen, Siva C. Obbineni, Sneha Kummetha, Raju S. Bapi
IEEE Trans. Affect. Comput.4
2019 StepEncog: A Convolutional LSTM Autoencoder for Near-Perfect fMRI Encoding
abstract
Learning a forward mapping that relates stimuli to the corresponding brain activation measured by functional magnetic resonance imaging (fMRI) is termed as estimating encoding models. Computational tractability usually forces current encoding as well as decoding solutions to typically consider only a small subset of voxels from the actual 3D volume of activation. Further, while reconstructing stimulus information from brain activation (brain decoding) has received wider attention, there have been only a few attempts at constructing encoding solutions in the extant neuro-imaging literature. In this paper, we present StepEncog, a convolutional LSTM autoencoder model trained on fMRI voxels. The model can predict the entire brain volume rather than a small subset of voxels, as presented in earlier research works. We argue that the resulting solution avoids the problem of devising encoding models based on a rule-based selection of informative voxels and the concomitant issue of wide spatial variability of such voxels across participants. The perturbation experiments indicate that the proposed deep encoder indeed learns to predict brain activations with high spatial accuracy. On challenging universal decoder imaging datasets, our model yielded encouraging results.
Subba Reddy Oota, Vijay Rowtula, Manish Gupta 0001, Raju S. Bapi
IJCNN4
2018 fMRI Semantic Category Decoding Using Linguistic Encoding of Word Embeddings
Subba Reddy Oota, Naresh Manwani, Raju S. Bapi
ICONIP (3)3
2017 Effects of Variable Response-Stimulus Interval (RSI) On Sequence Learning
Sneha Kummetha, Anuj Shukla, Raju S. Bapi
CogSci3
2017 Experimental and Computational Investigation of the Effect of Caffeine on Human Time Perception
Remya Sankar, Anuj Shukla, Raju S. Bapi
CogSci3
2017 A biologically inspired neuronal model of reward prediction error computation
abstract
The neurocomputational model described here proposes that two dimensions involved in computation of reward prediction errors i.e magnitude and time could be computed separately and later combined unlike traditional reinforcement learning models. The model is built on biological evidences and is able to reproduce various aspects of classical conditioning, namely, the progressive cancellation of the predicted reward, the predictive firing from conditioned stimuli, and delineation of early rewards by showing firing for sooner early rewards and not for early rewards that occur with a longer latency in accordance with biological data.
Pramod S. Kaushik, Maxime Carrere, Frédéric Alexandre, Raju S. Bapi
IJCNN4
2017 Metastability of cortical BOLD signals in maturation and senescence
abstract
We assess change in metastability to characterize age-effects on the dynamic repertoire of the functional networks at rest. Resting state fMRI signals from each subject (N=48) have been used and metastability is evaluated as the standard deviation of mean phase synchrony of BOLD signals across whole-brain as well as across known resting state networks. The results suggest that significant whole-brain metastability changes occur between middle to old age. We also demonstrate that static time-averaged FC largely undermines age-effects on the interaction between functional networks. Discriminant Function Analysis reveals existence of two different patterns of change in metastability, which maximally discriminates between two different processes of maturation and ageing.
Shruti Naik, Subba Reddy Oota, Arpan Banerjee, Dipanjan Roy, Raju S. Bapi
IJCNN5
2017 The art of scaling up : A computational account on action selection in basal ganglia
abstract
What makes a computational neuronal model `large scale'? Is it the number of neurons modeled? Or the number of brain regions modeled in a network? Most of the higher cognitive processes span across co-ordinated activity in a network of different brain areas. However at the same time, the basic information transfer takes place at a single neuron level, together with multiple other neurons. We explore modeling a neural system involving some areas of cortex, the basal ganglia (BG) and thalamus for the process of decision making, using a large-scale neural engineering framework, Nengo. Early results tend to replicate the known neural activity patterns as found in the previous action selection model by Guthrie et al. 2013, besides operating with a larger neuronal populations. The power of converting algorithms to efficiently weighed neural networks in Nengo (Stewart et al. 2009 and Bekolay et al. 2013) is exploited in this work. Crucial aspects in a computational model, like parameter tuning and detailed neural implementations, while moving from a simplistic to large-scale model, are studied.
Bhargav Teja Nallapu, Raju S. Bapi, Nicolas P. Rougier
IJCNN2
2016 Artificial bee colony algorithm for clustering: an extreme learning approach
Abobakr Khalil Alshamiri, Alok Singh 0001, Raju S. Bapi
Soft Comput.3
2014 Inter Subject Correlation of Brain Activity during Visuo-Motor Sequence Learning
Krishna P. Miyapuram, Ujjval Pamnani, Kenji Doya, Raju S. Bapi
ICONIP (1)4
2014 A survey of distance/similarity measures for categorical data
abstract
Similarity or distance between two objects plays a fundamental role in many data mining tasks like classification and clustering. Categorical data, unlike numeric data, conceptually is deficient of default ordering relations on the attribute values. This makes the task of devising similarity or distance metrics and data mining tasks such as classification and clustering of categorical data more challenging. In this paper we formulate a taxonomy of various distance or similarity measures used in conjunction with data whose attributes are categorical. We categorize the existing measures into two broad classes, namely, Context-free and Context-sensitive measures for categorical data. In addition, we suggest a taxonomy of the clustering approaches for categorical data. We also propose a hybrid approach for measuring similarity between objects. We make a relative comparison of the strengths and weaknesses of some of the similarity measures and point out future research directions.
Madhavi Alamuri, Raju S. Bapi, Atul Negi
IJCNN2
2014 Reinforcement learning and dopamine in the striatum: A modeling perspective
Shesharao M. Wanjerkhede, Raju S. Bapi, Vithal D. Mytri
Neurocomputing2
2013 The influence of spatial cueing on serial order visual memory
Rakesh Sengupta, Anvita Gopal, Prajit Basu, David Melcher, Raju S. Bapi
CogSci5
2010 Rule Extraction from Support Vector Machine Using Modified Active Learning Based Approach: An Application to CRM
M. A. H. Farquad, Vadlamani Ravi, Raju S. Bapi
KES (1)3
2010 Support vector regression based hybrid rule extraction methods for forecasting
M. A. H. Farquad, Vadlamani Ravi, Raju S. Bapi
Expert Syst. Appl.3
2009 A Computational study of pre-synaptic re-uptake of dopamine on phosphorylation of DARPP-32
abstract
The activity of mid brain dopamine (DA) neurons resembles the reward prediction error signal of the temporal difference (TD) model. Dopamine acting at D1 receptor activates adenylate cyclase through the coupling of D1 receptors with Gs/olf. The complex of adenylate cyclase activates and increases the intracellular levels of cAMP and thereby stimulating cAMP dependent protein kinase and increasing the state of phosphorylation of DARPP-32 (dopamine and cyclic AMP-regulated phosphoprotein, M, = 32,000). DARPP-32 plays an important role in autophosphorylation of calcium calmodulin dependent protein kinaseII (CaMKII) which is crucial for prolongation of long term potentiation (LTP). Here we verified the effect of pre-synaptic re-uptake of dopamine by dopamine transporter on the phosphorylation of DARPP-32 and found that the phosphorylation of DARPP-32 is unaffected by pre-synaptic re-uptake. Further we have seen that the DARPP-32 phosphorylation does not change with the change in the DA or D1R concentrations.
Shesharao M. Wanjerkhede, Raju S. Bapi
IJCNN2
2007 A Tighter Error Bound for Decision Tree Learning Using PAC Learnability
Chaithanya Pichuka, Raju S. Bapi, Chakravarthy Bhagvati, Arun K. Pujari, Bulusu Lakshmana Deekshatulu
IJCAI2
2007 Detection of Cognitive States from fMRI Data Using Machine Learning Techniques
Vishwajeet Singh, Krishna P. Miyapuram, Raju S. Bapi
IJCAI3
2007 Analysis of E.coli promoter recognition problem in dinucleotide feature space
abstract
MOTIVATION: Patterns in the promoter sequences within a species are known to be conserved but there exist many exceptions to this rule which makes the promoter recognition a complex problem. Although many complex feature extraction schemes coupled with several classifiers have been proposed for promoter recognition in the current literature, the problem is still open. RESULTS: A dinucleotide global feature extraction method is proposed for the recognition of sigma-70 promoters in Escherichia coli in this article. The positive data set consists of sigma-70 promoters with known transcription starting points which are part of regulonDB and promec databases. Four different kinds of negative data sets are considered, two of them biological sets (Gordon et al., 2003) and the other two synthetic data sets. Our results reveal that a single-layer perceptron using dinucleotide features is able to achieve an accuracy of 80% against a background of biological non-promoters and 96% for random data sets. A scheme for locating the promoter regions in a given genome sequence is proposed. A deeper analysis of the data set shows that there is a bifurcation of the data set into two distinct classes, a majority class and a minority class. Our results point out that majority class constituting the majority promoter and the majority non-promoter signal is linearly separable. Also the minority class is linearly separable. We further show that the feature extraction and classification methods proposed in the paper are generic enough to be applied to the more complex problem of eucaryotic promoter recognition. We present Drosophila promoter recognition as a case study. AVAILABILITY: http://202.41.85.117/htmfiles/faculty/tsr/tsr.html.
T. Sobha Rani, Durga Bhavani S., Raju S. Bapi
Bioinform.3
2007 Rough clustering of sequential data
Pradeep Kumar 0001, P. Radha Krishna 0001, Raju S. Bapi, Supriya Kumar De
Data Knowl. Eng.3
2006 Clustering using Similarity Upper Approximation
abstract
Rough set theory operates on an information system that consists of a set of objects. A core concept of rough set theory is that of equivalence between objects called indiscernibility. Indiscernibility reflects a total impossibility of distinguishing between objects, considering the available information. Considering a tolerance or similarity relation instead of an indiscernibility relation is quite relevant due to the existence of quantitative attributes in the information systems. Extending indiscernibility to tolerance relation results in weakening of some of the properties of the binary relation in terms of reflexivity, symmetry and transitivity. In this paper, we present a clustering technique using similarity relation with transitivity property being relaxed. The concept of similarity upper approximation has been used to form the initial family of cluster. A relationship based measure has been used to decide the belongingness of uncertain elements. We present an example to illustrate our proposed methodology. This promises to be a useful and interesting area of extension of the theory of rough sets.
Pradeep Kumar 0001, P. Radha Krishna 0001, Raju S. Bapi, Supriya Kumar De
FUZZ-IEEE3
2006 Hierarchical Chunking during Learning of Visuomotor Sequences
abstract
It is well known that learning a sequential skill involves chaining a number of primitive actions together into chunks. We describe three different experiments using an explicit visuomotor sequence learning paradigm called the m times n task. The m times n task enables hierarchical learning of sequences by presenting m elements of the sequence at a time (called the set). The entire sequence to be learned is composed of n such sets and is called a hyperset. In the first experiment, we showed the chunking phenomenon while learning a sequence as opposed to following randomly generated visual cues. We further explored the nature of chunking across sets using complex sequences in the second experiment. Finally, we investigated effector dependence of the chunking patterns in the third experiment. Our results point out the facilitating factors for chunk formation in visuomotor sequence learning.
Krishna P. Miyapuram, Raju S. Bapi, Chandrasekhar V. S. Pammi, Ahmed, Kenji Doya
IJCNN2
2006 MultisensorData Fusion Using Neural Networks
abstract
This paper presents a Hebbian learning based linear single-layer neural network based measurement fusion of multisensor data. The performance of the proposed unsupervised neural network algorithm is compared with traditional fusion methods based on Kalman filtering such as measurement fusion and state vector fusion. The experiments have been carried out using multisensor data obtained from different radars. The results demonstrate the viability of the proposed algorithm.
Narri Yadaiah, Lakshman Singh, Raju S. Bapi, V. Seshagiri Rao, Bulusu Lakshmana Deekshatulu, Atul Negi
IJCNN3
2005 Role of presynaptic reuptake on dopamine modulation of cortico-striatal activity in TD learning
abstract
It has been shown that midbrain dopamine neurons and the dopamine neuronal activity in the striatum mimic reward prediction error signal of the temporal difference learning (TDL) paradigm. James Houk proposed a theoretical model to explain the cellular basis of this dopamine activity. Our goal is to verify this model with simulations in GENESIS and CHEMESIS. This paper reports preliminary results from the simulations of this intracellular signaling scheme. The main result reported here is the influence of presynaptic reuptake on dopamine activity at the cortico-striatal synapse. It appears that the reuptake affects the rate of formation of D1/spl I.bar/DA receptor complex that in turn would have a bearing on the timing mechanisms operating in TDL.
Shesharao M. Wanjerkhede, Raju S. Bapi
IJCNN2
2005 Intrusion Detection System Using Sequence and Set Preserving Metric
Pradeep Kumar 0001, M. Venkateswara Rao, P. Radha Krishna 0001, Raju S. Bapi, Arijit Laha
ISI4
2004 State Estimation and Tracking Problems: A Comparison Between Kalman Filter and Recurrent Neural Networks
S. Kumar Chenna, Yogesh Kr. Jain, Himanshu Kapoor, Raju S. Bapi, Narri Yadaiah, Atul Negi, V. Seshagiri Rao, Bulusu Lakshmana Deekshatulu
ICONIP4
2004 Chunking Phenomenon in Complex Sequential Skill Learning in Humans
Chandrasekhar V. S. Pammi, Krishna P. Miyapuram, Raju S. Bapi, Kenji Doya
ICONIP3
2000 The Frontal Lobes and Executive Function
abstract
From the viewpoint of neural network modelling via nonlinear dynamical systems, it is useful to break down the overall executive function (EF) into smaller components that repeatedly occur in different tasks and in different combinations. The tasks whose models are described suggest three generic components that commonly occur, all related to aspects of pre-frontal function: 1) establishing links between working memory representations which could represent sensory stimuli and potential motor actions; 2) creating, learning, and deciding among high-level schemata that embody repeatable, but often flexible action sequences; and 3) incorporating affective evaluations of sensory events or potential motor plans and using these evaluations to guide actions. The network theory developed describes how the components of EF could arise from frontal system neuroanatomy.
John G. Taylor, Neill R. Taylor, Raju S. Bapi, Guido Bugmann, Daniel S. Levine 0001
IJCNN (1)3
1998 A Sequence Learning Architecture Based on Cortico-Basal Ganglionic Loops and Reinforcement Learning
Raju S. Bapi, Kenji Doya
ICONIP1
1997 Neuro-Resistive Grid Appraoch to Trainable Controllers: A Pole Balancing Example
Raju S. Bapi, Brendan D'Cruz, Guido Bugmann
Neural Comput. Appl.1
1994 Modeling the role of frontal lobes in sequential task performance. I. Basic structure and primacy effects
Raju S. Bapi, Daniel S. Levine 0001
Neural Networks1
1990 Networks modeling the involvement of the frontal lobes in learning and performance of flexible movement sequences
abstract
Network architectures for classifying spatiotemporal patterns are developed. These architectures combine the adaptive resonance architecture, which classifies spatial patterns, with the avalanche, which generates sequences. The primary spatiotemporal processing area is identified with all area of the corpus striatum. The prefrontal cortex is identified with higher-order controls of functions of this sequence-classifying area. Some of these controls make possible the formation of complex rules for sequence classification. Others cause competitive biases among sequence representations, favoring longer over shorter sequences to facilitate attention to a motor plan. The network models experimental results showing that frontal damage impairs the performance of flexible movement sequences but not of invariant movement sequences
Daniel S. Levine 0001, Raju S. Bapi
IJCNN2