VLDB 2026 Research / reviewers in the wild / expert
Subba Reddy Oota
dblp:190/1709 · also Oota Subbareddy
· DBLP profile ↗
29ranked-venue papers
18as first author
23since 2021 · last 2025
0000-0002-5975-622XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 15 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Large Language Models Are Human-Like Annotators
Mounika Marreddy, Subba Reddy Oota, Manish Gupta 0001 |
ECIR (5) | 2 |
| 2025 | Aligning Text/Speech Representations from Multimodal Models with MEG Brain Activity During ListeningabstractAlthough speech language models are expected to align well with brain language processing during speech comprehension, recent studies have found that they fail to capture brainrelevant semantics beyond low-level features.Surprisingly, text-based language models exhibit stronger alignment with brain language regions, as they better capture brain-relevant semantics.However, no prior work has examined the alignment effectiveness of text/speech representations from multimodal models.This raises several key questions: Can speech embeddings from such multimodal models capture brain-relevant semantics through cross-modal interactions?Which modality can take advantage of this synergistic multimodal understanding to improve alignment with brain language processing?Can text/speech representations from such multimodal models outperform unimodal models?To address these questions, we systematically analyze multiple multimodal models, extracting both text-and speech-based representations to assess their alignment with MEG brain recordings during naturalistic story listening.We find that text embeddings from both multimodal and unimodal models significantly outperform speech embeddings from these models.Specifically, multimodal text embeddings exhibit a peak around 200 ms, suggesting that they benefit from speech embeddings, with heightened activity during this time period.However, speech embeddings from these multimodal models still show a similar alignment compared to their unimodal counterparts, suggesting that they do not gain meaningful semantic benefits over text-based representations.These results highlight an asymmetry in cross-modal knowledge transfer, where the text modality benefits more from speech information, but not vice versa.We make the code publicly available 1 . Padakanti Srijith, Khushbu Pahwa, Radhika Mamidi, Raju S. Bapi, Manish Gupta 0001, Subba Reddy Oota |
EMNLP | 6 |
| 2025 | Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)abstractTransformer-based language models, though not explicitly trained to mimic brain recordings, have demonstrated surprising alignment with brain activity. Progress in these models—through increased size, instruction-tuning, and multimodality—has led to better representational alignment with neural data. Recently, a new class of instruction-tuned multimodal LLMs (MLLMs) have emerged, showing remarkable zero-shot capabilities in open-ended multimodal vision tasks. However, it is unknown whether MLLMs, when prompted with natural instructions, lead to better brain alignment and effectively capture instruction-specific representations. To address this, we first investigate the brain alignment, i.e., measuring the degree of predictivity of neural visual activity using text output response embeddings from MLLMs as participants engage in watching natural scenes. Experiments with 10 different instructions (like image captioning, visual question answering, etc.) show that MLLMs exhibit significantly better brain alignment than vision-only models and perform comparably to non-instruction-tuned multimodal models like CLIP. We also find that while these MLLMs are effective at generating high-quality responses suitable to the task-specific instructions, not all instructions are relevant for brain alignment. Further, by varying instructions, we make the MLLMs encode instruction-specific visual concepts related to the input image. This analysis shows that MLLMs effectively capture count-related and recognition-related concepts, demonstrating strong alignment with brain activity. Notably, the majority of the explained variance of the brain encoding models is shared between MLLM embeddings of image captioning and other instructions. These results indicate that enhancing MLLMs' ability to capture more task-specific information could allow for better differentiation between various types of instructions, and hence improve their precision in predicting brain responses. Subba Reddy Oota, Akshett Rai Jindal, Ishani Mondal, Khushbu Pahwa, Satya Sai Srinath Namburi, Manish Shrivastava 0001, Maneesh Kumar Singh 0001, Raju S. Bapi, Manish Gupta 0001 |
ICLR | 1 |
| 2025 | Multi-modal brain encoding models for multi-modal stimuliabstractDespite participants engaging in unimodal stimuli, such as watching images or silent videos, recent work has demonstrated that multi-modal Transformer models can predict visual brain activity impressively well, even with incongruent modality representations. This raises the question of how accurately these multi-modal models can predict brain activity when participants are engaged in multi-modal stimuli. As these models grow increasingly popular, their use in studying neural activity provides insights into how our brains respond to such multi-modal naturalistic stimuli, i.e., where it separates and integrates information across modalities through a hierarchy of early sensory regions to higher cognition (language regions). We investigate this question by using multiple unimodal and two types of multi-modal models—cross-modal and jointly pretrained—to determine which type of models is more relevant to fMRI brain activity when participants are engaged in watching movies (videos with audio). We observe that both types of multi-modal models show improved alignment in several language and visual regions. This study also helps in identifying which brain regions process unimodal versus multi-modal information. We further investigate the contribution of each modality to multi-modal alignment by carefully removing unimodal features one by one from multi-modal representations, and find that there is additional information beyond the unimodal embeddings that is processed in the visual and language regions. Based on this investigation, we find that while for cross-modal models, their brain alignment is partially attributed to the video modality; for jointly pretrained models, it is partially attributed to both the video and audio modalities. These findings serve as strong motivation for the neuro-science community to investigate the interpretability of these models for deepening our understanding of multi-modal information processing in brain. Subba Reddy Oota, Khushbu Pahwa, Mounika Marreddy, Maneesh Kumar Singh 0001, Manish Gupta 0001, Raju S. Bapi |
ICLR | 1 |
| 2025 | Brain-Informed Fine-Tuning for Improved Multilingual Understanding in Language ModelsabstractRecent studies have demonstrated that fine-tuning language models with brain data can improve their semantic understanding, although these findings have so far been limited to English. Interestingly, similar to the shared multilingual embedding space of pretrained multilingual language models, human studies provide strong evidence for a shared semantic system in bilingual individuals. Here, we investigate whether fine-tuning language models with bilingual brain data changes model representations in a way that improves them across multiple languages. To test this, we fine-tune monolingual and multilingual language models using brain activity recorded while bilingual participants read stories in English and Chinese. We then evaluate how well these representations generalize to the bilingual participants’ first language, their second language, and several other languages that the participants are not fluent in. We assess the fine-tuned language models on brain encoding performance and downstream NLP tasks. Our results show that bilingual brain-informed fine-tuned language models outperform their vanilla (pretrained) counterparts in both brain encoding performance and most downstream NLP tasks across multiple languages. These findings suggest that brain-informed fine-tuning improves multilingual understanding in language models, offering a bridge between cognitive neuroscience and NLP research. We make our code publicly available. Anuja Negi, Subba Reddy Oota, Anwar Nunez-Elizalde, Manish Gupta 0001, Fatma Deniz |
NeurIPS | 2 |
| 2024 | Speech language models lack important brain-relevant semanticsabstractDespite known differences between reading and listening in the brain, recent work has shown that text-based language models predict both text-evoked and speech-evoked brain activity to an impressive degree.This poses the question of what types of information language models truly predict in the brain.We investigate this question via a direct approach, in which we systematically remove specific lowlevel stimulus features (textual, speech, and visual) from language model representations to assess their impact on alignment with fMRI brain recordings during reading and listening.Comparing these findings with speech-based language models reveals starkly different effects of low-level features on brain alignment.While text-based models show reduced alignment in early sensory regions post-removal, they retain significant predictive power in late language regions.In contrast, speech-based models maintain strong alignment in early auditory regions even after feature removal but lose all predictive power in late language regions.These results suggest that speech-based models provide insights into additional information processed by early auditory regions, but caution is needed when using them to model processing in late language regions.We make our code publicly available.1 Subba Reddy Oota, Emin Çelik, Fatma Deniz, Mariya Toneva |
ACL (1) | 1 |
| 2024 | Modelling Cross-Situational Learning on Full Sentences in Few Shots with Simple RNNs
Xavier Hinaut, Subba Reddy Oota, Alexandre Variengien, Frédéric Alexandre |
CogSci | 2 |
| 2023 | Neural Architecture of SpeechabstractA vast literature on brain encoding has effectively harnessed deep neural network models for accurately predicting brain activations from visual or text stimuli. Unfortunately, there is not much work on brain encoding for speech stimuli. The few existing studies on brain encoding for speech stimuli transcribe speech to text and then leverage text-only models for encoding, thereby ignoring audio signals completely. However, recently several speech representation learning models have revolutionized the field of speech processing. Inspired by the recent progress on deep learning models for speech, we present a first systematic study on understanding human speech processing by probing neural speech models to predict both language and auditory brain region activations. In particular, we investigate 30 speech representation models grouped into four categories: (i) traditional feature engineering, (ii) generative, (iii) predictive, and (iv) contrastive, to study how these models encode the speech stimuli and align with human brain activity for the Moth Radio Hour fMRI (functional magnetic resonance imaging) dataset. We find that both contrastive (Wav2Vec2.0) and predictive models (HuBERT, Data2Vec) are very accurate. Specifically, Data2Vec aligns the best with both language and auditory brain regions among all investigated models. We make our code publicly available1. Subba Reddy Oota, Khushbu Pahwa, Mounika Marreddy, Manish Gupta 0001, Raju S. Bapi |
ICASSP | 1 |
| 2023 | GAE-ISUMM: Unsupervised Graph-based Summarization for Indian LanguagesabstractDocument summarization aims to create a precise and coherent summary of a text document. Many deep learning summarization models are developed mainly for English, often requiring a large training corpus and efficient pre-trained language models and tools. However, English summarization models for low-resource Indian languages are often limited by rich morphological variation, syntax, and semantic differences. In this paper, we propose GAE-ISUMM, an unsupervised Indic summarization model that extracts summaries from text documents. In particular, our proposed model, GAE-ISUMM uses Graph Autoencoder (GAE) to learn text representations and a document summary jointly. We also provide a manually-annotated Telugu summarization dataset TELSUM, to experiment with our model GAE-ISUMM. Further, we benchmark with the most publicly available Indian language summarization datasets to investigate the effectiveness of GAE-ISUMM. Our experiments of GAE-ISUMM on seven Indian languages make the following observations: (i) it is competitive or better than state-of-the-art results on all datasets, (ii) it reports benchmark results on TELSUM, and (iii) the inclusion of positional and cluster information in the proposed model improved the performance of summaries. We open-source our dataset and code11https://github.com/scsmuhio/Summarization. Lakshmi Sireesha Vakada, Anudeep Chaluvadi, Mounika Marreddy, Subba Reddy Oota, Radhika Mamidi |
IJCNN | 4 |
| 2023 | Speech Taskonomy: Which Speech Tasks are the most Predictive of fMRI Brain Activity?abstractInternational audience Subba Reddy Oota, Veeral Agarwal, Mounika Marreddy, Manish Gupta 0001, Raju S. Bapi |
INTERSPEECH | 1 |
| 2023 | MEG Encoding using Word Context Semantics in Listening StoriesabstractInternational audience Subba Reddy Oota, Nathan Trouvain, Frédéric Alexandre, Xavier Hinaut |
INTERSPEECH | 1 |
| 2023 | Joint processing of linguistic properties in brains and language modelsabstractLanguage models have been shown to be very effective in predicting brain recordings of subjects experiencing complex language stimuli. For a deeper understanding of this alignment, it is important to understand the correspondence between the detailed processing of linguistic information by the human brain versus language models. We investigate this correspondence via a direct approach, in which we eliminate information related to specific linguistic properties in the language model representations and observe how this intervention affects the alignment with fMRI brain recordings obtained while participants listened to a story. We investigate a range of linguistic properties (surface, syntactic, and semantic) and find that the elimination of each one results in a significant decrease in brain alignment. Specifically, we find that syntactic properties (i.e. Top Constituents and Tree Depth) have the largest effect on the trend of brain alignment across model layers. These findings provide clear evidence for the role of specific linguistic information in the alignment between brain and language models, and open new avenues for mapping the joint information processing in both systems. We make the code publicly available https://github.com/subbareddy248/lingprop-brain-alignment. Subba Reddy Oota, Manish Gupta 0001, Mariya Toneva |
NeurIPS | 1 |
| 2023 | WSNet: Towards An Effective Method for Wound Image SegmentationabstractMedical image segmentation is critical for effective computer-aided diagnosis and localization of ailments. Automated segmentation of wound regions from patient images can aid clinicians in measuring and managing chronic wounds and monitoring the wound healing trajectory. While there exists a plethora of work on general medical image segmentation, there is hardly any work on wound image analysis and segmentation. Existing methods are limited to segmenting a smaller subset of ulcers, such as foot ulcers, with no special processing for wound images. In this paper, we build segmentation models for eight different types of wound images. Wound image analysis is a challenging problem due to the lack of availability of extensive data (labeled or unlabeled), and annotation is also challenging due to the shortage of well-trained wound care clinicians. To handle these challenges, we contribute WoundSeg1, a large and diverse dataset of segmented wound images. Generic wound image segmentation is complex due to the heterogeneous appearance of wound area across images of similar wound types. We propose a novel image segmentation framework, WSNet, which leverages (a) wound-domain adaptive pretraining on a large unlabeled wound image collection and (b) a global-local architecture that utilizes full image and its patches to learn fine-grained details of heterogeneous wounds. On WoundSeg, we achieve a decent Dice score of 0.847. On existing AZH Woundcare and Medetec datasets, we establish a new state-of-the-art. Further, we show the impact of using segmentation for improving the accuracy of downstream tasks like wound area and volume prediction. Subba Reddy Oota, Vijay Rowtula, Shahid Saleem Mohammed, Minghsun Liu, Manish Gupta 0001 |
WACV | 1 |
| 2023 | Am I a Resource-Poor Language? Data Sets, Embeddings, Models and Analysis for four different NLP Tasks in Telugu LanguageabstractDue to the lack of a large annotated corpus, many resource-poor Indian languages struggle to reap the benefits of recent deep feature representations in Natural Language Processing (NLP) . Moreover, adopting existing language models trained on large English corpora for Indian languages is often limited by data availability, rich morphological variation, syntax, and semantic differences. In this paper, we explore the traditional to recent efficient representations to overcome the challenges of a low resource language, Telugu. In particular, our main objective is to mitigate the low-resource problem for Telugu. Overall, we present several contributions to a resource-poor language viz. Telugu. (i) a large annotated data (35,142 sentences in each task) for multiple NLP tasks such as sentiment analysis, emotion identification, hate-speech detection, and sarcasm detection, (ii) we create different lexicons for sentiment, emotion, and hate-speech for improving the efficiency of the models, (iii) pretrained word and sentence embeddings, and (iv) different pretrained language models for Telugu such as ELMo-Te , BERT-Te , RoBERTa-Te , ALBERT-Te , and DistilBERT-Te on a large Telugu corpus consisting of 8,015,588 sentences (1,637,408 sentences from Telugu Wikipedia and 6,378,180 sentences crawled from different Telugu websites). Further, we show that these representations significantly improve the performance of four NLP tasks and present the benchmark results for Telugu. We argue that our pretrained embeddings are competitive or better than the existing multilingual pretrained models: mBERT , XLM-R , and IndicBERT . Lastly, the fine-tuning of pretrained models show higher performance than linear probing results on four NLP tasks with the following F1-scores: Sentiment (68.72), Emotion (58.04), Hate-Speech (64.27), and Sarcasm (77.93). We also experiment on publicly available Telugu datasets (Named Entity Recognition, Article Genre Classification, and Sentiment Analysis) and find that our Telugu pretrained language models ( BERT-Te and RoBERTa-Te ) outperform the state-of-the-art system except for the sentiment task. We open-source our corpus, four different datasets, lexicons, embeddings, and code https://github.com/Cha14ran/DREAM-T. The pretrained Transformer models for Telugu are available at https://huggingface.co/ltrctelugu. Mounika Marreddy, Subba Reddy Oota, Lakshmi Sireesha Vakada, Venkata Charan Chinni, Radhika Mamidi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2022 | Deep Learning for Brain Encoding and Decoding
Subba Reddy Oota, Jashn Arora, Manish Gupta 0001, Raju S. Bapi, Mariya Toneva |
CogSci | 1 |
| 2022 | Long-Term Plausibility of Language Models and Neural Dynamics during Narrative Listening
Subba Reddy Oota, Frédéric Alexandre, Xavier Hinaut |
CogSci | 1 |
| 2022 | Multi-view and Cross-view Brain DecodingabstractCan we build multi-view decoders that can decode concepts from brain recordings corresponding to any view (picture, sentence, word cloud) of stimuli? Can we build a system that can use brain recordings to automatically describe what a subject is watching using keywords or sentences? How about a system that can automatically extract important keywords from sentences that a subject is reading? Previous brain decoding efforts have focused only on single view analysis and hence cannot help us build such systems. As a first step toward building such systems, inspired by Natural Language Processing literature on multi-lingual and cross-lingual modeling, we propose two novel brain decoding setups: (1) multi-view decoding (MVD) and (2) cross-view decoding (CVD). In MVD, the goal is to build an MV decoder that can take brain recordings for any view as input and predict the concept. In CVD, the goal is to train a model which takes brain recordings for one view as input and decodes a semantic vector representation of another view. Specifically, we study practically useful CVD tasks like image captioning, image tagging, keyword extraction, and sentence formation. Our extensive experiments lead to MVD models with ~0.68 average pairwise accuracy across view pairs, and also CVD models with ~0.8 average pairwise accuracy across tasks. Analysis of the contribution of different brain networks reveals exciting cognitive insights: (1) Models trained on picture or sentence view of stimuli are better MV decoders than a model trained on word cloud view. (2) Our extensive analysis across 9 broad regions, 11 language sub-regions and 16 visual sub-regions of the brain help us localize, for the first time, the parts of the brain involved in cross-view tasks like image captioning, image tagging, sentence formation and keyword extraction. We make the code publicly available. Subba Reddy Oota, Jashn Arora, Manish Gupta 0001, Raju S. Bapi |
COLING | 1 |
| 2022 | Visio-Linguistic Brain EncodingabstractBrain encoding aims at reconstructing fMRI brain activity given a stimulus. There exists a plethora of neural encoding models which study brain encoding for single mode stimuli: visual (pretrained CNNs) or text (pretrained language models). Few recent papers have also obtained separate visual and text representation models and performed late-fusion using simple heuristics. However, previous work has failed to explore the co-attentive multi-modal modeling for visual and text reasoning. In this paper, we systematically explore the efficacy of image and multi-modal Transformers for brain encoding. Extensive experiments on two popular datasets, BOLD5000 and Pereira, provide the following insights. (1) We find that VisualBERT, a multi-modal Transformer, significantly outperforms previously proposed single-mode CNNs, image Transformers as well as other previously proposed multi-modal models, thereby establishing new state-of-the-art. (2) The regions such as LPTG, LMTG, LIFG, and STS which have dual functionalities for language and vision, have higher correlation with multi-modal models which reinforces the fact that these models are good at mimicing the human brain behavior. (3) The supremacy of visio-linguistic models raises the question of whether the responses elicited in the visual regions are affected implicitly by linguistic processing even when passively viewing images. Future fMRI tasks can verify this computational insight in an appropriate experimental setting. We make our code publicly available. Subba Reddy Oota, Jashn Arora, Vijay Rowtula, Manish Gupta 0001, Raju S. Bapi |
COLING | 1 |
| 2022 | Multi-Task Text Classification using Graph Convolutional Networks for Large-Scale Low Resource LanguageabstractGraph Convolutional Networks (GCN) have achieved state-of-art results on single text classification tasks like sentiment analysis, emotion detection, etc. However, the performance is achieved by testing and reporting on resource-rich languages like English. Applying GCN for multi-task text classification is an unexplored area. Moreover, training a GCN or adopting an English GCN for Indian languages is often limited by data availability, rich morphological variation, syntax, and semantic differences. In this paper, we study the use of GCN for the Telugu language in single and multi-task settings for four natural language processing (NLP) tasks, viz. sentiment analysis (SA), emotion identification (EI), hate-speech (HS), and sarcasm detection (SAR). In order to evaluate the performance of GCN with one of the Indian languages, Telugu, we analyze the GCN based models with extensive experiments on four downstream tasks. In addition, we created an annotated Telugu dataset, TEL-NLP, for the four NLP tasks. Further, we propose a supervised graph reconstruction method, Multi-Task Text GCN (MT- Text GCN) on the Telugu that leverages to simultaneously (i) learn the low-dimensional word and sentence graph embeddings from word-sentence graph reconstruction using graph autoencoder (GAE) and (ii) perform multi-task text classification using these latent sentence graph embeddings. We argue that our proposed MT- Text GCN achieves significant improvements on TEL-NLP over existing Telugu pretrained word embeddings [1], multilingual pretrained Transformer models: mBERT [2], and XLM-R [3]. On TEL-NLP, we achieve a high Fl-score for four NLP tasks: SA (0.84), EI (0.55), HS (0.83) and SAR (0.66). Finally, we show our model's quantitative and qualitative analysis on the four NLP tasks in Telugu. We open-source our TEL-NLP dataset, pretrained models, and code11https://github.com/scsmuhio/MTGCN_Resources. Mounika Marreddy, Subba Reddy Oota, Lakshmi Sireesha Vakada, Venkata Charan Chinni, Radhika Mamidi |
IJCNN | 2 |
| 2022 | Multiple GraphHeat Networks for Structural to Functional Brain MappingabstractOver the last decade, there has been growing interest in learning the mapping from structural connectivity (SC) to functional connectivity (FC) of the brain. The spontaneous brain activity fluctuations during the resting-state as captured by functional MRI (rsfMRI) contain rich non-stationary dynamics over a relatively fixed structural connectome. Among the modeling approaches, graph diffusion-based methods with single and multiple diffusion kernels approximating static or dynamic functional connectivity have shown promise in predicting the FC given the SC. However, these methods are computationally expensive, not scalable, and fail to capture the complex dynamics underlying the whole process. Recently, deep learning methods such as GraphHeat networks along with graph diffusion have been shown to handle complex relational structures while preserving global information. In this paper, we propose multiple GraphHeat networks (M-GHN), a novel approach for mapping SC-FC. M-GHN enables us to model multiple heat kernel diffusion over the brain graph for approximating the complex Reaction Diffusion phenomenon. We argue that the proposed deep learning method overcomes the scalability and computational inefficiency issues but can still learn the SC-FC mapping successfully. Training and testing were done using the rsfMRI data of 100 participants from the human connectome project (HCP), and the results establish the viability of the proposed model. On the HCP dataset of 100 participants, the M-GHN achieves a high Pearson correlation of 0.747. Furthermore, experiments demonstrate that M-GHN outperforms the existing methods in learning the complex nature of human brain function. Subba Reddy Oota, Archi Yadav, Arpita Dash, Raju S. Bapi, Avinash Sharma 0001 |
IJCNN | 1 |
| 2022 | Neural Language Taskonomy: Which NLP Tasks are the most Predictive of fMRI Brain Activity?abstractSeveral popular Transformer based language models have been found to be successful for text-driven brain encoding. However, existing literature leverages only pretrained text Transformer models and has not explored the efficacy of task-specific learned Transformer representations. In this work, we explore transfer learning from representations learned for ten popular natural language processing tasks (two syntactic and eight semantic) for predicting brain responses from two diverse datasets: Pereira (subjects reading sentences from paragraphs) and Narratives (subjects listening to the spoken stories). Encoding models based on task features are used to predict activity in different regions across the whole brain. Features from coreference resolution, NER, and shallow syntax parsing explain greater variance for the reading activity. On the other hand, for the listening activity, tasks such as paraphrase generation, summarization, and natural language inference show better encoding performance. Experiments across all 10 task representations provide the following cognitive insights: (i) language left hemisphere has higher predictive brain activity versus language right hemisphere, (ii) posterior medial cortex, temporo-parieto-occipital junction, dorsal frontal lobe have higher correlation versus early auditory and auditory association cortex, (iii) syntactic and semantic tasks display a good predictive performance across brain regions for reading and listening stimuli resp. Subba Reddy Oota, Jashn Arora, Veeral Agarwal, Mounika Marreddy, Manish Gupta 0001, Raju S. Bapi |
NAACL-HLT | 1 |
| 2021 | Clickbait Detection in Telugu: Overcoming NLP Challenges in Resource-Poor Languages using Benchmarked TechniquesabstractClickbait headlines have become a nudge in social media and news websites. The methods to identify clickbaits are largely being developed for English. There is a need for the same in other languages as well with the increase in the usage of social media platforms in different languages. In this work, we present an annotated clickbait dataset of 112,657 headlines that can be used for building an automated clickbait detection system for Telugu, a resource-poor language. Our contribution in this paper includes (i) generation of the latest pre-trained language models, including RoBERTa, ALBERT, and ELECTRA trained on a large Telugu corpora of 8,015,588 sentences that we had collected, (ii) data analysis and benchmarking the performance of different approaches ranging from hand-crafted features to state-of-the-art models. We show that the pre-trained language models trained on Telugu outperform the existing pre-trained models viz. BERT-Mulingual-Case [1], XLM-MLM [2], and XLM-R [3] on clickbait task. On a large Telugu clickbait dataset of 112,657 samples, the Light Gradient Boosted Machines (LGBM) model achieves an F1-score of 0.94 for clickbait headlines. For Non-Clickbait headlines, F1-score of 0.93 is obtained which is similar to that of Clickbait class. We open-source our dataset, pre-trained models, and code ‘ We show that the pre-trained language models trained on Telugu outperform the existing pre-trained models viz. BERT-Mulingual-Case [1], XLM-MLM [2], and XLM-R [3] on clickbait task. On a large Telugu clickbait dataset of 112,657 samples, the Light Gradient Boosted Machines (LGBM) model achieves an F1-score of 0.94 for clickbait headlines. For Non-Clickbait headlines, F1-score of 0.93 is obtained which is similar to that of Clickbait class. We open-source our dataset, pre-trained models, and code11https://github.com/subbareddy248/Clickbait-Resources Mounika Marreddy, Subba Reddy Oota, Lakshmi Sireesha Vakada, Venkata Charan Chinni, Radhika Mamidi |
IJCNN | 2 |
| 2021 | HealTech - A System for Predicting Patient Hospitalization Risk and Wound Progression in Old PatientsabstractHow bad is my wound? How fast will the wound heal? Do I need to get hospitalized? Questions like these are critical for wound assessment, but challenging to answer. Given a wound image and patient attributes, our goal is to build models for two wound assessment tasks: (1) predicting if the patient needs hospitalization for the wound to heal, and (2) estimating wound progression, i.e., weeks to heal. The problem is challenging because wound progression and hospitalization risk depend on multiple factors that need to be inferred automatically from the given wound image. There exists no work which performs a rigorous study of wound assessment tasks considering multiple wound attributes inferred using a large dataset of wound images. We present HealTech, a two-stage wound assessment solution. The first stage predicts various wound attributes (like ulcer type, location, stage, etc.) from wound images, using deep neural networks. The second stage predicts (1) whether the wound would heal (using conventional in-house treatment) or not (needs hospitalization), and (2) the number of weeks to heal, using an evolutionary algorithm based stacked Light Gradient Boosted Machines (LGBM) model. On a large dataset of 125711 wound images, HealTech achieves a recall of 83 and a precision of 92 for wounds with the risk of hospitalization. For wounds that can be healed without hospitalization, precision and recall are as high as 99. Our wound progression model provides a mean absolute error of 3.3 weeks. Subba Reddy Oota, Vijay Rowtula, Shahid Saleem Mohammed, Jeffrey Galitz, Minghsun Liu, Manish Gupta 0001 |
WACV | 1 |
| 2019 | Towards Automated Evaluation of Handwritten AssessmentsabstractAutomated evaluation of handwritten answers has been a challenging problem for scaling the education system for many years. Speeding up the evaluation remains as the major bottleneck for enhancing the throughput of instructors. This paper describes an effective method for automatically evaluating the short descriptive handwritten answers from the digitized images. Our goal is to evaluate a student's handwritten answer by assigning an evaluation score that is comparable to the human-assigned scores. Existing works in this domain mainly focused on evaluating handwritten essays with handcrafted, non-semantic features. Our contribution is two-fold: 1) we model this problem as a self-supervised, feature-based classification problem, which can fine-tune itself for each question without any explicit supervision. 2) We introduce the usage of semantic analysis for auto-evaluation in handwritten text space using the combination of Information Retrieval and Extraction (IRE) and, Natural Language Processing (NLP) methods to derive a set of useful features. We tested our method on three datasets created from various domains, using the help of students of different age groups. Experiments show that our method performs comparably to that of human evaluators. Vijay Rowtula, Subba Reddy Oota, C. V. Jawahar |
ICDAR | 2 |
| 2019 | StepEncog: A Convolutional LSTM Autoencoder for Near-Perfect fMRI EncodingabstractLearning a forward mapping that relates stimuli to the corresponding brain activation measured by functional magnetic resonance imaging (fMRI) is termed as estimating encoding models. Computational tractability usually forces current encoding as well as decoding solutions to typically consider only a small subset of voxels from the actual 3D volume of activation. Further, while reconstructing stimulus information from brain activation (brain decoding) has received wider attention, there have been only a few attempts at constructing encoding solutions in the extant neuro-imaging literature. In this paper, we present StepEncog, a convolutional LSTM autoencoder model trained on fMRI voxels. The model can predict the entire brain volume rather than a small subset of voxels, as presented in earlier research works. We argue that the resulting solution avoids the problem of devising encoding models based on a rule-based selection of informative voxels and the concomitant issue of wide spatial variability of such voxels across participants. The perturbation experiments indicate that the proposed deep encoder indeed learns to predict brain activations with high spatial accuracy. On challenging universal decoder imaging datasets, our model yielded encouraging results. Subba Reddy Oota, Vijay Rowtula, Manish Gupta 0001, Raju S. Bapi |
IJCNN | 1 |
| 2018 | fMRI Semantic Category Decoding Using Linguistic Encoding of Word Embeddings
Subba Reddy Oota, Naresh Manwani, Raju S. Bapi |
ICONIP (3) | 1 |
| 2018 | Affect in Tweets using Experts Model
Subba Reddy Oota, Adithya Avvaru, Mounika Marreddy, Radhika Mamidi |
PACLIC | 1 |
| 2017 | Tag Me a Label with Multi-arm: Active Learning for Telugu Sentiment Analysis
Sandeep Sricharan Mukku, Subba Reddy Oota, Radhika Mamidi |
DaWaK | 2 |
| 2017 | Metastability of cortical BOLD signals in maturation and senescenceabstractWe assess change in metastability to characterize age-effects on the dynamic repertoire of the functional networks at rest. Resting state fMRI signals from each subject (N=48) have been used and metastability is evaluated as the standard deviation of mean phase synchrony of BOLD signals across whole-brain as well as across known resting state networks. The results suggest that significant whole-brain metastability changes occur between middle to old age. We also demonstrate that static time-averaged FC largely undermines age-effects on the interaction between functional networks. Discriminant Function Analysis reveals existence of two different patterns of change in metastability, which maximally discriminates between two different processes of maturation and ageing. Shruti Naik, Subba Reddy Oota, Arpan Banerjee, Dipanjan Roy, Raju S. Bapi |
IJCNN | 2 |