VLDB 2026 Research / reviewers in the wild / expert
Radhika Mamidi
dblp:134/6779
· DBLP profile ↗
32ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0003-0171-0816ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 10 since 2021Databases, data management, data science and information retrieval · 3Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Confabulations from ACL Publications (CAP): A Dataset for Scientific Hallucination Detection
Federica Gamba, Aman Sinha 0002, Timothee Mickus, Raúl Vázquez, Patanjali Bhamidipati, Claudio Savelli, Ahana Chattopadhyay, Laura A. Zanella, Yash Kankanampati, Binesh Arakkal Remesh, Aryan Ashok Chandramania, Chuyuan Li, Ioana Buhnila, Radhika Mamidi |
LREC | 15 |
| 2026 | EthiQuest: LLM-Powered Ethical Questionnaire Generation for Research Review
Ishank Kapania, Radhika Mamidi |
LREC | 2 |
| 2025 | Choose Your Words Wisely: Domain-Adaptive Masking Makes Language Models Learn Faster
Vanshpreet S. Kohli, Aaron Monis, Radhika Mamidi |
AIME (2) | 3 |
| 2025 | Aligning Text/Speech Representations from Multimodal Models with MEG Brain Activity During ListeningabstractAlthough speech language models are expected to align well with brain language processing during speech comprehension, recent studies have found that they fail to capture brainrelevant semantics beyond low-level features.Surprisingly, text-based language models exhibit stronger alignment with brain language regions, as they better capture brain-relevant semantics.However, no prior work has examined the alignment effectiveness of text/speech representations from multimodal models.This raises several key questions: Can speech embeddings from such multimodal models capture brain-relevant semantics through cross-modal interactions?Which modality can take advantage of this synergistic multimodal understanding to improve alignment with brain language processing?Can text/speech representations from such multimodal models outperform unimodal models?To address these questions, we systematically analyze multiple multimodal models, extracting both text-and speech-based representations to assess their alignment with MEG brain recordings during naturalistic story listening.We find that text embeddings from both multimodal and unimodal models significantly outperform speech embeddings from these models.Specifically, multimodal text embeddings exhibit a peak around 200 ms, suggesting that they benefit from speech embeddings, with heightened activity during this time period.However, speech embeddings from these multimodal models still show a similar alignment compared to their unimodal counterparts, suggesting that they do not gain meaningful semantic benefits over text-based representations.These results highlight an asymmetry in cross-modal knowledge transfer, where the text modality benefits more from speech information, but not vice versa.We make the code publicly available 1 . Padakanti Srijith, Khushbu Pahwa, Radhika Mamidi, Raju S. Bapi, Manish Gupta 0001, Subba Reddy Oota |
EMNLP | 3 |
| 2023 | Enhancing Code-mixed Text Generation Using Synthetic Data Filtering in Neural Machine TranslationabstractCode-Mixing 1 , the act of mixing two or more languages, is a common communicative phenomenon in multi-lingual societies.The lack of quality in code-mixed data is a bottleneck for NLP systems.On the other hand, Monolingual systems perform well due to ample high-quality data.To bridge the gap, creating coherent translations of monolingual sentences to their code-mixed counterparts can improve accuracy in code-mixed settings for NLP downstream tasks.In this paper, we propose a neural machine translation approach to generate high-quality code-mixed sentences by leveraging human judgements.We train filters based on human judgements to identify natural code-mixed sentences from a larger synthetically generated code-mixed corpus, resulting in a three-way silver parallel corpus between monolingual English, monolingual Indian language and code-mixed English with an Indian language.Using these corpora, we fine-tune multi-lingual encoder-decoder models viz, mT5 and mBART, for the translation task.Our results indicate that our approach of using filtered data for training outperforms the current systems for code-mixed generation in Hindi-English.Apart from Hindi-English, the approach performs well when applied to Telugu, a low-resource language, to generate Telugu-English code-mixed sentences. Dama Sravani, Radhika Mamidi |
CoNLL | 2 |
| 2023 | GAE-ISUMM: Unsupervised Graph-based Summarization for Indian LanguagesabstractDocument summarization aims to create a precise and coherent summary of a text document. Many deep learning summarization models are developed mainly for English, often requiring a large training corpus and efficient pre-trained language models and tools. However, English summarization models for low-resource Indian languages are often limited by rich morphological variation, syntax, and semantic differences. In this paper, we propose GAE-ISUMM, an unsupervised Indic summarization model that extracts summaries from text documents. In particular, our proposed model, GAE-ISUMM uses Graph Autoencoder (GAE) to learn text representations and a document summary jointly. We also provide a manually-annotated Telugu summarization dataset TELSUM, to experiment with our model GAE-ISUMM. Further, we benchmark with the most publicly available Indian language summarization datasets to investigate the effectiveness of GAE-ISUMM. Our experiments of GAE-ISUMM on seven Indian languages make the following observations: (i) it is competitive or better than state-of-the-art results on all datasets, (ii) it reports benchmark results on TELSUM, and (iii) the inclusion of positional and cluster information in the proposed model improved the performance of summaries. We open-source our dataset and code11https://github.com/scsmuhio/Summarization. Lakshmi Sireesha Vakada, Anudeep Chaluvadi, Mounika Marreddy, Subba Reddy Oota, Radhika Mamidi |
IJCNN | 5 |
| 2023 | A code-mixed task-oriented dialog dataset for medical domain
Suman Dowlagar, Radhika Mamidi |
Comput. Speech Lang. | 2 |
| 2023 | Am I a Resource-Poor Language? Data Sets, Embeddings, Models and Analysis for four different NLP Tasks in Telugu LanguageabstractDue to the lack of a large annotated corpus, many resource-poor Indian languages struggle to reap the benefits of recent deep feature representations in Natural Language Processing (NLP) . Moreover, adopting existing language models trained on large English corpora for Indian languages is often limited by data availability, rich morphological variation, syntax, and semantic differences. In this paper, we explore the traditional to recent efficient representations to overcome the challenges of a low resource language, Telugu. In particular, our main objective is to mitigate the low-resource problem for Telugu. Overall, we present several contributions to a resource-poor language viz. Telugu. (i) a large annotated data (35,142 sentences in each task) for multiple NLP tasks such as sentiment analysis, emotion identification, hate-speech detection, and sarcasm detection, (ii) we create different lexicons for sentiment, emotion, and hate-speech for improving the efficiency of the models, (iii) pretrained word and sentence embeddings, and (iv) different pretrained language models for Telugu such as ELMo-Te , BERT-Te , RoBERTa-Te , ALBERT-Te , and DistilBERT-Te on a large Telugu corpus consisting of 8,015,588 sentences (1,637,408 sentences from Telugu Wikipedia and 6,378,180 sentences crawled from different Telugu websites). Further, we show that these representations significantly improve the performance of four NLP tasks and present the benchmark results for Telugu. We argue that our pretrained embeddings are competitive or better than the existing multilingual pretrained models: mBERT , XLM-R , and IndicBERT . Lastly, the fine-tuning of pretrained models show higher performance than linear probing results on four NLP tasks with the following F1-scores: Sentiment (68.72), Emotion (58.04), Hate-Speech (64.27), and Sarcasm (77.93). We also experiment on publicly available Telugu datasets (Named Entity Recognition, Article Genre Classification, and Sentiment Analysis) and find that our Telugu pretrained language models ( BERT-Te and RoBERTa-Te ) outperform the state-of-the-art system except for the sentiment task. We open-source our corpus, four different datasets, lexicons, embeddings, and code https://github.com/Cha14ran/DREAM-T. The pretrained Transformer models for Telugu are available at https://huggingface.co/ltrctelugu. Mounika Marreddy, Subba Reddy Oota, Lakshmi Sireesha Vakada, Venkata Charan Chinni, Radhika Mamidi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2022 | cViL: Cross-Lingual Training of Vision-Language Models using Knowledge DistillationabstractVision-and-language tasks are gaining popularity in the research community, but the focus is still mainly on English. We propose a pipeline that utilizes English-only vision-language models to train a monolingual model for a target language. We propose to extend OSCAR+, a model which leverages object tags as anchor points for learning image-text alignments, to train on visual question answering datasets in different languages. We propose a novel approach to knowledge distillation to train the model in other languages using parallel sentences. Compared to other models that use the target language in the pretraining corpora, we can leverage an existing English model to transfer the knowledge to the target language using significantly lesser resources. We also release a large-scale visual question answering dataset in Japanese and Hindi language. Though we restrict our work to visual question answering, our model can be extended to any sequence-level classification task, and it can be extended to other languages as well. This paper focuses on two languages for the visual question answering task - Japanese and Hindi. Our pipeline outperforms the current state-of-the-art models by a relative increase of 4.4% and 13.4% respectively in accuracy. Kshitij Gupta 0001, Devansh Gautam, Radhika Mamidi |
ICPR | 3 |
| 2022 | Multi-Task Text Classification using Graph Convolutional Networks for Large-Scale Low Resource LanguageabstractGraph Convolutional Networks (GCN) have achieved state-of-art results on single text classification tasks like sentiment analysis, emotion detection, etc. However, the performance is achieved by testing and reporting on resource-rich languages like English. Applying GCN for multi-task text classification is an unexplored area. Moreover, training a GCN or adopting an English GCN for Indian languages is often limited by data availability, rich morphological variation, syntax, and semantic differences. In this paper, we study the use of GCN for the Telugu language in single and multi-task settings for four natural language processing (NLP) tasks, viz. sentiment analysis (SA), emotion identification (EI), hate-speech (HS), and sarcasm detection (SAR). In order to evaluate the performance of GCN with one of the Indian languages, Telugu, we analyze the GCN based models with extensive experiments on four downstream tasks. In addition, we created an annotated Telugu dataset, TEL-NLP, for the four NLP tasks. Further, we propose a supervised graph reconstruction method, Multi-Task Text GCN (MT- Text GCN) on the Telugu that leverages to simultaneously (i) learn the low-dimensional word and sentence graph embeddings from word-sentence graph reconstruction using graph autoencoder (GAE) and (ii) perform multi-task text classification using these latent sentence graph embeddings. We argue that our proposed MT- Text GCN achieves significant improvements on TEL-NLP over existing Telugu pretrained word embeddings [1], multilingual pretrained Transformer models: mBERT [2], and XLM-R [3]. On TEL-NLP, we achieve a high Fl-score for four NLP tasks: SA (0.84), EI (0.55), HS (0.83) and SAR (0.66). Finally, we show our model's quantitative and qualitative analysis on the four NLP tasks in Telugu. We open-source our TEL-NLP dataset, pretrained models, and code11https://github.com/scsmuhio/MTGCN_Resources. Mounika Marreddy, Subba Reddy Oota, Lakshmi Sireesha Vakada, Venkata Charan Chinni, Radhika Mamidi |
IJCNN | 5 |
| 2021 | Clickbait Detection in Telugu: Overcoming NLP Challenges in Resource-Poor Languages using Benchmarked TechniquesabstractClickbait headlines have become a nudge in social media and news websites. The methods to identify clickbaits are largely being developed for English. There is a need for the same in other languages as well with the increase in the usage of social media platforms in different languages. In this work, we present an annotated clickbait dataset of 112,657 headlines that can be used for building an automated clickbait detection system for Telugu, a resource-poor language. Our contribution in this paper includes (i) generation of the latest pre-trained language models, including RoBERTa, ALBERT, and ELECTRA trained on a large Telugu corpora of 8,015,588 sentences that we had collected, (ii) data analysis and benchmarking the performance of different approaches ranging from hand-crafted features to state-of-the-art models. We show that the pre-trained language models trained on Telugu outperform the existing pre-trained models viz. BERT-Mulingual-Case [1], XLM-MLM [2], and XLM-R [3] on clickbait task. On a large Telugu clickbait dataset of 112,657 samples, the Light Gradient Boosted Machines (LGBM) model achieves an F1-score of 0.94 for clickbait headlines. For Non-Clickbait headlines, F1-score of 0.93 is obtained which is similar to that of Clickbait class. We open-source our dataset, pre-trained models, and code ‘ We show that the pre-trained language models trained on Telugu outperform the existing pre-trained models viz. BERT-Mulingual-Case [1], XLM-MLM [2], and XLM-R [3] on clickbait task. On a large Telugu clickbait dataset of 112,657 samples, the Light Gradient Boosted Machines (LGBM) model achieves an F1-score of 0.94 for clickbait headlines. For Non-Clickbait headlines, F1-score of 0.93 is obtained which is similar to that of Clickbait class. We open-source our dataset, pre-trained models, and code11https://github.com/subbareddy248/Clickbait-Resources Mounika Marreddy, Subba Reddy Oota, Lakshmi Sireesha Vakada, Venkata Charan Chinni, Radhika Mamidi |
IJCNN | 5 |
| 2020 | Leveraging Multilingual Resources for Language Invariant Sentiment AnalysisabstractSentiment analysis is a widely researched NLP problem with state-of-the-art solutions capable of attaining human-like accuracies for various languages. However, these methods rely heavily on large amounts of labeled data or sentiment weighted language-specific lexical resources that are unavailable for low-resource languages. Our work attempts to tackle this data scarcity issue by introducing a neural architecture for language invariant sentiment analysis capable of leveraging various monolingual datasets for training without any kind of cross-lingual supervision. The proposed architecture attempts to learn language agnostic sentiment features via adversarial training on multiple resource-rich languages which can then be leveraged for inferring sentiment information at a sentence level on a low resource language. Our model outperforms the current state-of-the-art methods on the Multilingual Amazon Review Text Classification dataset [REF] and achieves significant performance gains over prior work on the low resource Sentiraama corpus [REF]. A detailed analysis of our research highlights the ability of our architecture to perform significantly well in the presence of minimal amounts of training data for low resource languages. Allen Antony, Arghya Bhattacharya, Jaipal Goud, Radhika Mamidi |
EAMT | 4 |
| 2020 | SentiInc: Incorporating Sentiment Information into Sentiment Transfer Without Parallel Data
Kartikey Pant, Yash Verma, Radhika Mamidi |
ECIR (2) | 3 |
| 2020 | Manovaad: A Novel Approach to Event Oriented Corpus Creation Capturing Subjectivity and FocusabstractIn today’s era of globalisation, the increased outreach for every event across the world has been leading to conflicting opinions, arguments and disagreements, often reflected in print media and online social platforms. It is necessary to distinguish factual observations from personal judgements in news, as subjectivity in reporting can influence the audience’s perception of reality. Several studies conducted on the different styles of reporting in journalism are essential in understanding phenomena such as media bias and multiple interpretations of the same event. This domain finds applications in fields such as Media Studies, Discourse Analysis, Information Extraction, Sentiment Analysis, and Opinion Mining. We present an event corpus Manovaad-v1.0 consisting of 1035 news articles corresponding to 65 events from 3 levels of newspapers viz., Local, National, and International levels. Using this novel format, we correlate the trends in the degree of subjectivity with the geographical closeness of reporting using a Bi-RNN model. We also analyse the role of background and focus in event reporting and capture the focus shift patterns within a global discourse structure for an event. We do this across different levels of reporting and compare the results with the existing work on discourse processing. Lalitha Kameswari, Radhika Mamidi |
LREC | 2 |
| 2020 | Annotated Corpus for Sentiment Analysis in Odia LanguageabstractGiven the lack of an annotated corpus of non-traditional Odia literature which serves as the standard when it comes sentiment analysis, we have created an annotated corpus of Odia sentences and made it publicly available to promote research in the field. Secondly, in order to test the usability of currently available Odia sentiment lexicon, we experimented with various classifiers by training and testing on the sentiment annotated corpus while using identified affective words from the same as features. Annotation and classification are done at sentence level as the usage of sentiment lexicon is best suited to sentiment analysis at this level. The created corpus contains 2045 Odia sentences from news domain annotated with sentiment labels using a well-defined annotation scheme. An inter-annotator agreement score of 0.79 is reported for the corpus. Gaurav Mohanty, Pruthwik Mishra, Radhika Mamidi |
LREC | 3 |
| 2020 | Dataset Creation and Evaluation of Aspect Based Sentiment Analysis in Telugu, a Low Resource LanguageabstractIn recent years, sentiment analysis has gained popularity as it is essential to moderate and analyse the information across the internet. It has various applications like opinion mining, social media monitoring, and market research. Aspect Based Sentiment Analysis (ABSA) is an area of sentiment analysis which deals with sentiment at a finer level. ABSA classifies sentiment with respect to each aspect to gain greater insights into the sentiment expressed. Significant contributions have been made in ABSA, but this progress is limited only to a few languages with adequate resources. Telugu lags behind in this area of research despite being one of the most spoken languages in India and an enormous amount of data being created each day. In this paper, we create a reliable resource for aspect based sentiment analysis in Telugu. The data is annotated for three tasks namely Aspect Term Extraction, Aspect Polarity Classification and Aspect Categorisation. Further, we develop baselines for the tasks using deep learning methods demonstrating the reliability and usefulness of the resource. Yashwanth Reddy Regatte, Rama Rohit Reddy Gangula, Radhika Mamidi |
LREC | 3 |
| 2020 | A Sentiwordnet Strategy for Curriculum Learning in Sentiment Analysis
Vijjini Anvesh Rao, Kaveri Anuranjana, Radhika Mamidi |
NLDB | 3 |
| 2018 | Resource Creation Towards Automated Sentiment Analysis in Telugu (a low resource language) and Integrating Multiple Domain Sources to Enhance Sentiment Prediction
Rama Rohit Reddy Gangula, Radhika Mamidi |
LREC | 2 |
| 2018 | From Humour to Hatred: A Computational Analysis of Off-Colour Humour
Vikram Ahuja, Radhika Mamidi, Navjyoti Singh |
NLPCC (2) | 2 |
| 2018 | Predicting the Genre and Rating of a Movie Based on its Synopsis
Varshit Battu, Vishal Batchu, Rama Rohit Reddy Gangula, Mohana Murali Krishna Reddy Dakannagari, Radhika Mamidi |
PACLIC | 5 |
| 2018 | Word Level Language Identification in English Telugu Code Mixed Data
Sunil Gundapu, Radhika Mamidi |
PACLIC | 2 |
| 2018 | Political Discourse Analysis : A Case Study of 2014 Andhra Pradesh State Assembly Election of Interpersonal Speech Choices
Lalitha Kameswari, Radhika Mamidi |
PACLIC | 2 |
| 2018 | Affect in Tweets using Experts Model
Subba Reddy Oota, Adithya Avvaru, Mounika Marreddy, Radhika Mamidi |
PACLIC | 4 |
| 2018 | Syllables for Sentence Classification in Morphologically Rich Languages
Madhuri Tummalapalli, Radhika Mamidi |
PACLIC | 2 |
| 2017 | Tag Me a Label with Multi-arm: Active Learning for Telugu Sentiment Analysis
Sandeep Sricharan Mukku, Subba Reddy Oota, Radhika Mamidi |
DaWaK | 3 |
| 2016 | A Karaka Dependency Based Dialog Act Tagging for Telugu Using Combination of LMs and HMM
Suman Dowlagar, Radhika Mamidi |
CICLing (1) | 2 |
| 2016 | Part-of-Speech Tagging for Code Mixed English-Telugu Social Media Data
Kovida Nelakuditi, Divya Sai Jitta, Radhika Mamidi |
CICLing (1) | 3 |
| 2016 | Shallow Parsing Pipeline - Hindi-English Code-Mixed Social Media TextabstractArnav Sharma, Sakshi Gupta, Raveesh Motlani, Piyush Bansal, Manish Shrivastava, Radhika Mamidi, Dipti M. Sharma. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Arnav Sharma, Raveesh Motlani, Piyush Bansal, Manish Shrivastava 0001, Radhika Mamidi, Dipti Misra Sharma |
HLT-NAACL | 6 |
| 2015 | Statistical Sandhi Splitter for Agglutinative Languages
Prathyusha Kuncham, Kovida Nelakuditi, Sneha Nallani, Radhika Mamidi |
CICLing (1) | 4 |
| 2015 | A Dialogue System for Telugu, a Resource-Poor Language
Mullapudi Ch. Sravanthi, Prathyusha Kuncham, Radhika Mamidi |
CICLing (2) | 3 |
| 2013 | A Novel Approach Towards Incorporating Context Processing Capabilities in NLIDB System
Arjun R. Akula, Rajeev Sangal, Radhika Mamidi |
IJCNLP | 3 |
| 2013 | Stance Classification in Online Debates by Recognizing Users' Intentions
Sarvesh Ranade, Rajeev Sangal, Radhika Mamidi |
SIGDIAL Conference | 3 |