Manish Shrivastava 0001

dblp:65/3881 · DBLP profile ↗
← Back
51ranked-venue papers
0as first author
19since 2021 · last 2025
0000-0001-8705-6637ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 45 · 18 since 2021Databases, data management, data science and information retrieval · 6 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2Theory of computation · 1
YearPublicationVenuePosition
2025 BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages
abstract
Shamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle, Terry Ruas, Meriem Beloucif, Christine de Kock, Nirmal Surange, Daniela Teodorescu, Ibrahim Said Ahmad, David Ifeoluwa Adelani, Alham Fikri Aji, Felermino D. M. A. Ali, Ilseyar Alimova, Vladimir Araujo, Nikolay Babakov, Naomi Baes, Ana-Maria Bucur, Andiswa Bukula, Guanqun Cao, Rodrigo Tufiño, Rendi Chevi, Chiamaka Ijeoma Chukwuneke, Alexandra Ciobotaru, Daryna Dementieva, Murja Sani Gadanya, Robert Geislinger, Bela Gipp, Oumaima Hourrane, Oana Ignat, Falalu Ibrahim Lawan, Rooweither Mabuya, Rahmad Mahendra, Vukosi Marivate, Alexander Panchenko, Andrew Piper, Charles Henrique Porto Ferreira, Vitaly Protasov, Samuel Rutunda, Manish Shrivastava, Aura Cristina Udrea, Lilian Diana Awuor Wanzare, Sophie Wu, Florian Valentin Wunderlich, Hanif Muhammad Zhafran, Tianhui Zhang, Yi Zhou, Saif M. Mohammad. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Shamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle, Terry Ruas, Meriem Beloucif, Christine de Kock, Nirmal Surange, Daniela Teodorescu, Ibrahim Said Ahmad, David Ifeoluwa Adelani, Alham Fikri Aji, Felermino D. M. A. Ali, Ilseyar Alimova, Vladimir Araujo, Nikolay Babakov, Naomi Baes, Ana-Maria Bucur, Andiswa Bukula, Guanqun Cao, Rodrigo Tufiño, Rendi Chevi, Chiamaka Ijeoma Chukwuneke, Alexandra Ciobotaru, Daryna Dementieva, Murja Sani Gadanya, Robert Geislinger, Bela Gipp, Oumaima Hourrane, Oana Ignat, Falalu Ibrahim Lawan, Rooweither Mabuya, Rahmad Mahendra, Vukosi Marivate, Alexander Panchenko, Andrew Piper, Charles Henrique Porto Ferreira, Vitaly Protasov, Samuel Rutunda, Manish Shrivastava 0001, Aura Cristina Udrea, Lilian Wanzare, Sophie Wu, Florian Valentin Wunderlich, Hanif Muhammad Zhafran, Tianhui Zhang, Yi Zhou 0019, Saif M. Mohammad
ACL (1)40
2025 Why should only High-Resource-Languages have all the fun? Pivot Based Evaluation in Low Resource Setting
abstract
Evaluating machine translation (MT) systems for low-resource languages has long been a challenge due to the limited availability of evaluation metrics and resources. As a result, researchers in this space have relied primarily on lexical-based metrics like BLEU, TER, and ChrF, which lack semantic evaluation. In this first-of-its-kind work, we propose a novel pivot-based evaluation framework that addresses these limitations; after translating low-resource language outputs into a related high-resource language, we leverage advanced neural and embedding-based metrics for more meaningful evaluation. Through a series of experiments using five low-resource languages: Assamese, Manipuri, Kannada, Bhojpuri, and Nepali, we demonstrate how this method extends the coverage of both lexical-based and embedding-based metrics, even for languages not directly supported by advanced metrics. Our results show that the differences between direct and pivot-based evaluation scores are minimal, proving that this approach is a viable and effective solution for evaluating translations in endangered and low-resource languages. This work paves the way for more inclusive, accurate, and scalable MT evaluation for underrepresented languages, marking a significant step forward in this under-explored area of research. The code and data will be made available at https://github.com/AnanyaCoder/PivotBasedEvaluation.
Ananya Mukherjee, Saumitra Yadav, Manish Shrivastava 0001
COLING3
2025 Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)
abstract
Transformer-based language models, though not explicitly trained to mimic brain recordings, have demonstrated surprising alignment with brain activity. Progress in these models—through increased size, instruction-tuning, and multimodality—has led to better representational alignment with neural data. Recently, a new class of instruction-tuned multimodal LLMs (MLLMs) have emerged, showing remarkable zero-shot capabilities in open-ended multimodal vision tasks. However, it is unknown whether MLLMs, when prompted with natural instructions, lead to better brain alignment and effectively capture instruction-specific representations. To address this, we first investigate the brain alignment, i.e., measuring the degree of predictivity of neural visual activity using text output response embeddings from MLLMs as participants engage in watching natural scenes. Experiments with 10 different instructions (like image captioning, visual question answering, etc.) show that MLLMs exhibit significantly better brain alignment than vision-only models and perform comparably to non-instruction-tuned multimodal models like CLIP. We also find that while these MLLMs are effective at generating high-quality responses suitable to the task-specific instructions, not all instructions are relevant for brain alignment. Further, by varying instructions, we make the MLLMs encode instruction-specific visual concepts related to the input image. This analysis shows that MLLMs effectively capture count-related and recognition-related concepts, demonstrating strong alignment with brain activity. Notably, the majority of the explained variance of the brain encoding models is shared between MLLM embeddings of image captioning and other instructions. These results indicate that enhancing MLLMs' ability to capture more task-specific information could allow for better differentiation between various types of instructions, and hence improve their precision in predicting brain responses.
Subba Reddy Oota, Akshett Rai Jindal, Ishani Mondal, Khushbu Pahwa, Satya Sai Srinath Namburi, Manish Shrivastava 0001, Maneesh Kumar Singh 0001, Raju S. Bapi, Manish Gupta 0001
ICLR6
2025 Analyzing (In)Abilities of SAEs via Formal Languages
abstract
Abhinav Menon, Manish Shrivastava, David Krueger, Ekdeep Singh Lubana. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Abhinav Menon, Manish Shrivastava 0001, David Krueger 0001, Ekdeep Singh Lubana
NAACL (Long Papers)2
2025 MAPWise: Evaluating Vision-Language Models for Advanced Map Queries
abstract
Srija Mukhopadhyay, Abhishek Rajgaria, Prerana Khatiwada, Manish Shrivastava, Dan Roth, Vivek Gupta. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Srija Mukhopadhyay, Abhishek Rajgaria, Prerana Khatiwada, Manish Shrivastava 0001, Dan Roth 0001, Vivek Gupta 0001
NAACL (Long Papers)4
2025 Non Idiomatic Conventionalised Expressions: A New Pain in the Neck?
Ganesh Katrapati, Manish Shrivastava 0001
PACLIC2
2025 From Human Judgements to Predictive Models: Unravelling Acceptability in Code-Mixed Sentences
abstract
Current computational approaches for analysing or generating code-mixed sentences do not explicitly model “naturalness” or “acceptability” of code-mixed sentences, but rely on training corpora to reflect distribution of acceptable code-mixed sentences. Modelling human judgement for the acceptability of code-mixed text can help in distinguishing natural code-mixed text and enable quality-controlled generation of code-mixed text. To this end, we construct Cline—a dataset containing human acceptability judgements for English-Hindi (en-hi) code-mixed text. Cline is the largest of its kind with 16,642 sentences, consisting of samples sourced from two sources: synthetically generated code-mixed text and samples collected from online social media. Our analysis establishes that popular code-mixing metrics such as CMI, Number of Switch Points, Burstines, which are used to filter/curate/compare code-mixed corpora have low correlation with human acceptability judgements, underlining the necessity of our dataset. Experiments using Cline demonstrate that simple Multilayer Perceptron (MLP) models when trained solely using code-mixing metrics as features are outperformed by fine-tuned pre-trained Multilingual Large Language Models (MLLMs). Specifically, among Encoder models XLM-Roberta and Bernice outperform IndicBERT across different configurations. Among Encoder-Decoder models, mBART performs better than mT5, however, Encoder-Decoder models are not able to outperform Encoder-only models. Decoder-only models perform the best when compared with all other MLLMS, with Llama 3.2 - 3B models outperforming similarly sized Qwen, Phi models. Comparison with zero and fewshot capabilitites of ChatGPT show that MLLMs fine-tuned on larger data outperform ChatGPT, providing scope for improvement in code-mixed tasks. Zero-shot transfer from English–Hindi to English-Telugu acceptability judgments using our model checkpoints proves superior to random baselines, enabling application to other code-mixed language pairs and providing further avenues of research. We publicly release our human-annotated dataset, trained checkpoints, code-mix corpus, and code for data generation and model training.
Prashant Kodali, Anmol Goel, Likhith Asapu, Vamshi Krishna Bonagiri, Anirudh Govil, Monojit Choudhury, Ponnurangam Kumaraguru, Manish Shrivastava 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.8
2024 TeClass: A Human-Annotated Relevance-based Headline Classification and Generation Dataset for Telugu
abstract
News headline generation is a crucial task in increasing productivity for both the readers and producers of news. This task can easily be aided by automated News headline-generation models. However, the presence of irrelevant headlines in scraped news articles results in sub-optimal performance of generation models. We propose that relevance-based headline classification can greatly aid the task of generating relevant headlines. Relevance-based headline classification involves categorizing news headlines based on their relevance to the corresponding news articles. While this task is well-established in English, it remains under-explored in low-resource languages like Telugu due to a lack of annotated data. To address this gap, we present TeClass, the first-ever human-annotated Telugu news headline classification dataset, containing 78,534 annotations across 26,178 article-headline pairs. We experiment with various baseline models and provide a comprehensive analysis of their results. We further demonstrate the impact of this work by fine-tuning various headline generation models using TeClass dataset. The headlines generated by the models fine-tuned on highly relevant article-headline pairs, showed about a 5 point increment in the ROUGE-L scores. To encourage future research, the annotated dataset as well as the annotation guidelines will be made publicly available.
Gopichand Kanumolu, Lokesh Madasu, Nirmal Surange, Manish Shrivastava 0001
LREC/COLING4
2023 TourismNLG: A Multi-lingual Generative Benchmark for the Tourism Domain
Sahil Manoj Bhatt, Sahaj Agarwal, Omkar Gurjar, Manish Gupta 0001, Manish Shrivastava 0001
ECIR (1)5
2023 Fine-grained Contract NER using instruction based mode
Hiranmai Sri Adibhatla, Pavan Baswani, Manish Shrivastava 0001
PACLIC3
2023 The WEAVE 2.0 Corpus: Role Labelled Synthetic Chemical Procedures from Patents with Chemical Named Entities
Shubhangi Dutta, Manish Shrivastava 0001, Prabhakar Bhimalapuram
PACLIC2
2023 Mukhyansh: A Headline Generation Dataset for Indic Languages
Lokesh Madasu, Gopichand Kanumolu, Nirmal Surange, Manish Shrivastava 0001
PACLIC4
2022 Diverse Multi-Answer Retrieval with Determinantal Point Processes
abstract
Often questions provided to open-domain question answering systems are ambiguous. Traditional QA systems that provide a single answer are incapable of answering ambiguous questions since the question may be interpreted in several ways and may have multiple distinct answers. In this paper, we address multi-answer retrieval which entails retrieving passages that can capture majority of the diverse answers to the question. We propose a re-ranking based approach using Determinantal point processes utilizing BERT as kernels. Our method jointly considers query-passage relevance and passage-passage correlation to retrieve passages that are both query-relevant and diverse. Results demonstrate that our re-ranking technique outperforms state-of-the-art method on the AmbigQA dataset.
Poojitha Nandigam, Nikhil Rayaprolu, Manish Shrivastava 0001
COLING3
2022 DocInfer: Document-level Natural Language Inference using Optimal Evidence Selection
abstract
We present DocInfer -a novel, end-to-end Document-level Natural Language Inference model that builds a hierarchical document graph enriched through inter-sentence relations (topical, entity-based, concept-based), performs paragraph pruning using the novel SubGraph Pooling layer, followed by optimal evidence selection based on REINFORCE algorithm to identify the most important context sentences for a given hypothesis.Our evidence selection mechanism allows it to transcend the input length limitation of modern BERT-like Transformer models while presenting the entire evidence together for inferential reasoning.We show this is an important property needed to reason on large documents where the evidence may be fragmented and located arbitrarily far from each other.Extensive experiments on popular corpora -DocNLI, ContractNLI, and ConTRoL datasets, and our new proposed dataset called CaseHoldNLI on the task of legal judicial reasoning, demonstrate significant performance gains of 8-12% over SOTA methods.Our ablation studies validate the impact of our model.Performance improvement of ∼ 3 -6% on annotation-scarce downstream tasks of fact verification, multiple-choice QA, and contract clause retrieval demonstrates the usefulness of DocInfer beyond primary NLI tasks.
Puneet Mathur, Gautam Kunapuli, Riyaz A. Bhat, Manish Shrivastava 0001, Dinesh Manocha, Maneesh Kumar Singh 0001
EMNLP4
2022 TeSum: Human-Generated Abstractive Summarization Corpus for Telugu
abstract
Expert human annotation for summarization is definitely an expensive task, and can not be done on huge scales. But with this work, we show that even with a crowd sourced summary generation approach, quality can be controlled by aggressive expert informed filtering and sampling-based human evaluation. We propose a pipeline that crowd-sources summarization data and then aggressively filters the content via: automatic and partial expert evaluation. Using this pipeline we create a high-quality Telugu Abstractive Summarization dataset (TeSum) which we validate with sampling-based human evaluation. We also provide baseline numbers for various models commonly used for summarization. A number of recently released datasets for summarization, scraped the web-content relying on the assumption that summary is made available with the article by the publishers. While this assumption holds for multiple resources (or news-sites) in English, it should not be generalised across languages without thorough analysis and verification. Our analysis clearly shows that this assumption does not hold true for most Indian language news resources. We show that our proposed filtration pipeline can even be applied to these large-scale scraped datasets to extract better quality article-summary pairs.
Ashok Urlana, Nirmal Surange, Pavan Baswani, Priyanka Ravva, Manish Shrivastava 0001
LREC5
2022 HashSet - A Dataset For Hashtag Segmentation
abstract
Hashtag segmentation is the task of breaking a hashtag into its constituent tokens. Hashtags often encode the essence of user-generated posts, along with information like topic and sentiment, which are useful in downstream tasks. Hashtags prioritize brevity and are written in unique ways - transliterating and mixing languages, spelling variations, creative named entities. Benchmark datasets used for the hashtag segmentation task - STAN, BOUN - are small and extracted from a single set of tweets. However, datasets should reflect the variations in writing styles of hashtags and account for domain and language specificity, failing which the results will misrepresent model performance. We argue that model performance should be assessed on a wider variety of hashtags, and datasets should be carefully curated. To this end, we propose HashSet, a dataset comprising of: a) 1.9k manually annotated dataset; b) 3.3M loosely supervised dataset. HashSet dataset is sampled from a different set of tweets when compared to existing datasets and provides an alternate distribution of hashtags to build and validate hashtag segmentation models. We analyze the performance of SOTA models for Hashtag Segmentation, and show that the proposed dataset provides an alternate set of hashtags to train and assess models.
Prashant Kodali, Akshala Bhatnagar, Naman Ahuja, Manish Shrivastava 0001, Ponnurangam Kumaraguru
LREC4
2022 Bilingual Tabular Inference: A Case Study on Indic Languages
abstract
Chaitanya Agarwal, Vivek Gupta, Anoop Kunchukuttan, Manish Shrivastava. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Chaitanya Agarwal, Vivek Gupta 0001, Anoop Kunchukuttan, Manish Shrivastava 0001
NAACL-HLT4
2022 Is My Model Using The Right Evidence? Systematic Probes for Examining Evidence-Based Tabular Reasoning
abstract
Abstract Neural models command state-of-the-art performance across NLP tasks, including ones involving “reasoning”. Models claiming to reason about the evidence presented to them should attend to the correct parts of the input while avoiding spurious patterns therein, be self-consistent in their predictions across inputs, and be immune to biases derived from their pre-training in a nuanced, context- sensitive fashion. Do the prevalent *BERT- family of models do so? In this paper, we study this question using the problem of reasoning on tabular data. Tabular inputs are especially well-suited for the study—they admit systematic probes targeting the properties listed above. Our experiments demonstrate that a RoBERTa-based model, representative of the current state-of-the-art, fails at reasoning on the following counts: it (a) ignores relevant parts of the evidence, (b) is over- sensitive to annotation artifacts, and (c) relies on the knowledge encoded in the pre-trained language model rather than the evidence presented in its tabular inputs. Finally, through inoculation experiments, we show that fine- tuning the model on perturbed data does not help it overcome the above challenges.
Vivek Gupta 0001, Riyaz A. Bhat, Atreya Ghosal, Manish Shrivastava 0001, Maneesh Kumar Singh 0001, Vivek Srikumar
Trans. Assoc. Comput. Linguistics4
2021 Topic Shift Detection for Mixed Initiative Response
abstract
Topic diversion occurs frequently with engaging open-domain dialogue systems like virtual assistants.The balance between staying on topic and rectifying the topic drift is important for a good collaborative system.In this paper, we present a model which uses a finetuned XLNet-base to classify the utterances pertaining to the major topic of conversation and those which are not, with a precision of 84%.We propose a preliminary study, classifying utterances into major, minor and offtopics, which further extends into a system initiative for diversion rectification.A case study was conducted where a system initiative is emulated as a response to the user going off-topic, mimicking a common occurrence of mixed initiative present in natural human-human conversation.This task of classifying utterances into those which belong to the major theme or not, would also help us in identification of relevant sentences for tasks like dialogue summarization and information extraction from conversations.
Konigari Rachna, Saurabh Ramola, Vijay Vardhan Alluri, Manish Shrivastava 0001
SIGDIAL4
2020 AbuseAnalyzer: Abuse Detection, Severity and Target Prediction for Gab Posts
abstract
While extensive popularity of online social media platforms has made information dissemination faster, it has also resulted in widespread online abuse of different types like hate speech, offensive language, sexist and racist opinions, etc.Detection and curtailment of such abusive content is critical for avoiding its psychological impact on victim communities, and thereby preventing hate crimes.Previous works have focused on classifying user posts into various forms of abusive behavior.But there has hardly been any focus on estimating the severity of abuse and the target.In this paper, we present a first of the kind dataset with 7,601 posts from Gab 1 which looks at online abuse from the perspective of presence of abuse, severity and target of abusive behavior.We also propose a system to address these tasks, obtaining an accuracy of ∼80% for abuse presence, ∼82% for abuse target prediction, and ∼65% for abuse severity prediction.
Mohit Chandra, Ashwin Pathak, Eesha Dutta, Paryul Jain, Manish Gupta 0001, Manish Shrivastava 0001, Ponnurangam Kumaraguru
COLING6
2020 Creation of Corpus and analysis in Code-Mixed Kannada-English Twitter data for Emotion Prediction
abstract
Emotion prediction is a critical task in the field of Natural Language Processing (NLP).There has been a significant amount of work done in emotion prediction for resource-rich languages.There has been work done on code-mixed social media corpus but not on emotion prediction of Kannada-English code-mixed Twitter data.In this paper, we analyze the problem of emotion prediction on corpus obtained from code-mixed Kannada-English extracted from Twitter annotated with their respective 'Emotion' for each tweet.We experimented with machine learning prediction models using features like Character N-Grams, Word N-Grams, Repetitive characters, and others on SVM and LSTM on our corpus, which resulted in an accuracy of 30% and 32%, respectively.
Appidi Abhinav Reddy, Vamshi Krishna Srirangam, Darsi Suhas, Manish Shrivastava 0001
COLING4
2020 Finding The Right One and Resolving it
abstract
One-anaphora has figured prominently in theoretical linguistic literature, but computational linguistics research on the phenomenon is sparse.Not only that, the long standing linguistic controversy between the determinative and the nominal anaphoric element one has propagated in the limited body of computational work on one-anaphora resolution, making this task harder than it is.In the present paper, we resolve this by drawing from an adequate linguistic analysis of the word one in different syntactic environments -once again highlighting the significance of linguistic theory in Natural Language Processing (NLP) tasks.We prepare an annotated corpus marking actual instances of one-anaphora with their textual antecedents, and use the annotations to experiment with state-of-the art neural models for one-anaphora resolution.Apart from presenting a strong neural baseline for this task, we contribute a gold-standard corpus, which is, to the best of our knowledge, the biggest resource on one-anaphora till date.
Payal Khullar, Arghya Bhattacharya, Manish Shrivastava 0001
CoNLL3
2020 MEE : An Automatic Metric for Evaluation Using Embeddings for Machine Translation
abstract
We propose MEE, an approach for automatic Machine Translation (MT) evaluation which leverages the similarity between embeddings of words in candidate and reference sentences to assess translation quality. Unigrams are matched based on their surface forms, root forms and meanings which aids to capture lexical, morphological and semantic equivalence. We perform experiments for MT from English to four Indian Languages (Telugu, Marathi, Bengali and Hindi) on a robust dataset comprising simple and complex sentences with good and bad translations. Further, it is observed that the proposed metric correlates better with human judgements than the existing widely used metrics.
Ananya Mukherjee, Hema Ala, Manish Shrivastava 0001, Dipti Misra Sharma
DSAA3
2020 Principle-to-Program: Neural Methods for Similar Question Retrieval in Online Communities
Muthusamy Chelliah, Manish Shrivastava 0001, Jaidam Ram Tej
ECIR (2)2
2020 Modeling ASR Ambiguity for Neural Dialogue State Tracking
abstract
Spoken dialogue systems typically use a list of top-N ASR hypotheses for inferring the semantic meaning and tracking the state of the dialogue. However ASR graphs, such as confusion networks (confnets), provide a compact representation of a richer hypothesis space than a top-N ASR list. In this paper, we study the benefits of using confusion networks with a state-of-the-art neural dialogue state tracker (DST). We encode the 2-dimensional confnet into a 1-dimensional sequence of embeddings using an attentional confusion network encoder which can be used with any DST system. Our confnet encoder is plugged into the state-of-the-art 'Global-locally Self-Attentive Dialogue State Tacker' (GLAD) model for DST and obtains significant improvements in both accuracy and inference time compared to using top-N ASR hypotheses.
Vaishali Pal, Fabien Guillot, Manish Shrivastava 0001, Jean-Michel Renders, Laurent Besacier
INTERSPEECH3
2020 NoEl: An Annotated Corpus for Noun Ellipsis in English
abstract
Ellipsis resolution has been identified as an important step to improve the accuracy of mainstream Natural Language Processing (NLP) tasks such as information retrieval, event extraction, dialog systems, etc. Previous computational work on ellipsis resolution has focused on one type of ellipsis, namely Verb Phrase Ellipsis (VPE) and a few other related phenomenon. We extend the study of ellipsis by presenting the No(oun)El(lipsis) corpus - an annotated corpus for noun ellipsis and closely related phenomenon using the first hundred movies of Cornell Movie Dialogs Dataset. The annotations are carried out in a standoff annotation scheme that encodes the position of the licensor, the antecedent boundary, and Part-of-Speech (POS) tags of the licensor and antecedent modifier. Our corpus has 946 instances of exophoric and endophoric noun ellipsis, making it the biggest resource of noun ellipsis in English, to the best of our knowledge. We present a statistical study of our corpus with novel insights on the distribution of noun ellipsis, its licensors and antecedents. Finally, we perform the tasks of detection and resolution of noun ellipsis with different classifiers trained on our corpus and report baseline results.
Payal Khullar, Kushal Majmundar, Manish Shrivastava 0001
LREC3
2019 Inductive Transfer Learning for Detection of Well-Formed Natural Language Search Queries
Bakhtiyar Syed, Vijayasaradhi Indurthi, Manish Gupta 0001, Manish Shrivastava 0001, Vasudeva Varma
ECIR (2)4
2018 Sentiment Analysis of Code-Mixed Languages Leveraging Resource Rich Languages
Nurendra Choudhary, Rajat Singh, Ishita Bindlish, Manish Shrivastava 0001
CICLing (2)4
2018 Emotions Are Universal: Learning Sentiment Based Representations of Resource-Poor Languages Using Siamese Networks
Nurendra Choudhary, Rajat Singh, Ishita Bindlish, Manish Shrivastava 0001
CICLing (2)4
2018 Contrastive Learning of Emoji-Based Representations for Resource-Poor Languages
Nurendra Choudhary, Rajat Singh, Ishita Bindlish, Manish Shrivastava 0001
CICLing (2)4
2018 Neural Network Architecture for Credibility Assessment of Textual Claims (Best Paper Award, First Place)
Nurendra Choudhary, Rajat Singh, Ishita Bindlish, Manish Shrivastava 0001
CICLing (2)4
2018 Automatic Normalization of Word Variations in Code-Mixed Social Media Text
Rajat Singh, Nurendra Choudhary, Manish Shrivastava 0001
CICLing (1)3
2018 Twitter corpus of Resource-Scarce Languages for Sentiment Analysis and Multilingual Emoji Prediction
abstract
In this paper, we leverage social media platforms such as twitter for developing corpus across multiple languages. The corpus creation methodology is applicable for resource-scarce languages provided the speakers of that particular language are active users on social media platforms. We present an approach to extract social media microblogs such as tweets (Twitter). In this paper, we create corpus for multilingual sentiment analysis and emoji prediction in Hindi, Bengali and Telugu. Further, we perform and analyze multiple NLP tasks utilizing the corpus to get interesting observations.
Nurendra Choudhary, Rajat Singh, Vijjini Anvesh Rao, Manish Shrivastava 0001
COLING4
2018 Transzaar: Empowers Human Translators
abstract
In this paper, we describe Transzaar - an AI powered tool that offers computer aided translation (CAT) functionality: pre-translation analysis, post-editing machine translated content, translation prediction, text aligning, extensive logging and integration with several machine translation (MT) systems. Transzaar aids a human translator to perform various language processing tasks, viz., Translation, Transliteration, Localization, and other kinds of Text Analysis tasks. Using Transzaar, human translators can post-edit the machine translated content, to improve fluency and accuracy of the translated content to match the naturalness of human translation while delivering with better turn-around time. Transzaar aids the process of post-editing, thereby increasing the productivity of human translators by 2-3 folds within couple of months of usage for certain language pairs. It collects feedback continuously, which helps the MT system to further learn and improve periodically, with the additional new generated data-set.
Rashid Ahmad 0003, Priyank Gupta, Nagaraju Vuppala, Sanket Kumar Pathak, Gagan Soni, Sravan Kumar, Manish Shrivastava 0001, Avinash K. Singh, Arbind K. Gangwar, Pawan Kumar 0001, Mukul K. Sinha
ICCSA (6)8
2018 Humor Detection in English-Hindi Code-Mixed Social Media Content : Corpus and Baseline System
Ankush Khandelwal, Sahil Swami, Syed Sarfaraz Akhtar, Manish Shrivastava 0001
LREC4
2018 Universal Dependency Parsing for Hindi-English Code-Switching
abstract
Irshad Bhat, Riyaz A. Bhat, Manish Shrivastava, Dipti Sharma. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Irshad Ahmad Bhat, Riyaz A. Bhat, Manish Shrivastava 0001, Dipti Misra Sharma
NAACL-HLT3
2018 Too Many Questions? What Can We Do? : Multiple Question Span Detection
Prathyusha Danda, Brij Mohan Lal Srivastava, Manish Shrivastava 0001
PACLIC3
2018 BoWLer: A neural approach to extractive text summarization
Pranav Dhakras, Manish Shrivastava 0001
PACLIC2
2017 The Unusual Suspects: Deep Learning Based Mining of Interesting Entity Trivia from Knowledge Graphs
abstract
Trivia is any fact about an entity which is interesting due to its unusualness, uniqueness or unexpectedness. Trivia could be successfully employed to promote user engagement in various product experiences featuring the given entity. A Knowledge Graph (KG) is a semantic network which encodes various facts about entities and their relationships. In this paper, we propose a novel approach called DBpedia Trivia Miner (DTM) to automatically mine trivia for entities of a given domain in KGs. The essence of DTM lies in learning an Interestingness Model (IM), for a given domain, from human annotated training data provided in the form of interesting facts from the KG. The IM thus learnt is applied to extract trivia for other entities of the same domain in the KG. We propose two different approaches for learning the IM - a) A Convolutional Neural Network (CNN) based approach and b) Fusion Based CNN (F-CNN) approach which combines both hand-crafted and CNN features. Experiments across two different domains - Bollywood Actors and Music Artists reveal that CNN automatically learns features which are relevant to the task and shows competitive performance relative to hand-crafted feature based baselines whereas F-CNN significantly improves the performance over the baseline approaches which use hand-crafted features alone. Overall, DTM achieves an F1 score of 0.81 and 0.65 in Bollywood Actors and Music Artists domains respectively.
Nausheen Fatma, Manoj Kumar Chinnakotla, Manish Shrivastava 0001
AAAI3
2017 Transition-Based Deep Input Linearization
abstract
Traditional methods for deep NLG adopt pipeline approaches comprising stages such as constructing syntactic input, predicting function words, linearizing the syntactic input and generating the surface forms.Though easier to visualize, pipeline approaches suffer from error propagation.In addition, information available across modules cannot be leveraged by all modules.We construct a transition-based model to jointly perform linearization, function word prediction and morphological generation, which considerably improves upon the accuracy compared to a pipelined baseline system.On a standard deep input linearization shared task, our system achieves the best results reported so far.
Ratish Puduppully, Yue Zhang 0004, Manish Shrivastava 0001
EACL (1)3
2017 Exploiting Morphological Regularities in Distributional Word Representations
abstract
We present an unsupervised, language agnostic approach for exploiting morphological regularities present in high dimensional vector spaces. We propose a novel method for generating embeddings of words from their morphological variants using morphological transformation operators. We evaluate this approach on MSR word analogy test set with an accuracy of 85% which is 12% higher than the previous best known system.
Arihant Gupta, Syed Sarfaraz Akhtar, Avijit Vajpayee, Arjit Srivastava, Madan Gopal Jhawar, Manish Shrivastava 0001
EMNLP6
2017 Improve performance of machine translation service using memcached
abstract
Sampark is machine translation system providing translations among nine pairs of Indian languages. Machine translation system is a class of natural language processing applications that is far more complex and highly compute intensive in nature. As the load on the deployed system increases, optimization becomes a challenge. Caching is one of the available options to improve the performance of a software system with increasing load which exhibit the characteristics of locality of reference. Sampark MT system being a natural language processing application exhibits this characteristic. Memcached is a well-known, simple, in-memory caching solution that has been applied to improve the performance of several distributed web applications in the past. This paper describes how memcached has been applied to improve the performance of Sampark machine translation service which is deployed on a large cluster of machines. By applying distributed caching to MT system the performance of the system has improved upto 40%.
Priyank Gupta, Rashid Ahmad 0003, Manish Shrivastava 0001, Pawan Kumar 0001, Mukul K. Sinha
ICCSA (7)3
2017 Significance of neural phonotactic models for large-scale spoken language identification
abstract
Language identification (LID) is vital frontend for spoken dialogue systems operating in diverse linguistic settings to reduce recognition and understanding errors. Existing LID systems which use low-level signal information for classification do not scale well due to exponential growth of parameters as the classes increase. They also suffer performance degradation due to the inherent variabilities of speech signal. In the proposed approach, we model the language-specific phonotactic information in speech using recurrent neural network for developing an LID system. The input speech signal is tokenized to phone sequences by using a common language-independent phone recognizer with varying phonetic coverage. We establish a causal relationship between phonetic coverage and LID performance. The phonotactics in the observed phone sequences are modeled using statistical and recurrent neural network language models to predict language-specific symbol from a universal phonetic inventory. Proposed approach is robust, computationally light weight and highly scalable. Experiments show that the convex combination of statistical and recurrent neural network language model (RNNLM) based phonotactic models significantly outperform a strong baseline system of Deep Neural Network (DNN) which is shown to surpass the performance of i-vector based approach for LID. The proposed approach outperforms the baseline models in terms of mean F1 score over 176 languages. Further we provide significant information-theoretic evidence to analyze the mechanism of the proposed approach.
Brij Mohan Lal Srivastava, Hari Krishna Vydana, Anil Kumar Vuppala, Manish Shrivastava 0001
IJCNN4
2016 Together we stand: Siamese Networks for Similar Question Retrieval
abstract
Community Question Answering (cQA) services like Yahoo! Answers 1 , Baidu Zhidao 2 , Quora 3 , StackOverflow 4 etc. provide a platform for interaction with experts and help users to obtain precise and accurate answers to their questions.The time lag between the user posting a question and receiving its answer could be reduced by retrieving similar historic questions from the cQA archives.The main challenge in this task is the "lexicosyntactic" gap between the current and the previous questions.In this paper, we propose a novel approach called "Siamese Convolutional Neural Network for cQA (SCQA)" to find the semantic similarity between the current and the archived questions.SCQA consist of twin convolutional neural networks with shared parameters and a contrastive loss function joining them.
Harish Yenala, Manoj Kumar Chinnakotla, Manish Shrivastava 0001
ACL (1)4
2016 Towards Sub-Word Level Compositions for Sentiment Analysis of Hindi-English Code Mixed Text
abstract
Sentiment analysis (SA) using code-mixed data from social media has several applications in opinion mining ranging from customer satisfaction to social campaign analysis in multilingual societies. Advances in this area are impeded by the lack of a suitable annotated dataset. We introduce a Hindi-English (Hi-En) code-mixed dataset for sentiment analysis and perform empirical analysis comparing the suitability and performance of various state-of-the-art SA methods in social media. In this paper, we introduce learning sub-word level representations in our LSTM (Subword-LSTM) architecture instead of character-level or word-level representations. This linguistic prior in our architecture enables us to learn the information about sentiment value of important morphemes. This also seems to work well in highly noisy text containing misspellings as shown in our experiments which is demonstrated in morpheme-level feature maps learned by our model. Also, we hypothesize that encoding this linguistic prior in the Subword-LSTM architecture leads to the superior performance. Our system attains accuracy 4-5% greater than traditional approaches on our dataset, and also outperforms the available system for sentiment analysis in Hi-En code-mixed text by 18%.
Ameya Prabhu, Manish Shrivastava 0001, Vasudeva Varma
COLING3
2016 Hand in Glove: Deep Feature Fusion Network Architectures for Answer Quality Prediction in Community Question Answering
abstract
Community Question Answering (cQA) forums have become a popular medium for soliciting direct answers to specific questions of users from experts or other experienced users on a given topic. However, for a given question, users sometimes have to sift through a large number of low-quality or irrelevant answers to find out the answer which satisfies their information need. To alleviate this, the problem of Answer Quality Prediction (AQP) aims to predict the quality of an answer posted in response to a forum question. Current AQP systems either learn models using - a) various hand-crafted features (HCF) or b) Deep Learning (DL) techniques which automatically learn the required feature representations. In this paper, we propose a novel approach for AQP known as - “Deep Feature Fusion Network (DFFN)” which combines the advantages of both hand-crafted features and deep learning based systems. Given a question-answer pair along with its metadata, the DFFN architecture independently - a) learns features from the Deep Neural Network (DNN) and b) computes hand-crafted features using various external resources and then combines them using a fully connected neural network trained to predict the final answer quality. DFFN is end-end differentiable and trained as a single system. We propose two different DFFN architectures which vary mainly in the way they model the input question/answer pair - DFFN-CNN uses a Convolutional Neural Network (CNN) and DFFN-BLNA uses a Bi-directional LSTM with Neural Attention (BLNA). Both these proposed variants of DFFN (DFFN-CNN and DFFN-BLNA) achieve state-of-the-art performance on the standard SemEval-2015 and SemEval-2016 benchmark datasets and outperforms baseline approaches which individually employ either HCF or DL based techniques alone.
Sai Praneeth Suggu, Kushwanth N. Goutham, Manoj Kumar Chinnakotla, Manish Shrivastava 0001
COLING4
2016 Transition-Based Syntactic Linearization with Lookahead Features
abstract
It has been shown that transition-based methods can be used for syntactic word ordering and tree linearization, achieving significantly faster speed compared with traditional best-first methods.State-of-the-art transitionbased models give competitive results on abstract word ordering and unlabeled tree linearization, but significantly worse results on labeled tree linearization.We demonstrate that the main cause for the performance bottleneck is the sparsity of SHIFT transition actions rather than heavy pruning.To address this issue, we propose a modification to the standard transition-based feature structure, which reduces feature sparsity and allows lookahead features at a small cost to decoding efficiency.Our model gives the best reported accuracies on all benchmarks, yet still being over 30 times faster compared with best-first-search.
Ratish Puduppully, Yue Zhang 0004, Manish Shrivastava 0001
HLT-NAACL3
2016 Shallow Parsing Pipeline - Hindi-English Code-Mixed Social Media Text
abstract
Arnav Sharma, Sakshi Gupta, Raveesh Motlani, Piyush Bansal, Manish Shrivastava, Radhika Mamidi, Dipti M. Sharma. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Arnav Sharma, Raveesh Motlani, Piyush Bansal, Manish Shrivastava 0001, Radhika Mamidi, Dipti Misra Sharma
HLT-NAACL5
2016 Mirror on the Wall: Finding Similar Questions with Deep Structured Topic Modeling
Manish Shrivastava 0001, Manoj Kumar Chinnakotla
PAKDD (2)2
2014 Do not do processing, when you can look up: Towards a Discrimination Net for WSD
abstract
The task of Word Sense Disambiguation (WSD) incorporates in its definition the role of 'context'.We present our work on the development of a tool which allows for automatic acquisition and ranking of 'context clues' for WSD.These clue words are extracted from the contexts of words appearing in a large monolingual corpus.These mined collection of contextual clues form a discrimination net in the sense that for targeted WSD, navigation of the net leads to the correct sense of a word given its context.Utilizing this resource we intend to develop efficient and light weight WSD based on look up and navigation of memoryresident knowledge base, thereby avoiding heavy computation which often prevents incorporation of any serious WSD in MT and search.The need for large quantities of sense marked data too can be reduced.
Diptesh Kanojia, Pushpak Bhattacharyya, Raj Dabre, Siddhartha Gunti, Manish Shrivastava 0001
GWC5
2006 Morphological Richness Offsets Resource Demand - Experiences in Constructing a POS Tagger for Hindi
Smriti Singh, Kuhoo Gupta, Manish Shrivastava 0001, Pushpak Bhattacharyya
ACL3