Mayank Jobanputra

dblp:220/8952 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0002-8802-2401ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Question answering and dialogue systems · 33% Language models and text generation · 33% Deep learning architectures and training · 33%
Software engineering, system software, and programming languages
1 paper
Software testing · 44% Software maintenance and evolution · 44% Empirical software engineering · 13%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › compositional generalization
length generalization
0.912025
Born a Transformer - Always a Transformer? On the Effect of Pretraining on Architectural Abilities · NeurIPS 2025
Natural language and speech › Question answering and dialogue systems
multi-hop reasoning
0.912025
Tree-of-Quote Prompting Improves Factuality and Attribution in Multi-Hop and Medical Reasoning · EMNLP 2025
Machine learning › Deep learning architectures and training › neural network expressivity
transformer expressivity
0.912025
Born a Transformer - Always a Transformer? On the Effect of Pretraining on Architectural Abilities · NeurIPS 2025
Software maintenance and evolution
code clone detection
0.612022
Mining Similar Methods for Test Adaptation · IEEE Trans. Software Eng. 2022
Software testing
test reuse
0.612022
Mining Similar Methods for Test Adaptation · IEEE Trans. Software Eng. 2022
Medical and health informatics › clinical decision-making
medical reasoning
0.312025
Tree-of-Quote Prompting Improves Factuality and Attribution in Multi-Hop and Medical Reasoning · EMNLP 2025
Empirical software engineering
mining software repositories
0.212022
Mining Similar Methods for Test Adaptation · IEEE Trans. Software Eng. 2022

Methods — techniques the papers use, named apart from their topics

tree-of-quote prompting · 1.7retrieval and copying tasks · 0.9mechanistic interpretability · 0.9fine-tuning · 0.9test mining · 0.6recommendation · 0.6
YearPublicationVenuePosition
2025 Tree-of-Quote Prompting Improves Factuality and Attribution in Multi-Hop and Medical Reasoning
abstract
Justin Xu, Yiming Li, Zizheng Zhang, Augustine Yui Hei Luk, Mayank Jobanputra, Samarth Oza, Ashley Murray, Meghana Reddy Kasula, Andrew Parker, David W Eyre. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Justin Xu, Zizheng Zhang, Augustine Yui Hei Luk, Mayank Jobanputra, Samarth Oza, Ashley Murray, Meghana Reddy Kasula, David Eyre 0001
EMNLP5
2025 Born a Transformer - Always a Transformer? On the Effect of Pretraining on Architectural Abilities
abstract
Transformers have theoretical limitations in modeling certain sequence-to-sequence tasks, yet it remains largely unclear if these limitations play a role in large-scale pretrained LLMs, or whether LLMs might effectively overcome these constraints in practice due to the scale of both the models themselves and their pretraining data. We explore how these architectural constraints manifest after pretraining by studying a family of *retrieval* and *copying* tasks inspired by Liu et al. [2024a]. We use a recently proposed framework for studying length generalization [Huang et al., 2025] to provide guarantees for each of our settings. Empirically, we observe an *induction-versus-anti-induction asymmetry*, where pretrained models are better at retrieving tokens to the right (induction) rather than the left (anti-induction) of a query token. This asymmetry disappears upon targeted fine-tuning if length-generalization is guaranteed by theory. Mechanistic analysis reveals that this asymmetry is connected to the differences in the strength of induction versus anti-induction circuits within pretrained transformers. We validate our findings through practical experiments on real-world tasks demonstrating reliability risks. Our results highlight that pretraining selectively enhances certain transformer capabilities, but does not overcome fundamental length-generalization limits.
Mayank Jobanputra, Yana Veitsman, Yash Raj Sarrof, Aleksandra Bakalova, Vera Demberg, Ellie Pavlick, Michael Hahn 0001
NeurIPS1
2024 Retrieval-Augmented Modular Prompt Tuning for Low-Resource Data-to-Text Generation
abstract
Data-to-text (D2T) generation describes the task of verbalizing data, often given as attribute-value pairs. While this task is relevant for many different data domains beyond the traditionally well-explored tasks of weather forecasting, restaurant recommendations, and sports reporting, a major challenge to the applicability of data-to-text generation methods is typically data sparsity. For many applications, there is extremely little training data in terms of attribute-value inputs and target language outputs available for training a model. Given the sparse data setting, recently developed prompting methods seem most suitable for addressing D2T tasks since they do not require substantial amounts of training data, unlike finetuning approaches. However, prompt-based approaches are also challenging, as a) the design and search of prompts are non-trivial; and b) hallucination problems may occur because of the strong inductive bias of these models. In this paper, we propose a retrieval-augmented modular prompt tuning () method, which constructs prompts that fit the input data closely, thereby bridging the domain gap between the large-scale language model and the structured input data. Experiments show that our method generates texts with few hallucinations and achieves state-of-the-art performance on a dataset for drone handover message generation.
Xudong Hong 0002, Mayank Jobanputra, Mattes Warning, Vera Demberg
LREC/COLING3
2022 Samanantar: The Largest Publicly Available Parallel Corpora Collection for 11 Indic Languages
abstract
Abstract We present Samanantar, the largest publicly available parallel corpora collection for Indic languages. The collection contains a total of 49.7 million sentence pairs between English and 11 Indic languages (from two language families). Specifically, we compile 12.4 million sentence pairs from existing, publicly available parallel corpora, and additionally mine 37.4 million sentence pairs from the Web, resulting in a 4× increase. We mine the parallel sentences from the Web by combining many corpora, tools, and methods: (a) Web-crawled monolingual corpora, (b) document OCR for extracting sentences from scanned documents, (c) multilingual representation models for aligning sentences, and (d) approximate nearest neighbor search for searching in a large collection of sentences. Human evaluation of samples from the newly mined corpora validate the high quality of the parallel sentences across 11 languages. Further, we extract 83.4 million sentence pairs between all 55 Indic language pairs from the English-centric parallel corpus using English as the pivot language. We trained multilingual NMT models spanning all these languages on Samanantar which outperform existing models and baselines on publicly available benchmarks, such as FLORES, establishing the utility of Samanantar. Our data and models are available publicly at Samanantar and we hope they will help advance research in NMT and multilingual NLP for Indic languages.
Gowtham Ramesh, Sumanth Doddapaneni, Aravinth Bheemaraj, Mayank Jobanputra, Raghavan AK, Ajitesh Sharma, Sujit Sahoo, Harshita Diddee, Mahalakshmi J, Divyanshu Kakwani, Aswin Pradeep, Srihari Nagaraj, Vivek Raghavan, Anoop Kunchukuttan, Mitesh Shantadevi Khapra
Trans. Assoc. Comput. Linguistics4
2022 Mining Similar Methods for Test Adaptation
abstract
Developers may choose to implement a library despite the existence of similar libraries, considering factors such as computational performance, language or platform dependency, accuracy, convenience, and completeness of an API. As a result, GitHub hosts several library projects that have overlaps in their functionalities. These overlaps have been of interest to developers from the perspective of code reuse or the preference of one implementation over the other. Through an empirical study, we explore the extent and nature of existence of these similarities in the library functions. We have further studied whether the similarity of functions across different libraries and their associated test suites can be leveraged to reveal defects in one another. We see scope for effectively using the mining of test suites from the perspective of revealing defects in a program or its documentation. Another noteworthy observation made in the study is that similar functions may exist across libraries implemented in the same language as well as in different languages. Identifying the challenges that lie in building a testing tool, we automate the entire process inMetallicus, a test mining and recommendation tool.Metallicusreturns a test suite for the given input of a query function and a template for its test suite. On a dataset of query functions taken from libraries implemented in Java or Python,Metallicusrevealed 46 defects.
Devika Sondhi, Mayank Jobanputra, Divya Rani, Salil Purandare, Rahul Purandare
IEEE Trans. Software Eng.2
2021 Detecting and Analyzing Collusive Entities on YouTube
abstract
YouTube sells advertisements on the posted videos, which in turn enables the content creators to monetize their videos. As an unintended consequence, this has proliferated various illegal activities such as artificial boosting of views, likes, comments, and subscriptions. We refer to such videos (gaining likes and comments artificially) and channels (gaining subscriptions artificially) as “collusive entities.” Detecting such collusive entities is an important yet challenging task. Existing solutions mostly deal with the problem of spotting fake views, spam comments, fake content, and so on, and oftentimes ignore how such fake activities emerge via collusion. Here, we collect a large dataset consisting of two types of collusive entities on YouTube— videos submitted to gain collusive likes and comment requests and channels submitted to gain collusive subscriptions. We begin by providing an in-depth analysis of collusive entities on YouTube fostered by various blackmarket services . Following this, we propose models to detect three types of collusive YouTube entities: videos seeking collusive likes, channels seeking collusive subscriptions, and videos seeking collusive comments. The third type of entity is associated with temporal information. To detect videos and channels for collusive likes and subscriptions, respectively, we utilize one-class classifiers trained on our curated collusive entities and a set of novel features. The SVM-based model shows significant performance with a true positive rate of 0.911 and 0.910 for detecting collusive videos and collusive channels, respectively. To detect videos seeking collusive comments, we propose CollATe , a novel end-to-end neural architecture that leverages time-series information of posted comments along with static metadata of videos. CollATe is composed of three components: metadata feature extractor (which derives metadata-based features from videos), anomaly feature extractor (which utilizes the time-series data to detect sudden changes in the commenting activity), and comment feature extractor (which utilizes the text of the comments posted during collusion and computes a similarity score between the comments). Extensive experiments show the effectiveness of CollATe (with a true positive rate of 0.905) over the baselines.
Hridoy Sankar Dutta, Mayank Jobanputra, Himani Negi, Tanmoy Chakraborty 0002
ACM Trans. Intell. Syst. Technol.2