Mayank Singh 0001

dblp:96/4770 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0001-8757-2623ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 9 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Beyond Monolingual Assumptions: A Survey on Code-Switched NLP in the Era of Large Language Models across Modalities
abstract
Amidst the rapid advances of large language models (LLMs), most LLMs still struggle with mixed-language inputs, limited Codeswitching (CSW) datasets, and evaluation biases, which hinder their deployment in multilingual societies.This survey provides the first comprehensive analysis of CSW-aware LLM research, reviewing 327 studies spanning five research areas, 15+ NLP tasks, 30+ datasets, and 80+ languages.We categorize recent advances by architecture, training strategy, and evaluation methodology, outlining how LLMs have reshaped CSW modeling and identifying the challenges that persist.The paper concludes with a roadmap that emphasizes the need for inclusive datasets, fair evaluation, and linguistically grounded models to achieve truly multilingual capabilities.1 * Work done while interning at IIT Gandhinagar.Maria Riveena Arul, Vigneshwaran Shanmugasundaram, S Rajalakshmi, Bharathi Raja Chakravarthi, and C. N
Rajvee Sheth, Samridhi Raj Sinha, Mahavir Patil, Himanshu Beniwal, Mayank Singh 0001
ACL (1)5
2025 Model Hubs and Beyond: Analyzing Model Popularity, Performance, and Documentation
abstract
With the massive surge in ML models on platforms like Hugging Face, users often lose track and struggle to choose the best model for their downstream tasks, frequently relying on model popularity indicated by download counts, likes, or recency. We investigate whether this popularity aligns with actual model performance and how the comprehensiveness of model documentation correlates with both popularity and performance. In our study, we evaluated a comprehensive set of 500 Sentiment Analysis models on Hugging Face. This evaluation involved massive annotation efforts, with human annotators completing nearly 80,000 annotations, alongside extensive model training and evaluation. Our findings reveal that model popularity does not necessarily correlate with performance. Additionally, we identify critical inconsistencies in model card reporting: approximately 80% of the models analyzed lack detailed information about the model, training, and evaluation processes. Furthermore, about 88% of model authors overstate their models' performance in the model cards. Based on our findings, we provide a checklist of guidelines for users to choose good models for downstream tasks.
Pritam Kadasi, Sriman Reddy, Srivathsa Vamsi Chaturvedula, Rudranshu Sen, Agnish Saha, Soumavo Sikdar, Sayani Sarkar, Suhani Mittal, Rohit Jindal, Mayank Singh 0001
ICWSM10
2024 How Robust Are the QA Models for Hybrid Scientific Tabular Data? A Study Using Customized Dataset
abstract
Question-answering (QA) on hybrid scientific tabular and textual data deals with scientific information, and relies on complex numerical reasoning. In recent years, while tabular QA has seen rapid progress, understanding their robustness on scientific information is lacking due to absence of any benchmark dataset. To investigate the robustness of the existing state-of-the-art QA models on scientific hybrid tabular data, we propose a new dataset, “SciTabQA”, consisting of 822 question-answer pairs from scientific tables and their descriptions. With the help of this dataset, we assess the state-of-the-art Tabular QA models based on their ability (i) to use heterogeneous information requiring both structured data (table) and unstructured data (text) and (ii) to perform complex scientific reasoning tasks. In essence, we check the capability of the models to interpret scientific tables and text. Our experiments show that “SciTabQA” is an innovative dataset to study question-answering over scientific heterogeneous data. We benchmark three state-of-the-art Tabular QA models, and find that the best F1 score is only 0.462.
Akash Ghosh, Venkata Sahith Bathini, Niloy Ganguly, Pawan Goyal 0002, Mayank Singh 0001
LREC/COLING5
2023 LineEX: Data Extraction from Scientific Line Charts
abstract
In this paper, we introduce LineEX that extracts data from scientific line charts. We adapt existing vision transformers and pose detection methods and showcase significant performance gains over existing SOTA baselines. We also propose a new loss function and present its effectiveness against existing loss functions. In addition, we synthetically created the largest line chart dataset comprising 430K images. The code is available at https: //github.com/Shiva-sankaran/LineEX.
Shivasankaran V. P, Muhammad Yusuf Hassan, Mayank Singh 0001
WACV3
2023 Tables to LaTeX: structure and content extraction from scientific tables
abstract
Scientific documents contain tables that list important information in a concise fashion. Structure and content extraction from tables embedded within PDF research documents is a very challenging task due to the existence of visual features like spanning cells and content features like mathematical symbols and equations. Most existing table structure identification methods tend to ignore these academic writing features. In this paper, we adapt the transformer-based language modeling paradigm for scientific table structure and content extraction. Specifically, the proposed model converts a tabular image to its corresponding LaTeX source code. Overall, we outperform the current state-of-the-art baselines and achieve an exact match accuracy of 70.35 and 49.69% on table structure and content extraction, respectively. Further analysis demonstrates that the proposed models efficiently identify the number of rows and columns, the alphanumeric characters, the LaTeX tokens, and symbols.
Pratik Kayal, Mrinal Anand, Mayank Singh 0001
Int. J. Document Anal. Recognit.4
2022 The Bull and the Bear: Summarizing Stock Market Discussions
abstract
Stock market investors debate and heavily discuss stock ideas, investing strategies, news and market movements on social media platforms. The discussions are significantly longer in length and require extensive domain expertise for understanding. In this paper, we curate such discussions and construct a first-of-its-kind of abstractive summarization dataset. Our curated dataset consists of 7888 Reddit posts and manually constructed summaries for 400 posts. We robustly evaluate the summaries and conduct experiments on SOTA summarization tools to showcase their limitations. We plan to make the dataset publicly available. The sample dataset is available here: https://dhyeyjani.github.io/RSMC
Dhyey Jani, Jay Shah, Devanshu Thakar, Varun Jain, Mayank Singh 0001
LREC6
2021 TabLeX: A Benchmark Dataset for Structure and Content Information Extraction from Scientific Tables
abstract
Information Extraction (IE) from the tables present in scientific articles is challenging due to complicated tabular representations and complex embedded text. This paper presents TabLeX, a large-scale benchmark dataset comprising table images generated from scientific articles. TabLeX consists of two subsets, one for table structure extraction and the other for table content extraction. Each table image is accompanied by its corresponding LATEX source code. To facilitate the development of robust table IE tools, TabLeX contains images in different aspect ratios and in a variety of fonts. Our analysis sheds light on the shortcomings of current state-of-the-art table extraction models and shows that they fail on even simple table images. Towards the end, we experiment with a transformer-based existing baseline to report performance scores. In contrast to the static benchmarks, we plan to augment this dataset with more complex and diverse tables at regular intervals.
Pratik Kayal, Mayank Singh 0001
ICDAR (2)3
2021 ICDAR 2021 Competition on Scientific Table Image Recognition to LaTeX
Pratik Kayal, Mrinal Anand, Mayank Singh 0001
ICDAR (4)4
2021 Quality Evaluation of the Low-Resource Synthetically Generated Code-Mixed Hinglish Text
abstract
In this shared task, we seek the participating teams to investigate the factors influencing the quality of the code-mixed text generation systems.We synthetically generate codemixed Hinglish sentences using two distinct approaches and employ human annotators to rate the generation quality.We propose two subtasks, quality rating prediction and annotators' disagreement prediction of the synthetic Hinglish dataset.The proposed subtasks will put forward the reasoning and explanation of the factors influencing the quality and human perception of the code-mixed text.
Mayank Singh 0001
INLG2
2021 Vartalaap: What Drives #AirQuality Discussions: Politics, Pollution or Pseudo-science?
abstract
Air pollution is a global challenge for cities across the globe. Understanding the public perception of air pollution can help policymakers engage better with the public and appropriately introduce policies. Accurate public perception can also help people to identify the health risks of air pollution and act accordingly. Unfortunately, current techniques for determining perception are not scalable: it involves surveying few hundred people with questionnaire-based surveys. Using the advances in natural language processing (NLP), we propose a more scalable solution called Vartalaap to gauge public perception of air pollution via the microblogging social network Twitter. We curated a dataset of more than 1.2M tweets discussing Delhi-specific air pollution. We find that (unfortunately) the public is supportive of unproven mitigation strategies to reduce pollution, thus risking their health due to a false sense of security. We also find that air quality is a year-long problem, but the discussions are not proportional to the level of pollution and spike up when pollution is more visible. The information required by Vartalaap is publicly available and, as such, it can be immediately applied to study different societal issues across the world.
Rishiraj Adhikary, Zeel B. Patel, Tanmay Srivastava, Nipun Batra 0001, Mayank Singh 0001, Udit Bhatia, Sarath Guttikunda
Proc. ACM Hum. Comput. Interact.5
2020 CovidExplorer: A Multi-faceted AI-based Search and Visualization Engine for COVID-19 Information
abstract
The entire world is engulfed in the fight against the COVID-19 pandemic, leading to a significant surge in research experiments, government policies, and social media discussions. A multi-modal information access and data visualization platform can play a critical role in supporting research aimed at understanding and developing preventive measures for the pandemic. In this paper, we present a multi-faceted AI-based search and visualization engine, CovidExplorer. Our system aims to help researchers understand current state-of-the-art COVID-19 research, identify research articles relevant to their domain, and visualize real-time trends and statistics of COVID-19 cases. In contrast to other existing systems, CovidExplorer also brings in India-specific topical discussions on social media to study different aspects of COVID-19. The system, demo video, and the datasets are available at http://covidexplorer.in.
Heer Ambavi, Kavita Vaishnaw, Udit Vyas, Abhisht Tiwari, Mayank Singh 0001
CIKM5
2020 NLPExplorer: Exploring the Universe of NLP Papers
Monarch Parmar, Naman Jain, Pranjali Jain, P. Jayakrishna Sahit, Soham Pachpande, Shruti Singh 0001, Mayank Singh 0001
ECIR (2)7
2020 Quantifying Nonrandomness in Evolving Networks
abstract
Complex systems have been successfully modeled as networks exhibiting the varying extent of randomness and nonrandomness. Network scientists contemplate randomness as one of the most desirable characteristics for real complex systems' efficient performance. However, the current methodologies for randomness (or nonrandomness) quantification are nontrivial. In this article, we empirically showcase severe limitations associated with the state-of-the-art graph spectral-based quantification approaches. Addressing these limitations led to the proposal of a novel spectrum-based methodology that leverages configuration models as a reference network to quantify the nonrandomness in a given candidate network. Besides, we derive mathematical formulations for demonstrating the dependence of nonrandomness on three structural properties: modularity, clustering, and the highest degree node's growth rate. We also introduce a novel graph signature (termed “cumulative spectral difference”) to visualize the nonrandomness in the network. Later, this article also discusses the relationship between the proposed nonrandomness measure and the diffusion affinity of networks. Toward the end, this article extensively discusses observations emerging from these signatures for both real-world and simulated networks.
Pradumn Kumar Pandey, Mayank Singh 0001
IEEE Trans. Comput. Soc. Syst.2
2019 Automated Early Leaderboard Generation from Comparative Tables
Mayank Singh 0001, Rajdeep Sarkar, Atharva Vyas, Pawan Goyal 0002, Animesh Mukherjee 0001, Soumen Chakrabarti
ECIR (1)1
2017 Relay-Linking Models for Prominence and Obsolescence in Evolving Networks
abstract
The rate at which nodes in evolving social networks acquire links (friends, citations) shows complex temporal dynamics. Preferential attachment and link copying models, while enabling elegant analysis, only capture rich-gets-richer effects, not aging and decline. Recent aging models are complex and heavily parameterized; most involve estimating 1-3 parameters per node. These parameters are intrinsic: they explain decline in terms of events in the past of the same node, and do not explain, using the network, where the linking attention might go instead. We argue that traditional characterization of linking dynamics are insufficient to judge the faithfulness of models. We propose a new temporal sketch of an evolving graph, and introduce several new characterizations of a network's temporal dynamics. Then we propose a new family of frugal aging models with no per-node parameters and only two global parameters. Our model is based on a surprising inversion or undoing of triangle completion, where an old node relays a citation to a younger follower in its immediate vicinity. Despite very few parameters, the new family of models shows remarkably better fit with real data. Before concluding, we analyze temporal signatures for various research communities yielding further insights into their comparative dynamics. To facilitate reproducible research, we shall soon make all the codes and the processed dataset available in the public domain.
Mayank Singh 0001, Rajdeep Sarkar, Pawan Goyal 0002, Animesh Mukherjee 0001, Soumen Chakrabarti
KDD1
2016 OCR++: A Robust Framework For Information Extraction from Scholarly Articles
abstract
This paper proposes OCR++, an open-source framework designed for a variety of information extraction tasks from scholarly articles including metadata (title, author names, affiliation and e-mail), structure (section headings and body text, table and figure headings, URLs and footnotes) and bibliography (citation instances and references). We analyze a diverse set of scientific articles written in English to understand generic writing patterns and formulate rules to develop this hybrid framework. Extensive evaluations show that the proposed framework outperforms the existing state-of-the-art tools by a large margin in structural information extraction along with improved performance in metadata and bibliography extraction tasks, both in terms of accuracy (around 50% improvement) and processing time (around 52% improvement). A user experience study conducted with the help of 30 researchers reveals that the researchers found this system to be very helpful. As an additional objective, we discuss two novel use cases including automatically extracting links to public datasets from the proceedings, which would further accelerate the advancement in digital libraries. The result of the framework can be exported as a whole into structured TEI-encoded documents. Our framework is accessible online at http://www.cnergres.iitkgp.ac.in/OCR++/home/.
Mayank Singh 0001, Barnopriyo Barua, Priyank Palod, Manvi Garg, Sidhartha Satapathy, Samuel Bushi, Kumar Ayush, Krishna Sai Rohith, Tulasi Gamidi, Pawan Goyal 0002, Animesh Mukherjee 0001
COLING1
2016 FeRoSA: A Faceted Recommendation System for Scientific Articles
Tanmoy Chakraborty 0002, Amrith Krishna, Mayank Singh 0001, Niloy Ganguly, Pawan Goyal 0002, Animesh Mukherjee 0001
PAKDD (2)3
2015 The Role Of Citation Context In Predicting Long-Term Citation Profiles: An Experimental Study Based On A Massive Bibliographic Text Dataset
abstract
The impact and significance of a scientific publication is measured mostly by the number of citations it accumulates over the years. Early prediction of the citation profile of research articles is a significant as well as challenging problem. In this paper, we argue that features gathered from the citation contexts of the research papers can be very relevant for citation prediction. Analyzing a massive dataset of nearly 1.5 million computer science articles and more than 26 million citation contexts, we show that average countX (number of times a paper is cited within the same article) and average citeWords (number of words within the citation context) discriminate between various citation ranges as well as citation categories. We use these features in a stratified learning framework for future citation prediction. Experimental results show that the proposed model significantly outperforms the existing citation prediction models by a margin of 8-10% on an average under various experimental settings. Specifically, the features derived from the citation context help in predicting long-term citation behavior.
Mayank Singh 0001, Vikas Patidar, Suhansanu Kumar, Tanmoy Chakraborty 0002, Animesh Mukherjee 0001, Pawan Goyal 0002
CIKM1