EDBT 2026 Demo / reviewers in the wild / expert
Mukund Rungta
dblp:213/1771
· DBLP profile ↗
6ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
1 paper |
Empirical software engineering · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Artificial intelligence
1 paper |
Information extraction and text analysis · 100% |
Topics — the 2 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Empirical software engineering
mining software repositories |
0.7 | 1 | 2023 | Forgotten Knowledge: Examining the Citational Amnesia in NLP · ACL (1) 2023 |
Computational social science and digital humanities
scientometrics |
0.6 | 1 | 2022 | Geographic Citation Gaps in NLP Research · EMNLP 2022 |
Methods — techniques the papers use, named apart from their topics
empirical analysis · 1.3meta-information extraction · 1.1citation network analysis · 1.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Alternate Preference Optimization for Unlearning Factual Knowledge in Large Language ModelsabstractMachine unlearning aims to efficiently eliminate the influence of specific training data, known as the forget set, from the model. However, existing unlearning methods for Large Language Models (LLMs) face a critical challenge: they rely solely on negative feedback to suppress responses related to the forget set, which often results in nonsensical or inconsistent outputs, diminishing model utility and posing potential privacy risks. To address this limitation, we propose a novel approach called Alternate Preference Optimization (AltPO), which combines negative feedback with in-domain positive feedback on the forget set. Additionally, we introduce new evaluation metrics to assess the quality of responses related to the forget set. Extensive experiments show that our approach not only enables effective unlearning but also avoids undesirable model behaviors while maintaining overall model performance. Anmol Reddy Mekala, Vineeth Dorna, Shreya Dubey, Abhishek Lalwani, David Koleczek, Mukund Rungta, Sadid A. Hasan, Elita A. Lobo |
COLING | 6 |
| 2024 | HiGen: Hierarchy-Aware Sequence Generation for Hierarchical Text ClassificationabstractHierarchical text classification (HTC) is a complex subtask under multi-label text classification, characterized by a hierarchical label taxonomy and data imbalance. The best-performing models aim to learn a static representation by combining document and hierarchical label information. However, the relevance of document sections can vary based on the hierarchy level, necessitating a dynamic document representation. To address this, we propose HiGen, a text-generation-based framework utilizing language models to encode dynamic text representations. We introduce a level-guided loss function to capture the relationship between text and label name semantics. Our approach incorporates a task-specific pretraining strategy, adapting the language model to in-domain knowledge and significantly enhancing performance for classes with limited examples. Furthermore, we present a new and valuable dataset called ENZYME, designed for HTC, which comprises articles from PubMed with the goal of predicting Enzyme Commission (EC) numbers. Through extensive experiments on the ENZYME dataset and the widely recognized WOS and NYT datasets, our methodology demonstrates superior performance, surpassing existing approaches while efficiently handling data and mitigating class imbalance. We release our code and dataset here: https://github.com/viditjain99/HiGen. Vidit Jain, Mukund Rungta, Yuchen Zhuang, Yue Yu 0001, Mu Gao, Jeffrey Skolnick, Chao Zhang 0014 |
EACL (1) | 2 |
| 2023 | Forgotten Knowledge: Examining the Citational Amnesia in NLPabstractCiting papers is the primary method through which modern scientific writing discusses and builds on past work.Collectively, citing a diverse set of papers (in time and area of study) is an indicator of how widely the community is reading.Yet there is little work looking at broad temporal patterns of citation.This work, systematically and empirically examines: How far back in time do we tend to go to cite papers?How has that changed over time, and what factors correlate with this citational attention/amnesia?We chose NLP as our domain of interest, and analyzed ∼71.5K papers to show and quantify several key trends in citation.Notably, ∼62% of cited papers are from the immediate five years prior to publication, whereas only ∼17% are more than ten years old.Furthermore, we show that the median age and age diversity of cited papers was steadily increasing from 1990 to 2014, but since then the trend has reversed, and current NLP papers have an all-time low temporal citation diversity.Finally, we show that unlike the 1990s, the highly cited papers in the last decade were also papers with the least citation diversity; likely contributing to the intense (and arguably harmful) recency focus.Code, data, and a demo are available at the project homepage.1 2 Janvijay Singh, Mukund Rungta, Diyi Yang, Saif M. Mohammad |
ACL (1) | 2 |
| 2022 | Geographic Citation Gaps in NLP ResearchabstractIn a fair world, people have equitable opportunities to education, to conduct scientific research, to publish, and to get credit for their work, regardless of where they live.However, it is common knowledge among researchers that a vast number of papers accepted at top NLP venues come from a handful of western countries and (lately) China; whereas, very few papers from Africa and South America get published.Similar disparities are also believed to exist for paper citation counts.In the spirit of "what we do not measure, we cannot improve", this work asks a series of questions on the relationship between geographical location and publication success (acceptance in top NLP venues and citation impact).We first created a dataset of 70,000 papers from the ACL Anthology, extracted their meta-information, and generated their citation network.We then show that not only are there substantial geographical disparities in paper acceptance and citation but also that these disparities persist even when controlling for a number of variables such as venue of publication and sub-field of NLP.Further, despite some steps taken by the NLP community to improve geographical diversity, we show that the disparity in publication metrics across locations is still on an increasing trend since the early 2000s.We release our code and dataset here Mukund Rungta, Janvijay Singh, Saif M. Mohammad, Diyi Yang |
EMNLP | 1 |
| 2020 | TransKP: Transformer based Key-Phrase ExtractionabstractIncreased connectivity has led to a sharp rise in the creation and availability of structured and unstructured text content, with millions of new documents being generated every minute. Key-phrase extraction is the process of finding the most important words and phrases which best capture the overall meaning and topics of a text document. Common techniques follow supervised or unsupervised methods for extractive or abstractive key-phrase extraction, but struggle to perform well and generalize to different datasets. In this paper, we follow a supervised, extractive approach and model the key-phrase extraction problem as a sequence labeling task. We utilize the power of transformers on sequential tasks and explore the effect of initializing the embedding layer of the model with pre-trained weights. We test our model on different standard key-phrase extraction datasets and our results significantly outperform all baselines as well as state-of-the-art scores on all the datasets. Mukund Rungta, Rishabh Kumar, Mehak Preet Dhaliwal, Hemant Tiwari, Vanraj Vala |
IJCNN | 1 |
| 2017 | Classification-Based Adaptive Web ScraperabstractWeb scraping is an important problem in computer science. The problem with the commonly-used position or structure-based web scraping tools is that they need to be manually reconfigured as soon as the structure of the web page changes. In this paper, we try to solve this problem of information extraction for web pages consisting of repetitive blocks. We extract these blocks and their constituent attributes, using a novel classification-based approach. Our approach gives high accuracy when used to extract product-offers from an offer-aggregator website. It is also highly adaptive to the changing structure of a website. Ujwal B. V. S, Bharat Gaind, Abhishek Kundu, Anusha Holla, Mukund Rungta |
ICMLA | 5 |