Mousumi Akter 0001

dblp:200/7260-1 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0003-2665-2077ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 60% Information extraction and text analysis · 17% Machine translation · 8%
Network and information security
1 paper
Privacy and data protection · 100%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
text summarization
1.422025
Benchmarking LLMs on Semantic Overlap Summarization · EMNLP 2025
Learning to Generate Overlap Summaries through Noisy Synthetic Data · EMNLP 2022
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
Benchmarking LLMs on Semantic Overlap Summarization · EMNLP 2025
Natural language and speech › Language models and text generation › text summarization
multi-document summarization
0.912025
Benchmarking LLMs on Semantic Overlap Summarization · EMNLP 2025
Natural language and speech › Language models and text generation › prompting
prompt sensitivity
0.912025
Benchmarking LLMs on Semantic Overlap Summarization · EMNLP 2025
Natural language and speech › Information extraction and text analysis › lexical semantics
word analogy
0.712023
On Evaluation of Bangla Word Analogies · EMNLP 2023
Natural language and speech › Information extraction and text analysis › distributional semantics
word embedding evaluation
0.712023
On Evaluation of Bangla Word Analogies · EMNLP 2023
Natural language and speech › Machine translation › machine translation evaluation
automatic evaluation metrics
0.612022
SEM-F1: an Automatic Way for Semantic Evaluation of Multi-Narrative Overlap Summaries at Scale · EMNLP 2022
Machine learning › Deep learning architectures and training
data augmentation
0.612022
Learning to Generate Overlap Summaries through Noisy Synthetic Data · EMNLP 2022
Natural language and speech › Language models and text generation › text summarization
summarization evaluation
0.612022
SEM-F1: an Automatic Way for Semantic Evaluation of Multi-Narrative Overlap Summaries at Scale · EMNLP 2022
Machine learning › Generative modeling
synthetic data generation
0.612022
Learning to Generate Overlap Summaries through Noisy Synthetic Data · EMNLP 2022
Privacy and data protection › privacy policy
privacy policy analysis
0.312025
Benchmarking LLMs on Semantic Overlap Summarization · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

prompting taxonomy · 1.7human evaluation · 1.7automatic metrics · 1.7word embeddings · 0.7seq-to-seq models · 0.6fine-tuning · 0.6SEM-F1 · 0.6ROUGE · 0.6
YearPublicationVenuePosition
2025 Benchmarking LLMs on Semantic Overlap Summarization
abstract
Semantic Overlap Summarization (SOS) is a multi-document summarization task focused on extracting the common information shared across alternative narratives which is a capability that is critical for trustworthy generation in domains such as news, law, and healthcare.We benchmark popular Large Language Models (LLMs) on SOS and introduce PrivacyPolicy-Pairs (3P), a new dataset of 135 high-quality samples from privacy policy documents, which complements existing resources and broadens domain coverage.Using the TELeR prompting taxonomy, we evaluate nearly one million LLM-generated summaries across two SOS datasets and conduct human evaluation on a curated subset.Our analysis reveals strong prompt sensitivity, identifies which automatic metrics align most closely with human judgments, and provides new baselines for future SOS research 1 .
John Salvador, Naman Bansal, Mousumi Akter 0001, Souvika Sarkar, Anupam Das 0008, Shubhra Kanti Karmaker Santu
EMNLP3
2025 LLMs as Meta-Reviewers' Assistants: A Case Study
abstract
Eftekhar Hossain, Sanjeev Kumar Sinha, Naman Bansal, R. Alexander Knipper, Souvika Sarkar, John Salvador, Yash Mahajan, Sri Ram Pavan Kumar Guttikonda, Mousumi Akter, Md. Mahadi Hassan, Matthew Freestone, Matthew C. Williams Jr., Dongji Feng, Santu Karmaker. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Eftekhar Hossain, Sanjeev Kumar Sinha, Naman Bansal, R. Alexander Knipper, Souvika Sarkar, John Salvador, Yash Mahajan, Sri Guttikonda, Mousumi Akter 0001, Md. Mahadi Hassan, Matthew Freestone, Matthew C. Williams Jr., Dongji Feng, Shubhra Kanti Karmaker Santu
NAACL (Long Papers)9
2023 On Evaluation of Bangla Word Analogies
abstract
This paper presents a benchmark dataset of Bangla word analogies for evaluating the quality of existing Bangla word embeddings.Despite being the 7 th largest spoken language in the world, Bangla is still a low-resource language and popular NLP models often struggle to perform well on Bangla data sets.Therefore, developing a robust evaluation set is crucial for benchmarking and guiding future research on improving Bangla word embeddings, which is currently missing.To address this issue, we introduce a new evaluation set of 16,678 unique word analogies in Bangla as well as a translated and curated version of the original Mikolov dataset (10,594 samples) in Bangla.Our experiments with different state-of-the-art embedding models reveal that current Bangla word embeddings struggle to achieve high accuracy on both data sets, demonstrating a significant gap in multilingual NLP research.
Mousumi Akter 0001, Souvika Sarkar, Shubhra Kanti Karmaker Santu
EMNLP1
2022 Rank-Aware Gain-Based Evaluation of Extractive Summarization
abstract
ROUGE has long been a popular metric for evaluating text summarization tasks as it eliminates time-consuming and costly human evaluations. However, ROUGE is not a fair evaluation metric for extractive summarization task as it is entirely based on lexical overlap. Additionally, ROUGE ignores the quality of the ranker for extractive summarization which performs the actual sentence/phrase extraction job. The main focus of the thesis is to design a nCG (normalized cumulative gain)-based evaluation metric for extractive summarization that is both rank-aware and semantic-aware (called Sem-nCG). One fundamental contribution of the work is that it demonstrates how we can generate more reliable semantic-aware ground truths for evaluating extractive summarization tasks without any additional human intervention. To the best of our knowledge, this work is the first of its kind. Preliminary experimental results demonstrate that the new Sem-nCG metric is indeed semantic-aware and also exhibits higher correlation with human judgement for single document summarization when single reference is considered.
Mousumi Akter 0001
CIKM1
2022 Semantic Overlap Summarization among Multiple Alternative Narratives: An Exploratory Study
abstract
In this paper, we introduce an important yet relatively unexplored NLP task called Semantic Overlap Summarization (SOS), which entails generating a single summary from multiple alternative narratives which can convey the common information provided by those narratives. As no benchmark dataset is readily available for this task, we created one by collecting 2,925 alternative narrative pairs from the web and then, went through the tedious process of manually creating 411 different reference summaries by engaging human annotators. As a way to evaluate this novel task, we first conducted a systematic study by borrowing the popular ROUGE metric from text-summarization literature and discovered that ROUGE is not suitable for our task. Subsequently, we conducted further human annotations to create 200 document-level and 1,518 sentence-level ground-truth overlap labels. Our experiments show that the sentence-wise annotation technique with three overlap labels, i.e., Absent (A), Partially-Present (PP), and Present (P), yields a higher correlation with human judgment and higher inter-rater agreement compared to the ROUGE metric.
Naman Bansal, Mousumi Akter 0001, Shubhra Kanti Karmaker Santu
COLING2
2022 SEM-F1: an Automatic Way for Semantic Evaluation of Multi-Narrative Overlap Summaries at Scale
abstract
Recent work has introduced an important yet relatively under-explored NLP task called Semantic Overlap Summarization (SOS) that entails generating a summary from multiple alternative narratives which conveys the common information provided by those narratives.Previous work also published a benchmark dataset for this task by collecting 2, 925 alternative narrative pairs from the web and manually annotating 411 different reference summaries by engaging human annotators.In this paper, we exclusively focus on the automated evaluation of the SOS task using the benchmark dataset.More specifically, we first use the popular ROUGE metric from text-summarization literature and conduct a systematic study to evaluate the SOS task.Our experiments discover that ROUGE is not suitable for this novel task and therefore, we propose a new sentencelevel precision-recall style automated evaluation metric, called SEM-F 1 (Semantic F 1 ).It is inspired by the benefits of the sentence-wise annotation technique using overlap labels reported by the previous work.Our experiments show that the proposed SEM-F 1 metric yields a higher correlation with human judgment and higher inter-rater agreement compared to the ROUGE metric.
Naman Bansal, Mousumi Akter 0001, Shubhra Kanti Karmaker Santu
EMNLP2
2022 Learning to Generate Overlap Summaries through Noisy Synthetic Data
abstract
Semantic Overlap Summarization (SOS) is a novel and relatively under-explored seq-to-seq task which entails summarizing common information from multiple alternate narratives.One of the major challenges for solving this task is the lack of existing datasets for supervised training.To address this challenge, we propose a novel data augmentation technique, which allows us to create large amount of synthetic data for training a seq-to-seq model that can perform the SOS task.Through extensive experiments using narratives from the news domain, we show that the models finetuned using the synthetic dataset provide significant performance improvements over the pre-trained vanilla summarization techniques and are close to the models fine-tuned on the golden training data; which essentially demonstrates the effectiveness of out proposed data augmentation technique for training seq-to-seq models on the SOS task.
Naman Bansal, Mousumi Akter 0001, Shubhra Kanti Karmaker Santu
EMNLP2
2017 Computing Aggregates Over Numeric Data with Personalized Local Differential Privacy
Mousumi Akter 0001, Tanzima Hashem
ACISP (2)1