Muskaan Singh

dblp:249/4543 · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
20since 2021 · last 2026
0009-0008-2638-7000ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 3 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Prosody as Supervision: Bridging the Non-Verbal-Verbal for Multilingual Speech Emotion Recognition
abstract
In this work, we introduce a paralinguisticsupervision paradigm for low-resource multilingual speech emotion recognition (LRM-SER) that leverages non-verbal vocalizations to exploit prosody-centric emotion cues.Unlike conventional SER systems that rely heavily on labeled verbal speech and suffer from poor cross-lingual transfer, our approach reformulates LRM-SER as non-verbal-to-verbal transfer, where supervision from a labeled non-verbal source domain is adapted to unlabeled verbal speech across multiple target languages.To this end, we propose NOVA-ARC, a geometry-aware framework that models affective structure in the Poincaré ball, discretizes paralinguistic patterns via a hyperbolic vector-quantized prosody codebook, and captures emotion intensity through a hyperbolic emotion lens.For unsupervised adaptation, NOVA-ARC performs optimal-transport-based prototype alignment between source emotion prototypes and target utterances, inducing soft supervision for unlabeled speech while being stabilized through consistency regularization.Experiments show that NOVA-ARC delivers the strongest performance under both non-verbal-to-verbal adaptation and the complementary verbal-to-verbal transfer setting, consistently outperforming Euclidean counterparts and strong SSL baselines.To the best of our knowledge, this work is the first to move beyond verbal-speech-centric supervision by introducing a non-verbal-to-verbal transfer paradigm for SER.
Girish, Mohd Mujtaba Akhtar, Muskaan Singh
ACL (1)3
2026 HDLSS Raman Spectroscopy Data Generation Using GANs and Genetic Algorithms
Thomas Poudevigne-Durance, Sahil Sharma 0001, Sayantan Tripathy, Ng Ka Wai, Muskaan Singh, Liam McDaid, Gerard L. Cote, Samuel B. Mabbott, Saugat Bhattacharyya
ICPRAM5
2026 MeetMulti-X: A benchmark analysis of scaling and prompting large language models on automatic minuting
abstract
The task of automatic minuting, i.e., capturing all the points from transcripts of multi-party meetings, presents considerable challenges owing to the spontaneous and complex nature of discussions. As organisations increasingly depend on meetings for decision-making, the need for efficient and optimised minuting has intensified, underscoring the shortcomings of manual note-taking due to cognitive overload and diverted participant engagement. This study systematically analyses the impact of scaling Large Language Models (LLMs) on automatic minuting , emphasising key factors including pretrained dataset size, model size, context length, and prompt length. The benchmark evaluation includes both quantitative and qualitative with 19 open-source models (from 77M to 70B parameters) and 4 closed-source models (over 1T parameters) across 4 meeting corpora and prompts. Our findings indicate that (1) models with less than 8B parameters offer a favorable trade-off between performance and efficiency, achieving comparable results to their larger counterparts. (2) scaling pretrained data size improves performance up to a threshold, beyond which gains diminish. (3) context length exhibits a non-linear effect, with optimal performance around 8K-16K tokens. (4) longer prompts consistently degrade output quality, highlighting the need for concise and well-structured prompting. To the best of our knowledge this is the first work, exploring scaling LLMS on automatic minuting. Code is available at https://anonymous.4open.science/r/MeetMultiX-9A36
Ashima Sood, Muskaan Singh, Bryan Gardiner, Joan Condell
Expert Syst. Appl.2
2025 Towards Machine Unlearning for Paralinguistic Speech Processing
Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar, Swarup Ranjan Behera, Vandana Rajan, Muskaan Singh, Arun Balaji Buduru, Rajesh Sharma 0002
INTERSPEECH7
2024 NeuRO: an application for code-switched autism detection in children
Mohd Mujtaba Akhtar, Girish, Orchid Chetia Phukan, Muskaan Singh
INTERSPEECH4
2024 ComFeAT: combination of neural and spectral features for improved depression detection
Orchid Chetia Phukan, Muskaan Singh, Arun Balaji Buduru, Rajesh Sharma 0002
INTERSPEECH4
2023 Decoding Neural Activity for Part-of-Speech Tagging (POS)
abstract
Decoding Part of Speech(POS) tagging directly from electroencephalography (EEG) signals whilst user overtly spoke (voiced speech) sentences could improve direct speech brain-computer interfaces (BCIs) using imagined or inner speech. To the best of our knowledge, earlier work uses machine learning approach using 74,953 sentences/tokens recorded in 75 EEG sessions. The tokens can be found in 4,479 phrases consisting of terms from the English Online treebank which contains the record of weblogs, newsgroups, reviews, and Yahoo Answers. The results demonstrated the feasibility of POS decoding from EEG based on word class, word frequency, and word length with accuracy of 71%, 86%, 89%, respectively. We believe that there is significant room for improvement with more advanced artificial intelligence. In this paper, we further extend the existing work with end-to-end transformers. Our results presents transformer model outperforms benchmark traditional ML results with +20% in length, +13% for the open vs closed class and +12% in frequency. In our empirical analysis, we find the decoding performance was better when using multi-electrode recordings as compared to single-electrode recordings.
Muskaan Singh, Saugat Bhattacharyya, Damien Coyle
SMC2
2023 An Automated Detection of Amyotrophic Lateral Sclerosis from Resting-State MEG Data Using 3D Deep Convolutional Neural Network
abstract
A novel 3D deep convolutional neural network (3D-CNN) model called MEGNet3D has been proposed in the paper. MEGNet3D is designed to differentiate between amyotrophic lateral sclerosis (ALS) and healthy individuals from their resting state (eyes open and eyes closed condition) sensor-level magnetoencephalography (MEG) data. The raw MEG data is initially transformed into their time-frequency representation, which are then used as inputs to MEGNet3D. Both magnetometer and gradiometer recordings have been investigated separately. The proposed model exhibits an accuracy of over 75% for most classification conditions. Thus, MEGNet3D is capable of handling high subject variability and shows that spectral-temporal representation of resting-state MEG data yields relevant neural markers related to the existence of ALS. Furthermore, it has also been observed resting state with eyes closed yields better classification accuracy as compared to the resting state with eyes open condition.
Kaniska Samanta, Sujit Roy, Véronique Marchand-Pauvert, Shirin Dora, Stéphanie Duguez, Muskaan Singh, Girijesh Prasad, Saugat Bhattacharyya
SMC6
2022 Multimodality for NLP-Centered Applications: Resources, Advances and Frontiers
abstract
With the development of multimodal systems and natural language generation techniques, the resurgence of multimodal datasets has attracted significant research interests, which aims to provide new information to enrich the representation of textual data. However, there remains a lack of a comprehensive survey for this task. To this end, we take the first step and present a thorough review of this research field. This paper provides an overview of a publicly available dataset with different modalities according to the applications. Furthermore, we discuss the new frontier and give our thoughts. We hope this survey of multimodal datasets can provide the community with quick access and a general picture of the multimodal dataset for specific Natural Language Processing (NLP) applications and motivates future researches. In this context, we release the collection of all multimodal datasets easily accessible here: https://github.com/drmuskangarg/Multimodal-datasets
Muskan Garg, Seema Wazarkar, Muskaan Singh, Ondrej Bojar
LREC3
2022 ELITR Minuting Corpus: A Novel Dataset for Automatic Minuting from Multi-Party Meetings in English and Czech
abstract
Taking minutes is an essential component of every meeting, although the goals, style, and procedure of this activity (“minuting” for short) can vary. Minuting is a rather unstructured writing activity and is affected by who is taking the minutes and for whom the intended minutes are. With the rise of online meetings, automatic minuting would be an important benefit for the meeting participants as well as for those who might have missed the meeting. However, automatically generating meeting minutes is a challenging problem due to a variety of factors including the quality of automatic speech recorders (ASRs), availability of public meeting data, subjective knowledge of the minuter, etc. In this work, we present the first of its kind dataset on Automatic Minuting. We develop a dataset of English and Czech technical project meetings which consists of transcripts generated from ASRs, manually corrected, and minuted by several annotators. Our dataset, AutoMin, consists of 113 (English) and 53 (Czech) meetings, covering more than 160 hours of meeting content. Upon acceptance, we will publicly release (aaa.bbb.ccc) the dataset as a set of meeting transcripts and minutes, excluding the recordings for privacy reasons. A unique feature of our dataset is that most meetings are equipped with more than one minute, each created independently. Our corpus thus allows studying differences in what people find important while taking the minutes. We also provide baseline experiments for the community to explore this novel problem further. To the best of our knowledge AutoMin is probably the first resource on minuting in English and also in a language other than English (Czech).
Anna Nedoluzhko, Muskaan Singh, Marie Hledíková, Tirthankar Ghosal, Ondrej Bojar
LREC2
2022 ALIGNMEET: A Comprehensive Tool for Meeting Annotation, Alignment, and Evaluation
abstract
Summarization is a challenging problem, and even more challenging is to manually create, correct, and evaluate the summaries. The severity of the problem grows when the inputs are multi-party dialogues in a meeting setup. To facilitate the research in this area, we present ALIGNMEET, a comprehensive tool for meeting annotation, alignment, and evaluation. The tool aims to provide an efficient and clear interface for fast annotation while mitigating the risk of introducing errors. Moreover, we add an evaluation mode that enables a comprehensive quality evaluation of meeting minutes. To the best of our knowledge, there is no such tool available. We release the tool as open source. It is also directly installable from PyPI.
Peter Polak, Muskaan Singh, Anna Nedoluzhko, Ondrej Bojar
LREC2
2022 An End-to-End Multilingual System for Automatic Minuting of Multi-Party Dialogues
Aakash Bhatnagar, Nidhir Bhavsar, Muskaan Singh, Petr Motlícek
PACLIC3
2022 HMIST: Hierarchical Multilingual Isometric Speech Translation using Multi-Task Learning Framework and it's influence on Automatic Dubbing
Nidhir Bhavsar, Aakash Bhatnagar, Muskaan Singh
PACLIC3
2022 Bio-Medical Multi-label Scientific Literature Classification using LWAN and Dual-attention module
Deepanshu Khanna, Aakash Bhatnagar, Nidhir Bhavsar, Muskaan Singh, Petr Motlícek
PACLIC4
2022 Automatic Minuting: A Pipeline Method for Generating Minutes from Multi-Party Meeting Proceedings
Kartik Shinde, Tirthankar Ghosal, Muskaan Singh, Ondrej Bojar
PACLIC3
2022 An Empirical Comparison of off-the-shelve Semantic Similarity methods for down-streaming Meeting Similarity
Aditya Upadhyay, Aakash Bhatnagar, Nidhir Bhavsar, Muskaan Singh
PACLIC4
2022 DeepCon: An End-to-End Multilingual Toolkit for Automatic Minuting of Multi-Party Dialogues
abstract
In this paper, we present our minuting tool DeepCon, an end-to-end toolkit for minuting the multiparty dialogues of meetings.It provides technological support for (multilingual) communication and collaboration, with a specific focus on Natural Language Processing (NLP) technologies: Automatic Speech Recognition (ASR), Machine Translation (MT), Automatic Minuting (AM), Topic Modelling (TM) and Named Entity Recognition (NER).To the best of our knowledge, there is no such tool available.Further, this tool follows a microservice architecture, and we release the tool as open-source, deployed on Amazon Web Services (AWS).We release our tool open-source here http://www.deepcon.in.
Aakash Bhatnagar, Nidhir Bhavsar, Muskaan Singh
SIGDIAL3
2021 ARGUABLY @ AI Debater-NLPCC 2021 Task 3: Argument Pair Extraction from Peer Review and Rebuttals
Guneet Singh Kohli, Prabsimran Kaur, Muskaan Singh, Tirthankar Ghosal, Prashant Singh Rana
NLPCC (2)3
2021 An Empirical Performance Analysis of State-of-the-Art Summarization Models for Automatic Minuting
Muskaan Singh, Tirthankar Ghosal, Ondrej Bojar
PACLIC1
2021 Improving neural machine translation for low-resource Indian languages using rule-based feature extraction
Muskaan Singh, Ravinder Kumar 0002, Inderveer Chana
Neural Comput. Appl.1
2020 A forefront to machine translation technology: deployment on the cloud as a service to enhance QoS parameters
Muskaan Singh, Ravinder Kumar 0002, Inderveer Chana
Soft Comput.1