Sarmistha Das 0001

dblp:310/6778-1 · DBLP profile ↗
← Back
11ranked-venue papers
10as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Talk, Snap, Complain: Validation-Aware Multimodal Expert Framework for Fine-Grained Customer Grievances
abstract
Existing approaches to complaint analysis largely rely on unimodal, short-form content such as tweets or product reviews. This work advances the field by leveraging multimodal, multi-turn customer support dialogues—where users often share both textual complaints and visual evidence (e.g., screenshots, product photos)—to enable fine-grained classification of complaint aspects and severity. We introduce VALOR, a Validation-Aware Learner with Expert Routing, tailored for this multimodal setting. It employs a multi-expert reasoning setup using large-scale generative models with Chain-of-Thought (CoT) prompting for nuanced decision-making. To ensure coherence between modalities, a semantic alignment score is computed and integrated into the final classification through a meta-fusion strategy. In alignment with the United Nations Sustainable Development Goals (UN SDGs), the proposed framework supports SDG 9 (Industry, Innovation and Infrastructure) by advancing AI-driven tools for robust, scalable, and context-aware service infrastructure. Further, by enabling structured analysis of complaint narratives and visual context, it contributes to SDG 12 (Responsible Consumption and Production) by promoting more responsive product design and improved accountability in consumer services. We evaluate VALOR on a curated multimodal complaint dataset annotated with fine-grained aspect and severity labels, showing that it consistently outperforms baseline models, especially in complex complaint scenarios where information is distributed across text and images. This study underscores the value of multimodal interaction and expert validation in practical complaint understanding systems.
Rishu Kumar Singh, Navneet Shreya, Sarmistha Das 0001, Apoorva Singh, Sriparna Saha 0001
AAAI3
2026 ExpertMix: Aspect and Severity Detection in Conversational Complaints
Sarmistha Das 0001, Apoorva Singh, Rishu Kumar Singh, Navneet Shreya, Sriparna Saha 0001
ECIR (1)1
2026 Let's Decipher the Origin: Towards Multimodal Aspect-Based Explainable Complaints in Finance
abstract
Financial grievances on social media increasingly feature multimodal content, combining text and images to express complex concerns. Previous text-only approaches often struggle to capture mixed emotional phenomena within posts, such as praise for credit card services alongside dissatisfaction with branch wait times. Prior research has relied on binary classifications, overlooking root aspect-specific complaints (e.g., categorizing card services as noncomplaint while branch services are complaint-related). To address the scarcity of multimodal data resources for this task, we developed the financial multimodal complaint (FMC) dataset, comprising over 6,000 tweets across nine financial domains, capturing cross-modal complaint dynamics. To demonstrate the practical use of the FMC dataset, we designed F-ACE, a multimodal, aspect-based, explainable complaint identification model that integrates BLIP-2, Sentence Transformer, and ResNet-152 via an intermodal fusion (IMF) framework, followed by a BART-Large-based causal extractor for nuanced aspect classification. Furthermore, we evaluate the proposed model’s comprehension and adaptability through the utilization of various domain-specific and cross-domain datasets alongside state-of-the-art models. Our contribution sets a preferable alternative for advancing research in aspect-aware, multimodal complaint analysis, enabling a deeper understanding of financial grievances in the digital age.
Sarmistha Das 0001, Anurag Deo, Harsha Vardhan Dasari, Sriparna Saha 0001, Alka Maurya
IEEE Trans. Comput. Soc. Syst.1
2025 When Words Can't Capture It All: Towards Video-Based User Complaint Text Generation with Multimodal Video Complaint Dataset
abstract
While there exists a lot of work on explainable complaint mining, articulating user concerns through text or video remains a significant challenge, often leaving issues unresolved. Users frequently struggle to express their complaints clearly in text but can easily upload videos depicting product defects (e.g., vague text such as 'worst product' paired with a 5-second video depicting a broken headphone with the right earcup). This paper formulates a new task in the field of complaint mining to aid the common users' need to write an expressive complaint, which is Complaint Description from Videos (CoD-V) (e.g., to help the above user articulate her complaint about the defective right earcup). To this end, we introduce ComVID, a video complaint dataset containing 1,175 complaint videos and the corresponding descriptions, also annotated with the emotional state of the complainer. Additionally, we present a new complaint retention (CR) evaluation metric that discriminates proposed (CoD-V) task against standard video summary generation and description task. To strengthen this initiative, we introduce a multimodal Retrieval-Augmented Generation (RAG) embedded VideoLLaMA2-7b model, designed to generate complaints while accounting for the user's emotional state. We conduct a comprehensive evaluation of several Video Language Models on several tasks (pre-trained and fine-tuned versions) with a range of established evaluation metrics, including METEOR, perplexity, and the Coleman-Liau readability score, among others. Our study lays the foundation for a new research direction to provide a platform for users to express complaints through video. Dataset and resources are available at: https://github.com/sarmistha-D/CoD-V.
Sarmistha Das 0001, R. E. Zera Marveen Lyngkhoi, Kirtan Jain, Vinayak Goyal, Sriparna Saha 0001, Manish Gupta 0001
CIKM1
2025 Fin-Ally: Pioneering the Development of an Advanced, Commonsense-Embedded Conversational AI for Money Matters
abstract
The exponential technological breakthrough of the FinTech industry has significantly enhanced user engagement through sophisticated advisory chatbots. However, large-scale fine-tuning of LLMs can occasionally yield unprofessional or flippant remarks, such as “With that money, you’re going to change the world,” which, though factually correct, can be contextually inappropriate and erode user trust. The scarcity of domain-specific datasets has led previous studies to focus on isolated components, such as reasoning-aware frameworks or the enhancement of human-like response generation. To address this research gap, we present Fin-Solution 2.O, an advanced solution that 1) introduces the multi-turn financial conversational dataset, Fin-Vault, and 2) incorporates a unified model, Fin-Ally, which integrates commonsense reasoning, politeness, and human-like conversational dynamics. Fin-Ally is powered by COMET-BART-embedded commonsense context and optimized with a Direct Preference Optimization (DPO) mechanism to generate human-aligned responses. The novel Fin-Vault dataset, consisting of 1,417 annotated multi-turn dialogues, enables Fin-Ally to extend beyond basic account management to provide personalized budgeting, real-time expense tracking, and automated financial planning. Our comprehensive results demonstrate that incorporating commonsense context enables language models to generate more refined, textually precise, and professionally grounded financial guidance, positioning this approach as a next-generation AI solution for the FinTech sector.
Sarmistha Das 0001, Priya Mathur, Ishani Sharma, Sriparna Saha 0001, Kitsuchart Pasupa, Alka Maurya
ECAI1
2025 Reasoning and Planning for Multimodal Large Language Models: A Multilingual and Cross-Domain Exploration
abstract
Recent advancements in Multimodal Large Language Models (MLLMs), coupled with the progress of reinforcement learning, have substantially enhanced reasoning and decision-making across modalities, including text, vision, audio, and video. This tutorial introduces the fundamental principles, methodologies, and practical applications of MLLM reasoning, with a particular emphasis on strengthening reasoning capabilities in multilingual and cross-domain settings. We further discuss the key challenges and limitations of current multimodal reasoning approaches, as well as future directions for advancing the field. By highlighting how MLLMs support enhanced reasoning and planning in cross-lingual and cross-domain contexts, this session aims to equip researchers and practitioners with the conceptual foundations and practical tools needed to effectively integrate MLLM reasoning into their work.
Sarmistha Das 0001, Akash Ghosh, Sriparna Saha 0001, Koustava Goswami, K. J. Joseph
ACM Multimedia1
2025 Unlocking Financial Insights: An advanced Multimodal Summarization with Multimodal Output Framework for Financial Advisory Videos
abstract
The dynamic propagation of social media has broadened the reach of financial advisory content through podcast videos, yet extracting insights from lengthy, multimodal segments (30-40 minutes) remains challenging. We introduce FASTER(Financial Advisory Summariser with Textual Embedded Relevant images), a modular framework that tackles three key challenges: (1) extracting modality-specific features, (2) producing optimized, concise summaries, and (3) aligning visual keyframes with associated textual points. FASTER employs BLIP-2 for semantic visual descriptions, OCR for textual patterns, and Whisper-based transcription with Speaker diarization as BOS features. A modified Direct Preference Optimization (DPO)-based loss function, equipped with BOS-specific fact-checking, ensures precision, relevance, and factual consistency against the human-aligned summary. A ranker-based retrieval mechanism further aligns keyframes with summarized content, enhancing interpretability and cross-modal coherence. To acknowledge data resource scarcity, we introduce Fin-APT, a dataset comprising 470 publicly accessible financial advisory pep-talk videos for robust multimodal research. Comprehensive cross-domain experiments confirm FASTER's strong performance, robustness, and generalizability when compared to Large Language Models (LLMs) and Vision-Language Models (VLMs). By establishing a new standard for multimodal summarization, FASTER makes financial advisory content more accessible and actionable, thereby opening new avenues for research.
Sarmistha Das 0001, R. E. Zera Marveen Lyngkhoi, Sriparna Saha 0001, Alka Maurya
ACM Multimedia1
2025 Deciphering the Complaint Aspects: Towards an Aspect-Based Complaint Identification Model with Video Complaint Dataset in Finance
abstract
In today's competitive marketing landscape, effective complaint management is crucial for customer service and business success. Video complaints, integrating text and image content, offer invaluable insights by addressing customer grievances and delineating product benefits and drawbacks. However, comprehending nuanced complaint aspects within vast daily multimodal financial data remains a formidable challenge. Addressing this gap, we have curated a proprietary multimodal video complaint dataset comprising 433 publicly accessible instances. Each instance is meticulously annotated at the utterance level, encompassing five distinct categories of financial aspects and their associated complaint labels. To support this en-deavour, we introduce Solution 3.0, a model designed for multimodal aspect-based complaint identification task. So-lution 3.0 is tailored to perform three key tasks: 1) handling multimodal features (audio and video), 2) facilitating mul-tilabel aspect classification, and 3) conducting multitasking for aspect classifications and complaint identification parallelly. Solution 3.0 utilizes a CLIP-based dual frozen encoder with an integrated image segment encoder for global feature fusion, enhanced by contextual attention (ISEC) to improve accuracy and efficiency. Our proposed framework surpasses current multimodal baselines, exhibiting superior performance across nearly all metrics by opening new ways to strengthen appropriate customer care initiatives and effectively assisting individuals in resolving their problems.
Sarmistha Das 0001, Basha Mujavarsheik, R. E. Zera Marveen Lyngkhoi, Sriparna Saha 0001, Alka Maurya
WACV1
2024 Negative Review or Complaint? Exploring Interpretability in Financial Complaints
abstract
In the financial service sector, customer service is the most critical tool for long-term business growth. A financial complaint detection (CD) system could aid in the identification of shortcomings in product features and service delivery. This could further ensure faster resolution of customer complaints and thereby help retain existing clients and attract new ones. Prior research has prioritized only complaint identification and prediction of the corresponding severity levels; the first aim is to categorize a textual element as a complaint or a noncompliant. The other attempts to classify complaints into several severity levels based on the degree of risk the complainant is willing to endure. Identifying the reason or source of a complaint in a text is a significant but underexplored area in natural language processing study. We propose an explainable complaint cause identification approach with a dyadic attention mechanism at the sentence and word levels, enabling it to give varying amounts of emphasis to more and less important information. As the first subtask, the model simultaneously trains CD, sentiment detection, and emotion recognition tasks. Afterwards, we identify the complaint’s cause and its severity level. To do this, the causal span annotations for complaint tweets are added to an existing financial complaints corpus. The findings suggest that conventional computing techniques can be adapted to solve extremely relevant new problems, generating novel opportunities for research1.
Sarmistha Das 0001, Apoorva Singh, Sriparna Saha 0001, Alka Maurya
IEEE Trans. Comput. Soc. Syst.1
2023 "Find the Table": A Contrastive Learning-based Approach with Faster RCNN for Establishing Tabular Entity Relationships
abstract
Financial industries rely on a variety of data to understand an organization's financial health and performance, and analysis of financial statements is used for decision-making and understanding business activities. Financial statements often have complex, unstructured formats, making it difficult to extract useful information for decision-making. Organizing and refining these data is crucial for effective analysis. Previous research in the field of table detection has primarily centred around object detection methods, with limited exploration of methods for cell-wise information extraction by identifying row and column entities. In order to address this research gap, this paper proposes a novel dataset called TERED (Tabular Entity Relationship Establishment Dataset) to train a model to identify relationships among elements in large financial tables, such as statements and balance sheets, using computer vision techniques which have yet to be fully explored in financial analysis. The dataset contains more than 10,000 tabular data in scanned image and pdf formats and is divided into 12 classes. We trained CLF-RCNN, a Contrastive Learning based Faster RCNN model which in turn is a state-of-the-art object detection model on this dataset and achieved an F1 score of 93 % for table detection and 74% for identifying tabular entities and relationships. Additionally, we introduced a new loss term influenced by contrastive learning that improves prediction performances with our developed algorithm to return the sequential order of the unordered predictions.
Sarmistha Das 0001, Tuhinangshu Gangopadhyay, Atulya Deep, Sriparna Saha 0001, Alka Maurya
IJCNN1
2023 Let the Model Make Financial Senses: A Text2Text Generative Approach for Financial Complaint Identification
Sarmistha Das 0001, Apoorva Singh, Raghav Jain, Sriparna Saha 0001, Alka Maurya
PAKDD (3)1