EDBT 2026 Demo / reviewers in the wild / expert
Abhisek Tiwari
dblp:302/7809
· DBLP profile ↗
16ranked-venue papers
10as first author
16since 2021 · last 2024
0000-0003-3460-1624ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 8 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | From Sights to Insights: Towards Summarization of Multimodal Clinical DocumentsabstractThe advancement of Artificial Intelligence is pivotal in reshaping healthcare, enhancing diagnostic precision, and facilitating personalized treatment strategies.One major challenge for healthcare professionals is quickly navigating through long clinical documents to provide timely and effective solutions.Doctors often struggle to draw quick conclusions from these extensive documents.To address this issue and save time for healthcare professionals, an effective summarization model is essential.Most current models assume the data is only textbased.However, patients often include images of their medical conditions in clinical documents.To effectively summarize these multimodal documents, we introduce EDI-Summ, an innovative Image-Guided Encoder-Decoder Model.This model uses modality-aware contextual attention on the encoder and an image cross-attention mechanism on the decoder, enhancing the BART base model to create detailed visual-guided summaries.We have tested our model extensively on three multimodal clinical benchmarks involving multimodal question and dialogue summarization tasks.Our analysis demonstrates that EDI-Summ outperforms state-of-the-art large language and vision-aware models in these summarization tasks. Akash Ghosh, Mohit Tomar, Abhisek Tiwari, Sriparna Saha 0001, Jatin Salve, Setu Sinha |
ACL (1) | 3 |
| 2024 | Seeing Is Believing! towards Knowledge-Infused Multi-modal Medical Dialogue GenerationabstractOver the last few years, artificial intelligence-based clinical assistance has gained immense popularity and demand in telemedicine, including automatic disease diagnosis. Patients often describe their signs and symptoms to doctors using visual aids, which provide vital evidence for identifying a medical condition. In addition to learning from our experiences, we learn from well-established theories/ knowledge. With the motivation of leveraging visual cues and medical knowledge, we propose a transformer-based, knowledge-infused multi-modal medical dialogue generation (KI-MMDG) framework. In addition, we present a discourse-aware image identifier (DII) that recognizes signs and their severity by leveraging the current conversation context in addition to the image of the signs. We first curate an empathy and severity-aware multi-modal medical dialogue (ES-MMD) corpus in English, which is annotated with intent, symptoms, and visual signs with severity information. Experimental results show the superior performance of the proposed KI-MMDG model over uni-modal and non-knowledge infused generative models, demonstrating the importance of visual signs and knowledge infusion in symptom investigation and diagnosis. We also observed that the DII model surpasses the existing state-of-the-art model by 7.84%, indicating the crucial significance of dialogue context for identifying a sign image surfaced during conversations. The code and dataset are available at https://github.com/NLP-RL/KI-MMDG. Abhisek Tiwari, Shreyangshu Bera, Preeti Verma, Jaithra Varma Manthena, Sriparna Saha 0001, Pushpak Bhattacharyya, Minakshi Dhar, Sarbajeet Tiwari |
LREC/COLING | 1 |
| 2024 | Action and Reaction Go Hand in Hand! a Multi-modal Dialogue Act Aided Sarcasm IdentificationabstractSarcasm primarily involves saying something but “meaning the opposite” or “meaning something completely different” in order to convey a particular tone or mood. In both the above cases, the “meaning” is reflected by the communicative intention of the speaker, known as dialogue acts. In this paper, we seek to investigate a novel phenomenon of analyzing sarcasm in the context of dialogue acts with the hypothesis that the latter helps to understand the former better. Toward this aim, we extend the multi-modal MUStARD dataset to enclose dialogue acts for each dialogue. To demonstrate the utility of our hypothesis, we develop a dialogue act-aided multi-modal transformer network for sarcasm identification (MM-SARDAC), leveraging interrelation between these tasks. In addition, we introduce an order-infused, multi-modal infusion mechanism into our proposed model, which allows for a more intuitive combined modality representation by selectively focusing on relevant modalities in an ordered manner. Extensive empirical results indicate that dialogue act-aided sarcasm identification achieved better performance compared to performing sarcasm identification alone. The dataset and code are available at https://github.com/mohit2b/MM-SARDAC. Mohit Tomar, Tulika Saha, Abhisek Tiwari, Sriparna Saha 0001 |
LREC/COLING | 3 |
| 2024 | Yes, This Is What I Was Looking For! Towards Multi-modal Medical Consultation Concern Summary Generation
Abhisek Tiwari, Shreyangshu Bera, Sriparna Saha 0001, Pushpak Bhattacharyya, Samrat Ghosh |
ECIR (3) | 1 |
| 2024 | An EcoSage Assistant: Towards Building A Multimodal Plant Care Dialogue Assistant
Mohit Tomar, Abhisek Tiwari, Tulika Saha, Prince Jha, Sriparna Saha 0001 |
ECIR (2) | 2 |
| 2023 | Experience and Evidence are the eyes of an excellent summarizer! Towards Knowledge Infused Multi-modal Clinical Conversation SummarizationabstractWith the advancement of telemedicine, both researchers and medical practitioners are working hand-in-hand to develop various techniques to automate various medical operations, such as diagnosis report generation. In this paper, we first present a multi-modal clinical conversation summary generation task that takes a clinician-patient interaction (both textual and visual information) and generates a succinct synopsis of the conversation. We propose a knowledge-infused, multi-modal, multi-tasking medical domain identification and clinical conversation summary generation (MM-CliConSummation) framework. It leverages an adapter to infuse knowledge and visual features and unify the fused feature vector using a gated mechanism. Furthermore, we developed a multi-modal, multi-intent clinical conversation summarization corpus annotated with intent, symptom, and summary. The extensive set of experiments, both quantitatively and qualitatively, led to the following findings: (a) critical significance of visuals, (b) more precise and medical entity preserving summary with additional knowledge infusion, and (c) a correlation between medical department identification and clinical synopsis generation. Furthermore, the dataset and source code are available at https://github.com/NLP-RL/MM-CliConSummation Abhisek Tiwari, Anisha Saha, Sriparna Saha 0001, Pushpak Bhattacharyya, Minakshi Dhar |
CIKM | 1 |
| 2023 | Local context is not enough! Towards Query Semantic and Knowledge Guided Multi-Span Medical Question AnsweringabstractMedical Question Answering (MedQA) is one of the most popular and significant tasks in developing healthcare assistants. When humans extract an answer to a question from a document, they first (a) understand the question itself in detail and (b) utilize relevant knowledge/experiences to determine the answer segments. In multi-span question answering, it becomes increasingly important to comprehend the query accurately and possess relevant knowledge, as the interrelationship among different answer segments is essential for achieving completeness. Motivated by this, we first propose a transformer-based query semantic and knowledge (QueSemKnow) guided multi-span question-answering model. The proposed QueSemKnow works in a two-phased manner; in the first stage, a multi-task model is proposed to extract query semantics: (i) intent identification and (ii) question type prediction. In the second stage, QueSemKnow selects a relevant subset of the knowledge graph as the underlying context/document and extracts answers depending on the semantic information extracted from the first stage and context. We build a multi-task query semantic extraction model for query intent and query type identification to investigate the co-relation among these tasks. Furthermore, we created a semantically aware medical question-answering corpus named QueSeMSpan MedQA wherein each question is annotated with its corresponding semantic information. The proposed model outperforms several baselines and existing state-of-the-art models by a large margin on multiple datasets, which firmly demonstrates the effectiveness of the human-inspired multi-span question-answering methodology. Abhisek Tiwari, Aman Bhansali, Sriparna Saha 0001, Pushpak Bhattacharyya, Preeti Verma, Minakshi Dhar |
ECAI | 1 |
| 2023 | Your tone speaks louder than your face! Modality Order Infused Multi-modal Sarcasm DetectionabstractFigurative language is an essential component of human communication, and detecting sarcasm in text has become a challenging yet highly popular task in natural language processing. As humans, we rely on a combination of visual and auditory cues, such as facial expressions and tone of voice, to comprehend a message. Our brains are implicitly trained to integrate information from multiple senses to form a complete understanding of the message being conveyed, a process known as multi-sensory integration. The combination of different modalities not only provides additional information but also amplifies the information conveyed by each modality in relation to the others. Thus, the infusion order of different modalities also plays a significant role in multimodal processing. In this paper, we investigate the impact of different modality infusion orders for identifying sarcasm in dialogues. We propose a modality order-driven module integrated into a transformer network, MO-Sarcation that fuses modalities in an ordered manner. Our model outperforms several state-of-the-art models by 1-3% across various metrics, demonstrating the crucial role of modality order in sarcasm detection. The obtained improvements and detailed analysis show that audio tone should be infused with textual content, followed by visual information to identify sarcasm efficiently. The code and dataset are available at https://github.com/mohit2b/MO-Sarcation. Mohit Tomar, Abhisek Tiwari, Tulika Saha, Sriparna Saha 0001 |
ACM Multimedia | 2 |
| 2023 | Towards personalized persuasive dialogue generation for adversarial task oriented dialogue setting
Abhisek Tiwari, Abhijeet Khandwe, Sriparna Saha 0001, Roshni R. Ramnani, Anutosh Maitra, Shubhashis Sengupta |
Expert Syst. Appl. | 1 |
| 2022 | Dr. Can See: Towards a Multi-modal Disease Diagnosis Virtual AssistantabstractArtificial Intelligence-based clinical decision support is gaining ever-growing popularity and demand in both the research and industry communities. One such manifestation is automatic disease diagnosis, which aims to assist clinicians in conducting symptom investigations and disease diagnoses. When we consult with doctors, we often report and describe our health conditions with visual aids. Moreover, many people are unacquainted with several symptoms and medical terms, such as mouth ulcer and skin growth. Therefore, visual form of symptom reporting is a necessity. Motivated by the efficacy of visual form of symptom reporting, we propose and build a novel end-to-end Multi-modal Disease Diagnosis Virtual Assistant (MDD-VA) using reinforcement learning technique. In conversation, users' responses are heavily influenced by the ongoing dialogue context, and multi-modal responses appear to be of no difference. We also propose and incorporate a Context-aware Symptom Image Identification module that leverages discourse context in addition to the symptom image for identifying symptoms effectively. Furthermore, we first curate a multi-modal conversational medical dialogue corpus in English that is annotated with intent, symptoms, and visual information. The proposed MDD-VA outperforms multiple uni-modal baselines in both automatic and human evaluation, which firmly establishes the critical role of symptom information provided by visuals . The dataset and code are available at https://github.com/NLP-RL/DrCanSee Abhisek Tiwari, Manisimha Manthena, Sriparna Saha 0001, Pushpak Bhattacharyya, Minakshi Dhar, Sarbajeet Tiwari |
CIKM | 1 |
| 2022 | Introducing Multi-modality in Persuasive Task Oriented Virtual Sales Agent
Aritra Raut, Abhisek Tiwari, Sriparna Saha 0001, Anutosh Maitra, Roshni R. Ramnani, Shubhashis Sengupta |
ICONIP (3) | 3 |
| 2022 | Symptoms are known by their companies: towards association guided disease diagnosis assistantabstractOver the last few years, dozens of healthcare surveys have shown a shortage of doctors and an alarming doctor-population ratio. With the motivation of assisting doctors and utilizing their time efficiently, automatic disease diagnosis using artificial intelligence is experiencing an ever-growing demand and popularity. Humans are known by the company they keep; similarly, symptoms also exhibit the association property, i.e., one symptom may strongly suggest another symptom's existence/non-existence, and their association provides crucial information about the suffering condition. The work investigates the role of symptom association in symptom investigation and disease diagnosis process. We propose and build a virtual assistant called Association guided Symptom Investigation and Diagnosis Assistant (A-SIDA) using hierarchical reinforcement learning. The proposed A-SIDDA converses with patients and extracts signs and symptoms as per patients' chief complaints and ongoing dialogue context. We infused association-based recommendations and critic into the assistant, which reinforces the assistant for conducting context-aware, symptom-association guided symptom investigation. Following the symptom investigation, the assistant diagnoses a disease based on the extracted signs and symptoms. The assistant then diagnoses a disease based on the extracted signs and symptoms. In addition to diagnosis accuracy, the relevance of inspected symptoms is critical to the usefulness of a diagnosis framework. We also propose a novel evaluation metric called Investigation Relevance Score (IReS), which measures the relevance of symptoms inspected during symptom investigation. The obtained improvements (Diagnosis success rate-5.36%, Dialogue length-1.16, Match rate-2.19%, Disease classifier-6.36%, IReS-0.3501, and Human score-0.66) over state-of-the-art methods firmly establish the crucial role of symptom association that gets uncovered by the virtual agent. Furthermore, we found that the association guided symptom investigation greatly increases human satisfaction, owing to its seamless topic (symptom) transition. Abhisek Tiwari, Tulika Saha, Sriparna Saha 0001, Pushpak Bhattacharyya, Shemim Begum, Minakshi Dhar, Sarbajeet Tiwari |
BMC Bioinform. | 1 |
| 2022 | A persona aware persuasive dialogue policy for dynamic and co-operative goal setting
Abhisek Tiwari, Tulika Saha, Sriparna Saha 0001, Shubhashis Sengupta, Anutosh Maitra, Roshni R. Ramnani, Pushpak Bhattacharyya |
Expert Syst. Appl. | 1 |
| 2022 | A knowledge infused context driven dialogue agent for disease diagnosis using hierarchical reinforcement learning
Abhisek Tiwari, Sriparna Saha 0001, Pushpak Bhattacharyya |
Knowl. Based Syst. | 1 |
| 2021 | Context Aware Joint Modeling of Domain Classification, Intent Detection and Slot Filling with Zero-Shot Intent Detection Approach
Neeti Priya, Abhisek Tiwari, Sriparna Saha 0001 |
ICONIP (3) | 2 |
| 2021 | Multi-Modal Dialogue Policy Learning for Dynamic and Co-operative Goal SettingabstractDeveloping an adequate and human-like virtual agent has been one of the primary applications of artificial intelligence. In the last few years, task-oriented dialogue systems have gained huge popularity because of their upsurging relevance and positive outcomes. In real-world, users may not always have a predefined and rigid task goal beforehand; they upgrade/downgrade/change their goal component dynamically depending upon their utility value and agent's serving capability. However, existing virtual agents fail to incorporate this dynamic behavior, leading to either unsuccessful task completion or an ungratified user experience. The paper presents an end to end multimodal dialogue system for dynamic and co-operative goal setting, which incorporates i) a multi-modal semantic state representation in policy learning to deal with multi-modal inputs, ii) a goal manager module in a traditional dialogue manager for handling dynamic and goal unavailability scenarios effectively, iii) an accumulative reward (task/persona/sentiment) for task success, personalized persuasion and user-adaptive behavior, respectively. The obtained experimental results and the comparisons with baselines firmly establish the need and efficacy of the proposed system. Abhisek Tiwari, Tulika Saha, Sriparna Saha 0001, Shubhashis Sengupta, Anutosh Maitra, Roshni R. Ramnani, Pushpak Bhattacharyya |
IJCNN | 1 |