Madhusudan Basak

dblp:129/5417 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
7since 2021 · last 2025
0009-0001-6762-9968ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Demo: FlyEnJoy: An Offline-First Mobile System for Trip Planning and In-Flight Social Interaction
abstract
Air travel often lacks cohesive digital tools to support both pre-trip organization and engaging in-flight experiences. We present FlyEnJoy, a mobile application with an offline-first architecture that integrates itinerary planning management, and in-flight social features to improve travel satisfaction. The system highlights local peer-to-peer networking for realtime in-flight interaction without the Internet and context-aware smart planning tools that adapt to each trip. This demo paper outlines the design and technical components of FlyEnJoy, including its offline-first design, peer-to-peer networking and context-aware smart planning approach. We also discuss insights from a traveler survey that motivated the need for such a platform and conclude with a demonstration that showcases the capabilities of FlyEnJoy in a live scenario.
Tanmoy Sen, Madhusudan Basak, Himel Dev, Michael Linden
MobiSys2
2024 Characterizing Information Seeking Events in Health-Related Social Discourse
abstract
Social media sites have become a popular platform for individuals to seek and share health information. Despite the progress in natural language processing for social media mining, a gap remains in analyzing health-related texts on social discourse in the context of events. Event-driven analysis can offer insights into different facets of healthcare at an individual and collective level, including treatment options, misconceptions, knowledge gaps, etc. This paper presents a paradigm to characterize health-related information-seeking in social discourse through the lens of events. Events here are board categories defined with domain experts that capture the trajectory of the treatment/medication. To illustrate the value of this approach, we analyze Reddit posts regarding medications for Opioid Use Disorder (OUD), a critical global health concern. To the best of our knowledge, this is the first attempt to define event categories for characterizing information-seeking in OUD social discourse. Guided by domain experts, we develop TREAT-ISE, a novel multilabel treatment information-seeking event dataset to analyze online discourse on an event-based framework. This dataset contains Reddit posts on information-seeking events related to recovery from OUD, where each post is annotated based on the type of events. We also establish a strong performance benchmark (77.4% F1 score) for the task by employing several machine learning and deep learning classifiers. Finally, we thoroughly investigate the performance and errors of ChatGPT on this task, providing valuable insights into the LLM's capabilities and ongoing characterization efforts.
Omar Sharif, Madhusudan Basak, Tanzia Parvin, Ava Scharfstein, Alphonso Bradham, Jacob T. Borodovsky, Sarah E. Lord, Sarah Masud Preum
AAAI2
2024 Explicit, Implicit, and Scattered: Revisiting Event Extraction to Capture Complex Arguments
abstract
Prior works formulate the extraction of eventspecific arguments as a span extraction problem, where event arguments are explicit -i.e.assumed to be contiguous spans of text in a document.In this study, we revisit this definition of Event Extraction (EE) by introducing two key argument types that cannot be modeled by existing EE frameworks.First, implicit arguments are event arguments which are not explicitly mentioned in the text, but can be inferred through context.Second, scattered arguments are event arguments that are composed of information scattered throughout the text.These two argument types are crucial to elicit the full breadth of information required for proper event modeling.To support the extraction of explicit, implicit, and scattered arguments, we develop a novel dataset, DiscourseEE, which includes 7,464 argument annotations from online health discourse.Notably, 51.2% of the arguments are implicit, and 17.4% are scattered, making Dis-courseEE a unique corpus for complex event extraction.Additionally, we formulate argument extraction as a text generation problem to facilitate the extraction of complex argument types.We provide a comprehensive evaluation of state-of-the-art models and highlight critical open challenges in generative event extraction.Our data and codebase are available at https://omar-sharif03.github.io/DiscourseEE.
Omar Sharif, Joseph Gatto, Madhusudan Basak, Sarah Masud Preum
EMNLP3
2024 Theme-Driven Keyphrase Extraction to Analyze Social Media Discourse
abstract
Social media platforms are vital resources for sharing self-reported health experiences, offering rich data on various health topics. Despite advancements in Natural Language Processing (NLP) enabling large-scale social media data analysis, a gap remains in applying keyphrase extraction to health-related content. Keyphrase extraction is used to identify salient concepts in social media discourse without being constrained by predefined entity classes. This paper introduces a theme-driven keyphrase extraction framework tailored for social media, a pioneering approach designed to capture clinically relevant keyphrases from user-generated health texts. Themes are defined as broad categories determined by the objectives of the extraction task. We formulate this novel task of theme-driven keyphrase extraction and demonstrate its potential for efficiently mining social media text for the use case of treatment for opioid use disorder. This paper leverages qualitative and quantitative analysis to demonstrate the feasibility of extracting actionable insights from social media data and efficiently extracting keyphrases using minimally supervised NLP models. Our contributions include the development of a novel data collection and curation framework for theme-driven keyphrase extraction and the creation of SuboxoPhrase, the first dataset of its kind comprising human-annotated keyphrases from a Reddit community. We also identify the scope of minimally supervised NLP models to extract keyphrases from social media data efficiently. Lastly, we found that a large language model (ChatGPT) outperforms unsupervised keyphrase extraction models, showcasing its efficacy in this task.
William Romano, Omar Sharif, Madhusudan Basak, Joseph Gatto, Sarah Masud Preum
ICWSM3
2023 Scope of Pre-trained Language Models for Detecting Conflicting Health Information
abstract
An increasing number of people now rely on online platforms to meet their health information needs. Thus identifying inconsistent or conflicting textual health information has become a safety-critical task. Health advice data poses a unique challenge where information that is accurate in the context of one diagnosis can be conflicting in the context of another. For example, people suffering from diabetes and hypertension often receive conflicting health advice on diet. This motivates the need for technologies which can provide contextualized, user-specific health advice. A crucial step towards contextualized advice is the ability to compare health advice statements and detect if and how they are conflicting. This is the task of health conflict detection (HCD). Given two pieces of health advice, the goal of HCD is to detect and categorize the type of conflict. It is a challenging task, as (i) automatically identifying and categorizing conflicts requires a deeper understanding of the semantics of the text, and (ii) the amount of available data is quite limited. In this study, we are the first to explore HCD in the context of pre-trained language models. We find that DeBERTa-v3 performs best with a mean F1 score of 0.68 across all experiments. We additionally investigate the challenges posed by different conflict types and how synthetic data improves a model's understanding of conflict-specific semantics. Finally, we highlight the difficulty in collecting real health conflicts and propose a human-in-the-loop synthetic data augmentation approach to expand existing HCD datasets. Our HCD training dataset is over 2x bigger than the existing HCD dataset and is made publicly available on Github.
Joseph Gatto, Madhusudan Basak, Sarah Masud Preum
ICWSM2
2023 HealthE: Recognizing Health Advice & Entities in Online Health Communities
abstract
The task of extracting and classifying entities is at the core of important Health-NLP systems such as misinformation detection, medical dialogue modeling, and patient-centric information tools. Granular knowledge of textual entities allows these systems to utilize knowledge bases, retrieve relevant information, and build graphical representations of texts. Unfortunately, most existing works on health entity recognition are trained on clinical notes, which are both lexically and semantically different from public health information found in online health resources or social media. In other words, existing health entity recognizers vastly under-represent the entities relevant to public health data, such as those provided by sites like WebMD. It is crucial that future Health-NLP systems be able to model such information, as people rely on online health advice for personal health management and clinically relevant decision making. In this work, we release a new annotated dataset, HealthE, which facilitates the large-scale analysis of online textual health advice. HealthE consists of 3,400 health advice statements with token-level entity annotations. Additionally, we release 2,256 health statements which are not health advice to facilitate health advice mining. HealthE is the first dataset with an entity-recognition label space designed for the modeling of online health advice. We motivate the need for HealthE by demonstrating the limitations of five widely-used health entity recognizers on HealthE, such as those offered by Google and Amazon. We additionally benchmark three pre-trained language models on our dataset as reference for future research. All data is made publicly available.
Joseph Gatto, Parker Seegmiller, Garrett Johnston, Madhusudan Basak, Sarah Masud Preum
ICWSM4
2022 Online Detection of Attentiveness of Students with Special Needs
abstract
In this COVID-19 pandemic era, students with Autism Spectrum Disorder (ASD) are struggling to adapt to classes in the online environment using Google Meet or Zoom. Failing to keep sustained attention in the class is a common problem for students with ASD. In face-to-face classes, teachers can track a student's behavior and activity to infer the student's attentiveness level and act accordingly. However, it becomes difficult for a teacher to monitor the attentiveness level of multiple students simultaneously on online platforms like Zoom. Detecting the attentiveness level of a student and notifying the teacher in an automated way can play a crucial role in improving the learning outcome. In this paper, we propose the first deep learning based attentiveness level prediction technique for students with ASD. Our model detects the behavior (e.g., unusual movement, gaze etc.) and activities from real-time videos and uses them as features to classify the attentiveness level as low, mid and high. Existing state-of-the-art techniques to detect the attentiveness level of typically developed students using gaze or facial expression cannot be trivially extended for students with ASD as they do not exhibit regular and consistent behavior. We collect video data belonging to different classes covering various types of activities over a long period, train our classifier, and run extensive experiments to validate the prediction performance. Our solution outperforms existing baselines by a large margin.
Khandker Aftarul Islam, Tanzima Hashem, Mohammed Eunus Ali, Tasin Ishmam, Aniruddha Ganguly, Madhusudan Basak, Sajida Rahman Danny
ACII6
2020 Not Low-Resource Anymore: Aligner Ensembling, Batch Filtering, and New Datasets for Bengali-English Machine Translation
abstract
Tahmid Hasan, Abhik Bhattacharjee, Kazi Samin, Masum Hasan, Madhusudan Basak, M. Sohel Rahman, Rifat Shahriyar. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Tahmid Hasan, Abhik Bhattacharjee, Kazi Samin, Masum Hasan, Madhusudan Basak, Mohammad Sohel Rahman, Rifat Shahriyar
EMNLP (1)5