Joseph Gatto

dblp:254/2754 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0001-7013-2445ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response Drafting
abstract
Parker Seegmiller, Joseph Gatto, Sarah E. Greer, Ganza Belise Isingizwe, Rohan Ray, Timothy E. Burdick, Sarah Masud Preum. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Parker Seegmiller, Joseph Gatto, Sarah E. Greer, Ganza Belise Isingizwe, Rohan Ray, Timothy E. Burdick, Sarah Masud Preum
ACL (1)2
2025 Follow-up Question Generation For Enhanced Patient-Provider Conversations
abstract
Joseph Gatto, Parker Seegmiller, Timothy E. Burdick, Inas S. Khayal, Sarah DeLozier, Sarah Masud Preum. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Joseph Gatto, Parker Seegmiller, Timothy E. Burdick, Inas Khayal, Sarah DeLozier, Sarah Masud Preum
ACL (1)1
2025 Document-Level Event-Argument Data Augmentation for Challenging Role Types
abstract
Event Argument Extraction (EAE) is a daunting information extraction problem -with significant limitations in few-shot cross-domain (FSCD) settings.A common solution to FSCD modeling is data augmentation.Unfortunately, existing augmentation methods are not wellsuited to a variety of real-world EAE contexts, including (i) modeling long documents (documents with over 10 sentences), and (ii) modeling challenging role types (i.e., event roles with little to no training data and semantically outlying roles).We introduce two novel LLMpowered data augmentation methods for generating extractive document-level EAE samples using zero in-domain training data.We validate the generalizability of our approach on four datasets -showing significant performance increases in low-resource settings.Our highest performing models provide a 13-pt increase in F1 score on zero-shot role extraction in FSCD evaluation.
Joseph Gatto, Omar Sharif, Parker Seegmiller, Sarah Masud Preum
ACL (1)1
2024 Explicit, Implicit, and Scattered: Revisiting Event Extraction to Capture Complex Arguments
abstract
Prior works formulate the extraction of eventspecific arguments as a span extraction problem, where event arguments are explicit -i.e.assumed to be contiguous spans of text in a document.In this study, we revisit this definition of Event Extraction (EE) by introducing two key argument types that cannot be modeled by existing EE frameworks.First, implicit arguments are event arguments which are not explicitly mentioned in the text, but can be inferred through context.Second, scattered arguments are event arguments that are composed of information scattered throughout the text.These two argument types are crucial to elicit the full breadth of information required for proper event modeling.To support the extraction of explicit, implicit, and scattered arguments, we develop a novel dataset, DiscourseEE, which includes 7,464 argument annotations from online health discourse.Notably, 51.2% of the arguments are implicit, and 17.4% are scattered, making Dis-courseEE a unique corpus for complex event extraction.Additionally, we formulate argument extraction as a text generation problem to facilitate the extraction of complex argument types.We provide a comprehensive evaluation of state-of-the-art models and highlight critical open challenges in generative event extraction.Our data and codebase are available at https://omar-sharif03.github.io/DiscourseEE.
Omar Sharif, Joseph Gatto, Madhusudan Basak, Sarah Masud Preum
EMNLP2
2024 Theme-Driven Keyphrase Extraction to Analyze Social Media Discourse
abstract
Social media platforms are vital resources for sharing self-reported health experiences, offering rich data on various health topics. Despite advancements in Natural Language Processing (NLP) enabling large-scale social media data analysis, a gap remains in applying keyphrase extraction to health-related content. Keyphrase extraction is used to identify salient concepts in social media discourse without being constrained by predefined entity classes. This paper introduces a theme-driven keyphrase extraction framework tailored for social media, a pioneering approach designed to capture clinically relevant keyphrases from user-generated health texts. Themes are defined as broad categories determined by the objectives of the extraction task. We formulate this novel task of theme-driven keyphrase extraction and demonstrate its potential for efficiently mining social media text for the use case of treatment for opioid use disorder. This paper leverages qualitative and quantitative analysis to demonstrate the feasibility of extracting actionable insights from social media data and efficiently extracting keyphrases using minimally supervised NLP models. Our contributions include the development of a novel data collection and curation framework for theme-driven keyphrase extraction and the creation of SuboxoPhrase, the first dataset of its kind comprising human-annotated keyphrases from a Reddit community. We also identify the scope of minimally supervised NLP models to extract keyphrases from social media data efficiently. Lastly, we found that a large language model (ChatGPT) outperforms unsupervised keyphrase extraction models, showcasing its efficacy in this task.
William Romano, Omar Sharif, Madhusudan Basak, Joseph Gatto, Sarah Masud Preum
ICWSM4
2023 Scope of Pre-trained Language Models for Detecting Conflicting Health Information
abstract
An increasing number of people now rely on online platforms to meet their health information needs. Thus identifying inconsistent or conflicting textual health information has become a safety-critical task. Health advice data poses a unique challenge where information that is accurate in the context of one diagnosis can be conflicting in the context of another. For example, people suffering from diabetes and hypertension often receive conflicting health advice on diet. This motivates the need for technologies which can provide contextualized, user-specific health advice. A crucial step towards contextualized advice is the ability to compare health advice statements and detect if and how they are conflicting. This is the task of health conflict detection (HCD). Given two pieces of health advice, the goal of HCD is to detect and categorize the type of conflict. It is a challenging task, as (i) automatically identifying and categorizing conflicts requires a deeper understanding of the semantics of the text, and (ii) the amount of available data is quite limited. In this study, we are the first to explore HCD in the context of pre-trained language models. We find that DeBERTa-v3 performs best with a mean F1 score of 0.68 across all experiments. We additionally investigate the challenges posed by different conflict types and how synthetic data improves a model's understanding of conflict-specific semantics. Finally, we highlight the difficulty in collecting real health conflicts and propose a human-in-the-loop synthetic data augmentation approach to expand existing HCD datasets. Our HCD training dataset is over 2x bigger than the existing HCD dataset and is made publicly available on Github.
Joseph Gatto, Madhusudan Basak, Sarah Masud Preum
ICWSM1
2023 HealthE: Recognizing Health Advice & Entities in Online Health Communities
abstract
The task of extracting and classifying entities is at the core of important Health-NLP systems such as misinformation detection, medical dialogue modeling, and patient-centric information tools. Granular knowledge of textual entities allows these systems to utilize knowledge bases, retrieve relevant information, and build graphical representations of texts. Unfortunately, most existing works on health entity recognition are trained on clinical notes, which are both lexically and semantically different from public health information found in online health resources or social media. In other words, existing health entity recognizers vastly under-represent the entities relevant to public health data, such as those provided by sites like WebMD. It is crucial that future Health-NLP systems be able to model such information, as people rely on online health advice for personal health management and clinically relevant decision making. In this work, we release a new annotated dataset, HealthE, which facilitates the large-scale analysis of online textual health advice. HealthE consists of 3,400 health advice statements with token-level entity annotations. Additionally, we release 2,256 health statements which are not health advice to facilitate health advice mining. HealthE is the first dataset with an entity-recognition label space designed for the modeling of online health advice. We motivate the need for HealthE by demonstrating the limitations of five widely-used health entity recognizers on HealthE, such as those offered by Google and Amazon. We additionally benchmark three pre-trained language models on our dataset as reference for future research. All data is made publicly available.
Joseph Gatto, Parker Seegmiller, Garrett Johnston, Madhusudan Basak, Sarah Masud Preum
ICWSM1