Jinho D. Choi

dblp:17/8156 · DBLP profile ↗
← Back
54ranked-venue papers
5as first author
25since 2021 · last 2026
0000-0003-2693-6934ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 49 · 5 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Towards Conversational Patient History-Taking: Voice-Interactive AI Agents for Pre-visit Dementia Diagnostic Interviews
abstract
Patient history-taking is a critical yet time-intensive component of clinical diagnosis, frequently hindered by time-constrained clinical visits. We present an LLM-based voice-interactive conversational system for conducting semi-structured diagnostic interviews with older adults suspected of Alzheimer’s disease and related dementias (ADRD). The system features conditional conversation branching over specialist-developed interview scripts and interaction adaptations tailored for older adults. In a within-subjects study with 30 participants from a cognitive neurology clinic, we compare two LLM prompting strategies through dialogue analysis, assess user experience, and evaluate symptom elicitation against routine specialist interviews. Our findings reveal that prompting strategy distinctly shapes both conversational dynamics and symptom coverage. Together with high sensitivity scores and positive user experience ratings, these results demonstrate the feasibility and clinical potential of LLM-based conversational agents for scalable, patient-centered history-taking in ADRD care.
Andrew G. Breithaupt, Nayoung Choi, James D. Finch, Jeanne M. Powell, Arin L. Nelson, Oz A. Alon, Howard J. Rosen, Jinho D. Choi
SIGDIAL8
2026 Tinker Tales: A Tangible Dialogue System for Child-AI Co-Creative Storytelling
abstract
Conversational AI agents are increasingly explored as creative partners, yet how conversation design shapes child–AI dialogue in co-creative settings remains underexplored. We present Tinker Tales, a tangible dialogue system for child–AI collaborative storytelling, in which educational frameworks—narrative development and social-emotional learning—are instantiated as conversation design, shaping how the agent engages children across four narrative stages. The system combines a physical storytelling board, NFC-embedded toys, and a mobile app mediating multimodal interaction through tangible manipulation and voice-based dialogue. We conducted a home-based user study with 10 children (ages 6–8) across two conversation design conditions varying in how the agent structured elaboration, with and without educational scaffolding. Our findings show that prompt framing shapes the form and consistency of children’s narrative contributions, structuring how they participate in co-creative dialogue with AI.
Nayoung Choi, Jiseung Hong, Peace Cyebukayire, Ikseon Choi, Jinho D. Choi
SIGDIAL5
2026 Building Task-Oriented Dialogue Systems via Instruction Guidance without Annotated Data
abstract
Task-oriented dialogue (TOD) systems conventionally rely on supervised fine-tuning over large datasets, an approach that is both resource-intensive and difficult to generalize across domains. We investigate whether large language models (LLMs) can serve as effective TOD agents without any fine-tuning, relying solely on in-context prompting and unstructured conversational logs. To this end, we propose a two-stage framework in which an LLM first induces structured procedural instructions from raw multi-turn dialogues, then leverages these instructions to generate goal-oriented interactions. An iterative refinement loop further improves instruction quality by evaluating intermediate dialogue outputs and propagating feedback to update the instructions. To address limitations inherent in existing evaluation protocols, we introduce an interactive evaluation framework centered on a constrained user simulator with access to ground-truth task goals. This design enables flexible assessment of task success beyond fixed dialogue trajectories, more faithfully reflecting the conditions of real-world deployment. Experiments demonstrate that the proposed approach produces coherent and task-effective dialogues without any annotated data. Using Gemma-3-27b-it as the backbone, our system achieves a dialogue state F1 of 86.3%, outperforming GALAXY (84.3%) and MARS (84.6%).
Henry Gao, Jinho D. Choi
SIGDIAL2
2026 PersonaKit (PK): A Plug-and-Play Platform for User Testing Diverse Roles in Simulated Full-Duplex Dialogue
abstract
As spoken dialogue systems take on diverse personas—authoritative instructors, uncooperative merchants, distracted workers—they need distinct, human-like turn-taking to stay immersive. Yet current full-duplex systems default to a rigid “always-yield” policy during overlap, undermining character consistency for non-submissive roles, and evaluating persona-specific alternatives through user studies demands substantial real-time engineering. We present PersonaKit (PK), an open-source, low-latency web platform that simulates full-duplex turn-taking pragmatics—via a cascade pipeline and prompt-level control rather than a jointly trained speak-while-listening model. Through intuitive JSON configurations, researchers define personas, specify probabilistic interruption handling (yield, hold, bridge, override), and auto-deploy comparative A/B surveys. An in-the-wild evaluation with 8 personas shows that PK is an extensible, end-to-end framework for studying complex sociolinguistic behaviors in next-generation spoken agents.
Hyunbae Jeon, Jinho D. Choi
SIGDIAL2
2026 TRUST: an large language model-based dialogue system for trauma understanding and structured assessments
abstract
OBJECTIVES: While large language models (LLMs) have been widely used to assist clinicians and support patients, no existing work has explored dialogue systems for standard diagnostic interviews and assessments. This study aims to bridge the gap in mental healthcare accessibility by developing an LLM-powered dialogue system that replicates clinician behavior. MATERIALS AND METHODS: We introduce TRUST, a framework of cooperative LLM modules capable of conducting formal diagnostic interviews and assessments for post-traumatic stress disorder (PTSD) following the Clinician-Administered PTSD Scale for DSM-5 (CAPS-5). To guide the generation of appropriate clinical responses, we propose a Dialogue Acts schema specifically designed for clinical interviews. Additionally, we develop a patient simulation approach based on real-life interview transcripts to replace time-consuming and costly manual testing by clinicians. RESULTS: A comprehensive set of evaluation metrics is designed to assess the dialogue system from both the agent and patient simulation perspectives. Expert evaluations by conversation and clinical specialists show that TRUST performs comparably to real-life clinical interviews. DISCUSSION: Our system performs with clinical quality approaching that of human clinicians, with room for future enhancements in communication styles and response appropriateness. CONCLUSIONS: Our TRUST framework shows its potential to facilitate mental healthcare availability.
Sichang Tu, Abigail Powers, Stephen Doogan, Jinho D. Choi
J. Am. Medical Informatics Assoc.4
2026 Generative Induction of Dialogue Task Schemas with Streaming Refinement and Simulated Interactions
abstract
Abstract In task-oriented dialogue (TOD) systems, Slot Schema Induction (SSI) is essential for automatically identifying key information slots from dialogue data without manual intervention. This paper presents a novel state-of-the-art (SotA) approach that formulates SSI as a text generation task, where a language model incrementally constructs and refines a slot schema over a stream of dialogue data. To develop this approach, we present a fully automatic LLM-based TOD simulation method that creates data with high-quality state labels for novel task domains. Furthermore, we identify issues in SSI evaluation due to data leakage and poor metric alignment with human judgment. We resolve these by creating new evaluation data using our simulation method with human guidance and correction, as well as designing improved evaluation metrics. These contributions establish a foundation for future SSI research and advance the SotA in dialogue understanding and system development.
James D. Finch, Yasasvi Josyula, Jinho D. Choi
Trans. Assoc. Comput. Linguistics3
2025 Finding A Voice: Exploring the Potential of African American Dialect and Voice Generation for Chatbots
abstract
As chatbots become integral to daily life, personalizing systems is key for fostering trust, engagement, and inclusivity.This study examines how linguistic similarity affects chatbot performance, focusing on integrating African American English (AAE) into virtual agents to better serve the African American community.We develop text-based and spoken chatbots using large language models and text-to-speech technology, then evaluate them with AAE speakers against standard English chatbots.Our results show that while text-based AAE chatbots often underperform, spoken chatbots benefit from an African American voice and AAE elements, improving performance and preference.These findings underscore the complexities of linguistic personalization and the dynamics between text and speech modalities, highlighting technological limitations that affect chatbots' AA speech generation and pointing to promising future research directions.
Sarah E. Finch, Ellie S. Paek, Ikseon Choi, Jinho D. Choi
ACL (1)4
2025 Reference-Aligned Retrieval-Augmented Question Answering over Heterogeneous Proprietary Documents
abstract
Proprietary corporate documents contain rich domain-specific knowledge, but their overwhelming volume and disorganized structure make it difficult even for employees to access the right information when needed. For example, in the automotive industry, vehicle crash-collision tests-each costing hundreds of thousands of dollars-produce highly detailed documentation. However, retrieving relevant content during decision-making remains time-consuming due to the scale and complexity of the material. While Retrieval-Augmented Generation (RAG)-based Question Answering (QA) systems offer a promising solution, building an internal RAG-QA system poses several challenges: (1) handling heterogeneous multi-modal data sources, (2) preserving data confidentiality, and (3) enabling traceability between each piece of information in the generated answer and its original source document. To address these, we propose a RAG-QA framework for internal enterprise use, consisting of: (1) a data pipeline that converts raw multi-modal documents into a structured corpus and QA pairs, (2) a fully on-premise, privacy-preserving architecture, and (3) a lightweight reference matcher that links answer segments to supporting content. Applied to the automotive domain, our system improves factual correctness (+1.79, +1.94), informativeness (+1.33, +1.16), and helpfulness (+1.08, +1.67) over a non-RAG baseline, based on 1-5 scale ratings from both human and LLM judge. The system was deployed internally for pilot testing and received positive feedback from employees.
Nayoung Choi, Grace Byun, Ellie S. Paek, Shinsun Lee, Jinho D. Choi
CIKM6
2025 Leveraging Explicit Reasoning for Inference Integration in Commonsense-Augmented Dialogue Models
abstract
Open-domain dialogue systems need to grasp social commonsense to understand and respond effectively to human users. Commonsense-augmented dialogue models have been proposed that aim to infer commonsense knowledge from dialogue contexts in order to improve response quality. However, existing approaches to commonsense-augmented dialogue rely on implicit reasoning to integrate commonsense inferences during response generation. In this study, we explore the impact of explicit reasoning against implicit reasoning over commonsense for dialogue response generation. Our findings demonstrate that separating commonsense reasoning into explicit steps for generating, selecting, and integrating commonsense into responses leads to better dialogue interactions, improving naturalness, engagement, specificity, and overall quality. Subsequent analyses of these findings unveil insights into the effectiveness of various types of commonsense in generating responses and the particular response traits enhanced through explicit reasoning for commonsense integration. Our work advances research in open-domain dialogue by achieving a new state-of-the-art in commonsense-augmented response generation.
Sarah E. Finch, Jinho D. Choi
COLING2
2024 Exploring the Impact of Human Evaluator Group on Chat-Oriented Dialogue Evaluation
abstract
Human evaluation has been widely accepted as the standard for evaluating chat-oriented dialogue systems. However, there is a significant variation in previous work regarding who gets recruited as evaluators. Evaluator groups such as domain experts, university students, and crowdworkers have been used to assess and compare dialogue systems, although it is unclear to what extent the choice of an evaluator group can affect results. This paper analyzes the evaluator group impact on dialogue system evaluation by testing 4 state-of-the-art dialogue systems using 4 distinct evaluator groups. Our analysis reveals a robustness towards evaluator groups for Likert evaluations that is not seen for Pairwise, with only minor differences observed when changing evaluator groups. Furthermore, two notable limitations to this robustness are observed, which reveal discrepancies between evaluators with different levels of chatbot expertise and indicate that evaluator objectivity is beneficial for certain dialogue metrics.
Sarah E. Finch, James D. Finch, Jinho D. Choi
LREC/COLING3
2024 Transforming Slot Schema Induction with Generative Dialogue State Inference
abstract
The challenge of defining a slot schema to represent the state of a task-oriented dialogue system is addressed by Slot Schema Induction (SSI), which aims to automatically induce slots from unlabeled dialogue data.Whereas previous approaches induce slots by clustering value spans extracted directly from the dialogue text, we demonstrate the power of discovering slots using a generative approach.By training a model to generate slot names and values that summarize key dialogue information with no prior task knowledge, our SSI method discovers high-quality candidate information for representing dialogue state.These discovered slotvalue candidates can be easily clustered into unified slot schemas that align well with humanauthored schemas.Experimental comparisons on the MultiWOZ and SGD datasets demonstrate that Generative Dialogue State Inference (GenDSI) outperforms the previous state-of-theart on multiple aspects of the SSI task.
James D. Finch, Boxin Zhao, Jinho D. Choi
SIGDIAL3
2024 Automating PTSD Diagnostics in Clinical Interviews: Leveraging Large Language Models for Trauma Assessments
abstract
Sichang Tu, Abigail Powers, Natalie Merrill, Negar Fani, Sierra Carter, Stephen Doogan, Jinho D. Choi. Proceedings of the 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2024.
Sichang Tu, Abigail Powers, Natalie Merrill, Negar Fani, Sierra Carter, Stephen Doogan, Jinho D. Choi
SIGDIAL7
2024 ConvoSense: Overcoming Monotonous Commonsense Inferences for Conversational AI
abstract
Abstract Mastering commonsense understanding and reasoning is a pivotal skill essential for conducting engaging conversations. While there have been several attempts to create datasets that facilitate commonsense inferences in dialogue contexts, existing datasets tend to lack in-depth details, restate information already present in the conversation, and often fail to capture the multifaceted nature of commonsense reasoning. In response to these limitations, we compile a new synthetic dataset for commonsense reasoning in dialogue contexts using GPT, ℂonvoSense, that boasts greater contextual novelty, offers a higher volume of inferences per example, and substantially enriches the detail conveyed by the inferences. Our dataset contains over 500,000 inferences across 12,000 dialogues with 10 popular inference types, which empowers the training of generative commonsense models for dialogue that are superior in producing plausible inferences with high novelty when compared to models trained on the previous datasets. To the best of our knowledge, ℂonvoSense is the first of its kind to provide such a multitude of novel inferences at such a large scale.
Sarah E. Finch, Jinho D. Choi
Trans. Assoc. Comput. Linguistics2
2023 Don't Forget Your ABC's: Evaluating the State-of-the-Art in Chat-Oriented Dialogue Systems
abstract
Despite tremendous advancements in dialogue systems, stable evaluation still requires human judgments producing notoriously high-variance metrics due to their inherent subjectivity.Moreover, methods and labels in dialogue evaluation are not fully standardized, especially for opendomain chats, with a lack of work to compare and assess the validity of those approaches.The use of inconsistent evaluation can misinform the performance of a dialogue system, which becomes a major hurdle to enhance it.Thus, a dimensional evaluation of chat-oriented opendomain dialogue systems that reliably measures several aspects of dialogue capabilities is desired.This paper presents a novel human evaluation method to estimate the rates of many dialogue system behaviors.Our method is used to evaluate four state-of-the-art open-domain dialogue systems and compared with existing approaches.The analysis demonstrates that our behavior method is more suitable than alternative Likert-style or comparative approaches for dimensional evaluation of these systems.
Sarah E. Finch, James D. Finch, Jinho D. Choi
ACL (1)3
2023 Towards Open-World Product Attribute Mining: A Lightly-Supervised Approach
abstract
We present a new task setting for attribute mining on e-commerce products, serving as a practical solution to extract open-world attributes without extensive human intervention.Our supervision comes from a high-quality seed attribute set bootstrapped from existing resources, and we aim to expand the attribute vocabulary of existing seed types, and also to discover any new attribute types automatically.A new dataset is created to support our setting, and our approach Amacer is proposed specifically to tackle the limited supervision.Especially, given that no direct supervision is available for those unseen new attributes, our novel formulation exploits selfsupervised heuristic and unsupervised latent attributes, which attains implicit semantic signals as additional supervision by leveraging product context.Experiments suggest that our approach surpasses various baselines by 12 F1, expanding attributes of existing types significantly by up to 12 times, and discovering values from 39% new types.Our data and code can be found at https://github.com/lxucs/woam.
Liyan Xu, Jingbo Shang, Jinho D. Choi
ACL (1)5
2023 FedTherapist: Mental Health Monitoring with User-Generated Linguistic Expressions on Smartphones via Federated Learning
abstract
Psychiatrists diagnose mental disorders via the linguistic use of patients.Still, due to data privacy, existing passive mental health monitoring systems use alternative features such as activity, app usage, and location via mobile devices.We propose FedTherapist, a mobile mental health monitoring system that utilizes continuous speech and keyboard input in a privacy-preserving way via federated learning.We explore multiple model designs by comparing their performance and overhead for FedTherapist to overcome the complex nature of on-device language model training on smartphones.We further propose a Context-Aware Language Learning (CALL) methodology to effectively utilize smartphones' large and noisy text for mental health signal sensing.Our IRBapproved evaluation of the prediction of selfreported depression, stress, anxiety, and mood from 46 participants shows higher accuracy of FedTherapist compared with the performance with non-language features, achieving 0.15 AU-ROC improvement and 8.21% MAE reduction.
Jaemin Shin 0005, Hyungjun Yoon, Seungjoo Lee, Yunxin Liu 0001, Jinho D. Choi, Sung-Ju Lee 0001
EMNLP6
2023 Aligning Speakers: Evaluating and Visualizing Text-based Speaker Diarization Using Efficient Multiple Sequence Alignment
abstract
This paper presents a novel evaluation approach to text-based speaker diarization (SD), tackling the limitations of traditional metrics that do not account for any contextual information in text. Two new metrics are proposed, Text-based Diarization Error Rate and Diarization F1, which perform utterance- and word-level evaluations by aligning tokens in reference and hypothesis transcripts. Our metrics encompass more types of errors compared to existing ones, allowing us to make a more comprehensive analysis in SD. To align tokens, a multiple sequence alignment algorithm is introduced that supports multiple sequences in the reference while handling high-dimensional alignment to the hypothesis using dynamic programming. Our work is packaged into two tools, align4d, providing an API for our alignment algorithm and TranscribeView for visualizing and evaluating SD errors, which can greatly aid in the creation of high-quality data, fostering the advancement of dialogue systems.
Jinho D. Choi
ICTAI3
2023 Leveraging Large Language Models for Automated Dialogue Analysis
abstract
Developing high-performing dialogue systems benefits from the automatic identification of undesirable behaviors in system responses.However, detecting such behaviors remains challenging, as it draws on a breadth of general knowledge and understanding of conversational practices.Although recent research has focused on building specialized classifiers for detecting specific dialogue behaviors, the behavior coverage is still incomplete and there is a lack of testing on real-world human-bot interactions.This paper investigates the ability of a state-of-the-art large language model (LLM), ChatGPT-3.5, to perform dialogue behavior detection for nine categories in real human-bot dialogues.We aim to assess whether ChatGPT can match specialized models and approximate human performance, thereby reducing the cost of behavior detection tasks.Our findings reveal that neither specialized models nor Chat-GPT have yet achieved satisfactory results for this task, falling short of human performance.Nevertheless, ChatGPT shows promising potential and often outperforms specialized detection models.We conclude with an in-depth examination of the prevalent shortcomings of ChatGPT, offering guidance for future research to enhance LLM capabilities.
Sarah E. Finch, Ellie S. Paek, Jinho D. Choi
SIGDIAL3
2023 Unleashing the True Potential of Sequence-to-Sequence Models for Sequence Tagging and Structure Parsing
abstract
Abstract Sequence-to-Sequence (S2S) models have achieved remarkable success on various text generation tasks. However, learning complex structures with S2S models remains challenging as external neural modules and additional lexicons are often supplemented to predict non-textual outputs. We present a systematic study of S2S modeling using contained decoding on four core tasks: part-of-speech tagging, named entity recognition, constituency, and dependency parsing, to develop efficient exploitation methods costing zero extra parameters. In particular, 3 lexically diverse linearization schemas and corresponding constrained decoding methods are designed and evaluated. Experiments show that although more lexicalized schemas yield longer output sequences that require heavier training, their sequences being closer to natural language makes them easier to learn. Moreover, S2S models using our constrained decoding outperform other S2S approaches using external resources. Our best models perform better than or comparably to the state-of-the-art for all 4 tasks, lighting a promise for S2S models to generate non-sequential structures.
Han He, Jinho D. Choi
Trans. Assoc. Comput. Linguistics2
2022 Zero-Shot Cross-Lingual Machine Reading Comprehension via Inter-sentence Dependency Graph
abstract
We target the task of cross-lingual Machine Reading Comprehension (MRC) in the direct zero-shot setting, by incorporating syntactic features from Universal Dependencies (UD), and the key features we use are the syntactic relations within each sentence. While previous work has demonstrated effective syntax-guided MRC models, we propose to adopt the inter-sentence syntactic relations, in addition to the rudimentary intra-sentence relations, to further utilize the syntactic dependencies in the multi-sentence input of the MRC task. In our approach, we build the Inter-Sentence Dependency Graph (ISDG) connecting dependency trees to form global syntactic relations across sentences. We then propose the ISDG encoder that encodes the global dependency graph, addressing the inter-sentence relations via both one-hop and multi-hop dependency paths explicitly. Experiments on three multilingual MRC datasets (XQuAD, MLQA, TyDiQA-GoldP) show that our encoder that is only trained on English is able to improve the zero-shot performance on all 14 test sets covering 8 languages, with up to 3.8 F1 / 5.2 EM improvement on-average, and 5.2 F1 / 11.2 EM on certain languages. Further analysis shows the improvement can be attributed to the attention on the cross-linguistically consistent syntactic path. Our code is available at https://github.com/lxucs/multilingual-mrc-isdg.
Liyan Xu, Xuchao Zhang, Bo Zong, Yanchi Liu, Wei Cheng 0002, Jingchao Ni, Liang Zhao 0002, Jinho D. Choi
AAAI9
2022 Automatic Generation of Large-scale Multi-turn Dialogues from Reddit
abstract
This paper presents novel methods to automatically convert posts and their comments from discussion forums such as Reddit into multi-turn dialogues. Our methods are generalizable to any forums; thus, they allow us to generate a massive amount of dialogues for diverse topics that can be used to pretrain language models. Four methods are introduced, Greedy_Baseline, Greedy_Advanced, Beam Search and Threading, which are applied to posts from 10 subreddits and assessed. Each method makes a noticeable improvement over its predecessor such that the best method shows an improvement of 36.3% over the baseline for appropriateness. Our best method is applied to posts from those 10 subreddits for the creation of a corpus comprising 10,098 dialogues (3.3M tokens), 570 of which are compared against dialogues in three other datasets, Blended Skill Talk, Daily Dialogue, and Topical Chat. Our dialogues are found to be more engaging but slightly less natural than the ones in the other datasets, while it costs a fraction of human labor and money to generate our corpus compared to the others. To the best of our knowledge, it is the first work to create a large multi-turn dialogue corpus from Reddit that can advance neural dialogue systems.
Daniil Huryn, William Hutsell, Jinho D. Choi
COLING3
2022 Modeling Task Interactions in Document-Level Joint Entity and Relation Extraction
abstract
We target on the document-level relation extraction in an end-to-end setting, where the model needs to jointly perform mention extraction, coreference resolution (COREF) and relation extraction (RE) at once, and gets evaluated in an entity-centric way.Especially, we address the two-way interaction between COREF and RE that has not been the focus by previous work, and propose to introduce explicit interaction namely Graph Compatibility (GC) that is specifically designed to leverage task characteristics, bridging decisions of two tasks for direct task interference.Our experiments are conducted on DocRED and DWIE; in addition to GC, we implement and compare different multi-task settings commonly adopted in previous work, including pipeline, shared encoders, graph propagation, to examine the effectiveness of different interactions.The result shows that GC achieves the best performance by up to 2.3/5.1 F1 improvement over the baseline.
Liyan Xu, Jinho D. Choi
NAACL-HLT2
2021 SMAT: An Attention-Based Deep Learning Solution to the Automation of Schema Matching
Bonggun Shin, Jinho D. Choi, Joyce C. Ho
ADBIS3
2021 The Stem Cell Hypothesis: Dilemma behind Multi-Task Learning with Transformer Encoders
abstract
Multi-task learning with transformer encoders (MTL) has emerged as a powerful technique to improve performance on closely-related tasks for both accuracy and efficiency while a question still remains whether or not it would perform as well on tasks that are distinct in nature.We first present MTL results on five NLP tasks, POS, NER, DEP, CON, and SRL, and depict its deficiency over single-task learning.We then conduct an extensive pruning analysis to show that a certain set of attention heads get claimed by most tasks during MTL, who interfere with one another to fine-tune those heads for their own objectives.Based on this finding, we propose the Stem Cell Hypothesis to reveal the existence of attention heads naturally talented for many tasks that cannot be jointly trained to create adequate embeddings for all of those tasks.Finally, we design novel parameter-free probes to justify our hypothesis and demonstrate how attention heads are transformed across the five tasks during MTL through label analysis.
Han He, Jinho D. Choi
EMNLP (1)2
2021 Boosting Cross-Lingual Transfer via Self-Learning with Uncertainty Estimation
abstract
Recent multilingual pre-trained language models have achieved remarkable zero-shot performance, where the model is only finetuned on one source language and directly evaluated on target languages.In this work, we propose a self-learning framework that further utilizes unlabeled data of target languages, combined with uncertainty estimation in the process to select high-quality silver labels.Three different uncertainties are adapted and analyzed specifically for the cross lingual transfer: Language Heteroscedastic/Homoscedastic Uncertainty (LEU/LOU), Evidential Uncertainty (EVI).We evaluate our framework with uncertainties on two cross-lingual tasks including Named Entity Recognition (NER) and Natural Language Inference (NLI) covering 40 languages in total, which outperforms the baselines significantly by 10 F1 on average for NER and 2.5 accuracy score for NLI.
Liyan Xu, Xuchao Zhang, Xujiang Zhao, Feng Chen 0001, Jinho D. Choi
EMNLP (1)6
2020 Incremental Sense Weight Training for In-Depth Interpretation of Contextualized Word Embeddings (Student Abstract)
abstract
We present a novel online algorithm that learns the essence of each dimension in word embeddings. We first mask dimensions determined unessential by our algorithm, apply the masked word embeddings to a word sense disambiguation task (WSD), and compare its performance against the one achieved by the original embeddings. Our results show that the masked word embeddings do not hurt the performance and can improve it by 3%.
Xinyi Jiang 0007, Zhengzhe Yang, Jinho D. Choi
AAAI3
2020 Automatic Text-Based Personality Recognition on Monologues and Multiparty Dialogues Using Attentive Networks and Contextual Embeddings (Student Abstract)
abstract
Previous works related to automatic personality recognition focus on using traditional classification models with linguistic features. However, attentive neural networks with contextual embeddings, which have achieved huge success in text classification, are rarely explored for this task. In this project, we have two major contributions. First, we create the first dialogue-based personality dataset, FriendsPersona , by annotating 5 personality traits of speakers from Friends TV Show through crowdsourcing. Second, we present a novel approach to automatic personality recognition using pre-trained contextual embeddings (BERT and RoBERTa) and attentive neural networks. Our models largely improve the state-of-art results on the monologue Essays dataset by 2.49%, and establish a solid benchmark on our FriendsPersona. By comparing results in two datasets, we demonstrate the challenges of modeling personality in multi-party dialogue.
Xianzhe Zhang, Jinho D. Choi
AAAI3
2020 Transformers to Learn Hierarchical Contexts in Multiparty Dialogue for Span-based Question Answering
abstract
We introduce a novel approach to transformers that learns hierarchical representations in multiparty dialogue.First, three language modeling tasks are used to pre-train the transformers, token-and utterance-level language modeling and utterance order prediction, that learn both token and utterance embeddings for better understanding in dialogue contexts.Then, multitask learning between the utterance prediction and the token span prediction is applied to finetune for span-based question answering (QA).Our approach is evaluated on the FRIENDSQA dataset and shows improvements of 3.8% and 1.4% over the two state-of-the-art transformer models, BERT and RoBERTa, respectively.
Changmao Li, Jinho D. Choi
ACL2
2020 Competence-Level Prediction and Resume & Job Description Matching Using Context-Aware Transformer Models
abstract
This paper presents a comprehensive study on resume classification to reduce the time and labor needed to screen an overwhelming number of applications significantly, while improving the selection of suitable candidates.A total of 6,492 resumes are extracted from 24,933 job applications for 252 positions designated into four levels of experience for Clinical Research Coordinators (CRC).Each resume is manually annotated to its most appropriate CRC position by experts through several rounds of triple annotation to establish guidelines.As a result, a high Kappa score of 61% is achieved for interannotator agreement.Given this dataset, novel transformer-based classification models are developed for two tasks: the first task takes a resume and classifies it to a CRC level (T1), and the second task takes both a resume and a job description to apply and predicts if the application is suited to the job (T2).Our best models using section encoding and multi-head attention decoding give results of 73.3% to T1 and 79.2% to T2.Our analysis shows that the prediction errors are mostly made among adjacent CRC levels, which are hard for even experts to distinguish, implying the practical value of our models in real HR platforms.
Changmao Li, Elaine Fisher, Rebecca Thomas, Steve Pittard, Vicki Stover Hertzberg, Jinho D. Choi
EMNLP (1)6
2020 Revealing the Myth of Higher-Order Inference in Coreference Resolution
abstract
This paper analyzes the impact of higher-order inference (HOI) on the task of coreference resolution.HOI has been adapted by almost all recent coreference resolution models without taking much investigation on its true effectiveness over representation learning.To make a comprehensive analysis, we implement an endto-end coreference system as well as four HOI approaches, attended antecedent, entity equalization, span clustering, and cluster merging, where the latter two are our original methods.We find that given a high-performing encoder such as SpanBERT, the impact of HOI is negative to marginal, providing a new perspective of HOI to this task.Our best model using cluster merging shows the Avg-F1 of 80.2 on the CoNLL 2012 shared task dataset in English.
Liyan Xu, Jinho D. Choi
EMNLP (1)2
2020 Towards Unified Dialogue System Evaluation: A Comprehensive Analysis of Current Evaluation Protocols
abstract
As conversational AI-based dialogue management has increasingly become a trending topic, the need for a standardized and reliable evaluation procedure grows even more pressing.The current state of affairs suggests various evaluation protocols to assess chat-oriented dialogue management systems, rendering it difficult to conduct fair comparative studies across different approaches and gain an insightful understanding of their values.To foster this research, a more robust evaluation protocol must be set in place.This paper presents a comprehensive synthesis of both automated and human evaluation methods on dialogue systems, identifying their shortcomings while accumulating evidence towards the most effective evaluation dimensions.A total of 20 papers from the last two years are surveyed to analyze three types of evaluation protocols: automated, static, and interactive.Finally, the evaluation dimensions used in these papers are compared against our expert evaluation on the system-user dialogue data collected from the Alexa Prize 2020.
Sarah E. Finch, Jinho D. Choi
SIGdial2
2020 Emora STDM: A Versatile Framework for Innovative Dialogue System Development
abstract
This demo paper presents Emora STDM (State Transition Dialogue Manager), a dialogue system development framework that provides novel workflows for rapid prototyping of chatbased dialogue managers as well as collaborative development of complex interactions.Our framework caters to a wide range of expertise levels by supporting interoperability between two popular approaches, state machine and information state, to dialogue management.Our Natural Language Expression package allows seamless integration of pattern matching, custom NLP modules, and database querying, that makes the workflows much more efficient.As a user study, we adopt this framework to an interdisciplinary undergraduate course where students with both technical and non-technical backgrounds are able to develop creative dialogue managers in a short period of time.
James D. Finch, Jinho D. Choi
SIGdial2
2019 The Pupil Has Become the Master: Teacher-Student Model-Based Word Embedding Distillation with Ensemble Learning
abstract
Recent advances in deep learning have facilitated the demand of neural models for real applications. In practice, these applications often need to be deployed with limited resources while keeping high accuracy. This paper touches the core of neural models in NLP, word embeddings, and presents an embedding distillation framework that remarkably reduces the dimension of word embeddings without compromising accuracy. A new distillation ensemble approach is also proposed that trains a high-efficient student model using multiple teacher models. In our approach, the teacher models play roles only during training such that the student model operates on its own without getting supports from the teacher models during decoding, which makes it run as fast and light as any single model. All models are evaluated on seven document classification datasets and show significant advantage over the teacher models for most cases. Our analysis depicts insightful transformation of word embeddings from distillation and suggests a future direction to ensemble approaches using neural models.
Bonggun Shin, Jinho D. Choi
IJCAI3
2019 FriendsQA: Open-Domain Question Answering on TV Show Transcripts
abstract
This paper presents FriendsQA, a challenging question answering dataset that contains 1,222 dialogues and 10,610 open-domain questions, to tackle machine comprehension on everyday conversations.Each dialogue, involving multiple speakers, is annotated with several types of questions regarding the dialogue contexts, and the answers are annotated with certain spans in the dialogue.A series of crowdsourcing tasks are conducted to ensure good annotation quality, resulting a high inter-annotator agreement of 81.82%.A comprehensive annotation analytics is provided for a deeper understanding in this dataset.Three state-of-the-art QA systems are experimented, R-Net, QANet, and BERT, and evaluated on this dataset.BERT in particular depicts promising results, an accuracy of 74.2% for answer utterance selection and an F1-score of 64.2% for answer span selection, suggesting that the FriendsQA task is hard yet has a great potential of elevating QA research on multiparty dialogue to another level.
Zhengzhe Yang, Jinho D. Choi
SIGdial2
2018 They Exist! Introducing Plural Mentions to Coreference Resolution and Entity Linking
abstract
This paper analyzes arguably the most challenging yet under-explored aspect of resolution tasks such as coreference resolution and entity linking, that is the resolution of plural mentions. Unlike singular mentions each of which represents one entity, plural mentions stand for multiple entities. To tackle this aspect, we take the character identification corpus from the SemEval 2018 shared task that consists of entity annotation for singular mentions, and expand it by adding annotation for plural mentions. We then introduce a novel coreference resolution algorithm that selectively creates clusters to handle both singular and plural mentions, and also a deep learning-based entity linking model that jointly handles both types of mentions through multi-task learning. Adjusted evaluation metrics are proposed for these tasks as well to handle the uniqueness of plural mentions. Our experiments show that the new coreference resolution and entity linking models significantly outperform traditional models designed only for singular mentions. To the best of our knowledge, this is the first time that plural mentions are thoroughly analyzed for these two resolution tasks.
Ethan Zhou, Jinho D. Choi
COLING2
2018 Building Universal Dependency Treebanks in Korean
Jayeol Chun, Na-Rae Han, Jena D. Hwang, Jinho D. Choi
LREC4
2018 Challenging Reading Comprehension on Daily Conversation: Passage Completion on Multiparty Dialog
abstract
Kaixin Ma, Tomasz Jurczyk, Jinho D. Choi. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Kaixin Ma, Tomasz Jurczyk, Jinho D. Choi
NAACL-HLT3
2017 Robust Coreference Resolution and Entity Linking on Dialogues: Character Identification on TV Show Transcripts
abstract
This paper presents a novel approach to character identification, that is an entity linking task that maps mentions to characters in dialogues from TV show transcripts.We first augment and correct several cases of annotation errors in an existing corpus so the corpus is clearer and cleaner for statistical learning.We also introduce the agglomerative convolutional neural network that takes groups of features and learns mention and mention-pair embeddings for coreference resolution.We then propose another neural model that employs the embeddings learned and creates cluster embeddings for entity linking.Our coreference resolution model shows comparable results to other state-of-the-art systems.Our entity linking model significantly outperforms the previous work, showing the F1 score of 86.76% and the accuracy of 95.30% for character identification.
Henry Y. Chen, Ethan Zhou, Jinho D. Choi
CoNLL3
2017 Classification of radiology reports using neural attention models
abstract
The electronic health record (EHR) contains a large amount of multi-dimensional and unstructured clinical data of significant operational and research value. Distinguished from previous studies, our approach embraces a double-annotated dataset and strays away from obscure “black-box” models to comprehensive deep learning models. In this paper, we present a novel neural attention mechanism that not only classifies clinically important findings. Specifically, convolutional neural networks (CNN) with attention analysis are used to classify radiology head computed tomography reports based on five categories that radiologists would account for in assessing acute and communicable findings in daily practice. The experiments show that our CNN attention models outperform non-neural models, especially when trained on a larger dataset. Our attention analysis demonstrates the intuition behind the classifier's decision by generating a heatmap that highlights attended terms used by the CNN model; this is valuable when potential downstream medical decisions are to be performed by human experts or the classifier information is to be used in cohort construction such as for epidemiological studies.
Bonggun Shin, Falgun H. Chokshi, Timothy Lee, Jinho D. Choi
IJCNN4
2016 Intrinsic and Extrinsic Evaluations of Word Embeddings
abstract
In this paper, we first analyze the semantic composition of word embeddings by cross-referencing their clusters with the manual lexical database, WordNet. We then evaluate a variety of word embedding approaches by comparing their contributions to two NLP tasks. Our experiments show that the word embedding clusters give high correlations to the synonym and hyponym sets in WordNet, and give 0.88% and 0.17% absolute improvements in accuracy to named entity recognition and part-of-speech tagging, respectively.
Michael Zhai, Johnny Tan, Jinho D. Choi
AAAI3
2016 SelQA: A New Benchmark for Selection-Based Question Answering
abstract
This paper presents a new selection-based question answering dataset, SelQA. The dataset consists of questions generated through crowdsourcing and sentence length answers that are drawn from the ten most prevalent topics in the English Wikipedia. We introduce a corpus annotation scheme that enhances the generation of large, diverse, and challenging datasets by explicitly aiming to reduce word co-occurrences between the question and answers. Our annotation scheme is composed of a series of crowdsourcing tasks with a view to more effectively utilize crowdsourcing in the creation of question answering datasets in various domains. Several systems are compared on the tasks of answer sentence selection and answer triggering, providing strong baseline results for future work to improve upon.
Tomasz Jurczyk, Michael Zhai, Jinho D. Choi
ICTAI3
2016 Dynamic Feature Induction: The Last Gist to the State-of-the-Art
abstract
We introduce a novel technique called dynamic feature induction that keeps inducing high dimensional features automatically until the feature space becomes 'more' linearly separable.Dynamic feature induction searches for the feature combinations that give strong clues for distinguishing certain label pairs, and generates joint features from these combinations.These induced features are trained along with the primitive low dimensional features.Our approach was evaluated on two core NLP tasks, part-of-speech tagging and named entity recognition, and showed the state-of-the-art results for both tasks, achieving the accuracy of 97.64 and the F1-score of 91.00 respectively, with about a 25% increase in the feature space.
Jinho D. Choi
HLT-NAACL1
2016 Character Identification on Multiparty Conversation: Identifying Mentions of Characters in TV Shows
abstract
This paper introduces a subtask of entity linking, called character identification, that maps mentions in multiparty conversation to their referent characters.Transcripts of TV shows are collected as the sources of our corpus and automatically annotated with mentions by linguistically-motivated rules.These mentions are manually linked to their referents through crowdsourcing.Our corpus comprises 543 scenes from two TV shows, and shows the inter-annotator agreement of κ = 79.96.For statistical modeling, this task is reformulated as coreference resolution, and experimented with a state-of-the-art system on our corpus.Our best model gives a purity score of 69.21 on average, which is promising given the challenging nature of this task and our corpus.
Jinho D. Choi
SIGDIAL Conference2
2015 It Depends: Dependency Parser Comparison Using A Web-based Evaluation Tool
abstract
Jinho D. Choi, Joel Tetreault, Amanda Stent. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Jinho D. Choi, Joel R. Tetreault, Amanda Stent
ACL (1)1
2015 Real-Time Community Question Answering: Exploring Content Recommendation and User Notification Strategies
abstract
Community-based Question Answering (CQA) services allow users to find and share information by interacting with others. A key to the success of CQA services is the quality and timeliness of the responses that users get. With the increasing use of mobile devices, searchers increasingly expect to find more local and time-sensitive information, such as the current special at a cafe around the corner. Yet, few services provide such hyper-local and time-aware question answering. This requires intelligent content recommendation and careful use of notifications (e.g., recommending questions to only selected users). To explore these issues, we developed RealQA, a real-time CQA system with a mobile interface, and performed two user studies: a formative pilot study with the initial system design, and a more extensive study with the revised UI and algorithms. The research design combined qualitative survey analysis and quantitative behavior analysis under different conditions. We report our findings of the prevalent information needs and types of responses users provided, and of the effectiveness of the recommendation and notification strategies on user experience and satisfaction. Our system and findings offer insights and implications for designing real-time CQA systems, and provide a valuable platform for future research.
Qiaoling Liu, Tomasz Jurczyk, Jinho D. Choi, Eugene Agichtein
IUI3
2015 Semantics-based Graph Approach to Complex Question-Answering
abstract
This paper suggests an architectural approach of representing knowledge graph for complex question-answering.There are four kinds of entity relations added to our knowledge graph: syntactic dependencies, semantic role labels, named entities, and coreference links, which can be effectively applied to answer complex questions.As a proof of concept, we demonstrate how our knowledge graph can be used to solve complex questions such as arithmetics.Our experiment shows a promising result on solving arithmetic questions, achieving the 3folds cross-validation score of 71.75%.
Tomasz Jurczyk, Jinho D. Choi
HLT-NAACL2
2015 Computational Exploration to Linguistic Structures of Future: Classification and Categorization
abstract
Aiming Ni, Jinho D. Choi, Jason Shepard, Phillip Wolff. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Student Research Workshop. 2015.
Aiming Ni, Jinho D. Choi, Jason Shepard, Phillip Wolff
HLT-NAACL2
2013 Transition-based Dependency Parsing with Selectional Branching
Jinho D. Choi, Andrew McCallum
ACL (1)1
2013 Dynamic Knowledge-Base Alignment for Coreference Resolution
Jiaping Zheng, Luke Vilnis, Sameer Singh 0001, Jinho D. Choi, Andrew McCallum
CoNLL4
2013 Towards comprehensive syntactic and semantic annotations of the clinical narrative
abstract
OBJECTIVE: To create annotated clinical narratives with layers of syntactic and semantic labels to facilitate advances in clinical natural language processing (NLP). To develop NLP algorithms and open source components. METHODS: Manual annotation of a clinical narrative corpus of 127 606 tokens following the Treebank schema for syntactic information, PropBank schema for predicate-argument structures, and the Unified Medical Language System (UMLS) schema for semantic information. NLP components were developed. RESULTS: The final corpus consists of 13 091 sentences containing 1772 distinct predicate lemmas. Of the 766 newly created PropBank frames, 74 are verbs. There are 28 539 named entity (NE) annotations spread over 15 UMLS semantic groups, one UMLS semantic type, and the Person semantic category. The most frequent annotations belong to the UMLS semantic groups of Procedures (15.71%), Disorders (14.74%), Concepts and Ideas (15.10%), Anatomy (12.80%), Chemicals and Drugs (7.49%), and the UMLS semantic type of Sign or Symptom (12.46%). Inter-annotator agreement results: Treebank (0.926), PropBank (0.891-0.931), NE (0.697-0.750). The part-of-speech tagger, constituency parser, dependency parser, and semantic role labeler are built from the corpus and released open source. A significant limitation uncovered by this project is the need for the NLP community to develop a widely agreed-upon schema for the annotation of clinical concepts and their relations. CONCLUSIONS: This project takes a foundational step towards bringing the field of clinical NLP up to par with NLP in the general domain. The corpus creation and NLP components provide a resource for research and application development that would have been previously impossible.
Daniel Albright, Arrick Lanfranchi, Anwen Fredriksen, William F. Styler IV, Colin Warner, Jena D. Hwang, Jinho D. Choi, Dmitriy Dligach, Rodney D. Nielsen, James H. Martin, Wayne H. Ward, Martha Palmer, Guergana K. Savova
J. Am. Medical Informatics Assoc.7
2012 Empty Argument Insertion in the Hindi PropBank
Ashwini Vaidya, Jinho D. Choi, Martha Palmer, Bhuvana Narasimhan
LREC2
2012 A corpus of full-text journal articles is a robust evaluation tool for revealing differences in performance of biomedical natural language processing tools
abstract
BACKGROUND: We introduce the linguistic annotation of a corpus of 97 full-text biomedical publications, known as the Colorado Richly Annotated Full Text (CRAFT) corpus. We further assess the performance of existing tools for performing sentence splitting, tokenization, syntactic parsing, and named entity recognition on this corpus. RESULTS: Many biomedical natural language processing systems demonstrated large differences between their previously published results and their performance on the CRAFT corpus when tested with the publicly available models or rule sets. Trainable systems differed widely with respect to their ability to build high-performing models based on this data. CONCLUSIONS: The finding that some systems were able to train high-performing models based on this corpus is additional evidence, beyond high inter-annotator agreement, that the quality of the CRAFT corpus is high. The overall poor performance of various systems indicates that considerable work needs to be done to enable natural language processing systems to work well when the input is full-text journal articles. The CRAFT corpus provides a valuable resource to the biomedical natural language processing community for evaluation and training of new models for biomedical full text publications.
Karin Verspoor, Kevin Cohen 0001, Arrick Lanfranchi, Colin Warner, Helen L. Johnson 0001, Christophe Roeder, Jinho D. Choi, Christopher S. Funk, Yuriy Malenkiy, Miriam Eckert, Nianwen Xue, William A. Baumgartner Jr., Michael Bada, Martha Palmer, Lawrence Hunter
BMC Bioinform.7
2010 Propbank Instance Annotation Guidelines Using a Dedicated Editor, Jubilee
Jinho D. Choi, Claire Bonial, Martha Palmer
LREC1
2010 Propbank Frameset Annotation Guidelines Using a Dedicated Editor, Cornerstone
Jinho D. Choi, Claire Bonial, Martha Palmer
LREC1