Hyundong Cho

dblp:263/6759 · also Hyundong Justin Cho · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Question answering and dialogue systems · 39% Language models and text generation · 35% Reinforcement learning · 16%
Human-computer interaction and pervasive computing
1 paper
Health and well-being technologies · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%
Databases, data mining, and information retrieval
1 paper
Web and social media mining · 100%

Topics — the 14 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
0.912025
Aligning Language Models with Demonstrated Feedback · ICLR 2025
Machine learning › Reinforcement learning
imitation learning
0.912025
Aligning Language Models with Demonstrated Feedback · ICLR 2025
Machine learning › Reinforcement learning › imitation learning
online imitation learning
0.912025
Aligning Language Models with Demonstrated Feedback · ICLR 2025
Natural language and speech › Language models and text generation
instruction tuning
0.812024
Speechworthy Instruction-tuned Language Models · EMNLP 2024
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue state tracking
0.712023
Continual Dialogue State Tracking via Example-Guided Question Answering · EMNLP 2023
Natural language and speech › Question answering and dialogue systems › dialogue generation
personalized dialogue generation
0.712023
RECAP: Retrieval-Enhanced Context-Aware Prefix Encoder for Personalized Dialogue Response Generation · ACL (1) 2023
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.712023
RECAP: Retrieval-Enhanced Context-Aware Prefix Encoder for Personalized Dialogue Response Generation · ACL (1) 2023
Computational social science and digital humanities › platform governance
online community moderation
0.712023
Analyzing Norm Violations in Live-Stream Chat · EMNLP 2023
Natural language and speech › Question answering and dialogue systems › dialogue generation
dialogue response generation
0.612022
Reflect, Not Reflex: Inference-Based Common Ground Improves Dialogue Response Quality · EMNLP 2022
Natural language and speech › Language models and text generation
prompting
0.612022
Reflect, Not Reflex: Inference-Based Common Ground Improves Dialogue Response Quality · EMNLP 2022
Natural language and speech › Question answering and dialogue systems › dialogue
common ground
0.412020
Grounding Conversations with Improvised Dialogues · ACL 2020
Natural language and speech › Question answering and dialogue systems › dialogue modeling
conversational grounding
0.412020
Grounding Conversations with Improvised Dialogues · ACL 2020
Natural language and speech › Question answering and dialogue systems › dialogue dataset
dialogue dataset construction
0.412020
Grounding Conversations with Improvised Dialogues · ACL 2020
Natural language and speech › Speech recognition and synthesis
spoken language generation
0.212024
Speechworthy Instruction-tuned Language Models · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

large language model assistants · 2.0content analysis · 1.3supervised fine-tuning · 0.9online imitation learning · 0.9large language model · 0.9RLHF · 0.9instruction tuning · 0.8hierarchical transformer retriever · 0.7example-guided question answering · 0.7context-aware prefix encoder · 0.7dataset annotation · 0.6
YearPublicationVenuePosition
2026 Uncovering Intervention Opportunities for Suicide Prevention with Language Model Assistants
abstract
Jaspreet Ranjit, Hyundong Justin Cho, Claire J. Smerdon, Yoonsoo Nam, Myles Phung, Jonathan May, John R. Blosnich, Swabha Swayamdipta. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jaspreet Ranjit, Hyundong Cho, Claire J. Smerdon, Yoonsoo Nam, Myles Phung, Jonathan May, John R. Blosnich, Swabha Swayamdipta
ACL (1)2
2025 NewsInterview: a Dataset and a Playground to Evaluate LLMs' Grounding Gap via Informational Interviews
abstract
Alexander Spangher, Michael Lu, Sriya Kalyan, Hyundong Justin Cho, Tenghao Huang, Weiyan Shi, Jonathan May. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Alexander Spangher, Michael Lu, Sriya Kalyan, Hyundong Cho, Tenghao Huang, Weiyan Shi 0001, Jonathan May
ACL (1)4
2025 Aligning Language Models with Demonstrated Feedback
abstract
Language models are aligned to emulate the collective voice of many, resulting in outputs that align with no one in particular. Steering LLMs away from generic output is possible through supervised finetuning or RLHF, but requires prohibitively large datasets for new ad-hoc tasks. We argue that it is instead possible to align an LLM to a specific setting by leveraging a very small number ($<10$) of demonstrations as feedback. Our method, Demonstration ITerated Task Optimization (DITTO), directly aligns language model outputs to a user's demonstrated behaviors. Derived using ideas from online imitation learning, DITTO cheaply generates online comparison data by treating users' demonstrations as preferred over output from the LLM and its intermediate checkpoints. We evaluate DITTO's ability to learn fine-grained style and task alignment across domains such as news articles, emails, and blog posts. Additionally, we conduct a user study soliciting a range of demonstrations from participants ($N=16$). Across our benchmarks and user study, we find that win-rates for DITTO outperform few-shot prompting, supervised fine-tuning, and other self-play methods by an average of 19\% points. By using demonstrations as feedback directly, DITTO offers a novel method for effective customization of LLMs.
Omar Shaikh, Michelle S. Lam, Joey Hejna, Yijia Shao, Hyundong Cho, Michael S. Bernstein, Diyi Yang
ICLR5
2024 Speechworthy Instruction-tuned Language Models
abstract
Hyundong Justin Cho, Nicolaas Paul Jedema, Leonardo F. R. Ribeiro, Karishma Sharma, Pedro Szekely, Alessandro Moschitti, Ruben Janssen, Jonathan May. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Hyundong Cho, Nicolaas Paul Jedema, Leonardo F. R. Ribeiro, Karishma Sharma, Pedro A. Szekely, Alessandro Moschitti, Ruben Janssen, Jonathan May
EMNLP1
2024 Can Language Model Moderators Improve the Health of Online Discourse?
abstract
Hyundong Cho, Shuai Liu, Taiwei Shi, Darpan Jain, Basem Rizk, Yuyang Huang, Zixun Lu, Nuan Wen, Jonathan Gratch, Emilio Ferrara, Jonathan May. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Hyundong Cho, Taiwei Shi, Darpan Jain, Basem Rizk, Zixun Lu, Nuan Wen, Jonathan Gratch, Emilio Ferrara, Jonathan May
NAACL-HLT1
2023 RECAP: Retrieval-Enhanced Context-Aware Prefix Encoder for Personalized Dialogue Response Generation
abstract
Endowing chatbots with a consistent persona is essential to an engaging conversation, yet it remains an unresolved challenge.In this work, we propose a new retrieval-enhanced approach for personalized response generation.Specifically, we design a hierarchical transformer retriever trained on dialogue domain data to perform personalized retrieval and a context-aware prefix encoder that fuses the retrieved information to the decoder more effectively.Extensive experiments on a real-world dataset demonstrate the effectiveness of our model at generating more fluent and personalized responses.We quantitatively evaluate our model's performance under a suite of human and automatic metrics and find it to be superior compared to state-of-the-art baselines on English Reddit conversations. 1
Hyundong Cho, Marjorie Freedman, Xuezhe Ma, Jonathan May
ACL (1)2
2023 Continual Dialogue State Tracking via Example-Guided Question Answering
abstract
Hyundong Cho, Andrea Madotto, Zhaojiang Lin, Khyathi Chandu, Satwik Kottur, Jing Xu, Jonathan May, Chinnadhurai Sankar. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Hyundong Cho, Andrea Madotto, Zhaojiang Lin, Khyathi Raghavi Chandu, Satwik Kottur, Jonathan May, Chinnadhurai Sankar
EMNLP1
2023 Analyzing Norm Violations in Live-Stream Chat
abstract
Jihyung Moon, Dong-Ho Lee, Hyundong Cho, Woojeong Jin, Chan Park, Minwoo Kim, Jonathan May, Jay Pujara, Sungjoon Park. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Jihyung Moon, Hyundong Cho, Woojeong Jin 0001, Chan Young Park, Jonathan May, Jay Pujara
EMNLP3
2022 Reflect, Not Reflex: Inference-Based Common Ground Improves Dialogue Response Quality
abstract
Human communication relies on common ground (CG), the mutual knowledge and beliefs shared by participants, to produce coherent and interesting conversations.In this paper, we demonstrate that current response generation (RG) models produce generic and dull responses in dialogues because they act reflexively, failing to explicitly model CG, both due to the lack of CG in training data and the standard RG training procedure.We introduce Reflect, a dataset that annotates dialogues with explicit CG (materialized as inferences approximating shared knowledge and beliefs) and solicits 9k diverse human-generated responses each following one common ground.Using Reflect, we showcase the limitations of current dialogue data and RG models: less than half of the responses in current data is rated as high quality (sensible, specific, and interesting) and models trained using this data have even lower quality, while most Reflect responses are judged high quality.Next, we analyze whether CG can help models produce better quality responses by using Reflect CG to guide RG models.Surprisingly, we find that simply prompting GPT3 to "think" about CG generates 30% more quality responses, showing promising benefits to integrating CG into the RG process. 1
Hyundong Cho, Pegah Jandaghi, Bill Y. Lin, Jay Pujara, Xiang Ren 0001
EMNLP2
2020 Grounding Conversations with Improvised Dialogues
abstract
Effective dialogue involves grounding, the process of establishing mutual knowledge that is essential for communication between people.Modern dialogue systems are not explicitly trained to build common ground, and therefore overlook this important aspect of communication.Improvisational theater (improv) intrinsically contains a high proportion of dialogue focused on building common ground, and makes use of the yes-and principle, a strong grounding speech act, to establish coherence and an actionable objective reality.We collect a corpus of more than 26,000 yes-and turns, transcribing them from improv dialogues and extracting them from larger, but more sparsely populated movie script dialogue corpora, via a bootstrapped classifier.We fine-tune chit-chat dialogue systems with our corpus to encourage more grounded, relevant conversation and confirm these findings with human evaluations.
Hyundong Cho, Jonathan May
ACL1