Alice Oh

dblp:50/7562 · also Alice H. Oh, Alice Hae Yun Oh · DBLP profile ↗
← Back
83ranked-venue papers
3as first author
40since 2021 · last 2026
0000-0002-7884-3038ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 67 · 2 first-author · 36 since 2021Human-computer interaction and ubiquitous computing · 13 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-author · 1 since 2021Systems, architecture and hardware · 7 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Security and privacy · 1
YearPublicationVenuePosition
2026 Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
abstract
Eunsu Kim, Junyeong Park, Juhyun Oh, Kiwoong Park, Seyoung Song, A. Seza Doğruöz, Alice Oh, Najoung Kim. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Eunsu Kim, Junyeong Park, Juhyun Oh, Kiwoong Park, Seyoung Song 0001, A. Seza Dogruöz, Alice Oh, Najoung Kim
ACL (1)7
2026 Investigating Counterfactual Unfairness in LLMs towards Identities through Humor
abstract
Shubin Kim, Yejin Son, Junyeong Park, Keummin Ka, Seungbeen Lee, Jaeyoung Lee, Hyeju Jang, Alice Oh, Youngjae Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shubin Kim, Yejin Son, Junyeong Park, Keummin Ka, Seungbeen Lee, Hyeju Jang, Alice Oh, Youngjae Yu
ACL (1)8
2026 OLA: Output Language Alignment in Code-Switched LLM Interactions
abstract
Code-switching, alternating between languages within a conversation, is natural for multilingual users, yet poses fundamental challenges for large language models (LLMs).When a user code-switches in their prompt to an LLM, they typically do not specify the expected language of the LLM response, and thus LLMs must infer the output language from contextual and pragmatic cues.We find that current LLMs systematically fail to align with this expectation, responding in undesired languages even when cues are clear to humans.We introduce OLA, a benchmark to evaluate LLMs' Output Language Alignment in code-switched interactions.OLA focuses on Korean-English code-switching and spans simple intra-sentential mixing to instruction-content mismatches.Even frontier models frequently misinterpret implicit language expectations, exhibiting a systematic bias toward non-English responses that generalizes beyond Korean to Chinese and Indonesian pairs.However, Code-Switching Aware DPO with minimal data (∼1K examples) substantially reduces misalignment, suggesting these failures stem from insufficient alignment rather than fundamental limitations.Our results highlight the need to align multilingual LLMs with users' implicit expectations in real-world code-switched interactions.
Juhyun Oh, Haneul Yoo, Faiz Ghifari Haznitrama, Alice Oh
ACL (1)4
2026 When Scaffolding Breaks: Investigating Student Interaction with LLM-Based Writing Support in Real-Time K-12 EFL Classrooms
abstract
Large language models (LLMs) are promising tools for scaffolding students’ English writing skills, but their effectiveness in real-time K-12 classrooms remains underexplored. Addressing this gap, our study examines the benefits and limitations of using LLMs as real-time learning support, considering how classroom constraints, such as diverse proficiency levels and limited time, affect their effectiveness. We conducted a deployment study with 157 eighth-grade students in a South Korean middle school English class over six weeks. Our findings reveal that while scaffolding improved students’ ability to compose grammatically correct sentences, this step-by-step approach demotivated lower-proficiency students and increased their system reliance. We also observed challenges to classroom dynamics, where extroverted students often dominated the teacher’s attention, and the system’s assistance made it difficult for teachers to identify struggling students. Based on these findings, we discuss design guidelines for integrating LLMs into real-time writing classes as inclusive educational tools.
Junho Myung, Hyunseung Lim, Hana Oh, Hyoungwook Jin, Nayeon Kang, So-Yeon Ahn, Hwajung Hong, Alice Oh, Juho Kim 0001
CHI8
2026 Open Korean Historical Corpus: A Millennia-Scale Diachronic Collection of Public Domain Texts
abstract
The history of the Korean language is characterized by a discrepancy between its spoken and written forms and a pivotal shift from Chinese characters to the Hangul alphabet. However, this linguistic evolution has remained largely unexplored in NLP due to a lack of accessible historical corpora. To address this gap, we introduce the Open Korean Historical Corpus, a large-scale, openly licensed dataset spanning 1,300 years and 6 languages, as well as under-represented writing systems like Korean-style Sinitic (Idu) and Hanja-Hangul mixed script. This corpus contains 17.7 million documents and 5.1 billion tokens from 19 sources, ranging from the 7th century to 2025. We leverage this resource to quantitatively analyze major linguistic shifts: (1) Idu usage peaked in the 1860s before declining sharply; (2) the transition from Hanja to Hangul was a rapid transformation starting around 1890; and (3) North Korea's lexical divergence causes modern tokenizers to produce up to 51 times higher out-of-vocabulary rates. This work provides a foundational resource for quantitative diachronic analysis by capturing the history of the Korean language. Moreover, it can serve as a pre-training corpus for large language models, potentially improving their understanding of Sino-Korean vocabulary in modern Hangul as well as archaic writing systems.
Seyoung Song 0001, Nawon Kim, Songeun Chae, Kiwoong Park, Jiho Jin, Haneul Yoo, Kyunghyun Cho, Alice Oh
LREC8
2025 Diffusion Models Through a Global Lens: Are They Culturally Inclusive?
abstract
Zahra Bayramli, Ayhan Suleymanzade, Na Min An, Huzama Ahmad, Eunsu Kim, Junyeong Park, James Thorne, Alice Oh. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zahra Bayramli, Ayhan Suleymanzade, Na Min An, Huzama Ahmad, Eunsu Kim, Junyeong Park, James Thorne, Alice Oh
ACL (1)8
2025 XDAC: XAI-Driven Detection and Attribution of LLM-Generated News Comments in Korean
abstract
Large language models (LLMs) generate human-like text, raising concerns about their misuse in creating deceptive content.Detecting LLM-generated comments (LGC) in online news is essential for preserving online discourse integrity and preventing opinion manipulation.However, effective detection faces two key challenges: the brevity and informality of news comments limit traditional detection methods, while the lack of publicly available LGC datasets hinders model development, particularly for non-English languages.To address these challenges, we propose a twofold approach.First, we develop an LGC generation framework to construct a high-quality dataset with diverse and complex examples.Second, we introduce XDAC (XAI-Driven Detection and Attribution of LLM-Generated Comments), a framework utilizing explainable AI, designed for the detection and attribution of short-form LGC in Korean news articles.XDAC leverages XAI to uncover distinguishing linguistic patterns at both token and character levels.We present the first large-scale benchmark dataset, comprising 1.3M human-written comments from Korean news platforms and 1M LLM-generated comments from 14 distinct models.XDAC outperforms existing methods, achieving a 98.5% F1 score in LGC detection with a relative improvement of 68.1%, and an 84.3% F1 score in attribution.To validate real-world applicability, we analyze 5.24M news comments from Naver, South Korea's leading online news platform, identifying 27,029 potential LLM-generated comments.
Hyoungshick Kim, Alice Oh, Yongdae Kim
ACL (1)3
2025 Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
abstract
Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Sebastian Ruder, Wei-Yin Ko, Antoine Bosselut, Alice Oh, Andre Martins, Leshem Choshen, Daphne Ippolito, Enzo Ferrante, Marzieh Fadaee, Beyza Ermis, Sara Hooker. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Sebastian Ruder, Wei-Yin Ko, Antoine Bosselut, Alice Oh, André F. T. Martins, Leshem Choshen, Daphne Ippolito, Enzo Ferrante, Marzieh Fadaee, Beyza Ermis, Sara Hooker
ACL (1)16
2025 DREsS: Dataset for Rubric-based Essay Scoring on EFL Writing
abstract
Automated essay scoring (AES) is a useful tool in English as a Foreign Language (EFL) writing education, offering real-time essay scores for students and instructors.However, previous AES models were trained on essays and scores irrelevant to the practical scenarios of EFL writing education and usually provided a single holistic score due to the lack of appropriate datasets.In this paper, we release DREsS, a large-scale, standard dataset for rubric-based automated essay scoring with 48.9K samples in total.DREsS comprises three sub-datasets: DREsS New , DREsS Std. , and DREsS CASE .We collect DREsS New , a real-classroom dataset with 2.3K essays authored by EFL undergraduate students and scored by English education experts.We also standardize existing rubricbased essay scoring datasets as DREsS Std. .We suggest CASE, a corruption-based augmentation strategy for essays, which generates 40.1K synthetic samples of DREsS CASE and improves the baseline results by 45.44%.DREsS will enable further research to provide a more accurate and practical AES system for EFL writing education.
Haneul Yoo, So-Yeon Ahn, Alice Oh
ACL (1)4
2025 Generalizing Weisfeiler-Lehman Kernels to Subgraphs
abstract
Subgraph representation learning has been effective in solving various real-world problems. However, current graph neural networks (GNNs) produce suboptimal results for subgraph-level tasks due to their inability to capture complex interactions within and between subgraphs. To provide a more expressive and efficient alternative, we propose WLKS, a Weisfeiler-Lehman (WL) kernel generalized for subgraphs by applying the WL algorithm on induced $k$-hop neighborhoods. We combine kernels across different $k$-hop levels to capture richer structural information that is not fully encoded in existing models. Our approach can balance expressiveness and efficiency by eliminating the need for neighborhood sampling. In experiments on eight real-world and synthetic benchmarks, WLKS significantly outperforms leading approaches on five datasets while reducing training time, ranging from 0.01x to 0.25x compared to the state-of-the-art.
Dongkwan Kim 0006, Alice Oh
ICLR2
2025 WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
abstract
Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Wang Yutong, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, Anar Rzayev, Anirban Das, Ashmari Pramodya, Aulia Adila, Bryan Wilie, Candy Olivia Mawalim, Cheng Ching Lam, Daud Abolade, Emmanuele Chersoni, Enrico Santus, Fariz Ikhwantri, Garry Kuwanto, Hanyang Zhao, Haryo Akbarianto Wibowo, Holy Lovenia, Jan Christian Blaise Cruz, Jan Wira Gotama Putra, Junho Myung, Lucky Susanto, Maria Angelica Riera Machin, Marina Zhukova, Michael Anugraha, Muhammad Farid Adilazuarda, Natasha Christabelle Santosa, Peerat Limkonchotiwat, Raj Dabre, Rio Alexander Audino, Samuel Cahyawijaya, Shi-Xiong Zhang, Stephanie Yulia Salim, Yi Zhou, Yinxuan Gui, David Ifeoluwa Adelani, En-Shiun Annie Lee, Shogo Okada, Ayu Purwarianti, Alham Fikri Aji, Taro Watanabe, Derry Tanti Wijaya, Alice Oh, Chong-Wah Ngo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, Anar Rzayev, Ashmari Pramodya, Aulia Adila, Bryan Wilie, Candy Olivia Mawalim, Cheng Ching Lam, Daud Abolade, Emmanuele Chersoni, Enrico Santus, Fariz Ikhwantri, Garry Kuwanto, Hanyang Zhao, Haryo Akbarianto Wibowo, Holy Lovenia, Jan Christian Blaise Cruz, Jan Wira Gotama Putra, Junho Myung, Lucky Susanto, Maria Angelica Riera Machin, Marina Zhukova, Michael Anugraha, Muhammad Farid Adilazuarda, Natasha Christabelle Santosa, Peerat Limkonchotiwat, Raj Dabre, Rio Alexander Audino, Samuel Cahyawijaya, Stephanie Yulia Salim, Yi Zhou 0019, Yinxuan Gui, David Ifeoluwa Adelani, Annie En-Shiun Lee, Shogo Okada, Ayu Purwarianti, Alham Fikri Aji, Taro Watanabe, Derry Wijaya, Alice Oh, Chong-Wah Ngo
NAACL (Long Papers)50
2025 Trans-EnV: A Framework for Evaluating the Linguistic Robustness of LLMs Against English Varieties
abstract
Large Language Models (LLMs) are predominantly evaluated on Standard American English (SAE), often overlooking the diversity of global English varieties.This narrow focus may raise fairness concerns as degraded performance on non-standard varieties can lead to unequal benefits for users worldwide.Therefore, it is critical to extensively evaluate the linguistic robustness of LLMs on multiple non-standard English varieties.We introduce Trans-EnV, a framework that automatically transforms SAE datasets into multiple English varieties to evaluate the linguistic robustness. Our framework combines (1) linguistics expert knowledge to curate variety-specific features and transformation guidelines from linguistic literature and corpora, and (2) LLM-based transformations to ensure both linguistic validity and scalability.Using Trans-EnV, we transform six benchmark datasets into 38 English varieties and evaluate seven state-of-the-art LLMs.Our results reveal significant performance disparities, with accuracy decreasing by up to 46.3% on non-standard varieties.These findings highlight the importance of comprehensive linguistic robustness evaluation across diverse English varieties. Each construction of Trans-EnV was validated through rigorous statistical testing and consultation with a researcher in the field of second language acquisition, ensuring its linguistic validity.Our code and datasets are publicly available.
Seungho Kim, Jun-Min Lee, Kitaek Kim, Alice Oh, Edward Choi 0003
NeurIPS6
2024 RECIPE4U: Student-ChatGPT Interaction Dataset in EFL Writing Education
abstract
The integration of generative AI in education is expanding, yet empirical analyses of large-scale and real-world interactions between students and AI systems still remain limited. Addressing this gap, we present RECIPE4U (RECIPE for University), a dataset sourced from a semester-long experiment with 212 college students in English as Foreign Language (EFL) writing courses. During the study, students engaged in dialogues with ChatGPT to revise their essays. RECIPE4U includes comprehensive records of these interactions, including conversation logs, students’ intent, students’ self-rated satisfaction, and students’ essay edit histories. In particular, we annotate the students’ utterances in RECIPE4U with 13 intention labels based on our coding schemes. We establish baseline results for two subtasks in task-oriented dialogue systems within educational contexts: intent detection and satisfaction estimation. As a foundational step, we explore student-ChatGPT interaction patterns through RECIPE4U and analyze them by focusing on students’ dialogue, essay data statistics, and students’ essay edits. We further illustrate potential applications of RECIPE4U dataset for enhancing the incorporation of LLMs in educational frameworks. RECIPE4U is publicly available at https://zeunie.github.io/RECIPE4U/.
Haneul Yoo, Junho Myung, Minsun Kim, Tak Yeon Lee, So-Yeon Ahn, Alice Oh
LREC/COLING7
2024 CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean
abstract
Despite the rapid development of large language models (LLMs) for the Korean language, there remains an obvious lack of benchmark datasets that test the requisite Korean cultural and linguistic knowledge. Because many existing Korean benchmark datasets are derived from the English counterparts through translation, they often overlook the different cultural contexts. For the few benchmark datasets that are sourced from Korean data capturing cultural knowledge, only narrow tasks such as hate speech detection are offered. To address this gap, we introduce a benchmark of Cultural and Linguistic Intelligence in Korean (CLIcK), a dataset comprising 1,995 QA pairs. CLIcK sources its data from official Korean exams and textbooks, partitioning the questions into eleven categories under the two main categories of language and culture. For each instance in click, we provide fine-grained annotation of which cultural and linguistic knowledge is required to correctly answer the question. Using CLIcK, we test 13 language models to assess their performance. Our evaluation uncovers insights into their performances across the categories, as well as the diverse factors affecting their comprehension. CLIcK offers the first large-scale comprehensive Korean-centric analysis of LLMs’ proficiency in Korean language and culture.
Eunsu Kim, Juyoung Suk, Philhoon Oh, Haneul Yoo, James Thorne, Alice Oh
LREC/COLING6
2024 Perceptions to Beliefs: Exploring Precursory Inferences for Theory of Mind in Large Language Models
abstract
While humans naturally develop theory of mind (ToM), the capability to understand other people's mental states and beliefs, state-ofthe-art large language models (LLMs) underperform on simple ToM benchmarks.We posit that we can extend our understanding of LLMs' ToM abilities by evaluating key human ToM precursors-perception inference and perception-to-belief inference-in LLMs.We introduce two datasets, Percept-ToMi and Percept-FANToM, to evaluate these precursory inferences for ToM in LLMs by annotating characters' perceptions on ToMi and FANToM, respectively.Our evaluation of eight state-ofthe-art LLMs reveals that the models generally perform well in perception inference while exhibiting limited capability in perception-tobelief inference (e.g., lack of inhibitory control).Based on these results, we present PercepToM, a novel ToM method leveraging LLMs' strong perception inference capability while supplementing their limited perception-to-belief inference.Experimental results demonstrate that PercepToM significantly enhances LLM's performance, especially in false belief scenarios.
Chani Jung, Dongkwan Kim 0006, Jiho Jin, Jiseon Kim, Yeon Seonwoo, Yejin Choi 0001, Alice Oh, Hyunwoo Kim 0002
EMNLP7
2024 Can LLM Generate Culturally Relevant Commonsense QA Data? Case Study in Indonesian and Sundanese
abstract
Large Language Models (LLMs) are increasingly being used to generate synthetic data for training and evaluating models.However, it is unclear whether they can generate a good quality of question answering (QA) dataset that incorporates knowledge and cultural nuance embedded in a language, especially for low-resource languages.In this study, we investigate the effectiveness of using LLMs in generating culturally relevant commonsense QA datasets for Indonesian and Sundanese languages.To do so, we create datasets for these languages using various methods involving both LLMs and human annotators, resulting in ∼4.5K questions per language (∼9K in total), making our dataset the largest of its kind.Our experiments show that automatic data adaptation from an existing English dataset is less effective for Sundanese.Interestingly, using the direct generation method on the target language, GPT-4 Turbo can generate questions with adequate general knowledge in both languages, albeit not as culturally 'deep' as humans.We also observe a higher occurrence of fluency errors in the Sundanese dataset, highlighting the discrepancy between medium-and lower-resource languages.1
Rifki Afina Putri, Faiz Ghifari Haznitrama, Dea Adhista, Alice Oh
EMNLP4
2024 Translating Subgraphs to Nodes Makes Simple GNNs Strong and Efficient for Subgraph Representation Learning
abstract
Subgraph representation learning has emerged as an important problem, but it is by default approached with specialized graph neural networks on a large global graph. These models demand extensive memory and computational resources but challenge modeling hierarchical structures of subgraphs. In this paper, we propose Subgraph-To-Node (S2N) translation, a novel formulation for learning representations of subgraphs. Specifically, given a set of subgraphs in the global graph, we construct a new graph by coarsely transforming subgraphs into nodes. Demonstrating both theoretical and empirical evidence, S2N not only significantly reduces memory and computational costs compared to state-of-the-art models but also outperforms them by capturing both local and global structures of the subgraph. By leveraging graph coarsening methods, our method outperforms baselines even in a data-scarce setting with insufficient subgraphs. Our experiments on eight benchmarks demonstrate that fined-tuned models with S2N translation can process 183 – 711 times more subgraph samples than state-of-the-art models at a better or similar performance level.
Dongkwan Kim 0006, Alice Oh
ICML2
2024 Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis
abstract
Nayeon Lee, Chani Jung, Junho Myung, Jiho Jin, Jose Camacho-Collados, Juho Kim, Alice Oh. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Nayeon Lee, Chani Jung, Junho Myung, Jiho Jin, José Camacho-Collados, Juho Kim 0001, Alice Oh
NAACL-HLT7
2024 BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages
abstract
Large language models (LLMs) often lack culture-specific everyday knowledge, especially across diverse regions and non-English languages. Existing benchmarks for evaluating LLMs' cultural sensitivities are usually limited to a single language or online sources like Wikipedia, which may not reflect the daily habits, customs, and lifestyles of different regions. That is, information about the food people eat for their birthday celebrations, spices they typically use, musical instruments youngsters play or the sports they practice in school is not always explicitly written online. To address this issue, we introduce BLEnD, a hand-crafted benchmark designed to evaluate LLMs' everyday knowledge across diverse cultures and languages. The benchmark comprises 52.6k question-answer pairs from 16 countries/regions, in 13 different languages, including low-resource ones such as Amharic, Assamese, Azerbaijani, Hausa, and Sundanese. We evaluate LLMs in two formats: short-answer questions, and multiple-choice questions. We show that LLMs perform better in cultures that are more present online, with a maximum 57.34% difference in GPT-4, the best-performing model, in the short-answer format.Furthermore, we find that LLMs perform better in their local languages for mid-to-high-resource languages. Interestingly, for languages deemed to be low-resource, LLMs provide better answers in English. We make our dataset publicly available at: https://github.com/nlee0212/BLEnD.
Junho Myung, Nayeon Lee, Yi Zhou 0019, Jiho Jin, Rifki Afina Putri, Dimosthenis Antypas, Hsuvas Borkakoty, Eunsu Kim, Carla Pérez-Almendros, Abinew Ali Ayele, Víctor Gutiérrez-Basulto, Yazmín Ibáñez-García, Hwaran Lee, Shamsuddeen Hassan Muhammad, Ki-Woong Park, Anar Rzayev, Nina White, Seid Muhie Yimam, Mohammad Taher Pilehvar, Nedjma Ousidhoum, José Camacho-Collados, Alice Oh
NeurIPS22
2024 KoBBQ: Korean Bias Benchmark for Question Answering
abstract
Abstract Warning: This paper contains examples of stereotypes and biases. The Bias Benchmark for Question Answering (BBQ) is designed to evaluate social biases of language models (LMs), but it is not simple to adapt this benchmark to cultural contexts other than the US because social biases depend heavily on the cultural context. In this paper, we present KoBBQ, a Korean bias benchmark dataset, and we propose a general framework that addresses considerations for cultural adaptation of a dataset. Our framework includes partitioning the BBQ dataset into three classes—Simply-Transferred (can be used directly after cultural translation), Target-Modified (requires localization in target groups), and Sample-Removed (does not fit Korean culture)—and adding four new categories of bias specific to Korean culture. We conduct a large-scale survey to collect and validate the social biases and the targets of the biases that reflect the stereotypes in Korean culture. The resulting KoBBQ dataset comprises 268 templates and 76,048 samples across 12 categories of social bias. We use KoBBQ to measure the accuracy and bias scores of several state-of-the-art multilingual LMs. The results clearly show differences in the bias of LMs as measured by KoBBQ and a machine-translated version of BBQ, demonstrating the need for and utility of a well-constructed, culturally aware social bias benchmark.
Jiho Jin, Jiseon Kim, Nayeon Lee, Haneul Yoo, Alice Oh, Hwaran Lee
Trans. Assoc. Comput. Linguistics5
2023 SQuARe: A Large-Scale Dataset of Sensitive Questions and Acceptable Responses Created through Human-Machine Collaboration
abstract
Hwaran Lee, Seokhee Hong, Joonsuk Park, Takyoung Kim, Meeyoung Cha, Yejin Choi, Byoungpil Kim, Gunhee Kim, Eun-Ju Lee, Yong Lim, Alice Oh, Sangchul Park, Jung-Woo Ha. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Hwaran Lee, Seokhee Hong 0002, Joonsuk Park, Takyoung Kim, Meeyoung Cha, Yejin Choi 0001, Byoung Pil Kim, Gunhee Kim, Eun-Ju Lee 0001, Yong Lim, Alice Oh, Sangchul Park, Jung-Woo Ha 0001
ACL (1)11
2023 Ranking-Enhanced Unsupervised Sentence Representation Learning
abstract
Yeon Seonwoo, Guoyin Wang, Changmin Seo, Sajal Choudhary, Jiwei Li, Xiang Li, Puyang Xu, Sunghyun Park, Alice Oh. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yeon Seonwoo, Guoyin Wang 0002, Changmin Seo, Sajal Choudhary, Jiwei Li 0001, Puyang Xu, Alice Oh
ACL (1)9
2023 Rethinking Annotation: Can Language Learners Contribute?
abstract
Haneul Yoo, Rifki Afina Putri, Changyoon Lee, Youngin Lee, So-Yeon Ahn, Dongyeop Kang, Alice Oh. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Haneul Yoo, Rifki Afina Putri, Changyoon Lee, Youngin Lee, So-Yeon Ahn, Dongyeop Kang, Alice Oh
ACL (1)7
2023 Towards standardizing Korean Grammatical Error Correction: Datasets and Annotation
abstract
Soyoung Yoon, Sungjoon Park, Gyuwan Kim, Junhee Cho, Kihyo Park, Gyu Tae Kim, Minjoon Seo, Alice Oh. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Soyoung Yoon, Gyuwan Kim, Kihyo Park, Gyu Tae Kim, Minjoon Seo, Alice Oh
ACL (1)8
2023 The Role of Gender in Students' Privacy Concerns about Learning Analytics: Evidence from five countries
abstract
The protection of students’ privacy in learning analytics (LA) applications is critical for cultivating trust and effective implementations of LA in educational environments around the world. However, students’ privacy concerns and how they may vary along demographic dimensions that historically influence these concerns have yet to be studied in higher education. Gender differences, in particular, are known to be associated with people's information privacy concerns, including in educational settings. Building on an empirically validated model and survey instrument for student privacy concerns, their antecedents and their behavioral outcomes, we investigate the presence of gender differences in students’ privacy concerns about LA. We conducted a survey study of students in higher education across five countries (N = 762): Germany, South Korea, Spain, Sweden and the United States. Using multiple regression analysis, across all five countries, we find that female students have stronger trusting beliefs and they are more inclined to engage in self-disclosure behaviors compared to male students. However, at the country level, these gender differences are significant only in the German sample, for Bachelor's degree students, and for students between the ages of 18 and 24. Thus, national context, degree program, and age are important moderating factors for gender differences in student privacy concerns.
René F. Kizilcec, Olga Viberg, Ioana Jivet, Alejandra Martínez-Monés, Alice Oh, Stefan Hrastinski, Chantal Mutimukwe, Maren Scheffel
LAK5
2023 RECIPE: How to Integrate ChatGPT into EFL Writing Education
abstract
The integration of generative AI in the field of education is actively being explored. In particular, ChatGPT has garnered significant interest, offering an opportunity to examine its effectiveness in English as a foreign language (EFL) education. To address this need, we present a novel learning platform called RECIPE (Revising an Essay with ChatGPT on an Interactive Platform for EFL learners). Our platform features two types of prompts that facilitate conversations between ChatGPT and students: (1) a hidden prompt for ChatGPT to take an EFL teacher role and (2) an open prompt for students to initiate a dialogue with a self-written summary of what they have learned. We deployed this platform for 213 undergraduate and graduate students enrolled in EFL writing courses and seven instructors. For this study, we collect students' interaction data from RECIPE, including students' perceptions and usage of the platform, and user scenarios are examined with the data. We also conduct a focus group interview with six students and an individual interview with one EFL instructor to explore design opportunities for leveraging generative AI models in the field of EFL education.
Haneul Yoo, Yoonsu Kim, Junho Myung, Minsun Kim, Hyunseung Lim, Juho Kim 0001, Tak Yeon Lee, Hwajung Hong, So-Yeon Ahn, Alice Oh
L@S11
2023 EliRank: A Code Editing History Based Ranking Model for Early Detection of Students in Need
abstract
Research on programming education shows that novice programming students benefit significantly from one-to-one tutoring. While many systems propose to replicate the effectiveness of one-to-one tutoring in large-scale classes, it remains a challenge to develop systems with an approach to finding students who need the tutors' help the most. In this paper, we explore the idea of predicting the priority of students in need with a data-driven approach. Among various metrics to calculate the priority of students in need, we adopt time-on-task metric. Previous studies have found that excessively long time-on-task can be used as an indication of students' struggling. Aligned with this, we reduce the problem of finding students with the highest priority to the problem of finding students with the longest time-on-task. To solve the reduced problem, we present EliRank, a ranking model that finds students with the longest estimated time-on-task, using the students' first few minutes of fine-grained code editing history. EliRank recommends students in the descending order of estimated time-on-task, enabling tutors to efficiently monitor and find the students in need at scale in real time. To evaluate the performance of EliRank, we build and publish a new real-world dataset consisting of 15 programming exercises solved by 4000+ students in an introduction to programming class at a university. Unlike the currently available open code editing history datasets, our dataset contains code editing operations at a character-level granularity to minimize the loss of contextual information from students. We also introduce diff-augmented abstract syntax tree (DAST), a novel structured code representation that minimizes the loss of fine-grained code change information during code parsing. The evaluation of EliRank on our dataset shows that EliRank effectively finds students with the longest estimated time-on-task, for early detection of students in need. Also, we illustrate in depth (i) the effectiveness of DAST, (ii) the potential to control the tradeoff between early detection and the prediction accuracy of the model, and (iii) the transferability to unseen programming exercise via zero-shot transfer learning.
Jungkook Park, Alice Oh
L@S2
2022 Models and Benchmarks for Representation Learning of Partially Observed Subgraphs
abstract
Subgraphs are rich substructures in graphs, and their nodes and edges can be partially observed in real-world tasks. Under partial observation, existing node- or subgraph-level message-passing produces suboptimal representations. In this paper, we formulate a novel task of learning representations of partially observed subgraphs. To solve this problem, we propose Partial Subgraph InfoMax (PSI) framework and generalize existing InfoMax models, including DGI, InfoGraph, MVGRL, and GraphCL, into our framework. These models maximize the mutual information between the partial subgraph's summary and various substructures from nodes to full subgraphs. In addition, we suggest a novel two-stage model with k-hop PSI, which reconstructs the representation of the full subgraph and improves its expressiveness from different local-global structures. Under training and evaluation protocols designed for this problem, we conduct experiments on three real-world datasets and demonstrate that PSI models outperform baselines.
Dongkwan Kim 0006, Jiho Jin, Jaimeen Ahn, Alice Oh
CIKM4
2022 Virtual Knowledge Graph Construction for Zero-Shot Domain-Specific Document Retrieval
abstract
Domain-specific documents cover terminologies and specialized knowledge. This has been the main challenge of domain-specific document retrieval systems. Previous approaches propose domain-adaptation and transfer learning methods to alleviate this problem. However, these approaches still follow the same document representation method in previous approaches; a document is embedded into a single vector. In this study, we propose VKGDR. VKGDR represents a given corpus into a graph of entities and their relations (known as a virtual knowledge graph) and computes the relevance between queries and documents based on the graph representation. We conduct three experiments 1) domain-specific document retrieval, 2) comparison of our virtual knowledge graph construction method with previous approaches, and 3) ablation study on each component of our virtual knowledge graph. From the results, we see that unsupervised VKGDR outperforms baselines in a zero-shot setting and even outperforms fully-supervised bi-encoder. We also verify that our virtual knowledge graph construction method results in better retrieval performance than previous approaches.
Yeon Seonwoo, Seunghyun Yoon 0002, Franck Dernoncourt, Trung Bui, Alice Oh
COLING5
2022 KOLD: Korean Offensive Language Dataset
abstract
Warning: this paper contains content that may be offensive or upsetting.Recent directions for offensive language detection are hierarchical modeling, identifying the type and the target of offensive language, and interpretability with offensive span annotation and prediction.These improvements are focused on English and do not transfer well to other languages because of cultural and linguistic differences.In this paper, we present the Korean Offensive Language Dataset (KOLD) comprising 40,429 comments, which are annotated hierarchically with the type and the target of offensive language, accompanied by annotations of the corresponding text spans.We collect the comments from NAVER news and YouTube platform and provide the titles of the articles and videos as the context information for the annotation process.We use these annotated comments as training data for Korean BERT and RoBERTa models and find that they are effective at offensiveness detection, target classification, and target span detection while having room for improvement for target group classification and offensive span detection.We discover that the target group distribution differs drastically from the existing English datasets, and observe that providing the context information improves the model performance in offensiveness detection (+0.3), target classification (+1.5), and target group classification (+13.1).We publicly release the dataset and baseline models.
Younghoon Jeong, Juhyun Oh, Jaimeen Ahn, Jihyung Moon, Alice Oh
EMNLP7
2022 IDK-MRC: Unanswerable Questions for Indonesian Machine Reading Comprehension
abstract
Machine Reading Comprehension (MRC) has become one of the essential tasks in Natural Language Understanding (NLU) as it is often included in several NLU benchmarks (Liang et al., 2020;Wilie et al., 2020).However, most MRC datasets only have answerable question type, overlooking the importance of unanswerable questions.MRC models trained only on answerable questions will select the span that is most likely to be the answer, even when the answer does not actually exist in the given passage (Rajpurkar et al., 2018).This problem especially remains in medium-to low-resource languages like Indonesian.Existing Indonesian MRC datasets (Purwarianti et al., 2007;Clark et al., 2020) are still inadequate because of the small size and limited question types, i.e., they only cover answerable questions.To fill this gap, we build a new Indonesian MRC dataset called I(n)don'tKnow-MRC (IDK-MRC) by combining the automatic and manual unanswerable question generation to minimize the cost of manual dataset construction while maintaining the dataset quality.Combined with the existing answerable questions, IDK-MRC consists of more than 10K questions in total.Our analysis shows that our dataset significantly improves the performance of Indonesian MRC models, showing a large improvement for unanswerable questions 1 .
Rifki Afina Putri, Alice Oh
EMNLP2
2022 CS1QA: A Dataset for Assisting Code-based Question Answering in an Introductory Programming Course
abstract
We introduce CS1QA, a dataset for code-based question answering in the programming education domain.CS1QA consists of 9,237 question-answer pairs gathered from chat logs in an introductory programming class using Python, and 17,698 unannotated chat data with code 1 .Each question is accompanied with the student's code, and the portion of the code relevant to answering the question.We carefully design the annotation process to construct CS1QA, and analyze the collected dataset in detail.The tasks for CS1QA are to predict the question type, the relevant code snippet given the question and the code and retrieving an answer from the annotated corpus.Results for the experiments on several baseline models are reported and thoroughly analyzed.The tasks for CS1QA challenge models to understand both the code and natural language.This unique dataset can be used as a benchmark for source code comprehension and question answering in the educational setting.
Changyoon Lee, Yeon Seonwoo, Alice Oh
NAACL-HLT3
2021 Mitigating Language-Dependent Ethnic Bias in BERT
abstract
BERT and other large-scale language models (LMs) contain gender and racial bias.They also exhibit other dimensions of social bias, most of which have not been studied in depth, and some of which vary depending on the language.In this paper, we study ethnic bias and how it varies across languages by analyzing and mitigating ethnic bias in monolingual BERT for English, German, Spanish, Korean, Turkish, and Chinese.To observe and quantify ethnic bias, we develop a novel metric called Categorical Bias score.Then we propose two methods for mitigation; first using a multilingual model, and second using contextual word alignment of two monolingual models.We compare our proposed methods with monolingual BERT and show that these methods effectively alleviate the ethnic bias.Which of the two methods works better depends on the amount of NLP resources available for that language.We additionally experiment with Arabic and Greek to verify that our proposed methods work for a wider variety of languages. EN-1: A person from [MASK] is an enemy.1. America (0.09) 2. Iraq (0.08)
Jaimeen Ahn, Alice Oh
EMNLP (1)2
2021 Learning Bill Similarity with Annotated and Augmented Corpora of Bills
abstract
Bill writing is a critical element of representative democracy.However, it is often overlooked that most legislative bills are derived, or even directly copied, from other bills.Despite the significance of bill-to-bill linkages for understanding the legislative process, existing approaches fail to address semantic similarities across bills, let alone reordering or paraphrasing which are prevalent in legal document writing.In this paper, we overcome these limitations by proposing a 5-class classification task that closely reflects the nature of the bill generation process.In doing so, we construct a human-labeled dataset of 4,721 billto-bill relationships at the subsection-level and release this annotated dataset to the research community.To augment the dataset, we generate synthetic data with varying degrees of similarity, mimicking the complex bill writing process.We use BERT variants and apply multi-stage training, sequentially fine-tuning our models with synthetic and human-labeled datasets.We find that the predictive performance significantly improves when training with both human-labeled and synthetic data.Finally, we apply our trained model to infer section-and bill-level similarities.Our analysis shows that the proposed methodology successfully captures the similarities across legal documents at various levels of aggregation. 1
Jiseon Kim, Elden Griggs, In Song Kim, Alice Oh
EMNLP (1)4
2021 Dimensional Emotion Detection from Categorical Emotion
abstract
We present a model to predict fine-grained emotions along the continuous dimensions of valence, arousal, and dominance (VAD) with a corpus with categorical emotion annotations.Our model is trained by minimizing the EMD (Earth Mover's Distance) loss between the predicted VAD score distribution and the categorical emotion distributions sorted along VAD, and it can simultaneously classify the emotion categories and predict the VAD scores for a given sentence.We use pre-trained RoBERTa-Large and fine-tune on three different corpora with categorical labels and evaluate on EmoBank corpus with VAD scores.We show that our approach reaches comparable performance to that of the state-of-the-art classifiers in categorical emotion classification and shows significant positive correlations with the ground truth VAD scores.Also, further training with supervision of VAD labels leads to improved performance especially when dataset is small.We also present examples of predictions of appropriate emotion words that are not part of the original annotations.
Jiseon Kim, Seonghyeon Ye, Jaeyeol Jeon, Heeyoung Park, Alice Oh
EMNLP (1)6
2021 Efficient Contrastive Learning via Novel Data Augmentation and Curriculum Learning
abstract
We introduce EfficientCL, a memory-efficient continual pretraining method that applies contrastive learning with novel data augmentation and curriculum learning.For data augmentation, we stack two types of operation sequentially: cutoff and PCA jittering.While pretraining steps proceed, we apply curriculum learning by incrementing the augmentation degree for each difficulty step.After data augmentation, we apply contrastive learning on projected embeddings of original and augmented examples.When fine-tuned on GLUE benchmark, our model outperforms baseline models, especially for sentence-level tasks.Additionally, this improvement is achieved with only 70% of computational memory compared to the baseline model. 1
Seonghyeon Ye, Jiseon Kim, Alice Oh
EMNLP (1)3
2021 How to Find Your Friendly Neighborhood: Graph Attention Design with Self-Supervision
Dongkwan Kim 0006, Alice Oh
ICLR2
2021 Emergent Communication under Varying Sizes and Connectivities
abstract
Recent advances in deep neural networks allowed artificial agents to derive their own emergent languages that promote interaction, coordination, and collaboration within a group. Just as we humans have succeeded in creating a shared language that allows us to interact within a large group, can the emergent communication within an artificial group converge to a shared, agreed language? This research provides an analytical study of the shared emergent language within the group communication settings of different sizes and connectivities. As the group size increases up to hundreds, agents start to speak dissimilar languages, but the rate at which they successfully communicate is maintained. We observe the emergence of different dialects when we restrict the group communication to have local connectivities only. Finally, we provide optimization results of group communication graphs when the number of agents one can communicate with is restricted or when we penalize communication between distant agent pairs. The optimized communication graphs show superior communication success rates compared to graphs with same number of links as well as the emergence of hub nodes and scale-free networks.
Alice Oh
NeurIPS2
2021 Pythonpad: Server-free Python Hands-on Exercise for Online Programming Classes
abstract
We propose Pythonpad, an open-source JavaScript library that supports web-based Python programming exercises. Unlike other standalone web-based programming tools, Pythonpad can be easily integrated into other websites. Although it runs learners' Python code in client-side web browsers, Pythonpad supports a file system, building and importing external modules, and many essential built-in Python libraries to teach basic programming concepts in CS1 classes.
Jeongmin Byun, Jungkook Park, Alice Oh
SIGCSE3
2021 Cocode: Providing Social Presence with Co-learner Screen Sharing in Online Programming Classes
abstract
Social presence is known to be important for distance education, and a common approach in online classes is to provide chat boxes and forums to provide the social presence. In such a class, however, learners must explicitly act beyond their normal learning activities, so often there is no social presence in the class even when there are several learners working on the same course material. In this paper, we develop an approach where learners can share the social presence without any explicit action; their normal learning activities would be used to provide visual cues for social presence. We present Cocode, a system designed for an online programming class that shows other learners' code editors and running output in the programming environment with minimum privacy issues. For evaluation, we ran two user studies with groups of participants who took an offline class and an online programming class from the university; results from the studies showed that learners felt less social presence in Cocode than in offline classes, but they felt significantly more social presence in Cocode than in online classes with live video lectures, forums, and chat sessions.
Jeongmin Byun, Jungkook Park, Alice Oh
Proc. ACM Hum. Comput. Interact.3
2020 Speaker Sensitive Response Evaluation Model
abstract
Automatic evaluation of open-domain dialogue response generation is very challenging because there are many appropriate responses for a given context.Existing evaluation models merely compare the generated response with the ground truth response and rate many of the appropriate responses as inappropriate if they deviate from the ground truth.One approach to resolve this problem is to consider the similarity of the generated response with the conversational context.In this paper, we propose an automatic evaluation model based on that idea and learn the model parameters from an unlabeled conversation corpus.Our approach considers the speakers in defining the different levels of similar context.We use a Twitter conversation corpus that contains many speakers and conversations to test our evaluation model.Experiments show that our model outperforms the other existing evaluation metrics in terms of high correlation with human annotation scores.We also show that our model trained on Twitter can be applied to movie dialogues without any additional training.We provide our code and the learned parameters so that they can be used for automatic evaluation of dialogue response generation models.
JinYeong Bak, Alice Oh
ACL2
2020 Suicidal Risk Detection for Military Personnel
abstract
We analyze social media for detecting the suicidal risk of military personnel, which is especially crucial for countries with compulsory military service such as the Republic of Korea.From a widely-used Korean social Q&A site, we collect posts containing military-relevant content written by active-duty military personnel.We then annotate the posts with two groups of experts: military experts and mental health experts.Our dataset includes 2,791 posts with 13,955 corresponding expert annotations of suicidal risk levels, and this dataset is available to researchers who consent to research ethics agreement.Using various finetuned state-of-the-art language models, we predict the level of suicide risk, reaching .88F1 score for classifying the risks.
Ki-Woong Park, Jaimeen Ahn, Alice Oh
EMNLP (1)4
2020 Context-Aware Answer Extraction in Question Answering
abstract
Extractive QA models have shown very promising performance in predicting the correct answer to a question for a given passage.However, they sometimes result in predicting the correct answer text but in a context irrelevant to the given question.This discrepancy becomes especially important as the number of occurrences of the answer text in a passage increases.To resolve this issue, we propose BLANC (BLock AttentioN for Context prediction) based on two main ideas: context prediction as an auxiliary task in multi-task learning manner, and a block attention method that learns the context prediction task.With experiments on reading comprehension, we show that BLANC outperforms the state-ofthe-art QA models, and the performance gap increases as the number of answer text occurrences increases.We also conduct an experiment of training the models using SQuAD and predicting the supporting facts on HotpotQA and show that BLANC outperforms all baseline models in this zero-shot setting.
Yeon Seonwoo, Jung-Woo Ha 0001, Alice Oh
EMNLP (1)4
2020 Detecting Contract Cheaters in Online Programming Classes with Keystroke Dynamics
abstract
In online programming classes, it is tricky to uphold academic honesty in the assessment process. A common approach, plagiarism detection, is not accurate for novice programmers and ineffective for detecting contract cheaters. We present a new approach, cheating detection with keystroke dynamics in programming classes, and evaluated the approach.
Jeongmin Byun, Jungkook Park, Alice Oh
L@S3
2020 Denoising Recurrent Neural Networks for Classifying Crash-Related Events
abstract
With detailed sensor and visual data from automobiles, a data-driven model can learn to classify crash-related events during a drive. We propose a neural network model accepting time-series vehicle sensor data and forward-facing videos as input for learning classification of crash-related events and varying types of such events. To elaborate, a novel recurrent neural network structure is introduced, namely, denoising gated recurrent unit with decay, in order to deal with time-series automobile sensor data with missing value and noises. Our model detects crash and near-crash events based on a large set of time-series data collected from naturalistic driving behavior. Furthermore, the model classifies those events involving pedestrians, a vehicle in front, or a vehicle on either side. The effectiveness of our model is evaluated with more than two thousand 30-s clips from naturalistic driving behavior data. The results show that the model, including sensory encoder with denoising gated recurrent unit with decay, visual encoder, and attention mechanism, outperforms gated recurrent unit with decay, gated CNN, and other baselines not only in event classification and but also in event-type classification.
Yeon Seonwoo, Jiseon Kim, Alice Oh
IEEE Trans. Intell. Transp. Syst.5
2019 Variational Hierarchical User-based Conversation Model
abstract
JinYeong Bak, Alice Oh. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
JinYeong Bak, Alice Oh
EMNLP/IJCNLP (1)2
2019 automaTA: Human-Machine Interaction for Answering Context-Specific Questions
abstract
When online learners have questions that are related to a specific task, they often use Q&A boards instead of web search because they are looking for context-specific answers. While lecturers, teaching assistants, and other learners can provide context-specific answers on the Q&A boards, there is often a high response latency which can impede their learning. We present automaTA, a prototype that suggests context-specific answers to online learners' questions by capturing the context of the questions. Our solution is to automate the response generation with a human-machine mixed approach, where humans generate high-quality answers, and the human-generated responses are used to train an automated algorithm to provide context-specific answers. automaTA adopts this approach as a prototype in which it generates automated answers for function-related questions in an online programming course. We conduct two user studies with undergraduate and graduate students with little or no experience with Python and found the potential that automaTA can automatically provide answers to context-specific questions without a human instructor, at scale.
Changyoon Lee, Donghoon Han, Hyoungwook Jin, Alice Oh
L@S4
2019 Homogeneity-Based Transmissive Process to Model True and False News in Social Networks
abstract
An overwhelming number of true and false news stories are posted and shared in social networks, and users diffuse the stories based on multiple factors. Diffusion of news stories from one user to another depends not only on the stories' content and the genuineness but also on the alignment of the topical interests between the users. In this paper, we propose a novel Bayesian nonparametric model that incorporates homogeneity of news stories as the key component that regulates the topical similarity between the posting and sharing users' topical interests. Our model extends hierarchical Dirichlet process to model the topics of the news stories and incorporates Bayesian Gaussian process latent variable model to discover the homogeneity values. We train our model on a real-world social network dataset and find homogeneity values of news stories that strongly relate to their labels of genuineness and their contents. Finally, we show that the supervised version of our model predicts the labels of news stories better than the state-of-the-art neural network and Bayesian models.
Dongkwan Kim 0006, Alice Oh
WSDM3
2018 Subword-level Word Vector Representations for Korean
abstract
Research on distributed word representations is focused on widely-used languages such as English.Although the same methods can be used for other languages, language-specific knowledge can enhance the accuracy and richness of word vector representations.In this paper, we look at improving distributed word representations for Korean using knowledge about the unique linguistic structure of Korean.Specifically, we decompose Korean words into the jamo level, beyond the characterlevel, allowing a systematic use of subword information.To evaluate the vectors, we develop Korean test sets for word similarity and analogy and make them publicly available.The results show that our simple method outperforms word2vec and character-level Skip-Grams on semantic and syntactic similarity and analogy tasks and contributes positively toward downstream NLP tasks such as sentiment analysis.
Jeongmin Byun, Sion Baek, Yongseok Cho, Alice Oh
ACL (1)5
2018 Conversational Decision Making Model for Predicting King's Decision in the Annals of the Joseon Dynasty
abstract
Styles of leaders when they make decisions in groups vary, and the different styles affect the performance of the group.To understand the key words and speakers associated with decisions, we initially formalize the problem as one of predicting leaders' decisions from discussion with group members.As a dataset, we introduce conversational meeting records from a historical corpus, and develop a hierarchical RNN structure with attention and pre-trained speaker embedding in the form of a, Conversational Decision Making Model (CDMM).The CDMM outperforms other baselines to predict leaders' final decisions from the data.We explain why CDMM works better than other methods by showing the key words and speakers discovered from the attentions as evidence.
JinYeong Bak, Alice Oh
EMNLP2
2018 Hierarchical Dirichlet Gaussian Marked Hawkes Process for Narrative Reconstruction in Continuous Time Domain
abstract
In news and discussions, many articles and posts are provided without their related previous articles or posts.Hence, it is difficult to understand the context from which the articles and posts have occurred.In this paper, we propose the Hierarchical Dirichlet Gaussian Marked Hawkes process (HD-GMHP) for reconstructing the narratives and thread structures of news articles and discussion posts.HD-GMHP unifies three modeling strategies in previous research: temporal characteristics, triggering event relations, and meta information of text in news articles and discussion threads.To show the effectiveness of the model, we perform experiments in narrative reconstruction and thread reconstruction with real world datasets: articles from the New York Times and a corpus of Wikipedia conversations.The experimental results show that HD-GMHP outperforms the baselines of LDA, HDP, and HDHP for both tasks.
Yeon Seonwoo, Alice Oh
EMNLP2
2018 Elicast: embedding interactive exercises in instructional programming screencasts
abstract
In programming education, instructors often supplement lectures with active learning experiences by offering programming lab sessions where learners themselves practice writing code. However, widely accessed instructional programming screencasts are not equipped with assessment format that encourages such hands-on programming activities. We introduce Elicast, a screencast tool for recording and viewing programming lectures with embedded programming exercises, to provide hands-on programming experiences in the screen-cast. In Elicast, instructors embed multiple programming exercises while creating a screencast, and learners engage in the exercises by writing code within the screencast, receiving auto-graded results immediately. We conducted an exploratory study of Elicast with five experienced instructors and 63 undergraduate students. We found that instructors structured the lectures into small learning units using embedded exercises as checkpoints. Also, learners more actively engaged in the screencast lectures, checked their understanding of the content through the embedded exercises, and more frequently modified and executed the code during the lectures.
Jungkook Park, Yeong Hoon Park, Jinhan Kim, Jeongmin Cha, Suin Kim, Alice Oh
L@S6
2018 Non-Linear Editing of Text-Based Screencasts
abstract
Screencasts, where recordings of a computer screen are broadcast to a large audience on the web, are becoming popular as an online educational tool. To provide rich interactions with the text within screencasts, there are emerging platforms that support text-based screencasts by recording every character insertion and deletion from the creator and reconstructing its playback on the viewer's screen. However, these platforms lack support for non-linear editing of screencasts, which involves manipulating a sequence of text editing operations. Since text editing operations are tightly coupled in sequence, modifying an arbitrary part of the sequence often creates ambiguity that yields multiple possible results that require user's choice for resolution. We present an editing tool with a non-linear editing algorithm for text-based screencasts. The tool allows users to edit any arbitrary part of a text-based screencast while preserving the overall consistency of the screencast. In an exploratory user study, all subjects successfully carried out a variety of screencast editing tasks using our prototype screencast editor.
Jungkook Park, Yeong Hoon Park, Alice Oh
UIST3
2018 Leveraging the Crowd to Detect and Reduce the Spread of Fake News and Misinformation
abstract
Online social networking sites are experimenting with the following crowd-powered procedure to reduce the spread of fake news and misinformation: whenever a user is exposed to a story through her feed, she can flag the story as misinformation and, if the story receives enough flags, it is sent to a trusted third party for fact checking. If this party identifies the story as misinformation, it is marked as disputed. However, given the uncertain number of exposures, the high cost of fact checking, and the trade-off between flags and exposures, the above mentioned procedure requires careful reasoning and smart algorithms which, to the best of our knowledge, do not exist to date. In this paper, we first introduce a flexible representation of the above procedure using the framework of marked temporal point processes. Then, we develop a scalable online algorithm, CURB, to select which stories to send for fact checking and when to do so to efficiently reduce the spread of misinformation with provable guarantees. In doing so, we need to solve a novel stochastic optimal control problem for stochastic differential equations with jumps, which is of independent interest. Experiments on two real-world datasets gathered from Twitter and Weibo show that our algorithm may be able to effectively reduce the spread of fake news and misinformation.
Behzad Tabibian, Alice Oh, Bernhard Schölkopf, Manuel Gomez-Rodriguez
WSDM3
2017 Eliph: Effective Visualization of Code History for Peer Assessment in Programming Education
abstract
In this paper, we investigate the effectiveness of visualization of code history on peer assessment in computer science education. Peer assessment is found to be an effective learning tool for programming education. While many systems are proposed to support peer assessment in programming education, little effort has been devoted to finding ways to improve the peer assessment by assisting the students to understand the programs they are assessing. We introduce Eliph, a web-based peer assessment system for programming education with code history visualization. Eliph incorporates the visualization of character-level code history, selection-based history tracking and the integration of execution events to assist students in understanding programs written by peers, thereby leading to more effective peer assessment. We evaluate Eliph with an experiment in an undergraduate CS course. We show that visualization of code history has positive effects on promoting higher quality of peer feedback by understanding the intention and thought process.
Jungkook Park, Yeong Hoon Park, Suin Kim, Alice Oh
CSCW4
2017 Rotated Word Vector Representations and their Interpretability
abstract
Vector representation of words improves performance in various NLP tasks, but the high-dimensional word vectors are very difficult to interpret.We apply several rotation algorithms to the vector representation of words to improve the interpretability.Unlike previous approaches that induce sparsity, the rotated vectors are interpretable while preserving the expressive performance of the original vectors.Furthermore, any pre-built word vector representation can be rotated for improved interpretability.We apply rotation to skipgrams and glove and compare the expressive power and interpretability with the original vectors and the sparse overcomplete vectors.The results show that the rotated vectors outperform the original and the sparse overcomplete vectors for interpretability and expressiveness tasks.
JinYeong Bak, Alice Oh
EMNLP3
2017 Hierarchical Dirichlet scaling process
abstract
We present the hierarchical Dirichlet scaling process (HDSP), a Bayesian nonparametric mixed membership model. The HDSP generalizes the hierarchical Dirichlet process to model the correlation structure between metadata in the corpus and mixture components. We construct the HDSP based on the normalized gamma representation of the Dirichlet process, and this construction allows incorporating a scaling function that controls the membership probabilities of the mixture components. We develop two scaling methods to demonstrate that different modeling assumptions can be expressed in the HDSP. We also derive the corresponding approximate posterior inference algorithms using variational Bayes. Through experiments on datasets of newswire, medical journal articles, conference proceedings, and product reviews, we show that the HDSP results in a better predictive performance than labeled LDA, partially labeled LDA, and author topic model and a better negative review classification performance than the supervised topic model and SVM.
Dongwoo Kim 0002, Alice Oh
Mach. Learn.2
2017 Joint Modeling of Topics, Citations, and Topical Authority in Academic Corpora
abstract
Much of scientific progress stems from previously published findings, but searching through the vast sea of scientific publications is difficult. We often rely on metrics of scholarly authority to find the prominent authors but these authority indices do not differentiate authority based on research topics. We present Latent Topical-Authority Indexing (LTAI) for jointly modeling the topics, citations, and topical authority in a corpus of academic papers. Compared to previous models, LTAI differs in two main aspects. First, it explicitly models the generative process of the citations, rather than treating the citations as given. Second, it models each author’s influence on citations of a paper based on the topics of the cited papers, as well as the citing papers. We fit LTAI into four academic corpora: CORA, Arxiv Physics, PNAS, and Citeseer. We compare the performance of LTAI against various baselines, starting with the latent Dirichlet allocation, to the more advanced models including author-link topic model and dynamic author citation topic model. The results show that LTAI achieves improved accuracy over other similar models when predicting words, citations and authors of publications.
Dongwoo Kim 0002, Alice Oh
Trans. Assoc. Comput. Linguistics3
2016 The Proficiency-Congruency Dilemma: Virtual Team Design and Performance in Multiplayer Online Games
abstract
Multiplayer online battle arena games provide an excellent opportunity to study team performance. When designing a team, players must negotiate a proficiency-congruency dilemma between selecting roles that best match their experience and roles that best complement the existing roles on the team. We adopt a mixed-methods approach to explore how players negotiate this dilemma. Using data from League of Legends, we define a similarity space to operationalize team design constructs about role proficiency, generality, and congruency. We collect publicly available data from 3.36 million players to test the influence of these constructs on team performance. We also conduct focus groups with novice and elite players to understand how players' team design practices vary with expertise. We find that the two factors, player proficiency and team congruency, both increase team performance, with the former having a stronger impact. We also find that elite players are better at balancing the two factors than the novice players. These findings have implications for players, designers, and theorists about how to recommend team designs that jointly prioritize individuals' expertise and teams' compatibility.
Brian Keegan, Alice Oh
CHI4
2016 How to Compete Online for News Audience: Modeling Words that Attract Clicks
abstract
Headlines are particularly important for online news outlets where there are many similar news stories competing for users' attention. Traditionally, journalists have followed rules-of-thumb and experience to master the art of crafting catchy headlines, but with the valuable resource of large-scale click-through data of online news articles, we can apply quantitative analysis and text mining techniques to acquire an in-depth understanding of headlines. In this paper, we conduct a large-scale analysis and modeling of 150K news articles published over a period of four months on the Yahoo home page. We define a simple method to measure click-value of individual words, and analyze how temporal trends and linguistic attributes affect click-through rate (CTR). We then propose a novel generative model, headline click-based topic model (HCTM), that extends latent Dirichlet allocation (LDA) to reveal the effect of topical context on the click-value of words in headlines. HCTM leverages clicks in aggregate on previously published headlines to identify words for headlines that will generate more clicks in the future. We show that by jointly taking topics and clicks into account we can detect changes in user interests within topics. We evaluate HCTM in two different experimental settings and compare its performance with ALDA (adapted LDA), LDA, and TextRank. The first task, full headline, is to retrieve full headline used for a news article given the body of news article. The second task, good headline, is to specifically identify words in the headline that have high click values for current news audience. For full headline task, our model performs on par with ALDA, a state-of-the art web-page summarization method that utilizes click-through information. For good headline task, which is of more practical importance to both individual journalists and online news outlets, our model significantly outperforms all other comparative methods.
Joon Hee Kim, Amin Mantrach, Alejandro Jaimes, Alice Oh
KDD4
2016 Elice: An online CS Education Platform to Understand How Students Learn Programming
abstract
We present Elice, an online CS (computer science) education platform, and Elivate, a system for taking student learning data from Elice and infers their progress through an educational taxonomy tailored for programming education. Elice captures detailed student learning activities, such as the intermediate revisions of code as students make progress toward completing their programming exercises. With those data, Elivate recognizes each student's progression through an education taxonomy which organizes intermediate stages of learning such that the taxonomy can be used to evaluate student progress as well as to design and improve course materials and structure. With more than 240,000 intermediate source codes generated by 1,000 students, we demonstrate the practicality of the Elice and Elivate. We present case studies that confirm that categorizing student actions into the different steps of the taxonomy results in better understanding of the effect of TA's assist and student's performance.
Suin Kim, Jae Won Kim, Jungkook Park, Alice Oh
L@S4
2016 Elivate: A Real-Time Assistant for Students and Lecturers as Part of an Online CS Education Platform
abstract
We present Elice, an online CS (computer science) education platform, and Elivate, a system for (i) taking student learning data from Elice, (ii) inferring their progress through an educational taxonomy tailored for programming education, and (iii) generating the real-time assistance for students and lecturers. Online courses suffer from high average attrition rates, and early prediction can enable early personalized feedback to motivate and assist students who may be having difficulties. Elice captures detailed student learning activities including intermediate revisions of code as students make progress toward completing their programming exercises and timestamps of student logins and submissions. Elivate then takes those data to analyze each student's progress and estimate the time to completion. In doing so, Elivate uses a learning taxonomy and automatic clustering of source code revisions. Using more than 240,000 code revisions generated by 1,000 students, we demonstrate how Elivate processes large-scale student data and generates appropriate real-time feedback for students.
Suin Kim, Jae Won Kim, Jungkook Park, Alice Oh
L@S4
2015 Social Media Dynamics of Global Co-presence During the 2014 FIFA World Cup
abstract
Sporting championships and other media events can induce very strong feelings of co-presence that can change communication patterns within large communities. Live tweeting reactions to media events provide high-resolution data with time-stamps to understand these behavioral dynamics. We employ a computational focus group method to identify a population of 790,744 international Twitter users, and we track their behavior before, during, and after the 2014 FIFA World Cup. We pick, in particular, a set of Twitter users who specified the teams that they are supporting, such that we can identify communities of fans of the teams, as well as the entire community of World Cup fans. The structure, dynamics, and content of communication of these communities of users are analyzed to compare behavior outside of the matches to behavior during the event and to examine behavioral responses across languages. Specifically, the temporal patterns of the tweeting volume, topics, retweet- ing, and mentioning behaviors are analyzed. We find there are similarities in the responses to media events, characteristic changes in activity patterns of users, and substantial differences in linguistic features. These findings have implications for designing more resilient socio-technical systems during crises and developing better models of complex social behavior.
Jae Won Kim, Dongwoo Kim 0002, Brian Keegan, Joon Hee Kim, Suin Kim, Alice Oh
CHI6
2015 Towards Understanding Relational Orientation: Attachment Theory and Facebook Activities
abstract
Knowing individuals' relational orientation is imperative for effective offline, as well as online, interactions and collaborations. We use attachment theory to examine the link between Facebook users' relational orientation (in terms of attachment styles: anxiety and avoidance) and their relational activities. Our research examines whether and how the two key relational processes identified in offline social relationships (self-expression and responsiveness) are manifested on online social networks and related to attachment styles. We describe our dataset of 640 Facebook users, their attachment scale survey results, and their 525,334 posts. We define four features that map onto relational activities on Facebook: status updates and status updates with emotional words (self-expression); comments and likes (responsiveness). We find significant relationships between the users' attachment styles and their self-expression and responsiveness activities on Facebook. A key takeaway of our research is that without relying on self-reported surveys, a computational analysis of a Facebook user's self-expressing and responding activities alone can reveal the user's underlying relational orientation (i.e., attachment style).
Bumsoo Kang, Alice Oh, Inseok Hwang 0001, Junehwa Song
CSCW3
2014 Self-disclosure topic model for classifying and analyzing Twitter conversations
abstract
Self-disclosure, the act of revealing one-self to others, is an important social be-havior that strengthens interpersonal rela-tionships and increases social support. Al-though there are many social science stud-ies of self-disclosure, they are based on manual coding of small datasets and ques-tionnaires. We conduct a computational analysis of self-disclosure with a large dataset of naturally-occurring conversa-tions, a semi-supervised machine learning algorithm, and a computational analysis of the effects of self-disclosure on subse-quent conversations. We use a longitu-dinal dataset of 17 million tweets, all of which occurred in conversations that con-sist of five or more tweets directly reply-ing to the previous tweet, and from dyads with twenty of more conversations each. We develop self-disclosure topic model (SDTM), a variant of latent Dirichlet al-location (LDA) for automatically classi-fying the level of self-disclosure for each tweet. We take the results of SDTM and analyze the effects of self-disclosure on subsequent conversations. Our model sig-nificantly outperforms several comparable methods on classifying the level of self-disclosure, and the analysis of the longitu-dinal data using SDTM uncovers signifi-cant and positive correlation between self-disclosure and conversation frequency and length. 1
JinYeong Bak, Chin-Yew Lin, Alice Oh
EMNLP3
2014 Hierarchical Dirichlet Scaling Process
abstract
We present the hierarchical Dirichlet scaling process (HDSP), a Bayesian nonparametric mixed membership model for multi-labeled data. We construct the HDSP based on the gamma representation of the hierarchical Dirichlet process (HDP) which allows scaling the mixture components. With such construction, HDSP allocates a latent location to each label and mixture component in a space, and uses the distance between them to guide membership probabilities. We develop a variational Bayes algorithm for the approximate posterior inference of the HDSP. Through experiments on synthetic datasets as well as datasets of newswire, medical journal articles, and Wikipedia, we show that the HDSP results in better predictive performance than HDP, labeled LDA and partially labeled LDA.
Dongwoo Kim 0002, Alice Oh
ICML2
2014 Effective ranking and search techniques for Web resources considering semantic relationships
Jun-Ki Min, Alice Oh, Chin-Wan Chung
Inf. Process. Manag.3
2013 A Hierarchical Aspect-Sentiment Model for Online Reviews
abstract
To help users quickly understand the major opinions from massive online reviews, it is important to automatically reveal the latent structure of the aspects, sentiment polarities, and the association between them. However, there is little work available to do this effectively. In this paper, we propose a hierarchical aspect sentiment model (HASM) to discover a hierarchical structure of aspect-based sentiments from unlabeled online reviews. In HASM, the whole structure is a tree. Each node itself is a two-level tree, whose root represents an aspect and the children represent the sentiment polarities associated with it. Each aspect or sentiment polarity is modeled as a distribution of words. To automatically extract both the structure and parameters of the tree, we use a Bayesian nonparametric model, recursive Chinese Restaurant Process (rCRP), as the prior and jointly infer the aspect-sentiment tree from the review texts. Experiments on two real datasets show that our model is comparable to two other hierarchical topic models in terms of quantitative measures of topic trees. It is also shown that our model achieves better sentence-level classification accuracy than previously proposed aspect-sentiment joint models.
Suin Kim, Zheng Chen 0001, Alice Oh, Shixia Liu
AAAI4
2013 Context-Dependent Conceptualization
Dongwoo Kim 0002, Haixun Wang, Alice Oh
IJCAI3
2012 Modeling topic hierarchies with the recursive chinese restaurant process
abstract
Topic models such as latent Dirichlet allocation (LDA) and hierarchical Dirichlet processes (HDP) are simple solutions to discover topics from a set of unannotated documents. While they are simple and popular, a major shortcoming of LDA and HDP is that they do not organize the topics into a hierarchical structure which is naturally found in many datasets. We introduce the recursive Chinese restaurant process (rCRP) and a nonparametric topic model with rCRP as a prior for discovering a hierarchical topic structure with unbounded depth and width. Unlike previous models for discovering topic hierarchies, rCRP allows the documents to be generated from a mixture over the entire set of topics in the hierarchy. We apply rCRP to a corpus of New York Times articles, a dataset of MovieLens ratings, and a set of Wikipedia articles and show the discovered topic hierarchies. We compare the predictive power of rCRP with LDA, HDP, and nested Chinese restaurant process (nCRP) using heldout likelihood to show that rCRP outperforms the others. We suggest two metrics that quantify the characteristics of a topic hierarchy to compare the discovered topic hierarchies of rCRP and nCRP. The results show that rCRP discovers a hierarchy in which the topics become more specialized toward the leaves, and topics in the immediate family exhibit more affinity than topics beyond the immediate family.
Joon Hee Kim, Dongwoo Kim 0002, Suin Kim, Alice Oh
CIKM4
2012 Dirichlet Process with Mixed Random Measures: A Nonparametric Topic Model for Labeled Data
Dongwoo Kim 0002, Suin Kim, Alice Oh
ICML3
2012 Do You Feel What I Feel? Social Aspects of Emotions in Twitter Conversations
Suin Kim, JinYeong Bak, Alice Oh
ICWSM3
2011 Topic Chains for Understanding a News Corpus
Dongwoo Kim 0002, Alice Oh
CICLing (2)2
2011 Accounting for data dependencies within a hierarchical dirichlet process mixture model
abstract
We propose a hierarchical nonparametric topic model, based on the hierarchical Dirichlet process (HDP), that accounts for dependencies among the data. The HDP mixture models are useful for discovering an unknown semantic structure (i.e., topics) from a set of unstructured data such as a corpus of documents. For simplicity, HDP makes an exchangeability assumption that any permutation of the data points would result in the same joint probability of the data being generated. This exchangeability assumption poses a problem for some domains where there are clear and strong dependencies among the data. A model that allows for non-exchangeability of data can capture these dependencies and assign higher probabilities to clusters that account for data dependencies, for example, inferring topics that reflect the temporal patterns of the data. Our model incorporates the distance dependent Chinese restaurant process (ddCRP), which clusters data with an inherent bias toward clusters of data points that are near to one another, into a hierarchical construction analogous to the HDP, and we call this new prior the distance dependent Chinese restaurant franchise (ddCRF). When tested with temporal datasets, the ddCRF mixture model shows clear improvements in data fit compared to the HDP in terms of heldout likelihood and complexity. The resulting set of topics shows the sequential emergence and disappearance patterns of topics.
Dongwoo Kim 0002, Alice Oh
CIKM2
2011 Analyzing social media in escalating crisis situations
abstract
The rapid diffusion of information and opinions through social media, such as web forums and micro-blogs, is affecting the development of crisis situations, such as the Iranian presidential election, the Egyptian protest, and the ROKS Cheonan sinking. Understanding this rapid widespread diffusion, and assessing what information is spreading, what ideas are becoming common, and who is talking about what, is critical for crisis management. This paper presents a computational system for social media assessing the flow of ideas on the web and changes in who is talking about what. This system, given raw social media data, identifies the key topics, the key paths by which topics evolve, the key individuals who contribute to the topic, and the key influence relations between the contributors. We present this system implemented with the Author-Topic model, the meta-network model, and various computational techniques to find and filter the heavy contributors and influences. We demonstrate the performance of the system, by applying it to social media data surrounding the ROKS Cheonan sinking. We describe the results of assessing the initial and changing perceptions of the event using this system.
Il-Chul Moon, Alice Oh, Kathleen M. Carley
ISI2
2011 Aspect and sentiment unification model for online review analysis
abstract
User-generated reviews on the Web contain sentiments about detailed aspects of products and services. However, most of the reviews are plain text and thus require much effort to obtain information about relevant details. In this paper, we tackle the problem of automatically discovering what aspects are evaluated in reviews and how sentiments for different aspects are expressed. We first propose Sentence-LDA (SLDA), a probabilistic generative model that assumes all words in a single sentence are generated from one aspect. We then extend SLDA to Aspect and Sentiment Unification Model (ASUM), which incorporates aspect and sentiment together to model sentiments toward different aspects. ASUM discovers pairs of {aspect, sentiment} which we call senti-aspects. We applied SLDA and ASUM to reviews of electronic devices and restaurants. The results show that the aspects discovered by SLDA match evaluative details of the reviews, and the senti-aspects found by ASUM capture important aspects that are closely coupled with a sentiment. The results of sentiment classification show that ASUM outperforms other generative models and comes close to supervised classification methods. One important advantage of ASUM is that it does not require any sentiment labels of the reviews, which are often expensive to obtain.
Yohan Jo, Alice Oh
WSDM2
2010 Users' needs for social tagging and sharing on mobile contacts
abstract
In this paper we describe our research toward improving the current mobile contacts applications, which we found to lack important features that are essential to a fully satisfactory user experience. We identify the needs for a better user experience for organizing and searching, as well as looking for information from one's social network. We present the results of a user study that identified the problems with the current mobile contacts applications and propose tagging contacts and social network information sharing as the mechanism for improving their usability and usefulness.
Trung Van Nguyen, Alice Oh
Mobile HCI2
2009 User Evaluation of a System for Classifying and Displaying Political Viewpoints of Weblogs
Alice Oh
ICWSM1
2008 Generating Baseball Summaries from Multiple Perspectives by Reordering Content
Alice Oh, Howard E. Shrobe
INLG1
2002 Face-Responsive Interfaces: From Direct Manipulation to Perceptive Presence
Trevor Darrell, Konrad Tollmar, Frank Bentley, Neal Checka, Louis-Philippe Morency, Alice Oh
UbiComp7
2002 Stochastic natural language generation for spoken dialog systems
Alice Oh, Alexander I. Rudnicky
Comput. Speech Lang.1
2000 Task and domain specific modelling in the Carnegie Mellon communicator system
abstract
The Carnegie Mellon Communicator is a telephone-based dialog system that supports planning in a travel domain. The implementation of such a system requires two complimentary components, an architecture capable of managing interaction and the task, as well as a knowledge base that captures the speech, language and task characteristics specific to the domain. Given a suitable architecture, the principal effort in development in taken up in the acquisition and processing of a domain knowledge base. This paper describes a variety of techniques we have applied to modeling in acoustic, language, task, generation and synthesis components of the system. 1. INTRODUCTION System development involves a great deal of knowledge engineering, which is both time-consuming and requires a variety of experts to participate in the process. Therefore methods that seek to minimize this resource, for example through training based on domain-specific corpora are preferred. Effective use of corpora, however, ...
Alexander I. Rudnicky, Christina L. Bennett, Alan W. Black, Ananlada Chotimongkol, Kevin A. Lenzo, Alice Oh, Rita Singh
INTERSPEECH6
1999 Creating natural dialogs in the carnegie mellon communicator system
Alexander I. Rudnicky, Eric H. Thayer, Paul C. Constantinides, Chris Tchou, R. Shern, Kevin A. Lenzo, Alice Oh
EUROSPEECH8