Han Kyul Kim

dblp:205/5686 · DBLP profile ↗
← Back
6ranked-venue papers in the field
5as first author
6since 2021 · last 2025
ORCID · conflict

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 6 (5 first)
YearPublicationVenuePosition
2025 Whose Line is it Anyway? Linguistic Attribution of Misinformation from Mainstream and Fringe LLMs
Han Kyul Kim, Ankur Garg, Hansea Kim, Eunjeong Joo, Andy Skumanich
IEEE Big Data1
2025 Performance Gap-Aware Distributionally Robust Optimization for Fair Deep Knowledge Tracing
Han Kyul Kim, Stephen Lu
IEEE Big Data1
2025 Understanding and Generating Student Questions with LLMs in Collaborative Learning
Han Kyul Kim, Shriniwas Nayak, Aleyeh Roknaldin, Stephen Lu
IEEE Big Data1
2024 Enhancing Predictive Fairness in Deep Knowledge Tracing with Sequence Inverted Data Augmentation
abstract
Ensuring fairness in predictive models is a critical challenge, particularly in sensitive domains such as education. This paper addresses the issue of predictive bias in Deep Knowledge Tracing (DKT), a widely used model in Intelligent Tutoring Systems (ITS) for modeling student learning and predicting future performance. While DKT and its successors have advanced the state-of-the-art in personalized learning, they remain vulnerable to biases that emerge due to distribution shifts in student data, often occurring at the end of academic cycles or semesters. In response to this challenge, we introduce a novel sequence-inverted data augmentation method that significantly enhances both predictive accuracy and fairness in DKT. Our approach generates synthetic student sequences representing diverse performance extremes, thereby bolstering the model’s robustness against distribution shifts and reducing predictive bias across gender groups.Through extensive experiments on a real-world, large-scale student dataset, we show that our method significantly outperforms the previously state-of-the-art Balanced-3 technique, which, despite its prominence in bias mitigation for educational contexts, proves ineffective in sequential prediction tasks like DKT. Unlike sampling-based methods, our proposed method achieves superior predictive fairness without sacrificing accuracy, maintaining stable performance even in the presence of severe distribution shifts.Our work is the first to introduce a bias mitigation approach tailored for KT models and to thoroughly evaluate the issue of predictive fairness using real student data. Previous research often addresses predictive fairness in a broad and generalized manner, overlooking the various factors that contribute to bias. In contrast, this paper delves into a specific and realistic factor influencing predictive fairness — distribution shifts in test data. This targeted approach discussed in this paper not only refines our understanding of predictive fairness in the context of KT but also establishes a new perspective on analyzing and mitigating bias. This targeted approach paves the way for more equitable AI solutions in education and offers valuable insights for other domains involving sequential prediction, such as time series analysis.
Han Kyul Kim, Stephen Lu
IEEE Big Data1
2024 Active Learning for Practical Misinformation Classification in Social Media: a Case Study on COVID-19
abstract
Misinformation on social media has become a significant societal issue, as these platforms increasingly serve as central hubs for human interaction. The recent COVID-19 pandemic vividly illustrated how the rapid spread of misinformation can lead to adverse personal and societal impacts, exacerbated by the ease with which information is generated, shared, and consumed on these platforms. Much of the previous research has focused on using machine learning algorithms to detect and identify misinformation in social media posts, primarily operating within a strict supervised learning framework where annotated datasets containing both accurate and misleading information are available.However, especially with the rise of generative AI, the task of annotating misinformation posts has become increasingly resource-intensive, requiring substantial domain-specific expertise. As a result, many of these algorithms struggle to adapt to the real-world environment, where data distribution and topics are constantly evolving on social media platforms. To bridge this gap, our paper evaluates the effectiveness of active learning approaches for classifying misinformation with a focus on COVID-19. We demonstrate that achieving a highly accurate classifier with a training dataset significantly smaller than those required in previous studies is possible. Furthermore, we introduce a novel uncertainty estimation method for active learning, the Embedding Vector Similarity (EVS) measure, which leverages the latent embedding space of large language models. This metric enhances the performance of active learning for misinformation classification, reducing the dependence on large, high-resource annotated datasets while maintaining the quality of the analysis for addressing this critical social issue.
Han Kyul Kim, Andy Skumanich
IEEE Big Data1
2024 Methods for Addressing the Societal Challenges of Malicious Narratives in Social Media Using AI/ML Analysis: a Case Study of Hate Speech
abstract
This paper highlights a methodology for analyzing and monitoring malicious communication in social media in order to mitigate the potential harm. There has been a deliberate "weaponization" of messaging through the use of social networks, including by politically oriented entities, both state-sponsored and privately run. The article identifies a use of AI/ML characterization of generalized "mal-info," a broad term which includes deliberate malicious narratives similar to hate speech, which adversely impacts society. A key point of the discussion is that this mal-info will dramatically increase in volume, and it will become essential for sharable quantifying tools to provide support for human expert intervention. Despite attempts to introduce moderation on major platforms like Facebook and X/Twitter, there are now established alternative social networks that offer completely unmoderated spaces. The paper presents an introduction to these platforms and a case study showing a qualitative and semi-quantitative analysis of representative examples of mal-info posts. The action examines several inflammatory terms using text analysis and, importantly, discusses the use of generative algorithms by one political agent in particular, providing some examples of the potential risks to society. This latter is of grave concern, and monitoring tools must be established. This paper presents a preliminary step to selecting relevant sources and setting a foundation for characterizing the mal-info, which must be addressed to ensure societal well-being. The AI/ML methods demonstrate a means for semi-quantitative signature capture. The impending use of "mal-GenAI" is presented. The main findings are: (1) we introduce specific politicized social channels of note; (2) we present viable indicative AI/ML modes of characterizing the output of these channels for analyzing and tracking mal-info; (3) we outline a case study with qualitative and semi-quantitative signatures of mal-info; (4) we flag the impending use of mal-GenAI in this context. It is critically important to develop mitigation strategies for mal-info, given the potential for societal harm.
Andy Skumanich, Han Kyul Kim
IEEE Big Data2