Xi Yang 0019

dblp:13/1520-19 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0003-0026-9096ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2025 A Generalized Apprenticeship Learning Framework for Capturing Evolving Student Pedagogical Strategies
Md. Mirajul Islam, Xi Yang 0019, Rajesh Debnath, Adittya Shoukarjya Saha, Min Chi
AIED (3)2
2025 THEMES: An Offline Apprenticeship Learning Framework for Evolving Reward Functions
abstract
Apprenticeship learning (AL) aims to induce decision-making policies by observing and imitating expert demonstrations.Existing AL approaches typically rely on online interactions and assume that the demonstrations follow a single reward function.Nevertheless, in real-world human-centric applications, policies are usually learned in an offline setting, with the demonstrations driven by multiple reward functions that evolve over time.To address these challenges, we introduce a novel AL framework: Time-aware Hierarchical EM Energy-based Sub-trajectory (THEMES) clustering.We evaluate the effectiveness of THEMES in two challenging human-centric domains -healthcare and education.Our experimental results across multiple datasets demonstrate that THEMES can accurately induce policies, outperforming competitive baselines and ablations, demonstrating its potential for tackling a broad range of complex, real-world human-centric tasks.
Xi Yang 0019, Md. Mirajul Islam, Min Chi
KDD (2)1
2024 Get a Head Start: On-Demand Pedagogical Policy Selection in Intelligent Tutoring
abstract
Reinforcement learning (RL) is broadly employed in human-involved systems to enhance human outcomes. Off-policy evaluation (OPE) has been pivotal for RL in those realms since online policy learning and evaluation can be high-stake. Intelligent tutoring has raised tremendous attentions as highly challenging when applying OPE to human-involved systems, due to that students' subgroups can favor different pedagogical policies and the costly procedure that policies have to be induced fully offline and then directly deployed to the upcoming semester. In this work, we formulate on-demand pedagogical policy selection (ODPS) to tackle the challenges for OPE in intelligent tutoring. We propose a pipeline, EduPlanner, as a concrete solution for ODPS. Our pipeline results in an theoretically unbiased estimator, and enables efficient and customized policy selection by identifying subgroups over both historical data and on-arrival initial logs. We evaluate our approach on the Probability ITS that has been used in real classrooms for over eight years. Our study shows significant improvement on learning outcomes of students with EduPlanner, especially for the ones associated with low-performing subgroups.
Xi Yang 0019, Min Chi
AAAI2
2024 A Generalized Apprenticeship Learning Framework for Modeling Heterogeneous Student Pedagogical Strategies
Md. Mirajul Islam, Xi Yang 0019, John Wesley Hostetter, Adittya Soukarjya Saha, Min Chi
EDM2
2024 On Trajectory Augmentations for Off-Policy Evaluation
abstract
In the realm of reinforcement learning (RL), off-policy evaluation (OPE) holds a pivotal position, especially in high-stake human-involved scenarios such as e-learning and healthcare. Applying OPE to these domains is often challenging with scarce and underrepresentative offline training trajectories. Data augmentation has been a successful technique to enrich training data. However, directly employing existing data augmentation methods to OPE may not be feasible, due to the Markovian nature within the offline trajectories and the desire for generalizability across diverse target policies. In this work, we propose an offline trajectory augmentation approach to specifically facilitate OPE in human-involved scenarios. We propose sub-trajectory mining to extract potentially valuable sub-trajectories from offline data, and diversify the behaviors within those sub-trajectories by varying coverage of the state-action space. Our work was empirically evaluated in a wide array of environments, encompassing both simulated scenarios and real-world domains like robotic control, healthcare, and e-learning, where the training trajectories include varying levels of coverage of the state-action space. By enhancing the performance of a variety of OPE methods, our work offers a promising path forward for tackling OPE challenges in situations where data may be limited or underrepresentative.
Qitong Gao, Xi Yang 0019, Song Ju, Miroslav Pajic, Min Chi
ICLR3
2024 Off-Policy Selection for Initiating Human-Centric Experimental Design
abstract
In human-centric applications like healthcare and education, the \textit{heterogeneity} among patients and students necessitates personalized treatments and instructional interventions. While reinforcement learning (RL) has been utilized in those tasks, off-policy selection (OPS) is pivotal to close the loop by offline evaluating and selecting policies without online interactions, yet current OPS methods often overlook the heterogeneity among participants. Our work is centered on resolving a \textit{pivotal challenge} in human-centric systems (HCSs): \textbf{\textit{how to select a policy to deploy when a new participant joining the cohort, without having access to any prior offline data collected over the participant?}} We introduce First-Glance Off-Policy Selection (FPS), a novel approach that systematically addresses participant heterogeneity through sub-group segmentation and tailored OPS criteria to each sub-group. By grouping individuals with similar traits, FPS facilitates personalized policy selection aligned with unique characteristics of each participant or group of participants. FPS is evaluated via two important but challenging applications, intelligent tutoring systems and a healthcare application for sepsis treatment and intervention. FPS presents significant advancement in enhancing learning outcomes of students and in-hospital care outcomes.
Xi Yang 0019, Qitong Gao, Song Ju, Miroslav Pajic, Min Chi
NeurIPS2
2023 Hierarchical Apprenticeship Learning for Disease Progression Modeling
abstract
Disease progression modeling (DPM) plays an essential role in characterizing patients' historical pathways and predicting their future risks. Apprenticeship learning (AL) aims to induce decision-making policies by observing and imitating expert behaviors. In this paper, we investigate the incorporation of AL-derived patterns into DPM, utilizing a Time-aware Hierarchical EM Energy-based Subsequence (THEMES) AL approach. To the best of our knowledge, this is the first study incorporating AL-derived progressive and interventional patterns for DPM. We evaluate the efficacy of this approach in a challenging task of septic shock early prediction, and our results demonstrate that integrating the AL-derived patterns significantly enhances the performance of DPM.
Xi Yang 0019, Min Chi
IJCAI1
2023 XAI to Increase the Effectiveness of an Intelligent Pedagogical Agent
abstract
We explore eXplainable AI (XAI) to enhance user experience and understand the value of explanations in AI-driven pedagogical decisions within an Intelligent Pedagogical Agent (IPA). Our real-time and personalized explanations cater to students' attitudes to promote learning. In our empirical study, we evaluate the effectiveness of personalized explanations by comparing three versions of the IPA: (1) personalized explanations and suggestions, (2) suggestions but no explanations, and (3) no suggestions. Our results show the IPA with personalized explanations significantly improves students' learning outcomes compared to the other versions.
John Wesley Hostetter, Cristina Conati, Xi Yang 0019, Mark Abdelshiheed, Tiffany Barnes, Min Chi
IVA3
2022 Mixing Backward- with Forward-Chaining for Metacognitive Skill Acquisition and Transfer
Mark Abdelshiheed, John Wesley Hostetter, Xi Yang 0019, Tiffany Barnes, Min Chi
AIED (1)3
2022 Student-Tutor Mixed-Initiative Decision-Making Supported by Deep Reinforcement Learning
Song Ju, Xi Yang 0019, Tiffany Barnes, Min Chi
AIED (1)2
2022 A Reinforcement Learning-Informed Pattern Mining Framework for Multivariate Time Series Classification
abstract
Multivariate time series (MTS) classification is a challenging and important task in various domains and real-world applications. Much of prior work on MTS can be roughly divided into neural network (NN)- and pattern-based methods. The former can lead to robust classification performance, but many of the generated patterns are challenging to interpret; while the latter often produce interpretable patterns that may not be helpful for the classification task. In this work, we propose a reinforcement learning (RL) informed PAttern Mining framework (RLPAM) to identify interpretable yet important patterns for MTS classification. Our framework has been validated by 30 benchmark datasets as well as real-world large-scale electronic health records (EHRs) for an extremely challenging task: sepsis shock early prediction. We show that RLPAM outperforms the state-of-the-art NN-based methods on 14 out of 30 datasets as well as on the EHRs. Finally, we show how RL informed patterns can be interpretable and can improve our understanding of septic shock progression.
Qitong Gao, Xi Yang 0019, Miroslav Pajic, Min Chi
IJCAI3
2021 Multi-series Time-aware Sequence Partitioning for Disease Progression Modeling
abstract
Electronic healthcare records (EHRs) are comprehensive longitudinal collections of patient data that play a critical role in modeling the disease progression to facilitate clinical decision-making. Based on EHRs, in this work, we focus on sepsis -- a broad syndrome that can develop from nearly all types of infections (e.g., influenza, pneumonia). The symptoms of sepsis, such as elevated heart rate, fever, and shortness of breath, are vague and common to other illnesses, making the modeling of its progression extremely challenging. Motivated by the recent success of a novel subsequence clustering approach: Toeplitz Inverse Covariance-based Clustering (TICC), we model the sepsis progression as a subsequence partitioning problem and propose a Multi-series Time-aware TICC (MT-TICC), which incorporates multi-series nature and irregular time intervals of EHRs. The effectiveness of MT-TICC is first validated via a case study using a real-world hand gesture dataset with ground-truth labels. Then we further apply it for sepsis progression modeling using EHRs. The results suggest that MT-TICC can significantly outperform competitive baseline models, including the TICC. More importantly, it unveils interpretable patterns, which sheds some light on better understanding the sepsis progression.
Xi Yang 0019, Yuan Zhang 0028, Min Chi
IJCAI1
2020 Student Subtyping via EM-Inverse Reinforcement Learning
Xi Yang 0019, Guojing Zhou, Michelle Taub, Roger Azevedo, Min Chi
EDM1
2020 PRIME: Block-Wise Missingness Handling for Multi-modalities in Intelligent Tutoring Systems
Xi Yang 0019, Yeo-Jin Kim, Michelle Taub, Roger Azevedo, Min Chi
MMM (2)1
2020 Improving Student-System Interaction Through Data-driven Explanations of Hierarchical Reinforcement Learning Induced Pedagogical Policies
abstract
Motivated by the recent advances of reinforcement learning and the traditional grounded Self Determination Theory (SDT), we explored the impact of hierarchical reinforcement learning (HRL) induced pedagogical policies and data-driven explanations of the HRL-induced policies on student experience in an Intelligent Tutoring System (ITS). We explored their impacts first independently and then jointly. Overall our results showed that 1) the HRL induced policies could significantly improve students' learning performance, and 2) explaining the tutor's decisions to students through data-driven explanations could improve the student-system interaction in terms of students' engagement and autonomy.
Guojing Zhou, Xi Yang 0019, Hamoon Azizsoltani, Tiffany Barnes, Min Chi
UMAP2
2019 Big, Little, or Both? Exploring the Impact of Granularity on Learning for Students with Different Incoming Competence
Guojing Zhou, Xi Yang 0019, Min Chi
CogSci2
2019 ATTAIN: Attention-based Time-Aware LSTM Networks for Disease Progression Modeling
abstract
Modeling patient disease progression using Electronic Health Records (EHRs) is critical to assist clinical decision making. Long-Short Term Memory (LSTM) is an effective model to handle sequential data, such as EHRs, but it encounters two major limitations when applied to EHRs: it is unable to interpret the prediction results and it ignores the irregular time intervals between consecutive events. To tackle these limitations, we propose an attention-based time-aware LSTM Networks (ATTAIN), to improve the interpretability of LSTM and to identify the critical previous events for current diagnosis by modeling the inherent time irregularity. We validate ATTAIN on modeling the progression of an extremely challenging disease, septic shock, by using real-world EHRs. Our results demonstrate that the proposed framework outperforms the state-of-the-art models such as RETAIN and T-LSTM. Also, the generated interpretative time-aware attention weights shed some lights on the progression behaviors of septic shock.
Yuan Zhang 0028, Xi Yang 0019, Julie S. Ivy, Min Chi
IJCAI2
2018 Time-aware Subgroup Matrix Decomposition: Imputing Missing Data Using Forecasting Events
abstract
Deep neural network models, especially Long Short Term Memory (LSTM), have shown great success in analyzing Electronic Health Records (EHRs) due to their ability to capture temporal dependencies in time series data. When applying the deep learning models to EHRs, we are generally confronted with two major challenges: high rate of missingness and time irregularity. Motivated by the original PACIFIER framework which utilized matrix decomposition for data imputation, we applied and further extended it by including three components: forecasting future events, a time-aware mechanism, and a subgroup basis approach. We evaluated the proposed framework with real-world EHRs which consists of 52,919 visits and 4,224,567 events on a task of early prediction of septic shock. We compared our work against multiple baselines including the original PACIFIER using both LSTM and Time-aware LSTM (T-LSTM). Experimental results showed that our proposed framework significantly outperformed all competitive baseline approaches. More importantly, the extracted interpretative latent patterns from subgroups could shed some lights for clinicians to discover the progression of septic shock patients.
Xi Yang 0019, Yuan Zhang 0028, Min Chi
IEEE BigData1