VLDB 2026 Research / reviewers in the wild / expert
Song Ju
dblp:244/6887
· DBLP profile ↗
9ranked-venue papers
5as first author
7since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | On Trajectory Augmentations for Off-Policy EvaluationabstractIn the realm of reinforcement learning (RL), off-policy evaluation (OPE) holds a pivotal position, especially in high-stake human-involved scenarios such as e-learning and healthcare. Applying OPE to these domains is often challenging with scarce and underrepresentative offline training trajectories. Data augmentation has been a successful technique to enrich training data. However, directly employing existing data augmentation methods to OPE may not be feasible, due to the Markovian nature within the offline trajectories and the desire for generalizability across diverse target policies. In this work, we propose an offline trajectory augmentation approach to specifically facilitate OPE in human-involved scenarios. We propose sub-trajectory mining to extract potentially valuable sub-trajectories from offline data, and diversify the behaviors within those sub-trajectories by varying coverage of the state-action space. Our work was empirically evaluated in a wide array of environments, encompassing both simulated scenarios and real-world domains like robotic control, healthcare, and e-learning, where the training trajectories include varying levels of coverage of the state-action space. By enhancing the performance of a variety of OPE methods, our work offers a promising path forward for tackling OPE challenges in situations where data may be limited or underrepresentative. Qitong Gao, Xi Yang 0019, Song Ju, Miroslav Pajic, Min Chi |
ICLR | 4 |
| 2024 | Off-Policy Selection for Initiating Human-Centric Experimental DesignabstractIn human-centric applications like healthcare and education, the \textit{heterogeneity} among patients and students necessitates personalized treatments and instructional interventions. While reinforcement learning (RL) has been utilized in those tasks, off-policy selection (OPS) is pivotal to close the loop by offline evaluating and selecting policies without online interactions, yet current OPS methods often overlook the heterogeneity among participants. Our work is centered on resolving a \textit{pivotal challenge} in human-centric systems (HCSs): \textbf{\textit{how to select a policy to deploy when a new participant joining the cohort, without having access to any prior offline data collected over the participant?}} We introduce First-Glance Off-Policy Selection (FPS), a novel approach that systematically addresses participant heterogeneity through sub-group segmentation and tailored OPS criteria to each sub-group. By grouping individuals with similar traits, FPS facilitates personalized policy selection aligned with unique characteristics of each participant or group of participants. FPS is evaluated via two important but challenging applications, intelligent tutoring systems and a healthcare application for sepsis treatment and intervention. FPS presents significant advancement in enhancing learning outcomes of students and in-hospital care outcomes. Xi Yang 0019, Qitong Gao, Song Ju, Miroslav Pajic, Min Chi |
NeurIPS | 4 |
| 2022 | Student-Tutor Mixed-Initiative Decision-Making Supported by Deep Reinforcement Learning
Song Ju, Xi Yang 0019, Tiffany Barnes, Min Chi |
AIED (1) | 1 |
| 2021 | Evaluating Critical Reinforcement Learning Framework in the Field
Song Ju, Guojing Zhou, Mark Abdelshiheed, Tiffany Barnes, Min Chi |
AIED (1) | 1 |
| 2021 | InferNet for Delayed Reinforcement Tasks: Addressing the Temporal Credit Assignment ProblemabstractRewards are the critical signals for Reinforcement Learning (RL) algorithms to learn the desired behavior in a sequential multi-step learning task. However, when these rewards are delayed and noisy in nature, the learning process becomes more challenging. The temporal Credit Assignment Problem (CAP) is a well-known and challenging task in AI. While RL, especially Deep RL, often works well with immediate rewards but may fail when rewards are delayed or noisy, or both. In this work, we propose delegating the CAP to a Neural Network-based algorithm named InferNet that explicitly learns to infer the immediate rewards from the delayed and noisy rewards. The effectiveness of InferNet was evaluated on three online RL tasks: a GridWorld, a CartPole, and 40 Atari games; and two offline RL tasks: GridWorld and a real-life Sepsis treatment task. The effectiveness of InferNet rewards is compared to that of immediate and delayed rewards in two settings: with and without noise. For the offline RL tasks, it is also compared to a strong baseline, InferGP [7]. Overall, our results show that InferNet is robust to delayed or noisy reward functions, and it could be used effectively for solving the temporal CAP in a wide range of RL tasks, when immediate rewards are not available or they are noisy. Markel Sanz Ausin, Hamoon Azizsoltani, Song Ju, Yeo-Jin Kim, Min Chi |
IEEE BigData | 3 |
| 2021 | To Reduce Healthcare Workload: Identify Critical Sepsis Progression Moments through Deep Reinforcement LearningabstractHealthcare systems are struggling with increasing workloads that adversely affect quality of care and patient outcomes. When clinical practitioners have to make countless medical decisions, they may not always able to make them consistently or spend time on them. In this work, we formulate clinical decision making as a reinforcement learning (RL) problem and propose a human-controlled machine-assisted (HC-MA) decision making framework whereby we can simultaneously give clinical practitioners (the humans) control over the decision-making process while supporting effective decision-making. In our HC-MA framework, the role of the RL agent is to nudge clinicians only if they make suboptimal decisions at critical moments. This framework is supported by a general Critical Deep RL (Critical-DRL) approach, which uses Long-Short Term Rewards (LSTRs) and Critical Deep Q-learning Networks (CriQNs). Critical-DRL’s effectiveness has been evaluated in both a GridWorld game and real-world datasets from two medical systems: a large health system in the northeast of USA, referred as NEMed and Mayo Clinic in Rochester, Minnesota, USA for septic patient treatment. We found that our Critical-DRL approach, by which decisions are made at critical junctures, is as effective as a fully executed DRL policy and moreover, it enables us to identify the critical moments in the septic treatment process, thus greatly reducing burden on medical decision-makers by allowing them to make critical clinical decisions without negatively impacting outcomes. Song Ju, Yeo Jin Kim, Markel Sanz Ausin, Maria E. Mayorga, Min Chi |
IEEE BigData | 1 |
| 2021 | Preparing Unprepared Students For Future Learning
Mark Abdelshiheed, Mehak Maniktala, Song Ju, Tiffany Barnes, Min Chi |
CogSci | 3 |
| 2020 | Pick the Moment: Identifying Critical Pedagogical Decisions Using Long-Short Term Rewards
Song Ju, Min Chi, Guojing Zhou |
EDM | 1 |
| 2019 | Identifying Critical Pedagogical Decisions through Adversarial Deep Reinforcement Learning
Song Ju, Guojing Zhou, Hamoon Azizsoltani, Tiffany Barnes, Min Chi |
EDM | 1 |