VLDB 2026 Research / reviewers in the wild / expert
Letian Chen
dblp:232/1880
· DBLP profile ↗
15ranked-venue papers
3as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RTMol: Rethinking Molecule-text Alignment in a Round-trip ViewabstractAligning molecular sequence representations (e.g., SMILES notations) with textual descriptions is critical for applications spanning drug discovery, materials design, and automated chemical literature analysis. Existing methodologies typically treat molecular captioning (molecule-to-text) and text-based molecular design (text-to-molecule) as separate tasks, relying on supervised fine-tuning or contrastive learning pipelines. These approaches face three key limitations: (i) conventional metrics like BLEU prioritize linguistic fluency over chemical accuracy, (ii) training datasets frequently contain chemically ambiguous narratives with incomplete specifications, and (iii) independent optimization of generation directions leads to bidirectional inconsistency. To address these issues, we propose RTMol, a bidirectional alignment framework that unifies molecular captioning and text-to-SMILES generation through self-supervised round-trip learning. The framework introduces novel round-trip evaluation metrics and enables unsupervised training for molecular captioning without requiring paired molecule-text corpora. Experiments demonstrate that RTMol enhances bidirectional alignment performance by up to 47% across various LLMs, establishing an effective paradigm for joint molecule-text understanding and generation. Letian Chen, Runhan Shi, Gufeng Yu |
AAAI | 1 |
| 2026 | Teaching the Teacher: Live Foundation Model and Augmented Reality Feedback for Human-to-Robot Skill TransferabstractDeploying robots in dynamic, human-populated environments will require techniques for adaptable robot skill acquisition that extend beyond pre-programmed functionality. Learning from demonstration (LfD) methods enable robots to learn skills from human-provided trajectories demonstrated in situ. However, prior work has shown non-expert end-users struggle to provide demonstrations that enable robots to perform complex, multi-step tasks, or to generalize skill knowledge beyond a specific environment and task context. This work enables robots to actively participate in the situated learning interaction by autonomously providing bespoke guidance in response to end-users' demonstrations, thus improving end-users' ability to teach robots useful skills via LfD. We introduce a novel LfD system integrating foundation model (FM)-based textual feedback and augmented reality (AR)-based visual feedback. The FM and AR feedbacks operate synergistically, with FM feedback helping users break tasks down effectively and with AR feedback allowing users to quickly evaluate how well demonstrations perform and generalize. This system provides targeted, actionable guidance throughout the demonstration process: it enhances users' ability to define, decompose, and demonstrate modular, repurposable skills capable of accomplishing complex tasks. We validate our system with a human-subjects experiment in which participants receive bespoke feedback as they teach a robot via kinesthetic demonstrations in a pair of robotic manipulation domains. From this study, we observe positive results demonstrating that the combination of AR and FM feedback improves the quality and generalizability of robot policies, compared to AR feedback alone, FM feedback alone, or a baseline system where learned skills can be played physically on the robot. Nina Moorman, Matthew B. Luebbers, Zhang Xi-Jia, Yee Ching (Marcus) Lau, Yixing Yao, Megan Langwasser, Zulfiqar Zaidi, Letian Chen, Sanne van Waveren, Matthew C. Gombolay |
HRI | 8 |
| 2026 | EPIC: multi-objective guided diffusion for epitope design in TCR-pMHC complexesabstractMOTIVATION: T cell receptor (TCR) recognition of peptide-major histocompatibility complex (pMHC) complexes is central to adaptive immunity, yet rational design of immunogenic epitopes remains elusive due to complex triplet binding constraints and data scarcity. No existing method can generate epitopes satisfying simultaneous requirements for antigenicity, MHC presentation, and TCR specificity. RESULTS: We present EPIC, a multi-objective diffusion framework that decomposes TCR-pMHC binding into three biologically grounded sub-tasks, enabling training-free gradient guidance without end-to-end retraining. By integrating ESM-based classifiers with a peptide diffusion generator, EPIC leverages heterogeneous immunological interaction datasets to generate diverse, context-aware epitopes. EPIC-designed top-three epitopes achieve lower predicted interface energies compared to ground-truth epitopes in 78.31% of test cases, while maintaining 80.1% sequence novelty and comparable structural confidence. Generated epitopes exhibit 100% uniqueness, high diversity (64.05%), and high antigenicity scores (0.4723). To our knowledge, EPIC is the first computational framework capable of de novo epitope design while explicitly integrating the triplet constraints of TCR-pMHC binding. This paradigm shift from discovery to design unlocks new potential for personalized cancer vaccines, precision adoptive T cell therapy, and rapid response to emerging infectious diseases. AVAILABILITY AND IMPLEMENTATION: The source code of EPIC is available at https://github.com/Octopus125/EPIC and archived on Zenodo (DOI: 10.5281/zenodo.18537646). Yueshan Huang, Gufeng Yu, Letian Chen, Haoyang Luan |
Bioinform. | 3 |
| 2026 | Data poisoning-based backdoor attacks against supervised learning rules of Spiking Neural Networks
Lingxin Jin, Wei Jiang 0016, Jinyu Zhan, Meiyu Lin, Letian Chen, Boran Quan, Lin Zuo, Xingzhi Zhou 0001, Maregu Assefa, Naoufel Werghi |
J. Syst. Archit. | 5 |
| 2025 | S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Model with Spatio-Temporal Visual RepresentationabstractThe latest advancements in multi-modal large language models (MLLMs) have spurred a strong renewed interest in end-to-end motion planning approaches for autonomous driving. Many end-to-end approaches rely on human annotations to learn intermediate perception and prediction tasks, while purely self-supervised approaches—which directly learn from sensor inputs to generate planning trajectories without human annotations—often underperform the state of the art. We observe a key gap in the input representation space: end-to-end approaches built on MLLMs are often pretrained with reasoning tasks in 2D image space rather than the native 3D space in which autonomous vehicles plan. To this end, we propose S4-Driver, a scalable self-supervised motion planning algorithm with spatio-temporal visual representation, based on the popular PaLI [9] multimodal large language model. S4-Driver uses a novel sparse volume strategy to seamlessly transform the strong visual representation of MLLMs from perspective view to 3D space without the need to finetune the vision encoder. This representation aggregates multi-view and multi-frame visual inputs and enables better prediction of planning trajectories in 3D space. To validate our method, we run experiments on both nuScenes and Waymo Open Motion Dataset (with in-house camera data). Results show that S4-Driver performs favorably against existing supervised multi-task approaches while requiring no human annotations. It also demonstrates great scalability when pretrained on large volumes of unannotated driving logs. Yichen Xie 0002, Runsheng Xu, Jyh-Jing Hwang, Katie Luo, Jingwei Ji, Hubert Lin, Letian Chen, Yiren Lu 0001, Zhaoqi Leng, Dragomir Anguelov, Mingxing Tan |
CVPR | 8 |
| 2025 | Generalized Behavior Learning from Diverse DemonstrationsabstractDiverse behavior policies are valuable in domains requiring quick test-time adaptation or personalized human-robot interaction. Human demonstrations provide rich information regarding task objectives and factors that govern individual behavior variations, which can be used to characterize \textit{useful} diversity and learn diverse performant policies.
However, we show that prior work that builds naive representations of demonstration heterogeneity fails in generating successful novel behaviors that generalize over behavior factors.
We propose Guided Strategy Discovery (GSD), which introduces a novel diversity formulation based on a learned task-relevance measure that prioritizes behaviors exploring modeled latent factors.
We empirically validate across three continuous control benchmarks for generalizing to in-distribution (interpolation) and out-of-distribution (extrapolation) factors that GSD outperforms baselines in novel behavior discovery by $\sim$21\%.
Finally, we demonstrate that GSD can generalize striking behaviors for table tennis in a virtual testbed while leveraging human demonstrations collected in the real world.
Code is available at https://github.com/CORE-Robotics-Lab/GSD. Varshith Sreeramdass, Rohan R. Paleja, Letian Chen, Sanne van Waveren, Matthew C. Gombolay |
ICLR | 3 |
| 2025 | ELEMENTAL: Interactive Learning from Demonstrations and Vision-Language Models for Reward Design in RoboticsabstractReinforcement learning (RL) has demonstrated compelling performance in robotic tasks, but its success often hinges on the design of complex, ad hoc reward functions. Researchers have explored how Large Language Models (LLMs) could enable non-expert users to specify reward functions more easily. However, LLMs struggle to balance the importance of different features, generalize poorly to out-of-distribution robotic tasks, and cannot represent the problem properly with only text-based descriptions. To address these challenges, we propose ELEMENTAL (intEractive LEarning froM dEmoNstraTion And Language), a novel framework that combines natural language guidance with visual user demonstrations to align robot behavior with user intentions better. By incorporating visual inputs, ELEMENTAL overcomes the limitations of text-only task specifications, while leveraging inverse reinforcement learning (IRL) to balance feature weights and match the demonstrated behaviors optimally. ELEMENTAL also introduces an iterative feedback-loop through self-reflection to improve feature, reward, and policy learning. Our experiment results demonstrate that ELEMENTAL outperforms prior work by 42.3% on task success, and achieves 41.3% better generalization in out-of-distribution tasks, highlighting its robustness in LfD. Letian Chen, Nina Moorman, Matthew C. Gombolay |
ICML | 1 |
| 2025 | Reaction Prediction via Interaction Modeling of Symmetric Difference Shingle SetsabstractChemical reaction prediction remains a fundamental challenge in organic chemistry, where existing machine learning models face two critical limitations: sensitivity to input permutations (molecule/atom orderings) and inadequate modeling of substructural interactions governing reactivity. These shortcomings lead to inconsistent predictions and poor generalization to real-world scenarios. To address these challenges, we propose ReaDISH, a novel reaction prediction model that learns permutation-invariant representations while incorporating interaction-aware features. It introduces two innovations: (1) symmetric difference shingle encoding, which extends the differential reaction fingerprint (DRFP) by representing shingles as continuous high-dimensional embeddings, capturing structural changes while eliminating order sensitivity; and (2) geometry-structure interaction attention, a mechanism that models intra- and inter-molecular interactions at the shingle level. Extensive experiments demonstrate that ReaDISH improves reaction prediction performance across diverse benchmarks. It shows enhanced robustness with an average improvement of 8.76\% on R$^2$ under permutation perturbations. Runhan Shi, Letian Chen, Gufeng Yu |
NeurIPS | 2 |
| 2025 | Design RNA with Specified Secondary Structures Using CDMsabstractAs RNA function is strongly tied to its secondary structure, designing RNA molecules with specified structures stands as a key challenge in computational biology. Existing approaches often either yield limited results or incur high computational costs. Building on the success of Denoising Diffusion Probabilistic Models (DDPMs) and their variants, including Conditional Diffusion Models (CDMs) in diverse fields such as image generation, we apply CDMs to generate RNA sequences matching specified secondary structures, using One-Hot or static vocabulary encoding, classifier-free guidance, and a Transformer denoiser to capture long-range dependencies. Experiments on the Rfam dataset show that our best model achieves a task-solving rate of 0.846, outperforming the current next-best method (≤ 0.648). Our models exhibit better performance with a simpler architecture, demonstrating the effectiveness of CDMs in sequence design problems. Future work may focus on optimizing the conditional encoding method and the architecture of the denoiser, as well as developing biological relevance metrics for generated sequences. Zhenran Xiao, Letian Chen, Yichong Li, Yang Yang 0030 |
SMC | 2 |
| 2024 | Enhancing Safety in Learning from Demonstration Algorithms via Control Barrier Function ShieldingabstractLearning from Demonstration (LfD) is a powerful method for non-roboticists end-users to teach robots new tasks, enabling them to customize the robot behavior. However, modern LfD techniques do not explicitly synthesize safe robot behavior, which limits the deployability of these approaches in the real world. To enforce safety in LfD without relying on experts, we propose a new framework, SElding with Control barrier fUnctions in inverse REinforcement learning (SECURE), which learns a customized Control Barrier Function (CBF) from end-users that prevents robots from taking unsafe actions while imposing little interference with the task completion. We evaluate SECURE in three sets of experiments. First, we empirically validate SECURE learns a high-quality CBF from demonstrations and outperforms conventional LfD methods on simulated robotic and autonomous driving tasks with improvements on safety by up to 100%. Second, we demonstrate that roboticists can leverage SECURE to outperform conventional LfD approaches on a real-world knife-cutting, meal-preparation task by 12.5% in task completion while driving the number of safety violations to zero. Finally, we demonstrate in a user study that non-roboticists can use SECURE to effectively teach the robot safe policies that avoid collisions with the person and prevent coffee from spilling. Yue Yang 0024, Letian Chen, Zulfiqar Zaidi, Sanne van Waveren, Arjun Krishna, Matthew C. Gombolay |
HRI | 2 |
| 2023 | The Effect of Robot Skill Level and Communication in Rapid, Proximate Human-Robot CollaborationabstractAs high-speed, agile robots become more commonplace, these robots will have the potential to better aid and collaborate with humans. However, due to the increased agility and functionality of these robots, close collaboration with humans can create safety concerns that alter team dynamics and degrade task performance. In this work, we aim to enable the deployment of safe and trustworthy agile robots that operate in proximity with humans. We do so by 1) Proposing a novel human-robot doubles table tennis scenario to serve as a testbed for studying agile, proximate human-robot collaboration and 2) Conducting a user-study to understand how attributes of the robot (e.g., robot competency or capacity to communicate) impact team dynamics, perceived safety, and perceived trust, and how these latent factors affect human-robot collaboration (HRC) performance. We find that robot competency significantly increases perceived trust (p < .001), extending skill-to-trust assessments in prior studies to agile, proximate HRC. Furthermore, interestingly, we find that when the robot vocalizes its intention to perform a task, it results in a significant decrease in team performance (p = .037) and perceived safety of the system (p = .009). Kin Man Lee, Arjun Krishna, Zulfiqar Zaidi, Rohan R. Paleja, Letian Chen, Erin Hedlund-Botti, Mariah Schrum, Matthew C. Gombolay |
HRI | 5 |
| 2023 | Learning Models of Adversarial Agent Behavior Under Partial ObservabilityabstractThe need for opponent modeling and tracking arises in several real-world scenarios, such as professional sports, video game design, and drug-trafficking interdiction. In this work, we present Graph based Adversarial Modeling with Mutual Information (GrAMMI) for modeling the behavior of an adversarial opponent agent. GrAMMI is a novel graph neural network (GNN) based approach that uses mutual information maximization as an auxiliary objective to predict the current and future states of an adversarial opponent with partial observability. To evaluate GrAMMI, we design two large-scale, pursuit-evasion domains inspired by real-world scenarios, where a team of heterogeneous agents is tasked with tracking and interdicting a single adversarial agent, and the adversarial agent must evade detection while achieving its own objectives. With the mutual information formulation, GrAMMI outperforms all baselines in both domains and achieves 31.68% higher log-likelihood on average for future adversarial state predictions across both domains. Sean Ye, Manisha Natarajan, Rohan R. Paleja, Letian Chen, Matthew C. Gombolay |
IROS | 5 |
| 2022 | A Hierarchical Coordination Framework for Joint Perception-Action Tasks in Composite Robot TeamsabstractWe propose a collaborative planning and control algorithm to enhance cooperation for composite teams of autonomous robots in dynamic environments. Composite robot teams are groups of agents that perform different tasks according to their respective capabilities in order to accomplish an overarching mission. Examples of such teams include groups of perception agents (can only sense) and action agents (can only manipulate) working together to perform disaster response tasks. Coordinating robots in a composite team is a challenging problem due to the heterogeneity in the robots’ characteristics and their tasks. Here, we propose a coordination framework for composite robot teams. The proposed framework consists of two hierarchical modules: First, A multiagent state-action-reward-time-state-action algorithm in multiagent partially observable semi-Markov decision process as the high-level decision-making module to enable perception agents to learn to surveil in an environment with an unknown number of dynamic targets and second, a low-level coordinated control and planning module that ensures probabilistically guaranteed support for action agents. Simulation and physical robot implementations of our algorithms on a multiagent robot testbed demonstrated the efficacy and feasibility of our coordination framework by reducing the overall operation times in a benchmark wildfire-fighting case study. Esmaeil Seraj, Letian Chen, Matthew C. Gombolay |
IEEE Trans. Robotics | 2 |
| 2020 | Joint Goal and Strategy Inference across Heterogeneous Demonstrators via Reward Network DistillationabstractReinforcement learning (RL) has achieved tremendous success as a general framework for learning how to make decisions. However, this success relies on the interactive hand-tuning of a reward function by RL experts. On the other hand, inverse reinforcement learning (IRL) seeks to learn a reward function from readily-obtained human demonstrations. Yet, IRL suffers from two major limitations: 1) reward ambiguity - there are an infinite number of possible reward functions that could explain an expert's demonstration and 2) heterogeneity - human experts adopt varying strategies and preferences, which makes learning from multiple demonstrators difficult due to the common assumption that demonstrators seeks to maximize the same reward. In this work, we propose a method to jointly infer a task goal and humans' strategic preferences via network distillation. This approach enables us to distill a robust task reward (addressing reward ambiguity) and to model each strategy's objective (handling heterogeneity). We demonstrate our algorithm can better recover task reward and strategy rewards and imitate the strategies in two simulated tasks and a real-world table tennis task. Letian Chen, Rohan R. Paleja, Muyleng Ghuy, Matthew C. Gombolay |
HRI | 1 |
| 2020 | Interpretable and Personalized Apprenticeship Scheduling: Learning Interpretable Scheduling Policies from Heterogeneous User DemonstrationsabstractResource scheduling and coordination is an NP-hard optimization requiring an efficient allocation of agents to a set of tasks with upper- and lower bound temporal and resource constraints. Due to the large-scale and dynamic nature of resource coordination in hospitals and factories, human domain experts manually plan and adjust schedules on the fly. To perform this job, domain experts leverage heterogeneous strategies and rules-of-thumb honed over years of apprenticeship. What is critically needed is the ability to extract this domain knowledge in a heterogeneous and interpretable apprenticeship learning framework to scale beyond the power of a single human expert, a necessity in safety-critical domains. We propose a personalized and interpretable apprenticeship scheduling algorithm that infers an interpretable representation of all human task demonstrators by extracting decision-making criteria via an inferred, personalized embedding non-parametric in the number of demonstrator types. We achieve near-perfect LfD accuracy in synthetic domains and 88.22\% accuracy on a planning domain with real-world data, outperforming baselines. Finally, our user study showed our methodology produces more interpretable and easier-to-use models than neural networks ($p < 0.05$). Rohan R. Paleja, Andrew Silva, Letian Chen, Matthew C. Gombolay |
NeurIPS | 3 |