EDBT 2026 Demo / reviewers in the wild / expert
Tong Mu
dblp:136/9877
· DBLP profile ↗
14ranked-venue papers
8as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 5 first-author · 6 since 2021Systems, architecture and hardware · 4 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 2 since 2021Computer networks · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Research on multi-platform heterogeneous rumor detection using federated learning and bidirectional graph attention mechanism
Shengcai Zhang, Tong Mu, Dezhi An |
Future Gener. Comput. Syst. | 2 |
| 2025 | Look Before You Leap: Using Serialized State Machine for Language Conditioned Robotic ManipulationabstractImitation learning frameworks for robotic manipulation have drawn attention in the recent development of language model grounded robotics. However, the success of the frameworks largely depends on the coverage of the demonstration cases: When the demonstration set does not include examples of how to act in all possible situations, the action may fail and can result in cascading errors. To solve this problem, we propose a framework that uses serialized Finite State Machine (FSM) to generate demonstrations and improve the success rate in manipulation tasks requiring a long sequence of precise interactions. To validate its effectiveness, we use environmentally evolving and long-horizon puzzles that require long sequential actions. Experimental results show that our approach achieves a success rate of up to 98% in these tasks, compared to the controlled condition using existing approaches, which only had a success rate of up to 60%, and, in some tasks, almost failed completely. The source code for this project can be accessed at https://imitate.finite-state.com/. Tong Mu, Mehran Armand |
IROS | 1 |
| 2024 | Rule Based Rewards for Language Model SafetyabstractReinforcement learning based fine-tuning of large language models (LLMs) on human preferences has been shown to enhance both their capabilities and safety behavior.
However, in cases related to safety, without precise instructions to human annotators, the data collected may cause the model to become overly cautious, or to respond in an undesirable style, such as being judgmental.
Additionally, as model capabilities and usage patterns evolve, there may be a costly need to add or relabel data to modify safety behavior.
We propose a novel preference modeling approach that utilizes AI feedback and only requires a small amount of human data.
Our method, Rule Based Rewards (RBR), uses a collection of rules for desired or undesired behaviors (e.g. refusals should not be judgmental) along with a LLM grader.
In contrast to prior methods using AI feedback, our method uses fine-grained, composable, LLM-graded few-shot prompts as reward directly in RL training, resulting in greater control, accuracy and ease of updating.
We show that RBRs are an effective training method, achieving an F1 score of 97.1, compared to a human-feedback baseline of 91.7, resulting in much higher safety-behavior accuracy through better balancing usefulness and safety. Tong Mu, Alec Helyar, Johannes Heidecke, Joshua Achiam, Andrea Vallone, Ian D. Kivlichan, Molly Lin, Alex Beutel, John Schulman, Lilian Weng |
NeurIPS | 1 |
| 2024 | Reinforcement-Learning-Assisted Multi-UAV Task Allocation and Path Planning for IIoTabstractExploring the widespread applications of unmanned aerial vehicles (UAVs) in Internet of Things has become a current research hotspot. In some tasks related to UAV-based environmental monitoring and transportation, the simultaneous consideration of UAV task allocation and path planning constitutes a category of joint optimization problems. This paper focuses on a warehouse cargo inspection scenario with multiple heterogeneous UAVs. In such scenarios, existing heuristic path finding algorithms that consider task allocation cannot make a good balance between solution time and solution quality. Therefore, in this paper, we propose a reinforcement learning assisted task allocation and conflict-free path framework to achieve better task allocation and path finding results. The framework uses a multiple traveling salesman transformation algorithm for task allocation and a multi-agent reinforcement learning (MARL) algorithm for conflict-free path finding. The path finding policy can be extended to larger-scale environments with more UAVs. We conduct the training of the path planning module and the verification of the overall framework in random environments. Simulation results show that our reinforcement learning assisted framework has a significant advantage over the existing algorithms in terms of solution time, solution quality and scalability. Guodong Zhao 0003, Tong Mu |
IEEE Internet Things J. | 3 |
| 2023 | MC-MLP: A Multiple Coordinate Frames MLP-Like Architecture for Vision
Zhimin Zhu, Tong Mu, Yuliang Yang, Mengyu Zhu |
ICANN (2) | 3 |
| 2023 | Simple Embodied Language Learning as a Byproduct of Meta-Reinforcement LearningabstractWhereas machine learning models typically learn language by directly training on language tasks (e.g., next-word prediction), language emerges in human children as a byproduct of solving non-language tasks (e.g., acquiring food). Motivated by this observation, we ask: can embodied reinforcement learning (RL) agents also indirectly learn language from non-language tasks? Learning to associate language with its meaning requires a dynamic environment with varied language. Therefore, we investigate this question in a multi-task environment with language that varies across the different tasks. Specifically, we design an office navigation environment, where the agent’s goal is to find a particular office, and office locations differ in different buildings (i.e., tasks). Each building includes a floor plan with a simple language description of the goal office’s location, which can be visually read as an RGB image when visited. We find RL agents indeed are able to indirectly learn language. Agents trained with current meta-RL algorithms successfully generalize to reading floor plans with held-out layouts and language phrases, and quickly navigate to the correct office, despite receiving no direct language supervision. Evan Zheran Liu, Sahaana Suri, Tong Mu, Allan Zhou, Chelsea Finn |
ICML | 3 |
| 2023 | A secure and lightweight cloud-centric intelligent medical system based on Internet of Medical Things
Tong Mu, Qiaochuan Ren, Bilin Shao, Genqing Bian |
J. Supercomput. | 1 |
| 2022 | Constraint Sampling Reinforcement Learning: Incorporating Expertise for Faster LearningabstractOnline reinforcement learning (RL) algorithms are often difficult to deploy in complex human-facing applications as they may learn slowly and have poor early performance. To address this, we introduce a practical algorithm for incorporating human insight to speed learning. Our algorithm, Constraint Sampling Reinforcement Learning (CSRL), incorporates prior domain knowledge as constraints/restrictions on the RL policy. It takes in multiple potential policy constraints to maintain robustness to misspecification of individual constraints while leveraging helpful ones to learn quickly. Given a base RL learning algorithm (ex. UCRL, DQN, Rainbow) we propose an upper confidence with elimination scheme that leverages the relationship between the constraints, and their observed performance, to adaptively switch among them. We instantiate our algorithm with DQN-type algorithms and UCRL as base algorithms, and evaluate our algorithm in four environments, including three simulators based on real data: recommendations, educational activity sequencing, and HIV treatment sequencing. In all cases, CSRL learns a good policy faster than baselines. Tong Mu, Georgios Theocharous, David T. Arbour, Emma Brunskill |
AAAI | 1 |
| 2022 | Factored DRO: Factored Distributionally Robust Policies for Contextual BanditsabstractWhile there has been extensive work on learning from offline data for contextual multi-armed bandit settings, existing methods typically assume there is no environment shift: that the learned policy will operate in the same environmental process as that of data collection. However, this assumption may limit the use of these methods for many practical situations where there may be distribution shifts. In this work we propose Factored Distributionally Robust Optimization (Factored-DRO), which is able to separately handle distribution shifts in the context distribution and shifts in the reward generating process. Prior work that either ignores potential shifts in the context, or considers them jointly, can lead to performance that is too conservative, especially under certain forms of reward feedback. Our Factored-DRO objective mitigates this by considering the shifts separately, and our proposed estimators are consistent and converge asymptotically. We also introduce a practical algorithm and demonstrate promising empirical results in environments based on real-world datasets, such as voting outcomes and scene classification. Tong Mu, Yash Chandak, Tatsunori B. Hashimoto, Emma Brunskill |
NeurIPS | 1 |
| 2022 | Attention-Based Lane Change and Crash Risk Prediction Model in HighwaysabstractLane change and crash risk prediction are critical technologies for autonomous driving. An attention-based LSTM model is proposed in this paper for lane change behavior prediction in highways, considering both the surrounding vehicles’ information around the target vehicle at the current time and the historical trajectory data of the vehicle, and it shows higher accuracy and better interpretability than other models. There are two sections in this prediction model, including a pre-judgment model based on the C4.5 decision tree and bagging ensemble learning, and a multi-step lane change prediction model of LSTM with attention mechanism. In addition, the model is applied in the NGSIM datasets to test practicability and accuracy, and the precision of left lane change and right lane change is 98% and 94% respectively 1 second before the vehicle changes lane. Finally, to judge the safety of vehicles in the driving process, a crash risk prediction model based on Time-To-Collision is proposed, which further verifies the effectiveness of the method. Zhen-Ni Li, Xing-Hui Huang, Tong Mu, Jiao Wang 0005 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Automatic Adaptive Sequencing in a Webgame
Tong Mu, Erik Andersen 0001, Emma Brunskill |
ITS | 1 |
| 2020 | Towards Suggesting Actionable Interventions for Wheel Spinning Students
Tong Mu, Andrea Jetten, Emma Brunskill |
EDM | 1 |
| 2018 | Combining adaptivity with progression ordering for intelligent tutoring systemsabstractLearning at scale (LAS) systems like Massive Open Online Classes (MOOCs) have hugely expanded access to high quality educational materials however, such material are frequently time and resource expensive to create. In this work we propose a new approach for automatically and adaptively sequencing practice activities for a particular learner and explore its application for foreign language learning. We evaluate our system through simulation and are in the process of running an experiment. Our simulation results suggest that such an approach may be significantly better than an expert system when there is high variability in the rate of learning among the students and if mastering prerequisites before advancing is important, and is likely to be no worse than an expert system if our generated curriculum approximately describes the necessary structure of learning in students. Tong Mu, Erik Andersen 0001, Emma Brunskill |
L@S | 1 |
| 2017 | Allocating Redundancy Between Erasure Coding and Channel Coding When Fading Channel Diversity Grows With Codeword LengthabstractA transmitter sends a packetized message over a fading channel using packet-level erasure coding and physical-layer channel coding of each resultant packet. Given an overall code rate, this paper finds the optimal rates of the erasure code and the channel code to minimize the transmit power required for a certain message error probability. This paper considers a practically important fading model in which the number of block fades in a transmitted channel codeword increases with the codeword length. Such a model applies, for example, in a time-varying channel with a fixed coherence time. The rate at which diversity grows with codeword length plays an important role in the optimization problem. If the diversity growth factor is large enough, then the erasure code plays a minor role, having an optimal rate that is essentially nondecreasing with decreasing overall rate. We prove analytically that, on a channel with linear growth in diversity, as overall rate decreases, the optimal erasure code rate eventually increases to its maximum possible value (e.g., a rate of 1 for an erasure code with no overhead). Additionally, we also consider the optimization problem of minimizing the message error probability given a transmit power. Numerical results again show that erasure coding is not necessary when overall code rates are sufficiently low. Sudarsan Vasista Srinivasan Ranganathan, Tong Mu, Richard D. Wesel |
IEEE Trans. Commun. | 2 |