VLDB 2026 Research / reviewers in the wild / expert
Xiao Ma 0006
dblp:35/573-6
· DBLP profile ↗
19ranked-venue papers
7as first author
13since 2021 · last 2025
0000-0001-5466-867XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ProcWorld: Benchmarking Large Model Planning in Reachability-Constrained EnvironmentsabstractWe introduce ProcWORLD, a large-scale benchmark for partially observable embodied spatial reasoning and long-term planning with large language models (LLM) and vision language models (VLM).ProcWORLD features a wide range of challenging embodied navigation and object manipulation tasks, covering 16 task types, 5,000 rooms, and over 10 million evaluation trajectories with diverse data distribution.ProcWORLD supports configurable observation modes, ranging from text-only descriptions to vision-only observations.It enables text-based actions to control the agent following language instructions.ProcWORLD has presented significant challenges for LLMs and VLMs: (1) active information gathering given partial observations for disambiguation; (2) simultaneous localization and decision-making by tracking the spatio-temporal state-action distribution; (3) constrained reasoning with dynamic states subject to physical reachability.Our extensive evaluation of 15 foundation models and 5 reasoning algorithms (with over 1 million rollouts) indicates larger models perform better.However, ProcWORLD remains highly challenging for existing state-of-the-art models and in-context learning methods due to constrained reachability and the need of combinatorial spatial reasoning. Xinghang Li, Zhengshen Zhang, Jirong Liu, Xiao Ma 0006, Hanbo Zhang, Tao Kong, Huaping Liu 0001 |
EMNLP | 5 |
| 2025 | CODA: Repurposing Continuous VAEs for Discrete Tokenization
Zanlin Ni, Yeguo Hua, Xiao Ma 0006, Gao Huang 0001 |
ICCV | 5 |
| 2025 | Decoupled Prioritized Resampling for Offline RLabstractOffline reinforcement learning (RL) is challenged by the distributional shift problem. To tackle this issue, existing works mainly focus on designing sophisticated policy constraints between the learned policy and the behavior policy. However, these constraints are applied equally to well-performing and inferior actions through uniform sampling, which might negatively affect the learned policy. In this article, we propose offline decoupled prioritized resampling (ODPR), which designs specialized priority functions for the suboptimal policy constraint issue in offline RL and employs unique decoupled resampling for training stability. Through theoretical analysis, we show that the distinctive priority functions induce a provable improved behavior policy by modifying the distribution of the original behavior policy, and when constrained to this improved policy, a policy-constrained offline RL algorithm is likely to yield a better solution. We provide two practical implementations to balance computation and performance: one estimates priorities based on a fit value network [advantage-based ODPR (ODPR-A)] and the other utilizes trajectory returns [return-based ODPR (ODPR-R)] for quick computation. As a highly compatible plug-and-play component, ODPR is evaluated with five prevalent offline RL algorithms: behavior cloning (BC), twin delayed deep deterministic policy gradient + BC (TD3 + BC), OnestepRL, conservative Q-learning (CQL), and implicit Q-learning (IQL). Our experiments confirm that both ODPR-A and ODPR-R significantly improve performance across all baseline methods. Moreover, ODPR-A can be effective in some challenging settings, i.e., without trajectory information. Code and pretrained weights are available at https://github.com/yueyang130/ODPR. Bingyi Kang, Xiao Ma 0006, Qisen Yang, Gao Huang 0001, Shiji Song, Shuicheng Yan |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Hierarchical Diffusion Policy for Kinematics-Aware Multi-Task Robotic ManipulationabstractThis paper introduces Hierarchical Diffusion Policy (HDP), a hierarchical agent for multi-task robotic manipulation. HDP factorises a manipulation policy into a hierarchical structure: a high-level task-planning agent which predicts a distant next-best end-effector pose (NBP), and a low-level goal-conditioned diffusion policy which generates optimal motion trajectories. The factorised policy representation allows HDP to tackle both long-horizon task planning while generating fine-grained low-level actions. To generate context-aware motion trajectories while satisfying robot kinematics constraints, we present a novel kinematics-aware goal-conditioned control agent, Robot Kinematics Diffuser (RK-Diffuser). Specifically, RK-Diffuser learns to generate both the end-effector pose and joint position trajectories, and distill the accurate but kinematics-unaware end-effector pose diffuser to the kinematics-aware but less accurate joint position diffuser via differentiable kinematics. Empirically, we show that HDP achieves a significantly higher success rate than the state-of-the-art methods in both simulation and real-world.11Code and videos are available in our project page. Xiao Ma 0006, Sumit Patidar, Iain Haughton, Stephen James |
CVPR | 1 |
| 2023 | Imitation Learning as State Matching via Differentiable PhysicsabstractExisting imitation learning (IL) methods such as inverse reinforcement learning (IRL) usually have a double-loop training process, alternating between learning a reward function and a policy and tend to suffer long training time and high variance. In this work, we identify the benefits of differentiable physics simulators and propose a new IL method, i.e., Imitation Learning as State Matching via Differentiable Physics (ILD), which gets rid of the double-loop design and achieves significant improvements in final performance, convergence speed, and stability. The proposed ILD incorporates the differentiable physics simulator as a physics prior into its computational graph for policy learning. ILD unrolls the dynamics by sampling actions from a parameterized policy and minimizing the distance between the expert trajectory and the agent trajectory. It back-propagates the gradient into the policy via temporal physics operators, which improves the transferability to unseen environments and yields higher final performance. ILD has a single-loop structure that stabilizes and speeds up training. It dynamically selects learning objectives for each state during optimization to simplify the complex optimization land-scape. Experiments show that ILD outperforms state-of-the-art methods in continuous control tasks with Brax, and can be applied to deformable object manipulation tasks, generalized to unseen configurations.11The link to the code: https://github.com/sail-sg/ILD Siwei Chen 0003, Xiao Ma 0006, Zhongwen Xu |
CVPR | 2 |
| 2023 | RPM: Generalizable Multi-Agent Policies for Multi-Agent Reinforcement Learning
Wei Qiu 0001, Xiao Ma 0006, Bo An 0001, Svetlana Obraztsova, Shuicheng Yan, Zhongwen Xu |
ICLR | 2 |
| 2023 | DaxBench: Benchmarking Deformable Object Manipulation with Differentiable Physics
Siwei Chen 0003, Yiqing Xu, Cunjun Yu, Linfeng Li 0001, Xiao Ma 0006, Zhongwen Xu, David Hsu |
ICLR | 5 |
| 2023 | DiffMimic: Efficient Motion Mimicking with Differentiable Physics
Jiawei Ren 0001, Cunjun Yu, Siwei Chen 0003, Xiao Ma 0006, Liang Pan, Ziwei Liu 0002 |
ICLR | 4 |
| 2023 | Mutual Information Regularized Offline Reinforcement LearningabstractThe major challenge of offline RL is the distribution shift that appears when out-of-distribution actions are queried, which makes the policy improvement direction biased by extrapolation errors. Most existing methods address this problem by penalizing the policy or value for deviating from the behavior policy during policy improvement or evaluation. In this work, we propose a novel MISA framework to approach offline RL from the perspective of Mutual Information between States and Actions in the dataset by directly constraining the policy improvement direction. MISA constructs lower bounds of mutual information parameterized by the policy and Q-values. We show that optimizing this lower bound is equivalent to maximizing the likelihood of a one-step improved policy on the offline dataset. Hence, we constrain the policy improvement direction to lie in the data manifold. The resulting algorithm simultaneously augments the policy evaluation and improvement by adding mutual information regularizations. MISA is a general framework that unifies conservative Q-learning (CQL) and behavior regularization methods (e.g., TD3+BC) as special cases. We introduce 3 different variants of MISA, and empirically demonstrate that tighter mutual information lower bound gives better offline RL performance. In addition, our extensive experiments show MISA significantly outperforms a wide range of baselines on various tasks of the D4RL benchmark, e.g., achieving 742.9 total points on gym-locomotion tasks. Our code is attached and will be released upon publication. Xiao Ma 0006, Bingyi Kang, Zhongwen Xu, Shuicheng Yan |
NeurIPS | 1 |
| 2023 | Efficient Diffusion Policies For Offline Reinforcement LearningabstractOffline reinforcement learning (RL) aims to learn optimal policies from offline datasets, where the parameterization of policies is crucial but often overlooked. Recently, Diffsuion-QL significantly boosts the performance of offline RL by representing a policy with a diffusion model, whose success relies on a parametrized Markov Chain with hundreds of steps for sampling. However, Diffusion-QL suffers from two critical limitations. 1) It is computationally inefficient to forward and backward through the whole Markov chain during training. 2) It is incompatible with maximum likelihood-based RL algorithms (e.g., policy gradient methods) as the likelihood of diffusion models is intractable. Therefore, we propose efficient diffusion policy (EDP) to overcome these two challenges. EDP approximately constructs actions from corrupted ones at training to avoid running the sampling chain. We conduct extensive experiments on the D4RL benchmark. The results show that EDP can reduce the diffusion policy training time from 5 days to 5 hours on gym-locomotion tasks. Moreover, we show that EDP is compatible with various offline RL algorithms (TD3, CRR, and IQL) and achieves new state-of-the-art on D4RL by large margins over previous methods. Bingyi Kang, Xiao Ma 0006, Tianyu Pang, Shuicheng Yan |
NeurIPS | 2 |
| 2023 | InsActor: Instruction-driven Physics-based CharactersabstractGenerating animation of physics-based characters with intuitive control has long been a desirable task with numerous applications. However, generating physically simulated animations that reflect high-level human instructions remains a difficult problem due to the complexity of physical environments and the richness of human language.
In this paper, we present $\textbf{InsActor}$, a principled generative framework that leverages recent advancements in diffusion-based human motion models to produce instruction-driven animations of physics-based characters.
Our framework empowers InsActor to capture complex relationships between high-level human instructions and character motions by employing diffusion policies for flexibly conditioned motion planning.
To overcome invalid states and infeasible state transitions in planned motions, InsActor discovers low-level skills and maps plans to latent skill sequences in a compact latent space.
Extensive experiments demonstrate that InsActor achieves state-of-the-art results on various tasks, including instruction-driven motion generation and instruction-driven waypoint heading. Notably, the ability of InsActor to generate physically simulated animations using high-level human instructions makes it a valuable tool, particularly in executing long-horizon tasks with a rich set of instructions. Our project page is available at [jiawei-ren.github.io/projects/insactor/index.html](https://jiawei-ren.github.io/projects/insactor/index.html) Jiawei Ren 0001, Cunjun Yu, Xiao Ma 0006, Liang Pan, Ziwei Liu 0002 |
NeurIPS | 4 |
| 2022 | Learning Latent Graph Dynamics for Visual Manipulation of Deformable ObjectsabstractManipulating deformable objects, such as ropes and clothing, is a long-standing challenge in robotics, because of their large degrees of freedom, complex non-linear dynamics, and self-occlusion in visual perception. The key difficulty is a suitable representation, rich enough to capture the object shape, dynamics for manipulation and yet simple enough to be estimated reliably from visual observations. This work aims to learn latent Graph dynamics for DefOrmable Object Manipulation (G-DOOM). G-DOOM approximates a deformable object as a sparse set of interacting keypoints, which are extracted automatically from images via unsupervised learning. It learns a graph neural network that captures abstractly the geometry and the interaction dynamics of the keypoints. To handle object self-occlusion, G-DOOM uses a recurrent neural network to track the keypoints over time and condition their interactions on the history. We then train the resulting recurrent graph dynamics model through contrastive learning in a high-fidelity simulator. For manipulation planning, G-DOOM reasons explicitly about the learned dynamics model through model-predictive control applied at each keypoint. Preliminary experiments of G-DOOM on a set of challenging rope and cloth manipulation tasks indicate strong performance, compared with state-of-the-art methods. Although trained in a simulator, G-DOOM transfers directly to a real robot for both rope and cloth manipulation11Demo video available online at https://youtu.be/oCfbNMx2sQI. Xiao Ma 0006, David Hsu, Wee Sun Lee |
ICRA | 1 |
| 2022 | Predicting Hot Events in the Early Period through Bayesian Model for Social NetworksabstractPredicting emerging hot events in an early stage is essential for various applications, including information dissemination mining, ads recommendation and etc. Existing techniques either require a long-term observation over the event or features that are expensive to extract. However, given limited data at the early stage of an emerging event, the temporal features of hot events and non-hot events are not distinctive enough yet. In this work, we introduce BEEP, a Bayesian perspective Early stage Event Prediction model, that tackles this dilemma. We formulate the hot event prediction problem by two Semi-Naive Bayes Classifiers, where we consider both the temporal features and structural features and perform distribution test for the selected features. Theoretical analysis and extensive empirical evaluations on two real datasets demonstrate the effectiveness of our methods. Zuowu Zheng, Xiaofeng Gao 0001, Xiao Ma 0006, Guihai Chen |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2020 | Particle Filter Recurrent Neural NetworksabstractRecurrent neural networks (RNNs) have been extraordinarily successful for prediction with sequential data. To tackle highly variable and multi-modal real-world data, we introduce Particle Filter Recurrent Neural Networks (PF-RNNs), a new RNN family that explicitly models uncertainty in its internal structure: while an RNN relies on a long, deterministic latent state vector, a PF-RNN maintains a latent state distribution, approximated as a set of particles. For effective learning, we provide a fully differentiable particle filter algorithm that updates the PF-RNN latent state distribution according to the Bayes rule. Experiments demonstrate that the proposed PF-RNNs outperform the corresponding standard gated RNNs on a synthetic robot localization dataset and 10 real-world sequence prediction datasets for text classification, stock price prediction, etc. Xiao Ma 0006, Péter Karkus, David Hsu, Wee Sun Lee |
AAAI | 1 |
| 2020 | Spatio-Temporal Graph Transformer Networks for Pedestrian Trajectory Prediction
Cunjun Yu, Xiao Ma 0006, Jiawei Ren 0001, Haiyu Zhao, Shuai Yi |
ECCV (12) | 2 |
| 2020 | Discriminative Particle Filter Reinforcement Learning for Complex Partial observations
Xiao Ma 0006, Péter Karkus, David Hsu, Wee Sun Lee |
ICLR | 1 |
| 2020 | Balanced Meta-Softmax for Long-Tailed Visual RecognitionabstractDeep classifiers have achieved great success in visual recognition. However, real-world data is long-tailed by nature, leading to the mismatch between training and testing distributions. In this paper, we show that the Softmax function, though used in most classification tasks, gives a biased gradient estimation under the long-tailed setup. This paper presents Balanced Softmax, an elegant unbiased extension of Softmax, to accommodate the label distribution shift between training and testing. Theoretically, we derive the generalization bound for multiclass Softmax regression and show our loss minimizes the bound. In addition, we introduce Balanced Meta-Softmax, applying a complementary Meta Sampler to estimate the optimal class sample rate and further improve long-tailed learning. In our experiments, we demonstrate that Balanced Meta-Softmax outperforms state-of-the-art long-tailed classification solutions on both visual recognition and instance segmentation tasks. Jiawei Ren 0001, Cunjun Yu, Shunan Sheng, Xiao Ma 0006, Haiyu Zhao, Shuai Yi, Hongsheng Li 0001 |
NeurIPS | 4 |
| 2017 | Trust-based time series data model for mobile crowdsensingabstractThe recent proliferation of mobile devices embedded with capable sensors, provides an opportunity to the popular concept of mobile crowdsensing. By studying the correlation of crowd-sensed data in both spatial and temporal dimensions, we can get a clear understanding of the intrinsic pattern of data in mobile crowdsensing, which is the basic for further data analysis, such as data filtering, smoothing and prediction. However, the crowd-sensed data are normally noise and unreliable due to the diverse mobility patterns and selfish behaviours of mobile users, making the classical data models in wireless sensor networks fail in this new context. In this paper, we propose a robust and reliable time series data model based on Dynamic Bayesian Network to describe the characteristics of the crowd-sensed data. The proposed data model can figure out the spatial and temporal correlation of data in the environment, where the data has high noise levels and mobile users are untrustworthy. We conduct extensive evaluations based on both simulation and a real-world data set. Our evaluation results show that our method successfully modeled the crowd-sensed time series data with effectiveness, efficiency and trustworthiness. Xiao Ma 0006, Zhenzhe Zheng 0001, Fan Wu 0006, Guihai Chen |
ICC | 1 |
| 2017 | BEEP: A Bayesian Perspective Early Stage Event Prediction Model for Online Social NetworksabstractIn recent years, predicting future hot events in online social networks is becoming increasingly meaningful in marketing, advertisement, and recommendation systems to support companies' strategy making. Currently, most prediction models require long-term observations over the event or depend a lot on other features which are expensive to extract. However, at the early stage of an event, the temporal features of hot events and non-hot events are not distinctive yet. Besides, given the small amount of available data, high noise and complex network structure, those state-of-art models are unable to give an accurate prediction at the very early stage of an event. Hence, we propose two Bayesian perspective models to handle this dilemma. We first mathematically define the hot event prediction problem and introduce the general early stage event prediction framework, then model the five selected features into several continuous distributions, and present two Semi-Naive Bayes Classifier based prediction models, BEEP and SimBEEP, which is the simplified version of BEEP. Extensive experiments on real dataset have demonstrated that our model significantly outperforms the baseline methods. Xiao Ma 0006, Xiaofeng Gao 0001, Guihai Chen |
ICDM | 1 |