VLDB 2026 Research / reviewers in the wild / expert
Zhou Xian
dblp:258/5020
· DBLP profile ↗
12ranked-venue papers
2as first author
10since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | One-Shot Video Imitation via Parameterized Symbolic Abstraction GraphsabstractLearning to manipulate dynamic and deformable objects from a single demonstration video holds great promise in terms of scalability. Previous approaches have predominantly focused on either replaying object relationships or actor trajectories. The former often struggles to generalize across diverse tasks, while the latter suffers from data inefficiency. Moreover, both methodologies encounter challenges in capturing invisible physical attributes, such as forces. In this paper, we propose to interpret video demonstrations through a series of Parameterized Symbolic Abstraction Graphs (PSAGs), where nodes represent objects and edges denote relationships between objects. We further ground geometric constraints through simulation to estimate non-geometric, visually imperceptible attributes. The augmented PSAGs are then applied in real robot experiments. Our approach has been validated across a range of tasks, such as Cutting Avocado, Cutting Vegetable, Pouring Liquid, Rolling Dough, and Slicing Pizza. We demonstrate successful generalization to novel objects with distinct visual and physical properties. For visualizations of the learned policies please check: https://www.jianrenw.com/PSAG/ Jianren Wang, Kangni Liu, Dingkun Guo, Zhou Xian, Christopher G. Atkeson |
ICRA | 4 |
| 2024 | DIFFTACTILE: A Physics-based Differentiable Tactile Simulator for Contact-rich Robotic ManipulationabstractWe introduce DIFFTACTILE, a physics-based differentiable tactile simulation system designed to enhance robotic manipulation with dense and physically accurate tactile feedback. In contrast to prior tactile simulators which primarily focus on manipulating rigid bodies and often rely on simplified approximations to model stress and deformations of materials in contact, DIFFTACTILE emphasizes physics-based contact modeling with high fidelity, supporting simulations of diverse contact modes and interactions with objects possessing a wide range of material properties. Our system incorporates several key components, including a Finite Element Method (FEM)-based soft body model for simulating the sensing elastomer, a multi-material simulator for modeling diverse object types (such as elastic, elastoplastic, cables) under manipulation, a penalty-based contact model for handling contact dynamics. The differentiable nature of our system facilitates gradient-based optimization for both 1) refining physical properties in simulation using real-world data, hence narrowing the sim-to-real gap and 2) efficient learning of tactile-assisted grasping and contact-rich manipulation skills. Additionally, we introduce a method to infer the optical response of our tactile sensor to contact using an efficient pixel-based neural module. We anticipate that DIFFTACTILE will serve as a useful platform for studying contact-rich manipulations, leveraging the benefits of dense tactile feedback and differentiable physics. Code and supplementary materials are available at the project website https://difftactile.github.io/. Zilin Si, Gu Zhang, Qingwei Ben, Branden Romero, Zhou Xian, Chuang Gan 0001 |
ICLR | 5 |
| 2024 | Thin-Shell Object Manipulations With Differentiable Physics SimulationsabstractIn this work, we aim to teach robots to manipulate various thin-shell materials.
Prior works studying thin-shell object manipulation mostly rely on heuristic policies or learn policies from real-world video demonstrations, and only focus on limited material types and tasks (e.g., cloth unfolding). However, these approaches face significant challenges when extended to a wider variety of thin-shell materials and a diverse range of tasks.
On the other hand, while virtual simulations are shown to be effective in diverse robot skill learning and evaluation, prior thin-shell simulation environments only support a subset of thin-shell materials, which also limits their supported range of tasks.
To fill in this gap, we introduce ThinShellLab - a fully differentiable simulation platform tailored for robotic interactions with diverse thin-shell materials possessing varying material properties, enabling flexible thin-shell manipulation skill learning and evaluation. Building on top of our developed simulation engine, we design a diverse set of manipulation tasks centered around different thin-shell objects. Our experiments suggest that manipulating thin-shell objects presents several unique challenges: 1) thin-shell manipulation relies heavily on frictional forces due to the objects' co-dimensional nature, 2) the materials being manipulated are highly sensitive to minimal variations in interaction actions, and 3) the constant and frequent alteration in contact pairs makes trajectory optimization methods susceptible to local optima, and neither standard reinforcement learning algorithms nor trajectory optimization methods (either gradient-based or gradient-free) are able to solve the tasks alone. To overcome these challenges, we present an optimization scheme that couples sampling-based trajectory optimization and gradient-based optimization, boosting both learning efficiency and converged performance across various proposed tasks. In addition, the differentiable nature of our platform facilitates a smooth sim-to-real transition. By tuning simulation parameters with a minimal set of real-world data, we demonstrate successful deployment of the learned skills to real-robot settings. ThinShellLab will be publicly available. Video demonstration and more information can be found on the project website https://vis-www.cs.umass.edu/ThinShellLab/. Yian Wang 0001, Juntian Zheng, Zhehuan Chen, Zhou Xian, Gu Zhang, Chuang Gan 0001 |
ICLR | 4 |
| 2024 | RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model FeedbackabstractReward engineering has long been a challenge in Reinforcement Learning (RL) research, as it often requires extensive human effort and iterative processes of trial-and-error to design effective reward functions. In this paper, we propose RL-VLM-F, a method that automatically generates reward functions for agents to learn new tasks, using only a text description of the task goal and the agent's visual observations, by leveraging feedbacks from vision language foundation models (VLMs). The key to our approach is to query these models to give preferences over pairs of the agent's image observations based on the text description of the task goal, and then learn a reward function from the preference labels, rather than directly prompting these models to output a raw reward score, which can be noisy and inconsistent. We demonstrate that RL-VLM-F successfully produces effective rewards and policies across various domains — including classic control, as well as manipulation of rigid, articulated, and deformable objects — without the need for human supervision, outperforming prior methods that use large pretrained models for reward generation under the same assumptions. Videos can be found on our project website: https://rlvlmf2024.github.io/ Yufei Wang 0007, Zhanyi Sun, Jesse Zhang, Zhou Xian, Erdem Biyik, David Held, Zackory Erickson |
ICML | 4 |
| 2024 | RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative SimulationabstractWe present RoboGen, a generative robotic agent that automatically learns diverse robotic skills at scale via generative simulation. RoboGen leverages the latest advancements in foundation and generative models. Instead of directly adapting these models to produce policies or low-level actions, we advocate for a generative scheme, which uses these models to automatically generate diversified tasks, scenes, and training supervisions, thereby scaling up robotic skill learning with minimal human supervision. Our approach equips a robotic agent with a self-guided propose-generate-learn cycle: the agent first proposes interesting tasks and skills to develop, and then generates simulation environments by populating pertinent assets with proper spatial configurations. Afterwards, the agent decomposes the proposed task into sub-tasks, selects the optimal learning approach (reinforcement learning, motion planning, or trajectory optimization), generates required training supervision, and then learns policies to acquire the proposed skill. Our fully generative pipeline can be queried repeatedly, producing an endless stream of skill demonstrations associated with diverse tasks and environments. Yufei Wang 0007, Zhou Xian, Tsun-Hsuan Wang, Yian Wang 0001, Katerina Fragkiadaki, Zackory Erickson, David Held, Chuang Gan 0001 |
ICML | 2 |
| 2024 | Gen2Sim: Scaling up Robot Learning in Simulation with Generative ModelsabstractGeneralist robot manipulators need to learn a wide variety of manipulation skills across diverse environments. Current robot training pipelines rely on humans to provide kinesthetic demonstrations or to program simulation environments and to code up reward functions for reinforcement learning. Such human involvement is an important bottleneck towards scaling up robot learning across diverse tasks and environments. We propose Generation to Simulation (Gen2Sim), a method for scaling up robot skill learning in simulation by automating generation of 3D assets, task descriptions, task decompositions and reward functions using large pre-trained generative models of language and vision. We generate 3D assets for simulation by lifting open-world 2D object-centric images to 3D using image diffusion models and querying LLMs to determine plausible physics parameters. Given URDF files of generated and human-developed assets, we chain-of-thought prompt LLMs to map these to relevant task descriptions, temporal decompositions, and corresponding python reward functions for reinforcement learning. We show Gen2Sim succeeds in learning policies for diverse long horizon tasks, where reinforcement learning with non temporally decomposed reward functions fails. Gen2Sim provides a viable path for scaling up reinforcement learning for robot manipulators in simulation, both by diversifying and expanding task and environment development, and by facilitating the discovery of reinforcement-learned behaviors through temporal task decomposition in RL. Our work contributes hundreds of simulated assets, tasks and demonstrations, taking a step towards fully autonomous robotic manipulation skill acquisition in simulation. Pushkal Katara, Zhou Xian, Katerina Fragkiadaki |
ICRA | 2 |
| 2024 | Architect: Generating Vivid and Interactive 3D Scenes with Hierarchical 2D InpaintingabstractCreating large-scale interactive 3D environments is essential for the development of Robotics and Embodied AI research. However, generating diverse embodied environments with realistic detail and considerable complexity remains a significant challenge. Current methods, including manual design, procedural generation, diffusion-based scene generation, and large language model (LLM) guided scene design, are hindered by limitations such as excessive human effort, reliance on predefined rules or training datasets, and limited 3D spatial reasoning ability. Since pre-trained 2D image generative models better capture scene and object configuration than LLMs, we address these challenges by introducing $\textit{Architect}$, a generative framework that creates complex and realistic 3D embodied environments leveraging diffusion-based 2D image inpainting. In detail, we utilize foundation visual perception models to obtain each generated object from the image and leverage pre-trained depth estimation models to lift the generated 2D image to 3D space. While there are still challenges that the camera parameters and scale of depth are still absent in the generated image, we address those problems by ''controlling'' the diffusion model by $\textit{hierarchical inpainting}$. Specifically, having access to ground-truth depth and camera parameters in simulation, we first render a photo-realistic image of only the background. Then, we inpaint the foreground in this image, passing the geometric cues to the inpainting model in the background, which informs the camera parameters.
This process effectively controls the camera parameters and depth scale for the generated image, facilitating the back-projection from 2D image to 3D point clouds. Our pipeline is further extended to a hierarchical and iterative inpainting process to continuously generate the placement of large furniture and small objects to enrich the scene. This iterative structure brings the flexibility for our method to generate or refine scenes from various starting points, such as text, floor plans, or pre-arranged environments. Experimental results demonstrate that $\textit{Architect}$ outperforms existing methods in producing realistic and complex environments, making it highly suitable for Embodied AI and robotics applications. Yian Wang 0001, Xiaowen Qiu, Jiageng Liu, Zhehuan Chen, Jiting Cai, Yufei Wang 0007, Tsun-Hsuan Wang, Zhou Xian, Chuang Gan 0001 |
NeurIPS | 8 |
| 2023 | SoftZoo: A Soft Robot Co-design Benchmark For Locomotion In Diverse Environments
Tsun-Hsuan Wang, Pingchuan Ma 0002, Andrew Spielberg, Zhou Xian, Josh Tenenbaum, Daniela Rus, Chuang Gan 0001 |
ICLR | 4 |
| 2023 | FluidLab: A Differentiable Environment for Benchmarking Complex Fluid Manipulation
Zhou Xian, Zhenjia Xu, Hsiao-Yu Fish Tung, Antonio Torralba 0001, Katerina Fragkiadaki, Chuang Gan 0001 |
ICLR | 1 |
| 2021 | HyperDynamics: Meta-Learning Object and Agent Dynamics with Hypernetworks
Zhou Xian, Shamit Lal, Hsiao-Yu Fish Tung, Emmanouil A. Platanios, Katerina Fragkiadaki |
ICLR | 1 |
| 2020 | Research on visualization planning method of distribution network based on graphical model integrationabstractHigh efficient video coding (HEVC) is a new video coding compression standard. HEVC adopts context-based adaptive binary arithmetic coding (CABAC) as the entropy coding scheme. In this paper, the overall architecture and efficiency of the main frequency are improved by the optimization of the input and output modules and the module optimization of the arithmetic coding CABAC hardware structure. In terms of input module optimization, four-level buffer input and residual coefficient transmission optimization are adopted; in terms of arithmetic coding module optimization, context model index pre-reading, pre-normalization look-up table and in-line serial stream output design are adopted so as to improve the overall efficiency of the architecture and the main frequency, reduce resource consumption, and achieve a high-frequency hardware architecture of the efficient coding pipeline. The combined results show that the pipeline can operate at 370MHz with 43.49K gates aiming at 90nm process. The processing rate and throughput can support real-time encoding of 1080P video under the general test conditions of the HEVC standard of 30 frames per second. Huang He, Zhou Xian, Guo Liang, Chang Hao, Ma Ning |
MSN | 2 |
| 2019 | Domain Randomization for Macromolecule Structure Classification and Segmentation in Electron Cyro-tomogramsabstractIt is crucial to study and understand cellular processes. In recent years, Cellular Electron CryoTomography (CECT) serves as a powerful 3D imaging tool to visualize spatial structure of macromolecules inside the cell. However, it is challenging to analyze the macromolecular structures in a systematic way due to nature of the structural complexity of subcellular components. Existing computational and deep learning based approaches suffer from limited scalability, discrimination ability and lack of accurate annotated CECT data. Training with cheap simulated data can alleviate this problem while facing new challenges of bridging the “reality gap” between synthetic training data and real testing data. In this paper, we tackle the tasks of macromolecule structure classification and segmentation in CECT images by adapting a simple but effective technique, domain randomization. We show that by combining deep neural models and domain randomization, we are able to achieve significant improvements of 35.21% and 46.34% in tasks of classification and semantic segmentation for real CECT data, comparing to the model trained only on syhthetic data that aims to faithfully reproduce real-world data distribution. Chengqian Che, Zhou Xian, Xin Gao 0001, Min Xu 0009 |
BIBM | 2 |