Xianyi Cheng

dblp:00/7782 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0001-8342-9459ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 6 since 2021Systems, architecture and hardware · 5 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CageCoOpt: Enhancing Manipulation Robustness through Caging-Guided Morphology and Policy Co-Optimization
abstract
Uncertainties in contact dynamics and object geometry remain significant barriers to robust robotic manipulation. Caging helps mitigate these uncertainties by constraining an object’s mobility without requiring precise contact modeling. Existing caging research often treats morphology and policy optimization as separate problems, overlooking their synergy. In this paper, we introduce CageCoOpt, a hierarchical framework that jointly optimizes manipulator morphology and control policy for robust caging-based manipulation. The framework employs reinforcement learning for policy optimization at the lower level and multitask Bayesian optimization for morphology optimization at the upper level. We incorporate a caging metric into both optimization levels to encourage caging configurations and thereby improve manipulation robustness. The evaluation consists of four manipulation tasks and demonstrates that co-optimizing morphology and policy improves task performance under uncertainties, establishing caging-guided co-optimization as a viable approach for robust manipulation.
Yifei Dong 0007, Shaohang Han, Xianyi Cheng, Werner Friedl, Rafael I. Cabral Muchacho, Máximo A. Roa, Jana Tumova, Florian T. Pokorny
IROS3
2024 WebArena: A Realistic Web Environment for Building Autonomous Agents
abstract
With advances in generative AI, there is now potential for autonomous agents to manage daily tasks via natural language commands. However, current agents are primarily created and tested in simplified synthetic environments, leading to a disconnect with real-world scenarios. In this paper, we build an environment for language-guided agents that is highly realistic and reproducible. Specifically, we focus on agents that perform tasks on the web, and create an environment with fully functional websites from four common domains: e-commerce, social forum discussions, collaborative software development, and content management. Our environment is enriched with tools (e.g., a map) and external knowledge bases (e.g., user manuals) to encourage human-like task-solving. Building upon our environment, we release a set of benchmark tasks focusing on evaluating the functional correctness of task completions. The tasks in our benchmark are diverse, long-horizon, and designed to emulate tasks that humans routinely perform on the internet. We experiment with several baseline agents, integrating recent techniques such as reasoning before acting. The results demonstrate that solving complex tasks is challenging: our best GPT-4-based agent only achieves an end-to-end task success rate of 14.41%, significantly lower than the human performance of 78.24%. These results highlight the need for further development of robust agents, that current state-of-the-art large language models are far from perfect performance in these real-life tasks, and that \ours can be used to measure such progress.\footnote{Code, data, environment reproduction instructions, video demonstrations are available in the supplementary.}
Shuyan Zhou, Frank F. Xu, Hao Zhu 0011, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon 0002, Graham Neubig
ICLR7
2023 OPECE: Optimal Placement of Edge Servers in Cloud Environment
Fengmei Chen, Shengjun Xue, Zheng Li 0026, Yachong Tian, Xianyi Cheng
GPC (2)6
2022 Contact Mode Guided Motion Planning for Quasidynamic Dexterous Manipulation in 3D
abstract
This paper presents Contact Mode Guided Manipulation Planning (CMGMP) for 3D quasistatic and quasi-dynamic rigid body motion planning in dexterous manipulation. The CMGMP algorithm generates hybrid motion plans including both continuous state transitions and discrete contact mode switches, without the need for pre-specified contact sequences or pre-designed motion primitives. The key idea is to use automatically enumerated contact modes of environment-object contacts to guide the tree expansions during the search. Contact modes automatically synthesize manipulation primitives, while the sampling-based planning framework sequences those primitives into a coherent plan. We test our algorithm on fourteen 3D manipulation tasks, and validate our models by executing some plans open-loop on a real robot-manipulator system11The video is available at https://youtu.be/JuLlliG3vGc.
Xianyi Cheng, Matthew T. Mason
ICRA1
2022 Extrinsic Dexterous Manipulation with a Direct-drive Hand: A Case Study
abstract
This paper explores a novel approach to dexterous manipulation, aimed at levels of speed, precision, robustness, and simplicity suitable for practical deployment. The enabling technology is a Direct-drive Hand (DDHand) comprising two fingers, two DOFs each, that exhibit high speed and a light touch. The test application is the dexterous manipulation of three small and irregular parts, moving them to a grasp suitable for a subsequent assembly operation, regardless of initial presentation. We employed four primitive behaviors that use ground contact as a “third finger”, prior to or during the grasp process: pushing, pivoting, toppling, and squeeze-grasping. In our experiments, each part was presented from 30 to 90 times randomly positioned in each stable pose. Success rates varied from 83% to 100%. The time to manipulate and grasp was 6.32 seconds on average, varying from 2.07 to 16 seconds. In some cases, performance was robust, precise, and fast enough for practical applications, but in other cases, pose uncertainty required time-consuming vision and arm motions. The paper concludes with a discussion of further improvements required to make the primitives robust, eliminate uncertainty, and reduce this dependence on vision and arm motion.
Arnav Gupta, Yuemin Mao, Ankit Bhatia, Xianyi Cheng, Jonathan King, Matthew T. Mason
IROS4
2021 Contact Mode Guided Sampling-Based Planning for Quasistatic Dexterous Manipulation in 2D
abstract
The discontinuities and multi-modality introduced by contacts make manipulation planning challenging. Many previous works avoid this problem by pre-designing a set of high-level motion primitives like grasping and pushing. However, such motion primitives are often not adequate to describe dexterous manipulation motions. In this work, we propose a method for dexterous manipulation planning at a more primitive level. The key idea is to use contact modes to guide the search in a sampling-based planning framework. Our method can automatically generate contact transitions and motion trajectories under the quasistatic assumption. In the experiments, this method sometimes generates motions that are often pre-designed as motion primitives, as well as dexterous motions that are more task-specific1.
Xianyi Cheng, Matthew T. Mason
ICRA1
2021 Efficient Contact Mode Enumeration in 3D
Xianyi Cheng, Matthew T. Mason
WAFR2
2019 Manipulation with Suction Cups Using External Contacts
Xianyi Cheng, Matthew T. Mason
ISRR1
2018 Sensor Selection and Stage & Result Classifications for Automated Miniature Screwdriving
abstract
Hundreds of billions of small screws are assembled in consumer electronics industry every year, yet reliably automating the screwdriving process remains one of the most challenging tasks. Two barriers to further adoption of robotic threaded fastening systems are system cost and technical challenges, especially for small screws. An affordable intelligent screwdriving system that can support online stage and result classification is the first step to bridge the gap. To this end, starting from a state transition graph of screwdriving processes and a labeled screwdriving dataset (1862 runs of M1.4 screws) on multiple sensor signals, we develop classification algorithms and perform sensor reduction. Fast and accurate result classifiers are developed using linear discriminant analysis, while a wrapper method for feature subset selection is used to identify the optimal feature subset and corresponding sensor signals to reduce cost. A stage classifier based on decision tree is developed using the optimal sensor subset. The stage classifier achieves high accuracy in realtime prediction of various stages when augmented with the state transition graph.
Xianyi Cheng, Zhenzhong Jia, Ankit Bhatia, Reuben M. Aronson, Matthew T. Mason
IROS1