Yunxuan Mao

dblp:257/4807 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-9243-7706ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Efficient Alignment of Unconditioned Action Prior for Language-Conditioned Pick and Place in Clutter
abstract
We study the task of language-conditioned pick and place in clutter, where a robot should grasp a target object in open clutter and move it to a specified place. Some approaches learn end-to-end policies with features from vision foundation models, requiring large datasets. Others combine foundation models in a zero-shot setting, suffering from cascading errors. In addition, they primarily leverage vision and language foundation models, focusing less on action priors. In this paper, we aim to develop an effective policy by integrating foundation priors from vision, language, and action. We propose A2, an action prior alignment method that aligns unconditioned action priors with 3D vision-language priors by learning one attention layer. The alignment formulation enables our policy to train with less data and preserve zero-shot generalization capabilities. We show that a shared policy for both pick and place actions enhances the performance for each task, and introduce a policy adaptation scheme to accommodate the multi-modal nature of actions. Extensive experiments in simulation and the real-world show that our policy achieves higher task success rates with fewer steps for both pick and place tasks in clutter, effectively generalizing to unseen objects and language instructions. Videos and codes are available at https://xukechun.github.io/papers/A2.
Kechun Xu, Xunlong Xia, Yunxuan Mao, Bing Deng, Jieping Ye, Rong Xiong, Yue Wang 0020
IEEE Trans Autom. Sci. Eng.5
2024 NGEL-SLAM: Neural Implicit Representation-based Global Consistent Low-Latency SLAM System
abstract
Neural implicit representations have emerged as a promising solution for providing dense geometry in Simultaneous Localization and Mapping (SLAM). However, existing methods in this direction fall short in terms of global consistency and low latency. This paper presents NGEL-SLAM to tackle the above challenges. To ensure global consistency, our system leverages a traditional feature-based tracking module that incorporates loop closure. Additionally, we maintain a global consistent map by representing the scene using multiple neural implicit fields, enabling quick adjustment to the loop closure. Moreover, our system allows for fast convergence through the use of octree-based implicit representations. The combination of rapid response to loop closure and fast convergence makes our system a truly low-latency system that achieves global consistency. Our system enables rendering high-fidelity RGB-D images, along with extracting dense and complete surfaces. Experiments on both synthetic and real-world datasets suggest that our system achieves state-of-the-art tracking and mapping accuracy while maintaining low latency.
Yunxuan Mao, Zhuqing Zhang, Yue Wang 0020, Rong Xiong, Yiyi Liao
ICRA1
2024 ν-DBA: Neural Implicit Dense Bundle Adjustment Enables Image-Only Driving Scene Reconstruction
abstract
The joint optimization of the sensor trajectory and 3D map is a crucial characteristic of bundle adjustment (BA), essential for autonomous driving. This paper presents ν-DBA, a novel framework implementing geometric dense bundle adjustment (DBA) using 3D neural implicit surfaces for map parametrization, which optimizes both the map surface and trajectory poses using geometric error guided by dense optical flow prediction. Additionally, we fine-tune the optical flow model with per-scene self-supervision to further improve the quality of the dense mapping. Our experimental results on multiple driving scene datasets demonstrate that our method achieves superior trajectory optimization and dense reconstruction accuracy. We also investigate the influences of photometric error and different neural geometric priors on the performance of surface reconstruction and novel view synthesis. Our method stands as a significant step towards leveraging neural implicit representations in dense bundle adjustment for more accurate trajectories and detailed environmental mapping.
Yunxuan Mao, Bingqi Shen, Rong Xiong, Yiyi Liao, Yue Wang 0020
IROS1
2021 Modeling and Control of an Untethered Magnetic Gripper
abstract
Small-scale robots have great potential in minimally invasive surgery (MIS). In this paper, we propose an untethered magnetic gripper with small scale and build a double-magnet model for it. The gripper is 4.3mm long and its maximum width is 4mm. It contains a spindle and two magnets, which can achieve precise control of orientation, position and open angle with external magnetic driven field. As a result, it can perform operations such as transporting medicines in confined and constrained environments. Modeling and analysis of the magnetic gripper have been carried out. Relationship between the open angle and external magnetic field has been established. Kinematics model of the gripper has been built. A 3-axis Helmholtz-Maxwell coil system has been established to generate the magnetic field, in which orientation and open angle can be controlled with uniform magnetic field while position can be controlled with gradient field. The proposed gripper have been validated with phantom experiments. An opened angle control error of 0.63° and direction control error of 1.1° have been obtained.
Yunxuan Mao, Sishen Yuan, Jiaole Wang, Jinmin Zhang, Shuang Song 0002
ICRA1