Boshi An

dblp:330/2184 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
3D vision · 41% Robot manipulation · 15% Multi-agent systems · 13%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › depth estimation
depth completion
0.812024
Bilateral Propagation Network for Depth Completion · CVPR 2024
Computer vision › 3D vision
depth estimation
0.812024
Bilateral Propagation Network for Depth Completion · CVPR 2024
Knowledge, reasoning and agents › Multi-agent systems
emergent communication
0.812024
Learning Multi-Object Positional Relationships via Emergent Communication · AAAI 2024
Machine learning › Generative modeling
image manipulation
0.812024
RGBManip: Monocular Image-based Robotic Manipulation through Active Object Pose Estimation · ICRA 2024
Robotics › Robot manipulation
affordance learning
0.712023
RLAfford: End-to-End Affordance Learning for Robotic Manipulation · ICRA 2023
Computer vision › 3D vision › human body modeling
contact map prediction
0.712023
RLAfford: End-to-End Affordance Learning for Robotic Manipulation · ICRA 2023
Robotics › Motion planning and robot control › robot learning › robotic reinforcement learning
reinforcement learning for manipulation
0.712023
RLAfford: End-to-End Affordance Learning for Robotic Manipulation · ICRA 2023
Computer vision › 3D vision › object pose estimation
6d object pose estimation
0.212024
RGBManip: Monocular Image-based Robotic Manipulation through Active Object Pose Estimation · ICRA 2024
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.212024
Learning Multi-Object Positional Relationships via Emergent Communication · AAAI 2024
Computer vision › Vision and language
multimodal fusion
0.212024
Bilateral Propagation Network for Depth Completion · CVPR 2024
Robotics › Robot manipulation
grasping
0.212023
RLAfford: End-to-End Affordance Learning for Robotic Manipulation · ICRA 2023

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.4referential game · 0.8multi-scale refinement · 0.8multi-layer perceptron · 0.8monocular camera · 0.8language transfer · 0.8image generator · 0.8bilateral filtering · 0.8visual affordance · 0.7end-to-end learning · 0.7
YearPublicationVenuePosition
2024 Learning Multi-Object Positional Relationships via Emergent Communication
abstract
The study of emergent communication has been dedicated to interactive artificial intelligence. While existing work focuses on communication about single objects or complex image scenes, we argue that communicating relationships between multiple objects is important in more realistic tasks, but understudied. In this paper, we try to fill this gap and focus on emergent communication about positional relationships between two objects. We train agents in the referential game where observations contain two objects, and find that generalization is the major problem when the positional relationship is involved. The key factor affecting the generalization ability of the emergent language is the input variation between Speaker and Listener, which is realized by a random image generator in our work. Further, we find that the learned language can generalize well in a new multi-step MDP task where the positional relationship describes the goal, and performs better than raw-pixel images as well as pre-trained image features, verifying the strong generalization ability of discrete sequences. We also show that language transfer from the referential game performs better in the new task than learning language directly in this task, implying the potential benefits of pre-training in referential games. All in all, our experiments demonstrate the viability and merit of having agents learn to communicate positional relationships between multiple objects through emergent communication.
Yicheng Feng, Boshi An, Zongqing Lu 0002
AAAI2
2024 Bilateral Propagation Network for Depth Completion
abstract
Depth completion aims to derive a dense depth map from sparse depth measurements with a synchronized color image. Current state-of-the-art (SOTA) methods are predominantly propagation-based, which work as an iterative refinement on the initial estimated dense depth. However, the initial depth estimations mostly result from direct applications of convolutional layers on the sparse depth map. In this paper, we present a Bilateral Propagation Network (BP-Net), that propagates depth at the earliest stage to avoid directly convolving on sparse data. Specifically, our approach propagates the target depth from nearby depth measurements via a non-linear model, whose coefficients are generated through a multi-layer perceptron conditioned on both radiometric difference and spatial distance. By integrating bilateral propagation with multi-modal fusion and depth refinement in a multi-scale framework, our BP-Net demonstrates outstanding performance on both indoor and outdoor scenes. It achieves SOTA on the NYUv2 dataset and ranks 1st on the KITTI depth completion benchmark at the time of submission. Experimental results not only show the effectiveness of bilateral propagation but also emphasize the significance of early-stage propagation in contrast to the refinement stage. Our code and trained models will be available on the project page.
Jie Tang 0015, Fei-Peng Tian, Boshi An, Jian Li 0003, Ping Tan 0002
CVPR3
2024 RGBManip: Monocular Image-based Robotic Manipulation through Active Object Pose Estimation
abstract
Robotic manipulation requires accurate perception of the environment, which poses a significant challenge due to its inherent complexity and constantly changing nature. In this context, RGB image and point-cloud observations are two commonly used modalities in visual-based robotic manipulation, but each of these modalities have their own limitations. Commercial point-cloud observations often suffer from issues like sparse sampling and noisy output due to the limits of the emission-reception imaging principle. On the other hand, RGB images, while rich in texture information, lack essential depth and 3D information crucial for robotic manipulation. To mitigate these challenges, we propose an image-only robotic manipulation framework that leverages an eye-on-hand monocular camera installed on the robot’s parallel gripper. By moving with the robot gripper, this camera gains the ability to actively perceive the object from multiple perspectives during the manipulation process. This enables the estimation of 6D object poses, which can be utilized for manipulation. While, obtaining images from more and diverse viewpoints typically improves pose estimation, it also increases the manipulation time. To address this trade-off, we employ a reinforcement learning policy to synchronize the manipulation strategy with active perception, achieving a balance between 6D pose accuracy and manipulation efficiency. Our experimental results in both simulated and real-world environments showcase the state-of-the-art effectiveness of our approach. We believe that our method will inspire further research on real-world-oriented robotic manipulation. See https://rgbmanip.github.io/ for more details.
Boshi An, Yiran Geng, Kai Chen 0028, Xiaoqi Li 0020, Qi Dou 0001, Hao Dong 0003
ICRA1
2023 RLAfford: End-to-End Affordance Learning for Robotic Manipulation
abstract
Learning to manipulate 3D objects in an interactive environment has been a challenging problem in Reinforcement Learning (RL). In particular, it is hard to train a policy that can generalize over objects with different semantic categories, diverse shape geometry and versatile functionality. In this study, we focused on the contact information in manipulation processes, and proposed a unified representation for critical interactions to describe different kinds of manipulation tasks. Specifically, we take advantage of the contact information generated during the RL training process and employ it as unified visual representation to predict contact map of interest. Such representation leads to an end-to-end learning framework that combined affordance based and RL based methods for the first time. Our unified framework can generalize over different types of manipulation tasks. Surprisingly, the effectiveness of such framework holds even under the multi-stage and multi-agent scenarios. We tested our method on eight types of manipulation tasks. Results showed that our methods outperform baseline algorithms, including visual affordance methods and RL methods, by a large margin on the success rate. The demonstration can be found at https://sites.google.com/view/rlafford/.
Yiran Geng, Boshi An, Yuanpei Chen, Yaodong Yang 0001, Hao Dong 0003
ICRA2