Minghao Gou

dblp:244/3288 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
6since 2021 · last 2023
0000-0003-1425-4231ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Robot manipulation · 49% 3D vision · 27% Deep learning architectures and training · 17%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Embedded and real-time systems · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot manipulation
grasping
1.632023
Target-referenced Reactive Grasping for Dynamic Objects · CVPR 2023
RGB Matters: Learning 7-DoF Grasp Poses on Monocular RGBD Images · ICRA 2021
GraspNet-1Billion: A Large-Scale Benchmark for General Object Grasping · CVPR 2020
Robotics › Robot manipulation › grasping › grasp detection
grasp pose estimation
0.922021
Graspness Discovery in Clutters for Fast and Accurate Grasp Detection · ICCV 2021
GraspNet-1Billion: A Large-Scale Benchmark for General Object Grasping · CVPR 2020
Machine learning › Deep learning architectures and training › data augmentation
instance-level augmentation
0.712023
InstaBoost++: Visual Coherence Principles for Unified 2D/3D Instance Level Data Augmentation · Int. J. Comput. Vis. 2023
Embedded and real-time systems › cyber-physical systems › robot systems
robotic manipulation
0.712023
AnyGrasp: Robust and Efficient Grasp Perception in Spatial and Temporal Domains · IEEE Trans. Robotics 2023
Computer vision › 3D vision
3d reconstruction
0.612022
A Real World Dataset for Multi-view 3D Reconstruction · ECCV (8) 2022
Computer vision › 3D vision › 3d reconstruction
multi-view reconstruction
0.612022
A Real World Dataset for Multi-view 3D Reconstruction · ECCV (8) 2022
Robotics › Robot manipulation › grasping
grasp detection
0.512021
Graspness Discovery in Clutters for Fast and Accurate Grasp Detection · ICCV 2021
Robotics › Robot manipulation › grasping
grasping in clutter
0.512021
Graspness Discovery in Clutters for Fast and Accurate Grasp Detection · ICCV 2021
Robotics › Robot manipulation › grasping › grasp detection
RGB-D grasp detection
0.512021
RGB Matters: Learning 7-DoF Grasp Poses on Monocular RGBD Images · ICRA 2021
Machine learning › Deep learning architectures and training › data augmentation
copy-paste augmentation
0.412019
InstaBoost: Boosting Instance Segmentation via Probability Map Guided Copy-Pasting · ICCV 2019
Machine learning › Deep learning architectures and training
data augmentation
0.412019
InstaBoost: Boosting Instance Segmentation via Probability Map Guided Copy-Pasting · ICCV 2019
Computer vision › Segmentation and scene understanding
instance segmentation
0.412019
InstaBoost: Boosting Instance Segmentation via Probability Map Guided Copy-Pasting · ICCV 2019
Robotics › Motion planning and robot control
robot learning
0.212023
Target-referenced Reactive Grasping for Dynamic Objects · CVPR 2023
Computer vision › 3D vision
point cloud processing
0.112020
GraspNet-1Billion: A Large-Scale Benchmark for General Object Grasping · CVPR 2020

Methods — techniques the papers use, named apart from their topics

visual coherence principles · 0.7grasp space tracking · 0.7dense supervision · 0.7deep learning · 0.7attentional graph neural network · 0.7neural network approximation · 0.5look-ahead searching · 0.5graspness model · 0.5encoder-decoder network · 0.5convolutional neural network · 0.5analytic search · 0.5decoupled direction and parameter learning · 0.4
YearPublicationVenuePosition
2023 Target-referenced Reactive Grasping for Dynamic Objects
abstract
Reactive grasping, which enables the robot to successfully grasp dynamic moving objects, is of great interest in robotics. Current methods mainly focus on the temporal smoothness of the predicted grasp poses but few consider their semantic consistency. Consequently, the predicted grasps are not guaranteed to fall on the same part of the same object, especially in cluttered scenes. In this paper, we propose to solve reactive grasping in a target-referenced setting by tracking through generated grasp spaces. Given a targeted grasp pose on an object and detected grasp poses in a new observation, our method is composed of two stages: 1) discovering grasp pose correspondences through an attentional graph neural network and selecting the one with the highest similarity with respect to the target pose; 2) refining the selected grasp poses based on target and historical information. We evaluate our method on a large-scale benchmark GraspNet-1Billion. We also collect 30 scenes of dynamic objects for testing. The results suggest that our method outperforms other representative methods. Furthermore, our real robot experiments achieve an average success rate of over 80 percent. Code and demos are available at: https://graspnet.net/reactive.
Jirong Liu, Ruo Zhang, Haoshu Fang, Minghao Gou, Hongjie Fang, Chenxi Wang 0003, Hengxu Yan, Cewu Lu
CVPR4
2023 InstaBoost++: Visual Coherence Principles for Unified 2D/3D Instance Level Data Augmentation
Jianhua Sun 0003, Haoshu Fang, Runzhong Wang, Minghao Gou, Cewu Lu
Int. J. Comput. Vis.5
2023 AnyGrasp: Robust and Efficient Grasp Perception in Spatial and Temporal Domains
abstract
As the basis for prehensile manipulation, it is vital to enable robots to grasp as robustly as humans. Our innate grasping system is prompt, accurate, flexible, and continuous across spatial and temporal domains. Few existing methods cover all these properties for robot grasping. In this article, we propose AnyGrasp for grasp perception to enable robots these abilities using a parallel gripper. Specifically, we develop a dense supervision strategy with real perception and analytic labels in the spatial–temporal domain. Additional awareness of objects' center-of-mass is incorporated into the learning process to help improve grasping stability. Utilization of grasp correspondence across observations enables dynamic grasp tracking. Our model can efficiently generate accurate, 7-DoF, dense, and temporally-smooth grasp poses and works robustly against large depth-sensing noise. Using AnyGrasp, we achieve a 93.3% success rate when clearing bins with over 300 unseen objects, which is on par with human subjects under controlled conditions. Over 900 mean-picks-per-hour is reported on a single-arm system. For dynamic grasping, we demonstrate catching swimming robot fish in the water.
Haoshu Fang, Chenxi Wang 0003, Hongjie Fang, Minghao Gou, Jirong Liu, Hengxu Yan, Wenhai Liu, Yichen Xie 0002, Cewu Lu
IEEE Trans. Robotics4
2022 A Real World Dataset for Multi-view 3D Reconstruction
Rakesh Shrestha, Siqi Hu, Minghao Gou
ECCV (8)3
2021 Graspness Discovery in Clutters for Fast and Accurate Grasp Detection
abstract
Efficient and robust grasp pose detection is vital for robotic manipulation. For general 6 DoF grasping, conventional methods treat all points in a scene equally and usually adopt uniform sampling to select grasp candidates. However, we discover that ignoring where to grasp greatly harms the speed and accuracy of current grasp pose detection methods. In this paper, we propose "graspness", a quality based on geometry cues that distinguishes graspable area in cluttered scenes. A look-ahead searching method is proposed for measuring the graspness and statistical results justify the rationality of our method. To quickly detect graspness in practice, we develop a neural network named graspness model to approximate the searching process. Extensive experiments verify the stability, generality and effectiveness of our graspness model, allowing it to be used as a plug-and-play module for different methods. A large improvement in accuracy is witnessed for various previous methods after equipping our graspness model. Moreover, we develop GSNet, an end-to-end network that incorporate our graspness model for early filtering of low quality predictions. Experiments on a large scale benchmark, GraspNet-1Billion, show that our method outperforms previous arts by a large margin (30 + AP) and achieves a high inference speed.
Chenxi Wang 0003, Haoshu Fang, Minghao Gou, Hongjie Fang, Cewu Lu
ICCV3
2021 RGB Matters: Learning 7-DoF Grasp Poses on Monocular RGBD Images
abstract
General object grasping is an important yet unsolved problem in the field of robotics. Most of the current methods either generate grasp poses with few DoF that fail to cover most of the success grasps, or only take the unstable depth image or point cloud as input which may lead to poor results in some cases. In this paper, we propose RGBD-Grasp, a pipeline that solves this problem by decoupling 7-DoF grasp detection into two sub-tasks where RGB and depth information are processed separately. In the first stage, an encoder-decoder like convolutional neural network Angle-View Net(AVN) is proposed to predict the SO(3) orientation of the gripper at every location of the image. Consequently, a Fast Analytic Searching(FAS) module calculates the opening width and the distance of the gripper to the grasp point. By decoupling the grasp detection problem and introducing the stable RGB modality, our pipeline alleviates the requirement for the high-quality depth image and is robust to depth sensor noise. We achieve state-of-the-art results on GraspNet-1Billion dataset compared with several baselines. Real robot experiments on a UR5 robot with an Intel Realsense camera and a Robotiq two-finger gripper show high success rates for both single object scenes and cluttered scenes. Our code and trained model are available at graspnet.net.
Minghao Gou, Haoshu Fang, Zhanda Zhu, Chenxi Wang 0003, Cewu Lu
ICRA1
2020 GraspNet-1Billion: A Large-Scale Benchmark for General Object Grasping
abstract
Object grasping is critical for many applications, which is also a challenging computer vision problem. However, for cluttered scene, current researches suffer from the problems of insufficient training data and the lacking of evaluation benchmarks. In this work, we contribute a large-scale grasp pose detection dataset with a unified evaluation system. Our dataset contains 97,280 RGB-D image with over one billion grasp poses. Meanwhile, our evaluation system directly reports whether a grasping is successful by analytic computation, which is able to evaluate any kind of grasp poses without exhaustively labeling ground-truth. In addition, we propose an end-to-end grasp pose prediction network given point cloud inputs, where we learn approaching direction and operation parameters in a decoupled manner. A novel grasp affinity field is also designed to improve the grasping robustness. We conduct extensive experiments to show that our dataset and evaluation system can align well with real-world experiments and our proposed network achieves the state-of-the-art performance. Our dataset, source code and models are publicly available at www.graspnet.net.
Haoshu Fang, Chenxi Wang 0003, Minghao Gou, Cewu Lu
CVPR3
2019 InstaBoost: Boosting Instance Segmentation via Probability Map Guided Copy-Pasting
abstract
Instance segmentation requires a large number of training samples to achieve satisfactory performance and benefits from proper data augmentation. To enlarge the training set and increase the diversity, previous methods have investigated using data annotation from other domain (e.g. bbox, point) in a weakly supervised mechanism. In this paper, we present a simple, efficient and effective method to augment the training set using the existing instance mask annotations. Exploiting the pixel redundancy of the background, we are able to improve the performance of Mask R-CNN for 1.7 mAP on COCO dataset and 3.3 mAP on Pascal VOC dataset by simply introducing random jittering to objects. Furthermore, we propose a location probability map based approach to explore the feasible locations that objects can be placed based on local appearance similarity. With the guidance of such map, we boost the performance of R101-Mask R-CNN on instance segmentation from 35.7 mAP to 37.9 mAP without modifying the backbone or network structure. Our method is simple to implement and does not increase the computational complexity. It can be integrated into the training pipeline of any instance segmentation model without affecting the training and inference efficiency. Our code and models have been released at https://github.com/GothicAi/InstaBoost.
Haoshu Fang, Jianhua Sun 0003, Runzhong Wang, Minghao Gou, Yong-Lu Li 0001, Cewu Lu
ICCV4