Long Zeng 0001

dblp:08/10533 · DBLP profile ↗
← Back
32ranked-venue papers
7as first author
24since 2021 · last 2025
0000-0002-3090-6319ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 16 · 1 first-author · 14 since 2021Systems, architecture and hardware · 12 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder
abstract
Diffusion models for garment-centric human generation from text or image prompts have garnered emerging attention for their great application potential. However, existing methods often face a dilemma: lightweight approaches, such as adapters, are prone to generate inconsistent textures; while finetune-based methods involve high training costs and struggle to maintain the generalization capabilities of pretrained diffusion models, limiting their performance across diverse scenarios. To address these challenges, we propose DreamFit, which incorporates a lightweight Anything-Dressing Encoder specifically tailored for the garment-centric human generation. DreamFit has three key advantages: (1) Lightweight training: with the proposed adaptive attention and LoRA modules, DreamFit significantly minimizes the model complexity to 83.4M trainable parameters. (2) Anything-Dressing: Our model generalizes surprisingly well to a wide range of (non-)garments, creative styles, and prompt instructions, consistently delivering high-quality results across diverse scenarios. (3) Plug-and-play: DreamFit is engineered for smooth integration with any community control plugins for diffusion models, ensuring easy compatibility and minimizing adoption barriers. To further enhance generation quality, DreamFit leverages pretrained large multi-modal models (LMMs) to enrich the prompt with fine-grained garment descriptions, thereby reducing the prompt gap between training and inference. We conduct comprehensive experiments on both 768 x 512 high-resolution benchmarks and in-the-wild images. DreamFit surpasses all existing methods, highlighting its state-of-the-art capabilities of garment-centric human generation.
Ente Lin, Xujie Zhang, Fuwei Zhao, Yuxuan Luo 0002, Long Zeng 0001, Xiaodan Liang
AAAI6
2025 Training-Free Point Cloud Recognition Based on Geometric and Semantic Information Fusion
abstract
The trend of employing training-free methods for point cloud recognition is becoming increasingly popular due to its significant reduction in computational resources and time costs. However, existing approaches are limited as they typically extract either geometric or semantic features. To address this limitation, we are the first to propose a novel training-free method that integrates both geometric and semantic features. For the geometric branch, we adopt a non-parametric strategy to extract geometric features. In the semantic branch, we leverage a model aligned with text features to obtain semantic features. Additionally, we introduce the GFE module to complement the geometric information of point clouds and the MFF module to improve performance in few-shot settings. Experimental results demonstrate that our method outperforms existing state-of-the-art training-free approaches on mainstream benchmark datasets, including ModelNet and ScanObiectNN.
Zhichao Liao, Xinghui Li, Long Zeng 0001
ICASSP6
2025 Constraint-Aware Feature Learning for Parametric Point Cloud
Ruiqi Lei, Zhichao Liao, Fengyuan Piao, Pingfa Feng, Long Zeng 0001
ICCV8
2025 GaussianRoom: Improving 3D Gaussian Splatting with SDF Guidance and Monocular Cues for Indoor Scene Reconstruction
abstract
Embodied intelligence requires precise reconstruction and rendering to simulate large-scale real-world data. Although 3D Gaussian Splatting (3DGS) has recently demonstrated high-quality results with real-time performance, it still faces challenges in indoor scenes with large, textureless regions, resulting in incomplete and noisy reconstructions due to poor point cloud initialization and underconstrained optimization. Inspired by the continuity of signed distance field (SDF), which naturally has advantages in modeling surfaces, we propose a unified optimization framework that integrates neural signed distance fields (SDFs) with 3DGS for accurate geometry reconstruction and real-time rendering. This framework incorporates a neural SDF field to guide the densification and pruning of Gaussians, enabling Gaussians to model scenes accurately even with poor initialized point clouds. Simultaneously, the geometry represented by Gaussians improves the efficiency of the SDF field by piloting its point sampling. Additionally, we introduce two regularization terms based on normal and edge priors to resolve geometric ambiguities in textureless areas and enhance detail accuracy. Extensive experiments in ScanNet and ScanNet++ show that our method achieves state-of-the-art performance in both surface reconstruction and novel view synthesis. Project page: https://xhd0612.github.io/GaussianRoom.github.io/
Haodong Xiang, Xinghui Li, Xiansong Lai, Wanting Zhang, Zhichao Liao, Long Zeng 0001, Xueping Liu 0003
ICRA7
2025 Diffusion Suction Grasping with Large-Scale Parcel Dataset
abstract
While recent advances in suction grasping have shown remarkable progress, significant challenges persist particularly in cluttered and complex parcel handling scenarios. Current approaches are limited by (1) the lack of comprehensive parcel-specific suction grasp datasets and (2) poor adaptability to diverse object properties, including size, geometry, and texture. We address these challenges through two main contributions. Firstly, we introduce the Parcel-Suction-Dataset, a large-scale synthetic dataset containing 25 thousand cluttered scenes with 410 million precision-annotated suction grasp poses, generated via our novel geometric sampling algorithm. Secondly, we propose Diffusion-Suction, a framework that innovatively reformulates suction grasp prediction as a conditional generation task using denoising diffusion probabilistic models. Our method iteratively refines random noise into suction grasping score through visual-conditioned guidance from point cloud observations, effectively learning spatial point-wise affordances from our synthetic dataset. Extensive experiments demonstrate that the simple yet efficient Diffusion-Suction achieves new state-of-the-art performance compared to previous models on both Parcel-Suction-Dataset and the public SuctionNet-1Billion benchmark. This work provides a robust foundation for advancing automated parcel handling systems in real-world applications.
Ding-Tao Huang, Debei Hua, Dongfang Yu, En-Te Lin, Liang-Hong Wang, Jin-Liang Hou, Long Zeng 0001
IROS8
2025 Data-Driven MPC with Data Selection for Flexible Cable-Driven Robotic Arms
abstract
Flexible cable-driven robotic arms (FCRAs) offer dexterous and compliant motion. Still, the inherent properties of cables, such as resilience, hysteresis, and friction, often lead to particular difficulties in modeling and control. This paper proposes a model predictive control (MPC) method that relies exclusively on input-output data, without a physical model, to improve the control accuracy of FCRAs. First, we develop an implicit model based on input-output data and integrate it into an MPC optimization framework. Second, a data selection algorithm (DSA) is introduced to filter the data that best characterize the system, thereby reducing the solution time per step to approximately 4 ms, which is an improvement of nearly 80%. Lastly, the influence of hyperparameters on tracking error is investigated through simulation. The proposed method has been validated on a real FCRA platform, including five-point positioning accuracy tests, a five-point response tracking test, and trajectory tracking for letter drawing. The results demonstrate that the average positioning accuracy is approximately 2.070 mm. Moreover, compared to the PID method with an average tracking error of 1.418°, the proposed method achieves an average tracking error of 0.541°.
Huayue Liang, Yanbo Chen 0001, Hongyang Cheng, Yanzhao Yu, Shoujie Li, Junbo Tan, Xueqian Wang 0001, Long Zeng 0001
SMC8
2025 Sketch123: Multi-spectral channel cross attention for sketch-based 3D generation via diffusion models
Zhentong Xu, Long Zeng 0001, Junli Zhao, Baodong Wang, Zhenkuan Pan 0001, Yong-Jin Liu 0001
Comput. Aided Des.2
2025 Imaginarium: Vision-guided High-Quality 3D Scene Layout Generation
abstract
Generating artistic and coherent 3D scene layouts is crucial in digital content creation. Traditional optimization-based methods are often constrained by cumbersome manual rules, while deep generative models face challenges in producing content with richness and diversity. Furthermore, approaches that utilize large language models frequently lack robustness and fail to accurately capture complex spatial relationships. To address these challenges, this paper presents a novel vision-guided 3D layout generation system. We first construct a high-quality asset library containing 2,037 scene assets and 147 3D scene layouts. Subsequently, we employ an image generation model to expand prompt representations into images, fine-tuning it to align with our asset library. We then develop a robust image parsing module to recover the 3D layout of scenes based on visual semantics and geometric information. Finally, we optimize the scene layout using scene graphs and overall visual semantics to ensure logical coherence and alignment with the images. Extensive user testing demonstrates that our algorithm significantly outperforms existing methods in terms of layout richness and quality. The code and dataset will be available at https://github.com/HiHiAllen/Imaginarium.
Qinghongbing Xie, Junsheng Yu, Yirui Guan, Zhongyuan Liu, Qijun Zhao, Ligang Liu 0001, Long Zeng 0001
ACM Trans. Graph.11
2025 PCKRF: Point Cloud Completion and Keypoint Refinement With Fusion Data for 6D Pose Estimation
abstract
Some robust point cloud registration approaches with controllable pose refinement magnitude, such as ICP and its variants, are commonly used to improve 6D pose estimation accuracy. However, the effectiveness of these methods gradually diminishes with the advancement of deep learning techniques and the enhancement of initial pose accuracy, primarily due to their lack of specific design for pose refinement. In this paper, we propose Point Cloud Completion and Keypoint Refinement with Fusion Data (PCKRF), a new pose refinement pipeline for 6D pose estimation. The pipeline consists of two steps. First, it completes the input point clouds via a novel pose-sensitive point completion network. The network uses both local and global features with pose information during point completion. Then, it registers the completed object point cloud with the corresponding target point cloud by our proposed Color supported Iterative KeyPoint (CIKP) method. The CIKP method introduces color information into registration and registers a point cloud around each keypoint to increase stability. The PCKRF pipeline can be integrated with existing popular 6D pose estimation methods, such as the full flow bidirectional fusion network, to further improve their pose estimation accuracy. Experiments demonstrate that our method exhibits superior stability compared to existing approaches when optimizing initial poses with relatively high precision. Notably, the results indicate that our method effectively complements most existing pose estimation techniques, leading to improved performance in most cases. Furthermore, our method achieves promising results even in challenging scenarios involving textureless and symmetrical objects.
Yiheng Han, Irvin Haozhe Zhan, Long Zeng 0001, Yu-Ping Wang 0001, Ran Yi 0002, Minjing Yu, Matthieu Lin, Jenny Sheng, Yong-Jin Liu 0001
IEEE Trans. Vis. Comput. Graph.3
2024 Fine-Detailed Neural Indoor Scene Reconstruction Using Multi-Level Importance Sampling And Multi-View Consistency
abstract
Recently, neural implicit 3D reconstruction in indoor scenarios has become popular due to its simplicity and impressive performance. Previous works could produce complete results leveraging monocular priors of normal or depth. However, they may suffer from over-smoothed reconstructions and long-time optimization due to unbiased sampling and inaccurate monocular priors. In this paper, we propose a novel neural implicit surface reconstruction method, named FD-NeuS, to learn fine-detailed 3D models using multi-level importance sampling strategy and multi-view consistency methodology. Specifically, we leverage segmentation priors to guide region-based ray sampling, and use piecewise exponential functions as weights to pilot 3 D points sampling along the rays, ensuring more attention on important regions. In addition, we introduce multi-view feature consistency and multi-view normal consistency as supervision and uncertainty respectively, which further improve the reconstruction of details. Extensive quantitative and qualitative results show that FD-NeuS outperforms existing methods in various scenes.
Xinghui Li, Yuchen Ji, Xiansong Lai, Wanting Zhang, Long Zeng 0001
ICIP5
2024 Mobile Robot Oriented Large-Scale Indoor Dataset for Dynamic Scene Understanding
abstract
Most existing robotic datasets capture static scene data and thus are limited in evaluating robots’ dynamic performance. To address this, we present a mobile robot oriented large-scale indoor dataset, denoted as THUD (Tsinghua University Dynamic) robotic dataset, for training and evaluating their dynamic scene understanding algorithms. Specifically, the THUD dataset construction is first detailed, including organization, acquisition, and annotation methods. It comprises both real-world and synthetic data, collected with a real robot platform and a physical simulation platform, respectively. Our current dataset includes 13 larges-scale dynamic scenarios, 90K image frames, 20M 2D/3D bounding boxes of static and dynamic objects, camera poses, and IMU. The dataset is still continuously expanding. Then, the performance of mainstream indoor scene understanding tasks, e.g. 3D object detection, semantic segmentation, and robot relocalization, is evaluated on our THUD dataset. These experiments reveal serious challenges for some robot scene understanding tasks in dynamic scenes. By sharing this dataset, we aim to foster and iterate new mobile robot algorithms quickly for robot actual working dynamic environment, i.e. complex crowded dynamic scenes.
Cong Tai, Fang-xing Chen, Wanting Zhang, Tao Zhang 0130, Xueping Liu 0003, Yong-Jin Liu 0001, Long Zeng 0001
ICRA8
2024 SD-Net: Symmetric-Aware Keypoint Prediction and Domain Adaptation for 6D Pose Estimation In Bin-picking Scenarios
abstract
Despite the success of 6D pose estimation in bin-picking scenarios, existing methods still struggle to produce accurate prediction results for symmetry objects in real-world scenarios. The primary bottlenecks include 1) the ambiguity in keypoints caused by object symmetries; and 2) the domain gap between real and synthetic data. To circumvent these problems, we propose a novel 6D pose estimation network with symmetric-aware keypoint prediction and self-training domain adaptation (SD-Net). SD-Net builds on point-wise keypoint regression and deep hough voting to perform reliable keypoint detection under clutter and occlusion. Specifically, at the keypoint prediction stage, we propose a robust 3D keypoint selection strategy considering the symmetry class of objects and equivalent keypoints, which facilitate locating 3D keypoints even in highly occluded scenes. Additionally, we build an effective filtering algorithm on predicted keypoints to dynamically eliminate multiple ambiguity and outlier key-point candidates. At the domain adaptation stage, we propose the self-training framework using a student-teacher training scheme. To carefully distinguish reliable predictions, we harness tailored heuristics for 3D geometry pseudo labelling based on semi-chamfer distance. On the public Siléane dataset, SD-Net achieves state-of-the-art results, obtaining an average precision of 96%. Testing learning and generalization abilities on public Parametric datasets, SD-Net is 8% higher than the state-of-the-art method.
Ding-Tao Huang, En-Te Lin, Lipeng Chen, Li-Fu Liu, Long Zeng 0001
IROS5
2024 A Two-Stage Reinforcement Learning Approach for Robot Navigation in Long-range Indoor Dense Crowd Environments
abstract
Safe and efficient mobility is vital for mobile robots navigating long-range indoor crowd environments, such as supermarkets, restaurants, and railway stations. Traditional path planning methods are challenged because of the high dynamics of pedestrians and constrained feasible regions. Existing long-range deep reinforcement learning (DRL) path planning methods often exhibit low success rates and driving speeds in long-range navigation tasks under crowded conditions. To overcome these issues, we propose a new two-stage DRL method, known as TSDRL, where the long-range navigation task is divided into subgoal generation (SG) and planning refinement (PR) stages. In the SG stage, the agent is trained to learn a decision-making policy to generate subgoals at each decision time to avoid dense crowds. In the PR stage, the agent learns a safer and more efficient planning policy based on each subgoal generated in the SG stage to improve the robot’s movement safety and speed. Simulated experiments show that our method outperforms traditional and long-range DRL path planning methods in terms of safety, efficiency, generalization, and robustness. Furthermore, we evaluate our approach using the Turtlebot2 platform in a real-world setting, demonstrating that the robot can navigate safely and efficiently while avoiding dense crowds.
Xinghui Jing, Fuhao Li, Tao Zhang 0130, Long Zeng 0001
IROS5
2024 GRID: Scene-Graph-based Instruction-driven Robotic Task Planning
abstract
Recent works have shown that Large Language Models (LLMs) can facilitate the grounding of instructions for robotic task planning. Despite this progress, most existing works have primarily focused on utilizing raw images to aid LLMs in understanding environmental information. However, this approach not only limits the scope of observation but also typically necessitates extensive multimodal data collection and large-scale models. In this paper, we propose a novel approach called Graph-based Robotic Instruction Decomposer (GRID), which leverages scene graphs instead of images to perceive global scene information and iteratively plan subtasks for a given instruction. Our method encodes object attributes and relationships in graphs through an LLM and Graph Attention Networks, integrating instruction features to predict subtasks consisting of pre-defined robot actions and target objects in the scene graph. This strategy enables robots to acquire semantic knowledge widely observed in the environment from the scene graph. To train and evaluate GRID, we establish a dataset construction pipeline to generate synthetic datasets for graph-based robotic task planning. Experiments have shown that our method outperforms GPT-4 by over 25.4% in subtask accuracy and 43.6% in task accuracy. Moreover, our method achieves a real-time speed of 0.11s per inference. Experiments conducted on datasets of unseen scenes and scenes with varying numbers of objects demonstrate that the task accuracy of GRID declined by at most 3.8%, showcasing its robust cross-scene generalization ability. We validate our method in both physical simulation and the real world. More details can be found on the project page https://jackyzengl.github.io/GRID.github.io/.
Zhe Ni, Xiaoxin Deng, Cong Tai, Xinyue Zhu, Qinghongbing Xie, Weihang Huang, Xiang Wu 0013, Long Zeng 0001
IROS8
2024 ParametricNet++: A 6DoF Pose Estimation Network with Sparse Keypoint Recovery for Parametric Shapes in Stacked Scenarios
abstract
Most industrial parts are designed from parametric shapes with the properties of diversity and uncertainty. We propose a 6DoF pose estimation network, ParametricNet++, based on pointwise regression and sparse keypoint recovery, which is extended from ParametricNet to include optimizations of keypoint selection and prediction. Keypoint selection optimization selects geometrically unique keypoints and keypoint groups to reduce the difficulty of scene keypoint prediction and template keypoint recovery. Keypoint prediction optimization predicts keypoints from rough to precise, which improves the accuracy of scene keypoint prediction and template keypoint recovery. Compared with other state-of-the-art methods, the average of APs of ParametricNet++ is improved by over 15% on the public Siléane dataset, and the average of mAPs is improved by 12% and 14% on L-dataset and G-dataset from Parametric dataset, respectively. In particular, ParametricNet++ outperforms our original ParametricNet by 5% for both learning and generalization ability evaluation on the Parametric dataset. The experimental results demonstrate that ParametricNet++ lays a solid foundation for robot grasping in industrial scenarios.
Wei Jie Lv, Long Zeng 0001
IROS5
2024 Freehand Sketch Generation from Mechanical Components
abstract
Drawing freehand sketches of mechanical components on multimedia devices for AI-based engineering modeling has become a new trend. However, its development is being impeded because existing works cannot produce suitable sketches for data-driven research. These works either generate sketches lacking a freehand style or utilize generative models not originally designed for this task resulting in poor effectiveness. To address this issue, we design a two-stage generative framework mimicking the human sketching behavior pattern, called MSFormer, which is the first time to produce humanoid freehand sketches tailored for mechanical components. The first stage employs Open CASCADE technology to obtain multi-view contour sketches from mechanical components, filtering perturbing signals for the ensuing generation process. Meanwhile, we design a view selector to simulate viewpoint selection tasks during human sketching for picking out information-rich sketches. The second stage translates contour sketches into freehand sketches by a transformer-based generator. To retain essential modeling features as much as possible and rationalize stroke distribution, we introduce a novel edge-constraint stroke initialization. Furthermore, we utilize a CLIP vision encoder and a new loss function incorporating the Hausdorff distance to enhance the generalizability and robustness of the model. Extensive experiments demonstrate that our approach achieves state-of-the-art performance for generating freehand sketches in the mechanical domain. Project page: https://mcfreeskegen.github.io/.
Zhichao Liao, Fengyuan Piao, Xinghui Li, Yue Ma 0033, Pingfa Feng, Heming Fang, Long Zeng 0001
ACM Multimedia8
2024 Refined-mask guided multi-stream blending network
Weijie Lv, Junyu Su, Long Zeng 0001
Multim. Tools Appl.6
2023 Edge-aware Neural Implicit Surface Reconstruction
abstract
Recently, neural implicit 3D reconstruction in indoor scenarios has achieved impressive performance. Utilizing the volume rendering method and neural implicit representation to learn 3D scenes, such per-scene optimization methods could reconstruct pretty complete models but also suffer from missing details and overly-smoothed reconstructions. In this paper, we propose a novel edge-aware neural implicit surface reconstruction method, named Ea-NeuS, to learn high-quality 3D models with fine details. Specifically, we use the edge of objects to locate the important areas, and propose a simple yet effective edge-guided ray-sampling strategy to learn the 3D models. The aforementioned edge information further guides the normal prior supervision, which helps reduce inaccurate optimization in detailed regions. We additionally use the visibility-aware sparse points to pilot the 3D points sampling along the rays and perform explicit supervision. As a result, our method achieves superior performance compared with existing methods on various scenes.
Xinghui Li, Yikang Ding, Xiansong Lai, Shihao Ren, Wensen Feng, Long Zeng 0001
ICME7
2023 Reinforcement Learning Based Pushing and Grasping Objects from Ungraspable Poses
abstract
Grasping an object when it is in an ungraspable pose is a challenging task, such as books or other large flat objects placed horizontally on a table. Inspired by human manipulation, we address this problem by pushing the object to the edge of the table and then grasping it from the hanging part. In this paper, we develop a model-free Deep Reinforcement Learning framework to synergize pushing and grasping actions. We first pre-train a Variational Autoencoder to extract high-dimensional features of input scenario images. One Proximal Policy Optimization algorithm with the common reward and sharing layers of Actor-Critic is employed to learn both pushing and grasping actions with high data efficiency. Experiments show that our one network policy can converge 2.5 times faster than the policy using two parallel networks. Moreover, the experiments on unseen objects show that our policy can generalize to the challenging case of objects with curved surfaces and off-center irregularly shaped objects. Lastly, our policy can be transferred to a real robot without fine-tuning by using CycleGAN for domain adaption and outperforms the push-to-wall baseline.
Hongzhuo Liang, Jianzhi Lyu, Long Zeng 0001, Pingfa Feng, Jianwei Zhang 0001
ICRA5
2023 Domain Adaptation on Point Clouds for 6D Pose Estimation in Bin-Picking Scenarios
abstract
Training with simulated data is a common approach in pose estimation research. However, a sim-to-real gap between clean simulated data and noisy real data will seriously weaken the generalization ability of the algorithm, especially for point clouds. To address this problem, this paper proposes a domain adaptive pose estimation network (DAPE-Net). For the feature extracted from the backbone, the network will conduct the real and simulation discrimination based on a feature discriminator, and complete the pose estimation by adversarial training. This makes the network pay more attention to the domain invariant features of simulation and real point clouds to complete domain adaptation. In our experiment, DAPE-Net improved the performance of pose estimation by 10%. To solve the problem that domain adaptation requires a small amount of real data, we propose a scheme that can semi-automatically collect real data in bin-picking scenarios for 6D pose estimation.
Wei Jie Lv, Long Zeng 0001
IROS5
2022 Class-Aware Contrastive Semi-Supervised Learning
abstract
Pseudo-label-based semi-supervised learning (SSL) has achieved great success on raw data utilization. However, its training procedure suffers from confirmation bias due to the noise contained in self-generated artificial labels. Moreover, the model's judgment becomes noisier in real-world applications with extensive out-of-distribution data. To address this issue, we propose a general method named Class-aware Contrastive Semi-Supervised Learning (CCSSL), which is a drop-in helper to improve the pseudo-label quality and enhance the model's robustness in the real-world setting. Rather than treating real-world data as a union set, our method separately handles reliable in-distribution data with class-wise clustering for blending into downstream tasks and noisy out-of-distribution data with image-wise contrastive for better generalization. Furthermore, by applying target reweighting, we successfully emphasize clean label learning and simultaneously reduce noisy label learning. Despite its simplicity, our proposed CCSSL has significant performance improvements over the state-of-the-art SSL methods on the standard datasets CIFAR100 [18] and STL10 [8]. On the real-world dataset Semi-iNat 2021 [27], we improve FixMatch [25] by 9.80% and CoMatch [19] by 3.18%. Code is available https://github.com/TencentYoutuResearch/Classification-SemiCLS.
Guannan Jiang, Yong Liu 0032, Feng Zheng 0001, Wei Zhang 0217, Chengjie Wang 0001, Long Zeng 0001
CVPR9
2022 PPR-Net++: Accurate 6-D Pose Estimation in Stacked Scenarios
abstract
Most supervised learning-based pose estimation methods for stacked scenes are trained on massive synthetic datasets. In most cases, the challenge is that the learned network on the training dataset is no longer optimal on the testing dataset. To address this problem, we propose a pose regression network PPR-Net++. It transforms each scene point into a point in the centroid space, followed by a clustering process and a voting process. In the training phase, a mapping function between the network’s critical parameter (i.e., the bandwidth of the clustering algorithm) and the compactness of the centroid distributions is obtained. This function is used to adapt the bandwidth between centroid distributions of two different domains. In addition, to further improve the pose estimation accuracy, the network also predicts the confidence of each point, based on its visibility and pose error. Only the points with high confidence have the right to vote for the final object pose. In experiments, our method is trained on the IPA synthetic dataset and compared with the state-of-the-art algorithm. When tested with the public synthetic Siléane dataset, our method is better in all eight objects, where five of them are improved by more than 5% in average precision (AP). On IPA real dataset, our method outperforms a large margin by 20%. This lays a solid foundation for robot grasping in industrial scenarios. Note to Practitioners—Our work is motivated by industrial product assembly based on robot grasping. The industrial parts are usually manufactured by numerical machines and piled in bins. Our method can estimate the poses of visible parts accurately. A pose of a part includes its centroid and spatial orientations. Combined with a depth camera, this algorithm allows an industrial robot to understand complex stacked scenes. We improve the pose estimation accuracy in order to assemble parts with robot grasping, without an additional pose adjuster. Our network can learn from a synthetic dataset and apply it to real-world data, without a significant accuracy drop. The synthetic dataset can be obtained easily by computer simulation programs, so the training data are sufficient. Experiments demonstrate that our method outperforms the state-of-the-art pose estimation approaches.
Long Zeng 0001, Wei Jie Lv, Zhi-Kai Dong, Yong-Jin Liu 0001
IEEE Trans Autom. Sci. Eng.1
2021 Efficient SE(3) Reachability Map Generation via Interplanar Integration of Intra-planar Convolutions
abstract
Convolution has been used for fast computation of reachability maps, but it has high computational costs when performing SE(3) convolution operations for general joint arrangements in industrial robots and 3D workspace. Its application is also limited to planar robots, 2D workspace, or robots with special spatial arrangements for joints. In this paper, we find that the SE(3) convolution can be decomposed into a set of SE(2) convolutions, which significantly reduces the computational complexity when computing the reachability map of high-DOF robotic manipulators in the 3D workspace. We also leverage GPU parallel computing and Fast Fourier transform to further accelerate the computation procedure. We demonstrate the time efficiency and quality of our approach using a set of numerical experiments for constructing reachability maps and also present a multi-robot plant phenotyping system that uses the computed reachability map for efficient viewpoint selection and path planning.
Yiheng Han, Jia Pan 0001, Mengfei Xia, Long Zeng 0001, Yong-Jin Liu 0001
ICRA4
2021 ParametricNet: 6DoF Pose Estimation Network for Parametric Shapes in Stacked Scenarios
abstract
Most industrial parts are parametric and their special properties are not fully explored yet. This paper proposes a new 6DoF pose estimation network for parametric shapes in stacked scenarios (ParametricNet). It treats a parametric shape, instead of a part object, as a category. The keypoints of individual instances are learned with point- wise regression and Hough voting scheme, from which specific parameter values are calculated. Then, the template keypoints are obtained based on the computed parameter values and the parametric shape templates. Finally, the 6DoF pose is estimated by least-square fitting between the individual instance’s and the template’s keypoints & centroid. On the public Siléane dataset, the average of APs of ParametricNet is 96%, compared with 82% for the state-of-the-art method. In addition, a new parametric dataset with four shape templates is constructed, in which the evaluated learning and generalization abilities of ParametricNet outperform the state-of-the-art methods. In particular, for the less symmetric shape, the mAP is improved by over 20%, which is an obvious improvement. Real-world experiments show that our method can grasp parametric shapes with unknown parameter values in stacked scenarios.
Long Zeng 0001, Wei Jie Lv, Yong-Jin Liu 0001
ICRA1
2020 An Efficient Pattern Design Method for Plush Toys Using Component-Based Templates
Yinghan Jin, Wanting Feng, Long Zeng 0001
Comput. Aided Des.3
2019 PPR-Net: Point-wise Pose Regression Network for Instance Segmentation and 6D Pose Estimation in Bin-picking Scenarios
abstract
Accurate object 6D pose estimation is a core task for robot bin-picking applications, especially when objects are randomly stacked with heavy occlusion. To address this problem, this paper proposes a simple but novel Point-wise Pose Regression Network (PPR-Net). For each point in the point cloud, the network regresses a 6D pose of the object instance that the point belongs to. We argue that the regressed poses of points from the same object instance should be located closely in pose space. Thus, these points can be clustered into different instances and their corresponding objects' 6D poses can be estimated simultaneously. In our experiments, PPR-Net outperforms the state-of-the-art approach by 15% - 41% in average precision when evaluated on the benchmark Siléane dataset. In addition, it also works well in real world robot bin-picking tasks.
Zhi-Kai Dong, Long Zeng 0001, Xingyao Yu, Houde Liu
IROS5
2019 Sketch-based Retrieval and Instantiation of Parametric Parts
Long Zeng 0001, Zhi-Kai Dong, Jia-yi Yu
Comput. Aided Des.1
2018 SpiderCNN: Deep Learning on Point Sets with Parameterized Convolutional Filters
Yifan Xu 0005, Tianqi Fan, Mingye Xu, Long Zeng 0001, Yu Qiao 0001
ECCV (8)4
2014 Sketch2Jewelry: Semantic feature modeling for sketch-based jewelry design
Long Zeng 0001, Yong-Jin Liu 0001, Jin Wang 0015, Matthew M. F. Yuen
Comput. Graph.1
2012 Q--Complex: Efficient non-manifold boundary representation with inclusion topology
Long Zeng 0001, Yong-Jin Liu 0001, Matthew M. F. Yuen
Comput. Aided Des.1
2012 Least squares quasi-developable mesh approximation
Long Zeng 0001, Yong-Jin Liu 0001, Matthew M. F. Yuen
Comput. Aided Geom. Des.1
2011 Efficient slicing procedure based on adaptive layer depth normal image
Long Zeng 0001, Lip Man-Lip Lai, Yuen-Hoo Lai, Matthew M. F. Yuen
Comput. Aided Des.1