Liu Liu 0012

dblp:74/7037-12 · DBLP profile ↗
← Back
41ranked-venue papers
9as first author
39since 2021 · last 2026
0000-0003-4218-8008ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 4 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 5 first-author · 20 since 2021Systems, architecture and hardware · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 SIAM: Towards Generalizable Articulated Object Modeling via Single Robot-Object Interaction
abstract
Articulated object modeling, which represents interconnected rigid bodies with their geometry, part segmentation, articulation tree, and physical properties, is crucial for robotic perception and manipulation. Recently existing methods like SAGCI leverage Interactive Perception (IP) to refine models through robot interaction. However, SAGCI suffers from prior-dependency (requiring initialization), neglects kinematic/dynamic constraints, and generates non-watertight meshes. To overcome these limitations, we propose SIAM, a novel framework for efficient and generalizable Single-Interaction Articulated Modeling. Given an initial point cloud, SIAM first enables minimal robot interaction to trigger object motion. It then precisely segments parts by analyzing point cloud differences pre- and post-interaction. For joint parameter estimation, we introduce an optimization incorporating novel kinematic energy constraints, enhancing physical consistency. Finally, we reconstruct a high-quality, topologically watertight mesh by learning 3D Gaussian Primitives from multi-view RGB-D observations under deformation. Extensive experiments on the PartNet-Mobility benchmark demonstrate state-of-the-art articulation modeling performance. Successful real-world deployment with an xArm robot further validates the framework's practicality and transferability. SIAM achieves accurate, prior-free modeling with significantly reduced interaction cost.
Li Zhang 0104, Yan Zhang 0053, Anran Huang, Liu Liu 0012, Dan Guo 0001
AAAI7
2026 Exploring Category-level Articulated Object Pose Tracking on SE(3) Manifolds
abstract
Articulated objects are prevalent in daily life and robotic manipulation tasks. However, compared to rigid objects, pose tracking for articulated objects remains an underexplored problem due to their inherent kinematic constraints. To address these challenges, this work proposes a novel point-pair-based pose tracking framework, termed PPF-Tracker. The proposed framework first performs quasi-canonicalization of point clouds in the SE(3) Lie group space, and then models articulated objects using Point Pair Features (PPF) to predict pose voting parameters by leveraging the invariance properties of SE(3). Finally, semantic information of joint axes is incorporated to impose unified kinematic constraints across all parts of the articulated object. PPF-Tracker is systematically evaluated on both synthetic datasets and real-world scenarios, demonstrating strong generalization across diverse and challenging environments. Experimental results highlight the effectiveness and robustness of PPF-Tracker in multi-frame pose tracking of articulated objects. We believe this work can foster advances in robotics, embodied intelligence, and augmented reality.
Xianhui Meng, Yukang Huo, Li Zhang 0104, Liu Liu 0012, Yan Zhong 0001, Pingrui Zhang, Cewu Lu, Jun Liu 0004
AAAI4
2026 Probing Effective and Efficient Category-Level Articulated Object Pose Perception
abstract
Category-level articulated object pose perception-encompassing both static pose estimation and dynamic pose tracking-is critical for embodied AI systems interacting with complex environments. Due to the inherent complexity and diverse motion structures of articulated objects, existing methods often exhibit limitations in adequately modeling kinematic constraints, handling self-occlusions, and meeting optimization requirements. Building upon EfficientCAPER (Yu et al., 2024), this work introduces CAPER++, a unified framework addressing these limitations through three key innovations: first, a joint-centric hierarchical model decomposes objects into a root part and constrained parts linked by joints, explicitly embedding kinematic constraints for geometrically consistent pose recovery. Second, an SE(3) manifold formulation leverages Lie algebra in the tangent space for singularity-free rotation representation and stable optimization, replacing error-prone direct regression. Third, for tracking, a proxy canonicalization strategy reformulates pose updates as SE(3) increment predictions relative to keyframes, enhanced by a dynamic keyframe mechanism to suppress drift. Extensive experiments on synthetic (ArtImage, PM-Videos), semi-synthetic (ReArtMix, ReArt-Videos), and real-world (RobotArm, RobotArm-Videos) benchmarks demonstrate state-of-the-art accuracy and robustness. CAPER++ achieves real-time inference (50 FPS) without post-processing, significantly advancing category-level articulated perception for real-world applications.
Li Zhang 0104, Xianhui Meng, Liu Liu 0012, Rujing Wang, Cewu Lu, Jun Liu 0004, Hong Zhang 0013
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 iClickSeg: Interactive click segmentation for zero-shot cross-category 3D part segmentation
Xing Yi, Liu Liu 0012, Qiupu Chen, Li Zhang 0104, Dan Guo 0001
Pattern Recognit.2
2026 Pre-Defined Keypoints Worth It: Multi-Modal Learning for Category-Level Articulated Objects Pose Estimation
abstract
Articulated objects play a vital role in daily interactions, but traditional RGB-based pose estimation methods often face challenges such as lighting variations and shadows. To address these limitations, we introduce PAGE, a novel Pre-defined keypoint-based framework for category-level articulation pose estimation via multi-modal AliGnmEnt. Our approach is motivated by the observation that the distance distribution between heuristically generated keypoints and visible points exhibits a divergent pattern, a phenomenon previously overlooked. To tackle this, we propose a customized unsupervised keypoint estimation method that enhances the stability and robustness of model predictions. Furthermore, to minimize mutual information redundancy between point clouds and RGB images, we design a geometry-color alignment module that fuses features after aligning the two modalities. This is followed by decoding the radius for each visible point and applying our proposal integration scoring strategy to predict keypoints. The framework ultimately outputs the per-part 6D pose of the articulated object.We conduct extensive experiments across diverse datasets, ranging from synthetic to real-world scenarios, demonstrating the robustness and superior performance of PAGE. This work holds significant promise for applications in robotics, embodied intelligence, and augmented reality. Codes and datasets are available at the website: https://sites.google.com/view/pageforart.
Li Zhang 0104, Liu Liu 0012, Rujing Wang, Yan Zhong 0001
IEEE Trans Autom. Sci. Eng.3
2026 Distilling Textual Priors From LLM to Efficient Image Fusion
abstract
Multi-modality image fusion aims to synthesize a single, comprehensive image from multiple source inputs. Traditional approaches, such as CNNs and GANs, offer efficiency but struggle to handle low-quality or complex inputs. Recent advances in text-guided methods leverage large model priors to overcome these limitations, but at the cost of significant computational overhead, both in memory and inference time. To address this challenge, we propose a novel framework for distilling large model priors, eliminating the need for text guidance during inference while dramatically reducing model size. Our framework utilizes a teacher-student architecture, where the teacher network incorporates large model priors and transfers this knowledge to a smaller student network via a tailored distillation process. Crucially, our experiments demonstrate that this knowledge transfer is the primary driver of performance gains, rather than mere architectural optimization. Additionally, we introduce a spatial-channel cross-fusion module to enhance the model’s ability to leverage textual priors across both spatial and channel dimensions. Our method achieves a favorable trade-off between computational efficiency and fusion quality. The distilled network, requiring only 10% of the parameters and inference time of the teacher network, retains 90% of its performance and outperforms existing SOTA methods. Extensive experiments demonstrate the effectiveness of our approach. Codes are available at https://github.com/Zirconium233/DTPF.
Xuanhua He, Ke Cao 0001, Liu Liu 0012, Li Zhang 0104, Man Zhou 0003, Jie Zhang 0033, Dan Guo 0001, Meng Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 R^2-Art: Category-Level Articulation Pose Estimation from Single RGB Image via Cascade Render Strategy
abstract
Human life is filled with articulated objects. Previous works for estimating the pose of category-level articulated objects rely on costly 3D point clouds or RGB-D images. In this paper, our goal is to estimate category-level articulation poses from a single RGB image, where we propose R2-Art, a novel category-level Articulation pose estimation framework from a single RGB image and a cascade Render strategy. Given an RGB image as input, R2-Art estimates per-part 6D pose for the articulation. Specifically, we design parallel regression branches tailored to generate camera-to-root translation and rotation. Using the predicted joint states, we perform PC prior transformation and deformation with a joint-centric modeling approach. For further refinement, a cascade render strategy is proposed for projecting the 3D deformed prior onto the 2D mask. Extensive experiments are provided to validate our R2-Art on various datasets ranging from synthetic datasets to real-world scenarios, demonstrating the superior performance and robustness of the R2-Art. We believe that this work has the potential to be applied in many fields including robotics, embodied intelligence, and augmented reality.
Li Zhang 0104, Yukang Huo, Yan Zhong 0001, Rujing Wang, Liu Liu 0012
AAAI8
2025 GaPT-DAR: Category-level Garments Pose Tracking via Integrated 2D Deformation and 3D Reconstruction
abstract
Garments are common in daily life and are important for embodied intelligence community. Current category-level garments pose tracking works focus on predicting point-wise canonical correspondence and learning a shape deformation in point cloud sequences. In this paper, motivated by the 2D warping space and shape prior, we propose GaPT-DAR, a novel category-level Garments Pose Tracking framework with integrated 2D Deformation And 3D Reconstruction function, which fully utilize 3D-2D projection and 2D-3D reconstruction to transform the 3D point-wise learning into 2D warping deformation learning. Specifically, GaPT-DAR firstly builds a Voting-based Project module that learns the optimal 3D-2D projection plane for maintaining the maximum orthogonal entropy during point projection. Next, a Garments Deformation module is designed in 2D space to explicitly model the garments warping procedure with deformation parameters. Finally, we build a Depth Reconstruction module to recover the 2D images into 3D warp field. We provide extensive experiments on VR-Folding dataset to evaluate our GaPT-DAR and the results show obvious improvements on most of the metrics compared to state-of-the-arts (i.e. Garment-Nets [8] and GarmentTracking [32]). More details are available at https://sites.google.com/view/gapt-dar.
Li Zhang 0104, Qiaojun Yu, Lixin Yang 0001, Yong-Lu Li 0001, Cewu Lu, Rujing Wang, Liu Liu 0012
CVPR9
2025 Generalizable Articulated Object Perception with Superpoints
abstract
Manipulating articulated objects with robotic arms is challenging due to the complex kinematic structure, which requires precise part segmentation for efficient manipulation. In this work, we introduce a novel superpoint-based perception method designed to improve part segmentation in 3D point clouds of articulated objects. We propose a learnable, part-aware superpoint generation technique that efficiently groups points based on their geometric and semantic similarities, resulting in clearer part boundaries. Furthermore, by leveraging the segmentation capabilities of the 2D foundation model SAM, we identify the centers of pixel regions and select corresponding superpoints as candidate query points. Integrating a query-based transformer decoder further enhances our method's ability to achieve precise part segmentation. Experimental results on the GAPartNet dataset show that our method outperforms existing state-of-the-art approaches in cross-category part segmentation, achieving AP50 scores of 77.9% for seen categories (4.4% improvement) and 39.3% for unseen categories (11.6% improvement), with superior results in 5 out of 9 part categories for seen objects and outperforming all previous methods across all part categories for unseen objects.
Qiaojun Yu, Ce Hao, Xibin Yuan, Li Zhang 0104, Liu Liu 0012, Yukang Huo, Cewu Lu
ICASSP5
2025 Towards Robust Category-level Articulation Pose Estimation via Integrated Differentiable Rendering
abstract
Accurate object pose estimation is crucial for embodied intelligence tasks such as manipulation, grasping, and human-robot interaction. However, due to the inherent characteristics of articulated objects, such as kinematic constraints and self-occlusion, pose estimation for articulated objects has remained a significant challenge. To address these issues, this paper proposes CAPED, an end-to-end robust Category-level Articulated object Pose Estimator integrated differentiable rendering. Given partial point cloud as input, CAPED outputs the per-part 6D pose for articulation. Specifically, with the proposed joint-centric modeling manner, CAPED firstly estimates the pose for the free part. Afterward, we canonicalize the input point cloud to estimate constrained parts’ poses by predicting the joint parameters and states as replacements. For further refinement, we propose a differentiable rendering scheme for pose optimization. Evaluations of the ArtImage and RobotArm datasets demonstrate that CAPED exhibits outstanding effectiveness and generalization in tasks ranging from synthetic data to real-world scenarios. We will publicly release the code.
Li Zhang 0104, Yukang Huo, Lin Wu 0001, Yanyan Wei, Harshal Suresh Shende, Liu Liu 0012, Linlin Ou
ICASSP9
2025 GASEM: Boosting Generalized and Actionable Parts Segmentation and Pose Estimation via Object Motion Perception
abstract
Category-level object understanding has progressed, but generalized part perception remains underexplored. This paper introduces GASEM, a framework for Generalizable and Actionable Parts GAPart Segmentation and pose Estimation via object Motion perception. GASEM utilizes point-wise motion data from observed point clouds and cross-perspective alignment to learn object motion using a scene flow model. It features a segmentation proposal architecture for GAPart segmentation and an Normalized Object Coordinate Space(NPCS) branch for pose estimation. Additionally, a reinforcement learning agent is trained for robust GAPart manipulation in both simulations and real-world environments. Experiments on the GAPartNet dataset show GASEM outperforms state-of-the-art methods. This work promises advancements in embodied intelligence applications like robot-object interaction and generalizable manipulation. Codes are available at https://github.com/Zirconium233/GASEM.
Liu Liu 0012, Li Zhang 0104, Yiming Tang 0001, Qi Wu 0007, Hao Wu 0040
ICME1
2025 UniAff: A Unified Representation of Affordances for Tool Usage and Articulation with Vision-Language Models
abstract
Previous studies on robotic manipulation are based on a limited understanding of the underlying 3D motion constraints and affordances. To address these challenges, we propose a comprehensive paradigm, termed UniAff, that integrates 3D object-centric manipulation and task understanding in a unified formulation. Specifically, we constructed a dataset labeled with manipulation-related key attributes, comprising 900 articulated objects from 19 categories and 600 tools from 12 categories. Furthermore, we leverage MLLMs to infer object-centric representations for manipulation tasks, including affordance recognition and reasoning about 3D motion constraints. Comprehensive experiments in both simulation and real-world settings indicate that UniAff significantly improves the generalization of robotic manipulation for tools and articulated objects. We hope that UniAff will serve as a general baseline for unified robotic manipulation tasks in the future. Images, videos, dataset and code are published on the project website at:https://sites.google.com/view/uni-aff/home.
Qiaojun Yu, Siyuan Huang 0004, Xibin Yuan, Zhengkai Jiang 0001, Ce Hao, Xin Li 0110, Haonan Chang, Junbo Wang 0004, Liu Liu 0012, Hongsheng Li 0001, Peng Gao 0007, Cewu Lu
ICRA9
2025 Pre-defined Keypoints Promote Category-level Articulation Pose Estimation via Multi-Modal Alignment
abstract
Articulations are essential in everyday interactions, yet traditional RGB-based pose estimation methods often struggle with issues such as lighting variations and shadows. To overcome these challenges, we propose a novel Pre-defined keypoint based framework for category-level articulation pose estimation via multi-modal Alignment, coined PAGE. Specifically, we first propose a customized keypoint estimation method, aiming to avoid the divergent distance pattern between heuristically generated keypoints and visible points. In addition, to reduce the mutual information redundancy between point clouds and RGB images, we design the geometry-color alignment, which fuses the features after aligning two modalities. This is followed by decoding the radius for each visible point, and applying our proposal integration scoring strategy to predict keypoints. Ultimately, the framework outputs the per-part 6D pose of the articulation. We conduct extensive experiments to evaluate PAGE across a variety of datasets, from synthetic to real-world scenarios, demonstrating its robustness and superior performance.
Li Zhang 0104, Liu Liu 0012, Yan Zhong 0001, Rujing Wang
IJCAI3
2025 Generalizable and Actionable Part Detection and Manipulation with SAM-rectified Segmentation and Iterative Pose Refinement
abstract
The ability to perform cross-category object perception and manipulation is highly desirable in building intelligent robots. One promising approach is to define the concept of Generalizable and Actionable Parts (GAParts), such as buttons and handles, on both seen and unseen object categories. However, the accurate cross-category perception of GAParts is still challenging due to the large inter-category object shape variations. To address this issue, we introduce SAMIR, a novel framework using SAM-rectified segmentation and Iterative pose Refinement for GAPart detection and manipulation. Firstly, we introduce a Segment Anything (SAM) segmentation prior to rectify the unconfident, fragmented GAPart instance proposals. Secondly, in addition to the zero-shot generalization of the SAM foundation model, we further finetune it with a lightweight adaptor model on our task dataset. Finally, we propose an iterative pose refinement procedure that improves the accuracy of GAPart pose estimation. Our perception experiments on GAPartNet dataset show that SAMIR consistently outperforms the baseline method on instance segmentation and pose estimation tasks. Our manipulation experiments in Sapien simulator illustrate that SAMIR leads to an improved manipulation success rate. We also deploy our method to a real robot for real-world manipulation. Our code and video are available at sites.google.com/view/samir-gapart.
Sucheng Qian, Li Zhang 0104, Yanyan Wei, Liu Liu 0012, Cewu Lu
IROS4
2025 ArtGS: 3D Gaussian Splatting for Interactive Visual-Physical Modeling and Manipulation of Articulated Objects
abstract
Articulated object manipulation remains a critical challenge in robotics due to the complex kinematic constraints and the limited physical reasoning of existing methods. In this work, we introduce ArtGS, a novel framework that extends 3D Gaussian Splatting (3DGS) by integrating visual-physical modeling for articulated object understanding and interaction. ArtGS begins with multi-view RGB-D reconstruction, followed by reasoning with a vision-language model (VLM) to extract semantic and structural information, particularly the articulated bones. Through dynamic, differentiable 3DGS-based rendering, ArtGS optimizes the parameters of the articulated bones, ensuring physically consistent motion constraints and enhancing the manipulation policy. By leveraging dynamic Gaussian splatting, cross-embodiment adaptability, and closed-loop optimization, ArtGS establishes a new framework for efficient, scalable, and generalizable articulated object modeling and manipulation. Experiments conducted in both simulation and real-world environments demonstrate that ArtGS significantly outperforms previous methods in joint estimation accuracy and manipulation success rates across a variety of articulated objects. Additional images and videos are available on the project website: sites.google.com/view/artgs.
Qiaojun Yu, Xibin Yuan, Dongzhe Zheng, Ce Hao, Yang You 0004, Yixing Chen 0008, Yao Mu 0001, Liu Liu 0012, Cewu Lu
IROS10
2025 DFGAP: Towards Depth-Free Cross-Category GAParts Perception via Uncertainty-Quantified Modeling
abstract
Cross-category object perception is one of the essential upstream tasks for generelizable robot object interaction and manipulation. Recently, an increasing number of researchers are focusing on investigating visual Generalizable and Actionable Parts understanding at cross-category level perception. However, these works are built upon the RGB-D or point cloud input, that relies on the depth information capture. Under the circumstances of limited depth camera performance, e.g. transparent or light absorbing material, perception algorithms that do not require depth information are urgently needed. In this paper, we propose DFGAP, a novel depth-free framework for RGB-based GAParts segmentation and pose estimation. Specifically, we independently model the ill-pose problems from the absence of depth for GAPart segmentation and pose estimation, by clearly quantifying the pixel-wise segmentation probability and relative depth. We reduce the uncertainty and benefit learning in these two tasks. The experimental results demonstrate the superior performance and robustness of our DFGAP. Our work provides a new research paradigm in GAParts perception. We believe that our work has the enormous potential to be applied in many areas of embodied AI system.
Xueyu Yuan, Jiangqi Song, Liu Liu 0012, Li Zhang 0104, Dan Guo 0001, Richang Hong, Meng Wang 0002
ACM Multimedia4
2025 REArtGS: Reconstructing and Generating Articulated Objects via 3D Gaussian Splatting with Geometric and Motion Constraints
abstract
Articulated objects, as prevalent entities in human life, their 3D representations play crucial roles across various applications. However, achieving both high-fidelity textured surface reconstruction and dynamic generation for articulated objects remains challenging for existing methods. In this paper, we present REArtGS, a novel framework that introduces additional geometric and motion constraints to 3D Gaussian primitives, enabling realistic surface reconstruction and generation for articulated objects. Specifically, given multi-view RGB images of arbitrary two states of articulated objects, we first introduce an unbiased Signed Distance Field (SDF) guidance to regularize Gaussian opacity fields, enhancing geometry constraints and improving surface reconstruction quality. Then we establish deformable fields for 3D Gaussians constrained by the kinematic structures of articulated objects, achieving unsupervised generation of surface meshes in unseen states. Extensive experiments on both synthetic and real datasets demonstrate our approach achieves high-quality textured surface reconstruction for given states, and enables high-fidelity surface generation for unseen states. Project site: https://sites.google.com/view/reartgs/home.
Liu Liu 0012, Zhou Linli, Anran Huang, Liangtu Song, Qiaojun Yu, Qi Wu 0007, Cewu Lu
NeurIPS2
2025 PanDiT: A Few-Step Diffusion Transformer for High-Fidelity and Efficient Pansharpening
abstract
Pansharpening plays a crucial role in remote sensing by fusing low-resolution multispectral (LRMS) images and high-resolution panchromatic (PAN) images to generate high-resolution multispectral (HRMS) images. While denoising diffusion models offer potential for high-fidelity image generation, their practical application in pansharpening has been severely hindered by huge computational costs from iterative sampling and naive conditioning strategies that struggle to fuse multi-modal information effectively. In this paper, we introduce PanDiT, a novel Diffusion Transformer framework designed to address these challenges. PanDiT is built on the core principle of a Decoupled Conditioning Mechanism, which explicitly disentangles and injects spatial and spectral guidance, and is engineered for practical, few-step inference. Our framework leverages a powerful Diffusion Transformer (DiT) backbone, where conditioning is achieved through two specialized injection blocks that capture spatial and time-frequency features. Crucially, by integrating an implicit sampling strategy, we accelerate the inference process to as few as two steps. Extensive experiments on multiple benchmark datasets demonstrate that PanDiT not only establishes a new state-of-the-art in fusion quality and achieves a good quality-efficiency trade-off. Code is available at https://github.com/para133/PanDIT.
Jiabin Fang, Ke Cao 0001, Xuanhua He, Jie Zhang 0033, Man Zhou 0003, Liu Liu 0012
IEEE Trans. Geosci. Remote. Sens.6
2024 KPA-Tracker: Towards Robust and Real-Time Category-Level Articulated Object 6D Pose Tracking
abstract
Our life is populated with articulated objects. Current category-level articulation estimation works largely focus on predicting part-level 6D poses on static point cloud observations. In this paper, we tackle the problem of category-level online robust and real-time 6D pose tracking of articulated objects, where we propose KPA-Tracker, a novel 3D KeyPoint based Articulated object pose Tracker. Given an RGB-D image or a partial point cloud at the current frame as well as the estimated per-part 6D poses from the last frame, our KPA-Tracker can effectively update the poses with learned 3D keypoints between the adjacent frames. Specifically, we first canonicalize the input point cloud and formulate the pose tracking as an inter-frame pose increment estimation task. To learn consistent and separate 3D keypoints for every rigid part, we build KPA-Gen that outputs the high-quality ordered 3D keypoints in an unsupervised manner. During pose tracking on the whole video, we further propose a keypoint-based articulation tracking algorithm that mines keyframes as reference for accurate pose updating. We provide extensive experiments on validating our KPA-Tracker on various datasets ranging from synthetic point cloud observation to real-world scenarios, which demonstrates the superior performance and robustness of the KPA-Tracker. We believe that our work has the potential to be applied in many fields including robotics, embodied intelligence and augmented reality. All the datasets and codes are available at https://github.com/hhhhhar/KPA-Tracker.
Liu Liu 0012, Anran Huang, Qi Wu 0007, Dan Guo 0001, Xun Yang 0001, Meng Wang 0001
AAAI1
2024 ICAF-4: An Integrated Framework of Category-level Articulated Object Perception and Manipulation for Embodied Intelligence
Li Zhang 0104, Qiankun Li 0004, Qi Wu 0007, Lin Wu 0001, Liu Liu 0012
BMVC6
2024 U-COPE: Taking a Further Step to Universal 9D Category-Level Object Pose Estimation
Li Zhang 0104, Weiqing Meng, Yan Zhong 0001, Jianming Du, Rujing Wang, Liu Liu 0012
ECCV (10)9
2024 GAMMA: Generalizable Articulation Modeling and Manipulation for Articulated Objects
abstract
Articulated objects like cabinets and doors are widespread in daily life. However, directly manipulating 3D articulated objects is challenging because they have diverse geometrical shapes, semantic categories, and kinetic constraints. Prior works mostly focused on recognizing and manipulating articulated objects with specific joint types. They can either estimate the joint parameters or distinguish suitable grasp poses to facilitate trajectory planning. Although these approaches have succeeded in certain types of articulated objects, they lack generalizability to unseen objects, which significantly impedes their application in broader scenarios. In this paper, we propose a novel framework of Generalizable Articulation Modeling and Manipulating for Articulated Objects (GAMMA), which learns both articulation modeling and grasp pose affordance from diverse articulated objects with different categories. In addition, GAMMA adopts adaptive manipulation to iteratively reduce the modeling errors and enhance manipulation performance. We train GAMMA with the PartNet-Mobility dataset and evaluate with comprehensive experiments in SAPIEN simulation and real-world Franka robot. Results show that GAMMA significantly outperforms SOTA articulation modeling and manipulation algorithms in unseen and cross-category articulated objects. Images, videos and codes are published on the project website at: sites.google.com/view/gamma-articulation.
Qiaojun Yu, Junbo Wang 0004, Wenhai Liu, Ce Hao, Liu Liu 0012, Lin Shao 0002, Cewu Lu
ICRA5
2024 RPMArt: Towards Robust Perception and Manipulation for Articulated Objects
abstract
Articulated objects are commonly found in daily life. It is essential that robots can exhibit robust perception and manipulation skills for articulated objects in real-world robotic applications. However, existing methods for articulated objects insufficiently address noise in point clouds and struggle to bridge the gap between simulation and reality, thus limiting the practical deployment in real-world scenarios. To tackle these challenges, we propose a framework towards Robust Perception and Manipulation for Articulated Objects (RPMArt), which learns to estimate the articulation parameters and manipulate the articulation part from the noisy point cloud. Our primary contribution is a Robust Articulation Network (RoArtNet) that is able to predict both joint parameters and affordable points robustly by local feature learning and point tuple voting. Moreover, we introduce an articulation-aware classification scheme to enhance its ability for sim-to-real transfer. Finally, with the estimated affordable point and articulation joint constraint, the robot can generate robust actions to manipulate articulated objects. After learning only from synthetic data, RPMArt is able to transfer zero-shot to real-world articulated objects. Experimental results confirm our approach’s effectiveness, with our framework achieving state-of-the-art performance in both noise-added simulation and real-world environments. Code, data and more results can be found on the project website at https://r-pmart.github.io.
Junbo Wang 0004, Wenhai Liu, Qiaojun Yu, Yang You 0004, Liu Liu 0012, Cewu Lu
IROS5
2024 Thermal-NeRF: Neural Radiance Fields from an Infrared Camera
abstract
In recent years, Neural Radiance Fields (NeRFs) have demonstrated significant potential in encoding highly-detailed 3D geometry and environmental appearance, positioning themselves as a promising alternative to traditional explicit representation for 3D scene reconstruction. However, the predominant reliance on RGB imaging presupposes ideal lighting conditions—a premise frequently unmet in robotic applications plagued by poor lighting or visual obstructions. This limitation overlooks the capabilities of infrared (IR) cameras, which excel in low-light detection and present a robust alternative under such adverse scenarios. To tackle these issues, we introduce Thermal-NeRF, the first method that estimates a volumetric scene representation in the form of a NeRF solely from IR imaging. By leveraging a thermal mapping and structural thermal constraint derived from the thermal characteristics of IR imaging, our method showcases unparalleled proficiency in recovering NeRFs in visually degraded scenes where RGB-based methods fall short. We conduct extensive experiments to demonstrate that Thermal-NeRF can achieve superior quality compared to existing methods. Furthermore, we contribute a dataset for IR-based NeRF applications, paving the way for future research in IR NeRF reconstruction, see https://github.com/Cerf-Volant425/Thermal-NeRF.
Tianxiang Ye, Qi Wu 0007, Junyuan Deng, Liu Liu 0012, Songpengcheng Xia, Wenxian Yu, Ling Pei
IROS5
2024 EfficientCAPER: An End-to-End Framework for Fast and Robust Category-Level Articulated Object Pose Estimation
abstract
Human life is populated with articulated objects. Pose estimation for category-level articulated objects is a significant challenge due to their inherent complexity and diverse kinematic structures. Current methods for this task usually meet the problems of insufficient consideration of kinematic constraints, self-occlusion, and optimization requirements. In this paper, we propose EfficientCAPER, an end-to-end Category-level Articulated object Pose EstimatoR, eliminating the need for optimization functions as post-processing and utilizing the kinematic structure for joint-centric pose modeling, thus enhancing the efficiency and applicability. Given a partial point cloud as input, the EfficientCAPER firstly estimates the pose for the free part of an articulated object using decoupled rotation representation. Next, we canonicalize the input point cloud to estimate constrained parts' poses by predicting the joint parameters and states as replacements. Evaluations on three diverse datasets, ArtImage, ReArtMix, and RobotArm, show EfficientCAPER's effectiveness and generalization ability to real-world scenarios. The framework exhibits excellent static pose estimation performance for articulated objects, contributing to the advancement of category-level pose estimation. Codes will be made publicly available.
Li Zhang 0104, Lin Wu 0001, Linlin Ou, Liu Liu 0012
NeurIPS6
2024 Rethinking 3D Convolution in $\ell_p$-norm Space
abstract
Convolution is a fundamental operation in the 3D backbone. However, under certain conditions, the feature extraction ability of traditional convolution methods may be weakened. In this paper, we introduce a new convolution method based on $\ell_p$-norm. For theoretical support, we prove the universal approximation theorem for $\ell_p$-norm based convolution, and analyze the robustness and feasibility of $\ell_p$-norms in 3D point cloud tasks. Concretely, $\ell_{\infty}$-norm based convolution is prone to feature loss. $\ell_2$-norm based convolution is essentially a linear transformation of the traditional convolution. $\ell_1$-norm based convolution is an economical and effective feature extractor. We propose customized optimization strategies to accelerate the training process of $\ell_1$-norm based Nets and enhance the performance. Besides, a theoretical guarantee is given for the convergence by \textit{regret} argument. We apply our methods to classic networks and conduct related experiments. Experimental results indicate that our approach exhibits competitive performance with traditional CNNs, with lower energy consumption and instruction latency.
Li Zhang 0104, Yan Zhong 0001, Zhe Min, RujingWang, Liu Liu 0012
NeurIPS6
2024 EACT-Det: An Efficient Adjusting Criss-cross windows Transformer Embedding Pyramid Networks for Similar Disease Detection
Fenmei Wang, Rujing Wang, Ziliang Huang, Shifeng Dong, Xiuzhen Wang, Qiong Zhou, Shijian Zheng, Liu Liu 0012
Multim. Tools Appl.8
2023 Category-Level Articulated Object 9D Pose Estimation via Reinforcement Learning
abstract
Human life is populated with articulated objects. Current category-level articulated object 9D pose estimation (Articulated Object 9D Pose Estimation, ArtOPE) methods usually meet the challenges of shared object representation requirement, kinematics-agnostic pose modeling and self-occlusions. In this paper, we propose a novel framework called Articulated object 9D Pose Estimation via Reinforcement Learning (ArtPERL), which formulates the category-level ArtOPE as a reinforcement learning problem. Given a point cloud or RGB-D image input, ArtPERL firstly retrieves the part-sensitive articulated object as reference point cloud, and then introduces a joint-centric pose modeling strategy that estimates 9D pose by fitting joint states via reinforced agent training. Finally, we further propose a pose optimization that refine the predicted 9D pose considering kinematic constraints. We evaluate our ArtPERL on various datasets ranging from synthetic point cloud to real-world multi-hinged object. Experiments demonstrate the superior performance and robustness of our ArtPERL. Our work provides a new perspective on category-level articulated object 9D pose estimation and has the potential to be applied in many fields, including robotics, augmented reality, and autonomous driving.
Liu Liu 0012, Jianming Du, Hao Wu 0040, Xun Yang 0001, Zhenguang Liu, Richang Hong, Meng Wang 0001
ACM Multimedia1
2022 AKB-48: A Real-World Articulated Object Knowledge Base
abstract
Human life is populated with articulated objects. A comprehensive understanding of articulated objects, namely appearance, structure, physical property, and semantics, will benefit many research communities. As current articulated object understanding solutions are usually based on synthetic object dataset with CAD models without physics properties, which prevent satisfied generalization from simulation to real-world applications in visual and robotics tasks. To bridge the gap, we present AKB-48: a large-scale Articulated object Knowledge Base which consists of 2,037 real-world 3D articulated object models of 48 categories. Each object is described by a knowledge graph ArtiKG. To build the AKB-48, we present a fast articulation knowledge modeling (FArM) pipeline, which can fulfill the ArtiKG for an articulated object within 10–15 minutes, and largely reduce the cost for object modeling in the real world. Using our dataset, we propose AKBNet, an integral pipeline for Category-level Visual Articulation Manipulation (C-VAM) task, in which we benchmark three sub-tasks, namely pose estimation, object reconstruction and manipulation. Dataset, codes, and models are publicly available at https://liuliu66.github.io/AKB-48.
Liu Liu 0012, Haoyuan Fu, Sucheng Qian, Qiaojun Yu, Cewu Lu
CVPR1
2022 OakInk: A Large-scale Knowledge Repository for Understanding Hand-Object Interaction
abstract
Learning how humans manipulate objects requires machines to acquire knowledge from two perspectives: one for understanding object affordances and the other for learning human's interactions based on the affordances. Even though these two knowledge bases are crucial, we find that current databases lack a comprehensive awareness of them. In this work, we propose a multi-modal and rich-annotated knowledge repository, OakInk, for visual and cognitive understanding of hand-object interactions. We start to collect 1,800 common household objects and annotate their affordances to construct the first knowledge base: Oak. Given the affordance, we record rich human interactions with 100 selected objects in Oak. Finally, we transfer the interactions on the 100 recorded objects to their virtual counterparts through a novel method: Tink. The recorded and transferred hand-object interactions constitute the second knowledge base: Ink. As a result, OakInk contains 50,000 distinct affordance-aware and intent-oriented hand-object interactions. We benchmark OakInk on pose estimation and grasp generation tasks. Moreover, we propose two practical applications of OakInk: intent-based interaction generation and handover generation. Our dataset and source code are publicly available at www.oakink.net.
Lixin Yang 0001, Kailin Li 0001, Xinyu Zhan 0001, Fei Wu 0001, Anran Xu 0003, Liu Liu 0012, Cewu Lu
CVPR6
2022 Towards densely clustered tiny pest detection in the wild environment
Jianming Du, Liu Liu 0012, Rui Li 0027, Lin Jiao, Chengjun Xie, Rujing Wang
Neurocomputing2
2022 A global activated feature pyramid network for tiny pest detection in the wild
Liu Liu 0012, Rujing Wang, Chengjun Xie, Rui Li 0027, Fangyuan Wang 0001
Mach. Vis. Appl.1
2022 When Pansharpening Meets Graph Convolution Network and Knowledge Distillation
abstract
In this article, we propose a novel graph convolutional network (GCN) for pansharpening, defined as GCPNet, which consists of three main modules: the spatial GCN module (SGCN), the spectral band GCN module (BGCN), and the atrous spatial pyramid module (ASPM). Specifically, due to the nature of GCN, the proposed SGCN and BGCN are capable of exploring the long-range relationship between the object and the global state in the spatial and spectral aspects, which benefits pansharpened results and has not been fully investigated before. In addition, the designed ASPM is equipped with multiscale atrous convolutions and learns richer local feature information, so as to cover the objects of different sizes in satellite images. To further enhance the representation of our proposed GCPNet, asynchronous knowledge distillation is introduced to provide compact features by heterogeneous task imitation in a teacher–student paradigm. In the paradigm, the teacher network acts as a variational autoencoder to extract compact features of the ground-truth MS images. The student network, devised for pansharpening, is trained with the assistance of the teacher network to transfer the important information of the expected ground-truth MS images. Extensive experimental results on different satellite datasets demonstrate that our proposed network outperforms the state-of-the-art methods both visually and quantitatively. The source code is released athttps://github.com/Keyu-Yan/GCPNet.
Man Zhou 0003, Liu Liu 0012, Chengjun Xie, Danfeng Hong
IEEE Trans. Geosci. Remote. Sens.3
2022 Toward Real-World Category-Level Articulation Pose Estimation
abstract
Human life is populated with articulated objects. Current Category-level Articulation Pose Estimation (CAPE) methods are studied under the single-instance setting with a fixed kinematic structure for each category. Considering these limitations, we aim to study the problem of estimating part-level 6D pose for multiple articulated objects with unknown kinematic structures in a single RGB-D image, and reform this problem setting for real-world environments and suggest a CAPE-Real (CAPER) task setting. This setting allows varied kinematic structures within a semantic category, and multiple instances to co-exist in an observation of real world. To support this task, we build an articulated model repository ReArt-48 and present an efficient dataset generation pipeline, which contains Fast Articulated Object Modeling (FAOM) and Semi-Authentic MixEd Reality Technique (SAMERT). Accompanying the pipeline, we build a large-scale mixed reality dataset ReArtMix and a real world dataset ReArtVal. Accompanying the CAPER problem and the dataset, we propose an effective framework that exploits RGB-D input to estimate part-level pose for multiple instances in a single forward pass. In our method, we introduce object detection from RGB-D input to handle the multi-instance problem and segment each instance into several parts. To address the unknown kinematic structure issue, we propose an Articulation Parsing Network to analyze the structure of detected instance, and also build a Pair Articulation Pose Estimation module to estimate per-part 6D pose as well as joint property from connected part pairs. Extensive experiments demonstrate that the proposed method can achieve good performance on CAPER, CAPE and instance-level Robot Arm pose estimation problems. We believe it could serve as a strong baseline for future research on the CAPER task. The datasets and codes in our work will be made publicly available.
Liu Liu 0012, Haoyuan Fu, Cewu Lu
IEEE Trans. Image Process.1
2021 OMAD: Object Model with Articulated Deformations for Pose Estimation and Retrieval
Liu Liu 0012, Haoyuan Fu, Cewu Lu
BMVC2
2021 Reinforcedet: Object Detection By Integrating Reinforcement Learning With Decoupled Pipeline
abstract
Recent object detection methods largely rely on numerous pre-defined anchors that suffer from huge computational cost and resource consumption. To solve this issue, we propose a low-memory deep reinforcement learning based anchor-free object detection approach, namely ReinforceDet, which computes few but accurate region proposals for detection. Specifically, the extracted feature maps are fed into a reinforcement learning network to localize objects as initial region proposals with our re-designed reward function and then adopt another neural network to refine them. To speed up this process in test phase, we decouple the two-branch CNN networks as light-head cascaded subnetworks, named IoU-net and bounding box net. Experimental results show that ReinforceDet could obtain the state-of-the-art performance with much lower compitational and memory cost.
Man Zhou 0003, Liu Liu 0012, Rujing Wang
ICIP2
2021 ReinforceNet: A reinforcement learning embedded object detection framework with region selection network
Man Zhou 0003, Rujing Wang, Chengjun Xie, Liu Liu 0012, Rui Li 0027, Fangyuan Wang 0001, Dengshan Li
Neurocomputing4
2021 Learning region-guided scale-aware feature selection for object detection
Liu Liu 0012, Rujing Wang, Chengjun Xie, Rui Li 0027, Fangyuan Wang 0001, Man Zhou 0003
Neural Comput. Appl.1
2021 Deep Learning Based Automatic Multiclass Wild Pest Monitoring Approach Using Hybrid Global and Local Activated Features
abstract
Specialized control of pests and diseases have been a high-priority issue for the agriculture industry in many countries. On account of automation and cost effectiveness, image analytic pest recognition systems are widely utilized in practical crops prevention applications. But due to powerless hand-crafted features, current image analytic approaches achieve low accuracy and poor robustness in practical large-scale multiclass pest detection and recognition. To tackle this problem, this article proposes a novel deep learning based automatic approach using hybrid and local activated features for pest monitoring. In the presented method, we exploit the global information from feature maps to build our global activated feature pyramid network to extract pests' highly discriminative features across various scales over both depth and position levels. It makes changes of depth or spatial sensitive features in pest images more visible during downsampling. Next, an improved pest localization module named local activated region proposal network is proposed to find the precise pest objects positions by augmenting contextualized and attentional information for feature completion and enhancement in local level. The approach is evaluated on our seven-year large-scale pest data-set containing 88.6 K images (16 types of pests) with 582.1 K manually labeled pest objects. The experimental results show that our solution performs over 75.03% mean average precision (mAP) in industrial circumstances, which outweighs two other state-of-the-art methods: Faster R-CNN with mAP up to 70% and feature pyramid network mAP up to 72%.
Liu Liu 0012, Chengjun Xie, Rujing Wang, Po Yang 0001, Sud Sudirman, Jie Zhang 0033, Rui Li 0027, Fangyuan Wang 0001
IEEE Trans. Ind. Informatics1
2020 FPHA-Afford: A Domain-Specific Benchmark Dataset for Occluded Object Affordance Estimation in Human-Object-Robot Interaction
abstract
In human-object-robot interactions, the recent explosion of standard datasets has offered promising opportunities for deep learning techniques in understanding the functionalities of object parts. But most of existing datasets are only suitable for the applications where objects are non-occluded or isolated during interaction while occlusion is a common challenge in practical object affordance estimation task. In this paper, we attempt to address this issue by introducing a new benchmark dataset named FPHA-Afford that is built upon the popular dataset FPHA. In FPHA-Afford, we adopt egocentric-view to pre-process the videos from FPHA and select part of the frames that contain objects under the strong occlusion of hand. To transfer the domain of FPHA into object affordance estimation task, all of the frames are re-annotated with pixel-level affordance masks. In total, our FPHA-Afford collects 61 videos containing 4.3K frames with 6.55K annotated affordance masks belonging to 9 classes. Some of state-of-the-art semantic segmentation architectures are explored and evaluated over FPHA-Afford. We believe the scale, diversity and novelty of our FPHA-Afford could offer great opportunities to researchers in the computer vision community and beyond. Our dataset and experiment code will be made publicly available on https://github.com/Hussainflr/FPHA-Afford
S. Muzamil Hussain, Liu Liu 0012, Cewu Lu
ICIP2
2019 Deep Learning based Automatic Approach using Hybrid Global and Local Activated Features towards Large-scale Multi-class Pest Monitoring
abstract
Monitoring pest in agriculture has been a high-priority issue all over the world. Computer vision techniques are widely utilized in practical crop pest prevention applications due to the rapid development of artificial intelligence technology. However, current deep learning image analytic approaches achieve low accuracy and poor robustness in agriculture pest monitoring task. This paper targets at this challenge by proposing a novel two-stage deep learning based automatic pest monitoring system with hybrid global and local activated feature. In this approach, a Global activated Feature Pyramid Network (GaFPN) is firstly proposed for extracting highly representative features of pests over both depth and spatial position activation levels. Then, an improved Local activated Region Proposal Network (LaRPN) augmenting contextual and attentional information is represented for precisely locating pest objects. Finally, we design a fully connected neural network to estimate the severity of input image under the detected pests. The experimental results on our 88.6K images dataset (with 16 types of common pests) show that our approach outweighs the state-of-the-art methods in industrial circumstances.
Liu Liu 0012, Rujing Wang, Chengjun Xie, Po Yang 0001, Sud Sudirman, Fangyuan Wang 0001, Rui Li 0027
INDIN1