VLDB 2026 Research / reviewers in the wild / expert
Yukang Huo
dblp:368/0108
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
0009-0001-9569-7028ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Category-level Articulated Object Pose Tracking on SE(3) ManifoldsabstractArticulated objects are prevalent in daily life and robotic manipulation tasks. However, compared to rigid objects, pose tracking for articulated objects remains an underexplored problem due to their inherent kinematic constraints. To address these challenges, this work proposes a novel point-pair-based pose tracking framework, termed PPF-Tracker. The proposed framework first performs quasi-canonicalization of point clouds in the SE(3) Lie group space, and then models articulated objects using Point Pair Features (PPF) to predict pose voting parameters by leveraging the invariance properties of SE(3). Finally, semantic information of joint axes is incorporated to impose unified kinematic constraints across all parts of the articulated object. PPF-Tracker is systematically evaluated on both synthetic datasets and real-world scenarios, demonstrating strong generalization across diverse and challenging environments. Experimental results highlight the effectiveness and robustness of PPF-Tracker in multi-frame pose tracking of articulated objects. We believe this work can foster advances in robotics, embodied intelligence, and augmented reality. Xianhui Meng, Yukang Huo, Li Zhang 0104, Liu Liu 0012, Yan Zhong 0001, Pingrui Zhang, Cewu Lu, Jun Liu 0004 |
AAAI | 2 |
| 2026 | Neural Radiance Field-Based Visual Rendering: A Comprehensive ReviewabstractNeural Radiance Field (NeRF) is a groundbreaking paradigm in neural implicit representations that revolutionized 3D reconstruction, rendering, and dynamic scene modeling. To address cross-domain fragmentation and unclear technical pathways, we present a systematic framework that surveys theoretical foundations, benchmark datasets, methodological advances, and application scenarios. We begin by analyzing NeRF's core mechanisms, including radiance field modeling and differentiable volume rendering, and by defining standardized evaluation benchmarks. Then we chart evolutionary pathways in model optimization, input adaptation, and dynamic scene modeling and analyze how key methods are linked. Furthermore, we provide task-specific insights that highlight migration bottlenecks and potential remedies across digital content creation, embodied perception, and other specialized domains. We also provide comprehensive references and forward-looking guidance for further theoretical refinements and cross-disciplinary deployment of NeRF-based technologies. Mingyuan Yao, Yukang Huo, Yang Ran, Qingbin Tian |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | R^2-Art: Category-Level Articulation Pose Estimation from Single RGB Image via Cascade Render StrategyabstractHuman life is filled with articulated objects. Previous works for estimating the pose of category-level articulated objects rely on costly 3D point clouds or RGB-D images. In this paper, our goal is to estimate category-level articulation poses from a single RGB image, where we propose R2-Art, a novel category-level Articulation pose estimation framework from a single RGB image and a cascade Render strategy. Given an RGB image as input, R2-Art estimates per-part 6D pose for the articulation. Specifically, we design parallel regression branches tailored to generate camera-to-root translation and rotation. Using the predicted joint states, we perform PC prior transformation and deformation with a joint-centric modeling approach. For further refinement, a cascade render strategy is proposed for projecting the 3D deformed prior onto the 2D mask. Extensive experiments are provided to validate our R2-Art on various datasets ranging from synthetic datasets to real-world scenarios, demonstrating the superior performance and robustness of the R2-Art. We believe that this work has the potential to be applied in many fields including robotics, embodied intelligence, and augmented reality. Li Zhang 0104, Yukang Huo, Yan Zhong 0001, Rujing Wang, Liu Liu 0012 |
AAAI | 3 |
| 2025 | Generalizable Articulated Object Perception with SuperpointsabstractManipulating articulated objects with robotic arms is challenging due to the complex kinematic structure, which requires precise part segmentation for efficient manipulation. In this work, we introduce a novel superpoint-based perception method designed to improve part segmentation in 3D point clouds of articulated objects. We propose a learnable, part-aware superpoint generation technique that efficiently groups points based on their geometric and semantic similarities, resulting in clearer part boundaries. Furthermore, by leveraging the segmentation capabilities of the 2D foundation model SAM, we identify the centers of pixel regions and select corresponding superpoints as candidate query points. Integrating a query-based transformer decoder further enhances our method's ability to achieve precise part segmentation. Experimental results on the GAPartNet dataset show that our method outperforms existing state-of-the-art approaches in cross-category part segmentation, achieving AP50 scores of 77.9% for seen categories (4.4% improvement) and 39.3% for unseen categories (11.6% improvement), with superior results in 5 out of 9 part categories for seen objects and outperforming all previous methods across all part categories for unseen objects. Qiaojun Yu, Ce Hao, Xibin Yuan, Li Zhang 0104, Liu Liu 0012, Yukang Huo, Cewu Lu |
ICASSP | 6 |
| 2025 | Towards Robust Category-level Articulation Pose Estimation via Integrated Differentiable RenderingabstractAccurate object pose estimation is crucial for embodied intelligence tasks such as manipulation, grasping, and human-robot interaction. However, due to the inherent characteristics of articulated objects, such as kinematic constraints and self-occlusion, pose estimation for articulated objects has remained a significant challenge. To address these issues, this paper proposes CAPED, an end-to-end robust Category-level Articulated object Pose Estimator integrated differentiable rendering. Given partial point cloud as input, CAPED outputs the per-part 6D pose for articulation. Specifically, with the proposed joint-centric modeling manner, CAPED firstly estimates the pose for the free part. Afterward, we canonicalize the input point cloud to estimate constrained parts’ poses by predicting the joint parameters and states as replacements. For further refinement, we propose a differentiable rendering scheme for pose optimization. Evaluations of the ArtImage and RobotArm datasets demonstrate that CAPED exhibits outstanding effectiveness and generalization in tasks ranging from synthetic data to real-world scenarios. We will publicly release the code. Li Zhang 0104, Yukang Huo, Lin Wu 0001, Yanyan Wei, Harshal Suresh Shende, Liu Liu 0012, Linlin Ou |
ICASSP | 4 |
| 2025 | Diff-Art: Category-level Articulation Pose Estimation via Conditional DiffusionabstractArticulated objects are prevalent in people’s daily lives, yet their diverse motion structures pose significant challenges for category-level pose estimation. To address this, in this work, we introduce Diff-Art, a method tailored for category-level articulation pose estimation using conditional diffusion. Given a partial point cloud as input, Diff-Art predicts the per-part 6D pose of the articulated object. Our approach incorporates a novel modeling strategy that exploits the unique kinematic constraints of articulated objects, effectively handling self-occlusion scenarios. Furthermore, we propose a re-scoring strategy to enhance the accuracy of 6D pose estimation within the conditional diffusion framework. Extensive experiments validate the effectiveness of Diff-Art, demonstrating its strong performance on both synthetic datasets and its ability to generalize to real-world scenarios. We believe this work holds significant potential for applications in embodied AI and robotics. Yukang Huo, Xianhui Meng, Li Zhang 0104, Yan Zhong 0001, Mingyuan Yao |
ICME | 1 |
| 2025 | FA-YOLO: Research On Efficient Feature Selection YOLO Improved Algorithm Based On FMDS and AGMF ModulesabstractOver the past few years, the YOLO series of models has emerged as one of the dominant methodologies in the realm of object detection. Many studies have advanced these baseline models by modifying their architectures, enhancing data quality, and developing new loss functions. However, current models still exhibit deficiencies in processing feature maps, such as overlooking the fusion of cross-scale features and a static fusion approach that lacks the capability for dynamic feature adjustment. To address these issues, this paper introduces an efficient Fine-grained Multi-scale Dynamic Selection Module (FMDS Module), which applies a more effective dynamic feature selection and fusion method on fine-grained multi-scale feature maps, significantly enhancing the detection accuracy of small, medium, and large-sized targets. Furthermore, this paper proposes an Adaptive Gated Multi-branch Focus Fusion Module (AGMF Module), which utilizes multiple parallel branches to perform complementary fusion of various features captured by the gated unit branch, FMDS Module branch, and TripletAttention branch. This approach further enhances the comprehensiveness, diversity, and integrity of feature fusion. This paper has integrated the FMDS Module, AGMF Module, into Yolov9 to develop a novel object detection model named FA-YOLO. Extensive experimental results show that under identical experimental conditions, FA-YOLO achieves an outstanding 66.1% mean Average Precision (mAP) on the PASCAL VOC 2007 dataset, representing 1.0% improvement over YOLOv9’s 65.1%. Additionally, the detection accuracies of FA-YOLO for small, medium, and large targets are 44.1%, 54.6%, and 70.8%, respectively, showing improvements of 2.0%, 3.1%, and 0.9% compared to YOLOv9’s 42.1%, 51.5%, and 69.9%. Yukang Huo, Mingyuan Yao, Qingbin Tian, Shulong Zhang, Jiayin Zhao |
IJCNN | 1 |
| 2025 | MPM-GS: Optimizing Sparse-View 3D Scene Reconstruction with Virtual View Rendering and Multimodal Regularizationabstract3D Gaussian Splatting is widely used in 3D reconstruction and has applications in novel view synthesis and scene generation. Recent work has addressed this problem by leveraging multi-view data for high-quality 3D scene reconstruction. Unfortunately, these approaches suffer from overfitting with sparse-view data due to insufficient view information, resulting in artifacts and aliasing. In contrast, we propose a novel method, MPM-GS, which incorporates monocular depth estimation, virtual views, and regularization techniques to address these challenges and improve reconstruction quality under sparse-view conditions. This fixes the overfitting and artifact issues, however, it does not solve all the limitations of sparse-view data in extreme edge cases. Consequently, we develop a novel densification strategy to optimize Gaussian point distributions and improve scene accuracy. While promising, this densification process is non-trivial, as it requires balancing efficiency with rendering fidelity. Therefore, we further optimize Gaussian points and generate new, representative points to enhance both accuracy and computational efficiency. We evaluate MPM-GS both qualitatively and quantitatively on the Tanks & Temples and LLFF datasets, achieving excellent rendering quality and fast rendering speed, particularly in novel view synthesis. Mingyuan Yao, Shulong Zhang, Yukang Huo, Jiayin Zhao, Yingyi Chen |
IJCNN | 3 |
| 2025 | PR-DETR: Extracting and utilizing prior knowledge for improved end-to-end object detection
Yukang Huo, Mingyuan Yao, Tonghao Wang, Qingbin Tian, Jiayin Zhao |
Image Vis. Comput. | 1 |