VLDB 2026 Research / reviewers in the wild / expert
Qi Ye 0001
dblp:19/6124-1
· DBLP profile ↗
44ranked-venue papers
3as first author
39since 2021 · last 2026
0000-0003-2285-3402ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 3 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 2 first-author · 20 since 2021Systems, architecture and hardware · 12 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Corrigendum to "LDM: Large tensorial SDF model for textured mesh generation" [Graphical Models, Volume 140, August 2025, 101271]
Rengan Xie, Xiaoliang Luo, Lvchun Wang, Qi Wang 0111, Qi Ye 0001, Wei Chen 0001, Wenting Zheng, Yuchi Huo |
Graph. Model. | 7 |
| 2026 | Improving Local Feature Matching by Entropy-Inspired Scale Adaptability and Flow-Endowed Local ConsistencyabstractRecent semi-dense image matching methods have achieved remarkable success, but two long-standing issues still impair their performance. At the coarse stage, the over-exclusion issue of their mutual nearest neighbor (MNN) matching layer makes them struggle to handle cases with scale difference between images. To this end, we comprehensively revisit the matching mechanism and make a key observation that the hint concealed in the score matrix can be exploited to indicate the scale ratio. Based on this, we propose a scale-aware matching module which is exceptionally effective but introduces negligible overhead. At the fine stage, we point out that existing methods neglect the local consistency of final matches, which undermines their robustness. To this end, rather than independently predicting the correspondence for each source pixel, we reformulate the fine stage as a cascaded flow refinement problem and introduce a novel gradient loss to encourage local consistency of the flow field. Extensive experiments demonstrate that our novel matching pipeline, with these proposed modifications, achieves robust and accurate matching performance on downstream tasks. Jiming Chen 0001, Qi Ye 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Toward Reasoning-Centric Video Object Segmentation via Multi-Modal Large Language ModelsabstractReferring Video Object Segmentation (RVOS) aims to segment the target objects specified in human instructions. Previous approaches typically rely on explicit human instructions that contain target categories or salient appearance descriptions. These approaches tend to fail when the instructions require temporal video understanding and complex relational reasoning. In this work, we present RViSeg, a reasoning-centric video object segmentation model that leverages the reasoning capability of Multi-modal Large Language Models (MLLM) to handle complex queries. The primary challenge lies in enabling MLLM to perform efficient pixel-level video perception. To tackle this challenge, we introduce a novel Spatial Token Merge (STM) module that consolidates lengthy video tokens into compact region-level clusters, while preserving essential spatial details. This structured representation enables MLLM to infer user intention by interleaving spatial and temporal visual information. Furthermore, we propose a Query-based Target Retrieval (QTR) module that utilizes learnable tokens as the target identity for mask prediction. By propagating these instance-specific tokens both intra-clip and inter-clip, our RViSeg effectively encodes object motion, ensuring spatio-temporal consistency in segmentation results. To facilitate training and evaluation, we construct InstructVideo, a single- and multiple-object reasoning video segmentation benchmark. Comprehensive experiments demonstrate the effectiveness of the proposed components. Yanyan Shao, Shuting He, Gengze Zhou, Qi Ye 0001, Xiufang Shi, Jiming Chen 0001, Qi Wu 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | DexRepNet++: Learning Dexterous Robotic Manipulation With Geometric and Spatial Hand-Object RepresentationsabstractRobotic dexterous manipulation is a challenging problem due to high degrees of freedom (DoFs) and complex contacts of multi-fingered robotic hands. Many existing deep reinforcement learning (DRL) based methods aim at improving sample efficiency in high-dimensional output action spaces. However, existing works often overlook the role of representations in achieving generalization of a manipulation policy in the complex input space during the hand-object interaction. In this paper, we propose DexRep, a novel hand-object interaction representation to capture object surface features and spatial relations between hands and objects for dexterous manipulation skill learning. Based on DexRep, policies are learned for three dexterous manipulation tasks, i.e. grasping, in-hand reorientation, bimanual handover, and extensive experiments are conducted to verify the effectiveness. In simulation, for grasping, the policy learned with 40 objects achieves a success rate of 87.9% on more than 5000 unseen objects of diverse categories, significantly surpassing existing work trained with thousands of objects; for the in-hand reorientation and handover tasks, the policies also boost the success rates and other metrics of existing hand-object representations by 20% to 40%. The grasp policies with DexRep are deployed to the real world under multi-camera and single-camera setups and demonstrate a small sim-to-real gap. Qingtao Liu, Zhengnan Sun, Haoming Li 0004, Gaofeng Li, Lin Shao 0002, Jiming Chen 0001, Qi Ye 0001 |
IEEE Trans. Robotics | 8 |
| 2025 | Hand-held Object Reconstruction from RGB Video with Dynamic InteractionabstractThis work aims to reconstruct the 3D geometry of a rigid object manipulated by one or both hands using monocular RGB video. Previous methods rely on Structure-from-Motion or hand priors to estimate relative motion between the object and camera, which typically assume textured objects or single-hand interactions. To accurately recover object geometry in dynamic interactions, we incorporate priors from 3D generation model into object pose estimation and propose semantic consistency constraints to solve the challenge of shape and texture discrepancy between the generated priors and observations. The poses are initialized, followed by joint optimization of the object poses and implicit neural representation. During optimization, a novel pose outlier voting strategy with inter-view consistency is proposed to correct large pose errors. Experiments on three datasets demonstrate that our method significantly outperforms the state-of-the-art in reconstruction quality for both single- and two-hand scenarios. Our project page: https://east-j.github.io/dynhor/ Shijian Jiang, Qi Ye 0001, Rengan Xie, Yuchi Huo, Jiming Chen 0001 |
CVPR | 2 |
| 2025 | A3GS: Arbitrary Artistic Style into Arbitrary 3D Gaussian Splatting
Zhiyuan Fang, Rengan Xie, Xuancheng Jin, Qi Ye 0001, Wei Chen 0001, Wenting Zheng, Rui Wang 0004, Yuchi Huo |
ICCV | 4 |
| 2025 | IntrinsicControlNet: Cross-Distribution Image Generation with Real and Unreal
Jiayuan Lu, Rengan Xie, Zhizhen Wu, Dianbing Xi, Qi Ye 0001, Rui Wang 0004, Hujun Bao, Yuchi Huo |
ICCV | 6 |
| 2025 | VTDexManip: A Dataset and Benchmark for Visual-tactile Pretraining and Dexterous Manipulation with Reinforcement LearningabstractVision and touch are the most commonly used senses in human manipulation. While leveraging human manipulation videos for robotic task pretraining has shown promise in prior works, it is limited to image and language modalities and deployment to simple parallel grippers. In this paper, aiming to address the limitations, we collect a vision-tactile dataset by humans manipulating 10 daily tasks and 182 objects. In contrast with the existing datasets, our dataset is the first visual-tactile dataset for complex robotic manipulation skill learning. Also, we introduce a novel benchmark, featuring six complex dexterous manipulation tasks and a reinforcement learning-based vision-tactile skill learning framework. 18 non-pretraining and pretraining methods within the framework are designed and compared to investigate the effectiveness of different modalities and pertaining strategies. Key findings based on our benchmark results and analyses experiments include: 1) Despite the tactile modality used in our experiments being binary and sparse, including it directly in the policy training boosts the success rate by about 20\% and joint pretraining it with vision gains a further 20\%. 2) Joint pretraining visual-tactile modalities exhibits strong adaptability in unknown tasks and achieves robust performance among all tasks. 3) Using binary tactile signals with vision is robust to viewpoint setting, tactile noise, and the binarization threshold, which facilitates to the visual-tactile policy to be deployed in reality. The dataset and benchmark are available at \url{https://github.com/LQTS/VTDexManip}. Qingtao Liu, Zhengnan Sun, Gaofeng Li, Jiming Chen 0001, Qi Ye 0001 |
ICLR | 6 |
| 2025 | UpViTaL: Unpaired Visual-Tactile Self-Supervised Representation Learning for Dexterous Robotic ManipulationabstractVisual and tactile pretraining have been extensively studied in dexterous robot manipulation tasks. However, existing methods typically require the simultaneous acquisition of visual and tactile data, making it difficult to utilize low-cost, unpaired visual-tactile datasets. Moreover, these methods often rely on tactile sensors to provide input data for reinforcement learning (RL) during the physical deployment of robotic dexterous hands, which highly increases deployment costs. To address these challenges, we propose UpViTaL, an unpaired visualtactile self-supervised representation learning method for RLbased robot dexterous manipulation. Specifically, we collect low-cost unpaired visual and tactile datasets for manipulation skill learning using a camera and tactile gloves on three robot manipulation tasks. The temporal tactile self-supervised representation learning module of UpViTaL is used to explore efficient tactile representations from time-series tactile data. In parallel, the visual pretraining module of UpViTaL helps to extract efficient visual representations from visual data. In addition, we fuse unpaired visual-tactile representations through an RL reward mechanism, which does not require robotic dexterous hands tactile sensors for practical deployment. We validate our approach on three dexterous robot manipulation tasks. Experimental results demonstrate that UpViTaL can efficiently learn robot manipulation skills. Compared to existing approaches for visual pretraining, our method significantly improves the success rate by more than 30%. Guwen Han, Qingtao Liu, Anjun Chen, Jiming Chen 0001, Qi Ye 0001 |
ICRA | 6 |
| 2025 | Temporal-Spatial Representation Fusion for Dexterous Manipulation Learning with Unpaired Visual-Action DataabstractSupervised behavioral cloning using robot visual-action data has been widely investigated in robot manipulation. However, these methods typically require simultaneous acquisition of visual and action data, which makes them difficult to utilize unpaired visual-action datasets: e.g. videos on Internet or action only data which has less privacy and security concerns. To take advantage of the action data without synchronized visual observation, we propose UnVALe, a novel dexterous robotic manipulation RL framework that utilizes action data without paired images to learn priors of human dexterous manipulation skills. Specifically, an LSTM-based network is designed to learn the temporal action prior by reconstructing the input trajectories, and a VAE network is designed to learn the spatial action prior by reconstructing the input action. Novel rewards are proposed to incorporate the priors into reinforcement learning, which encourages action output from RL polices to maintain low reconstruction errors in the LSTM and VAE networks. We perform extensive validation on three dexterous robot manipulation tasks. The experimental results show that UnVALe can effectively improve robot manipulation performance. Compared with existing visual pretraining methods, our method achieves a more than 30% increase in success rates. Guwen Han, Zhengnan Sun, Qingtao Liu, Anjun Chen, Huajin Chen, Rong Xiong, Jiming Chen 0001, Qi Ye 0001 |
IROS | 9 |
| 2025 | VTAO-BiManip: Masked Visual-Tactile-Action Pre-training with Object Understanding for Bimanual Dexterous ManipulationabstractBimanual dexterous manipulation remains a significant challenge in robotics due to the high DoFs of each hand and their coordination. Existing single-hand manipulation techniques often leverage human demonstrations to guide RL methods but fail to generalize to complex bimanual tasks involving multiple sub-skills. In this paper, we propose VTAO-BiManip, a novel framework that integrates visual-tactile-action pre-training with object understanding, aiming to enable human-like bimanual manipulation via curriculum reinforcement learning (RL). We improve prior learning by incorporating hand motion data, providing more effective guidance for dual-hand coordination. Our pretraining model predicts future actions as well as object pose and size using masked multimodal inputs, facilitating cross-modal regularization. To address the multi-skill learning challenge, we introduce a two-stage curriculum RL approach to stabilize training. We evaluate our method on a bimanual bottle-cap twisting task, demonstrating its effectiveness in both simulated and real-world environments. Our approach achieves a success rate that surpasses existing visual-tactile pretraining methods by over 20%. Zhengnan Sun, Zhaotai Shi, Qingtao Liu, Jiming Chen 0001, Qi Ye 0001 |
IROS | 7 |
| 2025 | RFMPose: Generative Category-level Object Pose Estimation via Riemannian Flow MatchingabstractWe introduce RFMPose, a novel generative framework for category-level 6D object pose estimation that learns deterministic pose trajectories through Riemannian Flow Matching (RFM). Existing discriminative approaches struggle with multi-hypothesis predictions (e.g., symmetry ambiguities) and often require specialized network architectures. RFMPose advances this paradigm through three key innovations:
(1) Ensuring geometric consistency via geodesic interpolation on Riemannian manifolds combined with bi-invariant metric constraints;
(2) Alleviating symmetry-induced ambiguities through Riemannian Optimal Transport for probability mass redistribution without ad-hoc design;
(3) Enabling end-to-end likelihood estimation through Hutchinson trace approximation, thereby eliminating auxiliary model dependencies.
Extensive experiments on the Omni6DPose demonstrate state-of-the-art performance of the proposed method, with significant improvements of $\textbf{+4.1}$ in $\mathrm{\textbf{IoU}_{25}}$ and $\textbf{+2.4}$ in $\textbf{5°2cm}$ metrics compared to prior generative approaches. Furthermore, the proposed RFM framework exhibits robust sim-to-real transfer capabilities and facilitates pose tracking extensions with minimal architectural adaptation. Wenzhe Ouyang, Qi Ye 0001, Zenglin Xu, Jiming Chen 0001 |
NeurIPS | 2 |
| 2025 | LDM: Large tensorial SDF model for textured mesh generationabstractPrevious efforts have managed to generate production-ready 3D assets from text or images. However, these methods primarily employ NeRF or 3D Gaussian representations, which are not adept at producing smooth, high-quality geometries required by modern rendering pipelines. In this paper, we propose LDM, a L arge tensorial S D F M odel, which introduces a novel feed-forward framework capable of generating high-fidelity, illumination-decoupled textured mesh from a single image or text prompts. We firstly utilize a multi-view diffusion model to generate sparse multi-view inputs from single images or text prompts, and then a transformer-based model is trained to predict a tensorial SDF field from these sparse multi-view image inputs. Finally, we employ a gradient-based mesh optimization layer to refine this model, enabling it to produce an SDF field from which high-quality textured meshes can be extracted. Extensive experiments demonstrate that our method can generate diverse, high-quality 3D mesh assets with corresponding decomposed RGB textures within seconds. The project code is available at https://github.com/rgxie/LDM . Rengan Xie, Xiaoliang Luo, Lvchun Wang, Qi Wang 0111, Qi Ye 0001, Wei Chen 0001, Wenting Zheng, Yuchi Huo |
Graph. Model. | 7 |
| 2025 | APPTracker+: Displacement Uncertainty for Occlusion Handling in Low-Frame-Rate Multiple Object Tracking
Qi Ye 0001, Wenhan Luo, Haizhou Ran, Zhiguo Shi 0001, Jiming Chen 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | Diffusion-Based Completion for Multirobot Active Scene Reconstruction Toward IoT ApplicationsabstractAutonomous reconstruction of unknown scenes using multiple robots acting as mobile Internet of Things (IoT) nodes becomes a fundamental capability for extensive IoT applications, such as environmental monitoring, and search and rescue. However, existing multi-robot autonomous reconstruction approaches still suffer from incomplete observations, redundant generated coverage tasks, and overly distant assigned tasks. To this end, we first utilize a diffusion-based model for object completion in autonomously reconstructed scenes using 3D Gaussian Splatting, and obtain Gaussian mixture model-based object uncertainties to guide the robots in generating more accurate scanning tasks that fill holes and enhance reconstruction quality. We then design an efficient task filtering mechanism that utilizes clustering frontiers and exploration tasks, as well as instances and reconstruction tasks, enabling the elimination of redundant coverage tasks. We also devise a task reassignment mechanism for robots based on the required travel costs to avoid unnecessary detours, which further improves scanning efficiency. Extensive experimental results show that our method exhibits higher reconstruction quality and superior planning efficiency compared to existing multi-robot autonomous reconstruction methods. Yang Xu 0042, Qi Ye 0001, Jiming Chen 0001 |
IEEE Internet Things J. | 4 |
| 2025 | Toward Weather-Robust 3D Human Body Reconstruction: Millimeter-Wave Radar-Based Dataset, Benchmark, and Multi-Modal Fusionabstract3D human reconstruction from RGB images achieves decent results in good weather conditions but degrades dramatically in rough weather. Complementarily, mmWave radars have been employed to reconstruct 3D human joints and meshes in rough weather. However, combining RGB and mmWave signals for weather-robust 3D human reconstruction is still an open challenge, given the sparse nature of mmWave and the vulnerability of RGB images. The limited research about the impact of missing points and sparsity features of mmWave data on reconstruction performance, as well as the lack of available datasets for paired mmWave-RGB data, further complicates the process of fusing the two modalities. To fill these gaps, we build up an automatic 3D body annotation system with multiple sensors to collect a large-scale mmWave dataset. The dataset consists of synchronized and calibrated mmWave radar point clouds and RGB(D) images under different weather conditions and skeleton/mesh annotations for humans in these scenes. With this dataset, we conduct a comprehensive analysis about the limitations of single-modality reconstruction and the impact of missing points and sparsity on the reconstruction performance. Based on the guidance of this analysis, we design ImmFusion, the first mmWave-RGB fusion solution to robustly reconstruct 3D human bodies in various weather conditions. Specifically, our ImmFusion consists of image and point backbones for token feature extraction and a Transformer module for token fusion. The image and point backbones refine global and local features from original data, and the Fusion Transformer Module aims for effective information fusion of two modalities by dynamically selecting informative tokens. Extensive experiments demonstrate that ImmFusion can efficiently utilize the information of two modalities to achieve robust 3D human body reconstruction in various weather environments. In addition, our method achieves superior accuracy compared to that of the state-of-the-art Transformer-based LiDAR-camera fusion methods. Anjun Chen, Kun Shi 0003, Yuchi Huo, Jiming Chen 0001, Qi Ye 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | CT-NeRF: Incremental Optimization of Neural Radiance Field and Camera Poses With Complex TrajectoryabstractNeural radiance field (NeRF) has achieved impressive results in high-quality 3D scene reconstruction. However, NeRF heavily relies on precise camera poses. While recent works like BARF have introduced camera pose optimization within NeRF, their applicability is limited to simple trajectory scenes. Existing methods struggle while tackling complex trajectories involving large rotations. To address this limitation, we propose CT-NeRF, an incremental reconstruction and optimization pipeline using only RGB images without pose and depth input. In this pipeline, we first propose a local-global bundle adjustment under a pose graph connecting neighboring frames to enforce the consistency between poses to escape the local minima caused by only pose consistency with the scene structure. Further, we instantiate the consistency between poses as a reprojection error constraint resulting from pixel-level correspondences between input image pairs. Through the incremental reconstruction, CT-NeRF enables the recovery of both camera poses and scene structure and is capable of handling scenes with complex trajectories. We evaluate the performance of CT-NeRF on two real-world datasets, NeRF-Buster and Free-Dataset, which feature complex trajectories. Results show CT-NeRF outperforms existing methods in novel view synthesis and pose estimation accuracy. Yunlong Ran, Yanxu Li, Qi Ye 0001, Yuchi Huo, Zhaopeng Cui, Zechun Bai, Jiming Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | AdaptiveFusion: Adaptive Multi-Modal Multi-View Fusion for 3D Human Body ReconstructionabstractRecent advancements in sensor technology and deep learning have led to significant progress in 3D human body reconstruction. However, most existing approaches rely on data from a specific sensor, which can be unreliable due to the inherent limitations of individual sensing modalities. Additionally, existing multi-modal fusion methods generally require customized designs based on the specific sensor combinations or setups, which limits the flexibility and generality of these methods. Furthermore, conventional point-image projection-based and Transformer-based fusion networks are susceptible to the influence of noisy modalities and sensor poses. To address these limitations and achieve robust 3D human body reconstruction in various conditions, we propose AdaptiveFusion, a generic adaptive multi-modal multi-view fusion framework that can effectively incorporate arbitrary combinations of uncalibrated sensor inputs. By treating different modalities from various viewpoints as equal tokens, and our handcrafted modality sampling module by leveraging the inherent flexibility of Transformer models, AdaptiveFusion is able to cope with arbitrary numbers of inputs and accommodate noisy modalities with only a single training network. Extensive experiments on large-scale human datasets demonstrate the effectiveness of AdaptiveFusion in achieving high-quality 3D human body reconstruction in various environments. In addition, our method achieves superior accuracy compared to state-of-the-art fusion methods. Anjun Chen, Kun Shi 0003, Yuchi Huo, Jiming Chen 0001, Qi Ye 0001 |
IEEE Trans. Multim. | 8 |
| 2024 | In-Hand 3D Object Reconstruction from a Monocular RGB VideoabstractOur work aims to reconstruct a 3D object that is held and rotated by a hand in front of a static RGB camera. Previous methods that use implicit neural representations to recover the geometry of a generic hand-held object from multi-view images achieved compelling results in the visible part of the object. However, these methods falter in accurately capturing the shape within the hand-object contact region due to occlusion. In this paper, we propose a novel method that deals with surface reconstruction under occlusion by incorporating priors of 2D occlusion elucidation and physical contact constraints. For the former, we introduce an object amodal completion network to infer the 2D complete mask of objects under occlusion. To ensure the accuracy and view consistency of the predicted 2D amodal masks, we devise a joint optimization method for both amodal mask refinement and 3D reconstruction. For the latter, we impose penetration and attraction constraints on the local geometry in contact regions. We evaluate our approach on HO3D and HOD datasets and demonstrate that it outperforms the state-of-the-art methods in terms of reconstruction surface quality, with an improvement of 52% on HO3D and 20% on HOD. Project webpage: https://east-j.github.io/ihor. Shijian Jiang, Qi Ye 0001, Rengan Xie, Yuchi Huo, Jiming Chen 0001 |
AAAI | 2 |
| 2024 | Context-Aware Integration of Language and Visual References for Natural Language TrackingabstractTracking by natural language specification (TNL) aims to consistently localize a target in a video sequence given a linguistic description in the initial frame. Existing methodologies perform language-based and template-based matching for target reasoning separately and merge the matching results from two sources, which suffer from tracking drift when language and visual templates missalign with the dynamic target state and ambiguity in the later merging stage. To tackle the issues, we propose a joint multi-modal tracking framework with 1) a prompt modulation module to leverage the complementarity between temporal visual templates and language expressions, enabling precise and context-aware appearance and linguistic cues, and 2) a unified target decoding module to integrate the multi-modal reference cues and executes the integrated queries on the search image to predict the target location in an end-to-end manner directly. This design ensures spatio-temporal consistency by leveraging historical visual information and introduces an integrated solution, generating predictions in a single step. Extensive experiments conducted on TNL2K, OTB-Lang, LaSOT, and RefCOCOg validate the efficacy of our proposed approach. The results demonstrate competitive performance against state-of-the-art methods for both tracking and grounding. Code is available at https://github.com/twotw02/QueryNLT Yanyan Shao, Shuting He, Qi Ye 0001, Yuchao Feng, Wenhan Luo, Jiming Chen 0001 |
CVPR | 3 |
| 2024 | TPGP: Temporal-Parametric Optimization with Deep Grasp Prior for Dexterous Motion PlanningabstractGrasping motion planning aims to find a feasible grasping trajectory in the configuration space given an input target grasp. While optimizing grasp motion with two or three-fingered grippers has been well studied, the study on natural grasp motion planning with a dexterous hand remains a very challenging problem due to the high dimensional working space. In this work, we propose a novel temporal-parametric grasp prior (TPGP) optimization method to simplify the difficulty of grasping trajectory optimization for the dexterous hand while maintaining smooth and natural properties of the grasping motion. Specifically, we formulate the discrete trajectory parameters into a temporal-based parameterization, where the prior constraint provided by a hand poser network, is introduced to ensure that hand pose is natural and reasonable throughout the trajectory. Finally, we present a joint target optimization strategy to enhance the target pose for more feasible trajectories. Extensive validations on two public datasets show that our method outperforms state-of-the-art methods regarding grasp motion on various metrics. Haoming Li 0004, Qi Ye 0001, Yuchi Huo, Qingtao Liu, Shijian Jiang, Jiming Chen 0001 |
ICRA | 2 |
| 2024 | InterRep: A Visual Interaction Representation for Robotic GraspingabstractRecently, pre-trained vision models have gained significant attention in motor control, showcasing impressive performance across diverse robotic learning tasks. While previous works predominantly concentrate on the significance of the pre-training phase, the equally important task of extracting more effective representations based on existing pre-trained visual models remains unexplored. To better leverage the representation capabilities of pre-trained models for robotic grasping, we propose InterRep, a novel interaction representation method that possesses not only the strengths of pre-trained models, known for their robustness in noisy environments and their proficiency in recognizing essential features, but also the capacity of capturing dynamic interaction details and local geometric features during the grasping process. Based on the novel representation, we introduce a deep reinforcement learning method to learn generalizable grasping policies. The experimental results demonstrate that our proposed representation outperforms the baselines in terms of both training speed and generalization. For the generalized grasping tasks with dexterous robotic hands, our method boasts a success rate nearly 20% higher than methods using the global features of the entire image from pre-trained models. In addition, our proposed representation method demonstrates promising performance when applied to a different robotic hand and task. It also exhibits excellent performance on real robots with a success rate of 70%. Qi Ye 0001, Qingtao Liu, Anjun Chen, Gaofeng Li, Jiming Chen 0001 |
ICRA | 2 |
| 2024 | CAMInterHand: Cooperative Attention for Multi-View Interactive Hand Pose and Mesh ReconstructionabstractInteractive hand mesh reconstruction from singleview images poses a significant challenge with the severe occlusion and depth ambiguity inherent in interactive hand gestures. Recent approaches that employ probabilistic models and tokenpruned techniques have shown decent results in multi-view human body reconstruction. Nevertheless, these methods have not fully utilized multi-scale semantic information from multiview images and are not applicable in scenarios involving severe occlusion during dual-hand interactions. Simultaneously, current single-view methods independently reconstruct the left and right hands, which are ineffective in enhancing the interaction between both hands. To address these challenges, we propose CAMInterHand, a cooperative attention-based method for multi-view interactive hand pose and mesh reconstruction. Specifically, CAMInterHand extracts local pyramid features and global vertex features from multi-scale feature maps of multi-view images, enabling the exploration of rich local semantic information and facilitating effective feature alignment. Furthermore, CAMInterHand employs the cooperative attention fusion module to fuse all features from multi-view images, enhancing interactions among vertices of dual hands within global and local contexts. We conduct extensive experiments on the large-scale multi-view dataset InterHand2.6M and CAMInterHand achieves a substantial performance improvement over existing methods for multi-view and single-view interactive hand reconstruction. Guwen Han, Qi Ye 0001, Anjun Chen, Jiming Chen 0001 |
ICRA | 2 |
| 2024 | Masked Visual-Tactile Pre-training for Robot ManipulationabstractRecent works on the pretraining for robot manipulation have demonstrated that representations learning from large human manipulation data can generalize well to new manipulation tasks and environments. However, these approaches mainly focus on human vision or natural language, neglecting tactile feedback. In this article, we make an attempt to explore how to pre-train a representation model for robotic manipulation using both human manipulation visual and tactile data. We develop a system for collecting visual and tactile data, featuring a cost-effective tactile glove to capture human tactile data and Hololens2 for capturing visual data. With this system, we collect a dataset of turning bottle caps. Furthermore, we introduce a novel visual-tactile fusion network and learning strategy M2VTP, with one key module to tokenize 20 sparse binary tactile signals sensing touch states for the learning of tactile context and the other key module applying the attention and mask mechanism to the interaction of visual and tactile tokens for visual-tactile representation learning. We utilize our dataset to pre-train the fusion model and embed the pre-trained model into a reinforcement learning framework for downstream tasks. Experimental results demonstrate that our pre-trained model significantly aids in learning manipulation skills. Compared to methods without pre-training, our approach achieves a success rate increase of over 60%. Additionally, when compared to current visual pre-training methods, our success rate exceeds them by more than 50%. Qingtao Liu, Qi Ye 0001, Zhengnan Sun, Gaofeng Li, Jiming Chen 0001 |
ICRA | 2 |
| 2024 | Autonomous Implicit Indoor Scene Reconstruction with Frontier ExplorationabstractImplicit neural representations have demonstrated significant promise for 3D scene reconstruction. Recent works have extended their applications to autonomous implicit reconstruction through the Next Best View (NBV) based method. However, the NBV method cannot guarantee complete scene coverage and often necessitates extensive viewpoint sampling, particularly in complex scenes. In the paper, we propose to 1) incorporate frontier-based exploration tasks for global coverage with implicit surface uncertainty-based reconstruction tasks to achieve high-quality reconstruction. and 2) introduce a method to achieve implicit surface uncertainty using color uncertainty, which reduces the time needed for view selection. Further with these two tasks, we propose an adaptive strategy for switching modes in view path planning, to reduce time and maintain superior reconstruction quality. Our method exhibits the highest reconstruction quality among all planning methods and superior planning efficiency in methods involving reconstruction tasks. We deploy our method on a UAV and the results show that our method can plan multi-task views and reconstruct a scene with high quality. Yanxu Li, Qi Ye 0001, Yunlong Ran, Jiming Chen 0001 |
ICRA | 4 |
| 2024 | Error-aware Sampling in Adaptive Shells for Neural Surface Reconstruction
Qi Wang 0111, Yuchi Huo, Qi Ye 0001, Rui Wang 0004, Hujun Bao |
IJCAI | 3 |
| 2024 | A Light-weight and Rapid Table Tennis Ball Trajectory Prediction Approaches towards Online Bouncing TaskabstractIt is essentially required to predict the ball’s flight trajectory accurately and timely for a robotic table tennis ball bouncing task. Existing solutions, which can be categorized into model-based and learning-based groups, both exhibits unpleasant disadvantages. For example, they often require to identify many dynamic parameters accurately or to collect extensive labeled data, which are generally very difficult or costly to achieve in real world. In this paper, we proposed a light-wight and rapid trajectory prediction approach for online table tennis bouncing tasks based on a simplified model. In the proposed approach, the ball’s flight poses are captured and estimated by a low-cost RGB-D camera. Then the ball’s landing position is predicted in advance by using a fitted 3D parabola. Compared with existing solutions, our proposed approach is lightweight and easy to deploy. In experiments, 66 flight trajectories of the ball are collected to serve as benchmark. The prediction errors for all landing positions are all less than 20mm, in which most of them are less than 10mm. In addition, the prediction can be achieved 141.7ms in advance, which is fast enough for the robotic arm to plan and move itself to the predicted landing point. Peisen Xu, Gaofeng Li, Qi Ye 0001, Jiming Chen 0001 |
RO-MAN | 3 |
| 2024 | Metaverse for the Energy Industry: Technologies, Applications, and SolutionsabstractThe Metaverse refers to the integration of physical and virtual realities, offering new possibilities for enhancing operations and services across various industries. However, its application in the energy sector is still in its nascent stage. The energy industry, crucial for the global economy and society, faces significant challenges due to its complex and risky nature, such as health, safety, and environmental (HSE) concerns, and the remote locations of extraction sites. Although some studies have explored the use of the Metaverse in this industry for data visualization, energy process modeling, and training, a comprehensive review of existing technologies and applications is lacking. This article addresses this gap by examining the potential of the Metaverse for the energy industry, using the Oil&Gas sector as a case study. We identify the essential technologies needed to create a realistic and immersive Metaverse experience and review the current literature on its industrial applications in the energy sector. Distinguishing our work from others, we present four practical case studies developed from our own experience in the Oil&Gas sector. These cases demonstrate the tangible benefits of Metaverse solutions in improving operational efficiency, reducing costs, and enhancing worker safety and productivity, highlighting the unique value of the Metaverse in addressing industry-specific challenges. Qi Ye 0001, Yunlong Ran, Jiaqi Zhan, Jiming Chen 0001, Youxian Sun |
IEEE Trans. Cybern. | 1 |
| 2023 | I2-SDF: Intrinsic Indoor Scene Reconstruction and Editing via Raytracing in Neural SDFsabstractIn this work, we present I2-SDF, a new method for intrinsic indoor scene reconstruction and editing using differentiable Monte Carlo raytracing on neural signed distance fields (SDFs). Our holistic neural SDF-based frame-work jointly recovers the underlying shapes, incident radiance and materials from multi-view images. We introduce a novel bubble loss for fine-grained small objects and error-guided adaptive sampling scheme to largely improve the reconstruction quality on large-scale indoor scenes. Further, we propose to decompose the neural radiance field into spatially-varying material of the scene as a neural field through surface-based, differentiable Monte Carlo raytracing and emitter semantic segmentations, which enables physically based and photorealistic scene relighting and editing applications. Through a number of qualitative and quantitative experiments, we demonstrate the superior quality of our method on indoor scene reconstruction, novel view synthesis, and scene editing compared to state-of-the-art baselines. Our project page is at https://jingsenzhu.github.io/i2-sdf. Jingsen Zhu, Yuchi Huo, Qi Ye 0001, Fujun Luan, Jifan Li, Dianbing Xi, Lisha Wang, Rui Tang 0015, Wei Hua 0002, Hujun Bao, Rui Wang 0004 |
CVPR | 3 |
| 2023 | Seal-3D: Interactive Pixel-Level Editing for Neural Radiance FieldsabstractWith the popularity of implicit neural representations, or neural radiance fields (NeRF), there is a pressing need for editing methods to interact with the implicit 3D models for tasks like post-processing reconstructed scenes and 3D content creation. While previous works have explored NeRF editing from various perspectives, they are restricted in editing flexibility, quality, and speed, failing to offer direct editing response and instant preview. The key challenge is to conceive a locally editable neural representation that can directly reflect the editing instructions and update instantly. To bridge the gap, we propose a new interactive editing method and system for implicit representations, called Seal-3D1, which allows users to edit NeRF models in a pixel-level and free manner with a wide range of NeRF-like backbone and preview the editing effects instantly. To achieve the effects, the challenges are addressed by our proposed proxy function mapping the editing instructions to the original space of NeRF models in the teacher model and a two-stage training strategy for the student model with local pretraining and global finetuning. A NeRF editing system is built to showcase various editing types. Our system can achieve compelling editing effects with an interactive speed of about 1 second. Jingsen Zhu, Qi Ye 0001, Yuchi Huo, Yunlong Ran, Jiming Chen 0001 |
ICCV | 3 |
| 2023 | F&F Attack: Adversarial Attack against Multiple Object Trackers by Inducing False Negatives and False PositivesabstractMulti-object tracking (MOT) aims to build moving trajectories for number-agnostic objects. Modern multi-object trackers commonly follow the tracking-by-detection strategy. Therefore, fooling detectors can be an effective solution but it usually requires attacks in multiple successive frames, resulting in low efficiency. Attacking association processes improves efficiency but may require model-specific design, leading to poor generalization. In this paper, we propose a novel False negative and False positive attack (F&F attack) mechanism: it perturbs the input image to erase original detections and to inject deceptive false alarms around original ones while integrating the association attack implicitly. The mechanism can produce effective identity switches against multi-object trackers by only fooling detectors in a few frames. To demonstrate the flexibility of the mechanism, we deploy it to three multi-object trackers (ByteTrack, SORT, and CenterTrack) which are enabled by two representative detectors (YOLOX and CenterNet). Comprehensive experiments on MOT17 and MOT20 datasets show that our method significantly outperforms existing attackers, revealing the vulnerability of the tracking-by-detection paradigm to detection attacks. Qi Ye 0001, Wenhan Luo, Kaihao Zhang, Zhiguo Shi 0001, Jiming Chen 0001 |
ICCV | 2 |
| 2023 | ImmFusion: Robust mmWave-RGB Fusion for 3D Human Body Reconstruction in All Weather Conditionsabstract3D human reconstruction from RGB images achieves decent results in good weather conditions but degrades dramatically in rough weather. Complementary, mmWave radars have been employed to reconstruct 3D human joints and meshes in rough weather. However, combining RGB and mmWave signals for robust all-weather 3D human reconstruction is still an open challenge, given the sparse nature of mmWave and the vulnerability of RGB images. In this paper, we present ImmFusion, the first mmWave-RGB fusion solution to reconstruct 3D human bodies in all weather conditions robustly. Specifically, our ImmFusion consists of image and point backbones for token feature extraction and a Transformer module for token fusion. The image and point backbones refine global and local features from original data, and the Fusion Transformer Module aims for effective information fusion of two modalities by dynamically selecting informative tokens. Extensive experiments on a large-scale dataset, mmBody, captured in various environments demonstrate that ImmFusion can efficiently utilize the information of two modalities to achieve a robust 3D human body reconstruction in all weather conditions. In addition, our method's accuracy is significantly superior to that of state-of-the-art Transformer-based LiDAR-camera fusion methods. Anjun Chen, Kun Shi 0003, Shaohao Zhu, Jiming Chen 0001, Yuchi Huo, Qi Ye 0001 |
ICRA | 9 |
| 2023 | Efficient View Path Planning for Autonomous Implicit ReconstructionabstractImplicit neural representations have shown promising potential for 3D scene reconstruction. Recent work applies it to autonomous 3D reconstruction by learning information gain for view path planning. Effective as it is, the computation of the information gain is expensive, and compared with that using volumetric representations, collision checking using the implicit representation for a 3D point is much slower. In the paper, we propose to 1) leverage a neural network as an implicit function approximator for the information gain field and 2) combine the implicit fine-grained representation with coarse volumetric representations to improve efficiency. Further with the improved efficiency, we propose a novel informative path planning based on a graph-based planner. Our method demonstrates significant improvements in the reconstruction quality and planning efficiency compared with autonomous reconstructions with implicit and explicit representations. We deploy the method on a real UAV and the results show that our method can plan informative views and reconstruct a scene with high quality. Yanxu Li, Yunlong Ran, Lincheng Li, Shibo He, Jiming Chen 0001, Qi Ye 0001 |
ICRA | 9 |
| 2023 | Contact2Grasp: 3D Grasp Synthesis via Hand-Object Contact Constraintabstract3D grasp synthesis generates grasping poses given an input object. Existing works tackle the problem by learning a direct mapping from objects to the distributions of grasping poses. However, because the physical contact is sensitive to small changes in pose, the high-nonlinear mapping between 3D object representation to valid poses is considerably non-smooth, leading to poor generation efficiency and restricted generality. To tackle the challenge, we introduce an intermediate variable for grasp contact areas to constrain the grasp generation; in other words, we factorize the mapping into two sequential stages by assuming that grasping poses are fully constrained given contact maps: 1) we first learn contact map distributions to generate the potential contact maps for grasps; 2) then learn a mapping from the contact maps to the grasping poses. Further, we propose a penetration-aware optimization with the generated contacts as a consistency constraint for grasp refinement. Extensive validations on two public datasets show that our method outperforms state-of-the-art methods regarding grasp generation on various metrics. Haoming Li 0004, Xinzhuo Lin, Yuchi Huo, Jiming Chen 0001, Qi Ye 0001 |
IJCAI | 7 |
| 2023 | DexRepNet: Learning Dexterous Robotic Grasping Network with Geometric and Spatial Hand-Object RepresentationsabstractRobotic dexterous grasping is a challenging problem due to the high degree of freedom (DoF) and complex contacts of multi-fingered robotic hands. Existing deep re-inforcement learning (DRL) based methods leverage human demonstrations to reduce sample complexity due to the high dimensional action space with dexterous grasping. However, less attention has been paid to hand-object interaction representations for high-level generalization. In this paper, we propose a novel geometric and spatial hand-object interaction representation, named DexRep, to capture object surface features and the spatial relations between hands and objects during grasping. DexRep comprises Occupancy Feature for rough shapes within sensing range by moving hands, Surface Feature for changing hand-object surface distances, and LocalGeo Feature for local geometric surface features most related to potential contacts. Based on the new representation, we propose a dexterous deep reinforcement learning method DexRepNet to learn a generalizable grasping policy. Experimental results show that our method outperforms baselines using existing representations for robotic grasping dramatically both in grasp success rate and convergence speed. It achieves a 93% grasping success rate on seen objects and higher than 80% grasping success rates on diverse objects of unseen categories in both simulation and real-world experiments. Qingtao Liu, Qi Ye 0001, Zhengnan Sun, Haoming Li 0004, Gaofeng Li, Lin Shao 0002, Jiming Chen 0001 |
IROS | 3 |
| 2023 | InterTracker: Discovering and Tracking General Objects Interacting with Hands in the WildabstractUnderstanding human interaction with objects is an important research topic for embodied Artificial Intelligence and identifying the objects that humans are interacting with is a primary problem for interaction understanding. Existing methods rely on frame-based detectors to locate interacting objects. However, this approach is subjected to heavy occlusions, background clutter, and distracting objects. To address the limitations, in this paper, we propose to leverage spatio-temporal information of hand-object interaction to track interactive objects under these challenging cases. Without prior knowledge of the general objects to be tracked like object tracking problems, we first utilize the spatial relation between hands and objects to adaptively discover the interacting objects from the scene. Second, the consistency and continuity of the appearance of objects between successive frames are exploited to track the objects. With this tracking formulation, our method also benefits from training on large-scale general object-tracking datasets. We further curate a video-level hand-object interaction dataset for testing and evaluation from 100DOH. The quantitative results demonstrate that our proposed method outperforms the state-of-the-art methods. Specifically, in scenes with continuous interaction with different objects, we achieve an impressive improvement of about 10% as evaluated using the Average Precision (AP) metric. Our qualitative findings also illustrate that our method can produce more continuous trajectories for interacting objects. Yanyan Shao, Qi Ye 0001, Wenhan Luo, Kaihao Zhang, Jiming Chen 0001 |
IROS | 2 |
| 2022 | mmBody Benchmark: 3D Body Reconstruction Dataset and Analysis for Millimeter Wave RadarabstractMillimeter Ware (mmWave) Radar is gaining popularity as it can work in adverse environments like smoke, rain, snow, poor lighting, etc. Prior work has explored the possibility of reconstructing 3D skeletons or meshes from the noisy and sparse mmWare Radar signals. However, it is unclear how accurately we can reconstruct the 3D body from the mmWave signals across scenes and how it performs compared with cameras, which are important aspects needed to be considered when either using mmWave radars alone or combining them with cameras. To answer these questions, an automatic 3D body annotation system is first designed and built up with multiple sensors to collect a large-scale dataset. The dataset consists of synchronized and calibrated mmWave radar point clouds and RGB(D) images in different scenes and skeleton/mesh annotations for humans in the scenes. With this dataset, we train state-of-the-art methods with inputs from different sensors and test them in various scenarios. The results demonstrate that 1) despite the noise and sparsity of the generated point clouds, the mmWave radar can achieve better reconstruction accuracy than the RGB camera but worse than the depth camera; 2) the reconstruction from the mmWave radar is affected by adverse weather conditions moderately while the RGB(D) camera is severely affected. Further, analysis of the dataset and the results shadow insights on improving the reconstruction from the mmWave radar and the combination of signals from different sensors. Anjun Chen, Shaohao Zhu, Yanxu Li, Jiming Chen 0001, Qi Ye 0001 |
ACM Multimedia | 6 |
| 2022 | APPTracker: Improving Tracking Multiple Objects in Low-Frame-Rate VideosabstractMulti-object tracking (MOT) in the scenario of low-frame-rate videos is a promising solution for deploying MOT methods on edge devices with limited computing, storage, power, and transmitting bandwidth. Tracking with a low frame rate poses particular challenges in the association stage as objects in two successive frames typically exhibit much quicker variations in locations, velocities, appearances, and visibilities than those in normal frame rates. In this paper, we observe severe performance degeneration of many existing association strategies caused by such variations. Though optical-flow-based methods like CenterTrack can handle the large displacement to some extent due to their large receptive field, the temporally local nature makes them fail to give correct displacement estimations of objects whose visibility flip within adjacent frames. To overcome the local nature of optical-flow-based methods, we propose an online tracking method by extending the CenterTrack architecture with a new head, named APP, to recognize unreliable displacement estimations. Then we design a two-stage association policy where displacement estimations or historical motion cues are leveraged in the corresponding stage according to APP predictions. Our method, with little additional computational overhead, shows robustness in preserving identities in low-frame-rate video sequences. Experimental results on public datasets in various low-frame-rate settings demonstrate the advantages of the proposed method. Wenhan Luo, Zhiguo Shi 0001, Jiming Chen 0001, Qi Ye 0001 |
ACM Multimedia | 5 |
| 2021 | Geometry-based Distance Decomposition for Monocular 3D Object DetectionabstractMonocular 3D object detection is of great significance for autonomous driving but remains challenging. The core challenge is to predict the distance of objects in the absence of explicit depth information. Unlike regressing the distance as a single variable in most existing methods, we propose a novel geometry-based distance decomposition to recover the distance by its factors. The decomposition factors the distance of objects into the most representative and stable variables, i.e. the physical height and the projected visual height in the image plane. Moreover, the decomposition maintains the self-consistency between the two heights, leading to robust distance prediction when both predicted heights are inaccurate. The decomposition also enables us to trace the causes of the distance uncertainty for different scenarios. Such decomposition makes the distance prediction interpretable, accurate, and robust. Our method directly predicts 3D bounding boxes from RGB images with a compact architecture, making the training and inference simple and efficient. The experimental results show that our method achieves the state-of-the-art performance on the monocular 3D Object Detection and Bird’s Eye View tasks of the KITTI dataset, and can generalize to images with different camera intrinsics1. Xuepeng Shi, Qi Ye 0001, Xiaozhi Chen, Chuangrong Chen, Zhixiang Chen 0003, Tae-Kyun Kim 0001 |
ICCV | 2 |
| 2020 | The Phong Surface: Efficient 3D Model Fitting Using Lifted Optimization
Jingjing Shen, Thomas J. Cashman 0001, Qi Ye 0001, Tim Hutton, Toby Sharp, Federica Bogo, Andrew W. Fitzgibbon, Jamie Shotton |
ECCV (1) | 3 |
| 2019 | Opening the Black Box: Hierarchical Sampling Optimization for Hand Pose EstimationabstractHand pose estimation, formulated as an inverse problem, is typically optimized by an energy function over pose parameters using a 'black box' image generation procedure, knowing little about either the relationships between the parameters or the form of the energy function. In this paper, we show significant improvement upon such black box optimization by exploiting high-level knowledge of the parameter structure and using a local surrogate energy function. Our new framework, hierarchical sampling optimization (HSO), consists of a sequence of discriminative predictors organized into a kinematic hierarchy. Each predictor is conditioned on its ancestors, and generates a set of samples over a subset of the pose parameters, with only one selected by the highly-efficient surrogate energy. The selected partial poses are concatenated to generate a full-pose hypothesis. Repeating the same process, several hypotheses are generated and the full energy function selects the best result. Under the same kinematic hierarchy, two methods based on decision forest and convolutional neural network are proposed to generate the samples and two optimization methods are studied when optimizing these samples. Experimental evaluations on three publicly available datasets show that our method is particularly impressive in low-compute scenarios where it significantly outperforms all other state-of-the-art methods. Danhang Tang, Qi Ye 0001, Shanxin Yuan, Jonathan Taylor 0001, Pushmeet Kohli, Cem Keskin, Tae-Kyun Kim 0001, Jamie Shotton |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Occlusion-Aware Hand Pose Estimation Using Hierarchical Mixture Density Network
Qi Ye 0001, Tae-Kyun Kim 0001 |
ECCV (10) | 1 |
| 2017 | BigHand2.2M Benchmark: Hand Pose Dataset and State of the Art AnalysisabstractIn this paper we introduce a large-scale hand pose dataset, collected using a novel capture method. Existing datasets are either generated synthetically or captured using depth sensors: synthetic datasets exhibit a certain level of appearance difference from real depth images, and real datasets are limited in quantity and coverage, mainly due to the difficulty to annotate them. We propose a tracking system with six 6D magnetic sensors and inverse kinematics to automatically obtain 21-joints hand pose annotations of depth maps captured with minimal restriction on the range of motion. The capture protocol aims to fully cover the natural hand pose space. As shown in embedding plots, the new dataset exhibits a significantly wider and denser range of hand poses compared to existing benchmarks. Current state-of-the-art methods are evaluated on the dataset, and we demonstrate significant improvements in cross-benchmark performance. We also show significant improvements in egocentric hand pose estimation with a CNN trained on the new dataset. Shanxin Yuan, Qi Ye 0001, Björn Stenger, Siddhant Jain, Tae-Kyun Kim 0001 |
CVPR | 2 |
| 2016 | Spatial Attention Deep Net with Partial PSO for Hierarchical Hybrid Hand Pose Estimation
Qi Ye 0001, Shanxin Yuan, Tae-Kyun Kim 0001 |
ECCV (8) | 1 |