VLDB 2026 Research / reviewers in the wild / expert
Hanjing Ye
dblp:317/8332
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0002-8522-3638ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FlowPlan: Zero-Shot Task Planning with LLM Flow Engineering for Robotic Instruction FollowingabstractRobotic instruction following tasks require seamless integration of visual perception, task planning, target localization, and motion execution. However, existing task planning methods for instruction following are either data-driven or underperform in zero-shot scenarios due to difficulties in grounding lengthy instructions into actionable plans under operational constraints. To address this, we propose FlowPlan, a structured multi-stage LLM workflow that elevates zero-shot pipeline and bridges the performance gap between zero-shot and data-driven in-context learning methods. By decomposing the planning process into modular stages—task information retrieval, language-level reasoning, symbolic-level planning, and logical evaluation—FlowPlan generates logically coherent action sequences while adhering to operational constraints and further extracts contextual guidance for precise instance-level target localization. Benchmarked on ALFRED and validated in real-world applications, our method achieves competitive performance relative to data-driven in-context learning methods and demonstrates adaptability across diverse environments. This work advances zero-shot task planning in robotic systems without reliance on labeled data. Project website: https://instruction-following-project.github.io/. Chao Tang 0001, Hanjing Ye, Hong Zhang 0013 |
IROS | 3 |
| 2025 | Monocular Person Localization under Camera Ego-MotionabstractLocalizing a person from a moving monocular camera is critical for Human-Robot Interaction (HRI). To estimate the 3D human position from a 2D image, existing methods either depend on the geometric assumption of a fixed camera or use a position regression model trained on datasets containing little camera ego-motion. These methods are vulnerable to fierce camera ego-motion, resulting in inaccurate person localization. We consider person localization as a part of a pose estimation problem. By representing a human with a four-point model, our method jointly estimates the 2D camera attitude and the person’s 3D location through optimization. Evaluations on both public datasets and real robot experiments demonstrate our method outperforms baselines in person localization accuracy. Our method is further implemented into a person-following system and deployed on an agile quadruped robot. Hanjing Ye, Hong Zhang 0013 |
IROS | 2 |
| 2025 | A Hierarchical Progressive Perception System for Autonomous Luggage Trolley CollectionabstractAdvancements in intelligent vehicle technologies and autonomous driving are now used for tasks like collecting luggage trolleys at airports. The robotic autonomous luggage trolley collection system employs robots to gather and transport scattered luggage trolleys. However, existing methods for detecting and locating luggage trolleys often fail when they are not fully visible. To address this, we introduce the hierarchical progressive perception system, which enhances the detection and localization of luggage trolleys under partial occlusion. The proposed system integrates the hierarchical process structure and progressive perception strategy. This innovative structure processes the luggage trolley’s position and orientation separately. It can accurately determine the luggage trolley’s position with just one well-detected keypoint and estimate its orientation when partially occluded. Once the luggage trolley’s initial pose is detected, the progressive perception strategy continuously refines this information until the robot begins grasping. The proposed system only needs RGB images for labeling and training, eliminating the need for complex data collection and annotation. The experiments on detection and localization demonstrate that the proposed system is more reliable under partial occlusion compared to existing methods. Its effectiveness and robustness have also been confirmed through practical tests in actual luggage trolley collection tasks. A website about this work is available at https://sites.google.com/view/robot-perception. Zhirui Sun, Jieting Zhao, Hanjing Ye, Jiankun Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2024 | GV-Bench: Benchmarking Local Feature Matching for Geometric Verification of Long-term Loop Closure DetectionabstractVisual loop closure detection is an important module in visual simultaneous localization and mapping (SLAM), which associates current camera observation with previously visited places. Loop closures correct drifts in trajectory estimation to build a globally consistent map. However, a false loop closure can be fatal, so verification is required as an additional step to ensure robustness by rejecting the false positive loops. Geometric verification has been a well-acknowledged solution that leverages spatial clues provided by local feature matching to find true positives. Existing feature matching methods focus on homography and pose estimation in long-term visual localization, lacking references for geometric verification. To fill the gap, this paper proposes a unified benchmark targeting geometric verification of loop closure detection under long-term conditional variations. Furthermore, we evaluate six representative local feature matching methods (handcrafted and learning-based) under the benchmark, with in-depth analysis for limitations and future directions. Jingwen Yu, Hanjing Ye, Jianhao Jiao, Ping Tan 0002, Hong Zhang 0013 |
IROS | 2 |
| 2024 | Human Orientation Estimation Under Partial ObservationabstractReliable Human Orientation Estimation (HOE) from a monocular image is critical for autonomous agents to understand human intention. Significant progress has been made in HOE under full observation. However, the existing methods easily make a wrong prediction under partial observation and give it an unexpectedly high confidence. To solve the above problems, this study first develops a method called Part-HOE that estimates orientation from the visible joints of a target person so that it is able to handle partial observation. Subsequently, we introduce a confidence-aware orientation estimation method, enabling more accurate orientation estimation and reasonable confidence estimation under partial observation. The effectiveness of our method is validated on both public and custom-built datasets, and it shows great accuracy and reliability improvement in partial observation scenarios. In particular, we show in real experiments that our method can benefit the robustness and consistency of the Robot Person Following (RPF) task. Jieting Zhao, Hanjing Ye, Hong Zhang 0013 |
IROS | 2 |
| 2023 | Robot Person Following Under Partial OcclusionabstractRobot person following (RPF) is a capability that supports many useful human-robot-interaction (HRI) applications. However, existing solutions to person following often as-sume full observation of the tracked person. As a consequence, they cannot track the person reliably under partial occlusion where the assumption of full observation is not satisfied. In this paper, we focus on the problem of robot person following under partial occlusion caused by a limited field of view of a monocular camera. Based on the key insight that it is possible to locate the target person when one or more of hislher joints are visible, we propose a method in which each visible joint contributes a location estimate of the followed person. Experiments on a public person-following dataset show that, even under partial occlusion, the proposed method can still locate the person more reliably than the existing SOTA methods. As well, the application of our method is demonstrated in real experiments on a mobile robot. Hanjing Ye, Jieting Zhao, Yaling Pan, Weinan Chen, Li He 0002, Hong Zhang 0013 |
ICRA | 1 |
| 2023 | Multi-Scale Point Octree Encoding Network for Point Cloud Based Place RecognitionabstractOver the past decades, point cloud-based place recognition has garnered significant attention. This research paper presents a pioneering approach, denoted as the Multi-scale Point Octree Encoding Network (MPOE-Net), designed to acquire a discriminative global descriptor for efficient retrieval of places. The key element of the MPOE-Net is the point octree encoding module, which adeptly captures local information for each point by considering its nearest and farthest neighbors. Further enhancing local relationships, a multi-transformer network is introduced, utilizing a novel grouped offset-attention mechanism. To amalgamate the multi-scale attention maps into a comprehensive global descriptor, a multi-NetVLAD layer is incorporated. Through rigorous experimentation across diverse benchmark datasets, our proposed method unequivocally outperforms existing techniques in the realm of point cloud-based place recognition tasks, achieving state-of-the-art results. Our code is released publicly at https://github.com/Zhilong-Tang/MPOE-Net. Zhilong Tang, Hanjing Ye, Hong Zhang 0013 |
IROS | 2 |
| 2022 | Keyframe Selection with Information Occupancy Grid Model for Long-term Data AssociationabstractAs the basics of Visual Simultaneous Localization And Mapping (VSLAM), keyframes play an essential role. In previous works, keyframes are selected according to a series of view change-based strategies for short-term data association (STDA). However, the texture enrichment of frames is always ignored, resulting in the failure of long-term data association (LTDA). In this paper, we propose an information enrichment selection strategy with an information occupancy grid model and a deep descriptor. Frame is expressed by a deep global descriptor for a statistical explainable abstraction, in which the texture enrichment is indicated. Based on the abstraction, an information occupancy grid model is established to measure the information enrichment and the potential LTDA ability. Evaluations on variant datasets are conducted, showing the advantage of our proposed method in terms of keyframe selection and tracking precision. Also, the statistical explainability of the deep descriptor is provided. The proposed keyframe selection strategy can improve LTDA and tracking precision, especially in situations with repeated observations and loop-closures. Weinan Chen, Hanjing Ye, Chao Tang 0001, Changfei Fu, Hong Zhang 0013 |
IROS | 2 |