VLDB 2026 Research / reviewers in the wild / expert
Yongliang Shi
dblp:34/5300
· DBLP profile ↗
22ranked-venue papers
2as first author
20since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 2 first-author · 18 since 2021Systems, architecture and hardware · 14 · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NeRF-Based Transparent Object Grasping Enhanced by Shape PriorsabstractTransparent object grasping remains a persistent challenge in robotics, largely due to the difficulty of acquiring precise 3D information. Conventional optical 3D sensors struggle to capture transparent objects, and machine learning methods are often hindered by their reliance on high-quality datasets. Leveraging NeRF's capability for continuous spatial opacity modeling, our proposed architecture integrates a NeRF-based approach for reconstructing the 3D information of transparent objects. Despite this, certain portions of the reconstructed 3D information may remain incomplete. To address these deficiencies, we introduce a shape-prior-driven completion mechanism, further refined by a geometric pose estimation method we have developed. This allows us to obtain a complete and reliable 3D information of transparent objects. Utilizing this refined data, we perform scene-level grasp prediction and deploy the results in real-world robotic systems. Experimental validation demonstrates the efficacy of our architecture, showcasing its capability to reliably capture 3D information of various transparent objects in cluttered scenes, and correspondingly, achieve high-quality, stable, and executable grasp predictions. Zixin Lin, Lvping Chen, Yongliang Shi, Gan Ma |
ICRA | 5 |
| 2025 | Robo-GS: A Physics Consistent Spatial-Temporal Model for Robotic Arm with Hybrid RepresentationabstractThe Real2Sim2Real (R2S2R) paradigm is critical for advancing robotic learning. Existing methods lack a comprehensive solution to accurately reconstruct real-world objects with both spatial representations and their associated physics attributes in the Real2Sim stage. We propose a Real2Sim pipeline to generate digital assets enabling high-fidelity simulation. We design a hybrid repre-sentation model that integrates mesh geometry, 3D Gaussian kernels, and physics attributes to enhance the representation of robotic arms in digital assets. This hybrid representation is implemented through a Gaussian-Mesh-Pixel binding technique, which establishes an isomorphic mapping between mesh vertices and the Gaussian model. This enables a fully differentiable rendering pipeline that can be optimized through numerical solvers, achieves high-fidelity rendering via Gaussian Splatting, and facilitates physically plausible simulation of the robotic arm's interaction with its environment through mesh geometry. With the digital assets, we propose a fully manipulable Real2Sim pipeline that standardizes coordinate systems and scales, ensuring the seamless integration of multiple components. To demonstrate its effectiveness, we include datasets covering various robotic manipulation tasks with their mesh reconstructions. Our model achieves state-of-the-art results in realistic rendering and mesh reconstruction quality for robotic applications. Our code and datasets will be made publicly available at robostudioapp.com. Haozhe Lou, Yurong Liu, Yike Pan, Yiran Geng, Jianteng Chen, Wenlong Ma 0006, Hengzhen Feng, Liyi Luo, Yongliang Shi |
ICRA | 12 |
| 2025 | OpenBench: A New Benchmark and Baseline for Semantic Navigation in Smart LogisticsabstractThe increasing demand for efficient last-mile delivery in smart logistics underscores the role of autonomous robots in enhancing operational efficiency and reducing costs. Traditional navigation methods, which depend on highprecision maps, are resource-intensive, while learning-based approaches often struggle with generalization in real-world scenarios. To address these challenges, this work proposes the Openstreetmap-enhanced oPen-air sEmantic Navigation (OPEN) system that combines foundation models with classic algorithms for scalable outdoor navigation. The system uses off-the-shelf OpenStreetMap (OSM) for flexible map representation, thereby eliminating the need for extensive pre-mapping efforts. It also employs Large Language Models (LLMs) to comprehend delivery instructions and Vision-Language Models (VLMs) for global localization, map updates, and house number recognition. To compensate the limitations of existing benchmarks that are inadequate for assessing last-mile delivery, this work introduces a new benchmark specifically designed for outdoor navigation in residential areas, reflecting the real-world challenges faced by autonomous delivery systems. Extensive experiments in simulated and real-world environments demonstrate the proposed system's efficacy in enhancing navigation efficiency and reliability. To facilitate further research, our code and benchmark are publicly available11https://ei-nav.github.io/OpenBench/. Dongjie Huo, Zehui Xu, Yongliang Shi, Yimin Yan, Yan Qiao 0004, Guyue Zhou |
ICRA | 4 |
| 2025 | Robust and High-Fidelity 3D Gaussian Splatting: Fusing Pose Priors and Geometry Constraints for Texture-Deficient Outdoor Scenesabstract3D Gaussian Splatting (3DGS) has emerged as a key rendering pipeline for digital asset creation due to its balance between efficiency and visual quality. To address the issues of unstable pose estimation and scene representation distortion caused by geometric texture inconsistency in large outdoor scenes with weak or repetitive textures, we approach the problem from two aspects: pose estimation and scene representation. For pose estimation, we leverage LiDAR-IMU Odometry to provide prior poses for cameras in large-scale environments. These prior pose constraints are incorporated into COLMAP’s triangulation process, with pose optimization performed via bundle adjustment. Ensuring consistency between pixel data association and prior poses helps maintain both robustness and accuracy. For scene representation, we introduce normal vector constraints and effective rank regularization to enforce consistency in the direction and shape of Gaussian primitives. These constraints are jointly optimized with the existing photometric loss to enhance the map quality. We evaluate our approach using both public and self-collected datasets. In terms of pose optimization, our method requires only one-third of the time while maintaining accuracy and robustness across both datasets. In terms of scene representation, the results show that our method significantly outperforms conventional 3DGS pipelines. Notably, on self-collected datasets characterized by weak or repetitive textures, our approach demonstrates enhanced visualization capabilities and achieves superior overall performance. Codes and data will be publicly available at https://github.com/justinyeah/normaljshape.git. Meijun Guo, Yongliang Shi, Caiyun Liu 0004, Yixiao Feng, Tinghai Yan, Weining Lu |
IROS | 2 |
| 2025 | Semi-distributed Cross-modal Air-Ground Relative LocalizationabstractEfficient, accurate, and flexible relative localization is crucial in air-ground collaborative tasks. However, current approaches for robot relative localization are primarily realized in the form of distributed multi-robot SLAM systems with the same sensor configuration, which are tightly coupled with the state estimation of all robots, limiting both flexibility and accuracy. To this end, we fully leverage the high capacity of Unmanned Ground Vehicle (UGV) to integrate multiple sensors, enabling a semi-distributed cross-modal air-ground relative localization framework. In this work, both the UGV and the Unmanned Aerial Vehicle (UAV) independently perform SLAM while extracting deep learning-based keypoints and global descriptors, which decouples the relative localization from the state estimation of all agents. The UGV employs a local Bundle Adjustment (BA) with LiDAR, camera, and an IMU to rapidly obtain accurate relative pose estimates. The BA process adopts sparse keypoint optimization and is divided into two stages: First, optimizing camera poses interpolated from LiDAR-Inertial Odometry (LIO), followed by estimating the relative camera poses between the UGV and UAV. Additionally, we implement an incremental loop closure detection algorithm using deep learning-based descriptors to maintain and retrieve keyframes efficiently. Experimental results demonstrate that our method achieves outstanding performance in both accuracy and efficiency. Unlike traditional multi-robot SLAM approaches that transmit images or point clouds, our method only transmits keypoint pixels and their descriptors, effectively constraining the communication bandwidth under 0.3 Mbps. Codes and data will be publicly available on https://github.com/Ascbpiac/cross-model-relative-localization.git. Weining Lu, Deer Bin, Lian Ma, Xiangyang Chen, Yixiao Feng, Zhouxian Jiang, Yongliang Shi |
IROS | 10 |
| 2025 | L-SNI: A Language-Driven Semantic Navigation System for Inspection TasksabstractFor inspection robots to achieve generalizability, stability, and ease of use, it is crucial that they understand natural language commands and navigate accurately to specified target objects. We propose L-SNI, a semantic navigation system adapted for inspection tasks, offering generalizability, robust stability, and practical ease of use. In the perception phase, L-SNI constructs a precise geometric depth map of the environment using LiDAR, while RGB images are employed to extract object categories, which are then combined with depth data to generate a semantic map. To enable the large language model (LLM) to interpret the environment, L-SNI encodes the 3D semantic map into a plain text representation. During single-task execution, L-SNI decodes human commands into inspection primitives using an LLM constrained by system initial prompts. These inspection primitives guide the robot’s low-level planner for task execution. To address the challenge of traditional 3D LiDAR localization and navigation systems in accurately positioning the robot around target objects during inspection tasks, we propose a target cost gradient to assist in optimizing the robot’s target point selection and attitude control in maps with semantic information. Upon reaching the target, L-SNI uses a visual language model (VLM) to describe the scene, which is simplified by the LLM into a user-friendly response. Through testing on 18 indoor scenes from the Matterport 3D dataset, L-SNI achieves a 46.9% improvement in Success Rate (SR) and a 58.3% increase in Success weighted by Path Length (SPL) over existing state-of-the-art (SOTA) solutions, while also demonstrating superior target image understanding. Moreover, it can be easily deployed on real-world robots without complex initialization. Jiawang Ma, Weichen Guo, Zinan Zhuang, Rongxiang Zeng, Yongliang Shi, Gang Ma 0008 |
IROS | 6 |
| 2025 | OPEN: Lightweight Map-Based Semantic Navigation for GPS-Free Last-Mile DeliveryabstractThe growing demand for efficient last-mile delivery highlights the need for autonomous robots to improve operational efficiency and reduce costs. Traditional navigation methods rely on high-precision maps, which are expensive to produce and maintain, while learning-based approaches often struggle to adapt to diverse real-world environments. To address these challenges, this paper presents OpenStreetMap-enhanced oPen-air sEmantic Navigation (OPEN), a system that combines foundation models with classic navigation algorithms to enable scalable outdoor navigation. By leveraging off-the-shelf OpenStreetMap (OSM), OPEN eliminates the need for extensive pre-mapping and provides a lightweight, readily available map representation. The system uses Large Language Models (LLMs) to interpret delivery instructions and Vision Language Models (VLMs) for global localization without relying on GPS, real-time map updates, and entrance recognition, ensuring robust navigation in complex environments. To further enhance adaptability, OPEN incorporates a local replanning method that dynamically adjusts waypoints in response to environmental changes and OSM inaccuracies. Since existing benchmarks do not adequately reflect the challenges of last-mile delivery, this work introduces a new benchmark designed for residential navigation. Experiments conducted in both simulated and real-world settings demonstrate that OPEN improves navigation accuracy, efficiency, and reliability, outperforming existing semantic navigation methods. To facilitate further research, the code and benchmark are publicly available. Dongjie Huo, Yongliang Shi, Yan Qiao 0004, Guyue Zhou |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Semantic-Independent Dynamic SLAM Based on Geometric Re-Clustering and Optical Flow ResidualsabstractDynamic objects pose significant challenges to the accuracy of state estimation and map quality in Simultaneous Localization and Mapping (SLAM). While current dynamic SLAM methods often rely on semantic information to detect specific movable objects, this dependency on pre-trained models and semantic priors can lead to false dynamic detections. This paper presents a novel semantic-independent dynamic SLAM method that detects truly moving regions, without being constrained by the classes or motion patterns of dynamic objects. We introduce a geometric re-clustering approach to improve object clustering by addressing the under- and over-segmentation caused by the K-Means algorithm. Next, instead of simply classifying entire clusters as dynamic or static, we propose a method to detect dynamic regions within each cluster based on dense optical flow residuals. This enables the detection of partial object movements, such as a seated person moving only his hands. Dynamic detection results are propagated across consecutive frames as dynamic priors for calculating optical flow residuals. Additionally, to enhance map quality, we address the mis-detection of slowly or intermittently moving objects through depth consistency checks applied over a larger time interval. Extensive evaluations on public datasets (TUM and Bonn) and real-world scenes show that our method outperforms state-of-the-art semantic-based methods in terms of localization accuracy and generalizability across various scenarios, particularly when facing unknown dynamic objects. Our method also achieves clean and dense reconstructions, demonstrating its potential for applications like robot navigation in dynamic environments. Hengbo Qi, Xuechao Chen, Zhangguo Yu, Yongliang Shi, Qingrui Zhao, Qiang Huang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Drone-assisted Road Gaussian Splatting with Cross-view Uncertainty
Saining Zhang, Baijun Ye, Xiaoxue Chen, Yuantao Chen, Zongzheng Zhang, Yongliang Shi, Hao Zhao 0002 |
BMVC | 7 |
| 2024 | Block-Map-Based Localization in Large-Scale EnvironmentabstractAccurate localization is an essential technology for the flexible navigation of robots in large-scale environments. Both SLAM-based and map-based localization will increase the computing load due to the increase in map size, which will affect downstream tasks such as robot navigation and services. To this end, we propose a localization system based on Block Maps (BMs) to reduce the computational load caused by maintaining large-scale maps. Firstly, we introduce a method for generating block maps and the corresponding switching strategies, ensuring that the robot can estimate the state in large-scale environments by loading local map information. Secondly, global localization according to Branch-and-Bound Search (BBS) in the 3D map is introduced to provide the initial pose. Finally, a graph-based optimization method is adopted with a dynamic sliding window that determines what factors are being marginalized whether a robot is exposed to a BM or switching to another one, which maintains the accuracy and efficiency of pose tracking. Comparison experiments are performed on publicly available large-scale datasets. Results show that the proposed method can track the robot pose even though the map scale reaches more than 6 kilometers, while efficient and accurate localization is still guaranteed on NCLT [6] and M2DGR [35]. Codes and data will be publicly available on https://github.com/YixFeng/blocklocalization. Yixiao Feng, Yongliang Shi, Yunlong Feng, Hao Zhao 0002, Guyue Zhou |
ICRA | 3 |
| 2024 | An Onboard Framework for Staircases Modeling Based on Point CloudsabstractThe detection of traversable regions on staircases and the physical modeling constitutes pivotal aspects of the mobility of legged robots. This paper presents an onboard framework tailored to the detection of traversable regions and the modeling of physical attributes of staircases by point cloud data. To mitigate the influence of illumination variations and the overfitting due to the dataset diversity, a series of data augmentations are introduced to enhance the training of the fundamental network. A curvature suppression cross-entropy(CSCE) loss is proposed to reduce the ambiguity of prediction on the boundary between traversable and non-traversable regions. Moreover, a measurement correction based on the pose estimation of stairs is introduced to calibrate the output of raw modeling that is influenced by tilted perspectives. Lastly, we collect a dataset pertaining to staircases and introduce new evaluation criteria. Through a series of rigorous experiments conducted on this dataset, we substantiate the superior accuracy and generalization capabilities of our proposed method. Codes, models, and datasets will be available at https://github.com/szturobotics/Stair-detection-and-modeling-project. Chun Qing, Rongxiang Zeng, Yongliang Shi, Gan Ma |
ICRA | 4 |
| 2024 | Camera Relocalization in Shadow-free Neural Radiance FieldsabstractCamera relocalization is a crucial problem in computer vision and robotics. Recent advancements in neural radiance fields (NeRFs) have shown promise in synthesizing photo-realistic images. Several works have utilized NeRFs for refining camera poses, but they do not account for lighting changes that can affect scene appearance and shadow regions, causing a degraded pose optimization process. In this paper, we propose a two-staged pipeline that normalizes images with varying lighting and shadow conditions to improve camera relocalization. We implement our scene representation upon a hash-encoded NeRF which significantly boosts up the pose optimization process. To account for the noisy image gradient computing problem in grid-based NeRFs, we further propose a re-devised truncated dynamic low-pass filter (TDLF) and a numerical gradient averaging technique to smoothen the process. Experimental results on several datasets with varying lighting conditions demonstrate that our method achieves state-of-the-art results in camera relocalization under varying lighting conditions. Code and data will be made publicly available. Shiyao Xu, Caiyun Liu 0004, Yuantao Chen, Zhenxin Zhu, Zike Yan, Yongliang Shi, Hao Zhao 0002, Guyue Zhou |
ICRA | 6 |
| 2024 | LIKO: LiDAR, Inertial, and Kinematic Odometry for Bipedal RobotsabstractHigh-frequency and accurate state estimation is crucial for biped robots. This paper presents a tightly-coupled LiDAR-Inertial-Kinematic Odometry (LIKO) for biped robot state estimation based on an iterated extended Kalman filter. Beyond state estimation, the foot contact position is also modeled and estimated. This allows for both position and velocity updates from kinematic measurement. Additionally, the use of kinematic measurement results in an increased output state frequency of about 1kHz. This ensures temporal continuity of the estimated state and makes it practical for control purposes of biped robots. We also announce a biped robot dataset consisting of LiDAR, inertial measurement unit (IMU), joint encoders, force/torque (F/T) sensors, and motion capture ground truth to evaluate the proposed method. The dataset is collected during robot locomotion, and our approach reached the best quantitative result among other LIO-based methods and biped robot state estimation algorithms. The dataset and source code will be available at https://github.com/Mr-Zqr/LIKO. Qingrui Zhao, Yongliang Shi, Xuechao Chen, Zhangguo Yu, Lianqiang Han, Zhenyuan Fu, Yuanxi Zhang, Qiang Huang 0002 |
ICRA | 3 |
| 2024 | Feasible Region Construction by Polygon Merging for Continuous Bipedal WalkingabstractFeasible regions for continuous walking must provide necessary information for footstep planning, including surrounding landing areas and details about obstacles to be avoided during foot swing. However, the current frame lacks sufficient information to construct a feasible region needed at the current moment due to knee occlusion. To this end, this paper uses polygon merging to construct an information-complete feasible region. This polygon merging refers to merging polygons from the current frame and a specific previous frame. Since the polygon is more concise and efficient than point cloud for environmental representation, construction can be completed quickly without GPU acceleration. Experiments show that the proposed method successfully constructs informative feasible regions within the allowed time frame, enabling the robot to navigate stairs. Xuechao Chen, Hengbo Qi, Qingqing Li 0004, Qingrui Zhao, Yongliang Shi, Zhangguo Yu, Lingxuan Zhao, Zhihong Jiang |
IROS | 6 |
| 2024 | Blending Distributed NeRFs with Tri-stage Robust Pose OptimizationabstractDue to the limited model capacity, leveraging distributed Neural Radiance Fields (NeRFs) for modeling extensive urban environments has become a necessity. However, current distributed NeRF registration approaches encounter aliasing artifacts, arising from discrepancies in rendering resolutions and suboptimal pose precision. These factors collectively deteriorate the fidelity of pose estimation within NeRF frameworks, resulting in occlusion artifacts during the NeRF blending stage. In this paper, we present a distributed NeRF system with tri-stage pose optimization. In the first stage, precise poses of images are achieved by bundle adjusting Mip-NeRF 360 with a coarse-to-fine strategy. In the second stage, we incorporate the inverting Mip-NeRF 360, coupled with the truncated dynamic low-pass filter, to enable the achievement of robust and precise poses, termed Frame2Model optimization. On top of this, we obtain a coarse transformation between NeRFs in different coordinate systems. In the third stage, we fine-tune the transformation between NeRFs by Model2Model pose optimization. After obtaining precise transformation parameters, we proceed to implement NeRF blending, showcasing superior performance metrics in both real-world and simulation scenarios. Codes and data will be publicly available at https://github.com/boilcy/Distributed-NeRF. Baijun Ye, Caiyun Liu 0004, Xiaoyu Ye, Yuantao Chen, Yuhai Wang, Zike Yan, Yongliang Shi, Hao Zhao 0002, Guyue Zhou |
IROS | 7 |
| 2024 | City-scale continual neural semantic mapping with three-layer sampling and panoptic representation
Yongliang Shi, Runyi Yang, Zirui Wu, Pengfei Li 0007, Caiyun Liu 0004, Hao Zhao 0002, Guyue Zhou |
Knowl. Based Syst. | 1 |
| 2023 | INT2: Interactive Trajectory Prediction at IntersectionsabstractMotion forecasting is an important component in autonomous driving systems. One of the most challenging problems in motion forecasting is interactive trajectory prediction, whose goal is to jointly forecasts the future trajectories of interacting agents. To this end, we present a large-scale interactive trajectory prediction dataset named INT2 for INTeractive trajectory prediction at INTersections. INT2 includes 612,000 scenes, each lasting 1 minute, containing up to 10,200 hours of data. The agent trajectories are auto-labeled by a high-performance offline temporal detection and fusion algorithm, whose quality is further inspected by human judges. Vectorized semantic maps and traffic light information are also included in INT2. Additionally, the dataset poses an interesting domain mismatch challenge. For each intersection, we treat rush-hour and non-rush-hour segments as different domains. We benchmark the best open-sourced interactive trajectory prediction method on INT2 and Waymo Open Motion, under in-domain and cross-domain settings. The dataset, code and models are publicly available at https://github.com/AIRDISCOVER/INT2. Zhijie Yan, Pengfei Li 0007, Zheng Fu, Shaocong Xu, Yongliang Shi, Xiaoxue Chen, Yuhang Zheng 0004, Yang Li 0178, Tianyu Liu 0008, Chuxuan Li, Nairui Luo, Zuoxu Wang, Yifeng Shi, Zhengxiao Han, Jirui Yuan, Jiangtao Gong, Guyue Zhou, Hang Zhao 0021, Hao Zhao 0002 |
ICCV | 5 |
| 2023 | LODE: Locally Conditioned Eikonal Implicit Scene Completion from Sparse LiDARabstractScene completion refers to obtaining dense scene representation from an incomplete perception of complex 3D scenes. This helps robots detect multi-scale obstacles and analyse object occlusions in scenarios such as autonomous driving. Recent advances show that implicit representation learning can be leveraged for continuous scene completion and achieved through physical constraints like Eikonal equations. However, former Eikonal completion methods only demonstrate results on watertight meshes at a scale of tens of meshes. None of them are successfully done for non-watertight LiDAR point clouds of open large scenes at a scale of thousands of scenes. In this paper, we propose a novel Eikonal formulation that conditions the implicit representation on localized shape priors which function as dense boundary value constraints, and demonstrate it works on SemanticKITTI and SemanticPOSS. It can also be extended to semantic Eikonal scene completion with only small modifications to the network architecture. With extensive quantitative and qualitative results, we demonstrate the benefits and drawbacks of existing Eikonal methods, which naturally leads to the new locally conditioned formulation. Notably, we improve IoU from 31.7% to 51.2% on SemanticKITTI and from 40.5% to 48.7% on SemanticPOSS. We extensively ablate our methods and demonstrate that the proposed formulation is robust to a wide spectrum of implementation hyper-parameters. Codes and models are publicly available at https://github.com/AIR-DISCOVER/LODE Pengfei Li 0007, Ruowen Zhao, Yongliang Shi, Hao Zhao 0002, Jirui Yuan, Guyue Zhou, Ya-Qin Zhang |
ICRA | 3 |
| 2023 | LATITUDE: Robotic Global Localization with Truncated Dynamic Low-pass Filter in City-scale NeRFabstractNeural Radiance Fields (NeRFs) have made great success in representing complex 3D scenes with high-resolution details and efficient memory. Nevertheless, current NeRF - based pose estimators have no initial pose prediction and are prone to local optima during optimization. In this paper, we present LATITUDE: Global Localization with Truncated Dynamic Low-pass Filter, which introduces a two-stage localization mechanism in city-scale NeRF. In place recognition stage, we train a regressor through images generated from trained NeRFs, which provides an initial value for global localization. In pose optimization stage, we minimize the residual between the observed image and rendered image by directly optimizing the pose on the tangent plane. To avoid falling into local optimum, we introduce a Truncated Dynamic Low-pass Filter (TDLF) for coarse-to-fine pose registration. We evaluate our method on both synthetic and real-world data and show its potential applications for high-precision navigation in large-scale city scenes. Codes and dataset will be publicly available at https://github.com/jike5/LATITUDE. Zhenxin Zhu, Yuantao Chen, Zirui Wu, Yongliang Shi, Chuxuan Li, Pengfei Li 0007, Hao Zhao 0002, Guyue Zhou |
ICRA | 5 |
| 2022 | TOIST: Task Oriented Instance Segmentation Transformer with Noun-Pronoun DistillationabstractCurrent referring expression comprehension algorithms can effectively detect or segment objects indicated by nouns, but how to understand verb reference is still under-explored. As such, we study the challenging problem of task oriented detection, which aims to find objects that best afford an action indicated by verbs like sit comfortably on. Towards a finer localization that better serves downstream applications like robot interaction, we extend the problem into task oriented instance segmentation. A unique requirement of this task is to select preferred candidates among possible alternatives. Thus we resort to the transformer architecture which naturally models pair-wise query relationships with attention, leading to the TOIST method. In order to leverage pre-trained noun referring expression comprehension models and the fact that we can access privileged noun ground truth during training, a novel noun-pronoun distillation framework is proposed. Noun prototypes are generated in an unsupervised manner and contextual pronoun features are trained to select prototypes. As such, the network remains noun-agnostic during inference. We evaluate TOIST on the large-scale task oriented dataset COCO-Tasks and achieve +10.7% higher $\rm{mAP^{box}}$ than the best-reported results. The proposed noun-pronoun distillation can boost $\rm{mAP^{box}}$ and $\rm{mAP^{mask}}$ by +2.6% and +3.6%. Codes and models are publicly available. Pengfei Li 0007, Beiwen Tian, Yongliang Shi, Xiaoxue Chen, Hao Zhao 0002, Guyue Zhou, Ya-Qin Zhang |
NeurIPS | 3 |
| 2020 | ReinforcedRimJump: Tangent-Based Shortest-Path Planning for Two-Dimensional MapsabstractPath planning under two-dimensional maps is a fundamental problem in mobile robotics and other real-world applications (unmanned vehicles, navigation applications for mobile phones, and so forth). However, traditional algorithms (graph searching, artificial potential field, genetic, and so forth) rely on grid-by-grid searching. Thus, these methods generally do not find the global optimal path, and as the map scale increases, their time cost increase sharply, except artificial potential field. A few algorithms that do not rely on grid-by-grid searching (rapidly-exploring random tree, visibility graph, and tangent graph) have special requirements for maps. Considering that the shortest path is composed of tangents between obstacles, in this paper, we propose a method called ReinforcedRimJump (RRJ) that does not rely on the point-by-point traversal but rather obtains the shortest path by finding the tangent multiple times between obstacles. The first improvement of this method is the precomputation of tangents, which causes the method to have a lower time cost than traditional methods. The second improvement of RRJ is edge segmentation, which allows RRJ to be used when the target is in the depression of the obstacle. To verify the theoretical advantages of RRJ, some comparative experiments under various maps are performed. The experimental results show that RRJ can always find the shortest path in the shortest time. Furthermore, the time cost of RRJ is insensitive to the map size compared to other methods. The experimental results presented herein demonstrate that RRJ meets the theoretical expectations. Zhuo Yao, Yongliang Shi, Mingzhu Li, Zhenshuo Liang, Qiang Huang 0002 |
IEEE Trans. Ind. Informatics | 3 |
| 2006 | Developing Methodologies of Knowledge Discovery and Data Mining to Investigate Metropolitan Land Use Evolution
Yongliang Shi, Rusong Wang |
PRICAI | 1 |