Jianmin Ji

dblp:16/1844 · also Jian-Min Ji · DBLP profile ↗
← Back
75ranked-venue papers
15as first author
49since 2021 · last 2026
0000-0002-1515-0402ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 53 · 11 first-author · 35 since 2021Systems, architecture and hardware · 23 · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-author · 12 since 2021Software engineering, systems software and programming languages · 7 · 3 first-author · 2 since 2021Theory of computation · 6 · 2 first-author · 1 since 2021Computer networks · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Vulnerability-Aware Robust Multimodal Adversarial Training
abstract
Multimodal learning has shown significant superiority on various tasks by integrating multiple modalities. However, the interdependencies among modalities increase the susceptibility of multimodal models to adversarial attacks. Existing methods mainly focus on attacks on specific modalities or indiscriminately attack all modalities. In this paper, we find that these approaches ignore the differences between modalities in their contribution to final robustness, resulting in suboptimal robustness performance. To bridge this gap, we introduce Vulnerability-Aware Robust Multimodal Adversarial Training (VARMAT), a probe-in-training adversarial training method that improves multimodal robustness by identifying the vulnerability of each modality. To be specific, VARMAT first explicitly quantifies the vulnerability of each modality, grounded in a first-order approximation of the attack objective (Probe). Then, we propose a targeted regularization term that penalizes modalities with high vulnerability, guiding robust learning while maintaining task accuracy (Training). We demonstrate the enhanced robustness of our method across multiple multimodal datasets involving diverse modalities. Finally, we achieve {12.73%, 22.21%, 11.19%} robustness improvement on three multimodal datasets, revealing a significant blind spot in multimodal adversarial training.
Junrui Zhang 0012, Jie Peng 0002, Chenjie Wang, Jianmin Ji, Tianlong Chen 0001
AAAI5
2026 Needle in a Haystack: Tracking UAVs from Massive Noise in Real-World 5G-A Base Station Data
Chengzhen Meng, Chenming He, Yidong Jiang, Xiaoran Fan, Dequan Wang, Jianmin Ji, Yanyong Zhang
MobiSys7
2025 Traffic Scenario Logic: A Spatial-Temporal Logic for Modeling and Reasoning of Urban Traffic Scenarios
abstract
Formal representations of traffic scenarios can be used to generate test cases for the safety verification of autonomous driving. However, most existing methods are limited to highway or highly simplified intersection scenarios due to the intricacy and diversity of traffic scenarios. In response, we propose Traffic Scenario Logic (TSL), which is a spatial-temporal logic designed for modeling and reasoning of urban pedestrian-free traffic scenarios. TSL provides a formal representation of the urban road network that can be derived from OpenDRIVE, i.e., the de facto industry standard of high-definition maps for autonomous driving, enabling the representation of a broad range of traffic scenarios without discretization approximations. We implemented the reasoning of TSL using Telingo, i.e., a solver for temporal programs based on Answer Set Programming, and tested it on different urban road layouts. Demonstrations show the effectiveness of TSL in test scenario generation and its potential value in areas like decision-making and control verification of autonomous driving. The code for TSL reasoning has been open-sourced.
Ruolin Wang, Yuejiao Xu, Jianmin Ji
AAAI3
2025 GraspCoT: Integrating Physical Property Reasoning for 6-DoF Grasping Under Flexible Language Instructions
abstract
Flexible instruction-guided 6-DoF grasping is a significant yet challenging task for real-world robotic systems. Existing methods utilize the contextual understanding capabilities of the large language models (LLMs) to establish mappings between expressions and targets, allowing robots to comprehend users' intentions in the instructions. However, the LLM's knowledge about objects' physical properties remains underexplored despite its tight relevance to grasping. In this work, we propose GraspCoT, a 6-DoF grasp detection framework that integrates a Chain-of-Thought (CoT) reasoning mechanism oriented to physical properties, guided by auxiliary question-answering (QA) tasks. Particularly, we design a set of QA templates to enable hierarchical reasoning that includes three stages: target parsing, physical property analysis, and grasp action selection. Moreover, GraspCoT presents a unified multimodal LLM architecture, which encodes multi-view observations of 3D scenes into 3D-aware visual tokens, and then jointly embeds these visual tokens with CoT-derived textual tokens within LLMs to generate grasp pose predictions. Furthermore, we present IntentGrasp, a large-scale benchmark that fills the gap in public datasets for multi-object grasp detection under diverse and indirect verbal commands. Extensive experiments on IntentGrasp demonstrate the superiority of our method, with additional validation in real-world robotic applications confirming its practicality. The code is available at https://github.com/cxmomo/GraspCoT.
Xiaomeng Chu, Jiajun Deng, Guoliang You, Jianmin Ji, Yanyong Zhang
ICCV6
2025 SpatialSplat: Efficient Semantic 3D from Sparse Unposed Images
abstract
A major breakthrough in 3D reconstruction is the feedforward paradigm to generate pixel-wise 3D points or Gaussian primitives from sparse, unposed images. To further incorporate semantics while avoiding the significant memory and storage costs of high-dimensional semantic features, existing methods extend this paradigm by associating each primitive with a compressed semantic feature vector. However, these methods have two major limitations: (a) the naively compressed feature compromises expressiveness, affecting the model's ability to capture fine-grained semantics, and (b) the pixel-wise primitive prediction introduces redundancy in overlapping areas, causing unnecessary memory overhead. To this end, we introduce \textbf{SpatialSplat}, a feedforward framework that produces redundancy-aware Gaussians and capitalizes on a dual-field semantic representation. Particularly, with the insight that primitives within the same instance exhibit high semantic consistency, we decompose the semantic representation into a coarse feature field that encodes uncompressed semantics with minimal primitives, and a fine-grained yet low-dimensional feature field that captures detailed inter-instance relationships. Moreover, we propose a selective Gaussian mechanism, which retains only essential Gaussians in the scene, effectively eliminating redundant primitives. Our proposed Spatialsplat learns accurate semantic information and detailed instances prior with more compact 3D Gaussians, making semantic 3D reconstruction more applicable. We conduct extensive experiments to evaluate our method, demonstrating a remarkable 60\% reduction in scene representation parameters while achieving superior performance over state-of-the-art methods. The code is available at https://github.com/shengyuuu/SpatialSplat.git
Yu Sheng, Jiajun Deng, Yu Zhang 0086, Bei Hua, Yanyong Zhang, Jianmin Ji
ICCV7
2025 CELLmap: Enhancing LiDAR SLAM Through Elastic and Lightweight Spherical Map Representation
abstract
SLAM is a fundamental capability of unmanned systems, with LiDAR-based SLAM gaining widespread adoption due to its high precision. Current SLAM systems can achieve centimeter-level accuracy within a short period. However, there are still several challenges when dealing with largescale mapping tasks including significant storage requirements and difficulty of reusing the constructed maps. To address this, we first design an elastic and lightweight map representation called CELLmap, composed of several CELLS, each representing the local map at the corresponding location. Then, we design a general backend including CELL-based bidirectional registration module and loop closure detection module to improve global map consistency. Our experiments have demonstrated that CELLmap can represent the precise geometric structure of large-scale maps of KITTI dataset using only about 60 MB. Additionally, our general backend achieves up to a 26.88% improvement over various LiDAR odometry methods.
Yifan Duan, Yao Li 0016, Guoliang You, Xiaomeng Chu, Jianmin Ji, Yanyong Zhang
ICRA6
2025 OG-Gaussian: Occupancy Based Street Gaussians for Autonomous Driving
abstract
Accurate and realistic 3D scene reconstruction enables the lifelike creation of autonomous driving simulation environments. With advancements in 3D Gaussian Splatting (3DGS), previous studies have applied it to reconstruct complex dynamic driving scenes. These methods typically require expensive LiDAR sensors and pre-annotated datasets of dynamic objects. To address these challenges, we propose OG-Gaussian, a novel approach that replaces LiDAR point clouds with Occupancy Grids (OGs) generated from surround-view camera images using Occupancy Prediction Network (ONet). Our method leverages the semantic information in OGs to separate dynamic vehicles from static street background, converting these grids into two distinct sets of initial point clouds for reconstructing both static and dynamic objects. Additionally, we estimate the trajectories and poses of dynamic objects through a learning-based approach, eliminating the need for complex manual annotations. Experiments on Waymo Open dataset demonstrate that OG-Gaussian is on par with the current state-of-the-art in terms of reconstruction quality and rendering speed, achieving an average PSNR of 35.13 and a rendering speed of 143 FPS, while significantly reducing computational costs and economic overhead.
Yedong Shen, Yifan Duan, Yilong Wu, Jianmin Ji, Yanyong Zhang, Huiqing Jin
ICRA7
2025 MT-PCR: Leveraging Modality Transformation for Large-Scale Point Cloud Registration with Limited Overlap
abstract
Large-scale scene point cloud registration with limited overlap is a challenging task due to computational load and constrained data acquisition. To tackle these issues, we propose a point cloud registration method, MT-PCR, based on Modality Transformation. MT-PCR leverages a Bird's Eye View (BEV) capturing the maximal overlap information to improve the accuracy and utilizes images to provide complementary spatial features. Specifically, MT-PCR converts 3D point clouds to BEV images and estimates correspondence by 2D image keypoints extraction and matching. Subsequently, the 2D correspondence estimates are then transformed back to 3D point clouds using inverse mapping. We have applied MT-PCR to Terrestrial Laser Scanning (TLS) and Aerial Laser Scanning (ALS) point cloud registration on the GrAco dataset, involving 8 low-overlap, square-kilometer scale registration scenarios. Experiments and comparisons with commonly used methods demonstrate that MT-PCR can achieve superior accuracy and robustness in large-scale scenes with limited overlap.
Yilong Wu, Yifan Duan, Yedong Shen, Jianmin Ji, Yanyong Zhang
ICRA6
2025 CAFE-AD: Cross-Scenario Adaptive Feature Enhancement for Trajectory Planning in Autonomous Driving
abstract
Imitation learning based planning tasks on the nuPlan dataset have gained great interest due to their potential to generate human-like driving behaviors. However, open-loop training on the nuPlan dataset tends to cause causal confusion during closed-loop testing, and the dataset also presents a longtail distribution of scenarios. These issues introduce challenges for imitation learning. To tackle these problems, we introduce CAFE-AD, a Cross-Scenario Adaptive Feature Enhancement for Trajectory Planning in Autonomous Driving method, designed to enhance feature representation across various scenario types. We develop an adaptive feature pruning module that ranks feature importance to capture the most relevant information while reducing the interference of noisy information during training. Moreover, we propose a cross-scenario feature interpolation module that enhances scenario information to introduce diversity, enabling the network to alleviate overfitting in dominant scenarios. We evaluate our method CAFEAD, on the challenging public nuPlan Test14-Hard closed-loop simulation benchmark. The results demonstrate that CAFEAD outperforms state-of-the-art methods including rule-based and hybrid planners, and exhibits the potential in mitigating the impact of long-tail distribution within the dataset. Additionally, we further validate its effectiveness in real-world environments. The code and models will be made available at https://github.com/AlniyatRui/CAFE-AD.
Junrui Zhang 0012, Chenjie Wang, Jie Peng 0002, Jianmin Ji, Yu Zhang 0086, Yanyong Zhang
ICRA5
2025 Improving Efficiency of Answer Set Planning with Rough Solutions from Large Language Models for Robotic Task Planning
abstract
Answer Set Programming (ASP) planning can be used to refine the rough solutions generated by Large Language Models (LLMs) to handle specific restrictions of actions, i.e., reconstruct the rough solutions to be executable, for robotic task planning. However, it is still challenging to efficiently solve ASP programs that have multiple variables with large domains, which prevents the above application of ASP planning from real-world task planning problems. In this paper, we consider how to reduce the domains of variables without losing possible solutions for ASP planning, while given these rough solutions from LLMs. Based on the above reduction, we introduce CLMASP, an approach that couples LLMs with ASP for robotic task planning. We evaluate CLMASP on the VirtualHome platform for common indoor tasks, demonstrating a significant improvement in the executable rate from under 10% to nearly 90% and reducing average ASP planning time from over 2 hours to under 5 seconds. Code is available at https://github.com/CLMASP/CLMASP.
Xinrui Lin, Yangfan Wu, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang
IJCAI6
2025 AAOPL: Automated Articulated Object Parameter Learning for Open-World Robotics
abstract
Articulated objects are ubiquitous in daily environments, and effective manipulation of these objects is essential for advancing open-world robotics. Existing approaches, which rely heavily on large-scale data collection or simulation, often face limitations in real-world applications, including issues with generalization and the sim-to-real gap. In this paper, we introduce the Automated Articulated Object Parameter Learning (AAOPL) framework, which autonomously learns the articulation parameters of real-world articulated objects through direct interaction. This approach enables robots to generate precise manipulation trajectories without relying on predefined object models or extensive human demonstration data. To accelerate the learning process, we develop Accelerated Single-Step Gradient (ASSG) algorithm, which efficiently refines the articulation parameters by leveraging real-time execution feedback. Experimental results demonstrate that AAOPL can learn accurate articulation parameters within 30 minutes (80 trials) and generate robust manipulation trajectories, outperforming baseline methods in terms of both task completion and force efficiency. Our approach eliminates the need for large-scale training datasets and can adapt to various articulated objects in real-world environments, offering a scalable solution for autonomous robotic manipulation in unstructured settings.
Ziyang Feng, Quecheng Qiu, Silong Zhang, Jianmin Ji
IROS4
2025 NaviDiffuser: Tackling Multi-Objective Robot Navigation by Weight Range Guided Diffusion Model
abstract
The data-driven paradigm has shown great potential in solving many decision-making tasks. In the robot navigation realm, it also sparked a new trend. People believe powerful data-driven methods can learn efficient and general navigation policies from a vast offline dataset. However, robot navigation tasks differ from common planning tasks and present unique challenges. It often involves multi-objective optimization to meet arbitrary and ever-changing human preferences. It should also overcome the short-sighted problem to obtain globally optimal performance. Furthermore, high planning frequency is needed to address real-time demands. These factors obstruct the application of data-driven methods in robot navigation. To address these challenges, we integrate one of the most powerful data-driven methods, the diffusion model, into robot navigation. Our proposed approach, NaviDiffuser, utilizes a novel classification label to guide the diffusion model in capturing the complex connections between navigation and human preferences. Its Transformer network backbone outputs action sequences to alleviate short-sightedness. It also includes special distillation skills to boost the planning speed and quality. We conduct experiments in both simulated and real-world scenarios to evaluate our approach. In these experiments, NaviDiffuser not only demonstrates an extremely high arrival rate but also adjusts its navigation policy to align with different human preferences.
Ziyang Feng, Quecheng Qiu, Jie Peng 0002, Jianmin Ji
IROS6
2025 Hierarchical Framework for Constrained Dual-Arm Cooperative Manipulation with Whole-Body Collision Avoidance
abstract
Dual-arm robotic systems hold great potential for complex bimanual tasks that require intricate and coordinated manipulation, such as holding and transporting a tray with a cup of coffee while navigating through cluttered environments. However, these tasks pose significant challenges due to the inherent closed-chain constraints between the arms and the object, as well as the need for real-time collision avoidance, especially in real-world applications. To address these challenges, we introduce a hierarchical framework that combines learning-based planning with classical control theory to ensure whole-body collision avoidance movement while maintaining the kinematic relationship. In addition, we present a novel, efficient, and cost-free data generation method specifically designed for dual-arm cooperative tasks, overcoming the lack of sufficient training data. Extensive experiments in both simulation and real-world scenarios demonstrate that our approach improves the success rate by 26.3% compared to existing planning methods and by 54.7% compared to end-to-end methods. These results highlight the advantages of our method in whole-body collision avoidance and environmental adaptability, making it a promising solution for dual-arm cooperative tasks.
Silong Zhang, Quecheng Qiu, Yingtai Ni, Yuecheng Shao, Ziyang Feng, Jianmin Ji
IROS6
2025 Ghost Points Matter: Far-Range Vehicle Detection with a Single mmWave Radar in Tunnel
abstract
Vehicle detection in tunnels is crucial for traffic monitoring and accident response, yet remains underexplored. In this paper, we develop mmTunnel, a millimeter-wave radar system that achieves far-range vehicle detection in tunnels. The main challenge here is coping with ghost points caused by multi-path reflections, which lead to severe localization errors and false alarms. Instead of merely removing ghost points, we propose correcting them to true vehicle positions by recovering their signal reflection paths, thus reserving more data points and improving detection performance, even in occlusion scenarios. However, recovering complex 3D reflection paths from limited 2D radar points is highly challenging. To address this problem, we develop a multi-path ray tracing algorithm that leverages the ground plane constraint and identifies the most probable reflection path based on signal path loss and spatial distance. We also introduce a curve-to-plane segmentation method to simplify tunnel surface modeling such that we can significantly reduce the computational delay and achieve real-time processing.
Chenming He, Chengzhen Meng, Xiaoran Fan, Dequan Wang, Haojie Ren, Jianmin Ji, Yanyong Zhang
MobiCom7
2025 UrgenGo: Urgency-Aware Transparent GPU Kernel Launching for Autonomous Driving
abstract
The rapid advancements in autonomous driving have introduced increasingly complex, real-time GPU-bound tasks critical for reliable vehicle operation. However, the proprietary nature of these autonomous systems and closed-source GPU drivers hinder fine-grained control over GPU executions, often resulting in missed deadlines that compromise vehicle performance. To address this, we present UrgenGo, a non-intrusive, urgency-aware GPU scheduling system that operates without access to application source code. UrgenGo implicitly prioritizes GPU executions through transparent kernel launch manipulation, employing task-level stream binding, delayed kernel launching, and batched kernel launch synchronization. We conducted extensive real-world evaluations in collaboration with a self-driving startup, developing 11 GPU-bound task chains for a realistic autonomous navigation application and implementing our system on a self-driving bus. Our results show a significant 61% reduction in the overall deadline miss ratio, compared to the state-of-the-art GPU scheduler that requires source code modifications.
Hanqi Zhu, Wuyang Zhang, Ziyang Tao, Xinrui Lin, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang
MobiCom7
2025 OA-DET3D: Embedding Object Awareness As A General Plug-in for Multi-Camera 3D Object Detection
Xiaomeng Chu, Jiajun Deng, Jianmin Ji, Yu Zhang 0086, Houqiang Li, Yanyong Zhang
Int. J. Comput. Vis.3
2025 Towards unified bijective image-text generation for text-to-image person re-identification
Xiaoguang Ma, Jianmin Ji, Honghu Pan
Knowl. Based Syst.4
2025 iLoc: An Adaptive, Efficient, and Robust Visual Localization System
abstract
In this article, we introduceiLoc, an innovative visual localization system designed to enhance the autonomy and adaptability of robotic agents in long-term and large-scale applications.iLocspecializes in: 1) extracting stable and consistent descriptors for place recognition, unaffected by changes in viewpoint and illumination; 2) performing swift and precise global relocalization to establish a robot's position within a large and complex environment; and 3) generating real-time tracking trajectories aligned with reference maps, ensuring continual orientation within known spaces. Distinctively,iLocincorporates a transformer-based learning module and an attention-enhanced recognition approach, enabling it to adapt to diverse environmental and viewpoint conditions.iLocleverages a coarse-to-fine global feature matching technique for enhanced localization and integrates robust state estimation combining visual odometry and loop closures through local refinement and pose graph optimization.iLocdemonstrates remarkable proficiency in place recognition, achieving localization over distances of up to 2 km within 0.5 s with average accuracy at 1 m. It maintains stable localization accuracy, even under variable conditions. Its versatile design allows integration across various environments, significantly broadening the scope of universal localization capabilities in robotics.iLocrepresents a substantial step forward in visual-based localization systems, delivering unparalleled speed and accuracy in place recognition. Its ability to adapt and respond to diverse environmental stimuli marks it as a crucial tool in advancing the field of robotic localization.
Peng Yin 0001, Jing Wang 0193, Ruohai Ge, Jianmin Ji, Yeping Hu, Huaping Liu 0001, Jianda Han
IEEE Trans. Robotics5
2024 SDAC: A Multimodal Synthetic Dataset for Anomaly and Corner Case Detection in Autonomous Driving
abstract
Nowadays, closed-set perception methods for autonomous driving perform well on datasets containing normal scenes. However, they still struggle to handle anomalies in the real world, such as unknown objects that have never been seen while training. The lack of public datasets to evaluate the model performance on anomaly and corner cases has hindered the development of reliable autonomous driving systems. Therefore, we propose a multimodal Synthetic Dataset for Anomaly and Corner case detection, called SDAC, which encompasses anomalies captured from multi-view cameras and the LiDAR sensor, providing a rich set of annotations for multiple mainstream perception tasks. SDAC is the first public dataset for autonomous driving that categorizes anomalies into object, scene, and scenario levels, allowing the evaluation under different anomalous conditions. Experiments show that closed-set models suffer significant performance drops on anomaly subsets in SDAC. Existing anomaly detection methods fail to achieve satisfactory performance, suggesting that anomaly detection remains a challenging problem. We anticipate that our SDAC dataset could foster the development of safe and reliable systems for autonomous driving.
Yu Zhang 0086, Yingqing Xia, Yanyong Zhang, Jianmin Ji
AAAI5
2024 BEVoxSeg: BEV-Voxel Representation for Fast and Accurate Camera-Based 3D Segmentation
abstract
Recent research has demonstrated the advantages of Bird’s-eye-view (BEV) representation in the field of 3D perception. However, due to the lack of height information, BEV representation alone is insufficient to accurately reconstruct the complete surrounding 3D scene. On the other hand, voxel representation excels in describing 3D structures, but their memory and computational cost pose challenges for fast inference. To tackle these limitations, we propose an innovative method dubbed BEVoxSeg, which leverages the computational efficiency of BEV methods while incorporating essential geometric information from voxel features. By combining the advantages from both representations, our approach achieved state-of-the-art results for LiDAR semantic segmentation on nuScenes and demonstrated a superior performance in the occupancy prediction tasks on Occ3D-nuScenes dataset.
Jianmin Ji, Yanyong Zhang
ICASSP4
2024 PathRL: An End-to-End Path Generation Method for Collision Avoidance via Deep Reinforcement Learning
abstract
Robot navigation using deep reinforcement learning (DRL) has shown great potential in improving the performance of mobile robots. Nevertheless, most existing DRL-based navigation methods primarily focus on training a policy that directly commands the robot with low-level controls, like linear and angular velocities, which leads to unstable speeds and unsmooth trajectories of the robot during the long-term execution. An alternative method is to train a DRL policy that outputs the navigation path directly. Then the robot can follow the generated path smoothly using sophisticated velocity-planning and path-following controllers, whose parameters are specified according to the hardware platform. However, two roadblocks arise for training a DRL policy that outputs paths: (1) The action space for potential paths often involves higher dimensions comparing to low-level commands, which increases the difficulties of training; (2) It takes multiple time steps to track a path instead of a single time step, which requires the path to predicate the interactions of the robot w.r.t. the dynamic environment in multiple time steps. This, in turn, amplifies the challenges associated with training. In response to these challenges, we propose PathRL, a novel DRL method that trains the policy to generate the navigation path for the robot. Specifically, we employ specific action space discretization techniques and tailored state space representation methods to address the associated challenges. Curriculum learning is employed to expedite the training process, while the reward function also takes into account the smooth transition between adjacent paths. In our experiments, PathRL achieves better success rates and reduces angular rotation variability compared to other DRL navigation methods, facilitating stable and smooth robot movement. We demonstrate the competitive edge of PathRL in both real-world scenarios and multiple challenging simulation environments.
Wenhao Yu 0010, Jie Peng 0002, Quecheng Qiu, Jianmin Ji
ICRA6
2024 OCC-VO: Dense Mapping via 3D Occupancy-Based Visual Odometry for Autonomous Driving
abstract
Visual Odometry (VO) plays a pivotal role in autonomous systems, with a principal challenge being the lack of depth information in camera images. This paper introduces OCC-VO, a novel framework that capitalizes on recent advances in deep learning to transform 2D camera images into 3D semantic occupancy, thereby circumventing the traditional need for concurrent estimation of ego poses and landmark locations. Within this framework, we utilize the TPV-Former to convert surround view cameras’ images into 3D semantic occupancy. Addressing the challenges presented by this transformation, we have specifically tailored a pose estimation and mapping algorithm that incorporates Semantic Label Filter, Dynamic Object Filter, and finally, utilizes Voxel PFilter for maintaining a consistent global semantic map. Evaluations on the Occ3D-nuScenes not only showcase a 20.6% improvement in Success Ratio and a 29.6% enhancement in trajectory accuracy against ORB-SLAM3, but also emphasize our ability to construct a comprehensive map. Our implementation is open-sourced and available at: https://github.com/USTCLH/OCC-VO.
Yifan Duan, Jianmin Ji, Yanyong Zhang
ICRA5
2024 CalibFormer: A Transformer-based Automatic LiDAR-Camera Calibration Network
abstract
The fusion of LiDARs and cameras has been increasingly adopted in autonomous driving for perception tasks. The performance of such fusion-based algorithms largely depends on the accuracy of sensor calibration, which is challenging due to the difficulty of identifying common features across different data modalities. Previously, many calibration methods involved specific targets and/or manual intervention, which has proven to be cumbersome and costly. Learning-based online calibration methods have been proposed, but their performance is barely satisfactory in most cases. These methods usually suffer from issues such as sparse feature maps, unreliable cross-modality association, inaccurate calibration parameter regression, etc. In this paper, to address these issues, we propose CalibFormer, an end-to-end network for automatic LiDAR-camera calibration. We aggregate multiple layers of camera and LiDAR image features to achieve high-resolution representations. A multi-head correlation module is utilized to identify correlations between features more accurately. Lastly, we employ transformer architectures to estimate accurate calibration parameters from the correlation information. Our method achieved a mean translation error of 0.8751cm and a mean rotation error of 0.0562° on the KITTI dataset, surpassing existing state-of-the-art methods and demonstrating strong robustness, accuracy, and generalization capabilities.
Yao Li 0016, Chengzhen Meng, Jianmin Ji, Yanyong Zhang
ICRA5
2024 NaviFormer: A Data-Driven Robot Navigation Approach via Sequence Modeling and Path Planning with Safety Verification
abstract
Reinforcement learning has shown great potential in improving the performance of robot navigation. In response to the increasing deployments of mobile robots within various scenarios, a data-driven paradigm of navigation approach with safety verification is preferred where one can train RL algorithms with large amounts of prior data, keep learning continuously, and ensure safe navigation in applications. Conventional end-to-end reinforcement learning navigation paradigms have encountered multiple challenges in meeting these demands. In this work, we introduce a novel robot navigation approach termed NaviFormer. This approach handles navigation tasks based on sequence modeling to obtain the data-driven ability. It also integrates rule-based verification for safety insurance. We conduct a series of experiments to validate the data-driven ability of our approach and to compare it with existing navigation methods. We also perform quantitative tests on a real-world robot platform, TurtleBot. The experimental results show our method’s outstanding data-driven ability and highlight its superior arrival rate and generalization compared to other state-of-the-art methods like the PPO-based navigation method.
Ziyang Feng, Quecheng Qiu, Yu'an Chen, Bei Hua, Jianmin Ji
ICRA6
2024 LDP: A Local Diffusion Planner for Efficient Robot Navigation and Collision Avoidance
abstract
The conditional diffusion model has been demonstrated as an efficient tool for learning robot policies, owing to its advancement to accurately model the conditional distribution of policies. The intricate nature of real-world scenarios, characterized by dynamic obstacles and maze-like structures, underscores the complexity of robot local navigation decision-making as a conditional distribution problem. Nevertheless, leveraging the diffusion model for robot local navigation is not trivial and encounters several under-explored challenges: (1) Data Urgency The complex conditional distribution in local navigation needs training data to include diverse policy in diverse real-world scenarios; (2) Myopic Observation Due to the diversity of the perception scenarios, diffusion decisions based on the local perspective of robots may prove suboptimal for completing the entire task, as they often lack foresight. In certain scenarios requiring detours, the robot may become trapped. To address these issues, our approach begins with an exploration of a diverse data generation mechanism that encompasses multiple agents exhibiting distinct preferences through target selection informed by integrated global-local insights. Then, based on this diverse training data, a diffusion agent is obtained, capable of excellent collision avoidance in diverse scenarios. Subsequently, we augment our Local Diffusion Planner, also known as LDP by incorporating global observations in a lightweight manner. This enhancement broadens the observational scope of LDP, effectively mitigating the risk of becoming ensnared in local optima and promoting more robust navigational decisions. Our experimental results demonstrated that the LDP outperforms other baseline algorithms in navigation performance, exhibiting enhanced robustness across diverse scenarios with different policy preferences and superior generalization capabilities for unseen scenarios. Moreover, we highlighted the competitive advantage of the LDP within real-world settings.
Wenhao Yu 0010, Jie Peng 0002, Junrui Zhang 0012, Yifan Duan, Jianmin Ji, Yanyong Zhang
IROS6
2024 CRPlace: Camera-Radar Fusion with BEV Representation for Place Recognition
abstract
The integration of complementary characteristics from camera and radar data has emerged as an effective approach in 3D object detection. However, such fusion-based methods remain unexplored for place recognition, an equally important task for autonomous systems. Given that place recognition relies on the similarity between a query scene and the corresponding candidate scene, the stationary background of a scene is expected to play a crucial role in the task. As such, current well-designed camera-radar fusion methods for 3D object detection can hardly take effect in place recognition because they mainly focus on dynamic foreground objects. In this paper, a background-attentive camera-radar fusion-based method, named CRPlace, is proposed to generate background-attentive global descriptors from multi-view images and radar point clouds for accurate place recognition. To extract stationary background features effectively, we design an adaptive module that generates the background-attentive mask by utilizing the camera BEV feature and radar dynamic points. With the guidance of a background mask, we devise a bidirectional cross-attention-based spatial fusion strategy to facilitate comprehensive spatial interaction between the background information of the camera BEV feature and the radar BEV feature. As the first camera-radar fusion-based place recognition network, CRPlace has been evaluated thoroughly on the nuScenes dataset. The results show that our algorithm outperforms a variety of baseline methods across a comprehensive set of metrics (recall@1 reaches 91.2%).
Shaowei Fu, Yifan Duan, Yao Li 0016, Chengzhen Meng, Jianmin Ji, Yanyong Zhang
IROS6
2024 MM-Gaussian: 3D Gaussian-based Multi-modal Fusion for Localization and Reconstruction in Unbounded Scenes
abstract
Localization and mapping are critical tasks for various applications such as autonomous vehicles and robotics. The challenges posed by outdoor environments present particular complexities due to their unbounded characteristics. In this work, we present MM-Gaussian, a LiDAR-camera multimodal fusion system for localization and mapping in unbounded scenes. Our approach is inspired by the recently developed 3D Gaussians, which demonstrate remarkable capabilities in achieving high rendering quality and fast rendering speed. Specifically, our system fully utilizes the geometric structure information provided by solid-state LiDAR to address the problem of inaccurate depth encountered when relying solely on visual solutions in unbounded, outdoor scenarios. Additionally, we utilize 3D Gaussian point clouds, with the assistance of pixel-level gradient descent, to fully exploit the color information in photos, thereby achieving realistic rendering effects. To further bolster the robustness of our system, we designed a relocalization module, which assists in returning to the correct trajectory in the event of a localization failure. Experiments conducted in multiple scenarios demonstrate the effectiveness of our method.
Yifan Duan, Yu Sheng, Jianmin Ji, Yanyong Zhang
IROS5
2024 FARFusion V2: A Geometry-based Radar-Camera Fusion Method on the Ground for Roadside Far-Range 3D Object Detection
abstract
Fusing the data of millimeter-wave Radar sensors and high-definition cameras has emerged as a viable approach to achieving precise 3D object detection for roadside traffic surveillance. For roadside perception systems, earlier studies have pointed out that it is better to perform the fusion on the 2D image plane than on the BEV plane (which is popular for on-car perception systems), especially when the perception range is large (e.g., >150m). Image-plane fusion requires critical transformations, like perspective projection from the Radar's BEV to the camera's 2D plane and reverse IPM. However, real-world issues like uneven terrain and sensor movement degrade these transformations' precision, impacting fusion effectiveness. To alleviate these issues, we propose a geometry-based Radar-camera fusion method on the ground, namely FARFusion V2. Specifically, we extend the ground-plane assumption in FARFusion[20] to support arbitrary shapes by formulating the ground height as an implicit representation based on geometric transformations. By incorporating the ground information, we can enhance Radar data with target height measurements. Consequently, we can thus project the enhanced Radar data onto the 2D plane to obtain more accurate depth information, thereby assisting the IPM process. A real-time parameterized transformation parameters estimation module is further introduced to refine the view transformation processes. Moreover, considering various measurement noises across these two sensors, we introduce an uncertainty-based depth fusion strategy into the 2D fusion process to maximize the probability of obtaining the optimal depth value. Extensive experiments are conducted on our collected roadside OWL benchmark, demonstrating the excellent localization capacity of FARFusion V2 in far-range scenarios. Our method achieves an average location accuracy of 0.771m when we extend the detection range up to 500m.
Yao Li 0016, Jiajun Deng, Yingjie Wang 0004, Xiaomeng Chu, Jianmin Ji, Yanyong Zhang
ACM Multimedia6
2024 Map++: Towards User-Participatory Visual SLAM Systems with Efficient Map Expansion and Sharing
abstract
Constructing precise 3D maps is crucial for the development of future map-based systems such as self-driving and navigation. However, generating these maps in complex environments, such as multi-level parking garages or shopping malls, remains a formidable challenge. In this paper, we introduce a participatory sensing approach that delegates map-building tasks to map users, thereby enabling cost-effective and continuous data collection. The proposed method harnesses the collective efforts of users, facilitating the expansion and ongoing update of the maps as the environment evolves.
Hanqi Zhu, Yifan Duan, Wuyang Zhang, Longfei Shangguan, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang
MobiCom7
2024 Enabling Tensor Language Model to Assist in Generating High-Performance Tensor Programs for Deep Learning
Yi Zhai 0005, Keyu Pan, Renwei Zhang, Shuo Liu 0019, Zichun Ye, Jianmin Ji, Jie Zhao 0002, Yu Zhang 0086, Yanyong Zhang
OSDI8
2023 TLP: A Deep Learning-Based Cost Model for Tensor Program Tuning
abstract
Tensor program tuning is a non-convex objective optimization problem, to which search-based approaches have proven to be effective. At the core of the search-based approaches lies the design of the cost model. Though deep learning-based cost models perform significantly better than other methods, they still fall short and suffer from the following problems. First, their feature extraction heavily relies on expert-level domain knowledge in hardware architectures. Even so, the extracted features are often unsatisfactory and require separate considerations for CPUs and GPUs. Second, a cost model trained on one hardware platform usually performs poorly on another, a problem we call cross-hardware unavailability.
Yi Zhai 0005, Yu Zhang 0086, Shuo Liu 0019, Xiaomeng Chu, Jie Peng 0002, Jianmin Ji, Yanyong Zhang
ASPLOS (2)6
2023 Bi-LRFusion: Bi-Directional LiDAR-Radar Fusion for 3D Dynamic Object Detection
abstract
LiDAR and Radar are two complementary sensing approaches in that LiDAR specializes in capturing an object's 3D shape while Radar provides longer detection ranges as well as velocity hints. Though seemingly natural, how to efficiently combine them for improved feature representation is still unclear. The main challenge arises from that Radar data are extremely sparse and lack height information. Therefore, directly integrating Radar features into LiDAR-centric detection networks is not optimal. In this work, we introduce a bi-directional LiDAR-Radar fusion framework, termed Bi-LRFusion, to tackle the challenges and improve 3D detection for dynamic objects. Technically, Bi-LRFusion involves two steps: first, it enriches Radar's local features by learning important details from the LiDAR branch to alleviate the problems caused by the absence of height information and extreme sparsity; second, it combines LiDAR features with the enhanced Radar features in a unified bird's-eye-view representation. We conduct extensive experiments on nuScenes and ORR datasets, and show that our Bi-LRFusion achieves state-of-the-art performance for detecting dynamic objects. Notably, Radar data in these two datasets have different formats, which demonstrates the generalizability of our method. Codes will be published.
Yingjie Wang 0005, Jiajun Deng, Yao Li 0016, Jinshui Hu, Cong Liu 0006, Yu Zhang 0086, Jianmin Ji, Wanli Ouyang, Yanyong Zhang
CVPR7
2023 P3O: Transferring Visual Representations for Reinforcement Learning via Prompting
abstract
It is important for deep reinforcement learning (DRL) algorithms to transfer their learned policies to new environments that have different visual inputs. In this paper, we introduce Prompt based Proximal Policy Optimization (P3O), a three-stage DRL algorithm that transfers visual representations from a target to a source environment by applying prompting. The process of P3O consists of three stages: pre-training, prompting, and predicting. In particular, we specify a prompt-transformer for representation conversion and propose a two-step training process to train the prompt-transformer for the target environment, while the rest of the DRL pipeline remains unchanged. We implement P3O and evaluate it on the OpenAI CarRacing video game. The experimental results show that P3O outperforms the state-of-the-art visual transferring schemes. In particular, P3O allows the learned policies to perform well in environments with different visual inputs, which is much more effective than retraining the policies in these environments.
Guoliang You, Xiaomeng Chu, Yifan Duan, Jie Peng 0002, Jianmin Ji, Yu Zhang 0086, Yanyong Zhang
ICME5
2023 Learning Complicated Navigation Skills from Limited Experience via Augmenting Offline Datasets
abstract
Deep reinforcement learning has yielded remarkable results in the field of robot navigation. Most of the existing RL-based methods tend to train the navigation policy with (1) a simulation environment in which the agent interacts and collects experience iteratively, (2) a shaped reward function that defines how a robot should reach the goal while avoiding collisions with obstacles. However, these methods suffer from several challenges, including the difficulty of generalizing the trained model to real-world scenarios, sub-optimal risk due to reward-shaping, and the inefficiency of data utilization. In this paper, we address these challenges by introducing the State & Goal-Relabel techniques based on Hindsight Experience Replay (HER), enabling the robot not only to learn success from failure but also to learn more complicated navigation skills from simple tasks. Instead of training only with limited real experiences, our approach aims to generate pseudo-experiences by relabeling both the local observation and target pose, cleverly improving the scale and quality of dataset, and using target-driven style to train a model with solid generalization ability. With experiences collected just in simple environments, our approach outperforms or performs comparably to classical and other learning-based methods, and generalizes well in physical environments.
Yu'an Chen, Jianmin Ji
ICTAI3
2023 Reinforcement Learning for Robot Navigation with Adaptive Forward Simulation Time (AFST) in a Semi-Markov Model
abstract
Deep reinforcement learning (DRL) algorithms have proven effective in robot navigation, especially in unknown environments, by directly mapping perception inputs into robot control commands. However, most existing methods ignore the local minimum problem in navigation and thereby cannot handle complex unknown environments. In this paper, we propose the first DRL-based navigation method modeled by a semi-Markov decision process (SMDP) with continuous action space, named Adaptive Forward Simulation Time (AFST), to overcome this problem. Specifically, we reduce the dimensions of the action space and improve the distributed proximal policy optimization (DPPO) algorithm for the specified SMDP problem by modifying its GAE to better estimate the policy gradient in SMDPs. Experiments in various unknown environments demonstrate the effectiveness of AFST.
Yu'an Chen, Ruosong Ye, Ziyang Tao, Hongjian Liu, Guangda Chen, Jie Peng 0002, Jun Ma 0034, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang
IROS9
2023 A²CoST: An ASP-based Avoidable Collision Scenario Testbench for Autonomous Vehicles
abstract
This paper addresses the challenge of generating safety-critical scenarios with multiple adversarial vehicles for testing autonomous vehicles. Such scenarios must be plausible and collision-avoidable while resulting in a collision with the vehicle-under-test. However, the tremendous number of scenarios and the low ratio of plausible scenarios makes previous methods squander primary resources on implausible scenarios, degenerating their efficiency. We propose a two-stage framework called the ASP-based Avoidable Collision Scenario Testbench (A²CoST) to overcome this obstacle and improve efficiency. In the former stage, we apply Answer Set Programming (ASP) for generating plausible logical scenarios. In the latter stage, we use a search algorithm to refine logical scenarios into safety-critical concrete scenarios. We also compute collision-free trajectories in these concrete scenarios while the vehicle-under-test fails to avoid the collision. We empirically show the A²CoST significantly decreases the time consumption for simple scenarios while still effectively generating complex critical scenarios. The comparison with real-world traffic data further demonstrates the value of A²CoST in generating plausible scenarios. The source codes of our method and the baselines are opened at https://github.com/Autonomous-Driving-Safety-Project/AACoST.
Ruolin Wang, Yuejiao Xu, Jie Peng 0002, Jianmin Ji
KR4
2023 CluB: Cluster Meets BEV for LiDAR-Based 3D Object Detection
abstract
Currently, LiDAR-based 3D detectors are broadly categorized into two groups, namely, BEV-based detectors and cluster-based detectors. BEV-based detectors capture the contextual information from the Bird's Eye View (BEV) and fill their center voxels via feature diffusion with a stack of convolution layers, which, however, weakens the capability of presenting an object with the center point. On the other hand, cluster-based detectors exploit the voting mechanism and aggregate the foreground points into object-centric clusters for further prediction. In this paper, we explore how to effectively combine these two complementary representations into a unified framework. Specifically, we propose a new 3D object detection framework, referred to as CluB, which incorporates an auxiliary cluster-based branch into the BEV-based detector by enriching the object representation at both feature and query levels. Technically, CluB is comprised of two steps. First, we construct a cluster feature diffusion module to establish the association between cluster features and BEV features in a subtle and adaptive fashion. Based on that, an imitation loss is introduced to distill object-centric knowledge from the cluster features to the BEV features. Second, we design a cluster query generation module to leverage the voting centers directly from the cluster branch, thus enriching the diversity of object queries. Meanwhile, a direction loss is employed to encourage a more accurate voting center for each cluster. Extensive experiments are conducted on Waymo and nuScenes datasets, and our CluB achieves state-of-the-art performance on both benchmarks.
Yingjie Wang 0005, Jiajun Deng, Yuenan Hou, Yao Li 0016, Yu Zhang 0086, Jianmin Ji, Wanli Ouyang, Yanyong Zhang
NeurIPS6
2023 Multi-Modal 3D Object Detection in Autonomous Driving: A Survey
Yingjie Wang 0005, Qiuyu Mao, Hanqi Zhu, Jiajun Deng, Yu Zhang 0086, Jianmin Ji, Houqiang Li, Yanyong Zhang
Int. J. Comput. Vis.6
2023 TrajMatch: Toward Automatic Spatio-Temporal Calibration for Roadside LiDARs Through Trajectory Matching
abstract
Recently, deploying sensors such as LiDARs on the roadside to monitor the passing traffic and assist autonomous vehicle perception has become popular. However, unlike autonomous vehicle systems, roadside sensor systems involve sensors from different subsystems, resulting in a lack of synchronization in both time and space between the sensors. Calibration is a critical technology that enables the central server to fuse data generated by different location infrastructures, which vastly improves sensing range and detection robustness. Regrettably, existing calibration algorithms frequently assume that LiDARs have significant overlap or that temporal calibration has already been achieved. However, since these assumptions do not always hold in real-world scenarios, the calibration results obtained from existing algorithms are frequently unsatisfactory. In this paper, we propose TrajMatch - the first system that can automatically calibrate roadside LiDARs in both time and space. The main idea is to automatically calibrate the sensors based on the result of the detection/tracking task, rather than relying on extracting special features. Furthermore, we propose a novel mechanism for evaluating calibration parameters that align with our algorithm, and we demonstrate its effectiveness through experiments. This mechanism can also guide parameter iterations for multiple calibrations, further enhancing the accuracy and efficiency of our calibration method. Finally, to evaluate the performance of TrajMatch, we collected two datasets, one simulated dataset LiDARnet-sim 1.0 and one real-world dataset. The experimental results show that TrajMatch can achieve a spatial calibration error of less than$10cm$and a temporal calibration error of less than$1.5ms$.
Haojie Ren, Sha Zhang 0002, Sugang Li, Yao Li 0016, Xinchen Li, Jianmin Ji, Yu Zhang 0086, Yanyong Zhang
IEEE Trans. Intell. Transp. Syst.6
2023 VPFNet: Improving 3D Object Detection With Virtual Point Based LiDAR and Stereo Data Fusion
abstract
It has been well recognized that fusing the complementary information from depth-aware LiDAR point clouds and semantic-rich stereo images would benefit 3D object detection. Nevertheless, it is non-trivial to explore the inherently unnatural interaction between sparse 3D points and dense 2D pixels. To ease this difficulty, the recent approaches generally project the 3D points onto the 2D image plane to sample the image data and then aggregate the data at the points. However, these approaches often suffer from the mismatch between the resolution of point clouds and RGB images, leading to sub-optimal performance. Specifically, taking the sparse points as the multi-modal data aggregation locations causes severe information loss for high-resolution images, which in turn undermines the effectiveness of multi-sensor fusion. In this paper, we presentVPFNet—a new architecture that cleverly aligns and aggregates the point cloud and image data at the “virtual” points. Particularly, with their density lying between that of the 3D points and 2D pixels, the virtual points can nicely bridge the resolution gap between the two sensors, and thus preserve more information for processing. Moreover, we also investigate the data augmentation techniques that can be applied to both point clouds and RGB images, as the data augmentation has made non-negligible contribution towards 3D object detectors to date. We have conducted extensive experiments on KITTI dataset, and have observed good performance compared to the state-of-the-art methods. Remarkably, ourVPFNetachieves 83.21% moderate$AP_{3D}$and 91.86% moderate$AP_{BEV}$on the KITTI test set. The network design also takes computation efficiency into consideration – we can achieve a FPS of 15 on a single NVIDIA RTX 2080Ti GPU.
Hanqi Zhu, Jiajun Deng, Yu Zhang 0086, Jianmin Ji, Qiuyu Mao, Houqiang Li, Yanyong Zhang
IEEE Trans. Multim.4
2022 Transferring Knowledge from Structure-aware Self-attention Language Model to Sequence-to-Sequence Semantic Parsing
abstract
Semantic parsing considers the task of mapping a natural language sentence into a target formal representation, where various sophisticated sequence-to-sequence (seq2seq) models have been applied with promising results. Generally, these target representations follow a syntax formalism that limits permitted forms. However, it is neither easy nor flexible to explicitly integrate this syntax formalism into a neural seq2seq model. In this paper, we present a structure-aware self-attention language model to capture structural information of target representations and propose a knowledge distillation based approach to incorporating the target language model into a seq2seq model, where grammar rules or sketches are not required in the training process. An ablation study shows that the proposed language model can notably improve the performance of the baseline model. The experiments show that our method achieves new state-of-the-art performance among neural approaches on four semantic parsing (ATIS, GEO) and Python code generation (Django, CoNaLa) tasks.
Jianmin Ji
COLING2
2022 Learning to Socially Navigate in Pedestrian-rich Environments with Interaction Capacity
abstract
Existing navigation policies for autonomous robots tend to focus on collision avoidance while ignoring human-robot interactions in social life. For instance, robots can pass along the corridor safer and easier if pedestrians notice them. Sounds have been considered as an efficient way to attract the attention of pedestrians, which can alleviate the freezing robot problem. In this work, we present a new deep reinforcement learning (DRL) based social navigation approach for autonomous robots to move in pedestrian-rich environments with interaction capacity. Most existing DRL based methods intend to train a general policy that outputs both navigation actions, i.e., expected robot's linear and angular velocities, and interaction actions, i.e., the beep action, in the context of reinforcement learning. Different from these methods, we intend to train the policy via both supervised learning and reinforcement learning. In specific, we first train an interaction policy in the context of supervised learning, which provides a better understanding of the social situation, then we use this interaction policy to train the navigation policy via multiple reinforcement learning algorithms. We evaluate our approach in various simulation environments and compare it to other methods. The experimental results show that our approach outperforms others in terms of the success rate. We also deploy the trained policy on a real-world robot, which shows a nice performance in crowded environments.
Quecheng Qiu, Shunyi Yao, Jing Wang 0193, Jun Ma 0034, Guangda Chen, Jianmin Ji
ICRA6
2022 PFilter: Building Persistent Maps through Feature Filtering for Fast and Accurate LiDAR-based SLAM
abstract
Simultaneous localization and mapping (SLAM) based on laser sensors has been widely adopted by mobile robots and autonomous vehicles. These SLAM systems are required to support accurate localization with limited computational resources. In particular, point cloud registration, i.e., the process of matching and aligning multiple LiDAR scans collected at multiple locations in a global coordinate framework, has been deemed as the bottleneck step in SLAM. In this paper, we propose a feature filtering algorithm, PFilter, that can filter out invalid features and can thus greatly alleviate this bottleneck. Meanwhile, the overall registration accuracy is also improved due to the carefully curated feature points. We integrate PFilter into the well-established scan-to-map LiDAR odometry framework, F-LOAM, and evaluate its performance on the KITTI dataset. The experimental results show that PFilter can remove about 48.4% of the points in the local feature map and reduce feature points in scan by 19.3% on average, which save 20.9% processing time per frame. In the mean time, we improve the accuracy by 9.4%.
Yifan Duan, Jie Peng 0002, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang
IROS4
2021 Towards an Online RRT-based Path Planning Algorithm for Ackermann-steering Vehicles
abstract
It is challenging to develop an online path planning algorithm for Ackermann-steering vehicles to find collision-free and kinematically-feasible paths, that is efficient for dense environments, adaptable to various environments, and suitable for environments with narrow passages. In this paper, we propose a kinematically constrained RRT-based path planning algorithm integrating with a trajectory parameter space (TP-space) with three novel improvements to meet the above requirements. In specific, we introduce a new way to choose candidate nodes to expand the tree for an RRT-based algorithm, which can significantly increase the success rate of the expansion and improve the efficiency of the algorithm. We also introduce a procedure to incrementally adjust the step size for the expansion, which enables the algorithm to automatically adapt to various environments. At last, we integrate rapidly-exploring random vines (RRV) with a TP-space to handle kinematic constraints and improve the performance of the algorithm to expand the tree through a narrow passage. We also prove that the algorithm is probabilistic complete and asymptotically near-optimal. An ablation study shows that all three improvements can notably improve the performance of the RRT-based path planning algorithm. We also evaluate the algorithm in various environments. The experimental results show that our algorithm achieves competitive performance compared with the state-of-the-art. The source code is available at https://github.com/PengJieb/fastbkrrt.
Jie Peng 0002, Yu'an Chen, Yifan Duan, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang
ICRA5
2021 DRQN-based 3D Obstacle Avoidance with a Limited Field of View
abstract
In this paper, we propose a map-based end-to-end DRL approach for three-dimensional (3D) obstacle avoidance in a partially observed environment, which is applied to achieve autonomous navigation for an indoor mobile robot using a depth camera with a narrow field of view. We first train a neural network with LSTM units in a 3D simulator of mobile robots to approximate the Q-value function in double DRQN. We also use a curriculum learning strategy to accelerate and stabilize the training process. Then we deploy the trained model to a real robot to perform 3D obstacle avoidance in its navigation. We evaluate the proposed approach both in the simulated environment and on a robot in the real world. The experimental results show that the approach is efficient and easy to be deployed, and it performs well for 3D obstacle avoidance with a narrow observation angle, which outperforms other existing DRL-based models by 15.5% on success rate.
Yu'an Chen, Guangda Chen, Lifan Pan, Jun Ma 0034, Yu Zhang 0086, Yanyong Zhang, Jianmin Ji
IROS7
2021 Crowd-Aware Robot Navigation for Pedestrians with Multiple Collision Avoidance Strategies via Map-based Deep Reinforcement Learning
abstract
It is challenging for a mobile robot to navigate through human crowds. Existing approaches usually assume that pedestrians follow a predefined collision avoidance strategy, like social force model (SFM) or optimal reciprocal collision avoidance (ORCA). However, their performances commonly need to be further improved for practical applications, where pedestrians follow multiple different collision avoidance strategies. In this paper, we propose a map-based deep reinforcement learning approach for crowd-aware robot navigation with various pedestrians. We use the sensor map to represent the environmental information around the robot, including its shape and observable appearances of obstacles. We also introduce the pedestrian map that specifies the movements of pedestrians around the robot. By applying both maps as inputs of the neural network, we show that a navigation policy can be trained to better interact with pedestrians following different collision avoidance strategies. We evaluate our approach under multiple scenarios both in the simulator and on an actual robot. The results show that our approach allows the robot to successfully interact with various pedestrians and outperforms compared methods in terms of the success rate.
Shunyi Yao, Guangda Chen, Quecheng Qiu, Jun Ma 0034, Jianmin Ji
IROS6
2021 Neighbor-Vote: Improving Monocular 3D Object Detection through Neighbor Distance Voting
abstract
As cameras are increasingly deployed in new application domains such as autonomous driving, performing 3D object detection on monocular images becomes an important task for visual scene understanding. Recent advances on monocular 3D object detection mainly rely on the "pseudo-LiDAR'' generation, which performs monocular depth estimation and lifts the 2D pixels to pseudo 3D points. However, depth estimation from monocular images, due to its poor accuracy, leads to inevitable position shift of pseudo-LiDAR points within the object. Therefore, the predicted bounding boxes may suffer from inaccurate location and deformed shape. In this paper, we present a novel neighbor-voting method that incorporates neighbor predictions to ameliorate object detection from severely deformed pseudo-LiDAR point clouds. Specifically, each feature point around the object forms their own predictions, and then the "consensus'' is achieved through voting. In this way, we can effectively combine the neighbors' predictions with local prediction and achieve more accurate 3D detection. To further enlarge the difference between the foreground region of interest (ROI) pseudo-LiDAR points and the background points, we also encode the ROI prediction scores of 2D foreground pixels into the corresponding pseudo-LiDAR points. We conduct extensive experiments on the KITTI benchmark to validate the merits of our proposed method. Our results on the bird's eye view detection outperform the state-of-the-art performance, especially for the "hard" level detection. The code is available at https://github.com/cxmomo/Neighbor-Vote.
Xiaomeng Chu, Jiajun Deng, Yao Li 0016, Zhenxun Yuan, Yanyong Zhang, Jianmin Ji, Yu Zhang 0086
ACM Multimedia6
2021 Combining Improvements for Exploiting Dependency Trees in Neural Semantic Parsing
Defeng Xie, Jianmin Ji
PRICAI (2)2
2021 Distributed Reinforcement Learning with Self-Play in Parameterized Action Space
abstract
Self-play has been shown to be effective to provide a proper training curriculum for a reinforcement learning agent in competitive multi-agent environments without direct supervision. However, its performance is still unstable for problems with sparse rewards, e.g., the scoring task with goalkeeper for robots in RoboCup soccer. It is challenging to solve these tasks in reinforcement learning, especially for those that require combining high-level actions with flexible control. To address these challenges, we introduce a distributed self-play training framework for an extended proximal policy optimization (PPO) algorithm that learns to act in parameterized action space and plays against a group of opponents, i.e., a league. Experiments on the domain of simulated RoboCup soccer show that, the approach is effective and learns more robust policies against various opponents compared to existing reinforcement learning methods. A demonstration video is available online at https://youtu.be/BuLli1vND4.
Jun Ma 0034, Shunyi Yao, Guangda Chen, Jiakai Song, Jianmin Ji
SMC5
2020 Multi-Robot Collision Avoidance with Map-based Deep Reinforcement Learning
abstract
Multi-robot collision avoidance in a communication-free environment is one of the key issues for mobile robotics and autonomous driving. In this paper, we propose a map-based deep reinforcement learning (DRL) approach for collision avoidance of multiple robots, where robots do not communicate with each other and only sense other robots' positions and the obstacles around them. We use the egocentric grid map of a robot to represent the environmental information around it, which can be easily generated by using multiple sensors or sensor fusion. The learned policy generated from the DRL model directly maps 3 frames of egocentric grid maps and the robot's relative local goal positions into low-level robot control commands. We first train a convolutional neural network for the navigation policy in a simulator of multiple mobile robots using proximal policy optimization (PPO). Then we deploy the trained model to real robots to perform collision avoidance in their navigation. We evaluate the approach with various scenarios both in the simulator and on three differential-drive mobile robots in the real world. Both qualitative and quantitative experiments show that our approach is efficient with a high success rate. The demonstration video can be found at https://youtu.be/jcLKlEXuFuk.
Shunyi Yao, Guangda Chen, Lifan Pan, Jun Ma 0034, Jianmin Ji
ICTAI5
2020 Lightweight Map-Enhanced 3D Object Detection and Tracking for Autonomous Driving
abstract
3D object detection and tracking are crucial to the real-time and accurate perception of the surrounding environment for autonomous driving. Recent approaches on 3D object detection and tracking have made great progress, thanks to the rapid development of deep learning models. Even though these models have achieved superior performance on specific datasets, the actual self-driving systems still cannot deal with real-world driving situations properly, especially in complicated scenarios like road intersections. With the development of vehicle-infrastructure cooperation technology, scene information such as map is considered to have great potential in alleviating these problems. In this paper, we explore the potential of solving corner cases in real driving scenarios through the cooperation between autonomous vehicles and map information. We propose a holistic approach that integrates and utilizes the map information in system following the tracking-by-detection paradigm. In order to ensure that the use of map information does not bring much overhead to detection and tracking, we propose a representation method for concise information extracted from rich map. We show that our framework can improve the detection and tracking accuracy with mild or no increase of latency. Specifically, in some cases, our results demonstrate a MOTA improvement of nearly 2% .
Shunhong Wang, Yu Zhang 0086, Yanyong Zhang, Jianmin Ji
Internetware5
2019 MRS-VPR: a multi-resolution sampling based global visual place recognition method
abstract
Place recognition and loop closure detection are challenging for long-term visual navigation tasks. SeqSLAM is considered to be one of the most successful approaches to achieve long-term localization under varying environmental conditions and changing viewpoints. SeqSLAM uses a brute-force sequential matching method, which is computationally intensive. In this work, we introduce a multi-resolution sampling-based global visual place recognition method (MRS-VPR), which can significantly improve the matching efficiency and accuracy in sequential matching. The novelty of this method lies in the coarse-to-fine searching pipeline and a particle filter-based global sampling scheme, that can balance the matching efficiency and accuracy in the long-term navigation task. Moreover, our model works much better than SeqSLAM when the testing sequence is over a much smaller time scale than the reference sequence. Our experiments demonstrate that MRSVPR is efficient in locating short temporary trajectories within long-term reference ones without compromising on the accuracy compared to SeqSLAM.
Peng Yin 0001, Rangaprasad Arun Srivatsan, Xueqian Li, Hongda Zhang, Lu Li 0018, Zhenzhong Jia, Jianmin Ji
ICRA9
2019 A Multi-Domain Feature Learning Method for Visual Place Recognition
abstract
Visual Place Recognition (VPR) is an important component in both computer vision and robotics applications, thanks to its ability to determine whether a place has been visited and where specifically. A major challenge in VPR is to handle changes of environmental conditions including weather, season and illumination. Most VPR methods try to improve the place recognition performance by ignoring the environmental factors, leading to decreased accuracy decreases when environmental conditions change significantly, such as day versus night. To this end, we propose an end-to-end conditional visual place recognition method. Specifically, we introduce the multi-domain feature learning method (MDFL) to capture multiple attribute-descriptions for a given place, and then use a feature detaching module to separate the environmental condition-related features from those that are not. The only label required within this feature learning pipeline is the environmental condition. Evaluation of the proposed method is conducted on the multi-season NORDLAND dataset, and the multi-weather GTAV dataset. Experimental results show that our method improves the feature robustness against variant environmental conditions.
Peng Yin 0001, Xueqian Li, Yingli Li, Rangaprasad Arun Srivatsan, Lu Li 0018, Jianmin Ji
ICRA8
2019 KDSL: a Knowledge-Driven Supervised Learning Framework for Word Sense Disambiguation
abstract
We propose KDSL, a new word sense disambiguation (WSD) framework that utilizes knowledge to automatically generate sense-labeled data for supervised learning. First, from WordNet, we automatically construct a semantic knowledge base called DisDict, which provides refined feature words that highlight the differences among word senses, i.e., synsets. Second, we automatically generate new sense-labeled data by DisDict from unlabeled corpora. Third, these generated data, together with manually labeled data and unlabeled data, are fed to a neural framework conducting supervised and unsupervised learning jointly to model the semantic relations among synsets, feature words and their contexts. The experimental results show that KDSL outperforms several representative state-of-the-art methods on various major benchmarks. Interestingly, it performs relatively well even when manually labeled data is unavailable, thus provides a potential solution for similar tasks in a lack of manual annotations.
Shangfei Wang, Jianmin Ji, Ruili Wang 0001
IJCNN5
2017 Well-founded operators for normal hybrid MKNF knowledge bases
abstract
Abstract Hybrid MKNF knowledge bases have been considered one of the dominant approaches to combining open world ontology languages with closed world rule-based languages. Currently, the only known inference methods are based on the approach of guess-and-verify, while most modern SAT/ASP solvers are built under the DPLL architecture. The central impediment here is that it is not clear what constitutes a constraint propagator, a key component employed in any DPLL-based solver. In this paper, we address this problem by formulating the notion of unfounded sets for non-disjunctive hybrid MKNF knowledge bases, based on which we propose and study two new well-founded operators. We show that by employing a well-founded operator as a constraint propagator, a sound and complete DPLL search engine can be readily defined. We compare our approach with the operator based on the alternating fixpoint construction by Knorr et al. (2011. Artificial Intelligence 175, 9, 1528–1554) and show that, when applied to arbitrary partial partitions, the new well-founded operators not only propagate more truth values but also circumvent the non-converging behavior of the latter. In addition, we study the possibility of simplifying a given hybrid MKNF knowledge base by employing a well-founded operator and show that, out of the two operators proposed in this paper, the weaker one can be applied for this purpose and the stronger one cannot. These observations are useful in implementing a grounder for hybrid MKNF knowledge bases, which can be applied before the computation of MKNF models.
Jianmin Ji, Fangfang Liu 0008, Jia-Huai You
Theory Pract. Log. Program.1
2016 Eliminating Disjunctions in Answer Set Programming by Restricted Unfolding
Jianmin Ji, Hai Wan, Kewen Wang 0001, Zhe Wang 0001
IJCAI1
2015 Splitting a Logic Program Revisited
abstract
Lifschitz and Turner introduced the notion of the splitting set and provided a method to divide a logic program into two parts. They showed that the task of computing the answer sets of the program can be converted into the tasks of computing the answer sets of these parts. However, the empty set and the set of all atoms are the only two splitting sets for many programs, then these programs cannot be divided by the splitting method. In this paper, we extend Lifschitz and Turner's splitting set theorem to allow the program to be split by an arbitrary set of atoms, while some new atoms may be introduced in the process. To illustrate the usefulness of the result, we show that for some typical programs the splitting process is efficient and the program simplification problem can be investigated using the concept of splitting.
Jianmin Ji, Hai Wan, Ziwei Huo, Zhenfeng Yuan
AAAI1
2015 On Elementary Loops and Proper Loops for Disjunctive Logic Programs
abstract
This paper proposes an alternative definition of elementary loops and extends the notion of proper loops for disjunctive logic programs. Different from normal logic programs, the computational complexities of recognizing elementary loops and proper loops for disjunctive programs are coNP-complete. To address this problem, we introduce weaker versions of both elementary loops and proper loops and provide polynomial time algorithms for identifying them respectively. On the other hand, based on the notion of elementary loops, the class of Head-Elementary-loop-Free (HEF) programs was presented, which can be turned into equivalent normal logic programs by shifting head atoms into bodies. However, the problem of recognizing an HEF program is coNP-complete. Then we present a subclass of HEF programs which generalizes the class of Head-Cycle-Free programs and provide a polynomial time algorithm to identify them. At last, some experiments show that both elementary loops and proper loops could be replaced by their weak versions in practice.
Jianmin Ji, Hai Wan, Peng Xiao 0009
AAAI1
2015 Simplifying A Logic Program Using Its Consequences
Jianmin Ji, Hai Wan, Ziwei Huo, Zhenfeng Yuan
IJCAI1
2015 On Forgetting Postulates in Answer Set Programming
Jianmin Ji, Jia-Huai You, Yisong Wang 0004
IJCAI1
2015 Discovering Classes of Strongly Equivalent Logic Programs with Negation as Failure in the Head
abstract
In this paper, we apply Fangzhen Lin’s methodology of computer aided theorem discovery to discover classes of strongly equivalent logic programs with negation as failure in the head. Specifically, with the help of computers, we discover exact conditions that capture the strong equivalence between small sets of rules, which have potential applications in the theory and practice of logic programming. In the experiment, we extend the previous approach to semi-automatically generate plausible conjectures. We also show that it is possible to divide the original problem in simpler cases and combine their solutions in order to obtain the solution of the original problem.
Jianmin Ji
KSEM1
2015 Multi-mode Natural Language Processing for human-robot interaction
abstract
As more and more open knowledge resources become available, it is interesting to explore opportunities of enhancing autonomous agents’ capacities by utilizing the knowledge in these resources, instead of hand-coding knowledge for agents. A major challenge towards this goal lies in the translation o f the open knowledge organized in multiple modes, unstructured or semi-structured, into the internal representations of agents. In this paper we present a set of multi-mode NLP techniques to formalize the open knowledge for autonomous agents. Two case studies are reported in which our robot, equipped with the multi-mode NLP techniques, succeeded in acquiring knowledge from the microwave oven manual and from the open knowledge database, OMICS, and solving problems that could not be solved before the robot acquired the knowledge. Experiments for evaluating the performance of our approach show that our approach is promising.
Jiongkun Xie, Jianmin Ji
Web Intell.3
2014 Elementary Loops Revisited
abstract
The notions of loops and loop formulas play an important role in answer set computation. However, there would be an exponential number of loops in the worst case. Gebser and Schaub characterized a subclass elementary loops and showed that they are sufficient for selecting answer sets from models of a logic program. This paper proposes an alternative definition of elementary loops and identify a subclass of elementary loops, called proper loops. By applying a special form of their loop formulas, proper loops are also sufficient for the SAT-based answer set computation. A polynomial algorithm to recognize a proper loop is given and shows that for certain logic programs, identifying all proper loops of a program is more efficient than that of elementary loops. Furthermore, we prove that, by considering the structure of the positive body-head dependency graph of a program, a large number of loops could be ignored for identifying proper loops. We provide another algorithm for identifying all proper loops of a program. The experiments show that, for certain programs whose dependency graphs consisting of sets of components that are densely connected inside and sparsely connected outside, the new algorithm is more efficient.
Jianmin Ji, Hai Wan, Peng Xiao 0009, Ziwei Huo, Zhanhao Xiao
AAAI1
2014 From Default and Autoepistemic Logics to Disjunctive Answer Set Programs via the Logic of GK
abstract
We show how the pure logic of GK can be embedded into disjunctive logic programming. The translation we present is polynomial, but not modular, and introduces new variables. The result can then be used to compute the extension/expansion semantics of default and autoepistemic logics using disjunctive ASP solvers.
Jianmin Ji, Hannes Strass
ECAI1
2014 A weighted causal theory for acquiring and utilizing open knowledge
Jianmin Ji
Int. J. Approx. Reason.1
2013 From Structured Task Instructions to Robot Task Plans
abstract
Abstract: For the purpose of allowing an autonomous robot to use task instructions for task planning, we present a for-malization for specifying structured task instructions and provide an approach for integrating these instructions with robot’s built-in knowledge to compute plans for open-ended tasks. We have implemented a prototype of the system. We also report a case study of the effectiveness of the approach. 1
Jianmin Ji
KEOD1
2013 Handling Open Knowledge for Service Robots
Jianmin Ji, Zhiqiang Sui, Jiongkun Xie
IJCAI2
2013 Turner's Logic of Universal Causation, Propositional Logic, and Logic Programming
Jianmin Ji, Fangzhen Lin
LPNMR1
2013 Toward open knowledge enabling for human-robot interaction
abstract
This paper presents an effort to enable robots to utilize open-source knowledge resources autonomously for human-robot interaction. The main challenges include how to extract knowledge in semi-structured and unstructured natural languages, how to make use of multiple types of knowledge in decision making, and how to identify the knowledge that is missing. A set of techniques for multi-mode natural language processing, integrated decision making, and open knowledge searching is proposed. The OK-KeJia robot prototype is implemented and evaluated, with special attention to two tests on 11,615 user tasks and 467 user desires. The experiments show that the overall performance improves remarkably due to the use of appropriate open knowledge.
Jiongkun Xie, Jianmin Ji, Zhiqiang Sui
J. Hum. Robot Interact.3
2013 Computing Loops with at Most One External Support Rule
abstract
A consequence of a logic program under answer set semantics is one that is true for all answer sets. This article considers using loop formulas to compute some of these consequences in order to increase the efficiency of answer set solvers. Since computing loop formulas are in general intractable, we consider only loops with either no external support or at most one external support, as their loop formulas are either unit or binary clauses. We show that for disjunctive logic programs, loop formulas of loops with no external support can be computed in polynomial time, and that an iterative procedure using unit propagation on these formulas and the program completion computes the well-founded models in the case of normal logic programs and the least fixed point of a simplification operator used by DLV for disjunctive logic programs. For loops with at most one external support, their loop formulas can be computed in polynomial time for normal logic programs, but are NP-hard for disjunctive programs. So for normal logic programs, we have a procedure similar to the iterative one for loops without any external support, but for disjunctive logic programs, we present a polynomial approximation algorithm. All these algorithms have been implemented, and our experiments show that for certain logic programs, the consequences computed by our algorithms can significantly speed up current ASP solvers cmodels, clasp, and DLV.
Jianmin Ji, Fangzhen Lin
ACM Trans. Comput. Log.2
2013 Computing Loops with at Most One External Support Rule for Basic Logic Programs with Arbitrary Constraint Atoms
Jianmin Ji, Fangzhen Lin, Jia-Huai You
Theory Pract. Log. Program.1
2012 Simulation Competitions on Domestic Robots
Jianmin Ji, Zhiqiang Sui, Guoqiang Jin, Jiongkun Xie
RoboCup1
2009 Computing Loops with at Most One External Support Rule for Disjunctive Logic Programs
Jianmin Ji, Fangzhen Lin
ICLP2
2009 Research Summary
Jianmin Ji
ICLP1
2008 Computing Loops with at Most One External Support Rule
Jianmin Ji, Fangzhen Lin
KR2