Yue Hu 0016

dblp:34/5808-16 · DBLP profile ↗
← Back
27ranked-venue papers
4as first author
27since 2021 · last 2026
0000-0002-8115-7020ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 10 since 2021Security and privacy · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards Autonomous UAV Visual Object Search in City Space: Benchmark and Agentic Methodology
abstract
Aerial Visual Object Search (AVOS) tasks in urban environments require Unmanned Aerial Vehicles (UAVs) to autonomously search for and identify target objects based on visual inputs without external guidance. Existing approaches struggle in complex urban environments due to redundant semantic processing, similar object ambiguity, and the exploration-exploitation dilemma. To advance research and support the AVOS task, we introduce CityAVOS, the first benchmark dataset for autonomous search of static urban objects. It features 2,420 tasks of varying difficulty across six object categories, designed to rigorously evaluate UAV search strategies. To solve the AVOS task, we also propose PRPSearcher (Perception-Reasoning-Planning Searcher), a novel agentic method powered by multi-modal large language models (MLLMs) that enables a UAV agent to think and reason like humans on visual cues when searching for objects. Specifically, PRPSearcher constructs three specialized maps: an object-centric dynamic semantic map enhancing spatial perception, a 3D cognitive map based on semantic "attraction" values for target reasoning, and a 3D uncertainty map for balanced exploration-exploitation search. Moreover, we propose a denoising mechanism to mitigate interference from similar objects and design an Inspiration Promote Thought prompting mechanism for adaptive action planning. Experimental results on CityAVOS demonstrate that PRPSearcher surpasses existing baselines in both success rate and search efficiency (on average: +37.69% SR, +28.96% SPL, -30.69% MSS, and -46.40% NE). Our work paves the way for future advances in embodied visual target search.
Yatai Ji, Zhengqiu Zhu, Beidan Liu, Chen Gao 0001, Sihang Qiu, Yue Hu 0016, Quanjun Yin
AAAI8
2026 CityCube: Benchmarking Cross-view Spatial Reasoning on Vision-Language Models in Urban Environments
abstract
Haotian Xu, Yue Hu, Zhengqiu Zhu, Chen Gao, Ziyou Wang, Junreng Rao, Wenhao Lu, Weishi Li, Quanjun Yin, Yong Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yue Hu 0016, Zhengqiu Zhu, Chen Gao 0001, Ziyou Wang, Junreng Rao, Wenhao Lu, Quanjun Yin, Yong Li 0008
ACL (1)2
2026 Natural Language-Guided Autonomous Agents for Counterterrorism Simulation via Deep Reinforcement Learning
Xinmeng Li, Kai Xu 0014, Yue Hu 0016, Miao Zhang 0037, Quanjun Yin
SIMULTECH3
2026 GeoNav: Empowering MLLMs with dual-scale geospatial reasoning for language-goal aerial navigation
Yue Hu 0016, Chen Gao 0001, Zhengqiu Zhu, Quanjun Yin
Pattern Recognit.2
2026 M2IST: Multi-Modal Interactive Side-Tuning for Efficient Referring Expression Comprehension
abstract
Referring expression comprehension (REC) is a vision-language task to locate a target object in an image based on a language expression. Fully fine-tuning general-purpose pre-trained vision-language foundation models for REC yields impressive performance but becomes increasingly costly. Parameter-efficient transfer learning (PETL) methods have shown strong performance with fewer tunable parameters. However, directly applying PETL to REC faces two challenges: (1) insufficient multi-modal interaction between pre-trained vision-language foundation models, and (2) high GPU memory usage due to gradients passing through the heavy vision-language foundation models. To this end, we present M2IST: Multi-Modal Interactive Side-Tuning with M3ISAs: Mixture of Multi-Modal Interactive Side-Adapters. During fine-tuning, we fix the pre-trained uni-modal encoders and update M3ISAs to enable efficient vision-language alignment for REC. Empirical results reveal that M2IST achieves better performance-efficiency trade-off than full fine-tuning and other PETL methods, requiring only 2.11% tunable parameters, 39.61% GPU memory, and 63.46% training time while maintaining competitive performance. Our code is released at https://github.com/xuyang-liu16/M2IST.
Xuyang Liu 0002, Ting Liu 0018, Siteng Huang, Yi Xin 0003, Yue Hu 0016, Long Qin 0004, Yuanyuan Wu 0001, Honggang Chen
IEEE Trans. Circuits Syst. Video Technol.5
2025 CityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City Space
abstract
Yong Zhao, Kai Xu, Zhengqiu Zhu, Yue Hu, Zhiheng Zheng, Yingfeng Chen, Yatai Ji, Chen Gao, Yong Li, Jincai Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Kai Xu 0014, Zhengqiu Zhu, Yue Hu 0016, Zhiheng Zheng, Yatai Ji, Chen Gao 0001, Yong Li 0008, Jincai Huang 0001
EMNLP4
2025 Query Rephrasing for Context Independence in Scene Knowledge-guided Visual Grounding
abstract
Scene Knowledge-guided Visual Grounding (SK-VG) aims to locate the specific object in an image that is referred to by an open-ended query, utilizing textual scene knowledge for guidance. Besides the grounding ability, SK-VG models are expected to infer more information of the target object based on the expatiatory contextual knowledge, given that the queries in SK-VG are often insufficient for the models to identify the object. Existing models perform well when the queries contain obvious visual clues but falter in hard situations where visual attributes are absent from the query. In this paper, we propose a Large Language Models (LLMs) based Inference framework for Scene Knowledge-guided Visual Grounding (LISK-VG) to address such a problem by rephrasing abstract queries with visually distinguishable object descriptions. Specifically, we explicitly decompose the SK-VG task into two steps, i.e., inference and grounding. The inference step takes advantage of the semantic understanding capabilities of LLMs to fully recognize the useful contextual information in the scene knowledge, which is then injected to the query. LISK-VG, in this way, transforms the originally context-dependent query into a more self-informative form that characterizes the target object mainly by the visual features on the given image. This process significantly facilitates the target localization by avoiding joint understanding of the veiled query and the long-winded knowledge text, which might be largely distractive for the grounding task. Then in the second step, the model performs precise grounding based on the converted query with an explicit object description. The experimental results conducted on the SK-VG dataset demonstrate that our LISK-VG significantly outperforms other approaches in terms of accuracy, particularly exhibiting substantial advantages on hard cases.
Xilong Qin, Haixiang Zhu, Wansen Wu, Yue Hu 0016
IJCNN5
2025 SwimVG: Step-Wise Multimodal Fusion and Adaption for Visual Grounding
abstract
Visual grounding aims to ground an image region through natural language, which heavily relies on cross-modal alignment. Most existing methods transfer visual/linguistic knowledge separately by fully fine-tuning uni-modal pre-trained models, followed by a simple stack of visual-language transformers for multimodal fusion. However, these approaches not only limit adequate interaction between visual and linguistic contexts, but also incur significant computational costs. Therefore, to address these issues, we explore a step-wise multimodal fusion and adaption framework, namely SwimVG. Specifically, SwimVG proposes step-wise multimodal prompts (Swip) and cross-modal interactive adapters (CIA) for visual grounding, replacing the cumbersome transformer stacks for multimodal fusion. Swip can improve the alignment between the vision and language representations step by step, in a token-level fusion manner. In addition, weight-level CIA further promotes multimodal fusion by cross-modal interaction. Swip and CIA are both parameter-efficient paradigms, and they fuse the cross-modal features from shallow to deep layers gradually. Experimental results on four widely-used benchmarks demonstrate that SwimVG achieves remarkable abilities and considerable benefits in terms of efficiency.
Liangtao Shi, Ting Liu 0018, Xiantao Hu, Yue Hu 0016, Quanjun Yin, Richang Hong
IEEE Trans. Multim.4
2024 DAP: Domain-Aware Prompt Learning for Vision-and-Language Navigation
abstract
Following language instructions to navigate in unseen environments is a challenging task for autonomous embodied agents. With strong representation capabilities, pretrained vision-and-language models are widely used in VLN. However, most of them are trained on web-crawled generalpurpose datasets, which incurs a considerable domain gap when used for VLN tasks. To address the problem, we propose a novel and model-agnostic Domain-Aware Prompt learning (DAP) framework. For equipping the pretrained models with specific object-level and scene-level cross-modal alignment in VLN tasks, DAP applies a low-cost prompt tuning paradigm to learn soft visual prompts for extracting in-domain image semantics. Specifically, we first generate a set of in-domain image-text pairs with the help of the CLIP model. Then we introduce soft visual prompts in the input space of the visual encoder in a pretrained model. DAP injects in-domain visual knowledge into the visual encoder of the pretrained model in an efficient way. Experimental results on both R2R and REVERIE show the superiority of DAP compared to existing state-of-the-art methods.
Ting Liu 0018, Yue Hu 0016, Wansen Wu, Youkai Wang, Kai Xu 0014, Quanjun Yin
ICASSP2
2024 DARA: Domain- and Relation-Aware Adapters Make Parameter-Efficient Tuning for Visual Grounding
abstract
Visual grounding (VG) is a challenging task to localize an object in an image based on a textual description. Recent surge in the scale of VG models has substantially improved performance, but also introduced a significant burden on computational costs during fine-tuning. In this paper, we explore applying parameter-efficient transfer learning (PETL) to efficiently transfer the pre-trained vision-language knowledge to VG. Specifically, we propose DARA, a novel PETL method comprising Domain-aware Adapters (DA Adapters) and Relation-aware Adapters (RA Adapters) for VG. DA Adapters first transfer intra-modality representations to be more fine-grained for the VG domain. Then RA Adapters share weights to bridge the relation between two modalities, improving spatial reasoning. Empirical results on widely-used benchmarks demonstrate that DARA achieves the best accuracy while saving numerous updated parameters compared to the full fine-tuning and other PETL methods. Notably, with only 2.13% tunable backbone parameters, DARA improves average accuracy by 0.81% across the three benchmarks compared to the baseline model. Our code is available at https://github.com/liuting20/DARA.
Ting Liu 0018, Xuyang Liu 0002, Siteng Huang, Honggang Chen, Quanjun Yin, Long Qin 0004, Yue Hu 0016
ICME8
2024 Enhancing Multimodal Sentiment Analysis via Learning from Large Language Model
abstract
Multimodal sentiment analysis (MSA) detects human sentiments by understanding data from multiple modalities, such as text and images. Existing research primarily strives for an effective multimodal fusion framework to derive informative representations. However, these methods neglect the necessity of exploiting external knowledge to aid in analyzing sentiments. As a result, the lack of external commonsense embarrasses these models when the opinion cues come in an implicit and obscure manner. To address the limitation, in this paper, we propose an Auxiliary Rationale Knowledge enhanced framework, namely ARK, which improves MSA models via learning from a multimodal large language model (MLLM). Specifically, based on text-image pairs, we employ Chain-of-Thought prompting to generate image descriptions and rationales from the MLLM as auxiliary knowledge, thus enriching the original samples with commonsense knowledge encoded within the MLLM. By combining the source text with image descriptions, we are able to effectively handle MSA through a Text+Text paradigm. In this paradigm, smaller pre-trained language models (LMs) can be tasked for sentiment classification via prompt-tuning. Besides, rationales are leveraged as additional supervision to facilitate the learning of reasoning abilities by LMs. Experimental results demonstrate that our proposed method outperforms current state-of-the-art approaches across four datasets. Our data and code are available at https://github.com/ningpang/ArkMSA.
Ning Pang, Wansen Wu, Yue Hu 0016, Kai Xu 0014, Quanjun Yin, Long Qin 0004
ICME3
2024 PANDA: Prompt-Based Context- and Indoor-Aware Pretraining for Vision and Language Navigation
Ting Liu 0018, Yue Hu 0016, Wansen Wu, Youkai Wang, Kai Xu 0014, Quanjun Yin
MMM (1)2
2024 ACT: Action-assoCiated and Target-Related Representations for Object Navigation
Youkai Wang, Yue Hu 0016, Wansen Wu, Ting Liu 0018, Yong Peng 0006
MMM (1)2
2024 Reinforcement learning from suboptimal demonstrations based on Reward Relabeling
Yong Peng 0006, Junjie Zeng 0004, Yue Hu 0016, Quanjun Yin
Expert Syst. Appl.3
2024 Vision-language navigation: a survey and taxonomy
Wansen Wu, Tao Chang, Xinmeng Li, Quanjun Yin, Yue Hu 0016
Neural Comput. Appl.5
2024 Visual Grounding With Dual Knowledge Distillation
abstract
Visual grounding is a task that seeks to predict the specific location of an object or region described by a linguistic expression within an image. Despite the recent success, existing methods still suffer from two problems. First, most methods use independently pre-trained unimodal feature encoders for extracting expressive feature embeddings, thus resulting in a significant semantic gap between unimodal embeddings and limiting the effective interaction of visual-linguistic contexts. Second, existing attention-based approaches equipped with the global receptive field have a tendency to neglect the local information present in the images. This limitation restricts the semantic understanding required to distinguish between referred objects and the background, consequently leading to inadequate localization performance. Inspired by the recent advance in knowledge distillation, in this paper, we propose a DUal knowlEdge disTillation (DUET) method for visual grounding models to bridge the cross-modal semantic gap and improve localization performance simultaneously. Specifically, we utilize the CLIP model as the teacher model to transfer the semantic knowledge to a student model, in which the vision and language modalities are linked into a unified embedding space. Besides, we design a self-distillation method for the student model to acquire localization knowledge by performing the region-level contrastive learning to make the predicted region close to the positive samples. To this end, this work further proposes a Semantics-Location Aware sampling mechanism to generate high-quality self-distillation samples. Extensive experiments on five datasets and ablation studies demonstrate the state-of-the-art performance of DUET and its orthogonality with different student models, thereby making DUET adaptable to a wide range of visual grounding architectures. Our code are available on DUET.
Wansen Wu, Meng Cao 0002, Yue Hu 0016, Yong Peng 0006, Long Qin 0004, Quanjun Yin
IEEE Trans. Circuits Syst. Video Technol.3
2024 MalPatch: Evading DNN-Based Malware Detection With Adversarial Patches
abstract
Static analysis is a crucial protection layer that enables modern antivirus systems to address the rampant proliferation of malware. These systems are increasingly relying on deep neural networks (DNNs) to automatically extract reliable features and achieve outstanding detection accuracy. Since DNNs are known to be vulnerable to adversarial examples, several studies have proposed practical evasion attacks to generate adversarial perturbations that can evade malware detectors. These attacks, however, require specific designs for the given input sample, prohibiting them from large-scale deployment. Therefore, it is more practical to generate sample-agnostic perturbations that do not involve recalculations regardless of the input malware sample. To this end, we leverage an adversarial patch attack, which is a special type of adversarial attack that dose not know the sample being modified during the attack construction process. In particular, we propose a new adversarial attack against malware detection systems called MalPatch. It locates the nonfunctional part of malware for adversarial patch injection to protect its executability while generating adversarial examples based on different strategies. The generated patch can be injected into any malware sample, fooling the detector into classifying it as benign. Experimental results demonstrate that MalPatch is effective under different attack settings. In the white-box setting, MalPatch achieves 69%-78% success rates against DNN detectors based on raw byte features and 47%-96% success rates against four grayscale detectors based on image features. In the black-box setting, the success rates of MalPatch against the same models reach 54%-74% and 27%-42%, respectively. We conclude by discussing several of its potential countermeasures and the generality of our approach.
Dazhi Zhan, Yexin Duan, Yue Hu 0016, Shize Guo, Zhisong Pan 0003
IEEE Trans. Inf. Forensics Secur.3
2023 PSP-Mal: Evading Malware Detection via Prioritized Experience-based Reinforcement Learning with Shapley Prior
abstract
With the widespread application of machine learning techniques in malware detection, researchers have proposed various adversarial attack methods to generate adversarial examples (AEs) of malware, thereby evading detection. Previous studies have shown that the reinforcement learning (RL) framework can enable black-box attacks by performing a sequence of function-preserving operations, which produces functional evasive malware samples. However, it is difficult to obtain the useful guidance and feedbacks from the environment for agent training in the black-box scenario, which results in the RL framework being unable to learn the effective evasion policy. In this paper, we propose the Shapley prior and establish a prior-guidance-based RL framework, namely PSP-Mal, to generate AEs against Portable Executable (PE) malware detectors. Our framework improves on existing methods in three aspects: 1) We explore feature effects of the black-box model by computing Shapley values and further propose the Shapley prior to represent the expected impact of operations. 2) A novel prioritized experience utilization mechanism is established regarding the Shapley prior guidance in the RL framework. 3) The actions are expanded into item-content pairs and we use the Thompson sampling to choose effective content, which helps to reduce randomness and ensure repeatability. We compare the attack performance of our framework with other methods, and experimental results demonstrate that our algorithm is more effective. The evasion rates of PSP-Mal against the LightGBM models trained on EMBER and SOREL-20M reach 76.88% and 72.03%, respectively.
Dazhi Zhan, Xin Liu 0042, Yue Hu 0016, Lei Zhang 0126, Shize Guo, Zhisong Pan 0003
ACSAC4
2023 Dynamic Multi-modal Prompting for Efficient Visual Grounding
Wansen Wu, Ting Liu 0018, Youkai Wang, Kai Xu 0014, Quanjun Yin, Yue Hu 0016
PRCV (7)6
2023 AMGmal: Adaptive mask-guided adversarial attack against malware detection with minimal perturbation
Dazhi Zhan, Yexin Duan, Yue Hu 0016, Lujia Yin, Zhisong Pan 0003, Shize Guo
Comput. Secur.3
2023 Towards robust CNN-based malware classifiers using adversarial examples generated based on two saliency similarities
Dazhi Zhan, Yue Hu 0016, Shize Guo, Zhisong Pan 0003
Neural Comput. Appl.2
2022 FedHiSyn: A Hierarchical Synchronous Federated Learning Framework for Resource and Data Heterogeneity
abstract
Federated Learning (FL) enables training a global model without sharing the decentralized raw data stored on multiple devices to protect data privacy. Due to the diverse capacity of the devices, FL frameworks struggle to tackle the problems of straggler effects and outdated models. In addition, the data heterogeneity incurs severe accuracy degradation of the global model in the FL training process. To address aforementioned issues, we propose a hierarchical synchronous FL framework, i.e., FedHiSyn. FedHiSyn first clusters all available devices into a small number of categories based on their computing capacity. After a certain interval of local training, the models trained in different categories are simultaneously uploaded to a central server. Within a single category, the devices communicate the local updated model weights to each other based on a ring topology. As the efficiency of training in the ring topology prefers devices with homogeneous resources, the classification based on the computing capacity mitigates the impact of straggler effects. Besides, the combination of the synchronous update of multiple categories and the device communication within a single category help address the data heterogeneity issue while achieving high accuracy. We evaluate the proposed framework based on MNIST, EMNIST, CIFAR10 and CIFAR100 datasets and diverse heterogeneous settings of devices. Experimental results show that FedHiSyn outperforms six baseline methods, e.g., FedAvg, SCAFFOLD, and FedAT, in terms of training accuracy and efficiency.
Yue Hu 0016, Miao Zhang 0037, Ji Liu 0003, Quanjun Yin, Yong Peng 0006, Dejing Dou
ICPP2
2022 FedGosp: A Novel Framework of Gossip Federated Learning for Data Heterogeneity
abstract
Federated learning (FL) provides the possibility to solve the problem of data privacy, but it suffers much from the data heterogeneity among different participants. Currently, some promising FL algorithms improve the effectiveness of learning under the non independent-and-identically-distributed (Non-IID) data settings. However, they require a large number of communication rounds between the server and clients for an acceptable accuracy. Inspired by the training paradigm of gossip learning, this paper proposes a new FL framework, named FedGosp. It first classifies the clients into different categories based on the model weights trained by the locally stored data. Then FedGosp utilizes the communication not only between clients and the server, but also between different classes of clients themselves. This training process enables instilling knowledge about various data distributions in the passed models. We evaluate the performance of FedGosp in multiple Non-IID settings on CIFAR10 and MNIST datasets, and compare it with the recently popular algorithms such as SCAFFOLD, FedAvg and FedProx. The experimental results show that FedGosp can improve the model accuracy by 6.53% and save 5.6 × communication costs at best compared to the second-ranked baseline.
Yue Hu 0016, Miao Zhang 0037, Li Li 0064, Tao Chang, Quanjun Yin
SMC2
2022 Who are the 'silent spreaders'?: contact tracing in spatio-temporal memory models
Yue Hu 0016, Budhitama Subagdja, Ah-Hwee Tan, Hiok Chai Quek, Quanjun Yin
Neural Comput. Appl.1
2022 Vision-Based Topological Mapping and Navigation With Self-Organizing Neural Networks
abstract
Spatial mapping and navigation are critical cognitive functions of autonomous agents, enabling one to learn an internal representation of an environment and move through space with real-time sensory inputs, such as visual observations. Existing models for vision-based mapping and navigation, however, suffer from memory requirements that increase linearly with exploration duration and indirect path following behaviors. This article presents e -TM, a self-organizing neural network-based framework for incremental topological mapping and navigation. e -TM models the exploration trajectories explicitly as episodic memory, wherein salient landmarks are sequentially extracted as "events" from streaming observations. A memory consolidation procedure then performs a playback mechanism and transfers the embedded knowledge of the environmental layout into spatial memory, encoding topological relations between landmarks. Fusion adaptive resonance theory (ART) networks, as the building block of the two memory modules, can generalize multiple input patterns into memory templates and, therefore, provide a compact spatial representation and support the discovery of novel shortcuts through inferences. For navigation, e -TM applies a transfer learning paradigm to integrate human demonstrations into a pretrained locomotion network for smoother movements. Experimental results based on VizDoom, a simulated 3-D environment, have shown that, compared to semiparametric topological memory (SPTM), a state-of-the-art model, e -TM reduces the time costs of navigation significantly while learning much sparser topological graphs.
Yue Hu 0016, Budhitama Subagdja, Ah-Hwee Tan, Quanjun Yin
IEEE Trans. Neural Networks Learn. Syst.1
2021 Interpretable Goal Recognition for Path Planning with ART Networks
abstract
Goal recognition for path planning is an important task of intention identification and situation awareness, requiring an observer to predict the goal of an evader given observations of its movements. While existing models based on planning or Markov Decision Process (MDP) show superior performance over traditional library based methods, they require much effort in model design and can hardly provide legible decision rules for their users. To make the system more user-friendly while preserving accuracy of goal inference, this paper proposes a novel self-organizing neural network based inference model, which learns compact rule sets through generalizing the streaming observations of an evader. More critically, the system manifests a high level of interpretability with the linguistic if-then rule base, making it easily comprehensible for human decision makers. We conducted extensive experiments on a large-scale real-world road network. Results show that the proposed model produces accuracy comparable to those of two state-of-the-art methods while uniquely providing legible inference rules and strong robustness against multiple goals with missing data.
Yue Hu 0016, Kai Xu 0014, Budhitama Subagdja, Ah-Hwee Tan, Quanjun Yin
IJCNN1
2021 Regarding Goal Bounding and Jump Point Search
abstract
Jump Point Search (JPS) is a well known symmetry-breaking algorithm that can substantially improve performance for grid-based optimal pathfinding. When the input grid is static further speedups can be obtained by combining JPS with goal bounding techniques such as Geometric Containers (instantiated as Bounding Boxes) and Compressed Path Databases. Two such methods, JPS+BB and Two-Oracle Path PlannING (Topping), are currently among the fastest known approaches for computing shortest paths on grids. The principal drawback for these algorithms is the overhead costs: each one requires an all-pairs precomputation step, the running time and subsequent storage costs of which can be prohibitive. In this work we consider an alternative approach where we precompute and store goal bounding data only for grid cells which are also jump points. Since the number of jump points is usually much smaller than the total number of grid cells, we can save up to orders of magnitude in preprocessing time and space. Considerable precomputation savings do not necessarily mean performance degradation. For a second contribution we show how canonical orderings, partial expansion strategies and enhanced intermediate pruning can be leveraged to improve online query performance despite a reduction in preprocessed data. The combination of faster preprocessing and stronger online reasoning leads to three new and highly performant algorithms: JPS+BB+ and Two-Oracle Pathfinding Search (TOPS) based on search, and Topping+ based on path extraction. We give a theoretical analysis showing that each method is complete and optimal. We also report convincing gains in a comprehensive empirical evaluation that includes almost all current and cutting-edge algorithms for grid-based pathfinding.
Yue Hu 0016, Daniel Harabor, Long Qin 0004, Quanjun Yin
J. Artif. Intell. Res.1