Yanfeng Hu

dblp:97/9617 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Computer networks · 5 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-grained dynamic feature resorter with curriculum bootstrapping for multimodal entity and relation extraction
Zhicong Lu, Li Jin 0001, Linhao Zhang, Kaiwen Wei, Qing Liu 0021, Yanfeng Hu
Neurocomputing8
2026 A Low-Cost Dual-Antenna Approach for GNSS Signal Authentication Against Constrained Distributed Spoofing
abstract
Owing to its cost-effectiveness and dependable accuracy, the Global Navigation Satellite System (GNSS) receiver has become an essential positioning utility across diverse navigation applications. However, its reliance on inherently insecure wireless signals renders it vulnerable to spoofing attacks, posing a significant threat to safety-critical operations. Most existing countermeasures predicate on the common hypothesis that all spoofing signals originate from a single direction, exhibiting performance degradation against distributed spoofing or mixed signal scenarios. Conventional array antennas offer excellent spatial resolution but suffer from high costs and calibration requirements. To address these limitations, this paper presents a novel low-cost dual-antenna approach that innovatively integrates dynamic baseline rotation with Pearson correlation analysis of C/Nosequences. This method extracts spatial signatures by exploiting motion-induced phase differences, enabling reliable detection of spoofing clusters from common directions without requiring specialized hardware or receiver modifications. Experimental results validate the core mechanism: when two satellite signals originate from the same direction, detection probability exceeds 99.5% with 0.01% false probability, providing principled verification for constrained distributed spoofing (where multiple spoofed signals emanate from a single source). The system's hardware independence (based on C/Nodata output) and compact design (baseline <1m) enable compatibility with vehicular and UAV platforms, where natural platform motion enhances baseline variability and spatial feature diversity. Simulations and real-world tests demonstrate efficacy against distributed-like spoofing scenarios, offering a practical, multipath-resistant solution for GNSS security enhancement.
Yanfeng Hu
IEEE Internet Things J.1
2026 A Novel OTFS-Based Massive Random Access Scheme in Cell-Free Massive MIMO Systems for High-Speed Mobility
Yanfeng Hu, Dongming Wang 0002, Xinjiang Xia, Jiamin Li 0001, Pengcheng Zhu 0001, Xiaohu You 0001
IEEE Trans. Mob. Comput.1
2026 ISAC for Cell-Free Massive MIMO: Cooperation and Sensing Information Fusion
abstract
To realize the potential of integrated sensing and communication (ISAC) in cell-free (CF) massive MIMO system, an ISAC framework is proposed in this paper. With numbers and locations of scatterers (targets) unknown, based on received downlink data signals from the transmit access point (tAP) antenna, multiple receive access points (rAPs) first estimate delays locally. In each outer iteration, by comparing the estimated delays with the delays of the paths through extracted scatterer locations, selected rAPs search out mismatched estimated delays. Through cooperation, the potential locations of the scatterers forming each candidate location set corresponding to each mismatched delay are obtained, and each set will be evaluated in sequence. Specifically, a joint evaluation algorithm is proposed, where the probability model for the joint evaluation problem is established based on delay and limited angular information. Under the expectation maximization (EM) framework, by fusing information from all the rAPs, the extracted scatterer locations and candidate locations are adjusted, and the evaluation results are estimated, which can be regarded as the global probability of scatterers exist at candidate locations. When the evaluation results of all the locations in a candidate location set are low, the corresponding delay will be discarded, otherwise, the candidate location with the highest evaluation result in the set will be extracted. The proposed framework avoids exhaustive search in different associations of scatterers and estimated delays, and is more flexible than extraction from all the candidate locations based on fixed thresholds. Then, based on a simplified model, the impact of system parameters on delay based location sensing accuracy is revealed with the theoretical analysis. Finally, based on the prototype system with CF radio access network (RAN) architecture, the proposed framework is experimentally validated.
Jie Ling 0003, Jing Jin 0007, Qixing Wang, Xinsheng Zhao, Jiamin Li 0001, Yanfeng Hu, Siying Lv, Dongming Wang 0002, Xiaohu You 0001, Jiangzhou Wang
IEEE Trans. Wirel. Commun.6
2025 SemStereo: Semantic-Constrained Stereo Matching Network for Remote Sensing
abstract
Semantic segmentation and 3D reconstruction are two fundamental tasks in remote sensing, typically treated as separate or loosely coupled tasks. Despite attempts to integrate them into a unified network, the constraints between the two heterogeneous tasks are not explicitly modeled, since the pioneering studies either utilize a loosely coupled parallel structure or engage in only implicit interactions, failing to capture the inherent connections. In this work, we explore the connections between the two tasks and propose a new network that imposes semantic constraints on the stereo matching task, both implicitly and explicitly. Implicitly, we transform the traditional parallel structure to a new cascade structure termed Semantic-Guided Cascade structure, where the deep features enriched with semantic information are utilized for the computation of initial disparity maps, enhancing semantic guidance. Explicitly, we propose a Semantic Selective Refinement (SSR) module and a Left-Right Semantic Consistency (LRSC) module. The SSR refines the initial disparity map under the guidance of the semantic map. The LRSC ensures semantic consistency between two views via reducing the semantic divergence after transforming the semantic map from one view to the other using the disparity map. Experiments on the US3D and WHU datasets demonstrate that our method achieves state-of-the-art performance for both semantic segmentation and stereo matching.
Chen Chen 0036, Liangjin Zhao, Yuanchun He, Yingxuan Long, Kaiqiang Chen, Zhirui Wang 0003, Yanfeng Hu, Xian Sun 0001
AAAI7
2025 SA-Occ: Satellite-Assisted 3D Occupancy Prediction in Real World
abstract
Existing vision-based 3D occupancy prediction methods are inherently limited in accuracy due to their exclusive reliance on street-view imagery, neglecting the potential benefits of incorporating satellite views. We propose SA-Occ, the first Satellite-Assisted 3D occupancy prediction model, which leverages GPS & IMU to integrate historical yet readily available satellite imagery into real-time applications, effectively mitigating limitations of ego-vehicle perceptions, involving occlusions and degraded performance in distant regions. To address the core challenges of cross-view perception, we propose: 1) Dynamic-Decoupling Fusion, which resolves inconsistencies in dynamic regions caused by the temporal asynchrony between satellite and street views; 2) 3D-Proj Guidance, a module that enhances 3D feature extraction from inherently 2D satellite imagery; and 3) Uniform Sampling Alignment, which aligns the sampling density between street and satellite views. Evaluated on Occ3D-nuScenes, SA-Occ achieves state-of-the-art performance, especially among single-frame methods, with a 39.05% mIoU (a 6.97% improvement), while incurring only 6.93 ms of additional latency per frame. Our code and newly curated dataset are available at https://github.com/chenchen235/SA-Occ.
Chen Chen 0036, Zhirui Wang 0003, Taowei Sheng, Yundu Li, Peirui Cheng, Luning Zhang, Kaiqiang Chen, Yanfeng Hu, Xue Yang 0005, Xian Sun 0001
ICCV9
2025 Hetrify: Efficient Verification of Heterogeneous Programs on RISC-V
abstract
The heterogeneous nature of contemporary software, comprising components like closed-source libraries, embedded assembly snippets, and modules written in multiple programming languages, leads to significant verification challenges. Currently, there are no mature and available methods to effectively address such problems. To bridge this gap, we propose a verification approach capable of effectively verifying heterogeneous programs. This approach is universally applicable. It theoretically supports the verification of any heterogeneous program that can be compiled into binary code, without being constrained by any specific programming language. The approach begins by compiling the entire program or its unverifiable segments into binary format. Under guarantees of semantic equivalence, these binaries are converted into verifiable C code, which can then be verified using existing C verification tools. Based on the RISC-V architecture, we developed the Hetrify tool to implement this verification approach. The tool is supported by rigorous mathematical proofs to ensure operational semantic equivalence between the converted C programs and their original counterparts. To validate our approach, we conducted verification experiments on 130 programs, including 100 assembly programs and 30 large heterogeneous programs with missing critical function source code, demonstrating the effectiveness of our approach.
Yiwei Li 0006, Liangze Yin, Wei Dong 0006, Yanfeng Hu
ICSE5
2025 U-MERE: Unconstrained Multimodal Entity and Relation Extraction with Collaborative Modeling and Order-Sensitive Optimization
abstract
Existing multimodal entity and relation extraction tasks primarily focus on text-to-text or text-to-visual entity relations, overlooking real-world complexities involving visual-to-text and visual-to-visual cases, thus failing to capture the richer semantic structures in complex cross-modal interactions. To address the limitations, we propose a new task, Unconstrained Multimodal Entity and Relation Extraction (U-MERE), which jointly extracts arbitrary visual and textual entities, and their relations from image-text pairs. To accomplish U-MERE, we construct UMERE-Bench, a benchmark with over 9,000 samples that comprehensively covers four cross-modal entity relation directions and three task settings. Given the difficulty of jointly modeling diverse directions of cross-modal entity relations, we introduce Collaborative Modeling and Order-Sensitive (CMOS), which collaboratively guides large vision-language models (LVLMs) to decompose task complexity and mitigates generation order bias from fixed target relation sequences. CMOS employs small models to generate candidate entities, guiding LVLMs to capture key information and jointly optimizes multiple feasible relation orderings to reduce order dependency. Additionally, we design a Multimodal Order-aware Matching (MOM) evaluation method to align predictions with ground truth for precise assessment. Experimental results reveal that current LVLMs show limited performance on U-MERE, underscoring its inherent challenges, while CMOS consistently achieves superior performance across multiple advanced LVLMs, demonstrating its effectiveness and generalization capability. The dataset and code will be available in https://github.com/jiaweidoris/U-MERE.
Li Jin 0001, Kaiwen Wei, Yuying Shang, Nayu Liu, Zhicong Lu, Qing Liu 0021, Linhao Zhang, Yanfeng Hu
ACM Multimedia10
2025 Beyond Test Cases: Multi-Agent Collaboration for Detecting Errors in Full-Score Code Implementations
abstract
Automated evaluation of programming code on online platforms often relies on predefined test cases.However, due to limited test coverage, many programs receive full marks despite violating intended specifications.We present Maveric, a framework that combines large language models (LLMs) with formal verification to more rigorously assess code correctness.Maveric consists of four agents: a template generator that derives formal specifications from problem descriptions, a consistency checker that validates semantic alignment, a code analyzer that detects potential defects and synthesizes counterexamples, and a counterexample validator that formally verifies their validity.We evaluated Maveric on 100 full-score code submissions from 10 real-world programming tasks sourced from a widely used online education platform.Manual review identified 32 with functional defects.Maveric accurately detected 31 of these with no false positives, completing the evaluation of each program in under one minute.In contrast, LLM-only methods detected 25 defects but yielded 6 false positives, while formal verification alone found 23 and suffered frequent timeouts.Importantly, all defects reported by Maveric were supported by verifiable counterexamples, confirming their semantic violations.These results demonstrate Maveric's effectiveness and practicality for automated program evaluation in educational settings.
Yiwei Li 0006, Yanfeng Hu, Liangze Yin, Wei Dong 0006
SEKE3
2025 A Hybrid Preamble Scheme for Massive Random Access in High-Mobility MIMO-OTFS System
abstract
Orthogonal Time Frequency Space (OTFS) modulation can effectively enhance transmission efficiency in doubly selective channels and is considered as one of the modulation techniques for next generation wireless communication. To meet the demands of machine-type communication in high-speed scenarios, this paper conducts research on a massive random access scheme based on OTFS modulation. Specifically, the transmission and reception model of multiple-input multiple-output (MIMO) OTFS signals is analyzed, where channel estimation is formulated as a block-sparse signal recovery problem. Therefore, based on existing superimposed and embedded pilot schemes, we propose a novel hybrid preamble scheme. This scheme utilizes the superimposed preamble and embedded preamble to achieve rough active user detection (AUD) and precise AUD, respectively, enabling accurate detection and channel estimation while supporting a large number of device access. On this basis, this paper employs a generalized approximate message passing with pattern-coupled sparse Bayesian learning (GAMP-PCSBL) algorithm that can capture the block sparsity characteristics of the channel matrix and achieve accurate estimation. The results of numerical simulations verify the effectiveness and superiority of the proposed scheme.
Yanfeng Hu, Dongming Wang 0002, Pengzhe Xin, Yunxiang Guo, Jie Ling 0003, Xinjiang Xia
VTC2025-Fall1
2025 Two-Stage Learning Approach for Semantic-Aware Task Scheduling in Container-Based Clouds
abstract
Container-based task scheduling is critical for ensuring a reliable, flexible and cost-effective cloud computing mode. However, in different business cloud systems, state-of-the-art scheduling models are not as effective as those in the simulated world due to the sparsity issues associated with sample sizes and features. Herein, we propose a novel containerized task scheduling framework (SA2CTS) based on reinforcement learning (RL) that incorporates cross-modal contrastive learning (CL) loss. This framework optimizes the scheduler's understanding of the container-based cloud state in RL by adding a pretraining stage, promoting accurate scheduling action inference. Specifically, we design a two-stage learning pipeline. The initial stage involves pretraining the model on a large collection of aligned image-text pairs to extract fine-grained scheduling affinity features, and the high-level semantic representations of scheduling tasks are learned in the multimodal space. In the second stage, we fine-tune the pretrained model with multisource cluster feedback, i.e., build a mapping from state representations to scheduling actions through the RL paradigm, achieving task-oriented and semantic-aware scheduling. The experimental results obtained on three large-scale production cluster datasets substantiate that the proposed SA2CTS method can provide average convergence efficiency and resource utilization improvements of 17.57% and 10.42%, respectively, over the state-of-the-art RL scheduling methods.
Lilu Zhu, Yanfeng Hu, Yang Wang 0015
IEEE Trans. Cloud Comput.3
2025 View-Based Knowledge-Augmented Multimodal Semantic Understanding for Optical Remote Sensing Images
abstract
Optical remote sensing (RS) images serve as a pivotal source of geographic information. Due to the continuous development of deep learning technology, the evolving demands for multisource optical RS of the public shifted from recognition and acquisition of explicit features to comprehension and application of the fine-grained semantics and relationships implied in images. To address this challenge, we propose a semantic-augmented approach integrated multiview knowledge graph for a comprehensive understanding of optical RS images (RSMVKF). The RSMVKF delves into the structured representations of external knowledge from different human-like cognitive views and further explores the discovery ability of high-level features on the basis of multiple modalities and granularities. Specifically, the RSMVKF consists of two stages. First, we guide a large language model (LLM) to condense relevant knowledge from lengthy external knowledge passages and generate a view-level knowledge graph (RS-VKG). Then, an asymmetric multimodal contrastive network model (RS-M2CL) is designed to investigate efficient semantic augmentation. In this way, two types of contrastive loss functions, cross-modal and cross-granularity, are adopted to improve the understanding of implicit knowledge. The experimental results demonstrate that the RSMVKF greatly improves several perception tasks and reasoning tasks with rich features in optical RS imagery. In particular, in perception tasks such as fine-grained object detection and k-nearest neighbor (KNN) retrieval, the RSMVKF yields enhancements of 6.7% and 8.1%, respectively. In addition, in knowledge-driven reasoning tasks such as RS image captioning (RSCP), RS visual grounding (RSVG), and RS visual question answering (RSVQA), the RSMVKF demonstrates superior performance with margins of 8.9%, 5.3%, and 11.4%, respectively.
Lilu Zhu, Xiaolu Su, Jiaxuan Tang, Yanfeng Hu
IEEE Trans. Geosci. Remote. Sens.4
2025 A Multi-Objective Path Planning Framework in Off-Road Environments Based on Deep Reinforcement Learning
abstract
Efficient and safe path planning is a key navigation requirement for intelligent vehicles, particularly in off-road environments characterized by complex terrain. Traditional methods focus mostly on minimizing travel time or path length, neglecting the potential for vehicle failure due to complex terrain and dynamic conditions. This limitation is further exacerbated under varying meteorological influences. Here, we propose a heuristic multiobjective path planning (AC-MOPP) framework based on deep reinforcement learning (DRL). This framework overcomes the shortcomings of traditional methods by incorporating multiple objectives and accounting for meteorological variations while reducing learning costs through DRL algorithms. Specifically, we define a traversability map environment, an RL agent, driving actions, and path evaluation methods to construct the path planning model. Using the actor–critic algorithm, we integrate meteorology-driven heuristic rules and a prioritized experience replay method to assist the agent in adjusting its path, accelerating the convergence process, and reducing its learning costs. Finally, comparative experiments are conducted, and the results demonstrate that the AC-MOPP outperforms the standard DRL and state-of-the-art non-DRL algorithms in terms of planning efficiency and convergence stability. Moreover, AC-MOPP achieves better navigation performance than do commercial platforms such as Google Maps, Bing Maps, and OpenStreetMap. This work offers a robust navigation solution for intelligent vehicles under various off-road conditions.
Lilu Zhu, Yanfeng Hu
IEEE Trans. Intell. Transp. Syst.4
2024 RustPruner: A Program Slicing Tool for Rust Programs
abstract
In the realm of analyzing Rust programs, traditional full-scale analytical approaches are often rendered impractical due to considerable time and performance expenditures.Despite the potential for program slicing technologies to drastically curtail these expenses, the vast majority of existing slicing tools lack compatibility with the Rust language.This study introduces RustPruner, a specialized tool designed for slicing Rust code.RustPruner commences by constructing a goto-style control flow graph (CFG) through rigorous control flow analysis, comprehensively mapping out all feasible execution paths.Subsequently, it leverages a backward slicing algorithm, integrating the outcomes of program dependency analysis and abstract syntax tree (AST) analysis, to generate precise program slices.By innovatively adapting slicing technology to Rust, RustPruner facilitates comprehensive, low-level analysis of Rust's distinctive features.The generated slices undergo rigorous compilation and verification procedures, thereby enhancing the efficiency of the analysis process.The effectiveness of RustPruner has been validated within some crucial system modules of operational operating systems, which significantly improves both the efficiency and accuracy of the verification.
Yanfeng Hu, Ruiyu Zhang, Liangze Yin
SEKE1
2024 Decentralized Massive Access Random Scheme in User-Centric Cell-Free Massive MIMO System
abstract
This paper considers a decentralized massive random access scheme applicable to a user-centric cell-free massive MIMO system. In this system, a large number of user equipment (UEs) select access points (APs) based on the quality of channel, and adjacent APs can achieve lossless data exchange with a finite data volume. Each AP is equipped with an Edge Distributed Unit (EDU) capable of independently performing Maximum Likelihood (ML) estimation of the large-scale fading coefficients (LSFC) for the UEs associated with. Subsequently, the APs collectively assess the activity of associated UEs based on the estimated LSFC vectors obtained from adjacent APs, and employ a heuristic method to identify primary interference sources. On this basis, each AP's EDU uses the approximate message passing - sparse Bayesian learning (AMP-SBL) algorithm for channel estimation (CE). The proposed scheme concentrates on the associated user set, reducing computational complexity and signaling overhead while maintaining accurate performance. Numerical simulations validate the effectiveness and superiority of the approach presented in this paper.
Yanfeng Hu, Dongming Wang 0002, Xinjiang Xia, Xiaohu You 0001
WCNC1
2024 Active Detection and Channel Estimation Schemes for Massive Random Access in User-Centric Cell-Free Massive MIMO System
abstract
The demand for higher transmission efficiency and denser user access has been put forth by the next generation of wireless communication systems. To cater to the future communication development, this article focuses on massive random access schemes under the user-centric cell-free massive multiple-input-multiple-output (MIMO) architecture. For uplink transmission, a data frame structure is designed to enable active user detection (AUD), channel estimation (CE), and data transmission. The association between access points (APs) and user equipment (UEs) is presented to facilitate an user-centric cell-free scalable architecture. In this article, a maximum likelihood (ML)-based method is proposed for AUD to obtain the set of active UEs. By setting appropriate thresholds and combining the UE-AP association, accurate active detection results can be obtained. CE can be accomplished with lower computational complexity by utilizing the detected active UE set in AUD module. Specifically, the generalized approximate message passing-based sparse Bayesian learning with Dirichlet process (GAMP-DP-SBL) is adopted as the CE algorithm, leveraging the spatial aggregation and dispersion characteristics of APs to enhance the estimation accuracy. Building upon GAMP-DP-SBL algorithm, a clustered algorithm (GAMP-CDP-SBL) is proposed to reduce the scale of the sensing matrix and improve the accuracy of CE for associated active UEs. Moreover, to enhance system scalability, decentralized AUD and CE algorithms are proposed in this article. Simulation results under various parameter settings and different scenarios exhibit the superior performance of the proposed scheme.
Yanfeng Hu, Qingtian Wang, Dongming Wang 0002, Xinjiang Xia, Xiaohu You 0001
IEEE Internet Things J.1
2024 Cross-Modal Contrastive Learning With Spatiotemporal Context for Correlation-Aware Multiscale Remote Sensing Image Retrieval
abstract
Optical satellites are the most popular observation platforms for humans viewing Earth. Driven by rapidly developing multisource optical remote sensing technology, content-based remote sensing image retrieval (CBRSIR), which aims to retrieve images of interest using extracted visual features, faces new challenges derived from large data volumes, complex feature information, and various spatiotemporal resolutions. Most previous works delve into optical image representation and transformation to the semantic space of retrieval via supervised or unsupervised learning. These retrieval methods fail to fully leverage geospatial information, especially spatiotemporal features, which can improve the accuracy and efficiency to some extent. In this article, we propose a cross-modal contrastive learning method (CCLS2T) to maximize the mutual information of multisource remote sensing platforms for correlation-aware retrieval. Specifically, we develop an asymmetric dual-encoder architecture with a vision encoder that operates on multiscale visual inputs, and a lightweight text encoder that reconstructs spatiotemporal embeddings and adopts an intermediate contrastive objective on representations from unimodal encoders. Then, we add a hash layer to transform the deep fusion features into compact hash index codes. In addition, CCLS2T exploits the prompt template (R2STFT) for multisource remote sensing retrieval to address the text heterogeneity of metadata files and the hierarchical semantic tree (RSHST) to address the feature sparsification of semantic-aware indexing structures. The experimental results on three optical remote sensing datasets substantiate that the proposed CCLS2T can improve retrieval performance by 11.64% and 9.91% compared with many existing hash learning methods and server-side retrieval engines, respectively, in typical optical remote sensing retrieval scenarios.
Lilu Zhu, Yang Wang 0056, Yanfeng Hu, Xiaolu Su, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 Analysis of RRU Correlation Performance in Full Spectrum Uplink Cell-free RAN Systems
abstract
In recent years, with the improvement of mobile communication network performance, cell-free mMIMO with fully integrated and distributed processing has been widely studied, and wireless access networks have also become a widely studied topic in the academic community. This article analyzes the spectral efficiency of a new full spectrum cell-free wireless access network architecture with multiple EDUs and derives the performance upper and lower bounds of its traversal achievable rate. The full set formula and distributed processing can be used as two special cases; Secondly, the article further considers the distribution of users and large-scale fading models and studies the location distribution of RRUs. It is concluded that a uniform distribution of RRUs is beneficial for user traversal and speed improvement, and RRUs in multiple EDUs need to be as intertwined as possible, which is different from traditional multi-node clustering centralized collaborative processing; The article further proposes an modified genetic algorithms (GA), which simulates the performance of a new full spectrum cellfree wireless network with pilot pollution. The performance is compared with that of multi-RRU clustering group collaboration processing through Monte Carlo simulation.
Dongming Wang 0002, Yunxiang Guo, Yanfeng Hu, Jie Ling 0003, Baiping Xiong
VTC Fall4
2023 A priority-aware scheduling framework for heterogeneous workloads in container-based cloud
Lilu Zhu, Yanfeng Hu
Appl. Intell.4
2023 A heuristic multi-objective task scheduling framework for container-based clouds via actor-critic reinforcement learning
Lilu Zhu, Feng Wu 0001, Yanfeng Hu, Xinmei Tian 0001
Neural Comput. Appl.3
2017 Aircraft Recognition Based on Landmark Detection in Remote Sensing Images
abstract
Aircraft type recognition of remote sensing images is critical both in civil and military applications. In this letter, we propose a novel landmark-based aircraft recognition method which is highly accurate and efficient. First, we propose a new idea to address the aircraft type recognition problem by aircraft's landmark detection. Its advantages are two folds. On the one hand, it needs fewer labeled data and alleviates the work of human annotation. On the other hand, a trained model has strong expansibility because it can be used for any type of aircraft that not contained in training data set without retraining. Then, we use a variant of a convolutional neural network called vanilla network for all landmarks regression at the same time. Therefore, it can avoid bad local minimum effectively by encoding the geometric constraints among landmarks implicitly. To handle aircrafts in myriads of poses, rotation jittering is used for data augmentation in preprocessing and multicrop fusion is used in postprocessing. Thus, an 80% reduction in error rate could be reached. Finally, we use the landmark template matching to recognize the aircraft. Our method shows a competitive performance both in accuracy and efficiency.
Kun Fu 0001, Siyue Wang, Jiawei Zuo, Yuhang Zhang 0006, Yanfeng Hu
IEEE Geosci. Remote. Sens. Lett.6
2008 Contextual Models for Automatic Building Extraction in High Resolution Remote Sensing Image Using Object-Based Boosting Method
abstract
Many traditional target extraction methods encountered new challenges as the spatial resolution is increasing quickly. For the purpose of extracting buildings in that circumstance, a new method combing both the object-based approach and boosting algorithm is proposed in this paper. The method associates segmentation with recognition by constructing a hierarchical object network, which effectively improves the problem of detecting targets with a modifiable sliding window existed in other methods. And some useful features are selected automatically to train a validate classifier. Then the label confidence of each object is computed using contextual models to complete the extraction procedure. Competitive results for both multiform and complicated buildings demonstrate the precision, robustness and effectiveness of the proposed method.
Xian Sun 0001, Kun Fu 0001, Hui Long, Yanfeng Hu, Lun Cai
IGARSS (2)4