VLDB 2026 Research / reviewers in the wild / expert
Changhong Fu 0001
dblp:117/4527
· DBLP profile ↗
68ranked-venue papers
23as first author
43since 2021 · last 2025
0000-0002-9897-6022ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 54 · 17 first-author · 33 since 2021Systems, architecture and hardware · 36 · 13 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EdgeSR: Reparameterization-Driven Fast Thermal Super-Resolution for Edge Electro-Optical DeviceabstractSuper-resolution (SR) can greatly promote the development of edge electro-optical (EO) devices. However, most existing SR models struggle to simultaneously achieve effective thermal reconstruction and real-time inference on edge EO devices with limited computing resources. To address these issues, this work proposes a novel fast thermal SR model (EdgeSR) for edge EO devices. Specifically, reparameterized scale-integrated convolutions (RepSConv) are proposed to deeply explore high-frequency features, incorporating multi-scale information and enhancing the scale-awareness of the backbone during the training phase. Furthermore, an inter-active reparameterization module (IRM), combining historical high-frequency with low-frequency information, is introduced to guide the extraction of high-frequency features, ultimately boosting the high-quality reconstruction of thermal images. Edge EO deployment-oriented reparameterization (EEDR) is designed to reparameterize all modules into standard convolutions that are hardware-friendly for edge EO devices and onboard real-time inference. Additionally, a new benchmark for thermal SR on cityscapes (CS-TSR) is built. The experimental results on this benchmark show that, compared to state-of-the-art lightweight SR networks, EdgeSR delivers superior reconstruction quality and faster inference speed on edge EO devices. In real-world applications, EdgeSR exhibits robust performance on edge EO devices, making it suitable for real-world deployment. The code and demo is available at https://github.com/vision4robotics/EdgeSR. Changhong Fu 0001, Zijie Zhang 0006, Haobo Zuo |
IROS | 1 |
| 2025 | EdgeSpotter: Multi-Scale Dense Text Spotting for Industrial Panel MonitoringabstractText spotting for industrial panels is a key task for intelligent monitoring. However, achieving efficient and accurate text spotting for complex industrial panels remains challenging due to issues such as cross-scale localization and ambiguous boundaries in dense text regions. Moreover, most existing methods primarily focus on representing a single text shape, neglecting a comprehensive exploration of multi-scale feature information across different texts. To address these issues, this work proposes a novel multi-scale dense text spotter for edge AI-based vision system (EdgeSpotter) to achieve accurate and robust industrial panel monitoring. Specifically, a novel Transformer with efficient mixer is developed to learn the interdependencies among multi-level features, integrating multi-layer spatial and semantic cues. In addition, a new feature sampling with Catmull-Rom splines is designed, which explicitly encodes the shape, position, and semantic information of text, thereby alleviating missed detections and reducing recognition errors caused by multi-scale or dense text regions. Furthermore, a new benchmark dataset for industrial panel monitoring (IPM) is constructed. Extensive qualitative and quantitative evaluations on this challenging benchmark dataset validate the superior performance of the proposed method in different challenging panel monitoring tasks. Finally, practical tests based on the self-designed edge AI-based vision system demonstrate the practicality of the method. The code and demo are available at https://github.com/vision4robotics/EdgeSpotter. Changhong Fu 0001, Haobo Zuo, Liangliang Yao |
IROS | 1 |
| 2025 | AnyTSR: Any-Scale Thermal Super-Resolution for UAVabstractThermal imaging can greatly enhance the application of intelligent unmanned aerial vehicles (UAV) in challenging environments. However, the inherent low resolution of thermal sensors leads to insufficient details and blurred boundaries. Super-resolution (SR) offers a promising solution to address this issue, while most existing SR methods are designed for fixed-scale SR. They are computationally expensive and inflexible in practical applications. To address above issues, this work proposes a novel any-scale thermal SR method (AnyTSR) for UAV within a single model. Specifically, a new image encoder is proposed to explicitly assign specific feature code to enable more accurate and flexible representation. Additionally, by effectively embedding coordinate offset information into the local feature ensemble, an innovative any-scale upsampler is proposed to better understand spatial relationships and reduce artifacts. Moreover, a novel dataset (UAV-TSR), covering both land and water scenes, is constructed for thermal SR tasks. Experimental results demonstrate that the proposed method consistently outperforms state-of-the-art methods across all scaling factors as well as generates more accurate and detailed high-resolution images. The code is located at https://github.com/vision4robotics/AnyTSR. Changhong Fu 0001, Zijie Zhang 0006, Haobo Zuo, Liangliang Yao |
IROS | 2 |
| 2025 | U-Snake: A Small-Sized Smart Underwater Snake RobotabstractWith the rapid development of AI chips, underwater snake robots hold significant promise for navigating complex underwater environments, offering unique advantages in exploration, monitoring, and inspection tasks due to their flexible body and high mobility. However, existing underwater snake robots predominantly employ bulky mechanical configurations with expensive manufacturing costs, resulting in excessive power consumption and limited operational endurance with standard batteries, which impede their widespread adoption and limit their operational flexibility. Moreover, most path following methods used in underwater snake robots inadequately account for the dynamic changes in path curvature, leading to serious tracking error in scenarios involving sharp turns or complex environments, which does not address the demands of more intricate trajectories. To address above issues, this work introduces the U-Snake, a small-sized smart underwater snake robot with a simple lightweight structure and a highly maneuverable controller, adapted to various complicated path following tasks. In particular, each joint of U-Snake is designed to be small and lightweight to achieve higher spatial utilization, which is covered by a convenient and efficient 3D printed waterproof casing to achieve robust water resistance. In addition, a path following method based on curvature of the path is designed to achieve high performance in various complicated trajectories. Furthermore, an integrated controller combines the method with the kinematics and dynamics models of U-Snake, enabling precise following of straight and curved paths. The experimental results demonstrate that the proposed control structure effectively guides U-Snake to follow the desired path. Haobo Zuo, Changhong Fu 0001 |
IROS | 3 |
| 2025 | Lattice Boltzmann Model for Learning Real-World Pixel DynamicityabstractThis work proposes the Lattice Boltzmann Model (LBM) to learn real-world pixel dynamicity for visual tracking.
LBM decomposes visual representations into dynamic pixel lattices and solves pixel motion states through collision-streaming processes.
Specifically, the high-dimensional distribution of the target pixels is acquired through a multilayer predict-update network to estimate the pixel positions and visibility. The predict stage formulates lattice collisions among the spatial neighborhood of target pixels and develops lattice streaming within the temporal visual context. The update stage rectifies the pixel distributions with online visual representations. Compared with existing methods, LBM demonstrates practical applicability in an online and real-time manner, which can efficiently adapt to real-world visual tracking tasks. Comprehensive evaluations of real-world point tracking benchmarks such as TAP-Vid and RoboTAP validate LBM's efficiency. A general evaluation of large-scale open-world object tracking benchmarks such as TAO, BFT, and OVT-B further demonstrates LBM's real-world practicality. Guangze Zheng 0001, Shijie Lin, Haobo Zuo, Si Si, Ming-Shan Wang, Changhong Fu 0001, Jia Pan 0001 |
NeurIPS | 6 |
| 2025 | AnyTSR++: Prompt-Oriented Any-Scale Thermal Super-Resolution for Unmanned Aerial VehicleabstractThermal imaging significantly augments the operational capabilities of intelligent unmanned aerial vehicles (UAVs) in complex environments. However, due to the limited resolution of onboard thermal sensors, thermal images captured by UAV suffer from insufficient detail and blurred object boundaries, thereby limiting their practicality. Although super-resolution (SR) provides a promising solution to this issue, existing any-scale SR methods adopt identical feature representations across all scales, lacking the ability to adaptively adjust features according to varying scale requirements, leading to suboptimal SR results. This issue becomes more pronounced in asymmetric scale SR, where the resolution differs significantly along different directions. To address these limitations, a novel prompt-oriented any-scale thermal SR method (AnyTSR++) is proposed for UAV. Specifically, a new image encoder is introduced to explicitly assign any-scale prompt, enabling more precise and adaptive feature representation. Furthermore, an innovative any-scale upsampler is designed by refining the coordinate offset and the local feature ensemble, enhancing spatial awareness and reducing artifacts. Additionally, a novel dataset (UAV-TSR++) comprising 24,000 images covering both land and water surface scenes is constructed to facilitate the community to conduct thermal SR research. Experimental results demonstrate that AnyTSR++ consistently outperforms state-of-the-art methods across both symmetric and asymmetric scaling factors, producing higher-resolution images with greater accuracy and more fine-grained details. The source code and new dataset are located at https://github.com/vision4robotics/AnyTSR++. Changhong Fu 0001, Zijie Zhang 0006, Haobo Zuo, Liangliang Yao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | NetTrack: Tracking Highly Dynamic Objects with a NetabstractThe complex dynamicity of open-world objects presents non-negligible challenges for multi-object tracking (MOT), often manifested as severe deformations, fast motion, and occlusions. Most methods that solely depend on coarse-grained object cues, such as boxes and the overall appearance of the object, are susceptible to degradation due to distorted internal relationships of dynamic objects. To address this problem, this work proposes Net Track, an efficient, generic, and affordable tracking framework to introduce fine-grained learning that is robust to dynamicity. Specifically, N etTrack constructs a dynamicity-aware association with a fine-grained Net, leveraging point-level visual cues. Correspondingly, a fine-grained sampler and matching method have been incorporated. Furthermore, NetTrack learns object-text correspondence for fine-grained localization. To evaluate MOT in extremely dynamic open-world scenarios, a bird flock tracking (BFT) dataset is constructed, which exhibits high dynamicity with diverse species and open-world scenarios. Comprehensive evaluation on BFT validates the effectiveness of fine-grained learning on object dynamicity, and thorough transfer experiments on challenging open-world benchmarks, i.e., TAO, TAO-OW, AnimalTrack, and GMOT-40, validate the strong generalization ability of NetTrack even without finetuning. Guangze Zheng 0001, Shijie Lin, Haobo Zuo, Changhong Fu 0001, Jia Pan 0001 |
CVPR | 4 |
| 2024 | Progressive Representation Learning for Real-Time UAV TrackingabstractVisual object tracking has significantly promoted autonomous applications for unmanned aerial vehicles (UAVs). However, learning robust object representations for UAV tracking is especially challenging in complex dynamic environments, when confronted with aspect ratio change and occlusion. These challenges severely alter the original information of the object. To handle the above issues, this work proposes a novel progressive representation learning framework for UAV tracking, i.e., PRL-Track. Specifically, PRL-Track is divided into coarse representation learning and fine representation learning. For coarse representation learning, two innovative regulators, which rely on appearance and semantic information, are designed to mitigate appearance interference and capture semantic information. Furthermore, for fine representation learning, a new hierarchical modeling generator is developed to intertwine coarse object representations. Exhaustive experiments demonstrate that the proposed PRL-Track delivers exceptional performance on three authoritative UAV tracking benchmarks. Real-world tests indicate that the proposed PRL-Track realizes superior tracking performance with 42.6 frames per second on the typical UAV platform equipped with an edge smart camera. The code, model, and demo videos are available at https://github.com/vision4robotics/PRL-Track. Changhong Fu 0001, Xiang Lei, Haobo Zuo, Liangliang Yao, Guangze Zheng 0001, Jia Pan 0001 |
IROS | 1 |
| 2024 | Prompt-Driven Temporal Domain Adaptation for Nighttime UAV TrackingabstractNighttime UAV tracking under low-illuminated scenarios has achieved great progress by domain adaptation (DA). However, previous DA training-based works are deficient in narrowing the discrepancy of temporal contexts for UAV trackers. To address the issue, this work proposes a prompt-driven temporal domain adaptation training framework to fully utilize temporal contexts for challenging nighttime UAV tracking, i.e., TDA. Specifically, the proposed framework aligns the distribution of temporal contexts from daytime and nighttime domains by training the temporal feature generator against the discriminator. The temporal-consistent discriminator progressively extracts shared domain-specific features to generate coherent domain discrimination results in the time series. Additionally, to obtain high-quality training samples, a prompt-driven object miner is employed to precisely locate objects in unannotated nighttime videos. Moreover, a new benchmark for long-term nighttime UAV tracking is constructed. Exhaustive evaluations on both public and self-constructed nighttime benchmarks demonstrate the remarkable performance of the tracker trained in TDA framework, i.e., TDA-Track. Real-world tests at nighttime also show its practicality. The code and demo videos are available at https://github.com/vision4robotics/TDA-Track. Changhong Fu 0001, Yiheng Wang 0001, Liangliang Yao, Guangze Zheng 0001, Haobo Zuo, Jia Pan 0001 |
IROS | 1 |
| 2024 | Intelligent Fish Detection System with Similarity-Aware TransformerabstractFish detection in water-land transfer has significantly contributed to the fishery. However, manual fish detection in crowd-collaboration performs inefficiently and expensively, involving insufficient accuracy. To further enhance the water-land transfer efficiency, improve detection accuracy, and reduce labor costs, this work designs a new type of lightweight and plug-and-play edge intelligent vision system to automatically conduct fast fish detection with high-speed camera. Moreover, a novel similarity-aware vision Transformer for fast fish detection (FishViT) is proposed to onboard identify every single fish in a dense and similar group. Specifically, a novel similarity-aware multi-level encoder is developed to enhance multi-scale features in parallel, thereby yielding discriminative representations for varying-size fish. Additionally, a new soft-threshold attention mechanism is introduced, which not only effectively eliminates background noise from images but also accurately recognizes both the edge details and overall features of different similar fish. 85 challenging video sequences with high framerate and high-resolution are collected to establish a benchmark from real fish water-land transfer scenarios. Exhaustive evaluation conducted with this challenging benchmark has proved the robustness and effectiveness of FishViT with over 80 FPS. Real work scenario tests validate the practicality of the proposed method. The code and demo video are available at https://github.com/vision4robotics/FishViT. Shengchen Li, Haobo Zuo, Changhong Fu 0001 |
IROS | 3 |
| 2024 | Conditional Generative Denoiser for Nighttime UAV TrackingabstractState-of-the-art (SOTA) visual object tracking methods have significantly enhanced the autonomy of unmanned aerial vehicles (UAVs). However, in low-light conditions, the presence of irregular real noise from the environments severely degrades the performance of these SOTA methods. Moreover, existing SOTA denoising techniques often fail to meet the real-time processing requirements when deployed as plug-and-play denoisers for UAV tracking. To address this challenge, this work proposes a novel conditional generative denoiser (CG-Denoiser), which breaks free from the limitations of traditional deterministic paradigms and generates the noise conditioning on the input, subsequently removing it. To better align the input dimensions and accelerate inference, a novel nested residual Transformer conditionalizer is developed. Furthermore, an innovative multi-kernel conditional refiner is designed to pertinently refine the denoised output. Extensive experiments show that CGDenoiser promotes the tracking precision of the SOTA tracker by 18.18% on DarkTrack2021 whereas working 5.8 times faster than the second well-performed denoiser. Real-world tests with complex challenges also prove the effectiveness and practicality of CGDenoiser. Code, video demo and supplementary proof for CGDenoier are now available at: https://github.com/vision4robotics/CGDenoiser. Yucheng Wang 0004, Changhong Fu 0001, Kunhan Lu, Liangliang Yao, Haobo Zuo |
IROS | 2 |
| 2024 | Enhancing Nighttime UAV Tracking with Light Distribution SuppressionabstractVisual object tracking has boosted extensive intelligent applications for unmanned aerial vehicles (UAVs). However, the state-of-the-art (SOTA) enhancers for nighttime UAV tracking always neglect the uneven light distribution in low-light images, inevitably leading to excessive enhancement in scenarios with complex illumination. To address these issues, this work proposes a novel enhancer, i.e., LDEnhancer, enhancing nighttime UAV tracking with light distribution suppression. Specifically, a novel image content refinement module is developed to decompose the light distribution information and image content information in the feature space, allowing for the targeted enhancement of the image content information. Then this work designs a new light distribution generation module to capture light distribution effectively. The features with light distribution information and image content information are fed into the different parameter estimation modules, respectively, for the parameter map prediction. Finally, leveraging two parameter maps, an innovative interweave iteration adjustment is proposed for the collaborative pixel-wise adjustment of low-light images. Additionally, a challenging nighttime UAV tracking dataset with uneven light distribution, namely NAT2024-2, is constructed to provide a comprehensive evaluation, which contains 40 challenging sequences with over 74K frames in total. Experimental results on the authoritative UAV benchmarks and the proposed NAT2024-2 demonstrate that LDEnhancer outperforms other SOTA low-light enhancers for nighttime UAV tracking. Furthermore, real-world tests on a typical UAV platform with an NVIDIA Orin NX confirm the practicality and efficiency of LDEnhancer. The code is available at https: //github.com/vision4robotics/LDEnhancer. Liangliang Yao, Changhong Fu 0001, Yiheng Wang 0001, Haobo Zuo, Kunhan Lu |
IROS | 2 |
| 2024 | DaDiff: Domain-aware Diffusion Model for Nighttime UAV TrackingabstractDomain adaptation is an inspiring solution to the misalignment issue of day/night image features for nighttime UAV tracking. However, the one-step adaptation paradigm is inadequate in addressing the prevalent difficulties posed by low-resolution (LR) objects when viewed from the UAVs at night, owing to the blurry edge contour and limited detail information. Moreover, these approaches struggle to perceive LR objects disturbed by nighttime noise. To address these challenges, this work proposes a novel progressive alignment paradigm, named domain-aware diffusion model (DaDiff), aligning nighttime LR object features to the daytime by virtue of progressive and stable generations. The proposed DaDiff includes an alignment encoder to enhance the detail information of nighttime LR objects, a tracking-oriented layer designed to achieve close collaboration with tracking tasks, and a successive distribution discriminator presented to distinguish different feature distributions at each diffusion timestep successively. Furthermore, an elaborate nighttime UAV tracking benchmark is constructed for LR objects, namely NUT-LR, consisting of 100 annotated sequences. Exhaustive experiments have demonstrated the robustness and feature alignment ability of the proposed DaDiff. The source code and video demo are available at https://github.com/vision4robotics/DaDiff. Haobo Zuo, Changhong Fu 0001, Guangze Zheng 0001, Liangliang Yao, Kunhan Lu, Jia Pan 0001 |
IROS | 2 |
| 2024 | Spatial Reliability Enhanced Correlation Filter: An Efficient Approach for Real-Time UAV TrackingabstractTraditional discriminative correlation filter (DCF) has received great popularity due to its high computational efficiency. However, the lightweight framework of DCF cannot promise robust performance when the tracker faces appearance variations within the background. These unpredictable appearance variations always distract the filter. Most existing DCF-based trackers either utilize deep convolutional features or incorporate additional constraints to elevate tracking robustness. Despite some improvements, both of them hamper the tracking speed and can only roughly alleviate the distractions of appearance variations. In this paper, a novel spatial reliability enhanced learning strategy is proposed to handle the problems aforementioned. By monitoring the variation of response produced in detection phase, a dynamic reliability map is generated to indicate the reliability of each background subregion. Then, label adjustment is conducted to repress the distractions of these unreliable areas. Compared with the conventional way of constraint where a new term is always added to realize the desired goal, label adjustment is simultaneously more efficient and effective. Moreover, to promise the accuracy and dependability of the reliability map, an adaptively updated response pool recording reliable historical response values is proposed. Extensive and exhaustive experiments on three challenging unmanned aerial vehicle (UAV) benchmarks, i.e., UAV123@10fps, DTB70 and UAVDT, which totally include 243 video sequences, validate the superiority of the proposed method against other state-of-the-art trackers and exhibit a remarkable generality in a variety of scenarios. Meanwhile, the tracking speed of 65.2FPS on a cheap CPU makes it suitable for real-time UAV applications. Changhong Fu 0001, Fangqiang Ding, Yiming Li 0003, Geng Lu |
IEEE Trans. Multim. | 1 |
| 2023 | PVT++: A Simple End-to-End Latency-Aware Visual Tracking FrameworkabstractVisual object tracking is essential to intelligent robots. Most existing approaches have ignored the online latency that can cause severe performance degradation during real-world processing. Especially for unmanned aerial vehicles (UAVs), where robust tracking is more challenging and onboard computation is limited, the latency issue can be fatal. In this work, we present a simple framework for end-to-end latency-aware tracking, i.e., end-to-end predictive visual tracking (PVT++). Unlike existing solutions that naively append Kalman Filters after trackers, PVT++ can be jointly optimized, so that it takes not only motion information but can also leverage the rich visual knowledge in most pre-trained tracker models for robust prediction. Besides, to bridge the training-evaluation domain gap, we propose a relative motion factor, empowering PVT++ to generalize to the challenging and complex UAV tracking scenes. These careful designs have made the small-capacity lightweight PVT++ a widely effective solution. Additionally, this work presents an extended latency-aware evaluation benchmark for assessing an any-speed tracker in the online setting. Empirical results on a robotic platform from the aerial perspective show that PVT++ can achieve significant performance gain on various trackers and exhibit higher accuracy than prior solutions, largely mitigating the degradation brought by latency. Our code is public at https: //github.com/Jaraxxus-Me/PVT_pp.git. Bowen Li 0007, Ziyuan Huang 0003, Junjie Ye 0004, Yiming Li 0003, Sebastian A. Scherer, Hang Zhao 0021, Changhong Fu 0001 |
ICCV | 7 |
| 2023 | Continuity-Aware Latent Interframe Information Mining for Reliable UAV TrackingabstractUnmanned aerial vehicle (UAV) tracking is crucial for autonomous navigation and has broad applications in robotic automation fields. However, reliable UAV tracking remains a challenging task due to various difficulties like frequent occlusion and aspect ratio change. Additionally, most of the existing work mainly focuses on explicit information to improve tracking performance, ignoring potential interframe connections. To address the above issues, this work proposes a novel framework with continuity-aware latent interframe information mining for reliable UAV tracking, i.e., ClimRT. Specifically, a new efficient continuity-aware latent interframe information mining network (ClimNet) is proposed for UAV tracking, which can generate highly-effective latent frame between two adjacent frames. Besides, a novel location-continuity Transformer (LCT) is designed to fully explore continuity-aware spatial-temporal information, thereby markedly enhancing UAV tracking. Extensive qualitative and quantitative experiments on three authoritative aerial benchmarks strongly validate the robustness and reliability of ClimRT in UAV tracking performance. Furthermore, real-world tests on the aerial platform validate its practicability and effectiveness. The code and demo materials are released at https://github.com/vision4robotics/ClimRT. Changhong Fu 0001, Mutian Cai, Sihang Li 0001, Kunhan Lu, Haobo Zuo, Chongjun Liu |
ICRA | 1 |
| 2023 | SGDViT: Saliency-Guided Dynamic Vision Transformer for UAV TrackingabstractVision-based object tracking has boosted extensive autonomous applications for unmanned aerial vehicles (UAVs). However, the dynamic changes in flight maneuver and viewpoint encountered in UAV tracking pose significant difficulties, e.g., aspect ratio change, and scale variation. The conventional cross-correlation operation, while commonly used, has limitations in effectively capturing perceptual similarity and incorporates extraneous background information. To mitigate these limitations, this work presents a novel saliency-guided dynamic vision Transformer (SGDViT) for UAV tracking. The proposed method designs a new task-specific object saliency mining network to refine the cross-correlation operation and effectively discriminate foreground and background information. Additionally, a saliency adaptation embedding operation dynamically generates tokens based on initial saliency, thereby reducing the computational complexity of the Transformer architecture. Finally, a lightweight saliency filtering Transformer further refines saliency information and increases the focus on appearance information. The efficacy and robustness of the proposed approach have been thoroughly assessed through experiments on three widely-used UAV tracking benchmarks and real-world scenarios, with results demonstrating its superiority. The source code and demo videos are available at https://github.com/vision4robotics/SGDViT. Liangliang Yao, Changhong Fu 0001, Sihang Li 0001, Guangze Zheng 0001, Junjie Ye 0004 |
ICRA | 2 |
| 2023 | An Open-Source Robotic Chinese Chess PlayerabstractConsumer robots can accompany children growing up, improving their abilities while playing and entertaining. This paper presents an open-source, practical, low-cost robotic Chinese chess player. The proposed system includes an elaborate mechanical structure, a simple kinematic solution, a novel robot operating system, real-time and accurate chess recognition. Regarding its mechanical design, it combines a magnetism structure and mechanical cam drive, while the overall system has just three servo motors. At the same time, its control strategy is simple and effective. Furthermore, a lightweight robot message communication mechanism, entitled TinyROS, is developed for computing resource-limited embedded chips. Concerning the recognition process, our CNNbased object detector determines chess and achieves accurate identification. As a result, our robotic Chinese chess player is exquisite and easy for large-scale promotion while improving users' chess skills. Aiming to facilitate future consumer robot research and popularize customer robots, the model's mechanical and software design and the TinyROS protocol are open-sourced at https://github.com/Star-Robot/chinese-chess-robot. Shan An, Guangfu Che, Jinghao Guo, Konstantinos A. Tsintotas, Fukai Zhang, Junjie Ye 0004, Changhong Fu 0001, Haogang Zhu, Hong Zhang 0013 |
IROS | 9 |
| 2023 | Towards Real-World Visual Tracking With Temporal ContextsabstractVisual tracking has made significant improvements in the past few decades. Most existing state-of-the-art trackers 1) merely aim for performance in ideal conditions while overlooking the real-world conditions; 2) adopt the tracking-by-detection paradigm, neglecting rich temporal contexts; 3) only integrate the temporal information into the template, where temporal contexts among consecutive frames are far from being fully utilized. To handle those problems, we propose a two-level framework (TCTrack) that can exploit temporal contexts efficiently. Based on it, we propose a stronger version for real-world visual tracking, i.e., TCTrack++. It boils down to two levels: features and similarity maps. Specifically, for feature extraction, we propose an attention-based temporally adaptive convolution to enhance the spatial features using temporal information, which is achieved by dynamically calibrating the convolution weights. For similarity map refinement, we introduce an adaptive temporal transformer to encode the temporal knowledge efficiently and decode it for the accurate refinement of the similarity map. To further improve the performance, we additionally introduce a curriculum learning strategy. Also, we adopt online evaluation to measure performance in real-world conditions. Exhaustive experiments on 8 well-known benchmarks demonstrate the superiority of TCTrack++. Real-world tests directly verify that TCTrack++ can be readily used in real-world applications. Ziang Cao, Ziyuan Huang 0003, Liang Pan, Shiwei Zhang 0001, Ziwei Liu 0002, Changhong Fu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Scale-Aware Siamese Object Tracking for Vision-Based UAM ApproachingabstractIn many industrial applications of unmanned aerial manipulator (UAM), visual approaching the object is crucial to subsequent manipulating. In comparison with the widely-studied manipulating, the key to efficient vision-based UAM approaching, i.e., UAM object tracking, is still limited. Since traditional model-based UAM tracking is costly and cannot track arbitrary objects, an intuitive solution is to introduce state-of-the-art model-free Siamese trackers from the visual tracking field. Although Siamese tracking is most suitable for the onboard embedded processors, severe object scale variation in UAM tracking brings formidable challenges. To address these problems, this work proposes a novel model-free scale-aware Siamese tracker (SiamSA). Specifically, a scale attention network is proposed to emphasize scale awareness in feature processing. A scale-aware anchor proposal network is designed to achieve anchor proposing. Besides, two novel UAM tracking benchmarks are first recorded. Comprehensive experiments on benchmarks validate the effectiveness of SiamSA. Furthermore, real-world tests also confirm practicality for industrial UAM approaching tasks with high efficiency and robustness. Guangze Zheng 0001, Changhong Fu 0001, Junjie Ye 0004, Bowen Li 0007, Geng Lu, Jia Pan 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | All-Day Object Tracking for Unmanned Aerial VehicleabstractUnmanned aerial vehicle (UAV) has facilitated a wide range of real-world applications and attracted extensive research in the mobile computing field. Specially, developing real-time robust visual onboard trackers for all-day aerial maneuver can remarkably broaden the scope of intelligent deployment of UAV. However, prior tracking methods have merely focused on robust tracking in the well-illuminated scenes, while ignoring trackers’ capabilities to be deployed in the dark. In darkness, the conditions can be more complex and harsh, easily posing inferior robust tracking or even tracking failure. To this end, this work proposes a novel discriminative correlation filter-based tracker with illumination adaptive and anti-dark capability, namely ADTrack. ADTrack firstly exploits image illuminance information to enable adaptability of the model to the given light condition. Then, by virtue of an efficient enhancer, ADTrack carries out image pretreatment where a target aware mask is generated. Benefiting from the mask, ADTrack aims to solve a novel dual regression problem where dual filters are online trained with mutual constraint. Besides, this work also constructs a UAV nighttime tracking benchmark UAVDark135. Exhaustive experiments on authoritative benchmarks and onboard tests are implemented to validate the superiority and robustness of ADTrack in all-day conditions. Bowen Li 0007, Changhong Fu 0001, Fangqiang Ding, Junjie Ye 0004, Fuling Lin |
IEEE Trans. Mob. Comput. | 2 |
| 2022 | TCTrack: Temporal Contexts for Aerial TrackingabstractTemporal contexts among consecutive frames are far from being fully utilized in existing visual trackers. In this work, we present TCTrack11https://github.com/vision4robotics/TCTrack, a comprehensive framework to fully exploit temporal contexts for aerial tracking. The temporal contexts are incorporated at two levels: the extraction of features and the refinement of similarity maps. Specifically, for feature extraction, an online temporally adaptive convolution is proposed to enhance the spatial features using temporal information, which is achieved by dynamically calibrating the convolution weights according to the previous frames. For similarity map refinement, we propose an adaptive temporal transformer, which first effectively encodes temporal knowledge in a memory-efficient way, before the temporal knowledge is decoded for accurate adjustment of the similarity map. TCTrack is effective and efficient: evaluation on four aerial tracking benchmarks shows its impressive performance; real-world UAV tests show its high speed of over 27 FPS on NVIDIA Jetson AGX Xavier. Ziang Cao, Ziyuan Huang 0003, Liang Pan, Shiwei Zhang 0001, Ziwei Liu 0002, Changhong Fu 0001 |
CVPR | 6 |
| 2022 | Unsupervised Domain Adaptation for Nighttime Aerial TrackingabstractPrevious advances in object tracking mostly reported on favorable illumination circumstances while neglecting performance at nighttime, which significantly impeded the development of related aerial robot applications. This work instead develops a novel unsupervised domain adaptation framework for nighttime aerial tracking (named UDAT). Specifically, a unique object discovery approach is provided to generate training patches from raw nighttime tracking videos. To tackle the domain discrepancy, we employ a Transformer-based bridging layer post to the feature extractor to align image features from both domains. With a Transformer day/night feature discriminator, the day-time tracking model is adversarially trained to track at night. Moreover, we construct a pioneering benchmark namely NAT2021 for unsupervised domain adaptive night-time tracking, which comprises a test set of 180 manually annotated tracking sequences and a train set of over 276k unlabelled nighttime tracking frames. Exhaustive experiments demonstrate the robustness and domain adaptability of the proposed framework in nighttime aerial tracking. The code and benchmark are available at https://github.com/vision4robotics/UDAT. Junjie Ye 0004, Changhong Fu 0001, Guangze Zheng 0001, Danda Pani Paudel, Guang Chen 0001 |
CVPR | 2 |
| 2022 | Ad2Attack: Adaptive Adversarial Attack on Real-Time UAV TrackingabstractVisual tracking is adopted to extensive unmanned aerial vehicle (UAV)-related applications, which leads to a highly demanding requirement on the robustness of UAV trackers. However, adding imperceptible perturbations can easily fool the tracker and cause tracking failures. This risk is often overlooked and rarely researched at present. Therefore, to help increase awareness of the potential risk and the robustness of UAV tracking, this work proposes a novel adaptive adversarial attack approach, i.e., Ad2Attack, against UAV object tracking. Specifically, adversarial examples are generated online during the resampling of the search patch image, which leads trackers to lose the target in the following frames. Ad2Attack is composed of a direct downsampling module and a super-resolution upsampling module with adaptive stages. A novel optimization function is proposed for balancing the imperceptibility and efficiency of the attack. Comprehensive experiments on several well-known benchmarks and real-world conditions show the effectiveness of our attack method, which dramatically reduces the performance of the most advanced Siamese trackers. Changhong Fu 0001, Sihang Li 0001, Xinnan Yuan, Junjie Ye 0004, Ziang Cao, Fangqiang Ding |
ICRA | 1 |
| 2022 | Siamese Object Tracking for Vision-Based UAM Approaching with Pairwise Scale-Channel AttentionabstractAlthough the manipulating of the unmanned aerial manipulator (UAM) has been widely studied, vision-based UAM approaching, which is crucial to the subsequent manipulating, generally lacks effective design. The key to the visual UAM approaching lies in object tracking, while current UAM tracking typically relies on costly model-based methods. Besides, UAM approaching often confronts more severe object scale variation issues, which makes it inappro-priate to directly employ state-of-the-art model-free Siamese-based methods from the object tracking field. To address the above problems, this work proposes a novel Siamese network with pairwise scale-channel attention (SiamSA) for vision-based UAM approaching. Specifically, SiamSA consists of a pairwise scale-channel attention network (PSAN) and a scale-aware anchor proposal network (SA-APN). PSAN acquires valuable scale information for feature processing, while SA-APN mainly attaches scale awareness to anchor proposing. Moreover, a new tracking benchmark for UAM approaching, namely UAMT100, is recorded with 35K frames on a flying UAM platform for evaluation. Exhaustive experiments on the benchmarks and real-world tests validate the efficiency and practicality of SiamSA with a promising speed. Both the code and UAMT100 benchmark are now available at https://github.com/vision4robotics/SiamSA. Guangze Zheng 0001, Changhong Fu 0001, Junjie Ye 0004, Bowen Li 0007, Geng Lu, Jia Pan 0001 |
IROS | 2 |
| 2022 | HighlightNet: Highlighting Low-Light Potential Features for Real-Time UAV TrackingabstractLow-light environments have posed a formidable challenge for robust unmanned aerial vehicle (UAV) tracking even with state-of-the-art (SOTA) trackers since the poten-tial image features are hard to extract under adverse light conditions. Besides, due to the low visibility, accurate online selection of the object also becomes extremely difficult for human monitors to initialize UAV tracking in ground con-trol stations. To solve these problems, this work proposes a novel enhancer, i.e., HighlightNet, to light up potential objects for both human operators and UAV trackers. By employing Transformer, HighlightNet can adjust enhancement parameters according to global features and is thus adaptive for the illumination variation. Pixel-level range mask is introduced to make HighlightNet more focused on the enhancement of the tracking object and regions without light sources. Furthermore, a soft truncation mechanism is built to prevent background noise from being mistaken for crucial features. Evaluations on image enhancement benchmarks demonstrate HighlightNet has advantages in facilitating human perception. Experiments on the public UAVDark135 benchmark show that HightlightNet is more suitable for UAV tracking tasks than other state-of-the-art (SOTA) low-light enhancers. In addition, real-world tests on a typical UAV platform verify HightlightNet's practicability and efficiency in nighttime aerial tracking-related applications. The code and demo videos are available at https://github.com/vision4robotics/HighlightNet. Changhong Fu 0001, Haolin Dong, Junjie Ye 0004, Guangze Zheng 0001, Sihang Li 0001, Jilin Zhao |
IROS | 1 |
| 2022 | Local Perception-Aware Transformer for Aerial TrackingabstractTransformer-based visual object tracking has been utilized extensively. However, the Transformer structure is lack of enough inductive bias. In addition, only focusing on encoding the global feature does harm to modeling local details, which restricts the capability of tracking in aerial robots. Specifically, with local-modeling to global-search mechanism, the proposed tracker replaces the global encoder by a novel local-recognition encoder. In the employed encoder, a local-recognition attention and a local element correction network are carefully designed for reducing the global redundant information interference and increasing local inductive bias. Meanwhile, the latter can model local object details precisely under aerial view through detail-inquiry net. The proposed method achieves competitive accuracy and robustness in several authoritative aerial benchmarks with 316 sequences in total. The proposed tracker's practicability and efficiency have been validated by the real-world tests. The source code is available at https://github.com/vision4robotics/LPAT. Changhong Fu 0001, Weiyu Peng, Sihang Li 0001, Junjie Ye 0004, Ziang Cao |
IROS | 1 |
| 2022 | End-to-End Feature Decontaminated Network for UAV TrackingabstractObject feature pollution is one of the burning issues in vision-based UAV tracking, commonly caused by occlusion, fast motion, and illumination variation. Due to the contaminated information in the polluted object features, most trackers fail to precisely estimate the object location and scale. To address the above disturbing issue, this work proposes a novel end-to-end feature decontaminated network for efficient and effective UAV tracking, i.e., FDNT. FDNT mainly includes two modules: a decontaminated downsampling network and a decontaminated upsampling network. The former reduces the interference information of the feature pollution and enhanced the expression of the object location information with two asymmetric convolution branches. The latter restores the object scale information with the super-resolution technology-based low-to-high encoder, achieving a further decontamination effect. Moreover, a novel pooling distance loss is carefully developed to assist the decontaminated downsampling network in concentrating on the critical regions with the object information. Exhaustive experiments on three well-known benchmarks validate the effectiveness of FDNT, especially on the sequences with feature pollution. In addition, real-world tests show the efficiency of FDNT with 31.4 frames per second. The code and demo videos are available at https://github.com/vision4robotics/FDNT. Haobo Zuo, Changhong Fu 0001, Sihang Li 0001, Junjie Ye 0004, Guangze Zheng 0001 |
IROS | 2 |
| 2022 | Onboard Real-Time Aerial Tracking With Efficient Siamese Anchor Proposal NetworkabstractObject tracking approaches based on the Siamese network have demonstrated their huge potential in the remote sensing field recently. Nevertheless, due to the limited computing resource of aerial platforms and special challenges in aerial tracking, most existing Siamese-based methods can hardly meet the real-time and state-of-the-art performance simultaneously. Consequently, a novel Siamese-based method is proposed in this work for onboard real-time aerial tracking, i.e., SiamAPN. The proposed method is a no-prior two-stage method, i.e., Stage-1 for proposing adaptive anchors to enhance the ability of object perception and Stage-2 for fine-tuning the proposed anchors to obtain accurate results. Distinct from the traditional predefined anchors, the proposed anchors can adapt automatically to the tracking object. Besides, the internal information of adaptive anchors is utilized to feedback SiamAPN for enhancing the object perception. Attributing to the feature fusion network, different semantic information is integrated, enriching the information flow that is significant for robust aerial tracking. In the end, the regression and multiclassification operation refine the proposed anchors meticulously. Comprehensive evaluations on three well-known aerial tracking benchmarks have proven the superior performance of the presented approach. Moreover, to verify the practicability of the proposed method, SiamAPN is implemented onboard a typical embedded aerial tracking platform to conduct the real-world evaluations on specific aerial tracking scenarios, e.g., fast motion, long-term tracking, and low resolution. The results have demonstrated the efficiency and accuracy of the proposed approach, with a processing speed of over 30 frames/s. In addition, the image sequences in the real-world evaluations are collected and annotated as a new aerial tracking benchmark, i.e., UAVTrack112. Changhong Fu 0001, Ziang Cao, Yiming Li 0003, Junjie Ye 0004, Chen Feng 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | DeconNet: End-to-End Decontaminated Network for Vision-Based Aerial TrackingabstractVision-based aerial tracking has proven enormous potential in the field of remote sensing recently. However, challenges such as occlusion, fast motion, and illumination variation remain crucial issues for realistic aerial tracking applications. These challenges, frequently occurring from the aerial perspectives, can easily cause object feature pollution. With the contaminated object features, the credibility of trackers is prone to be substantially degraded. To address this issue, this work proposes a novel end-to-end decontaminated network, i.e., DeconNet, to alleviate object feature pollution efficiently and effectively. DeconNet mainly consists of downsampling and upsampling phases. Specifically, the decontaminated downsampling network first decreases the polluted object information with two convolution branches, enhancing the object location information. Subsequently, the decontaminated upsampling network applies the super-resolution technology to restore the object scale and shape information, with the low-to-high (LTH) encoder for further decontamination. In addition, the pooling distance (PD) loss function is carefully designed to improve the decontamination effect of the decontaminated downsampling network. Comprehensive evaluations on four well-known aerial tracking benchmarks validate the effectiveness of DeconNet. Especially, the proposed tracker has superior performance on the sequences with feature pollution. Besides, real-world tests on an aerial platform have proven the efficiency of DeconNet with 30.6 fps. Haobo Zuo, Changhong Fu 0001, Sihang Li 0001, Junjie Ye 0004, Guangze Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | ReCF: Exploiting Response Reasoning for Correlation Filters in Real-Time UAV TrackingabstractObject tracking is a fundamental task for the visual perception system on the intelligent unmanned aerial vehicle (UAV). The high efficiency of correlation filter (CF) based trackers has advanced the widespread development of online UAV object tracking. This kind of method can effectively train a filter to discriminate the target from the background. However, most CF-based methods require a fixed label function over all the previous samples, leading to over-fitting and filter degradation, especially in complex drone scenarios. To address this problem, a novel adaptive response reasoning approach is proposed for CF learning. It can leverage temporal information in filter training and significantly promote the robustness of the tracker. Specifically, the proposed response reasoning method goes beyond the standard response consistency requirement and constructs an auxiliary label of the current sample. Besides, it helps learn a generic relationship between the previous and current filters, thereby realizing self-regulated filter updating and enhancing the discriminability of the filter. Extensive experiments on four well-known challenging UAV tracking benchmarks with 278 videos sequences show that the presented method yields superior results to 40 state-of-the-art trackers with real-time performance on a single CPU, which is suitable for UAV online tracking missions. Fuling Lin, Changhong Fu 0001, Yujie He 0002, Weijiang Xiong |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | HiFT: Hierarchical Feature Transformer for Aerial TrackingabstractMost existing Siamese-based tracking methods execute the classification and regression of the target object based on the similarity maps. However, they either employ a single map from the last convolutional layer which degrades the localization accuracy in complex scenarios or separately use multiple maps for decision making, introducing intractable computations for aerial mobile platforms. Thus, in this work, we propose an efficient and effective hierarchical feature transformer (HiFT) for aerial tracking. Hierarchical similarity maps generated by multi-level convolutional layers are fed into the feature transformer to achieve the interactive fusion of spatial (shallow layers) and semantics cues (deep layers). Consequently, not only the global contextual information can be raised, facilitating the target search, but also our end-to-end architecture with the transformer can efficiently learn the interdependencies among multi-level features, thereby discovering a tracking-tailored feature space with strong discriminability. Comprehensive evaluations on four aerial benchmarks have proven the effectiveness of HiFT. Real-world tests on the aerial platform have strongly validated its practicability with a real-time speed. Our code is available at https://github.com/vision4robotics/HiFT. Ziang Cao, Changhong Fu 0001, Junjie Ye 0004, Bowen Li 0007, Yiming Li 0003 |
ICCV | 2 |
| 2021 | Siamese Anchor Proposal Network for High-Speed Aerial TrackingabstractIn the domain of visual tracking, most deep learning-based trackers highlight the accuracy but casting aside efficiency. Therefore, their real-world deployment on mobile platforms like the unmanned aerial vehicle (UAV) is impeded. In this work, a novel two-stage Siamese network-based method is proposed for aerial tracking, i.e., stage-1 for high-quality anchor proposal generation, stage-2 for refining the anchor proposal. Different from anchor-based methods with numerous pre-defined fixed-sized anchors, our no-prior method can 1) increase the robustness and generalization to different objects with various sizes, especially to small, occluded, and fast-moving objects, under complex scenarios in light of the adaptive anchor generation, 2) make calculation feasible due to the substantial decrease of anchor numbers. In addition, compared to anchor-free methods, our framework has better performance owing to refinement at stage-2. Comprehensive experiments on three benchmarks have proven the superior performance of our approach, with a speed of ∼200 frames/s. Changhong Fu 0001, Ziang Cao, Yiming Li 0003, Junjie Ye 0004, Chen Feng 0002 |
ICRA | 1 |
| 2021 | Online Recommendation-based Convolutional Features for Scale-Aware Visual TrackingabstractIn this paper, we develop an online learning-based visual tracking framework that can optimize the target model and estimate the scale variation for object tracking. We propose a recommender-based tracker, which is capable of selecting the representative convolutional neural network (CNN) layers and feature maps autonomously. In addition, the proposed recommender computes the weights of these layers and feature maps. A discriminative target percept of each recommended layer is reconstructed by the weighted sum of the recommended feature maps. Then the target model of the correlation filter is updated by the weighted sum of the target percepts. Thus, a sub-network is extracted from the pre-trained CNN backbone for the tracking process of a specific target. To deal with scale changes, we propose a spatiotemporal-based min-channel method to estimate the target size variation directly from CNN features. Experimental results on 50 benchmark datasets and video data from a rescue drone demonstrate that the proposed tracker is quite competitive with the state-of-the-art CNN-based trackers in terms of accuracy, scale adaptation, and robustness for UAV-related applications. Ran Duan 0002, Changhong Fu 0001, Kostas Alexis, Erdal Kayacan |
ICRA | 2 |
| 2021 | ADTrack: Target-Aware Dual Filter Learning for Real-Time Anti-Dark UAV TrackingabstractPrior correlation filter (CF)-based tracking methods for unmanned aerial vehicles (UAVs) have virtually focused on tracking in the daytime. However, when the night falls, the trackers will encounter more harsh scenes, which can easily lead to tracking failure. In this regard, this work proposes a novel tracker with anti-dark function (ADTrack). The proposed method integrates an efficient and effective low-light image enhancer into a CF-based tracker. Besides, a target-aware mask is simultaneously generated by virtue of image illumination variation. The target-aware mask can be applied to jointly train a target-focused filter that assists the context filter for robust tracking. Specifically, ADTrack adopts dual regression, where the context filter and the target-focused filter restrict each other for dual filter learning. Exhaustive experiments are conducted on typical dark sceneries benchmark, consisting of 37 typical night sequences from authoritative benchmarks, i.e., UAVDark, and our newly constructed benchmark UAVDark70. The results have shown that ADTrack favorably outperforms other state-of-the-art trackers and achieves a real-time speed of 34 frames/s on a single CPU, greatly extending robust UAV tracking to night scenes. Bowen Li 0007, Changhong Fu 0001, Fangqiang Ding, Junjie Ye 0004, Fuling Lin |
ICRA | 2 |
| 2021 | Mutation Sensitive Correlation Filter for Real-Time UAV Tracking with Adaptive Hybrid LabelabstractUnmanned aerial vehicle (UAV) based visual tracking has been confronted with numerous challenges, e.g., object motion and occlusion. These challenges generally introduce unexpected mutations of target appearance and result in tracking failure. However, prevalent discriminative correlation filter (DCF) based trackers are insensitive to target mutations due to a predefined label, which concentrates on merely the centre of the training region. Meanwhile, appearance mutations caused by occlusion or similar objects usually lead to the inevitable learning of wrong information. To cope with appearance mutations, this paper proposes a novel DCF-based method to enhance the sensitivity and resistance to mutations with an adaptive hybrid label, i.e., MSCF. The ideal label is optimized jointly with the correlation filter and remains temporal consistency. Besides, a novel measurement of mutations called mutation threat factor (MTF) is applied to correct the label dynamically. Considerable experiments are conducted on widely used UAV benchmarks. The results indicate that the performance of MSCF tracker surpasses other 26 state-of-the- art DCF-based and deep-based trackers. With a real-time speed of ~38 frames/s, the proposed approach is sufficient for UAV tracking commissions. Guangze Zheng 0001, Changhong Fu 0001, Junjie Ye 0004, Fuling Lin, Fangqiang Ding |
ICRA | 2 |
| 2021 | Real-Time Monocular Human Depth Estimation and Segmentation on Embedded SystemsabstractEstimating a scene’s depth to achieve collision avoidance against moving pedestrians is a crucial and fundamental problem in the robotic field. This paper proposes a novel, low complexity network architecture for fast and accurate human depth estimation and segmentation in indoor environments, aiming to applications for resource-constrained platforms (including battery-powered aerial, micro-aerial, and ground vehicles) with a monocular camera being the primary perception module. Following the encoder-decoder structure, the proposed framework consists of two branches, one for depth prediction and another for semantic segmentation. Moreover, network structure optimization is employed to improve its forward inference speed. Exhaustive experiments on three self-generated datasets prove our pipeline’s capability to execute in real-time, achieving higher frame rates than contemporary state-of-the-art frameworks (114.6 frames per second on an NVIDIA Jetson Nano GPU with TensorRT) while maintaining comparable accuracy. Shan An, Fangru Zhou, Haogang Zhu, Changhong Fu 0001, Konstantinos A. Tsintotas |
IROS | 5 |
| 2021 | SiamAPN++: Siamese Attentional Aggregation Network for Real-Time UAV TrackingabstractRecently, the Siamese-based method has stood out from multitudinous tracking methods owing to its state-of-the-art (SOTA) performance. Nevertheless, due to various special challenges in UAV tracking, e.g., severe occlusion and fast motion, most existing Siamese-based trackers hardly combine superior performance with high efficiency. To this concern, in this paper, a novel attentional Siamese tracker (SiamAPN++) is proposed for real-time UAV tracking. By virtue of the attention mechanism, we conduct a special attentional aggregation network (AAN) consisting of self-AAN and cross-AAN for raising the representation ability of features eventually. The former AAN aggregates and models the self-semantic interdependencies of the single feature map via spatial and channel dimensions. The latter aims to aggregate the cross-interdependencies of two different semantic features including the location information of anchors. In addition, the anchor proposal network based on dual features is proposed to raise its robustness of tracking objects with various scales. Experiments on two well-known authoritative benchmarks are conducted, where SiamAPN++ outperforms its baseline SiamAPN and other SOTA trackers. Besides, real-world tests onboard a typical embedded platform demonstrate that SiamAPN++ achieves promising tracking results with real-time speed. Ziang Cao, Changhong Fu 0001, Junjie Ye 0004, Bowen Li 0007, Yiming Li 0003 |
IROS | 2 |
| 2021 | DarkLighter: Light Up the Darkness for UAV TrackingabstractRecent years have witnessed the fast evolution and promising performance of the convolutional neural network (CNN)-based trackers, which aim at imitating biological visual systems. However, current CNN-based trackers can hardly generalize well to low-light scenes that are commonly lacked in the existing training set. In indistinguishable night scenarios frequently encountered in unmanned aerial vehicle (UAV) tracking-based applications, the robustness of the state-of-the-art (SOTA) trackers drops significantly. To facilitate aerial tracking in the dark through a general fashion, this work proposes a low-light image enhancer namely DarkLighter, which dedicates to alleviate the impact of poor illumination and noise iteratively. A lightweight map estimation network, i.e., ME-Net, is trained to efficiently estimate illumination maps and noise maps jointly. Experiments are conducted with several SOTA trackers on numerous UAV dark tracking scenes. Exhaustive evaluations demonstrate the reliability and universality of DarkLighter, with high efficiency. Moreover, DarkLighter has further been implemented on a typical UAV system. Real-world tests at night scenes have verified its practicability and dependability. Junjie Ye 0004, Changhong Fu 0001, Guangze Zheng 0001, Ziang Cao, Bowen Li 0007 |
IROS | 2 |
| 2021 | Learning dynamic regression with automatic distractor repression for real-time UAV tracking
Changhong Fu 0001, Fangqiang Ding, Yiming Li 0003, Chen Feng 0002 |
Eng. Appl. Artif. Intell. | 1 |
| 2021 | Learning Temporary Block-Based Bidirectional Incongruity-Aware Correlation Filters for Efficient UAV Object TrackingabstractIn the field of UAV object tracking, correlation filter based approaches have received lots of attention due to their computational efficiency. The methods learn filters by the ridge regression and generate response maps to distinguish the specified target from the background. An ideal filter can predict the object's position in a new frame, and in turn, can backtrack the object in the past frames. However, the neglect of tracking reversibility in most methods limits the potential of using inter-frame information to improve performance. In this work, a novel bidirectional incongruity-aware correlation filter is presented based on the nature of tracking reversibility. The proposed method incorporates the response-based bidirectional incongruity, which represents the gap between the filters' discriminative difference in the forward and backward tracking perspective caused by object appearance changes. It enables the filter not only to inherit the discriminability from previous filters but also to enhance the generalization capability to unpredictable appearance variations in upcoming frames. Moreover, a temporary block-based strategy is introduced to empower the filter accommodate more drastic object appearance changes and make more effective use of inter-frame information. Comprehensive experiments are conducted on three challenging UAV tracking benchmarks, including UAV123@10fps, DTB70, and UAVDT. Experimental results indicate that the proposed method has superior performance compared with the other 34 state-of-the-art trackers. Our approach permits real-time performance at ~46.8 FPS on a single CPU and is suitable for UAV online tracking applications. Fuling Lin, Changhong Fu 0001, Yujie He 0002, Fuyu Guo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Disruptor-Aware Interval-Based Response Inconsistency for Correlation Filters in Real-Time Aerial TrackingabstractAerial object tracking approaches based on discriminative correlation filter (DCF) have attracted wide attention in the tracking community due to their impressive progress recently. Many studies introduce temporal regularization into the DCF-based framework to achieve a more robust appearance model and further enhance the tracking performance. However, existing temporal regularization approaches usually utilize the information of two consecutive frames, which are not robust enough due to limited information. Although some methods attempt to incorporate abundant training samples and generally improve the tracking performance, these improvements are at the expense of significantly increased computing consumption. Besides, most existing methods introduce historical information directly without denoising, which means that background noises are also introduced into the filter training and may degrade the tracking accuracy. To tackle the drawbacks mentioned earlier, this work proposes a novel aerial object tracking approach to exploit disruptor-aware interval-based response inconsistency, i.e., IBRI tracker. The proposed method is able to incorporate historical interval information by utilizing responses in the filter training process, thereby obtaining a robust tracking performance while maintaining the real-time speed. Moreover, to reduce the disruptions caused by similar object, partial occlusion, and other challenging scenes, a novel disruptor-aware scheme based on response bucketing is introduced to detect the disruptor and enforce a spatial penalty for the disruptive area around the tracked object. Exhausted experiments on multiple well-known challenging aerial tracking benchmarks demonstrate the accuracy and robustness of the proposed IBRI tracker against other 35 state-of-the-art trackers. With a real-time speed of ~32 frames/s on a single CPU, the proposed approach can be applied for typical aerial platforms to achieve aerial visual object tracking efficiently. Changhong Fu 0001, Junjie Ye 0004, Juntao Xu, Yujie He 0002, Fuling Lin |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Intermittent Contextual Learning for Keyfilter-Aware UAV Object Tracking Using Deep Convolutional FeatureabstractVisual tracking, one of the most favorable multimedia applications, has been widely used in unmanned aerial vehicle (UAV) for civil infrastructure monitoring, aerial cinematography, autonomous navigation, etc. Most existing trackers utilize deep convolutional feature to enhance tracking robustness in scenarios of various appearance variation. However, they commonly neglect speed which is crucial for UAV with restricted calculation resources. In this work, a novel correlation filter-based keyfilter-aware tracker with a new intermittent context learning strategy is proposed to efficiently and effectively alleviate the problems of background clutter, deficient description, occlusion, illumination change, etc. Specifically, context information is utilized to empower the filter higher discriminating ability through response repression of the omnidirectional context patches. Furthermore, keyfilter is produced from the periodically selected keyframe. The latest produced keyfilter is used to restrain the current filter's corrupted changes. Most importantly, context learning of correlation filter is implemented intermittently to fully increase the tracking efficiency. This intermittent learning strategy can ensure every filter maintain context awareness owing to the restriction of keyfilter, periodically enhancing the context awareness. Substantial experiments on three challenging UAV benchmarks totally with 213 image sequences have shown that our tracker surpasses the state-of-the-art results, and exhibits a remarkable generality in short-term and long-term UAV tracking tasks as well as a variety of challenging attributes. Yiming Li 0003, Changhong Fu 0001, Ziyuan Huang 0003, Yinqiang Zhang, Jia Pan 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | AutoTrack: Towards High-Performance Visual Tracking for UAV With Automatic Spatio-Temporal RegularizationabstractMost existing trackers based on discriminative correlation filters (DCF) try to introduce predefined regularization term to improve the learning of target objects, e.g., by suppressing background learning or by restricting change rate of correlation filters. However, predefined parameters introduce much effort in tuning them and they still fail to adapt to new situations that the designer did not think of. In this work, a novel approach is proposed to online automatically and adaptively learn spatio-temporal regularization term. Spatially local response map variation is introduced as spatial regularization to make DCF focus on the learning of trust-worthy parts of the object, and global response map variation determines the updating rate of the filter. Extensive experiments on four UAV benchmarks have proven the superiority of our method compared to the state-of-the-art CPU- and GPU-based trackers, with a speed of ~60 frames per second running on a single CPU. Our tracker is additionally proposed to be applied in UAV localization. Considerable tests in the indoor practical scenarios have proven the effectiveness and versatility of our localization method. The code is available at https://github.com/vision4robotics/AutoTrack. Yiming Li 0003, Changhong Fu 0001, Fangqiang Ding, Ziyuan Huang 0003, Geng Lu |
CVPR | 2 |
| 2020 | Keyfilter-Aware Real-Time UAV Object TrackingabstractCorrelation filter-based tracking has been widely applied in unmanned aerial vehicle (UAV) with high efficiency. However, it has two imperfections, i.e., boundary effect and filter corruption. Several methods enlarging the search area can mitigate boundary effect, yet introducing undesired background distraction. Existing frame-by-frame context learning strategies for repressing background distraction nevertheless lower the tracking speed. Inspired by keyframe-based simultaneous localization and mapping, keyfilter is proposed in visual tracking for the first time, in order to handle the above issues efficiently and effectively. Keyfilters generated by periodically selected keyframes learn the context intermittently and are used to restrain the learning of filters, so that 1) context awareness can be transmitted to all the filters via keyfilter restriction, and 2) filter corruption can be repressed. Compared to the state-of-the-art results, our tracker performs better on two challenging benchmarks, with enough speed for UAV real-time applications. Yiming Li 0003, Changhong Fu 0001, Ziyuan Huang 0003, Yinqiang Zhang, Jia Pan 0001 |
ICRA | 2 |
| 2020 | Training-Set Distillation for Real-Time UAV Object TrackingabstractCorrelation filter (CF) has recently exhibited promising performance in visual object tracking for unmanned aerial vehicle (UAV). Such online learning method heavily depends on the quality of the training-set, yet complicated aerial scenarios like occlusion or out of view can reduce its reliability. In this work, a novel time slot-based distillation approach is proposed to efficiently and effectively optimize the training-set's quality on the fly. A cooperative energy minimization function is established to score the historical samples adaptively. To accelerate the scoring process, frames with high confident tracking results are employed as the keyframes to divide the tracking process into multiple time slots. After the establishment of a new slot, the weighted fusion of the previous samples generates one key-sample, in order to reduce the number of samples to be scored. Besides, when the current time slot exceeds the maximum frame number, which can be scored, the sample with the lowest score will be discarded. Consequently, the training-set can be efficiently and reliably distilled. Comprehensive tests on two well-known UAV benchmarks prove the effectiveness of our method with real-time speed on single CPU. Changhong Fu 0001, Fuling Lin, Yiming Li 0003, Peng Lu 0003 |
ICRA | 2 |
| 2020 | BiCF: Learning Bidirectional Incongruity-Aware Correlation Filter for Efficient UAV Object TrackingabstractCorrelation filters (CFs) have shown excellent performance in unmanned aerial vehicle (UAV) tracking scenarios due to their high computational efficiency. During the UAV tracking process, viewpoint variations are usually accompanied by changes in the object and background appearance, which poses a unique challenge to CF-based trackers. Since the appearance is gradually changing over time, an ideal tracker can not only forward predict the object position but also backtrack to locate its position in the previous frame. There exist response-based errors in the reversibility of the tracking process containing the information on the changes in appearance. However, some existing methods do not consider the forward and backward errors based on while using only the current training sample to learn the filter. For other ones, the applicants of considerable historical training samples impose a computational burden on the UAV. In this work, a novel bidirectional incongruity-aware correlation filter (BiCF) is proposed. By integrating the response-based bidirectional incongruity error into the CF, BiCF can Efficiently learn the changes in appearance and suppress the inconsistent error. Extensive experiments on 243 challenging sequences from three UAV datasets (UAV123, UAVDT, and DTB70) are conducted to demonstrate that BiCF favorably outperforms other 25 state-of-the-art trackers and achieves a real-time speed of 45.4 FPS on a single CPU, which can be applied in UAV Efficiently. Fuling Lin, Changhong Fu 0001, Yujie He 0002, Fuyu Guo |
ICRA | 2 |
| 2020 | DR2Track: Towards Real-Time Visual Tracking for UAV via Distractor Repressed Dynamic RegressionabstractVisual tracking has yielded promising applications with unmanned aerial vehicle (UAV). In literature, the advanced discriminative correlation filter (DCF) type trackers generally distinguish the foreground from the background with a learned regressor which regresses the implicit circulated samples into a fixed target label. However, the predefined and unchanged regression target results in low robustness and adaptivity to uncertain aerial tracking scenarios. In this work, we exploit the local maximum points of the response map generated in the detection phase to automatically locate current distractors1. By repressing the response of distractors in the regressor learning, we can dynamically and adaptively alter our regression target to leverage the tracking robustness as well as adaptivity. Substantial experiments conducted on three challenging UAV benchmarks demonstrate both excellent performance and extraordinary speed (~50fps on a cheap CPU) of our tracker. Changhong Fu 0001, Fangqiang Ding, Yiming Li 0003, Chen Feng 0002 |
IROS | 1 |
| 2020 | Learning Consistency Pursued Correlation Filters for Real-Time UAV TrackingabstractCorrelation filter (CF)-based methods have demonstrated exceptional performance in visual object tracking for unmanned aerial vehicle (UAV) applications, but suffer from the undesirable boundary effect. To solve this issue, spatially regularized correlation filters (SRDCF) proposes the spatial regularization to penalize filter coefficients, thereby significantly improving the tracking performance. However, the temporal information hidden in the response maps is not considered in SRDCF, which limits the discriminative power and the robustness for accurate tracking. This work proposes a novel approach with dynamic consistency pursued correlation filters, i.e., the CPCF tracker. Specifically, through a correlation operation between adjacent response maps, a practical consistency map is generated to represent the consistency level across frames. By minimizing the difference between the practical and the scheduled ideal consistency map, the consistency level is constrained to maintain temporal smoothness, and rich temporal information contained in response maps is introduced. Besides, a dynamic constraint strategy is proposed to further improve the adaptability of the proposed tracker in complex situations. Comprehensive experiments are conducted on three challenging UAV benchmarks, i.e., UAV123@10FPS, UAVDT, and DTB70. Based on the experimental results, the proposed tracker favorably surpasses the other 25 state-of-the-art trackers with real-time running speed (~43FPS) on a single CPU. Changhong Fu 0001, Xiaoxiao Yang, Juntao Xu, Changjing Liu, Peng Lu 0003 |
IROS | 1 |
| 2020 | Augmented Memory for Correlation Filters in Real-Time UAV TrackingabstractThe outstanding computational efficiency of discriminative correlation filter (DCF) fades away with various complicated improvements. Previous appearances are also gradually forgotten due to the exponential decay of historical views in traditional appearance updating scheme of DCF framework, reducing the model's robustness. In this work, a novel tracker based on DCF framework is proposed to augment memory of previously appeared views while running at real-time speed. Several historical views and the current view are simultaneously introduced in training to allow the tracker to adapt to new appearances as well as memorize previous ones. A novel rapid compressed context learning is proposed to increase the discriminative ability of the filter efficiently. Substantial experiments on UAVDT and UAV123 datasets have validated that the proposed tracker performs competitively against other 26 top DCF and deep-based trackers with over 40 FPS on CPU. Yiming Li 0003, Changhong Fu 0001, Fangqiang Ding, Ziyuan Huang 0003, Jia Pan 0001 |
IROS | 2 |
| 2020 | Automatic Failure Recovery and Re-Initialization for Online UAV Tracking with Joint Scale and Aspect Ratio OptimizationabstractCurrent unmanned aerial vehicle (UAV) visual tracking algorithms are primarily limited with respect to: (i) the kind of size variation they can deal with, (ii) the implementation speed which hardly meets the real-time requirement. In this work, a real-time UAV tracking algorithm with powerful size estimation ability is proposed. Specifically, the overall tracking task is allocated to two 2D filters: (i) translation filter for location prediction in the space domain, (ii) size filter for scale and aspect ratio optimization in the size domain. Besides, an efficient two-stage re-detection strategy is introduced for long-term UAV tracking tasks. Large-scale experiments on four UAV benchmarks demonstrate the superiority of the presented method which has computation feasibility on a low-cost CPU. Fangqiang Ding, Changhong Fu 0001, Yiming Li 0003, Chen Feng 0002 |
IROS | 2 |
| 2020 | Towards Robust Visual Tracking for Unmanned Aerial Vehicle with Tri-Attentional Correlation FiltersabstractObject tracking has been broadly applied in unmanned aerial vehicle (UAV) tasks in recent years. However, existing algorithms still face difficulties such as partial occlusion, clutter background, and other challenging visual factors. Inspired by the cutting-edge attention mechanisms, a novel object tracking framework is proposed to leverage multi-level visual attention. Three primary attention, i.e., contextual attention, dimensional attention, and spatiotemporal attention, are integrated into the training and detection stages of correlation filter-based tracking pipeline. Therefore, the proposed tracker is equipped with robust discriminative power against challenging factors while maintaining high operational efficiency in UAV scenarios. Quantitative and qualitative experiments on two well-known benchmarks with 173 challenging UAV video sequences demonstrate the effectiveness of the proposed framework. The proposed tracking algorithm favorably outperforms 12 state-of-the-art methods, yielding 4.8% relative gain in UAVDT and 8.2% relative gain in UAV123@10fps against the baseline tracker while operating at the speed of ~28 frames per second. Yujie He 0002, Changhong Fu 0001, Fuling Lin, Yiming Li 0003, Peng Lu 0003 |
IROS | 2 |
| 2020 | Robust multi-kernelized correlators for UAV tracking with adaptive context analysis and dynamic weighted filters
Changhong Fu 0001, Yujie He 0002, Fuling Lin, Weijiang Xiong |
Neural Comput. Appl. | 1 |
| 2020 | Surrounding-aware correlation filter for UAV tracking with selective spatial regularization
Changhong Fu 0001, Weijiang Xiong, Fuling Lin, Yufeng Yue |
Signal Process. | 1 |
| 2020 | Object Saliency-Aware Dual Regularized Correlation Filter for Real-Time Aerial TrackingabstractSpatial regularization has been proved as an effective method for alleviating the boundary effect and boosting the performance of a discriminative correlation filter (DCF) in aerial visual object tracking. However, existing spatial regularization methods usually treat the regularizer as a supplementary term apart from the main regression and neglect to regularize the filter involved in the correlation operation. To address the aforementioned issue, this article introduces a novel object saliency-aware dual regularized correlation filter, i.e., DRCF. Specifically, the proposed DRCF tracker suggests a dual regularization strategy to directly regularize the filter involved with the correlation operation inside the core of the filter generating ridge regression. This allows the DRCF tracker to suppress the boundary effect and consequently enhance the performance of the tracker. Furthermore, an efficient method based on a saliency detection algorithm is employed to generate the dual regularizers dynamically and provide the regularizers with online adjusting ability. This enables the generated dynamic regularizers to automatically discern the object from the background and actively regularize the filter to accentuate the object during its unpredictable appearance changes. By the merits of the dual regularization strategy and the saliency-aware dynamical regularizers, the proposed DRCF tracker performs favorably in terms of suppressing the boundary effect, penalizing the irrelevant background noise coefficients and boosting the overall performance of the tracker. Exhaustive evaluations on 193 challenging video sequences from multiple well-known challenging aerial object tracking benchmarks validate the accuracy and robustness of the proposed DRCF tracker against 27 other state-of-the-art methods. Meanwhile, the proposed tracker can perform real-time aerial tracking applications on a single CPU with sufficient speed of 38.4 frames/s. Changhong Fu 0001, Juntao Xu, Fuling Lin, Fuyu Guo, Tingcong Liu, Zhijun Zhang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Learning Aberrance Repressed Correlation Filters for Real-Time UAV TrackingabstractTraditional framework of discriminative correlation filters (DCF) is often subject to undesired boundary effects. Several approaches to enlarge search regions have been already proposed in the past years to make up for this shortcoming. However, with excessive background information, more background noises are also introduced and the discriminative filter is prone to learn from the ambiance rather than the object. This situation, along with appearance changes of objects caused by full/partial occlusion, illumination variation, and other reasons has made it more likely to have aberrances in the detection process, which could substantially degrade the credibility of its result. Therefore, in this work, a novel approach to repress the aberrances happening during the detection process is proposed, i.e., aberrance repressed correlation filter (ARCF). By enforcing restriction to the rate of alteration in response maps generated in the detection phase, the ARCF tracker can evidently suppress aberrances and is thus more robust and accurate to track objects. Considerable experiments are conducted on different UAV datasets to perform object tracking from an aerial view, i.e., UAV123, UAVDT, and DTB70, with 243 challenging image sequences containing over 90K frames to verify the performance of the ARCF tracker and it has proven itself to have outperformed other 20 state-of-the-art trackers based on DCF and deep-based frameworks with sufficient speed for real-time applications. Ziyuan Huang 0003, Changhong Fu 0001, Yiming Li 0003, Fuling Lin, Peng Lu 0003 |
ICCV | 2 |
| 2019 | Boundary Effect-Aware Visual Tracking for UAV with Online Enhanced Background Learning and Multi-Frame Consensus VerificationabstractDue to implicitly introduced periodic shifting of limited searching area, visual object tracking using correlation filters often has to confront undesired boundary effect. As boundary effect severely degrade the quality of object model, it has made it a challenging task for unmanned aerial vehicles (UAV) to perform robust and accurate object following. Traditional hand-crafted features are also not precise and robust enough to describe the object in the viewing point of UAV. In this work, a novel tracker with online enhanced background learning is specifically proposed to tackle boundary effects. Real background samples are densely extracted to learn as well as update correlation filters. Spatial penalization is introduced to offset the noise introduced by exceedingly more background information so that a more accurate appearance model can be established. Meanwhile, convolutional features are extracted to provide a more comprehensive representation of the object. In order to mitigate changes of objects' appearances, multi-frame technique is applied to learn an ideal response map and verify the generated one in each frame. Exhaustive experiments were conducted on 100 challenging UAV image sequences and the proposed tracker has achieved state-of-the-art performance. Changhong Fu 0001, Ziyuan Huang 0003, Yiming Li 0003, Ran Duan 0002, Peng Lu 0003 |
IROS | 1 |
| 2019 | Visual tracking with online structural similarity-based weighted multiple instance learning
Changhong Fu 0001, Ran Duan 0002, Erdal Kayacan |
Inf. Sci. | 1 |
| 2017 | Similarity-based non-singleton fuzzy logic control for improved performance in UAVsabstractAs non-singleton fuzzy logic controllers (NSFLCs) are capable of capturing input uncertainties, they have been effectively used to control and navigate unmanned aerial vehicles (UAVs) recently. To further enhance the capability to handle the input uncertainty for the UAV applications, a novel NSFLC with the recently introduced similarity-based inference engine, i.e., Sim-NSFLC, is developed. In this paper, a comparative study in a 3D trajectory tracking application has been carried out using the aforementioned Sim-NSFLC and the NSFLCs with the standard as well as centroid composition-based inference engines, i.e., Sta-NSFLC and Cen-NSFLC. All the NSFLCs are developed within the robot operating system (ROS) using the C++ programming language. Extensive ROS Gazebo simulation-based experiments show that the Sim-NSFLCs can achieve better control performance for the UAVs in comparison with the Sta-NSFLCs and Cen-NSFLCs under different input noise levels. Changhong Fu 0001, Andriy Sarabakha, Erdal Kayacan, Christian Wagner 0002, Robert Ivor John, Jonathan M. Garibaldi |
FUZZ-IEEE | 1 |
| 2017 | Double-input interval type-2 fuzzy logic controllers: Analysis and designabstractA significant number of investigations of type-1 and type-2 fuzzy logic controllers have revealed their exceptional ability to capture uncertainties in complex and nonlinear systems, particularly in real-time control applications. However, regardless of being type-1 or type-2, fuzzy logic controller design is still a complicated task due to the lack of a closed form solution of the output and an interpretable relationship between the control output and fuzzy logic controller design parameters, such as center or width of the membership functions. To simplify the design procedure further, we think every attempt to obtain such interpretable relationships is worthwhile. Accordingly, this paper aims to design a double-input interval type-2 fuzzy PID controller and obtain interpretable relationships between the input and the output of the controller. Thereafter, we deploy the novel design for the control of a Y6 coaxial tricopter unmanned aerial vehicle. Simulation results, which are realised in robot operating system (ROS) using C++ and Gazebo environment, are found to tally with the theoretical analysis and claims in the paper. Andriy Sarabakha, Changhong Fu 0001, Erdal Kayacan |
FUZZ-IEEE | 2 |
| 2017 | Tracking-recommendation-detection: A novel online target modeling for visual tracking
Ran Duan 0002, Changhong Fu 0001, Erdal Kayacan |
Eng. Appl. Artif. Intell. | 2 |
| 2016 | A comparative study on the control of quadcopter UAVs by using singleton and non-singleton fuzzy logic controllersabstractFuzzy logic controllers (FLCs) have extensively been used for the autonomous control and guidance of unmanned aerial vehicles (UAVs) due to their capability of handling uncertainties and delivering adequate control without the need for a precise, mathematical system model which is often either unavailable or highly costly to develop. Despite the fact that non-singleton FLCs (NSFLCs) have shown more promising performance in several applications when compared to their singleton counterparts (SFLCs), most of UAV applications are still realized by using SFLCs. In this paper, we explore the potential of both standard and the recently introduced centroid based NSFLCs, i.e., Sta-NSFLC and Cen-NSFLC, for the control of a quadcopter UAV under various input noise conditions using different levels of fuzzifier, and a comparative study has been conducted using the three aforementioned FLCs. We present a series of simulation-based experiments, the simulation results show that the control performances of NSFLCs are better than those of SFLC, and the Cen-NSFLC outperforms the Sta-NSFLC especially under highly noisy conditions. Changhong Fu 0001, Andriy Sarabakha, Erdal Kayacan, Christian Wagner 0002, Robert Ivor John, Jonathan M. Garibaldi |
FUZZ-IEEE | 1 |
| 2016 | A performance evaluation of detectors and descriptors for UAV visual trackingabstractThis paper is made up of a series of performance evaluations of computer vision algorithms, namely detectors and descriptors. The OpenCV 3.1 implementations of these algorithms were used for these evaluations. The main purpose behind these evaluations was to determine the best algorithms to use for a UAV guidance system. Bruce Cowan, Nursultan Imanberdiyev, Changhong Fu 0001, Yiqun Dong, Erdal Kayacan |
ICARCV | 3 |
| 2016 | RRT-based 3D path planning for formation landing of quadrotor UAVsabstractThis paper discusses the formation landing problem of quadrotor UAVs, which is considered as a UAV leader-follower problem, avoiding static obstacles. Rapidly-exploring random tree algorithm is used to generate the path for the leader UAV firstly. In particular, specifics of tree-grow including nodes selection, parent node connection, feasible and optimal path generation are explained. Given the leader UAV position, path finding for the follower UAV is conducted to avoid both static obstacles and the leader quadrotor. Based on the intensive simulations, which are conducted in ROS-Gazebo environment, the proposed framework is considered to be applicable in real-time formation landing of quadrotor UAVs. Yiqun Dong, Changhong Fu 0001, Erdal Kayacan |
ICARCV | 2 |
| 2016 | Autonomous navigation of UAV by using real-time model-based reinforcement learningabstractAutonomous navigation in an unknown or uncertain environment is one of the challenging tasks for unmanned aerial vehicles (UAVs). In order to address this challenge, it is necessary to have sophisticated high level control methods that can learn and adapt themselves to changing conditions. One of the most promising frameworks for such a purpose is reinforcement learning. In this paper, a novel model-based reinforcement learning algorithm, TEXPLORE, is developed as a high level control method for autonomous navigation of UAVs. The developed approach has been extensively tested with a quadcopter UAV in ROS-Gazebo environment. The experimental results show that our method is able to learn an efficient trajectory in a few iterations and perform actions in real-time. Moreover, we show that our approach significantly outperforms Q-learning based method. To the best of our knowledge, this is the first time that TEXPLORE has been developed to achieve autonomous navigation of UAVs. Nursultan Imanberdiyev, Changhong Fu 0001, Erdal Kayacan, I-Ming Chen 0001 |
ICARCV | 2 |
| 2016 | Recommended keypoint-aware tracker: Adaptive real-time visual tracking using consensus feature prior rankingabstractThis paper deals with the problem of historical feature selection for appearance model update in feature-based tracking. In particular, we convert the feature selection procedure into a ranking process where the top-N keypoint features are ranked based on the tracking histories. To the best of our knowledge, for the first time in this paper, a consensus feature prior (CFP) recommendation system is proposed that allows us to learn and update the appearance model online within a limited model size. Furthermore, the ranking scores obtained from the proposed recommendation system also provide a conviction of recovering the tracking after its failure. Extensive experiments (more than 600,000 frames) have been done by strictly following the Visual Tracking Benchmark v1.0 protocol. The results demonstrate that our method outperforms most of the state-of-art trackers both in terms of speed and accuracy. Ran Duan 0002, Changhong Fu 0001, Erdal Kayacan, Danda Pani Paudel |
ICIP | 2 |
| 2016 | Recoverable recommended keypoint-aware visual tracking using coupled-layer appearance modellingabstractObject tracking over image sequences plays an remarkably crucial role in several computer vision applications, interalia, automated video surveillance, unmanned aerial vehicles and 3D reconstruction. In this paper, a novel, accurate, robust and recoverable real-time feature-based tracking framework is presented. The appearance modelling consists of a local and global layer. We propose a recommended keypoint-aware (RKA) tracker, which is fast and accurate, for the former, while the latter employs support vector machine (SVM) to determine the object and background, so that the RKA tracker can be recovered under possible target losing circumstances. Furthermore, the RKA tracker converts the tracking problem into the ranking of samples which provides a score of tracking confidence. Therefore, the priority switching between the local layer and global layer dependent upon the score becomes valid. Extensive experiments have been done by strictly following the visual tracking benchmark v1.0 protocol. The results demonstrate that the proposed novel method outperforms the state-of-the-art trackers in terms of robustness, speed and accuracy. Ran Duan 0002, Changhong Fu 0001, Erdal Kayacan |
IROS | 2 |
| 2014 | Robust real-time vision-based aircraft tracking from Unmanned Aerial VehiclesabstractAircraft tracking plays a key and important role in the Sense-and-Avoid system of Unmanned Aerial Vehicles (UAVs). This paper presents a novel robust visual tracking algorithm for UAVs in the midair to track an arbitrary aircraft at real-time frame rates, together with a unique evaluation system. This visual algorithm mainly consists of adaptive discriminative visual tracking method, Multiple-Instance (MI) learning approach, Multiple-Classifier (MC) voting mechanism and Multiple-Resolution (MR) representation strategy, that is called Adaptive M3tracker, i.e. AM3. In this tracker, the importance of test sample has been integrated to improve the tracking stability, accuracy and real-time performances. The experimental results show that this algorithm is more robust, efficient and accurate against the existing state-of-art trackers, overcoming the problems generated by the challenging situations such as obvious appearance change, variant surrounding illumination, partial aircraft occlusion, blur motion, rapid pose variation and onboard mechanical vibration, low computation capacity and delayed information communication between UAVs and Ground Station (GS). To our best knowledge, this is the first work to present this tracker for solving online learning and tracking freewill aircraft/intruder in the UAVs. Changhong Fu 0001, Adrian Carrio, Miguel A. Olivares-Méndez, Ramón A. Suárez Fernández, Pascual Campoy Cervera |
ICRA | 1 |