EDBT 2026 Demo / reviewers in the wild / expert
Yuanjie Dang
dblp:212/1578
· DBLP profile ↗
23ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0002-8302-1338ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Scene-Aware Meta-learning Framework for Robust Photovoltaic Power Forecasting
Yihan Yu, Yuanjie Dang, Peng Chen 0008, Yilong Zhang 0001, Ronghua Liang |
ICIC (16) | 2 |
| 2026 | Learning Decoupled Features With Perceptual Distillation for Blind Image Quality AssessmentabstractExisting Blind Image Quality Assessment (BIQA) approaches typically employ subjective scores as optimization targets to train the model, aiming for results consistent with human judgments. Such judgments are derived from a comprehensive analysis of complex distortions and diverse semantics from images, whereas subjective scores represent the overall quality. This poses a significant challenge for a single model to learn diverse perceptual cues under weak supervision. To address this, we propose a Decoupled Feature Learning (DFL) framework that learns compact global content-aware and local distortion-aware features in a disentangled modeling for BIQA. Our key insight is to leverage global-local input pairs to decompose content-aware and distortion-aware cues entangled in distorted images, and aggregate decoupled perceptual features into a single network. We design a perceptual knowledge distillation strategy that progressively guides the student from fragmented representations to build local-to-global correspondences by distilling self-supervised semantic knowledge, while incorporating the Just-Noticeable-Difference (JND) model to highlight the transfer of perceptually sensitive content features. Finally, we introduce a local distortion-guided attention module to model synergistic effects of different perceptual features from the student for quality evaluation. Extensive experiments on eight benchmark datasets demonstrate the superior performance of the proposed model over the state-of-the-arts. In addition, the DFL framework is flexibly used to improve the perception ability of other Transformer variants. The code is released at https://github.com/JianjunXiang/DFT. Jianjun Xiang, Yuanjie Dang, Peng Chen 0008, Ronghua Liang, Weisi Lin |
IEEE Trans. Image Process. | 2 |
| 2025 | Spatial Continuity-Aware OCT Fingerprint Reconstruction Using Iterative Feature EnhancementabstractOptical coherence tomography (OCT) is a non-invasive imaging technique capable of acquiring depth information up to 1-3mm beneath the skin surface, including the stratum corneum and viable epidermis regions. This technique allows for the reconstruction of internal and external fingerprint images from grayscale data. However, existing fingerprint extraction methods heavily rely on contour features and current 2D approaches overlook the spatial continuity of biometric features in OCT images. Therefore, this paper proposes a novel iterative algorithm for internal and external fingerprint extraction from OCT images. This algorithm incorporates the spatial continuity of OCT slice images and an iterative feature enhancement module during the prediction phase to improve segmentation continuity. Additionally, a soft label technique is employed to reduce contour dependence and mitigate interference from noise and anomaly interference. Qualitative and quantitative experiments demonstrate significant improvements in segmentation accuracy with higher fingerprint quality, validating the effectiveness of the proposed approach. Yilong Zhang 0001, Xuanbing Chen, Shengming Zhu, Haohao Sun, Haixia Wang 0002, Jian Liu 0053, Yuanjie Dang, Ronghua Liang, Peng Chen 0008 |
IJCB | 7 |
| 2025 | FDDet: Frequency-Decoupling for Boundary Refinement in Temporal Action Detection
Xinnan Zhu, Yicheng Zhu, Tixin Chen, Yuanjie Dang |
ICIC (6) | 5 |
| 2025 | SonarPoint: Weak-Heterogeneity Awareness Object Detection Network for 3D Sonar Point CloudabstractUnderwater target detection is primarily achieved through two methods: optical imaging and underwater sonar. 3D sonar, as the most advanced underwater detection technology, is characterized by strong penetration and long scanning distance, making it more suitable for tasks such as deep-sea exploration, murky water detection, and long-distance target identification. However, acquiring underwater sonar images is challenging, and there is no open-source 3D sonar dataset. Traditional three-dimensional target detection methods typically require highquality data and face significant challenges when dealing with weak heterogeneous sonar point clouds caused by high noise, low resolution, and occlusions. To address the aforementioned issues, we first propose a novel fuzzy decoupling module that differs from traditional foreground-background segmentation. This module simultaneously extracts valuable information about the target and its surrounding environment, mitigating the reduction in heterogeneity caused by noise and sonar side lobes. To achieve efficient fusion and capture global information after fuzzy decoupling, a multi-hop Mamba seamless adaptive decoupling point is introduced. It effectively enhances the connection between the two decoupled parts. To address missing and occlusion problems, a second-stage refinement based on Markov prediction is proposed. This low-cost approach, in contrast to using the original point cloud for contour and detail completion, enriches target boundary information. To validate our method, we have designed a practical 3D sonar imaging system and tested it through lake-based experiments. We have collected extensive raw data from Qiandao Lake and conducted annotation work. Through qualitative and quantitative experiments, our method outperforms the most advanced methods by 11.4%. Tiancheng Cai, Peng Chen 0008, Weibo Mao, Yingtian Hu, Yilong Zhang 0001, Yuanjie Dang, Ronghua Liang, Xiang Tian 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | A Spatial-Aware Temporal Modeling Network for Imitation Learning-Based Drone NavigationabstractImitation learning-based drone autonomous navigation has attracted significant attention due to the ability of leveraging deep neural networks to learn the control policy from human pilot demonstrations. However, most current studies generate the control command using only a single image, overlooking the semantic information embedded in the sequential input images. While some reinforcement learning-based methods have explored the temporal modeling of sequential input images, they often overlook the spatial relations between frames and vectorize 2D information of each image into a 1D feature. In this paper, we propose a novel imitation learning-based method, termed the spatial-aware temporal modeling network (SATMN), for autonomous drone navigation using sequential images as input. Specifically, we introduce a spatial-temporal-separated modeling mechanism to extract low-resolution spatial features from original images and then perceive spatial-temporal relations among these 2D features. SATMN preserves the spatial information of each 2D image feature during temporal modeling and enables real-time onboard computing on a drone. To validate the effectiveness of the proposed method, we design a compact quadrotor platform capable of autonomous navigation using SATMN, entirely powered by onboard computing devices. Comprehensive and reproducible experiments on public datasets demonstrate the superior performance of our method compared to existing approaches. Tianwei Yu, Yuanjie Dang, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Generate anomalies from normal: a partial pseudo-anomaly augmented approach for video anomaly detection
Yuanjie Dang, Jiangyun Chen, Peng Chen 0008, Nan Gao 0001, Ruohong Huan |
Vis. Comput. | 1 |
| 2024 | Task-Agnostic Self-Distillation for Few-Shot Action Recognition
Yuanjie Dang, Peng Chen 0008, Ronghua Liang, Nan Gao 0001, Ruohong Huan, Xiaofei He 0001 |
IJCAI | 2 |
| 2024 | Semantic-Aware and Quality-Aware Interaction Network for Blind Video Quality AssessmentabstractCurrent state-of-the-art video quality assessment (VQA) models typically integrate various perceptual features to comprehensively represent video quality degradation. These models either directly concatenate features or fuse different perceptual scores while ignoring the domain gaps between cross-aware features, thus failing to adequately learn the correlations and interactions between different perceptual features. To this end, we analyze the independent effects and information gaps of quality-and semantic-aware features on video quality. Based on an analysis of the spatial and temporal differences between two aware features, we propose a semantic-Aware and quality-Aware Interaction Network (A2INet) for blind VQA. For spatial gaps, we introduce a cross-aware guided interaction module to enhance the interaction between semantic-and quality-aware features in a local-to-global manner. Considering temporal discrepancies, we design a cross-aware temporal modeling module to further perceive temporal content variation and quality saliency information, and perceptual features are regressed into quality score by a temporal network and a temporal pooling. Extensive experiments on six benchmark VQA datasets show that our model achieves state-of-the-art performance, and ablation studies further validate the effectiveness of each module. We also present a simple video sampling strategy to balance the effectiveness and efficiency of the model. The code for the proposed method will be released at https://github.com/JianjunXiang/A2INet. Jianjun Xiang, Yuanjie Dang, Peng Chen 0008, Ronghua Liang, Ruohong Huan, Nan Gao 0001 |
ACM Multimedia | 2 |
| 2024 | Saliency-Guided Fine-Grained Temporal Mask Learning for Few-Shot Action RecognitionabstractTemporal relation modeling is one of the core aspects of few-shot action recognition. Most previous works mainly focus on temporal relation modeling based on coarse-level actions, without considering the atomic action details and fine-grained temporal information. This oversight represents a significant limitation in this task. Specifically, coarse-level temporal relation modeling can make the few-shot models overfit in high-discrepancy temporal context, and ignore the low-discrepancy but high-semantic relevance action details in the video. To address these issues, we propose a saliency-guided fine-grained temporal mask learning method that models the temporal atomic action relation for few-shot action recognition in a finer manner. First, to model the comprehensive temporal relations of video instances, we design a temporal mask learning architecture to automatically search for the best matching of each atomic action snippet. Next, to exploit the low-discrepancy atomic action features, we introduce a saliency-guided temporal mask module to adaptively locate and excavate the atomic action information. After that, the few-shot predictions can be obtained by feeding the embedded rich temporal-relation features to a common feature matcher. Extensive experimental results on standard datasets demonstrate our method's superior performance compared to existing state-of-the-art methods. Yuanjie Dang, Peng Chen 0008, Ruohong Huan, Ronghua Liang |
ACM Multimedia | 2 |
| 2024 | Focus on Subtle Actions: Semantic and Saliency Knowledge Co-Propagation Method for Weakly-Supervised Temporal Action Localization
Yuanjie Dang, Haoyu Shou, Peng Chen 0008, Nan Gao 0001, Ruohong Huan, Yilong Zhang 0001 |
PRCV (10) | 1 |
| 2024 | Learning Reliable Dense Pseudo-Labels for Point-Level Weakly-Supervised Action LocalizationabstractAbstract Point-level weakly-supervised temporal action localization aims to accurately recognize and localize action segments in untrimmed videos, using only point-level annotations during training. Current methods primarily focus on mining sparse pseudo-labels and generating dense pseudo-labels. However, due to the sparsity of point-level labels and the impact of scene information on action representations, the reliability of dense pseudo-label methods still remains an issue. In this paper, we propose a point-level weakly-supervised temporal action localization method based on local representation enhancement and global temporal optimization. This method comprises two modules that enhance the representation capacity of action features and improve the reliability of class activation sequence classification, thereby enhancing the reliability of dense pseudo-labels and strengthening the model’s capability for completeness learning. Specifically, we first generate representative features of actions using pseudo-label feature and calculate weights based on the feature similarity between representative features of actions and segments features to adjust class activation sequence. Additionally, we maintain the fixed-length queues for annotated segments and design a action contrastive learning framework between videos. The experimental results demonstrate that our modules indeed enhance the model’s capability for comprehensive learning, particularly achieving state-of-the-art results at high IoU thresholds. Yuanjie Dang, Guozhu Zheng, Peng Chen 0008, Nan Gao 0001, Ruohong Huan, Ronghua Liang |
Neural Process. Lett. | 1 |
| 2024 | A Distributed and Parallel Accelerator Design for 3-D Acoustic Imaging on FPGA-Based Systemsabstract3-D imaging sonar is crucial in the exploration of marine resources, and the development of portable device with high imaging quality and high real-time performance is the general trend. However, traditional framework methods are limited by the huge amount of computation brought by high-quality imaging, making it difficult to implement in engineering. To address this issue, we develop 3-D real-time sonar system in an algorithm-hardware co-designed way. An ultrawideband distributed and parallel subarray beamforming algorithm (UWBDPS) is proposed for 3-D acoustic imaging. This is a multi-stage array time-frequency beamforming method under a distributed parallel computing architecture. Based on this, we propose field-programmable gate array (FPGA)-based accelerator. It divides a large sonar receiving planar array into multiple parallel subarrays, and complete the beamforming in two stages, which can reduce the calculation load and speeds up 3-D imaging. For engineering implementation, we optimized the sparseness of the planar transducer array, with a sparse rate as high as 97.7%. The experimental results show that the calculation amount of the proposed UWB-DPS algorithm is reduced to 1/5.7 of the traditional framework algorithm, the imaging performance is effectively improved, and the FPGA-based accelerator outperforms the CPU software implementation by 935×. Weibo Mao, Peng Chen 0008, Yingtian Hu, Haoran Liang 0001, Yuanjie Dang, Ronghua Liang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | Multi-Level Objective Alignment Transformer for Fine-Grained Oral Panoramic X-Ray Report GenerationabstractAutomatically generated oral panoramic X-ray report is highly beneficial for improving the efficiency of dental diagnosis. However, recent solutions adopt holistic methods, resulting in a cursory description of the oral condition. This may lead to reports lacking details, such as specific sites or lesion contours. Therefore, we propose a Multi-Level objective Alignment Transformer(MLAT) network, which integrates all tooth and disease objects into a positional alignment graph to extract fine-grained object-level features. Specifically, we introduce a novel Object-Level Collaborative Encoder (OLCE) module, which uses a positional alignment graph to construct object relationships. OLCE enhances object-level feature extraction by eliminating interference information between pathologically unrelated objects. In addition, we build a high-quality panoramic X-ray image-report dataset consisting of 562 sets of images and reports labeled by 13 experienced dental specialists. Experiments on the collected dataset show that the proposed MLAT significantly outperforms the state-of-the-art baselines by more than 5% in 4 different metrics, including BLEUs, Meteor, Rouge, and BERTScore. Nan Gao 0001, Renyuan Yao, Ronghua Liang, Peng Chen 0008, Tianshuang Liu, Yuanjie Dang |
IEEE Trans. Multim. | 6 |
| 2024 | Pseudo Light Field Image and 4D Wavelet-Transform-Based Reduced-Reference Light Field Image Quality AssessmentabstractReduced-reference light field image (LFI) quality assessment (RR LFIQA) automatically assesses image quality with only partial information about the reference LFI is available. Existing RR LFIQA has difficulty extracting effective RR information and perceptual features to represent the LFI quality. In this article, we propose an RR LFIQA model based on pseudo LFI (PLFI) and four-dimensional (4D) wavelet transform. To extract RR information related to LFI perceptual quality, a PLFI is created as the RR information of the LFI using a view synthesis algorithm. Considering that the high-dimensional characteristics of the PLFI, 4D wavelet transform is used to decompose the original and distorted PLFIs. The 4D wavelet transform essentially performs a continuous 1D wavelet transform for the 4D signal to enable the local 4D structure of the PLFIs to be characterized effectively in the 4D wavelet domain. A novel spatial-angular weighting strategy is proposed to describe the importance of each location for quality evaluation, to further improve the performance of the proposed method. Experimental results on four benchmark datasets show that the proposed model performs better than the representative 2DIQA and LFIQA models. Jianjun Xiang, Peng Chen 0008, Yuanjie Dang, Ronghua Liang, Gangyi Jiang |
IEEE Trans. Multim. | 3 |
| 2024 | Discriminative Action Snippet Propagation Network for Weakly Supervised Temporal Action LocalizationabstractWeakly supervised temporal action localization (WTAL) aims to classify and localize actions in untrimmed videos with only video-level labels. Recent studies have attempted to obtain more accurate temporal boundaries by exploiting latent action instances in ambiguous snippets or propagating representative action features. However, empirically handcrafted ambiguous snippet extraction and the imprecise alignment of representative snippet propagation lead to challenges in modeling the completeness of actions for these methods. In this article, we propose a Discriminative Action Snippet Propagation Network (DASP-Net) to accurately discover ambiguous snippets in videos and propagate discriminative instance-level features throughout the video for improving action completeness. Specifically, we introduce a novel discriminative feature propagation module for capturing the global contextual attention and propagating the action concept across the whole video by perceiving the discriminative action snippets with instance information from the same video. Simultaneously, we incorporate denoised pseudo-labels as supervision, where we correct the controversial prediction based on the feature space distribution during training, thereby alleviating false detection caused by noise background features. Furthermore, we design an ambiguous feature mining module, which maximizes the feature affinity information of action and background in ambiguous snippets to generate more accurate latent action and background snippets and learns more precise action instance boundaries through contrastive learning of action and background snippets. Extensive experiments show that DASP-Net achieves state-of-the-art results on THUMOS14 and ActivityNet1.2 datasets. Yuanjie Dang, Chunxia Huang, Peng Chen 0008, Nan Gao 0001, Ronghua Liang, Ruohong Huan |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | STAN: Spatio-Temporal Alignment Network for No-Reference Video Quality Assessment
Zhengyi Yang 0008, Yuanjie Dang, Jianjun Xiang, Peng Chen 0008 |
ICANN (3) | 2 |
| 2023 | Spatial-angular Quality-aware Representation Learning for Blind Light Field Image Quality AssessmentabstractBlind light field image quality assessment (BLFIQA) remains a challenging task in deep learning due to the unique spatial-angular structure of light field images (LFIs) and the lack of large-scale labeled data for training. In this work, we propose a novel BLFIQA method using spatial-angular quality-aware representation learning in a self-supervised learning manner. Visual content and distortion type are important factors affecting the perceived quality of LFIs. In our observation, the band-pass transform maps of LFIs with the same distortion type exhibit similar Gaussian distributions. Thus, we learn spatial-angular quality-aware representations by minimizing the distance in the embedding space between the luminance map and the band-pass transform map of the same LFI. To implement spatial-angular quality-aware representations of LFI, we also build a large-scale unlabeled dataset containing 40k distorted LFIs with different distortion types and visual content. Further, we propose a fusion-separation-fusion network (FSFNet) to extract features for representing the intrinsic spatial-angular structure of the LFI. After pre-training on the unlabeled dataset using the proposed self-supervised learning, the FSFNet is employed for downstream BLFIQA tasks and achieves good performance. Experimental results show that our proposed method outperforms seventeen state-of-the-art models on the Win5-LID, NBU-LF1.0 and LFDD datasets, and achieves 3.78%, 6.61% and 4.06% SRCC improvements, respectively. The code and dataset will be publicly available in https://github.com/JianjunXiang/SSL_and_FSFNet. Jianjun Xiang, Yuanjie Dang, Peng Chen 0008, Ronghua Liang, Ruohong Huan |
ACM Multimedia | 2 |
| 2023 | Multi-Speed Global Contextual Subspace Matching for Few-Shot Action RecognitionabstractFew-shot action recognition (FSAR) aims to classify unseen query actions into categories represented by a few labeled support videos. Most current FSAR methods adopt the frame-level matching mechanism that requires continuous actions to be represented by a fixed number of frame features. However, this could compromise the completeness of the contextual video information and make it difficult to handle video features of varying frame sampling speeds. In this paper, we propose a multi-speed global contextual subspace matching (MGCSM) method that generates global contextual action subspace representations from videos containing different numbers of frames to preserve contextual semantic information. Specifically, we propose to obtain the scale-agnostic information of embedding video features using a global contextual aggregation (GCA) module and then generate the discriminative action subspace representation with an action subspace generation (ASG) module. Furthermore, we introduce a multi-speed subspace matching (MSM) mechanism that generates a multi-speed classification score by integrating the similarities between query videos and support subspaces of varying sampling speeds. The proposed method is embedding-agnostic and can be combined with most mainstream embedding networks without model re-designs. Comprehensive and reproducible experiments on standard datasets demonstrate our method's superior performance compared to existing state-of-the-art methods. Tianwei Yu, Peng Chen 0008, Yuanjie Dang, Ruohong Huan, Ronghua Liang |
ACM Multimedia | 3 |
| 2023 | Path-Analysis-Based Reinforcement Learning Algorithm for Imitation FilmingabstractImitation filming has been applied to autonomous filming by mimicking human operators. To imitate the operation of cameramen when filming multiple human actions, existing methods plan the camera motion through time series prediction or train multiple models to handle a particular style in a specific situation. As a result, these methods require various settings to adapt to different scenarios. In this work, we overcome such limitations and propose an end-to-end imitation learning framework for drone cinematography systems. The framework consists of two main components: (1) an efficient motion feature extraction module for generating a compact motion feature space, (2) a path-analysis-based reinforcement learning (PABRL) algorithm for imitating multiple filming styles from demonstrations and incorporating aesthetical features for improved perspective shots. Our PABRL method is based on the actor–critic network, which regards multiple human motion variables, camera translations, and image composition as inputs and then outputs an aesthetical filming strategy related to the subject motion. In addition, we propose an attention mechanism and a long–short-term rewarding function to enhance the motion feature space and the integrity of the generated trajectory, respectively. Extensive experimental results in simulated and real outdoor environments demonstrate that compared with state-of-the-art methods, our method can achieve 69.8% higher performance in terms of trajectory planning accuracy while successfully incorporating aesthetical features into the captured videos. Yuanjie Dang, Chong Huang 0005, Peng Chen 0008, Ronghua Liang, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Multim. | 1 |
| 2022 | One-Shot Imitation Drone Filming of Human Motion VideosabstractImitation learning has recently been applied to mimic the operation of a cameraman in existing autonomous camera systems. To imitate a certain demonstration video, existing methods require users to collect a significant number of training videos with a similar filming style. Because the trained model is style-specific, it is challenging to generalize the model to imitate other videos with a different filming style. To address this problem, we propose a framework that we term "one-shot imitation filming", which can imitate a filming style by "seeing" only a single demonstration video of the target style without style-specific model training. This is achieved by two key enabling techniques: 1) filming style feature extraction, which encodes sequential cinematic characteristics of a variable-length video clip into a fixed-length feature vector; and 2) camera motion prediction, which dynamically plans the camera trajectory to reproduce the filming style of the demo video. We implemented the approach with a deep neural network and deployed it on a 6 degrees of freedom (DOF) drone system by first predicting the future camera motions, and then converting them into the drone's control commands via an odometer. Our experimental results on comprehensive datasets and showcases exhibit that the proposed approach achieves significant improvements over conventional baselines, and our approach can mimic the footage of an unseen style with high fidelity. Chong Huang 0005, Yuanjie Dang, Peng Chen 0008, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Fast Depth Prediction and Obstacle Avoidance on a Monocular Drone Using Probabilistic Convolutional Neural NetworkabstractRecent studies employ advanced deep convolutional neural networks (CNNs) for monocular depth perception, which can hardly run efficiently on small drones that rely on low/middle-grade GPU(e.g. TX2 and 1050Ti) for computation. In addition, the methods which can effectively and efficiently produce probabilistic depth prediction with a measure of model confidence have not been well studied. The lack of such a method could yield erroneous, sometimes fatal, decisions in drone applications (e.g. selecting a waypoint in a region with a large depth yet a low estimation confidence). This paper presents a real-time onboard approach for monocular depth prediction and obstacle avoidance with a lightweight probabilistic CNN (pCNN), which will be ideal for use in a lightweight energy-efficient drone. For each video frame, our pCNN can efficiently predict its depth map and the corresponding confidence. The accuracy of our lightweight pCNN is greatly boosted by integrating sparse depth estimation from a visual odometry into the network for guiding dense depth and confidence inference. The estimated depth map is transformed into Ego Dynamic Space (EDS) by embedding both dynamic motion constraints of a drone and the confidence values into the spatial depth map. Traversable waypoints are automatically computed in EDS based on which appropriate control inputs for the drone are produced. Extensive experimental results on public datasets demonstrate that our depth prediction method runs at 12Hz and 45Hz on TX2 and 1050Ti GPU respectively, which is 1.8X~5.6X faster than the state-of-the-art methods and achieves better depth estimation accuracy. We also conducted experiments of obstacle avoidance in both simulated and real environments to demonstrate the superiority of our method to the baseline methods. Xin Yang 0008, Yuanjie Dang, Hongcheng Luo, Yuesheng Tang, Chunyuan Liao, Peng Chen 0008, Kwang-Ting Cheng |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2018 | Real-Time Object Tracking on a Drone With Multi-Inertial Sensing DataabstractReal-time object tracking on a drone under a dynamic environment has been a challenging issue for many years, with existing approaches using off-line calculation or powerful computation units on board. This paper presents a new lightweight real-time onboard object tracking approach with multi-inertial sensing data, wherein a highly energy-efficient drone is built based on the Snapdragon flight board of Qualcomm. The flight board uses a digital signal processor core of the Snapdragon 801 processor to realize PX4 autopilot, an open-source autopilot system oriented toward inexpensive autonomous aircraft. It also uses an ARM core to realize Linux, robot operating systems, open-source computer vision library, and related algorithms. A lightweight moving object detection algorithm is proposed that extracts feature points in the video frame using the oriented FAST and rotated binary robust independent elementary features algorithm and adapts a local difference binary algorithm to construct the image binary descriptors. The K-nearest neighbor method is then used to match the image descriptors. Finally, an object tracking method is proposed that fuses inertial measurement unit data, global positioning system data, and the moving object detection results to calculate the relative position between coordinate systems of the object and the drone. All the algorithms are run on the Qualcomm platform in real time. Experimental results demonstrate the superior performance of our method over the state-of-the-art visual tracking method. Peng Chen 0008, Yuanjie Dang, Ronghua Liang, Wei Zhu 0006, Xiaofei He 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |