EDBT 2026 Demo / reviewers in the wild / expert
Yuqi Han
dblp:28/2360
· DBLP profile ↗
35ranked-venue papers
8as first author
30since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 9 since 2021Computer networks · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | One size doesn't fit all: Divide-and-conquer detector for UAV images
Yuqi Han, Xiaojing Yang, Xin Zhang 0093, Zengdi Bao |
Expert Syst. Appl. | 2 |
| 2026 | Indoor ranging localization algorithm using LOSAPs for smartphones
Lu Huang 0001, Hongji Cao, Yinsong Zhang, Jingxue Bi, Huidong Lei, Yequn Wei, Guofeng Jing, Yuqi Han, Kaikai Qiao |
Comput. Networks | 9 |
| 2026 | AR-HUD interaction interface Kansei design based on Pythagorean fuzzy rough CRITIC and TCN-Transformer-LSTM method
Tianxiong Wang, Yuqi Han |
Expert Syst. Appl. | 4 |
| 2026 | UTA-Sign: Unsupervised thermal video augmentation via event-assisted traffic signage sketching
Yuqi Han, Songqian Zhang, Weijian Su, Jin-Li Suo, Qiang Zhang 0008 |
Pattern Recognit. | 1 |
| 2026 | Seq-IF: Sequentially Consistent Infrared-Visible Video Fusion Under Time-Varying Illumination for Perception EnhancementabstractInfrared–visible image fusion leverages the complementary strengths of both modalities to enhance visual perception in challenging environments. While image-level fusion has achieved promising results, extending it to video remains challenging due to temporal illumination variations, brightness flickering, and visual inconsistency caused by motion under non-uniform illumination. To address these issues, we propose Seq-IF, a sequential fusion framework that ensures visual consistency and structural fidelity when processing video sequences. The framework comprises a static–dynamic decoupling module for robust foreground–background separation. For the background, fusion is performed by selecting the frame with the highest contrast to ensure clarity and stability. For the dynamic objects, the pixel intensity is adaptively adjusted across consecutive frames by a lightweight MLP-based illumination-consistency fine-tuning module that performs online adaptation and dynamically optimizes brightness in response to scene changes. Later we introduce a spatial–frequency fusion module integrating multi-scale encoder and edge-guided decoder to ensure structural consistency. Extensive experiments demonstrate that the fusion results produced by Seq-IF outperform baseline methods in terms of clarity, detail preservation, and illumination stability, achieving smoother temporal transitions and enhanced perceptual quality. Furthermore, the effectiveness of Seq-IF is validated on downstream tasks, including object detection and optical flow estimation, highlighting its applicability to real-world scenarios. Yuqi Han, Zhihui Zheng, Weijian Su, Mingkai Wei, Liang Zhang 0031, Jin-Li Suo, Qiang Zhang 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | HeRIF: A Mixture-of-Experts Framework for Infrared and Visible Image Fusion with Heterogeneous Resolutions
Songqian Zhang, Weijian Su, Yuqi Han, Jin-Li Suo, Qiang Zhang 0008 |
PRCV (5) | 4 |
| 2025 | Satellite Attitude Estimation Based on Whole-to-Part DecouplingabstractSatellite attitude estimation is a key technology in on-orbit servicing missions. However, occlusions on the satellite’s surface in space introduce local omissions in satellite images, resulting in incomplete extraction of global satellite features. This causes significant biases in the mapping from features to attitude. Besides, there is a coupling between the parameters in rotation representations such as Euler angles and quaternions, making it difficult to improve the accuracy of each attitude parameter simultaneously. Existing methods focus on improving feature extraction performance and lack an evaluation of the relationship between satellite structure and attitude. To address these issues, we proposed a satellite attitude estimation method based on whole-to-part decoupling to separate features and reconstruct rotation representation, balancing attitude estimation biases. Specifically, we constructed a multibranch mapping with feature decoupling to select robust local features that contribute to attitude estimation, mitigating the impact of occlusion on feature extraction. Meanwhile, we designed an explicit rotation representation to decouple attitude parameters and match the relevant mappings for each parameter, reducing the estimation biases. The experimental results on public datasets demonstrate that the proposed method outperforms existing methods. Chenwei Deng, Yuqi Han, Zhuokai Li |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | Decoupling-Based Cross-Layer Connection Removal for Compact Remote Sensing Image ClassificationabstractCross-layer connections can couple different layers of neural networks in diverse paths, improving the classification accuracy of remote sensing images. However, some structures are unnecessary for inference and not conducive to be deployed on starboard/airborne platforms. Roughly removing them lengthens the backpropagation path and causes the gradient vanishing problem, degrading the performance of existing end-to-end training based lightweighting methods. To address this issue, we propose a decoupling-based cross-layer connection removal framework, which decomposes the complete deep network into multiple shallower stages based on the density of connections. By optimizing each stage independently, the backpropagation path is shortened and the gradient vanishing problem is expected to be alleviated. Moreover, to approximate the original network in each stage, we perform concret reconstruction error suppression on branches, layers and channels respectively. Combining these strategies, redundant structures are removed while maintaining performance. Extensive experiments on multiple public datasets and hardware platforms demonstrate that this compact architecture achieves substantially improved deployability and inference efficiency at the cost of only marginal accuracy degradation. Yuqi Han, Chenwei Deng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | A Multivehicle Tracking Method for Video-SAR With Reliable Foreground-Background Motion Feature CompensationabstractMost of the existing video synthetic aperture radar (ViSAR) vehicle multitarget tracking (MTT) methods only perform interframe association based on the idea of appearance modeling, and are not closely integrated with the ViSAR moving target imaging characteristics, resulting in limited accuracy improvement of existing MTT methods. ViSAR moving targets have the characteristics of individual similarity, time-varying appearance, and background pseudo-motion, which have a great impact on tracking performance. In this regard, we propose a multivehicle tracking method for ViSAR with reliable foreground-background motion feature compensation (RFBMFC). Specifically, in order to improve the distinguishability of individual features, the spatial-temporal semantic sparse alignment (STSSA) module with intraframe and interframe context key information aggregation and interaction is constructed in the feature extraction stage, which can generate more accurate dense optical flow to enhance the detection and association of foreground targets. In order to improve the tracking continuity of foreground targets with time-varying appearance, the shadow-observation-state mining (SOSM) module is designed in the interframe association stage, which can cluster targets under different appearance states and adaptively restore lost target trajectories. In addition, the background motion fast compensation (BMFC) module is designed, which can learn background motion estimation and correct the trajectory prediction error of foreground targets in an end-to-end self-supervised manner to improve the MTT accuracy under camera motion. Tests on datasets captured by Sandia National Laboratories (SNL) and Beijing Institute of Radio Measurement (BIRM) show that RFBMFC outperforms many representative MTT methods. Compared with the suboptimal method, RFBMFC improves the multiobject tracking accuracy (MOTA) by 1.10% on the SNL data, and by 5.00% on the BIRM data, verifying the effectiveness of RFBMFC. Jianzhi Hong, Taoyang Wang, Yuqi Han, Weicheng Di, Tiancheng Dong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Multi-Target Tracking for Satellite Videos Guided by Spatial-Temporal Proximity and Topological RelationshipsabstractThe features of moving targets in satellite videos are sparse and similar, resulting in two major challenges for multi-target tracking: detection losses and association errors. The rich spatial-temporal proximity and topological relationships among targets in satellite videos can indicate the feature enhancement of small targets and generate discriminative individual descriptions of targets, which helps to improve the accuracy of multi-target tracking. In this article, we propose a novel Multi-target Tracking method for satellite videos guided by Spatial-Temporal Proximity and Topological Relationships (MTT-STPTR). Specifically, a spatial-temporal relationship sparse attention (STRSA) module is constructed in the feature extraction stage to accurately enhance the feature expression of small targets by capturing the cross-frame semantics relevant to the target areas. In addition, a joint feature matching (JFM) module is designed in the interframe association stage, which constructs a novel similarity measurement method of star-shaped topological structures and uses it to measure the similarity of multidimensional features of target individuals, thereby alleviating association errors caused by dense individuals with similar features. Moreover, a novel multiple granularity spatial-temporal contrastive learning (MGSTCL) module is designed to promote a balanced optimization of detection and association tasks for multicategory targets in satellite videos. Experiments conducted on two public datasets, VISO and AIR-MOT, demonstrate that MTT-STPTR outperforms existing state-of-the-art multi-target tracking methods in terms of the multiobject-tracking accuracy (MOTA) and identification$F1$-score (IDF1), indicating its effectiveness. Jianzhi Hong, Taoyang Wang, Yuqi Han |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Enhanced Aircraft Detection in Compressed Remote Sensing Images Using CMSFF-YOLOv8abstractObject detection is an important part of remote sensing image analysis to accurately identify key features in Earth observation data while minimizing resource usage. However, challenges such as multi-scale object representation and the identification of small-sized objects, particularly in compressed images, persist. To tackle this problem, we propose a state-of-the-art detector called Compressed Multiscale Feature Fusion YOLOv8 (CMSFF-YOLOv8) that employs a four-channel DWT coefficient image representation for preprocessing. This method processes the four coefficients of the 3-DWT result, arranging them into a 4-channel stacked image representation that retains the important spatial-frequency characteristics of the information. The improved preprocessing architecture design employs frequency domain filtering to retain important edge information, adaptive histogram equalization to adjust the intensity distribution locally, and multi-scale feature transfer to enable the extraction of important edge features at different scales. Squeeze-and-Excitation Convolution (SEConv) for channel-wise attention and Multi-Granularity Enhanced Feature Aggregation (MGEFA) for multi-level feature extraction of DWT coefficients are integrated into the backbone of the CMSFF-YOLOv8 architecture, enabling the selective prioritization of structural and edge features in complex settings. Spatial Pyramid Pooling Fast (SPPF) combined with shallow feature fusion networks (SFFNs) enhances small object detection. A custom neck design with upsampling facilitates multi-scale feature fusion. The result is mapped back to the LL3 channel in the compressed domain, reducing computational complexity while retaining critical global context. The proposed CMSFF-YOLOv8 detector, combined with a four-channel image representation and advanced preprocessing, demonstrates an efficient and accurate solution for object identification in compressed-domain remote sensing applications, achieving an F1-score of 0.8473, a recall of 0.82165, a precision of 0.8746, and mAP values of 0.90895 at IoU = 0.5 and 0.60063 at IoU = 0.75 on the HRPlanesV2 dataset. Sharan Thapa, Yuqi Han, Baojun Zhao, Senlin Luo |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | PAPooling: Graph-based Position Adaptive Aggregation of Local Geometry in Point CloudsabstractFine-grained geometry, obtained through the assimilation of localized point features, is crucial in the realms of object recognition and scene comprehension within point cloud contexts. Traditional point cloud backbones predominantly utilize max pooling for the amalgamation of local features, a process that tends to overlook spatial interrelations among points, consequently leading to the potential loss of fine-grained geometric details. To overcome this limitation, we introduce an innovative operation termed Position Adaptive Pooling (PAPooling), which is designed to amalgamate local features while sensitively considering the spatial positions of points. This is achieved by employing a graph-based representation to explicitly model the spatial relationships of points. PAPooling involves two principal components: first, the local graph construction , which establishes a local graph for a set of points by linking a central point with its adjacent points, thereby transforming pairwise relative positions into channel-specific attention weights; second, the attentive feature aggregation , which adeptly takes into account the contribution of each node and simulates the inter-node relationships within the local graph, effectively extracting representations of local features through a Graph Convolution Network (GCN). PAPooling’s simplicity and efficacy make it a versatile addition to widely used point-based backbones such as PointNet++ and DGCNN, offering a plug-and-play solution. Comprehensive experimental analysis demonstrates PAPooling’s enhanced capability in capturing local geometry, contributing significantly across a spectrum of applications including 3D shape classification, part segmentation, scene segmentation, and corruption defense, all with minimal computational increase. Code will be public at https://github.com/Roywangj/PAPooling/ . Jie Wang 0097, Tingfa Xu, Liqiang Song, Lihe Ding, Peng Jiang 0013, Yuqi Han, Jianan Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | Fast-Response Edge Caching Scheme for Graph DataabstractBy deploying distributed storage space on edge servers, mobile edge networks significantly enhance computation and transmission efficiency for wireless tasks. The selection of an appropriate caching policy not only optimizes bandwidth utilization but also alleviates network congestion. Given the intricate connectivity and vast data volume in edge computing, coupled with users’ demand for rapid response times, proposing a cache solution that closely matches data attributes becomes imperative to enhance overall efficiency. In this paper, we introduce RECG, a high-speed edge caching scheme designed specifically for graph data, leveraging the intricate data connectivity. RECG generates query graphs from edge servers to ensure swift and accurate identification of popular nodes. Additionally, we introduce a rapid hot-spot propagation partitioning technique to optimize the partitioning process, increasing the hit rate of partitioned subgraphs in the cache while reducing runtime. Our experimental evaluation, conducted on real-world datasets, compares RECG with specialized edge caching graph partitioning algorithms such as LGPE and other baseline algorithms. The thorough experimental results demonstrate the advantages of the proposed algorithm in terms of cache hit rate and processing delay. Pengfei Wang 0013, Yuqi Han, Feiye Ye, Qiang Zhang 0008 |
IEEE Trans. Netw. | 3 |
| 2025 | Super-NeRF: View-Consistent Detail Generation for NeRF Super-ResolutionabstractThe neural radiance field (NeRF) achieved remarkable success in modeling 3D scenes and synthesizing high-fidelity novel views. However, existing NeRF-based methods focus more on making full use of high-resolution images to generate high-resolution novel views, but less considering the generation of high-resolution details given only low-resolution images. In analogy to the extensive usage of image super-resolution, NeRF super-resolution is an effective way to generate low-resolution-guided high-resolution 3D scenes and holds great potential applications. Up to now, such an important topic is still under-explored. In this article, we propose a NeRF super-resolution method, named Super-NeRF, to generate high-resolution NeRF from only low-resolution inputs. Given multi-view low-resolution images, Super-NeRF constructs a multi-view consistency-controlling super-resolution module to generate various view-consistent high-resolution details for NeRF. Specifically, an optimizable latent code is introduced for each input view to control the generated reasonable high-resolution 2D images satisfying view consistency. The latent codes of each low-resolution image are optimized synergistically with the target Super-NeRF representation to utilize the view consistency constraint inherent in NeRF construction. We verify the effectiveness of Super-NeRF on synthetic, real-world, and even AI-generated NeRFs. Super-NeRF achieves state-of-the-art NeRF super-resolution performance on high-resolution detail generation and cross-view consistency. Yuqi Han, Tao Yu 0007, Xiaohang Yu, Di Xu 0012, Binge Zheng, Zonghong Dai, Changpeng Yang, Yuwang Wang, Qionghai Dai |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | ImmersiveNeRF: Hybrid Radiance Fields for Unbounded Immersive Light Field ReconstructionabstractThis article proposes a hybrid radiance field representation for unbounded immersive light field reconstruction which supports high-quality rendering and aggressive view extrapolation. The key idea is to first formally separate the foreground and the background and then adaptively balance learning of them during the training process. To fulfill this goal, we represent the foreground and background as two separate radiance fields with two different spatial mapping strategies. We further propose an adaptive sampling strategy and a segmentation regularizer for more clear segmentation and robust convergence. Finally, we contribute a novel immersive light field dataset, named THUImmersive, with the potential to achieve much larger space 6DoF immersive rendering effects compared with existing datasets, by capturing multiple neighboring viewpoints for the same scene, to stimulate the research and AR/VR applications in the immersive light field domain. Extensive experiments demonstrate the strong performance of our method for unbounded immersive light field reconstruction. Xiaohang Yu, Haoxiang Wang 0006, Yuqi Han, Lei Yang 0045, Tao Yu 0007, Qionghai Dai |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | Enhancing few-shot object detection through mixing and separating tuning strategies
Zhengquan Piao, Fuyong Feng, Ruina Dang, Yuqi Han |
Vis. Comput. | 6 |
| 2024 | Feature-Based Knowledge Distillation for Infrared Small Target DetectionabstractInfrared small target detection is an extremely challenging task because of its low resolution, noise interference, and the weak thermal signal of small targets. Nevertheless, despite these difficulties, there is a growing interest in this field due to its significant application value in areas such as military, security, and unmanned aerial vehicles. In light of this, we propose a feature-based knowledge distillation method(IRKD) for infrared small target detection which can efficiently transfer detailed knowledge to students. The key idea behind IRKD is to assign varying importance to the features of teachers and students in different areas during the distillation process. Treating all features equally would negatively impact the distillation results. Therefore, We have developed a Unified Channel-Spatial Attention(UCSA) module that adaptively enhancing the crucial learning areas within the features. The experimental results show that compared to other knowledge distillation methods, our student detector achieved significant improvement in mean average precision(mAP). For example, IRKD improves ResNet50 based RetinaNet from 50.5% to 59.1% mAP, and improver ResNet50 based FCOS from 53.7% to 56.6% on Roboflow. Jinglei Xue, Jianan Li 0001, Yuqi Han, Chenwei Deng, Tingfa Xu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Satellite Component Detection Based Upon Searchable Multiview Features MappingabstractSatellite component detection (SCD) is one of the key technologies for on-orbit service (OOS) in space exploration. Unlike detection tasks on the ground, the single illumination in the space environment often leads to significant shadow occlusion when capturing satellite components from multiple viewing angles, resulting in the loss of component information. Existing component detection methods utilize this information to estimate their high-dimensional feature distributions and lack descriptions of the correlations between multiview features. To address this issue, we proposed an SCD based upon searchable multiview features mapping, which describes the distribution of satellite components by capturing both the discrimination between positive and negative samples and the commonality of intraclass components. Specifically, to enhance the connection of multiview features, orthogonal subspaces for different classes are constructed by the improved cosine distance between multiview features which is performed as a subspace metric to represent the discrimination and is used to cluster the mapped features. Besides, considering inherent errors in the mapping process, a loose constraint of the metric is introduced to search for common features, maintaining the local sparsity of multiview features. In the proposed dataset and the public dataset of satellite components, the proposed method achieved mean average precisions (mAP) of 0.878 and 0.862, respectively, outperforming existing state-of-the-art methods. Chenwei Deng, Yuqi Han |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | OPAL: Occlusion Pattern Aware Loss for Unsupervised Light Field Disparity EstimationabstractLight field disparity estimation is an essential task in computer vision. Currently, supervised learning-based methods have achieved better performance than both unsupervised and optimization-based methods. However, the generalization capacity of supervised methods on real-world data, where no ground truth is available for training, remains limited. In this paper, we argue that unsupervised methods can achieve not only much stronger generalization capacity on real-world data but also more accurate disparity estimation results on synthetic datasets. To fulfill this goal, we present the Occlusion Pattern Aware Loss, named OPAL, which successfully extracts and encodes general occlusion patterns inherent in the light field for calculating the disparity loss. OPAL enables: i) accurate and robust disparity estimation by teaching the network how to handle occlusions effectively and ii) significantly reduced network parameters required for accurate and efficient estimation. We further propose an EPI transformer and a gradient-based refinement module for achieving more accurate and pixel-aligned disparity estimation results. Extensive experiments demonstrate our method not only significantly improves the accuracy compared with SOTA unsupervised methods, but also possesses stronger generalization capacity on real-world data compared with SOTA supervised methods. Last but not least, the network training and inference efficiency are much higher than existing learning-based methods. Our code will be made publicly available. Jiayin Zhao, Jingyao Wu 0003, Chao Deng 0005, Yuqi Han, Haoqian Wang, Tao Yu 0007 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Learning a Robust Topological Relationship for Online Multiobject Tracking in UAV ScenariosabstractMany existing multiobject tracking (MOT) methods tend to model each object’s feature individually. However, under acute viewpoint variation and occlusion, there may exist significant differences between the current and historical features of objects, which easily leads to object loss. To alleviate these issues, the topological relationships (i.e., geometric shapes formed by objects) should be modeled as a supplement to individual object features to maintain stability. In this article, we propose a novel MOT framework, which consists of a frame graph and association graph, to leverage the topological relationships both spatially and temporally. Technically, the frame graph models distance and angle among objects to resist viewpoint change, while the association graph utilizes the interframe temporal consistency of topological features to recover occluded objects. Extensive experiments on mainstream datasets demonstrate the effectiveness. Chenwei Deng, Jiapeng Wu, Yuqi Han, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | GradQuant: Low-Loss Quantization for Remote-Sensing Object DetectionabstractConvolutional neural network based methods have shown remarkable performance in remote sensing object detection. However, their deployment on resource-limited embedded devices is hindered by their high computational complexity. Neural network quantization methods have been proven effective in compressing and accelerating CNN models by clipping outlier activations and utilizing low-precision values to represent weights and clipped activations. Nonetheless, the clipping of outlier activations leads to distortion of object local features. Furthermore, the lack of enhanced overall feature mining exacerbates the degradation of detection accuracy. To address the limitations above, we propose an innovative clipping-free quantization method called GradQuant, which mitigates model’s quantization accuracy loss caused by clipping outlier activations and the lack of overall feature mining. Specifically, a bounded activation function (sigmoid-weighted tanh, SiTanh) is carefully designed to ensure that object features are represented within a limited range without clipping. On the basis of this, an activation substitute training (AST) method is co-designed to prompt models to focus more on non-outlier object features instead of outlier-like local ones. Extensive experiments on public remote-sensing datasets demonstrate the effectiveness of GradQuant method compared with other state-of-the-art quantization methods. Chenwei Deng, Yuqi Han, Donglin Jing, Hong Zhang 0018 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Toward Hierarchical Adaptive Alignment for Aerial Object Detection in Remote Sensing ImagesabstractThe aerial objects tend to distribute with a major variation on the scale and arbitrary orientations in remote sensing images. To meet such characteristics of the aerial object, most of the existing anchor-based detectors rely on preset anchors with variable scales, angles, and aspect ratios, which leads to the misalignment of the selection of candidate regions, the extraction of object features, and the label assignment of preset boxes, interfering the performance of the detector. To address this issue, we propose a Hierarchical Adaptive Alignment Network (HAA-Net). Specifically, we first design the Region Refinement Module (RRM), Feature Alignment module (FAM), and Potential Label Assignment Module (PLAM) to alleviate the misalignment of the region, feature, and label levels respectively; furthermore, we use the gradient equalization strategy to jointly optimize these modules at different levels, so that the whole network can be fully trained to significantly improve detection performance. Extensive experiments demonstrate that our approach can achieve superior performance in three common aerial object datasets (e.g., DOTA, HRSC2016, and UCAS-AOD) when compared with state-of-the-art detectors. Chenwei Deng, Donglin Jing, Yuqi Han, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Domain Adaptive Remote Sensing Scene Recognition via Semantic Relationship Knowledge TransferabstractScene recognition has attracted rising attentions of many researchers in the remote sensing fields, owing to the rapidly advancing of remote sensing devices in recent years. However, images obtained from various sensors dominate diverse sensor-specific characteristics, which will dramatically weaken the model transferability trained on a source data domain to a different target domain on account of the domain shift issues. To mitigate the domain discrepancy, most existing methods attend to align the cross-domain distributions. While the valuable knowledge of semantic relationships between different scenes is generally overlooked, and the underlying correlation across scenes cannot be fully discovered. For the sake of tackling this challenge, we propose an adaptive remote sensing scene recognition network, which can successfully transfer both the discriminative knowledge and cross-scene relationship from source to target. Specifically, in this paper, we acquire sensor-invariant representations in an adversarial manner and realize fine-grained conditional distribution alignment contrastively. In such a way, the tremendous domain gap can be mitigated to a large extent, and the discriminative and well-matched representations will be derived favorably. In addition, we explicitly construct class-wise relationship distributions belonging to two domains respectively and minimize their divergence to conduct semantic relationship knowledge transfer (SRKT), for the purpose of sufficiently unearthing the intrinsic semantic relative structures that can prompt generality of the model in the target domain. Finally, we conduct multiple experiments on representative multi-domain remote sensing benchmarks, and the extensive experimental results demonstrate the superiority of our proposed approach. Shuang Li 0008, Chi Harold Liu, Yuqi Han, Hao Shi 0006, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Intelli-AR Preloading: A Learning Approach to Proactive Hologram Transmissions in Mobile ARabstractMobile augmented reality (AR), which integrates virtual objects (i.e., holographic contents) with 3-D real environments in real time, has been rapidly gaining popularity in the last five years. The delivery mechanisms of these holographic contents to mobile AR devices, however, are rarely investigated. To combat bandwidth limitations that preclude providing holographic contents to user devices on-demand, in this article, we propose the intelligent AR (Intelli-AR) preloading algorithm to improve transmission efficiency in the edge-assisted network, in which edge servers proactively transmit holographic contents to the devices. Without user devices’ future motion trajectories, the Intelli-AR preloading algorithm models the user devices’ motion trajectories as Markov decision process (MDP) and adaptively learns the optimal preloading policy. The Intelli-AR preloading is decomposed into two parts and separately deployed on the edge server and the user devices to reduce the computation complexity. The Intelli-AR solution improves the ratio of successful preloading by 11.52% compared to the best baseline in the practical data set when the users’ motion trajectories tend to be more random, and by 21.97% compared to the best baseline in the data set which is synthesized from a real-life mobile AR environment. Yuqi Han, Rui Wang 0001, Jun Wu 0006, Maria Gorlatova |
IEEE Internet Things J. | 1 |
| 2022 | FAR-Net: Fast Anchor Refining for Arbitrary-Oriented Object DetectionabstractCompared with natural images, targets in remote-sensing images are often distributed with more flexible orientation, aspect ratio, and scale. Thus, anchor-based algorithms often employ plenty of preset anchors to encode the above-mentioned attributes in object detection tasks. However, they often suffer from the following issues: 1) significant computational burden caused by dense-sampling anchors; 2) serious background interference since many anchors only cover small parts of the actual target; and 3) feature misalignment between the targets with the preset anchors due to the absence of the most discriminant features for target extraction. Therefore, in this letter, a fast anchor refining network (FAR-Net) is advocated to address the remaining issues for arbitrary-oriented object detection in the remote-sensing field. To be specific, a rotation alignment module (RAM) and balanced regression loss function (BR-loss) are carefully designed in the FAR-Net. The RAM is capable of generating high-quality anchors based on a refinement convolution and adaptively aligning the convolutional features by complying with the anchor boxes to reduce redundant calculation. The BR-loss is designed by employing a balanced loss function to prevent misaligned anchors from causing major gradient descents, thereby achieving a more stable network training procedure. Extensive experiments on public remote-sensing datasets (HRSC2016 and UCAS-AOD) demonstrate the excellent detection performance of our algorithm in comparison with numerous existing detectors. Chenwei Deng, Donglin Jing, Yuqi Han, Shuliang Wang 0001, Hongshuo Wang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Effectiveness Guided Cross-Modal Information Sharing for Aligned RGB-T Object DetectionabstractIntegrating multi-modal data can significantly increase detection performance in a complex scene by introducing additional targets' information. However, most of the existing multi-modal detectors separately extract the features from the respective modalities without regarding the correlation between the modalities. Considering the spatial correlation across different modalities for aligned multi-modal data, we attempt to exploit such correlation to share target's information across different modalities, thereby enhancing the targets' feature representation capability. To this end, in this letter, we propose an Effectiveness Guided Cross-Modal Information Sharing Network (ECISNet) for aligned multi-modal data, which can still accurately detect objects when a modality fails. Specifically, the Cross-Modal Information Sharing (CIS) module is proposed to enhance the feature extraction capability by sharing information about targets across different modalities. Afterward, considering that the failed modality may interfere with other modalities when sharing information, we designed a Modal Effectiveness Guiding (MEG) module that guides the CIS module to exclude the interference of failed modalities. Extensive experiments on three latest multi-modal detection datasets demonstrate that ECISNet outperforms relevant state-of-the-art detection algorithms. Zijia An, Chunlei Liu 0001, Yuqi Han |
IEEE Signal Process. Lett. | 3 |
| 2022 | RB-Net: Training Highly Accurate and Efficient Binary Neural Networks With Reshaped Point-Wise Convolution and Balanced ActivationabstractIn this paper, we find that the conventional convolution operation becomes the bottleneck for extremely efficient binary neural networks (BNNs). To address this issue, we open up a new direction by introducing a reshaped point-wise convolution (RPC) to replace the conventional one to build BNNs. Specifically, we conduct a point-wise convolution after rearranging the spatial information into depth, with which at least$2.25\times $computation reduction can be achieved. Such an efficient RPC allows us to explore more powerful representational capacity of BNNs under a given computation complexity budget. Moreover, we propose to use a balanced activation (BA) to adjust the distribution of the scaled activations after binarization, which enables significant performance improvement of BNNs. After integrating RPC and BA, the proposed network, dubbed as RB-Net, strikes a good trade-off between accuracy and efficiency, achieving superior performance with lower computational cost against the state-of-the-art BNN methods. Specifically, our RB-Net achieves 66.8% Top-1 accuracy with ResNet-18 backbone on ImageNet, exceeding the state-of-the-art Real-to-Binary Net (65.4%) by 1.4% while achieving more than$3\times $reduction (52M vs. 165M) in computational complexity. Chunlei Liu 0001, Wenrui Ding, Peng Chen 0037, Bohan Zhuang, Yufeng Wang 0004, Yang Zhao 0019, Baochang Zhang 0001, Yuqi Han |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2021 | Robust Shadow Tracking for Video SARabstractVideo synthetic aperture radar (Video SAR) has sparked lots of research attention since it provides the ability of continuously observing and tracking the object of interest. Instead of tracking a target directly, in Video SAR, it is preferred to track its shadow due to the target's instable backscattering characteristic and defocused patterns. However, the current tracking algorithms do not work well since they always suffer from the interference caused by surrounding clutters and cannot locate the target precisely. To address these issues, we propose a robust shadow tracking algorithm for Video SAR in this letter. We equip our tracker with the spatial-temporal (ST) information and saliency-based detection mechanism (SD) against the distractions and background clutters. The proposed optimization formula could be solved efficiently using the alternating direction method of multipliers (ADMM) technique and fully carried in the Fourier domain at a low computational burden. Experiments have demonstrated that the proposed algorithm performs favorably against other state-of-the-art methods. Baojun Zhao, Yuqi Han, Hongshuo Wang, Linbo Tang, Taoyang Wang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | Learning Dynamic Spatial-Temporal Regularization for UAV Object TrackingabstractWith the wide vision and high flexibility, unmanned aerial vehicle (UAV) has been widely used into object tracking in recent years. However, its limited computing capability poses a great challenges to tracking algorithms. On the other hand, Discriminative Correlation Filter (DCF) based trackers have attracted great attention due to their computational efficiency and superior accuracy. Many studies introduce spatial and temporal regularization into the DCF framework to achieve a more robust appearance model and further enhance the tracking performance. However, such algorithms generally set fixed spatial or temporal regularization parameters, which lack flexibility and adaptability under cluttered and challenging scenarios. To tackle such issue, in this letter, we propose a novel DCF tracking model by introducing dynamic spatial regularization weight, which encourage the filter focuses on more reliable region during training stage. Furthermore, our method could optimize the spatial and temporal regularization weight simultaneously using Alternative Direction Method of Multiplies (ADMM) technique method, where each sub-problem has closed-form solution. Through the joint optimization, our tracker could not only suppress the potential distractors but also construct robust target appearance on the basis of reliable historical information. Experiments on two UAV benchmarks have demonstrated that our tracker performs favorably against other state-of-the-art algorithms. Chenwei Deng, Shuangcheng He, Yuqi Han, Boya Zhao |
IEEE Signal Process. Lett. | 3 |
| 2021 | Cache Placement Optimization in Mobile Edge Computing Networks With Unaware Environment - An Extended Multi-Armed Bandit ApproachabstractCaching high-frequency reuse contents at the edge servers in the mobile edge computing (MEC) network omits the part of backhaul transmission and further releases the pressure of data traffic. However, how to efficiently decide the caching contents for edge servers is still an open problem, which refers to the cache capacity of edge servers, the popularity of each content, and the wireless channel quality during transmission. In this paper, we discuss the influence of unknown user density and popularity of content on the cache placement solution at the edge server. Specifically, towards the implementation of the cache placement solution in the practical network, there are two problems needing to be solved. First, the estimation of unknown users’ preference needs a huge amount of records of users’ previous requests. Second, the overlapping serving regions among edge servers cause the wrong estimation of users’ preference, which hinders the individual decision of caching placement. To address the first issue, we propose a learning-based solution to adaptively optimize the cache placement policy without any previous knowledge of the user density and the popularity of the contents. We develop the extended multi-armed bandit (Extended MAB), which combines the generalized global bandit (GGB) and Standard Multi-armed bandit (MAB), to iteratively estimate both a global parameter, i.e., the user density, and individual parameters, i.e., the popularity of each content. For the second problem, a multi-agent Extended MAB based solution is presented to avoid the mis-estimation of parameters and achieve the decentralized cache placement policy. The proposed solution determines the primary time slot and secondary time slot for each edge server. The edge servers estimate expected satisfied user number of caching a content with the overlap information and determine the cache placement solution. The proposed strategies are proven to achieve the bounded regret according to the mathematical analysis. Extensive simulations verify the optimality of the proposed strategies when comparing with baselines. Yuqi Han, Lihua Ai, Rui Wang 0001, Jun Wu 0006, Dian Liu, Haoqi Ren |
IEEE Trans. Wirel. Commun. | 1 |
| 2020 | High-Performance Visual Tracking With Extreme Learning Machine FrameworkabstractIn real-time applications, a fast and robust visual tracker should generally have the following important properties: 1) feature representation of an object that is not only efficient but also has a good discriminative capability and 2) appearance modeling which can quickly adapt to the variations of foreground and backgrounds. However, most of the existing tracking algorithms cannot achieve satisfactory performance in both of the two aspects. To address this issue, in this paper, we advocate a novel and efficient visual tracker by exploiting the excellent feature learning and classification capabilities of an emerging learning technique, that is, extreme learning machine (ELM). The contributions of the proposed work are as follows: 1) motivated by the simplicity and learning ability of the ELM autoencoder (ELM-AE), an ELM-AE-based feature extraction model is presented, and this model can provide a compact and discriminative representation of the inputs efficiently and 2) due to the fast learning speed of an ELM classifier, an ELM-based appearance model is developed for feature classification, and is able to rapidly distinguish the object of interest from its surroundings. In addition, in order to cope with the visual changes of the target and its backgrounds, the online sequential ELM is used to incrementally update the appearance model. Plenty of experiments on challenging image sequences demonstrate the effectiveness and robustness of the proposed tracker. Chenwei Deng, Yuqi Han, Baojun Zhao |
IEEE Trans. Cybern. | 2 |
| 2019 | Spatial-Temporal Context-Aware TrackingabstractDiscriminative correlation filters (DCFs) have recently achieved competitive performance in visual tracking benchmarks. However, most of the existing DCF trackers only consider the spatial features of the target and could hardly benefit from the inter-frame and historical information, which may degrade the tracking performance when occlusion and deformation occurs. To tackle the above-mentioned issues, in this letter, by introducing the temporal constrain into the DCF tracker, we advocate our spatial-temporal context-aware tracker. Through jointly modeling the spatial context and historical target information, our tracker could not only adapt the appearance change but also maintain a relatively stable filter due to the small target variation between inter-frames. Furthermore, we show that the proposed objective formula could be directly solved using the Alternating Direction Method of Multipliers (ADMM) technique with low computational cost. Experiments on the large-scale benchmark demonstrate that the proposed trackers perform favorably against other state-of-the-art methods. Yuqi Han, Chenwei Deng, Boya Zhao, Baojun Zhao |
IEEE Signal Process. Lett. | 1 |
| 2019 | State-Aware Anti-Drift Object TrackingabstractCorrelation filter (CF) based trackers have aroused increasing attentions in visual tracking field due to the superior performance on several datasets while maintaining high running speed. For each frame, an ideal filter is trained in order to discriminate the target from its surrounding background. Considering that the target always undergoes external and internal interference during tracking procedure, the trained tracker should not only have the ability to judge the current state when failure occurs, but also to resist the model drift caused by challenging distractions. To this end, we present a State-aware Anti-drift Tracker (SAT) in this paper, which jointly model the discrimination and reliability information in filter learning. Specifically, global context patches are incorporated into filter training stage to better distinguish the target from backgrounds. Meanwhile, a color-based reliable mask is learned to encourage the filter to focus on more reliable regions suitable for tracking. We show that the proposed optimization problem could be efficiently solved using Alternative Direction Method of Multipliers and fully carried out in Fourier domain. Furthermore, a Kurtosis-based updating scheme is advocated to reveal the tracking condition as well as guarantee a high-confidence template updating. Extensive experiments are conducted on OTB-100 and UAV-20L datasets to compare the SAT tracker with other relevant state-of-the-art methods. Both quantitative and qualitative evaluations further demonstrate the effectiveness and robustness of the proposed work. Yuqi Han, Chenwei Deng, Baojun Zhao, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2017 | Adaptive feature representation for visual trackingabstractRobust feature representation plays significant role in visual tracking. However, it remains a challenging issue, since many factors may affect the experimental performance. The existing method, which combine different features by setting them equally with the fixed weight, could hardly solve the issues, due to the different statistical properties of different features across various of scenarios and attributes. In this paper, by exploiting the internal relationship among these features, we develop a robust method to construct a more stable feature representation. More specifically, we utilize a co-training paradigm to formulate the intrinsic complementary information of multi-feature template into the efficient correlation filter framework. We test our approach on challenging sequences with illumination variation, scale variation, deformation etc. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods favorably. Yuqi Han, Chenwei Deng, Zengshuo Zhang, Jiatong Li 0004, Baojun Zhao |
ICIP | 1 |
| 2016 | Robust Uncoded Video Transmission under Practical Channel EstimationabstractThis research solves the performance degradation problem of uncoded video system when the multiplicative noise caused by channel estimation error is not negligible. Through extensive analysis, we find that transmitting a specific part of source data by multiple channel use can decrease the mean end-to-end distortion and stabilize the video quality. However, under limited power and bandwidth constraints, allocating multiple channel use to one part of data means dropping another part of data. How to allocate the channel use to highly diverse video coefficients is an NP-hard problem. For the sake of practical implementation, a greedy iterative algorithm is proposed to achieve the optimal trade off between distortion decrease by multiple channel use and distortion increase from data dropping. Extensive simulations validate our analysis result. The proposed algorithm achieve 1.99~8.63dB gain compared to existing schemes without considering the multiplicative noise. Hao Cui 0001, Dian Liu, Yuqi Han, Jun Wu 0006 |
GLOBECOM | 3 |