You He 0002

dblp:90/3779-2 · DBLP profile ↗
← Back
63ranked-venue papers
0as first author
55since 2021 · last 2026
0000-0002-6111-340XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 23 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 16 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 One Aligned LLM to Serve Them All: A Transfer Recipe for Training VLMs without Visual-Language Re-Alignment
Jiazuo Yu 0001, Yunzhi Zhuge, Lu Zhang 0053, Wei Zhou 0021, Dong Wang 0004, Huchuan Lu, You He 0002
Int. J. Comput. Vis.8
2026 CDTFusion: Crossing Domain and Task for Infrared and Visible Image Fusion
abstract
Infrared and visible images present different domains that hinder the fusion process, thereby losing texture details. Besides, the low-level fusion and subsequent high-level segmentation appear cross-task feature gap that impedes their mutual promotion, causing blurred object edges. Addressing the above issues, this paper proposes a novel infrared and visible image fusion method that simultaneously crosses domain and task. First, a swap image translation strategy is built to transfer the features of visible and infrared images into an adaptive domain. Meanwhile, a global-local constraint is introduced to achieve overall domain space transfer, and shorten their feature distance. Second, a task interaction & query module is designed to explore the cross-task feature interactive relationship, which is then used as a bridge to realize the gradient backpropagation. Thus, a fine-grained mapping from the segmentation feature to fusion feature is obtained. Extensive experiments demonstrate that the proposed method exhibits superior fusion and segmentation performance than the state-of-the-art methods.
Wenda Zhao 0003, You He 0002, Huchuan Lu
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Global Structure-aware and Feature-augmented Graph Neural Network for Heterophilic Graphs
abstract
Graph Neural Networks (GNNs) have been widely used across various fields under the homophily assumption that connected nodes are similar. However, in heterophilic graphs, where connected nodes tend to have dissimilar features, existing GNNs still face some limitations. From the perspective of structure, shallow GNNs could not capture the high-order node information, whereas deep GNNs may suffer from the over-smoothing problem. From the perspective of feature, the useful information of high-order similar nodes is often weakened by low-order dissimilar nodes in the feature update phase. To address the above problems, we propose a Global Structure-aware and Feature-augmented Graph Neural Network (GSF-GNN) to alleviate the limitations from the perspectives of structure and feature. Specifically, from the structure perspective, we design a Structure-based Global Propagation (SGP) module to establish global connections among nodes and adaptively adjust edge weights for message propagation. From the feature perspective, we introduce a Feature-augmented Compensatory Update (FCU) module, which employs a multi-view feature updating mechanism to enhance node features from different perspectives. Our theoretical analysis formally demonstrates the effectiveness of GSF-GNN in heterophilic graphs. Experiments on heterophilic and homophilic benchmark datasets validate the effectiveness of GSF-GNN across various graph structures. Moreover, GSF-GNN achieves stable performance across multiple layers and effectively alleviates the over-smoothing problem. Our codes are available on https://github.com/huijieliu2023/GSF-GNN .
Huijie Liu 0001, Shulan Ruan, Qi Liu 0003, Mingyue Cheng 0004, Zhenya Huang, Yu Liu 0005, Enhong Chen, You He 0002
ACM Trans. Inf. Syst.8
2025 ReNeg: Learning Negative Embedding with Reward Guidance
abstract
In text-to-image (T2I) generation applications, negative embeddings have proven to be a simple yet effective approach for enhancing generation quality. Typically, these negative embeddings are derived from user-defined negative prompts, which, while being functional, are not necessarily optimal. In this paper, we introduce ReNeg, an end-to-end method designed to learn improved Negative embeddings guided by a Reward model. We employ a reward feedback learning framework and integrate classifier-free guidance (CFG) into the training process, which was previously utilized only during inference, thus enabling the effec tive learning of negative embeddings. We also propose two strategies for learning both global and per-sample negative embeddings. Extensive experiments show that the learned negative embedding significantly outperforms null-text and handcrafted counterparts, achieving substantial improvements in human preference alignment. Additionally, the negative embedding learned within the same text embedding space exhibits strong generalization capabilities. For example, using the same CLIP text encoder, the negative embedding learned on SD1.5 can be seamlessly transferred to text-to-image or even text-to-video models such as ControlNet, ZeroScope, and VideoCrafter2, resulting in consistent performance improvements across the board. Code is available at https://github.com/AMD-AIG-AIMA/ReNeg.
Xiaomin Li 0001, Yixuan Liu 0004, Takashi Isobe, Xu Jia 0012, Qinpeng Cui, Dong Zhou 0003, Dong Li 0025, You He 0002, Huchuan Lu, Zhongdao Wang, Emad Barsoum
CVPR8
2025 TrackFusion: Enhancing Multi-Object Tracking With Temporal Trajectory Modeling and Frame-Integrated Detection
abstract
Although MOTIP is the SOTA multi-object tracking method, there are still some issues that limit its performance. First, MOTIP still has defects in temporal information modeling, which leads to the failure to fully utilize the historical information of the tracked target and affects the correlation performance of the model. Second, in MOT, objects in consecutive video frames usually have temporal continuity and spatial consistency. Therefore, the object information of the previous frame can effectively assist the detection of the current frame. However, MOTIP performs independent detection between each frame, which does not fully utilize the correlation information between frames, resulting in suboptimal model performance. To address the above problems, we propose TrackFusion, which optimizes model performance from the perspective of trajectory modeling and inter-frame joint detection. First, we extract embeddings in video sequences through a Transformer-based detector, then combine the embeddings of the same object in different frames into sequences and input them into the trajectory modeling module for sequence association. This strategy effectively enhances the association ability. Thanks to these improvements, TrackFusion’s HOTA on the DanceTrack test set reached 68.6%, an increase of 1.1% compared to MOTIP’s 67.5%.
Shuai Liu 0009, Bingyang Wang, Jiaojiao Dai, Jinqing Qi, Huchuan Lu, You He 0002
ICASSP7
2025 Towards Survivability in Complex Motion Scenarios: RGB-Event Object Tracking via Historical Trajectory Prompting
abstract
Event data has recently emerged as a valuable complement to object tracking, offering dense temporal resolution and a high dynamic range. However, existing RGB-Event trackers struggle with targets exhibiting complex motion trajectories, where RGB features alone fail to provide sufficient discrimination. To address this, we propose EventTPT, an innovative RGB-Event tracking framework that leverages pivotal prompts embedded in historical trajectories for enhanced tracking. Specifically, EventTPT integrates the trajectories of multiple adjacent frames into a single event image using a time-weighted aggregation and subsequently inputs this as a visual prompt into the tracker for current frame locating. A cross-modal adaptive fusion module is further designed for object perception in scenarios with photometric inconsistency. Additionally, we introduce EventUAV, a novel and challenging RGB-Event tracking benchmark featuring objects with intricate motion dynamics and poor visibility in RGB-only modalities. Extensive experiments demonstrate that EventTPT surpasses state-of-the-art trackers on EventUAV and achieves competitive performance on other benchmarks (e.g., COESOT and VisEvent), underscoring its strong generalizability and robustness for resilient robotic vision systems. The code can be found at https://github.com/xiawenhao2022/EventTPT.
Wenhao Xia, Jiawen Zhu 0003, Jinqing Qi, You He 0002, Xu Jia 0012
ICRA5
2025 FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
abstract
Multi-modal Large Language Models (MLLMs) have shown remarkable capabilities across a wide range of vision-language tasks. However, due to the restricted input resolutions, MLLMs face significant challenges in precisely understanding and localizing visual details in high-resolution images---particularly when dealing with extra-small objects embedded in cluttered contexts. To address this issue, we propose FineRS, a two-stage MLLM-based reinforcement learning framework for jointly reasoning and segmenting extremely small objects within high-resolution scenes. FineRS adopts a coarse-to-fine pipeline comprising Global Semantic Exploration (GSE) and Localized Perceptual Refinement (LPR). Specifically, GSE performs instruction-guided reasoning to generate a textural response and a coarse target region, while LPR refines this region to produce an accurate bounding box and segmentation mask. To couple the two stages, we introduce a locate-informed retrospective reward, where LPR's outputs are used to optimize GSE for more robust coarse region exploration. Additionally, we present FineRS-4k, a new dataset for evaluating MLLMs on attribute-level reasoning and pixel-level segmentation on subtle, small-scale targets in complex high-resolution scenes. Experimental results on FineRS-4k and public datasets demonstrate that our method consistently outperforms state-of-the-art MLLM-based approaches on both instruction-guided segmentation and visual reasoning tasks.
Lu Zhang 0053, Jiazuo Yu 0001, Haomiao Xiong, Ping Hu 0001, Yunzhi Zhuge, Huchuan Lu, You He 0002
NeurIPS7
2025 Communication-Efficient Collaborative Perception with Semantic and Statistical Compression
Yuankun Zeng, Zhi Li 0057, Shulan Ruan, Yu Liu 0005, You He 0002
PRCV (11)6
2025 Self-calibrated region-level regression for crowd counting
Jiawen Zhu 0003, Wenda Zhao 0003, You He 0002, Huchuan Lu
Sci. China Inf. Sci.3
2025 MoE-Adapters++: Toward More Efficient Continual Learning of Vision-Language Models Via Dynamic Mixture-of-Experts Adapters
abstract
In this paper, we first propose MoE-Adapters, a parameter-efficient training framework to alleviate long-term forgetting issues in incremental learning with Vision-Language Models (VLM). Our MoE-Adapters leverages incrementally added routers to activate and integrate exclusive expert adapters from a pre-defined static expert set, enabling the pre-trained CLIP to efficiently adapt to new tasks. To preserve the zero-shot capability of VLM, a Distribution Discriminative Auto-Selector (DDAS) is introduced that automatically routes in-distribution and out-of-distribution inputs to the MoE-Adapters and the original CLIP, respectively. However, relying on a static expert set and a separate distribution selector can lead to parameter redundancy and increased training complexity. In response, we further extend an MoE-Adapters++ framework by introducing dynamic MoE-adapters, which allows experts to be adaptively involved during the continual learning process. Additionally, a Latent Embedding Auto-Selector (LEAS) is proposed that incorporates distribution selection within CLIP to create a more unified architecture. Extensive experiments across diverse settings demonstrate that the proposed method consistently surpasses previous state-of-the-art approaches while concurrently improving training efficiency.
Jiazuo Yu 0001, Zichen Huang 0004, Yunzhi Zhuge, Lu Zhang 0053, Ping Hu 0001, Dong Wang 0004, Huchuan Lu, You He 0002
IEEE Trans. Pattern Anal. Mach. Intell.8
2025 FreeFusion: Infrared and Visible Image Fusion via Cross Reconstruction Learning
abstract
Existing fusion methods empirically design elaborate fusion losses to retain the specific features from source images. Since image fusion has no ground truth, the hand-crafted losses may not make the fused images cover all the vital features, and then affect the performance of the high-level tasks. Here, there are two main challenges: domain discrepancy among source images and semantic mismatch at different-level tasks. This paper proposes an infrared and visible image fusion via cross reconstruction learning, which doesn't using any hand-crafted fusion losses, but prompts the network to adaptively fuse complementary information of source images. Firstly, we design a cross reconstruction learning model that decouples the fusion features to reconstruct another-modality source image. Thus, the fusion network is forced to learn the domain-adaptive representations of two modal features, which enables their domain alignment in a latent space. Secondly, we propose a dynamic interactive fusion strategy that builds a correlation matrix between fusion features and object semantic features to overcome the semantic mismatch. Further, we enhance the strong correlation features and suppress the weak correlation features to improve the interactive ability. Extensive experiments on three datasets demonstrate the superior fusion performance compared to the state-of-the-art methods, concurrently facilitating the segmentation accuracy.
Wenda Zhao 0003, Hengshuai Cui, You He 0002, Huchuan Lu
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Remote Sensing Image Generation via Object Text Decoupling
abstract
Remote sensing images usually reveal various objects with complex structures and different locations within vast ground area backgrounds. That leads to a major challenge for conventional generative models in handling remote sensing objects with correct shapes and clear textures. Integrating additional object-level controls can be a potential solution to improve generation quality, yet previous approaches inject the object-related conditions by specifying their locations, causing a limitation in object layout in generated results. To enable high object fidelity, high layout diversity and object customizable generation for remote sensing images, we propose a remote sensing image generation via object text decoupling, namely OTD-GAN. OTD-GAN takes advantage of the inherent text-to-image generation procedure and adaptively integrates the decoupled textual representations of visual objects into the global captions, thus achieving object-level controls without layout restrictions. Specifically, we design an object text decoupling module to predict a semantically consistent textual representation for each object. By decoupling the textual representation into a class invariant part and an object specific part, the converted representation is able to catch general semantic for similar objects as well as differentiated details for individual objects. After that, we use an object text semantic enhancement module to fuse the obtained object text representations with the global captions to enrich the object-related semantic within the textual modality. As a result, the generator will benefit from the object conditions and reinforce the generation quality while remaining flexibility to create diverse layouts. Extensive experiments on remote sensing image-caption datasets including NWPU-Captions and RSICD demonstrate that our method achieves leading performance compared to existing state-of-the-art approaches.
Wenda Zhao 0003, Zhepu Zhang, Fan Zhao 0005, You He 0002, Huchuan Lu
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 RVMamba: Selective Text-Vision Mamba for Referring Video Object Segmentation
abstract
Existing RVOS methods typically employ Transformers to model global cross-modal, temporal-spatial correspondences, but their quadratic complexity limits deployment on resource-constrained devices. To overcome this limitation, Mamba offers a sequence modeling framework with linear computational complexity. We introduceRVMamba, which utilizes weight modulation to selectively update hidden states across text-frame sequences, enabling effective linguistic context propagation, and a learning-based scanning strategy to efficiently capture spatio-temporal dependencies with linear memory consumption. Extensive experiments demonstrate thatRVMambaachieves state-of-the-art performance on public benchmarks, with significantly reduced memory growth, offering an efficient and scalable solution for long video processing.
Zhenyu Chen 0001, Jiawen Zhu 0003, Lu Zhang 0053, Ping Hu 0001, Yunzhi Zhuge, Huchuan Lu, You He 0002
IEEE Signal Process. Lett.7
2025 MoBox: Enhancing Video Object Segmentation With Motion-Augmented Box Supervision
abstract
We propose MoBox, a low-cost solution for semi-supervised video object segmentation that requires only bounding boxes as manual annotations for training. Built upon a mature semi-supervised video object segmentation network, we redesign the training losses and employ a more stringent training strategy. Specifically, we introduce a well-designed constraint term that enhances traditional spatial projection by simultaneously leveraging the projections of both the ground-truth box and the predicted mask across two axes, rather than evaluating discrepancies along the x-axis and y-axis independently. To harness the intrinsic properties of videos, considering the underlying correspondence between motion represented by optical flow and the original image, we incorporate motion coherence information into the color consistency loss as supplementary information and propose a motion discrepancy loss to obtain accurate boundaries. Additionally, to mitigate the ambiguity of weak supervision, we further introduce the pseudo strict constraint during training, which significantly improves model performance. Our approach yields competitive scores on popular benchmarks, achieving a$\mathcal {J}\& \mathcal {F}$score of 78.6 on the DAVIS 2017 validation set and an Overall score of 78.0 on the YouTube-VOS 2018 validation set. These results highlight the efficacy of MoBox, demonstrating that the semi-supervised video object segmentation model can be effectively trained using only motion-augmented box supervision and intrinsic information of videos.
Xiaomin Li 0001, Dezhuang Li, Mengmeng Ge 0002, Xu Jia 0012, You He 0002, Huchuan Lu
IEEE Trans. Circuits Syst. Video Technol.6
2025 Rotation-Invariant Knowledge Distillation for Remote Sensing Object Detection
abstract
Detecting small-rotated objects in remote sensing remains a challenging task due to feature dilution and insufficient rotation invariance. Feature dilution arises when small object features are overwhelmed by background noise and progressively lost as network depth increases. Meanwhile, the lack of rotation invariance stems from the fixed nature of convolution, which struggles to handle arbitrary orientations. To address these challenges, we propose a rotation-invariant knowledge distillation, a visual-language models (VLMs) driven knowledge distillation framework tailored for optimizing small-rotated object detection in remote sensing. Our method introduces two novel components:Enhanced-Consistency Feature Distillation(ECFD) andRotation-Invariant Feature Distillation(RIFD). ECFD mitigates feature dilution by aligning consistent language representations from VLMs with cross-depth features, ensuring consistent small-rotated object representation across different depths. RIFD enhances rotation invariance by leveraging VLMs to distill robust rotational knowledge into detectors, aligning positive and negative language features with detector features to reduce sensitivity to orientation changes and mitigate class confusion. Without introducing additional computational overhead during inference, our method significantly improves the performance of remote sensing object detectors. Extensive experiments on public remote sensing datasets with complex scenes demonstrate the state-of-the-art results. Code is available at https://github.com/Shower-Lee9527/CRKD.
Feiyi Li, Xiao Zhang 0050, Wenda Zhao 0003, You He 0002
IEEE Trans. Geosci. Remote. Sens.5
2025 Weakly Supervised Cross Mixer for Infrared and Visible Image Fusion
abstract
Remote sensing infrared and visible image fusion aims to integrate information from multiple source images to enhance visual representation and support high-level visual tasks. However, due to the lack of ground truth supervision, most existing methods depend on predefined fusion-specific loss functions. Such manual definitions of cross-modal feature representations are often incomplete, potentially leading to information loss and reduced fusion performance. Complete information retention in fusion results can significantly benefit high-level tasks. Motivated by this, we propose a novel approach that utilizes weakly-supervised segmentation guidance for comprehensive information representation ability, eliminating the need for fusion-specific losses. Firstly, a self-supervised reconstruction model is proposed to obtain a decoder with robust cross-modal feature representation, which can adaptively reconstruct images based on cross-modal features. Then we design a weakly-supervised feature mining module to capture fused features with complete cross-modal information representation under the guidance of the segmentation task. The finial fusion result is adaptively reconstructed from the fused features by the cross-modal adaptive decoder without using any constraints. Extensive experiments demonstrate that the proposed method effectively mitigates information loss and achieves superior performance in both fusion and segmentation tasks compared to the state-of-the-art methods. Model and code are available at https://github.com/wangwenbo26/WSCM.
Wenda Zhao 0003, You He 0002
IEEE Trans. Geosci. Remote. Sens.4
2025 CADDN: A Content-Aware Downsampling-Based Detection Method for Small Objects in Remote Sensing Images
abstract
A key issue of existing deep-learning-based object detection methods in remote sensing images is that they often struggle to differentiate the background and small object regions due to multi-level downsampling operations therein. Downsampling operations help extract high-level semantic features but result in excessive loss of spatial features of small objects. In this paper, we propose a new small object detector using multispectral remote sensing images, named content-aware downsampling-based detection network (CADDN), where we newly design a content-aware downsampling-based module (CADM). Unlike conventional downsampling operations that apply uniform downsampling parameters across the entire feature map, CADM adaptively assigns higher weights to feature elements that are critical for distinguishing objects from the background, and this assignment is guided by the contextual awareness of object locations during the downsampling process. Experiments based on multispectral remote sensing images with small ships and vehicles demonstrate that CADM can accurately identify and preserve the locations of important object-related features, and CADDN correspondingly achieves superior small object detection performance than state-of-the-art methods.
Linping Zhang, Yu Liu 0005, Xueqian Wang 0002, You He 0002, Gang Li 0008, Chang Liu 0053, Zhizhuo Jiang, Yang Liu 0119
IEEE Trans. Geosci. Remote. Sens.4
2025 Diverse Text-Prompt Generation for Remote Sensing Image Classification
abstract
Inadequate remote sensing image training data usually makes remote sensing image classification models achieve low accuracy. Thus, we propose a diverse text-prompt generation learning (DPL) method. The context optimization (CoOp) model transfers the feature extraction capabilities of the CLIP model to downstream tasks with learnable text prompts. However, the small number of samples in remote sensing images can easily lead to noise. In contrast, DPL introduces a diverse text-prompt generation structure. Due to the diversity of multiple prompts, noise generated by inadequate samples is suppressed. Moreover, in order to keep the prompts diverse, we propose a prompt diversity loss. This loss pulls the prompts away from each other, which suppresses the noise generated by the limited training samples. Extensive experiments show that our method achieves superior performance than the existing methods on DOTA, HRRSD, and NWPU VHR-10 datasets. The model and code are available athttps://github.com/LvXiangzhu/DPL.
Wenda Zhao 0003, Xiangzhu Lv, Ruikun He, Fan Zhao 0005, You He 0002
IEEE Trans. Geosci. Remote. Sens.6
2025 Cross-Domain Few-Shot Remote Sensing Object Classification via Triplet Relation-Aware Metric
abstract
In real-world scenarios, peculiar remote sensing categories are difficult to collect on account of high cost and technical requirements. Moreover, there exists domain distribution gap among different datasets. Existing methods leverage inter-class and intra-class relations to enhance feature representation. Since remote images are shot from top to bottom, there is little difference between classes. Thus, such distance constraint only forms decision boundary between different classes. This paper proposes a triplet relation-aware metric for cross-domain few-shot remote sensing object classification, where the triplet relation-aware metric adjusts the distances among three kinds of inter-instance relations (i.e., same instance, same class and different class relations) to obtain a precise and effective feature representation. Especially, the distance of the same instance is regarded as a distance coordinate origin to guide distance metric learning. In this way, we constitute richer feature relations to promote representation learning in the source domain. Concretely, this procedure is optimized by the supervision of the designed relation-aware soft label based on the distance coordinate origin. Then, we align the triplet relation-aware metric between source domain and pseudo domain generated by the proposed episode style adversarial attack, thereby obtaining a domain-invariant feature representation. Extensive experiments on five widely-used remote sensing datasets demonstrate the superior performance of the proposed method compared with the state of the arts. Code is available at: https://github.com/jackhdpbl/TRAM.
Ruikun He, Wenda Zhao 0003, You He 0002
IEEE Trans. Image Process.4
2025 MaskTrack: Auto-Labeling and Stable Tracking for Video Object Segmentation
abstract
Video object segmentation (VOS) has witnessed notable progress due to the establishment of video training datasets and the introduction of diverse, innovative network architectures. However, video mask annotation is a highly intricate and labor-intensive task, as meticulous frame-by-frame comparisons are needed to ascertain the positions and identities of targets in the subsequent frames. Current VOS benchmarks often annotate only a few instances in each video to save costs, which, however, hinders the model's understanding of the complete context of the video scenes. To simplify video annotation and achieve efficient dense labeling, we introduce a zero-shot auto-labeling strategy based on the segment anything model (SAM), enabling it to densely annotate video instances without access to any manual annotations. Moreover, although existing VOS methods demonstrate improving performance, segmenting long-term and complex video scenes remains challenging due to the difficulties in stably discriminating and tracking instance identities. To this end, we further introduce a new framework, MaskTrack, which excels in long-term VOS and also exhibits significant performance advantages in distinguishing instances in complex videos with densely packed similar objects. We conduct extensive experiments to demonstrate the effectiveness of the proposed method and show that without introducing image datasets for pretraining, it achieves excellent performance on both short-term (86.2% in YouTube-VOS val) and long-term (68.2% in LVOS val) VOS benchmarks. Our method also surprisingly demonstrates strong generalization ability and performs well in visual object tracking (VOT) (65.6% in VOTS2023) and referring VOS (RVOS) (65.2% in Ref YouTube VOS) challenges.
Zhenyu Chen 0001, Lu Zhang 0053, Ping Hu 0001, Huchuan Lu, You He 0002
IEEE Trans. Neural Networks Learn. Syst.5
2025 Last-Iterate Convergence to Approximate Nash Equilibria in Multiplayer Imperfect Information Games
abstract
Imperfect information and multiple players are the two common features of real-world games. However, few of the existing game-theoretic methods are applicable to multiplayer imperfect information games (IIGs) when it comes to finding Nash equilibria. Moreover, the commonly used methods that rely on average-iterate convergence are not conducive to deep reinforcement learning (DRL), which is widely applied to large-scale problems, as it is costly to preserve average policies under function approximation. To deal with these problems, we construct a continuous-time dynamic named imperfect-information exponential-decay score-based learning (IESL) by considering the concept of Nash distribution [a type of quantal response equilibrium (QRE)] in IIGs. Theoretically, we prove the last-iterate convergence of IESL to approximate Nash equilibria in multiplayer IIGs under the assumption of individual concavity. Empirically, we verify that IESL converges in six poker scenarios, with the ultimate NashConv lower than that of the comparative methods (including counterfactual regret minimization (CFR), replicator dynamics (RDs), and their variants) in multiplayer Leduc hold'em. When compared with the existing equilibrium-finding algorithms in multiplayer normal-form games (NFGs), IESL also demonstrates a more stable performance. In addition, we observe a trade-off between the difficulty of IESL's last-iterate convergence and the NashConv of the convergent policies, which aligns with our convergence analysis based on the hypomonotonicity of the game.
Runyu Lu, Yuanheng Zhu, Dongbin Zhao, Yu Liu 0005, You He 0002
IEEE Trans. Neural Networks Learn. Syst.5
2024 Multi-Agent Collaborative Perception via Motion-Aware Robust Communication Network
abstract
Collaborative perception allows for information sharing between multiple agents, such as vehicles and infrastructure, to obtain a comprehensive view of the environment through communication and fusion. Current research on multi-agent collaborative perception systems often assumes ideal communication and perception environments and neglects the effect of real-world noise such as pose noise, motion blur, and perception noise. To address this gap, in this paper, we propose a novel motion-aware robust communication network (MRCNet) that mitigates noise interference and achieves accurate and robust collaborative perception. MRCNet consists of two main components: multi-scale robust fusion (MRF) addresses pose noise by developing cross-semantic multi-scale enhanced aggregation to fuse features of different scales, while motion enhanced mechanism (MEM) captures motion context to compensate for information blurring caused by moving objects. Experimental results on popular collaborative 3D object detection datasets demonstrate that MRCNet outperforms competing methods in noisy scenarios with improved perception performance using less bandwidth. Our code will be released at https://github.com/IndigoChildren/collaborative-perception-MRCNet.
Shixin Hong, Yu Liu 0005, Zhi Li 0057, You He 0002
CVPR5
2024 Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters
abstract
Continual learning can empower vision-language models to continuously acquire new knowledge, without the need for access to the entire historical dataset. However, mitigating the performance degradation in large-scale models is non-trivial due to (i) parameter shifts throughout life-long learning and (ii) significant computational burdens associated with full-model tuning. In this work, we present a parameter-efficient continual learning framework to alleviate long-term forgetting in incremental learning with vision-language models. Our approach involves the dynamic expansion of a pre-trained CLIP model, through the integration of Mixture-of-Experts (MoE) adapters in response to new tasks. To preserve the zero-shot recognition capability of vision-language models, we further introduce a Distribution Discriminative Auto-Selector (DDAS) that automatically routes in-distribution and out-of-distribution inputs to the MoE Adapter and the original CLIP, respectively. Through extensive experiments across various settings, our proposed method consistently outperforms previous state-of-the-art approaches while concurrently reducing parameter training burdens by 60%. Our code locates at https://github.com/JiazuoYu/MoE-Adapters4CL
Jiazuo Yu 0001, Yunzhi Zhuge, Lu Zhang 0053, Ping Hu 0001, Dong Wang 0004, Huchuan Lu, You He 0002
CVPR7
2024 ReDiffuser: Reliable Decision-Making Using a Diffuser with Confidence Estimation
abstract
The diffusion model has demonstrated impressive performance in offline reinforcement learning. However, non-deterministic sampling in diffusion models can lead to unstable performance. Furthermore, the lack of confidence measurements makes it difficult to evaluate the reliability and trustworthiness of the sampled decisions. To address these issues, we present ReDiffuser, which utilizes confidence estimation to ensure reliable decision-making. We achieve this by learning a confidence function based on Random Network Distillation. The confidence function measures the reliability of sampled decisions and contributes to quantitative recognition of reliable decisions. Additionally, we integrate the confidence function into task-specific sampling procedures to realize adaptive-horizon planning and value-embedded planning. Experiments show that the proposed ReDiffuser achieves state-of-the-art performance on standard offline RL datasets.
Nantian He, Zhi Li 0057, Yu Liu 0005, You He 0002
ICML5
2024 MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models
abstract
Existing pretrained text-to-video (T2V) models have demonstrated impressive abilities in generating realistic videos with basic motion or camera movement. However, these models exhibit significant limitations when generating intricate, human-centric motions. Current efforts primarily focus on fine-tuning models on a small set of videos containing a specific motion. They often fail to effectively decouple motion and the appearance in the limited reference videos, thereby weakening the modeling capability of motion patterns. To this end, we propose MoTrans, a customized motion transfer method enabling video generation of similar motion in new context. Specifically, we introduce a multimodal large language model (MLLM)-based recaptioner to expand the initial prompt to focus more on appearance and an appearance injection module to adapt appearance prior from video frames to the motion modeling process. These complementary multimodal representations from recaptioned prompt and video frames promote the modeling of appearance and facilitate the decoupling of appearance and motion. In addition, we devise a motion-specific embedding for further enhancing the modeling of the specific motion. Experimental results demonstrate that our method effectively learns specific motion pattern from singular or multiple reference videos, performing favorably against existing methods in customized video generation.
Xiaomin Li 0001, Xu Jia 0012, Haiwen Diao, Mengmeng Ge 0002, You He 0002, Huchuan Lu
ACM Multimedia7
2024 LLMs Can Evolve Continually on Modality for X-Modal Reasoning
Jiazuo Yu 0001, Haomiao Xiong, Lu Zhang 0053, Haiwen Diao, Yunzhi Zhuge, Lanqing Hong, Dong Wang 0004, Huchuan Lu, You He 0002, Long Chen 0016
NeurIPS9
2024 Learning depth-aware decomposition for single image dehazing
Yumeng Kang, Lu Zhang 0053, Ping Hu 0001, Yu Liu 0005, Huchuan Lu, You He 0002
Comput. Vis. Image Underst.6
2024 Efficient Adaptive Feature Fusion Network for Remote-Sensing Image Super-Resolution
abstract
Image super-resolution is a fundamental low-level vision task aimed at recovering high-resolution images with fine details. Deep learning has significantly enhanced the performance of super-resolution techniques for remote sensing imagery. However, increasing the depth of networks and the size of their parameters has resulted in substantial computational and storage burdens. To address this challenge, we propose an adaptive approach that learns both local and global information for each region. We introduce a lightweight hybrid model named the Efficient Adaptive Feature Fusion Network, which combines CNNs and Transformers to fully exploit the texture information in remote sensing images. This model leverages local details and long-range dependencies within images in an adaptive manner to achieve superior super-resolution. Specifically, a set of Transformers is employed to model the self-similarity between pixels and perform dense texture pattern predictions at each pixel, while a set of CNNs captures local details within the images. The computed global and local features serve as inputs to the proposed Adaptive Contextual Fusion Block, which learns to fuse local and global information across different regions to generate robust image super-resolution features. We conduct extensive experimental evaluations of the proposed method on the UCMerced and AID datasets, demonstrating its outstanding performance in terms of PSNR and SSIM metrics. Comprehensive experiments validate the effectiveness of our approach, showing that the proposed method achieves an excellent balance between performance and complexity.
Shuai Hao 0007, Shuai Liu 0009, Xu Jia 0012, Huchuan Lu, You He 0002
IEEE Signal Process. Lett.5
2024 Center-Wise Feature Consistency Learning for Long-Tailed Remote Sensing Object Recognition
abstract
Long-tailed distribution of remote sensing data generally limits the object recognition performance of deep neural networks. We notice that too many samples from head class will induce the neural network to learn features of tail class samples being biased towards the head. To solve this, we propose a novel center-wise feature consistency learning (CFCL) mechanism for long-tailed remote sensing object recognition. Firstly, we implement a head-tail center feature generation procedure that builds two teacher models to extract the knowledge from the head class and tail class samples respectively, so as to avoid the extracted tail class features being affected by the head classes. Secondly, a center-wise feature consistency learning strategy is introduced, which distills the central feature of each class to a student model, thereby making the classification boundaries more prominent. Especially, the central feature is estimated by referring to the features which are correctly classified by the teacher models, thus the inaccurate knowledge is abandoned. Extensive experiments on widely-adopted remote sensing recognition datasets including FGSC-23, DIOR, xView and HRSC2016 demonstrate that our method achieves superior performance compared to the state-of-the-art approaches.Code and data are available at: https://github.com/wdzhao123/CWFC.
Wenda Zhao 0003, Zhepu Zhang, Jiani Liu 0004, Yu Liu 0005, You He 0002, Huchuan Lu
IEEE Trans. Geosci. Remote. Sens.5
2024 Attacking Defocus Detection With Blur-Aware Transformation for Defocus Deblurring
abstract
Previous fully-supervised defocus deblurring has made significant progress. However, training such deep models requires abundant paired ground truth, which is expensive and error-prone. This paper makes an attempt to train a defocus deblurring model without using paired ground truth and any other unpaired data. Related reblur-to-deblur schemes generally use physics-based reblur or GAN-based reblur, suffering from the robustness of blur kernel and hallucination generated by GAN. Besides, the domain gap between the realistic blurred image and reblurred image hinders deblurring performance. Addressing these challenges, we propose a weakly-supervised defocus deblurring framework via defocus detection attack. On one hand, we build a focused area detection attack (FADA) to enforce the focused area to reblur, thereby reversing its detection result by a pretrained defocus blur detection network. Moreover, we introduce a blur-aware transfer modulated from the defocused region to help FADA render a robust reblurred region. On the other hand, we implement a defocused region detection attack to guide the realistic blurred region to deblur in the process of training deblurring network with simulated-paired areas. Extensive experiments on three widely-used datasets verify the effectiveness of our framework.
Wenda Zhao 0003, Fei Wei, You He 0002, Huchuan Lu
IEEE Trans. Multim.5
2024 Deformable Dynamic Sampling and Dynamic Predictable Mask Mining for Image Inpainting
abstract
Existing image inpainting methods often produce artifacts that are caused by using vanilla convolution layers as building blocks that treat all image regions equally and generate holes at random locations with equal probability. This design does not differentiate the missing regions and valid regions in inference and does not consider the predictability of missing regions in training. To address these issues, we propose a deformable dynamic sampling (DDS) mechanism which is built on deformable convolutions (DCs), and a constraint is proposed to avoid the deformably sampled elements falling into the corrupted regions. Furthermore, to select both valid sample locations and suitable kernels dynamically, we equip DCs with content-aware dynamic kernel selection (DKS). In addition, to further encourage the DDS mechanism to find meaningful sampling locations, we propose to train the inpainting model with mined predictable regions as holes. During training, we jointly train a mask generator with the inpainting network to generate hole masks dynamically for each training sample. Thus, the mask generator can find large yet predictable missing regions as a better alternative to random masks. Extensive experiments demonstrate the advantages of our method over state-of-the-art methods qualitatively and quantitatively.
Cai Cai, Yu Zeng 0001, Shu Yang 0004, Xu Jia 0012, Huchuan Lu, You He 0002
IEEE Trans. Neural Networks Learn. Syst.6
2024 Defocus Blur Detection Attack via Mutual-Referenced Feature Transfer
abstract
Benefiting from deep learning, defocus blur detection (DBD) has made prominent progress. Existing DBD methods generally study multiscale and multilevel features to improve performance. In this article, from a different perspective, we explore to generate confrontational images to attack DBD network. Based on the observation that defocus area and focus region in an image can provide mutual feature reference to help improve the quality of the confrontational image, we propose a novel mutual-referenced attack framework. Firstly, we design a divide-and-conquer perturbation image generation model, where the focus region attack image and defocus area attack image are generated respectively. Then, we integrate mutual-referenced feature transfer (MRFT) models to improve attack performance. Comprehensive experiments are provided to verify the effectiveness of our method. Moreover, related applications of our study are presented, e.g., sample augmentation to improve DBD and paired sample generation to boost defocus deblurring.
Wenda Zhao 0003, Fei Wei, You He 0002, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.5
2024 Confusion Region Mining for Crowd Counting
abstract
Existing works mainly focus on crowd and ignore the confusion regions which contain extremely similar appearance to crowd in the background, while crowd counting needs to face these two sides at the same time. To address this issue, we propose a novel end-to-end trainable confusion region discriminating and erasing network called CDENet. Specifically, CDENet is composed of two modules of confusion region mining module (CRM) and guided erasing module (GEM). CRM consists of basic density estimation (BDE) network, confusion region aware bridge and confusion region discriminating network. The BDE network first generates a primary density map, and then the confusion region aware bridge excavates the confusion regions by comparing the primary prediction result with the ground-truth density map. Finally, the confusion region discriminating network learns the difference of feature representations in confusion regions and crowds. Furthermore, GEM gives the refined density map by erasing the confusion regions. We evaluate the proposed method on four crowd counting benchmarks, including ShanghaiTech Part_A, ShanghaiTech Part_B, UCF_CC_50, and UCF-QNRF, and our CDENet achieves superior performance compared with the state-of-the-arts.
Jiawen Zhu 0003, Wenda Zhao 0003, Libo Yao, You He 0002, Maodi Hu, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.4
2023 Style-Content Metric Learning for Multidomain Remote Sensing Object Recognition
abstract
Previous remote sensing recognition approaches predominantly perform well on the training-testing dataset. However, due to large style discrepancies not only among multidomain datasets but also within a single domain, they suffer from obvious performance degradation when applied to unseen domains. In this paper, we propose a style-content metric learning framework to address the generalizable remote sensing object recognition issue. Specifically, we firstly design an inter-class dispersion metric to encourage the model to make decision based on content rather than the style, which is achieved by dispersing predictions generated from the contents of both positive sample and negative sample and the style of input image. Secondly, we propose an intra-class compactness metric to force the model to be less style-biased by compacting classifier's predictions from the content of input image and the styles of positive sample and negative sample. Lastly, we design an intra-class interaction metric to improve model's recognition accuracy by pulling in classifier's predictions obtained from the input image and positive sample. Extensive experiments on four datasets show that our style-content metric learning achieves superior generalization performance against the state-of-the-art competitors. Code and model are available at: https://github.com/wdzhao123/TSCM.
Wenda Zhao 0003, Ruikai Yang, Yu Liu 0005, You He 0002
AAAI4
2023 MetaFusion: Infrared and Visible Image Fusion via Meta-Feature Embedding from Object Detection
abstract
Fusing infrared and visible images can provide more texture details for subsequent object detection task. Conversely, detection task furnishes object semantic information to improve the infrared and visible image fusion. Thus, a joint fusion and detection learning to use their mutual promotion is attracting more attention. However, the feature gap between these two different-level tasks hinders the progress. Addressing this issue, this paper proposes an infrared and visible image fusion via meta-feature embedding from object detection. The core idea is that meta-feature embedding model is designed to generate object semantic features according to fusion network ability, and thus the semantic features are naturally compatible with fusion features. It is optimized by simulating a meta learning. Moreover, we further implement a mutual promotion learning between fusion and detection tasks to improve their performances. Comprehensive experiments on three public datasets demonstrate the effectiveness of our method. Code and model are available at: https://github.com/wdzhao123/MetaFusion.
Wenda Zhao 0003, Shigeng Xie, Fan Zhao 0005, You He 0002, Huchuan Lu
CVPR4
2023 Image Super-Resolution with Implicit Texture Pattern Modulation
abstract
Image super-resolution is one of the classical low-level vision tasks with the purpose of restoring a high-resolution image with fine details. Being aware of texture patterns with an image would benefit super-resolution performance a lot. However, it would be difficult to predict texture patterns for each region because of lack of annotations with different kinds of categories. In this work, we propose to implicitly model texture information with each region and take that as prior to promote super-resolution performance. In order to fully explore texture patterns, a hybrid model of convolutional neural networks and transformers is proposed. It is able to take advantage of both local and long range dependencies within an image to super-resolve an image. Specifically, a set of transformers are employed to model self-similarity among pixels within an image and to make dense texture patterns prediction at each pixel. The computed texture pattern representation then works as condition to modulate convolution-based residual blocks. In this way texture patterns could be integrated into the CNNs to obtain powerful features for image super-resolution. Extensive experiments on several benchmark datasets demonstrate its favorable performance against state-of-the-art methods and show its potential as a generic design for hybrid of transformers and CNNs.
Shuai Hao 0007, Xu Jia 0012, You He 0002, Huchuan Lu
ICME4
2023 Video Diffusion Models with Local-Global Context Guidance
abstract
Diffusion models have emerged as a powerful paradigm in video synthesis tasks including prediction, generation, and interpolation. Due to the limitation of the computational budget, existing methods usually implement conditional diffusion models with an autoregressive inference pipeline, in which the future fragment is predicted based on the distribution of adjacent past frames. However, only the conditions from a few previous frames can't capture the global temporal coherence, leading to inconsistent or even outrageous results in long-term video prediction. In this paper, we propose a Local-Global Context guided Video Diffusion model (LGC-VD) to capture multi-perception conditions for producing high-quality videos in both conditional/unconditional settings. In LGC-VD, the UNet is implemented with stacked residual blocks with self-attention units, avoiding the undesirable computational cost in 3D Conv. We construct a local-global context guidance strategy to capture the multi-perceptual embedding of the past fragment to boost the consistency of future prediction. Furthermore, we propose a two-stage training strategy to alleviate the effect of noisy frames for more stable predictions. Our experiments demonstrate that the proposed method achieves favorable performance on video prediction, interpolation, and unconditional video generation. We release code at https://github.com/exisas/LGC-VD.
Lu Zhang 0053, Yu Liu 0005, Zhizhuo Jiang, You He 0002
IJCAI5
2023 A Simple Baseline for Open-World Tracking via Self-training
abstract
Open-World Tracking (OWT) presents a challenging yet emerging problem, aiming to track every object of any category. Different from traditional Multi-Object Tracking (MOT), OWT needs to additionally track targets beyond predefined categories in the training set. To address the problem, we propose a simple baseline, SimOWT. We simplify the recently proposed OWT algorithm by streamlining the association module and accelerating the inference speed. By leveraging the self-training paradigm, SimOWT can distinguish unknown-class targets from the background, fully unleashing the potential of TAO-OW dataset. Furthermore, we enhance SimOWT from the perspectives of Pseudo Boxes Merging and Re-Weighting, thereby discovering more targets belonging to unknown classes and reducing the sensitivity of the model to low-quality pseudo-labels. Benefiting from the proposed approaches, SimOWT demonstrates a significant improvement in tracking performance on unknown classes. Moreover, the comprehensive experiments on the TAO-OW benchmark demonstrate that our model outperforms the state-of-the-art OWT method, OWTB, with an absolute gain of 11.2% OWTA and 16.4% detection recall respectively on unknown classes. The code is released at https://github.com/22109095/SimOWT.
Bingyang Wang, Tanlin Li, Jiannan Wu, Yi Jiang 0009, Huchuan Lu, You He 0002
ACM Multimedia6
2023 Persymmetric adaptive detection of range-spread targets in subspace interference plus Gaussian clutter
Tao Jian, Yu Liu 0005, You He 0002, Cong'an Xu, Zikeng Xie
Sci. China Inf. Sci.4
2023 Frequency-Adaptive Learning for SAR Ship Detection in Clutter Scenes
abstract
Convolutional neural networks (CNNs) have been widely applied in the context of ship detection in synthetic aperture radar (SAR) images, but the detection performance is still not ideal in scenarios with clutter interference. Mining frequency-domain information to suppress the sea clutter in SAR ship detection has attracted wide attention. However, existing frequency-domain ship detection methods do not process frequency-domain information adaptively, which results in the degradation of ship detection performance. To overcome this problem, this article proposes a novel deep learning network called YOLO-FA. YOLO-FA contains the proposed frequency attention module (FAM), which can process frequency-domain information of SAR images adaptively. The proposed method can suppress the sea clutter in the SAR images with the help of frequency-domain information. We evaluate the proposed method YOLO-FA on two datasets, i.e., the high-resolution SAR images’ dataset (HRSID) and SAR ship detection dataset (SSDD). Compared with the baseline method YOLOv5 and the existing commonly used methods, YOLO-FA achieves state-of-the-art detection performance on both the datasets.
Linping Zhang, Yu Liu 0005, Wenda Zhao 0003, Xueqian Wang 0002, Gang Li 0008, You He 0002
IEEE Trans. Geosci. Remote. Sens.6
2023 Weakly Correlated Distillation for Remote Sensing Object Recognition
abstract
Remote sensing object labels require high specialization, resulting in a limited number of labeled samples. Without large labeled samples to support training, general remote sensing object recognition models have limited accuracy. Addressing this issue, this paper proposes a weakly correlated distillation learning framework for remote sensing object recognition with small number of samples. Benefitting from large-scale natural image datasets, many recognition models achieve superior feature extraction capabilities. Thus, we use them as backbones to build teacher models, and then fine-tune the teacher models with a small-scale remote sensing dataset. However, due to the limited number of remote sensing samples, the teacher models may produce noisy features that reduce the performance of the student model. Therefore, we propose a weakly correlated distillation method that selects the weakly correlated features from teacher models to distill the student. Since the weakly correlated features contain different noise distributions which can be mutually suppressed, thereby improving the performance of the student. Extensive experiments on three widely-used datasets of DOTA, HRRSD and NWPU VHR-10 demonstrate the superior performance of our method compared with the state of the arts. Code is available at: https://github.com/wdzhao123/WCD.
Wenda Zhao 0003, Xiangzhu Lv, Yu Liu 0005, You He 0002, Huchuan Lu
IEEE Trans. Geosci. Remote. Sens.5
2023 Nowhere to Disguise: Spot Camouflaged Objects via Saliency Attribute Transfer
abstract
Both salient object detection (SOD) and camouflaged object detection (COD) are typical object segmentation tasks. They are intuitively contradictory, but are intrinsically related. In this paper, we explore the relationship between SOD and COD, and then borrow successful SOD models to detect camouflaged objects to save the design cost of COD models. The core insight is that both SOD and COD leverage two aspects of information: object semantic representations for distinguishing object and background, and context attributes that decide object category. Specifically, we start by decoupling context attributes and object semantic representations from both SOD and COD datasets through designing a novel decoupling framework with triple measure constraints. Then, we transfer saliency context attributes to the camouflaged images through introducing an attribute transfer network. The generated weakly camouflaged images can bridge the context attribute gap between SOD and COD, thereby improving the SOD models' performances on COD datasets. Comprehensive experiments on three widely-used COD datasets verify the ability of the proposed method. Code and model are available at: https://github.com/wdzhao123/SAT.
Wenda Zhao 0003, Shigeng Xie, Fan Zhao 0005, You He 0002, Huchuan Lu
IEEE Trans. Image Process.4
2023 Full-Scene Defocus Blur Detection With DeFBD+ via Multi-Level Distillation Learning
abstract
Existing defocus blur detection (DBD) methods generally perform well on a single type of unfocused blur scene (e.g., foreground focus), thereby suffering from the performance degradation for the other types of unfocused blur scenes. In this paper, we present the first exploration on full-scene DBD, and propose a separate-and-combine framework to achieve excellent performance for diverse defocus blur scenes. We firstly structure full-scene DBD dataset (named as DeFBD+) through collecting more types of unfocused blur scenes (e.g., background focus, full focus and full out of focus) with pixel-level annotations. Then, to avoid performance degradation caused by mutual interference from local feature representation and global content perception, we implement a pixel-level DBD network and an image-level DBD classification network to learn these two abilities separately. After that, we propose an isomeric distillation mechanism to combine these two abilities. Extensive experiments show that the proposed approach achieves superior performance compared with state-of-the-art methods.
Wenda Zhao 0003, Fei Wei, You He 0002, Huchuan Lu
IEEE Trans. Multim.4
2022 Video Object Segmentation via Structural Feature Reconfiguration
Zhenyu Chen 0001, Ping Hu 0001, Lu Zhang 0053, Huchuan Lu, You He 0002, Maodi Hu
ACCV (7)5
2022 Multi-Object Tracking Meets Moving UAV
abstract
Multi-object tracking in unmanned aerial vehicle (UAV) videos is an important vision task and can be applied in a wide range of applications. However, conventional multi-object trackers do not work well on UAV videos due to the challenging factors of irregular motion caused by moving camera and view change in 3D directions. In this paper, we propose a UAVMOT network specially for multi-object tracking in UAV views. The UAVMOT introduces an ID feature update module to enhance the object's feature association. To better handle the complex motions under UAV views, we develop an adaptive motion filter module. In addition, a gradient balanced focal loss is used to tackle the imbalance categories and small objects detection problem. Experimental results on the VisDrone2019 and UAVDT datasets demonstrate that the proposed UAVMOT achieves considerable improvement against the state-of-the-art tracking methods on UAV videos.
Shuai Liu 0009, Xin Li 0034, Huchuan Lu, You He 0002
CVPR4
2022 United Defocus Blur Detection and Deblurring via Adversarial Promoting Learning
Wenda Zhao 0003, Fei Wei, You He 0002, Huchuan Lu
ECCV (30)3
2022 Prospects for multi-agent collaboration and gaming: challenge, technology, and application
abstract
In this study, we presented the prospects for multi-agent system research with a special focus on agent collaboration and gaming tasks. We briefly introduced some open issues and task challenges from three major perspectives: the multi-agent environment, collaboration, and gaming. Then we provided a related outlook for the technology directions that may create some research challenge insights. Finally, we discussed the outlook for the multi-agent collaboration and gaming application areas.
Yu Liu 0005, Zhi Li 0057, Zhizhuo Jiang, You He 0002
Frontiers Inf. Technol. Electron. Eng.4
2022 Robust STAP Detection Based on Volume Cross-Correlation Function in Heterogeneous Environments
abstract
The performance of moving target detection in heterogeneous environments with the traditional space-time adaptive processing (STAP) may degrade when the real clutter environments deviate from the prior assumption on the clutter distribution. In this letter, a new detector for STAP applications based on volume cross-correlation function (VCF), namely VCF-STAP, is proposed to achieve robust performance of moving target detection in heterogeneous environments. In the new VCF-STAP, the VCF is used to form a distance measure between the sample signal subspace and the target subspace without modeling the clutter distribution. Then, a new robust STAP detection statistic is constructed using this distance measure. Simulation and experimental results show that the proposed VCF-STAP achieves robust performance of moving target detection in heterogeneous environments, especially it achieves much superior detection performance compared with existing STAP methods when the real clutter environments do not satisfy their prior assumptions. Besides, it is also shown that VCF-STAP has the constant false alarm rate (CFAR) property.
Zhizhuo Jiang, You He 0002, Gang Li 0008, Xiao-Ping Zhang 0002
IEEE Geosci. Remote. Sens. Lett.2
2022 Center-Boundary Dual Attention for Oriented Object Detection in Remote Sensing Images
abstract
Recently, anchor-free object detectors have shown promising performance in oriented object detection on remote sensing images. However, the objects in remote sensing images always have large variations in arbitrary orientations, sizes, and aspect ratios, which makes the existing anchor-free methods hard to obtain satisfactory results. In this article, we propose a novel anchor-free detector, center-boundary dual attention (CBDA) network (CBDA-Net), for fast and accurate oriented object detection on remote sensing images. In CBDA-Net, we construct a CBDA module, which utilizes a dual attention mechanism to extract attention features on the center and boundary regions of objects. The CBDA module can learn more essential features for rotating objects and reduce the interference from complex background. Besides, to resolve the influence of object aspect ratio on angle errors, we propose an aspect ratio weighted angle loss (arwLoss), where diffident penalties are assigned on the angle loss based on the aspect ratios of objects. This loss construction is effective in improving the detection accuracy of oriented objects, especially for slender objects. We conduct extensive experiments on two publish benchmarks, i.e., DOTA and HRSC2016. The experimental results demonstrate that our CBDA-Net achieves favorable performance against other anchor-free state of the arts with a real-time speed of 50 FPS.
Shuai Liu 0009, Lu Zhang 0053, Huchuan Lu, You He 0002
IEEE Trans. Geosci. Remote. Sens.4
2022 Teaching Teachers First and Then Student: Hierarchical Distillation to Improve Long-Tailed Object Recognition in Aerial Images
abstract
Remote sensing data distribution generally exposes the long-tail characteristic. This will limit the object recognition performance of existing deep models when they are trained with such unbalanced data. In this paper, we propose a novel hierarchical distillation framework to address the long-tailed object recognition in aerial images. Firstly, we notice that not only student model should learn feature representations from teachers, but also teacher models should learn feature representations from each other. Therefore, we build hierarchical teacher-wise distillation to improve the feature representations of the teacher models trained with middle and tail data, which is achieved by distilling the feature representations of the teacher model trained with head data. Secondly, we notice that the feature representations of the middle and tail classes can not be effectively distilled from the teacher to the student, since too little middle and tail data can be used to learn. Thus, we propose self-calibrated sampling learning that enforces the student to strengthen the learning of the middle and tail data, thereby improving the student’ feature learning ability. Extensive experiments on two widely-used DOTA and FGSC-23 datasets demonstrate superior performance of the proposed method compared with state-of-the-art methods. Model and code are publicly available at: https://github.com/wdzhao123/T2FTS.
Wenda Zhao 0003, Jiani Liu 0004, Yu Liu 0005, Fan Zhao 0005, You He 0002, Huchuan Lu
IEEE Trans. Geosci. Remote. Sens.5
2022 Diversity Consistency Learning for Remote-Sensing Object Recognition With Limited Labels
abstract
Annotating remote sensing object recognition needs high professionalism, and thus limited labeled samples are available. Suffering from this, general remote sensing object recognition methods are facing low recognition accuracy. Addressing this issue, this paper proposes a diversity consistency learning for remote sensing object recognition with limited labels. Specifically, diversity generation model is designed as a teacher model to generate diverse results, which is trained with labeled samples. Then, round consistency distillation model is introduced to distill the knowledge of diverse pseudo labels to a student network, which is trained with unlabeled samples. Especially, diverse pseudo labels are generated by the well-trained diversity generation model, which can improve recognition accuracy since diverse pseudo label errors can cancel each other out. Extensive experiments on two widely-used datasets of FS23 and HRSC2016 demonstrate the superior performance of our method compared with the state of the arts.
Wenda Zhao 0003, Tingting Tong, Fan Zhao 0005, You He 0002, Huchuan Lu
IEEE Trans. Geosci. Remote. Sens.5
2022 Feature Balance for Fine-Grained Object Classification in Aerial Images
abstract
Fine-grained object classification (FGOC) focuses on identifying subcategories of objects, which is crucial in military and civilian. Existing FGOC methods primarily focus on high-resolution aerial images, limiting their application on low-resolution (LR) FGOC that is a more realistic setting, especially on resource-constrained satellite devices. It is more challenging to deal with LR FGOC since objects’ details are blurred or missing. Addressing this issue, we make the first attempt to explore LR FGOC and propose a novel pipeline based on two technical insights: 1) feature balance strategy discriminatively integrates super-resolution weak and strong detailed presentations into coarse features of LR aerial images, achieving a feature balance to avoid that the weak detailed presentations are inhibited by the strong ones and 2) iterative interaction mechanism alternately refines feature details of the discriminative ship regions and optimizes the performance of FGOC. Moreover, we build a low-resolution fine-grained object (LFS) dataset to promote further study and evaluation. Extensive experiments on the proposed LFS dataset and the other three object datasets of DOTA, FS23, and HRSC2016 demonstrate that our method outperforms state-of-the-art algorithms. Dataset and code are publicly available athttps://github.com/wdzhao123/FBNet.
Wenda Zhao 0003, Tingting Tong, Libo Yao, Yu Liu 0005, Cong'an Xu, You He 0002, Huchuan Lu
IEEE Trans. Geosci. Remote. Sens.6
2022 A Novel Smooth Variable Structure Filter for Target Tracking Under Model Uncertainty
abstract
Model uncertainty is a serious challenge for robustness of tracking algorithms in radar systems. The smooth variable structure filter (SVSF) achieves error-bounded estimations for target state by scaling the magnitude of kinematic modeling error and accordingly performing a flexible switching strategy for the correction gain. However, the SVSF, without any smoothing functions, suffers from undesired chattering phenomenon since the measurement noise causes random disturbance to the identification of actual level of uncertainties, leading to obvious deterioration of tracking accuracy. In this paper, we present a new switching function for SVSF, i.e. the hyperbolic tangent function, for effective chattering suppression. Then we propose a new algorithm named as the Tanh-SVSF, which reformulates the correction gain with the new switching function, to improve the estimation accuracy for target state. A mathematical definition of SVSF chattering is proposed to quantify the chattering amplitude. It is demonstrated that the new switching function exerts a nonlinear compressing effect on the likelihood of measurement innovation and substantially reduces the disturbance of measurement noise, leading to elimination of the chattering problem. The stability of the Tanh-SVSF is analyzed, based on a proposed stability theorem and the numerical exhaustion strategy. Finally, the proposed method is tested on a simulated vehicle tracking scenario and real-world radar data from the Oxford Radar RobotCar Dataset, and shows superior performance over existing SVSF formulations and the Kalman filter, in view of tracking accuracy, track continuity and the proposed chattering indicator.
Yaowen Li, Gang Li 0008, Yu Liu 0005, Xiao-Ping Zhang 0002, You He 0002
IEEE Trans. Intell. Transp. Syst.5
2021 Polar Ray: A Single-stage Angle-free Detector for Oriented Object Detection in Aerial Images
abstract
Oriented bounding boxes are widely used for object detection in aerial images. Existing oriented object detection methods typically follow the general object detection paradigm by adding an extra rotation angle on the horizontal bounding boxes. However, the angular periodicity incurs the difficulty in angle regression and rotation sensitivity on bounding boxes. In this paper, we propose a new anchor-free oriented object detector, Polar Ray Network (PRNet), where object keypoints are represented by polar coordinates without angle regression. Our PRNet learns a set of polar rays from the object center to boundary with predefined equal-distributed angles. We introduce a dynamic PointConv module to optimize the regression of polar ray by incorporating object corner features. Furthermore, a classification feature guidance module is presented to improve the classification accuracy by incorporating more spatial contents from polar rays. Experimental results on two public datasets, i.e., DOTA and HRSC2016, demonstrate that the proposed PRNet significantly outperforms existing anchor-free detectors, and shows highly competitiveness with the state-of-the-art two-stage anchor-based methods.
Shuai Liu 0009, Lu Zhang 0053, Shuai Hao 0007, Huchuan Lu, You He 0002
ACM Multimedia5
2021 Defocus Blur Detection via Boosting Diversity of Deep Ensemble Networks
abstract
Existing defocus blur detection (DBD) methods usually explore multi-scale and multi-level features to improve performance. However, defocus blur regions normally have incomplete semantic information, which will reduce DBD's performance if it can't be used properly. In this paper, we address the above problem by exploring deep ensemble networks, where we boost diversity of defocus blur detectors to force the network to generate diverse results that some rely more on high-level semantic information while some ones rely more on low-level information. Then, diverse result ensemble makes detection errors cancel out each other. Specifically, we propose two deep ensemble networks (e.g., adaptive ensemble network (AENet) and encoder-feature ensemble network (EFENet)), which focus on boosting diversity while costing less computation. AENet constructs different light-weight sequential adapters for one backbone network to generate diverse results without introducing too many parameters and computation. AENet is optimized only by the self- negative correlation loss. On the other hand, we propose EFENet by exploring the diversity of multiple encoded features and ensemble strategies of features (e.g., group-channel uniformly weighted average ensemble and self-gate weighted ensemble). Diversity is represented by encoded features with less parameters, and a simple mean squared error loss can achieve the superior performance. Experimental results demonstrate the superiority over the state-of-the-arts in terms of accuracy and speed. Codes and models are available at: https://github.com/wdzhao123/DENets.
Wenda Zhao 0003, Xueqing Hou, You He 0002, Huchuan Lu
IEEE Trans. Image Process.3
2020 Unsupervised Video Object Segmentation with Joint Hotspot Tracking
Lu Zhang 0053, Jianming Zhang 0001, Zhe Lin 0001, Radomír Mech, Huchuan Lu, You He 0002
ECCV (14)6
2020 Fully distributed variational Bayesian non-linear filter with unknown measurement noise in sensor networks
Yu Liu 0005, Jun Liu 0040, Cong'an Xu, Gang Li 0008, You He 0002
Sci. China Inf. Sci.5
2020 Segmentation based rotated bounding boxes prediction and image synthesizing for object detection of high resolution aerial images
Lijun Wang 0001, Huchuan Lu, You He 0002
Neurocomputing4
2020 Towards Weakly-Supervised Focus Region Detection via Recurrent Constraint Network
abstract
Recent state-of-the-art methods on focus region detection (FRD) rely on deep convolutional networks trained with costly pixel-level annotations. In this study, we propose a FRD method that achieves competitive accuracies but only uses easily obtained bounding box annotations. Box-level tags provide important cues of focus regions but lose the boundary delineation of the transition area. A recurrent constraint network (RCN) is introduced for this challenge. In our static training, RCN is jointly trained with a fully convolutional network (FCN) through box-level supervision. The RCN can generate a detailed focus map to locate the boundary of the transition area effectively. In our dynamic training, we iterate between fine-tuning FCN and RCN with the generated pixel-level tags and generate finer new pixel-level tags. To boost the performance further, a guided conditional random field is developed to improve the quality of the generated pixel-level tags. To promote further study of the weakly supervised FRD methods, we construct a new dataset called FocusBox, which consists of 5000 challenging images with bounding box-level labels. Experimental results on existing datasets demonstrate that our method not only yields comparable results than fully supervised counterparts but also achieves a faster speed.
Wenda Zhao 0003, Xueqing Hou, You He 0002, Huchuan Lu
IEEE Trans. Image Process.4
2019 ROI Pooled Correlation Filters for Visual Tracking
abstract
The ROI (region-of-interest) based pooling method performs pooling operations on the cropped ROI regions for various samples and has shown great success in the object detection methods. It compresses the model size while preserving the localization accuracy, thus it is useful in the visual tracking field. Though being effective, the ROI-based pooling operation is not yet considered in the correlation filter formula. In this paper, we propose a novel ROI pooled correlation filter (RPCF) algorithm for robust visual tracking. Through mathematical derivations, we show that the ROI-based pooling can be equivalently achieved by enforcing additional constraints on the learned filter weights, which makes the ROI-based pooling feasible on the virtual circular samples. Besides, we develop an efficient joint training formula for the proposed correlation filter algorithm, and derive the Fourier solvers for efficient model training. Finally, we evaluate our RPCF tracker on OTB-2013, OTB-2015 and VOT-2017 benchmark datasets. Experimental results show that our tracker performs favourably against other state-of-the-art trackers.
Yuxuan Sun 0003, Dong Wang 0004, You He 0002, Huchuan Lu
CVPR4
2019 CapSal: Leveraging Captioning to Boost Semantics for Salient Object Detection
abstract
Detecting salient objects in cluttered scenes is a big challenge. To address this problem, we argue that the model needs to learn discriminative semantic features for salient objects. To this end, we propose to leverage captioning as an auxiliary semantic task to boost salient object detection in complex scenarios. Specifically, we develop a CapSal model which consists of two sub-networks, the Image Captioning Network (ICN) and the Local-Global Perception Network (LGPN). ICN encodes the embedding of a generated caption to capture the semantic information of major objects in the scene, while LGPN incorporates the captioning embedding with local-global visual contexts for predicting the saliency map. ICN and LGPN are jointly trained to model high-level semantics as well as visual saliency. Extensive experiments demonstrate the effectiveness of image captioning in boosting the performance of salient object detection. In particular, our model performs significantly better than the state-of-the-art methods on several challenging datasets of complex scenarios.
Lu Zhang 0053, Jianming Zhang 0001, Zhe Lin 0001, Huchuan Lu, You He 0002
CVPR5
2019 Fast Video Object Segmentation via Dynamic Targeting Network
abstract
We propose a new model for fast and accurate video object segmentation. It consists of two convolutional neural networks, a Dynamic Targeting Network (DTN) and a Mask Refinement Network (MRN). DTN locates the object by dynamically focusing on regions of interest surrounding the target object. The target region is predicted by DTN via two sub-streams, Box Propagation (BP) and Box Re-identification (BR). The BP stream is faster but less effective at objects with large deformation or occlusion. The BR stream performs better in difficult scenarios at a higher computation cost. We propose a Decision Module (DM) to adaptively determine which sub-stream to use for each frame. Finally, MRN is exploited to predict segmentation within the target region. Experimental results on two public datasets demonstrate that the proposed model significantly outperforms existing methods without online training in both accuracy and efficiency, and is comparable to online training-based methods in accuracy with an order of magnitude faster speed.
Lu Zhang 0053, Zhe Lin 0001, Jianming Zhang 0001, Huchuan Lu, You He 0002
ICCV5
2018 A Bi-Directional Message Passing Model for Salient Object Detection
abstract
Recent progress on salient object detection is beneficial from Fully Convolutional Neural Network (FCN). The saliency cues contained in multi-level convolutional features are complementary for detecting salient objects. How to integrate multi-level features becomes an open problem in saliency detection. In this paper, we propose a novel bi-directional message passing model to integrate multi-level features for salient object detection. At first, we adopt a Multi-scale Context-aware Feature Extraction Module (MCFEM) for multi-level feature maps to capture rich context information. Then a bi-directional structure is designed to pass messages between multi-level features, and a gate function is exploited to control the message passing rate. We use the features after message passing, which simultaneously encode semantic information and spatial details, to predict saliency maps. Finally, the predicted results are efficiently combined to generate the final saliency map. Quantitative and qualitative experiments on five benchmark datasets demonstrate that our proposed model performs favorably against the state-of-the-art methods under different evaluation metrics.
Lu Zhang 0053, Ju Dai, Huchuan Lu, You He 0002, Gang Wang 0012
CVPR4