Song Tang 0001

dblp:181/8826-1 · DBLP profile ↗
← Back
34ranked-venue papers
11as first author
25since 2021 · last 2026
0000-0003-2635-1872ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 8 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 10 since 2021Systems, architecture and hardware · 5 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Progressive category prototype optimization for black-box domain adaptation
Lihua Zhou, Song Tang 0001, Yan Gan, Mao Ye 0001
Neurocomputing3
2026 Source-Free Domain Adaptive Object Detection with semantics compensation
Song Tang 0001, Jiuzheng Yang, Mao Ye 0001, Yan Gan, Xiatian Zhu
Pattern Recognit.1
2026 From Point to Flow: Enhancing Unsupervised Domain Adaptation With Flow Classification
abstract
Unsupervised domain adaptation aims to transfer knowledge from a labeled source domain to an unlabeled target domain. Existing methods, whether based on distribution matching or self-supervised learning, often focus solely on classifying individual source samples, potentially overlooking discriminative information. To address this limitation, we propose FlowUDA, a novel plugin method that enhances existing UDA frameworks by constructing semantically invariant flows from individual source samples to corresponding target samples, forming cross-domain trajectories. By leveraging a diffusion network guided by ordinary differential equations, FlowUDA ensures these flows preserve the topological structure of the source domain, maintaining their distinguishability. Our method then classifies these flows by sampling points along them and transferring labels from source samples, effectively capturing spatial relationships between domains. In essence, FlowUDA transforms the traditional point-based classification on individual source samples into flow-based classification on flows, allowing the model to learn richer, more discriminative features that bridge the gap between source and target domains. Extensive experiments on standard benchmarks demonstrate that integrating FlowUDA into existing UDA methods leads to notable performance gains, highlighting its effectiveness in addressing domain shift challenges.
Lihua Zhou, Mao Ye 0001, Nianxin Li, Song Tang 0001, Xu-Qian Fan, Lei Deng 0001, Zhen Lei 0001, Xiatian Zhu
IEEE Trans. Circuits Syst. Video Technol.4
2026 Long-Short Match for Lost Control in UAV Multi-Object Tracking
abstract
Multi-Object Tracking (MOT) in Unmanned Aerial Vehicles (UAV) aims to continuously and stably detect and track objects in videos captured by UAVs. In existing MOT tracking-by-detection schemes, the tracker with a fixed step size is always employed, and a fixed length of past tracking information is input to the tracker to guide position prediction. However, the limited prediction range of a single-scale tracker leads to frequent tracking losses, and limited historical information also reduces tracking accuracy. To address these limitations, we propose a novel Long-Short Match (LSMTrack) tracking method. The key idea is to use long and short trackers and maintain a long-term motion state to improve tracking performance, thus reducing the likelihood of entering the lost status. To this end, a new Mamba-based tracker and a long-short match strategy are proposed. For long and short trackers, the same architecture is used based on Mamba. Unlike the previous Mamba-based approach, the proposed tracker maintains a long-term state while updating the state and making position predictions in each time step, so we call it a step Mamba tracker. Meanwhile, we devise a long-short match strategy at the inference stage to integrate long and short trackers, and design a lost control operation which updates the long-term states using historical state values. In this way, the matching probability and the inference efficiency are guaranteed. Experimental results on two UAV MOT datasets confirm the state-of-the-art performance. Specifically, the best results are achieved in terms of two popular MOTA and IDF1 tracking evaluation metrics.
Zi-Zhuang Zou, Mao Ye 0001, Luping Ji, Lihua Zhou, Song Tang 0001, Yan Gan, Shuai Li 0005
IEEE Trans. Multim.5
2025 Self-Prompting Analogical Reasoning for UAV Object Detection
abstract
Unmanned Aerial Vehicle Object Detection (UAVOD) presents unique challenges due to varying altitudes, dynamic backgrounds, and the small size of objects. Traditional detection methods often struggle with these challenges, as they typically rely on visual feature only and fail to extract the semantic relations between the objects. To address these limitations, we propose a novel approach named Self-Prompting Analogical Reasoning (SPAR). Our method utilizes the vision-language model (CLIP) to generate context-aware prompts based on image feature, providing rich semantic information that guides analogical reasoning. SPAR includes two main modules: self-prompting and analogical reasoning. Self-prompting module based on learnable description and CLIP-text encoder generates context-aware prompt by combining specific image feature; then an objectness prompt score map is produced by computing the similarity between pixel-level features and context-aware prompt. With this score map, multi-scale image features are enhanced and pixel-level features are chosen for graph construction. While for analogical reasoning module, graph nodes consists of category-level prompt nodes and pixel-level image feature nodes. Analogical inference is based graph convolution. Under the guidance of category-level nodes, different-scale object features have been enhanced, which helps achieve more accurate detection of challenging objects. Extensive experiments illustrate that SPAR outperforms traditional methods, offering a more robust and accurate solution for UAVOD.
Nianxin Li, Mao Ye 0001, Lihua Zhou, Song Tang 0001, Yan Gan, Zizhuo Liang, Xiatian Zhu
AAAI4
2025 Pseudo Visible Feature Fine-Grained Fusion for Thermal Object Detection
abstract
Thermal object detection is a critical task in various fields, such as surveillance and autonomous driving. Current state-of-the-art (SOTA) models always leverage a prior Thermal-To-Visible (T2V) translation model to obtain visible spectrum information, followed by a cross-modality aggregation module to fuse information from both modalities. However, this fusion approach does not fully exploit the complementary visible spectrum information beneficial for thermal detection. To address this issue, we propose a novel cross-modal fusion method called Pseudo Visible Feature Fine-Grained Fusion (PFGF). Specifically, a graph is constructed with nodes generated from multi-level thermal features and pseudo-visual latent features produced by the T2V model. Each level of features corresponds to a subgraph. An Inter-Mamba block is proposed to perform cross-modality fusion between nodes at the lowest level; while a Cascade Knowledge Integration (CKI) strategy is used to fuse low-level fused information to high-level subgraphs in a cascade manner. After several iterations of graph node updating, each subgraph outputs an aggregated feature to the detection head respectively. Unlike previous cross-modal fusion methods, our approach explicitly models high-level relationships between cross-modal data, effectively fusing different granularity information. Experimental results demonstrate that our method achieves SOTA detection performance. Code is available at https://github.com/liting1018/PFGF.
Mao Ye 0001, Tianwen Wu, Nianxin Li, Shuaifeng Li, Song Tang 0001, Luping Ji
CVPR6
2025 Domain Adaptive Diabetic Retinopathy Grading with Model Absence and Flowing Data
abstract
Domain shift (the difference between source and target domains) poses a significant challenge in clinical applications, e.g., Diabetic Retinopathy (DR) grading. Despite considering certain clinical requirements, like source data privacy, conventional transfer methods are predominantly model-centered and often struggle to prevent model-targeted attacks. In this paper, we address a challenging Online Model-aGnostic Domain Adaptation (OMG-DA) setting, driven by the demands of clinical environments. This setting is characterized by the absence of the model and the flow of target data. To tackle the new challenge, we propose a novel approach, Generative Unadversarial ExampleS (GUES), which enables adaptation from a data-centric perspective. Specifically, we first theoretically reformulate conventional perturbation optimization in a generative way—learning a perturbation generation function with a latent input variable. During model instantiation, we leverage a Variational AutoEncoder to express this function. The encoder with the reparameterization trick predicts the latent input, whilst the decoder is responsible for the generation. Furthermore, the saliency map is selected as pseudo-perturbation labels. Because it not only captures potential lesions but also theoretically provides an upper bound on the function input, enabling the identification of the latent variable. Extensive experiments on DR benchmarks with both frozen pre-trained models and trainable models demonstrate the superiority of GUES, showing robustness even with small batch size. The source code and data are available at https://github.com/tntek/GUES.
Wenxin Su, Song Tang 0001, Xiaojing Yi, Mao Ye 0001, Chunxiao Zu, Xiatian Zhu
CVPR2
2025 Proxy Denoising for Source-Free Domain Adaptation
abstract
Source-Free Domain Adaptation (SFDA) aims to adapt a pre-trained source model to an unlabeled target domain with no access to the source data. Inspired by the success of large Vision-Language (ViL) models in many applications, the latest research has validated ViL's benefit for SFDA by using their predictions as pseudo supervision. However, we observe that ViL's supervision could be noisy and inaccurate at an unknown rate, potentially introducing additional negative effects during adaption. To address this thus-far ignored challenge, we introduce a novel Proxy Denoising (__ProDe__) approach. The key idea is to leverage the ViL model as a proxy to facilitate the adaptation process towards the latent domain-invariant space. Concretely, we design a proxy denoising mechanism to correct ViL's predictions. This is grounded on a proxy confidence theory that models the dynamic effect of proxy's divergence against the domain-invariant space during adaptation. To capitalize the corrected proxy, we further derive a mutual knowledge distilling regularization. Extensive experiments show that ProDe significantly outperforms the current state-of-the-art alternatives under both conventional closed-set setting and the more challenging open-set, partial-set, generalized SFDA, multi-target, multi-source, and test-time settings. Our code and data are available at https://github.com/tntek/source-free-domain-adaptation.
Song Tang 0001, Wenxin Su, Yan Gan, Mao Ye 0001, Jianwei Zhang 0001, Xiatian Zhu
ICLR1
2025 Multimodal Causal Reasoning for UAV Object Detection
abstract
Unmanned Aerial Vehicle (UAV) object detection faces significant challenges due to complex environmental conditions and different imaging conditions. These factors introduce significant changes in scale and appearance, particularly for small objects that occupy limited pixels and exhibit limited information, complicating detection tasks. To address these challenges, we propose a Multimodel Causal Reasoning framework based on YOLO backbone for UAV Object Detection (MCR-UOD). The key idea is to use the backdoor adjustment to discover the condition-invariant object representation for easy detection. Specifically, the YOLO backbone is first adjusted to incorporate the pre-trained vision-language model. The original category labels are replaced with semantic text prompts, and the detection head is replaced with text-image contrastive learning. Based on this backbone, our method consists of two parts. The first part, named language guided region exploration, discovers the regions with high probability of object existence using text embeddings based on vision-language model such as CLIP. Another part is the backdoor adjustment casual reasoning module, which constructs a confounder dictionary tailored to different imaging conditions to capture global image semantics and derives a prior probability distribution of shooting conditions. During causal inference, we use the confounder dictionary and the prior to intervene on local instance features, disentangling condition variations, and obtaining condition-invariant representations. Experimental results on several public datasets confirm the state-of-the-art performance of our approach. The code, data and models will be released upon publication of this paper.
Nianxin Li, Mao Ye 0001, Lihua Zhou, Shuaifeng Li, Song Tang 0001, Luping Ji, Ce Zhu
NeurIPS5
2025 Few-shot medical image segmentation with high-fidelity prototypes
abstract
Few-shot Semantic Segmentation (FSS) aims to adapt a pretrained model to new classes with as few as a single labeled training sample per class. Despite the prototype based approaches have achieved substantial success, existing models are limited to the imaging scenarios with considerably distinct objects and not highly complex background, e.g., natural images. This makes such models suboptimal for medical imaging with both conditions invalid. To address this problem, we propose a novel D etail S elf-refined P rototype Net work ( DSPNet ) to construct high-fidelity prototypes representing the object foreground and the background more comprehensively. Specifically, to construct global semantics while maintaining the captured detail semantics, we learn the foreground prototypes by modeling the multimodal structures with clustering and then fusing each in a channel-wise manner. Considering that the background often has no apparent semantic relation in the spatial dimensions, we integrate channel-specific structural information under sparse channel-aware regulation. Extensive experiments on three challenging medical image benchmarks show the superiority of DSPNet over previous state-of-the-art methods. The code and data are available at https://github.com/tntek/DSPNet . • A novel prototypical FSS approach DSPNet that enhances prototypes’ self-representation. • A class prototype self-refining method FSPA integrating the cluster prototypes. • A background prototype self-refining method BCMA coding channel-specific structure.
Song Tang 0001, Shaxu Yan, Xiaozhi Qi, Mao Ye 0001, Jianwei Zhang 0001, Xiatian Zhu
Medical Image Anal.1
2025 Learning economically for Chinese word segmentation: tuning pretrained model via active learning and N-gram preference
Zhiyuan Ma 0001, Jiwei Qin, Song Tang 0001, Jinpeng Mi
Neural Comput. Appl.3
2025 Adaptive Surveillance Video Compression With Background Hyperprior
abstract
Neural surveillance video compression methods have demonstrated significant improvements over traditional video compression techniques. In current surveillance video compression frameworks, the first frame in a Group of Pictures (GOP) is usually compressed fully as an I frame, and the subsequent P frames are compressed by referencing this I frame at Low Delay P (LDP) encoding mode. However, this compression approach overlooks the utilization of background information, which limits its adaptability to different scenarios. In this paper, we propose a novel Adaptive Surveillance Video Compression framework based on background hyperprior, dubbed as ASVC. This background hyperprior is related with side information to assist in coding both the temporal and spatial domains. Our method mainly consists of two components. First, the background information from a GOP is extracted, modeled as hyperprior and is compressed by exiting methods. Then these hyperprior is used as side information to compress both I frames and P frames. ASVC effectively captures the temporal dependencies in the latent representations of surveillance videos by leveraging background hyperprior for auxiliary video encoding. The experimental results demonstrate that applying ASVC to traditional and learning based methods significantly improves performance.
Song Tang 0001, Mao Ye 0001
IEEE Signal Process. Lett.2
2025 Intelligent Event-Triggered H∞ Load Frequency Control for Power Systems With Multiple-Resource Delays
abstract
This paper considers intelligent event-triggered${H}_{\infty }$Load Frequency Control (LFC) for a new kind of multi-area power system. Firstly, delay factors are introduced into the prime mover, engine, and energy storage unit, i.e, multiple-resource delays, and the impact of these delays on the dynamic behavior of system is discussed for the first time, making the model more realistic and accurate. Secondly, a new lemma is proposed, which introduces free variables, extending to the scenarios with time delays. Combining this lemma with the event-triggered mechanism, the control scheme is designed, which is featured by two characteristics. On the one hand, in the Lyapunov functional, looped terms on neighbor triggering instants are considered, reduce the conservatism of the obtained criteria by capturing the information of triggered instants. Meanwhile, a mandatory triggering mechanism is introduced to break the possible dead loop caused by long-term no-triggering. On the other hand, a genetic algorithm (GA) is employed to optimize the parameters of the trigger threshold, encouraging efficient control within the context of LFC and delays. The effectiveness of the proposed approach is confirmed by a typical numerical example with four case studies.
Jinnan Luo, Kaibo Shi, Song Tang 0001, Ruimei Zhang, Jun Cheng 0004, Ju H. Park 0001
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 LESEP: Boosting Adversarial Transferability via Latent Encoding and Semantic Embedding Perturbations
abstract
Transferability and imperceptibility of adversarial examples are pivotal for assessing the efficacy of black-box attacks. While diffusion models have been employed to generate adversarial examples, leveraging their advanced image generation capability to enhance transferability and imperceptibility, these methods typically focus only on perturbing the image or latent space. They often ignore the critical role of semantic information in the denoising process, thereby impeding the improvement of the transferability of adversarial examples. Furthermore, the modification of high-level semantics inevitably introduces image blurring. This degradation in visual quality makes the adversarial examples more susceptible to detection. To overcome the above limitations, we are the first to utilize image latent encoding and semantic embedding perturbations to enhance the performance of adversarial attacks. Then, the LESEP method is proposed. In the LESEP framework, we first apply image latent encoding attack to achieve deception of the target model. Second, the semantic embedding attack enhances the transferability of adversarial examples. Additionally, we utilize the image restoration technique to guarantee the high imperceptibility of the crafted adversarial examples. Through comprehensive experiments on diverse datasets, different network architectures and defense methods, we have demonstrated that the LESEP method achieves outstanding transferability and imperceptibility while displaying strong robustness.
Yan Gan, Chengqian Wu, Deqiang Ouyang, Song Tang 0001, Mao Ye 0001, Tao Xiang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Source-Free Domain Adaptation with Frozen Multimodal Foundation Model
abstract
Source-Free Domain Adaptation (SFDA) aims to adapt a source model for a target domain, with only access to unlabeled target training data and the source model pretrained on a supervised source domain. Relying on pseudo labeling and/or auxiliary supervision, conventional methods are inevitably error-prone. To mitigate this limitation, in this work we for the first time explore the potentials of off-the-shelf vision-language (ViL) multimodal models (e.g., CLIP) with rich whilst heterogeneous knowledge. We find that directly applying the ViL model to the target domain in a zero-shot fashion is unsatisfactory, as it is not specialized for this particular task but largely generic. To make it task specific, we propose a novel Distilling multImodal Foundation mOdel (DIFO) approach. Specifically, DIFO alternates between two steps during adaptation: (i) Customizing the ViL model by maximizing the mutual information with the target model in a prompt learning manner, (ii) Distilling the knowledge of this customized ViL model to the target model. For more fine-grained and reliable distillation, we further introduce two effective regularization terms, namely most-likely category encouragement and predictive consistency. Extensive experiments show that DIFO significantly outperforms the state-of-the-art alternatives. Code is here.
Song Tang 0001, Wenxin Su, Mao Ye 0001, Xiatian Zhu
CVPR1
2024 Adversarial Experts Model for Black-box Domain Adaptation
abstract
Black-box domain adaptation treats the source domain model as a black box. During the transfer process, the only available information about the target domain is the noisy labels output by the black-box model. This poses significant challenges for domain adaptation. Conventional approaches typically tackle the black-box noisy label problem from two aspects: self-knowledge distillation and pseudo-label denoising, both achieving limited performance due to limited knowledge information. To mitigate this issue, we explore the potential of off-the-shelf vision-language (ViL) multimodal models with rich semantic information for black-box domain adaptation by introducing an Adversarial Experts Model (AEM). Specifically, our target domain model is designed as one feature extractor and two classifiers, trained over two stages: In the knowledge transferring stage, with a shared feature extractor, the black-box source model and the ViL model act as two distinct experts for joint knowledge contribution, guiding the learning of one classifier each. While contributing their respective knowledge, the experts are also updated due to their own limitation and bias. In the adversarial alignment stage, to further distill expert knowledge to the target domain model, adversarial learning is conducted between the feature extractor and the two classifiers. A new consistency-max loss function is proposed to measure two classifier consistency and further improve classifier prediction certainty. Extensive experiments on multiple datasets demonstrate the effectiveness of our approach. Code is available at https://github.com/singinger/AEM.
Siying Xiao, Mao Ye 0001, Qichen He, Shuaifeng Li, Song Tang 0001, Xiatian Zhu
ACM Multimedia5
2024 Cloud Object Detector Adaptation by Integrating Different Source Knowledge
abstract
We propose to explore an interesting and promising problem, Cloud Object Detector Adaptation (CODA), where the target domain leverages detections provided by a large cloud model to build a target detector. Despite with powerful generalization capability, the cloud model still cannot achieve error-free detection in a specific target domain. In this work, we present a novel Cloud Object detector adaptation method by Integrating different source kNowledge (COIN). The key idea is to incorporate a public vision-language model (CLIP) to distill positive knowledge while refining negative knowledge for adaptation by self-promotion gradient direction alignment. To that end, knowledge dissemination, separation, and distillation are carried out successively. Knowledge dissemination combines knowledge from cloud detector and CLIP model to initialize a target detector and a CLIP detector in target domain. By matching CLIP detector with the cloud detector, knowledge separation categorizes detections into three parts: consistent, inconsistent and private detections such that divide-and-conquer strategy can be used for knowledge distillation. Consistent and private detections are directly used to train target detector; while inconsistent detections are fused based on a consistent knowledge generation network, which is trained by aligning the gradient direction of inconsistent detections to that of consistent detections, because it provides a direction toward an optimal target detector. Experiment results demonstrate that the proposed COIN method achieves the state-of-the-art performance.
Shuaifeng Li, Mao Ye 0001, Lihua Zhou, Nianxin Li, Siying Xiao, Song Tang 0001, Xiatian Zhu
NeurIPS6
2024 Shooting condition insensitive unmanned aerial vehicle object detection
Jinzong Cui, Mao Ye 0001, Xiatian Zhu, Song Tang 0001
Expert Syst. Appl.5
2024 Source-Free Domain Adaptation via Target Prediction Distribution Searching
abstract
Abstract Existing Source-Free Domain Adaptation (SFDA) methods typically adopt the feature distribution alignment paradigm via mining auxiliary information (eg., pseudo-labelling, source domain data generation). However, they are largely limited due to that the auxiliary information is usually error-prone whilst lacking effective error-mitigation mechanisms. To overcome this fundamental limitation, in this paper we propose a novel Target Prediction Distribution Searching (TPDS) paradigm. Theoretically, we prove that in case of sufficient small distribution shift, the domain transfer error could be well bounded. To satisfy this condition, we introduce a flow of proxy distributions that facilitates the bridging of typically large distribution shift from the source domain to the target domain. This results in a progressive searching on the geodesic path where adjacent proxy distributions are regularized to have small shift so that the overall errors can be minimized. To account for the sequential correlation between proxy distributions, we develop a new pairwise alignment with category consistency algorithm for minimizing the adaptation errors. Specifically, a manifold geometry guided cross-distribution neighbour search is designed to detect the data pairs supporting the Wasserstein distance based shift measurement. Mutual information maximization is then adopted over these pairs for shift regularization. Extensive experiments on five challenging SFDA benchmarks show that our TPDS achieves new state-of-the-art performance. The code and datasets are available at https://github.com/tntek/TPDS .
Song Tang 0001, An Chang, Fabian Zhang, Xiatian Zhu, Mao Ye 0001, Changshui Zhang
Int. J. Comput. Vis.1
2024 Source-free domain adaptation with Class Prototype Discovery
Lihua Zhou, Nianxin Li, Mao Ye 0001, Xiatian Zhu, Song Tang 0001
Pattern Recognit.5
2024 Illumination Distribution-Aware Thermal Pedestrian Detection
abstract
Pedestrian detection is an important task in computer vision, which is also an important part of intelligent transportation systems. For privacy protection, thermal images are widely used in pedestrian detection problems. However, thermal pedestrian detection is challenging due to the significant effect of temperature variation on the illumination of images and that fine-grained illumination annotations are difficult to be acquired. The existing methods have attempted to exploit coarse-grained day/night labels, which however even hampers the model performance. In this work, we introduce a novel idea of regressing conditional thermal-visible feature distribution, dubbed as Illumination Distribution-Aware adaptation (IDA). The key idea is to predict the conditional visible feature distribution given a thermal image, subject to their pre-computed joint distribution. Specifically, we first estimate the thermal-visible feature joint distribution by constructing feature co-occurrence matrices, offering a conditional probability distribution for any given thermal image. With this pairing information, we then form a conditional probability distribution regression task for model optimization. Critically, as a model agnostic strategy, this allows the visible feature knowledge to be transferred to the thermal counterpart implicitly for learning more discriminating feature representation. Experiment results show that our method outperforms the prior art methods, which use extra illumination annotations. Besides, as a plug-in, our method can averagely reduce about 2% MR on KAIST dataset, and improve about 1% mAP on FLIR-aligned and Autonomous Vehicles datasets without extra calculation for test. Code is available athttps://github.com/HaMeow-lst1/IDA.
Mao Ye 0001, Luping Ji, Song Tang 0001, Yan Gan, Xiatian Zhu
IEEE Trans. Intell. Transp. Syst.4
2024 Progressive Source-Aware Transformer for Generalized Source-Free Domain Adaptation
abstract
Source-free domain adaptation (SFDA) tends to forget the source domain, suffering from limitations in real-world scenarios. Recently, generalized source-free domain adaptation (GSFDA) problem naturally emerges, aiming for good performance on both target and source domains. The existing methods attempt to retain model parameters associated with the source domain to prevent such forgetting. However, this strategy is not conducive to improving cross-domain performance on the target domain, prioritizing mitigating forgetting on the source domain. This article introduces a Progressive Source-Aware Transformer approach for GSFDA, dubbed PSAT-GDA. Our core idea is to enforce the domain adaptation process to remember the source domain by imposing source guidance, offering a target domain-centric anti-forgetting mechanism. Specifically, for each epoch, a Transformer-based deep network is adapted to do domain alignment like the traditional SFDA method, because the transformer working on the image patch sequence helps to reduce image noise caused by domain shift. Meanwhile, another Transformer is designed to generate source guidance supervising domain alignment. By augmenting target sample and mining the source information from the historical models before current epoch, source injected feature group is constructed. Based on the Transformer mechanism, the attention block can select useful source information for each target sample. From it, we devise neighbour-based and augmentation-based regularizations to shape the source guidance. Experiments on three challenging datasets show that our method can achieve evident cross-domain improvement on the target domains. Also, it can mitigate forgetting on all domains after adapting to single or multiple target domains.
Song Tang 0001, Yuji Shi, Mao Ye 0001, Changshui Zhang, Jianwei Zhang 0001
IEEE Trans. Multim.1
2023 Weakly Supervised Referring Expression Grounding via Target-Guided Knowledge Distillation
abstract
Weakly supervised referring expression grounding aims to train a model without the manual labels between image regions and referring expressions during the training phase. Current predominant models often adopt deep structures to reconstruct the region-expression correspondence. A crucial deficiency of the existing approaches lies in that these models neglect to exploit potential valuable information to further improve their grounding performance. To address this issue, we leverage knowledge distillation as a unique scheme to excavate and transfer helpful information for acquiring a better model. Specifically, we propose a target-guided knowledge distillation framework that accounts for region-expression pairs reconstruction and matching. We reactivate the target-related prediction information learned by a pre-trained teacher model and transfer the target-related prediction knowledge from the teacher to guide the training process and boost the performance of the student model. We conduct extensive experiments on three benchmark datasets, i.e., RefCOCO, RefCOCO+, and RefCOCOg. Without bells and whistles, our approach achieves state-of-the-art results on several splits of benchmark datasets. The implementation codes and trained models are available at: https://github.com/dami23/WREG_KD.
Jinpeng Mi, Song Tang 0001, Zhiyuan Ma 0001, Qingdu Li, Jianwei Zhang 0001
ICRA2
2022 Semantic consistency learning on manifold for source data-free unsupervised domain adaptation
Song Tang 0001, Yan Zou, Jianzhi Lyu, Mao Ye 0001, Shouming Zhong, Jianwei Zhang 0001
Neural Networks1
2021 Model Adaptation through Hypothesis Transfer with Gradual Knowledge Distillation
abstract
The ability to adapt their perception to changing environments is a core characterization of intelligent robots. At present, Unsupervised Domain Adaptation (UDA) methods are used to address this problem where the adaptation task is formulated as a transfer problem from a well-described scenario (source domain) to a new scenario (target domain). In order to implement the domain adaptation, these methods require access to the source data for achieving the distribution matching between both domains. However, in many real-world applications, the source data is inaccessible and only a source model pre-trained on the source domain is available during the transfer process. Therefore, the traditional UDA methods cannot support the challenging setting. This paper developed a new hypothesis transfer method to achieve model adaptation with gradual knowledge distillation. Specifically, we first prepare a source model through training a deep network on the labeled source domain by supervised learning. Then, we transfer the source model to the unlabeled target domain by self-training. To implement gradual knowledge distillation, we sliced the self-training into several epochs and then used the soft pseudo-labels from the latest epoch to guide the current epoch. In this process, the soft labels were generated by a semantic fusion on a proposed geometry of the neighborhood. To regulate the self-training, we developed a new objective constructed on the neighborhood. Experiments on three benchmarks have confirmed the state-of-the-art results of our method.
Song Tang 0001, Yuji Shi, Zhiyuan Ma 0001, Jianzhi Lyu, Qingdu Li, Jianwei Zhang 0001
IROS1
2019 PointNetGPD: Detecting Grasp Configurations from Point Sets
abstract
In this paper, we propose an end-to-end grasp evaluation model to address the challenging problem of localizing robot grasp configurations directly from the point cloud. Compared to recent grasp evaluation metrics that are based on handcrafted depth features and a convolutional neural network (CNN), our proposed PointNetGPD is lightweight and can directly process the 3D point cloud that locates within the gripper for grasp evaluation. Taking the raw point cloud as input, our proposed grasp evaluation network can capture the complex geometric structure of the contact area between the gripper and the object even if the point cloud is very sparse. To further improve our proposed model, we generate a large-scale grasp dataset with 350k real point cloud and grasps with the YCB object set for training. The performance of the proposed model is quantitatively measured both in simulation and on robotic hardware. Experiments on object grasping and clutter removal show that our proposed model generalizes well to novel objects and outperforms state-of-the-art methods. Code and video are available at https://lianghongzhuo.github.io/PointNetGPD.
Hongzhuo Liang, Xiaojian Ma 0001, Shuang Li 0014, Michael Görner, Song Tang 0001, Bin Fang 0003, Fuchun Sun 0001, Jianwei Zhang 0001
ICRA5
2019 Visual Domain Adaptation Exploiting Confidence-Samples
abstract
Domain adaptation methods are used to address a problem, in which train scenario (source domain) and test scenario (target domain) are different. The existing methods mainly perform adaptation via reducing domain discrepancy from the view of a probability distribution. However, the idea of probability distribution matching always leads to a complex optimization process. Thereby these methods are difficult to apply in some scenario like online application or fast perception in dynamic environments. In this paper, we propose a new and simple domain adaptation method that utilizes confidence- samples to facilitate the classifier training on the target domain. Here, the confidence-samples are a subset of the target samples, and they have very credibly predicted labels. In order to detect the samples, a Category Similarity Collaborative Representation (CSCR) is first developed, by which the raw labels of all target samples are predicted using the smallest projection error according to the law of category. After this, the confidence score of the raw predicted labels is evaluated by the energy context information of CSCR. Finally, the target samples with a high confidence score are selected. Because of the linearity of CSCR, our method avoids complex optimization for matching the probability distribution. Empirical studies on a standard dataset demonstrate the advantages of our method.
Song Tang 0001, Yunfeng Ji, Jianzhi Lyu, Jinpeng Mi, Qingdu Li, Jianwei Zhang 0001
IROS1
2019 Boosting VLAD with weighted fusion of local descriptors for image retrieval
Qingjie Zhao, Jimmy T. Mbelwa, Song Tang 0001, Jianwei Zhang 0001
Multim. Tools Appl.5
2019 Adaptive pedestrian detection by predicting classifier
Song Tang 0001, Mao Ye 0001, Pei Xu 0009, Xudong Li 0001
Neural Comput. Appl.1
2019 Weighted two-step aggregated VLAD for image retrieval
Qingjie Zhao, Jimmy T. Mbelwa, Song Tang 0001, Jianwei Zhang 0001
Vis. Comput.4
2017 A Method of Pedestrian Re-identification Based on Multiple Saliency Features
Cailing Wang, Yechao Xu, Guangwei Gao, Song Tang 0001, Xiaoyuan Jing
ICONIP (6)4
2017 Movie Recommendation via BLSTM
Song Tang 0001, Zhiyong Wu 0001, Kang Chen 0001
MMM (2)1
2017 Accurate object detection using memory-based models in surveillance scenes
Xudong Li 0001, Mao Ye 0001, Yiguang Liu, Feng Zhang 0052, Song Tang 0001
Pattern Recognit.6
2016 Memory-based object detection in surveillance scenes
abstract
Object detection is a significant step of intelligent video surveillance. The existing methods achieve the goals by technically designing or learning special features and detection models. Conversely, we propose a method to simulate the mechanism of memory and prediction in our brain. Firstly, a fix-sized window is slid on a static image to generate sequences. Then, a convolutional neural network extracts the sequence features. Finally, a long short-term memory receives these sequence features in proper order to memorize and recognize the sequential patterns. Our contributions are 1) a memory-based classification model in which both of feature learning and sequence learning are integrated subtly, and 2) a memory-based prediction model which is specially designed to predict the potential object locations in the surveillance scene. Compared with the state-of-the-art methods, our method obtains the best performance on three surveillance datasets. Our method may give some new insights on object detection researches.
Xudong Li 0001, Mao Ye 0001, Feng Zhang 0052, Song Tang 0001
ICME5