EDBT 2026 Demo / reviewers in the wild / expert
Song Tang 0001
dblp:181/8826-1
· DBLP profile ↗
34ranked-venue papers
11as first author
25since 2021 · last 2026
0000-0003-2635-1872ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 8 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 10 since 2021Systems, architecture and hardware · 5 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Progressive category prototype optimization for black-box domain adaptation
Lihua Zhou, Song Tang 0001, Yan Gan, Mao Ye 0001 |
Neurocomputing | 3 |
| 2026 | Source-Free Domain Adaptive Object Detection with semantics compensation
Song Tang 0001, Jiuzheng Yang, Mao Ye 0001, Yan Gan, Xiatian Zhu |
Pattern Recognit. | 1 |
| 2026 | From Point to Flow: Enhancing Unsupervised Domain Adaptation With Flow ClassificationabstractUnsupervised domain adaptation aims to transfer knowledge from a labeled source domain to an unlabeled target domain. Existing methods, whether based on distribution matching or self-supervised learning, often focus solely on classifying individual source samples, potentially overlooking discriminative information. To address this limitation, we propose FlowUDA, a novel plugin method that enhances existing UDA frameworks by constructing semantically invariant flows from individual source samples to corresponding target samples, forming cross-domain trajectories. By leveraging a diffusion network guided by ordinary differential equations, FlowUDA ensures these flows preserve the topological structure of the source domain, maintaining their distinguishability. Our method then classifies these flows by sampling points along them and transferring labels from source samples, effectively capturing spatial relationships between domains. In essence, FlowUDA transforms the traditional point-based classification on individual source samples into flow-based classification on flows, allowing the model to learn richer, more discriminative features that bridge the gap between source and target domains. Extensive experiments on standard benchmarks demonstrate that integrating FlowUDA into existing UDA methods leads to notable performance gains, highlighting its effectiveness in addressing domain shift challenges. Lihua Zhou, Mao Ye 0001, Nianxin Li, Song Tang 0001, Xu-Qian Fan, Lei Deng 0001, Zhen Lei 0001, Xiatian Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Long-Short Match for Lost Control in UAV Multi-Object TrackingabstractMulti-Object Tracking (MOT) in Unmanned Aerial Vehicles (UAV) aims to continuously and stably detect and track objects in videos captured by UAVs. In existing MOT tracking-by-detection schemes, the tracker with a fixed step size is always employed, and a fixed length of past tracking information is input to the tracker to guide position prediction. However, the limited prediction range of a single-scale tracker leads to frequent tracking losses, and limited historical information also reduces tracking accuracy. To address these limitations, we propose a novel Long-Short Match (LSMTrack) tracking method. The key idea is to use long and short trackers and maintain a long-term motion state to improve tracking performance, thus reducing the likelihood of entering the lost status. To this end, a new Mamba-based tracker and a long-short match strategy are proposed. For long and short trackers, the same architecture is used based on Mamba. Unlike the previous Mamba-based approach, the proposed tracker maintains a long-term state while updating the state and making position predictions in each time step, so we call it a step Mamba tracker. Meanwhile, we devise a long-short match strategy at the inference stage to integrate long and short trackers, and design a lost control operation which updates the long-term states using historical state values. In this way, the matching probability and the inference efficiency are guaranteed. Experimental results on two UAV MOT datasets confirm the state-of-the-art performance. Specifically, the best results are achieved in terms of two popular MOTA and IDF1 tracking evaluation metrics. Zi-Zhuang Zou, Mao Ye 0001, Luping Ji, Lihua Zhou, Song Tang 0001, Yan Gan, Shuai Li 0005 |
IEEE Trans. Multim. | 5 |
| 2025 | Self-Prompting Analogical Reasoning for UAV Object DetectionabstractUnmanned Aerial Vehicle Object Detection (UAVOD) presents unique challenges due to varying altitudes, dynamic backgrounds, and the small size of objects. Traditional detection methods often struggle with these challenges, as they typically rely on visual feature only and fail to extract the semantic relations between the objects. To address these limitations, we propose a novel approach named Self-Prompting Analogical Reasoning (SPAR). Our method utilizes the vision-language model (CLIP) to generate context-aware prompts based on image feature, providing rich semantic information that guides analogical reasoning. SPAR includes two main modules: self-prompting and analogical reasoning. Self-prompting module based on learnable description and CLIP-text encoder generates context-aware prompt by combining specific image feature; then an objectness prompt score map is produced by computing the similarity between pixel-level features and context-aware prompt. With this score map, multi-scale image features are enhanced and pixel-level features are chosen for graph construction. While for analogical reasoning module, graph nodes consists of category-level prompt nodes and pixel-level image feature nodes. Analogical inference is based graph convolution. Under the guidance of category-level nodes, different-scale object features have been enhanced, which helps achieve more accurate detection of challenging objects. Extensive experiments illustrate that SPAR outperforms traditional methods, offering a more robust and accurate solution for UAVOD. Nianxin Li, Mao Ye 0001, Lihua Zhou, Song Tang 0001, Yan Gan, Zizhuo Liang, Xiatian Zhu |
AAAI | 4 |
| 2025 | Pseudo Visible Feature Fine-Grained Fusion for Thermal Object DetectionabstractThermal object detection is a critical task in various fields, such as surveillance and autonomous driving. Current state-of-the-art (SOTA) models always leverage a prior Thermal-To-Visible (T2V) translation model to obtain visible spectrum information, followed by a cross-modality aggregation module to fuse information from both modalities. However, this fusion approach does not fully exploit the complementary visible spectrum information beneficial for thermal detection. To address this issue, we propose a novel cross-modal fusion method called Pseudo Visible Feature Fine-Grained Fusion (PFGF). Specifically, a graph is constructed with nodes generated from multi-level thermal features and pseudo-visual latent features produced by the T2V model. Each level of features corresponds to a subgraph. An Inter-Mamba block is proposed to perform cross-modality fusion between nodes at the lowest level; while a Cascade Knowledge Integration (CKI) strategy is used to fuse low-level fused information to high-level subgraphs in a cascade manner. After several iterations of graph node updating, each subgraph outputs an aggregated feature to the detection head respectively. Unlike previous cross-modal fusion methods, our approach explicitly models high-level relationships between cross-modal data, effectively fusing different granularity information. Experimental results demonstrate that our method achieves SOTA detection performance. Code is available at https://github.com/liting1018/PFGF. Mao Ye 0001, Tianwen Wu, Nianxin Li, Shuaifeng Li, Song Tang 0001, Luping Ji |
CVPR | 6 |
| 2025 | Domain Adaptive Diabetic Retinopathy Grading with Model Absence and Flowing DataabstractDomain shift (the difference between source and target domains) poses a significant challenge in clinical applications, e.g., Diabetic Retinopathy (DR) grading. Despite considering certain clinical requirements, like source data privacy, conventional transfer methods are predominantly model-centered and often struggle to prevent model-targeted attacks. In this paper, we address a challenging Online Model-aGnostic Domain Adaptation (OMG-DA) setting, driven by the demands of clinical environments. This setting is characterized by the absence of the model and the flow of target data. To tackle the new challenge, we propose a novel approach, Generative Unadversarial ExampleS (GUES), which enables adaptation from a data-centric perspective. Specifically, we first theoretically reformulate conventional perturbation optimization in a generative way—learning a perturbation generation function with a latent input variable. During model instantiation, we leverage a Variational AutoEncoder to express this function. The encoder with the reparameterization trick predicts the latent input, whilst the decoder is responsible for the generation. Furthermore, the saliency map is selected as pseudo-perturbation labels. Because it not only captures potential lesions but also theoretically provides an upper bound on the function input, enabling the identification of the latent variable. Extensive experiments on DR benchmarks with both frozen pre-trained models and trainable models demonstrate the superiority of GUES, showing robustness even with small batch size. The source code and data are available at https://github.com/tntek/GUES. Wenxin Su, Song Tang 0001, Xiaojing Yi, Mao Ye 0001, Chunxiao Zu, Xiatian Zhu |
CVPR | 2 |
| 2025 | Proxy Denoising for Source-Free Domain AdaptationabstractSource-Free Domain Adaptation (SFDA) aims to adapt a pre-trained source model to an unlabeled target domain with no access to the source data. Inspired by the success of large Vision-Language (ViL) models in many applications, the latest research has validated ViL's benefit for SFDA by using their predictions as pseudo supervision. However, we observe that ViL's supervision could be noisy and inaccurate at an unknown rate, potentially introducing additional negative effects during adaption. To address this thus-far ignored challenge, we introduce a novel Proxy Denoising (__ProDe__) approach. The key idea is to leverage the ViL model as a proxy to facilitate the adaptation process towards the latent domain-invariant space. Concretely, we design a proxy denoising mechanism to correct ViL's predictions. This is grounded on a proxy confidence theory that models the dynamic effect of proxy's divergence against the domain-invariant space during adaptation. To capitalize the corrected proxy, we further derive a mutual knowledge distilling regularization. Extensive experiments show that ProDe significantly outperforms the current state-of-the-art alternatives under both conventional closed-set setting and the more challenging open-set, partial-set, generalized SFDA, multi-target, multi-source, and test-time settings. Our code and data are available at https://github.com/tntek/source-free-domain-adaptation. Song Tang 0001, Wenxin Su, Yan Gan, Mao Ye 0001, Jianwei Zhang 0001, Xiatian Zhu |
ICLR | 1 |
| 2025 | Multimodal Causal Reasoning for UAV Object DetectionabstractUnmanned Aerial Vehicle (UAV) object detection faces significant challenges due to complex environmental conditions and different imaging conditions. These factors introduce significant changes in scale and appearance, particularly for small objects that occupy limited pixels and exhibit limited information, complicating detection tasks. To address these challenges, we propose a Multimodel Causal Reasoning framework based on YOLO backbone for UAV Object Detection (MCR-UOD). The key idea is to use the backdoor adjustment to discover the condition-invariant object representation for easy detection. Specifically, the YOLO backbone is first adjusted to incorporate the pre-trained vision-language model. The original category labels are replaced with semantic text prompts, and the detection head is replaced with text-image contrastive learning. Based on this backbone, our method consists of two parts. The first part, named language guided region exploration, discovers the regions with high probability of object existence using text embeddings based on vision-language model such as CLIP. Another part is the backdoor adjustment casual reasoning module, which constructs a confounder dictionary tailored to different imaging conditions to capture global image semantics and derives a prior probability distribution of shooting conditions. During causal inference, we use the confounder dictionary and the prior to intervene on local instance features, disentangling condition variations, and obtaining condition-invariant representations. Experimental results on several public datasets confirm the state-of-the-art performance of our approach. The code, data and models will be released upon publication of this paper. Nianxin Li, Mao Ye 0001, Lihua Zhou, Shuaifeng Li, Song Tang 0001, Luping Ji, Ce Zhu |
NeurIPS | 5 |
| 2025 | Few-shot medical image segmentation with high-fidelity prototypesabstractFew-shot Semantic Segmentation (FSS) aims to adapt a pretrained model to new classes with as few as a single labeled training sample per class. Despite the prototype based approaches have achieved substantial success, existing models are limited to the imaging scenarios with considerably distinct objects and not highly complex background, e.g., natural images. This makes such models suboptimal for medical imaging with both conditions invalid. To address this problem, we propose a novel D etail S elf-refined P rototype Net work ( DSPNet ) to construct high-fidelity prototypes representing the object foreground and the background more comprehensively. Specifically, to construct global semantics while maintaining the captured detail semantics, we learn the foreground prototypes by modeling the multimodal structures with clustering and then fusing each in a channel-wise manner. Considering that the background often has no apparent semantic relation in the spatial dimensions, we integrate channel-specific structural information under sparse channel-aware regulation. Extensive experiments on three challenging medical image benchmarks show the superiority of DSPNet over previous state-of-the-art methods. The code and data are available at https://github.com/tntek/DSPNet . • A novel prototypical FSS approach DSPNet that enhances prototypes’ self-representation. • A class prototype self-refining method FSPA integrating the cluster prototypes. • A background prototype self-refining method BCMA coding channel-specific structure. Song Tang 0001, Shaxu Yan, Xiaozhi Qi, Mao Ye 0001, Jianwei Zhang 0001, Xiatian Zhu |
Medical Image Anal. | 1 |
| 2025 | Learning economically for Chinese word segmentation: tuning pretrained model via active learning and N-gram preference
Zhiyuan Ma 0001, Jiwei Qin, Song Tang 0001, Jinpeng Mi |
Neural Comput. Appl. | 3 |
| 2025 | Adaptive Surveillance Video Compression With Background HyperpriorabstractNeural surveillance video compression methods have demonstrated significant improvements over traditional video compression techniques. In current surveillance video compression frameworks, the first frame in a Group of Pictures (GOP) is usually compressed fully as an I frame, and the subsequent P frames are compressed by referencing this I frame at Low Delay P (LDP) encoding mode. However, this compression approach overlooks the utilization of background information, which limits its adaptability to different scenarios. In this paper, we propose a novel Adaptive Surveillance Video Compression framework based on background hyperprior, dubbed as ASVC. This background hyperprior is related with side information to assist in coding both the temporal and spatial domains. Our method mainly consists of two components. First, the background information from a GOP is extracted, modeled as hyperprior and is compressed by exiting methods. Then these hyperprior is used as side information to compress both I frames and P frames. ASVC effectively captures the temporal dependencies in the latent representations of surveillance videos by leveraging background hyperprior for auxiliary video encoding. The experimental results demonstrate that applying ASVC to traditional and learning based methods significantly improves performance. Song Tang 0001, Mao Ye 0001 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Intelligent Event-Triggered H∞ Load Frequency Control for Power Systems With Multiple-Resource DelaysabstractThis paper considers intelligent event-triggered${H}_{\infty }$Load Frequency Control (LFC) for a new kind of multi-area power system. Firstly, delay factors are introduced into the prime mover, engine, and energy storage unit, i.e, multiple-resource delays, and the impact of these delays on the dynamic behavior of system is discussed for the first time, making the model more realistic and accurate. Secondly, a new lemma is proposed, which introduces free variables, extending to the scenarios with time delays. Combining this lemma with the event-triggered mechanism, the control scheme is designed, which is featured by two characteristics. On the one hand, in the Lyapunov functional, looped terms on neighbor triggering instants are considered, reduce the conservatism of the obtained criteria by capturing the information of triggered instants. Meanwhile, a mandatory triggering mechanism is introduced to break the possible dead loop caused by long-term no-triggering. On the other hand, a genetic algorithm (GA) is employed to optimize the parameters of the trigger threshold, encouraging efficient control within the context of LFC and delays. The effectiveness of the proposed approach is confirmed by a typical numerical example with four case studies. Jinnan Luo, Kaibo Shi, Song Tang 0001, Ruimei Zhang, Jun Cheng 0004, Ju H. Park 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | LESEP: Boosting Adversarial Transferability via Latent Encoding and Semantic Embedding PerturbationsabstractTransferability and imperceptibility of adversarial examples are pivotal for assessing the efficacy of black-box attacks. While diffusion models have been employed to generate adversarial examples, leveraging their advanced image generation capability to enhance transferability and imperceptibility, these methods typically focus only on perturbing the image or latent space. They often ignore the critical role of semantic information in the denoising process, thereby impeding the improvement of the transferability of adversarial examples. Furthermore, the modification of high-level semantics inevitably introduces image blurring. This degradation in visual quality makes the adversarial examples more susceptible to detection. To overcome the above limitations, we are the first to utilize image latent encoding and semantic embedding perturbations to enhance the performance of adversarial attacks. Then, the LESEP method is proposed. In the LESEP framework, we first apply image latent encoding attack to achieve deception of the target model. Second, the semantic embedding attack enhances the transferability of adversarial examples. Additionally, we utilize the image restoration technique to guarantee the high imperceptibility of the crafted adversarial examples. Through comprehensive experiments on diverse datasets, different network architectures and defense methods, we have demonstrated that the LESEP method achieves outstanding transferability and imperceptibility while displaying strong robustness. Yan Gan, Chengqian Wu, Deqiang Ouyang, Song Tang 0001, Mao Ye 0001, Tao Xiang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Source-Free Domain Adaptation with Frozen Multimodal Foundation ModelabstractSource-Free Domain Adaptation (SFDA) aims to adapt a source model for a target domain, with only access to unlabeled target training data and the source model pretrained on a supervised source domain. Relying on pseudo labeling and/or auxiliary supervision, conventional methods are inevitably error-prone. To mitigate this limitation, in this work we for the first time explore the potentials of off-the-shelf vision-language (ViL) multimodal models (e.g., CLIP) with rich whilst heterogeneous knowledge. We find that directly applying the ViL model to the target domain in a zero-shot fashion is unsatisfactory, as it is not specialized for this particular task but largely generic. To make it task specific, we propose a novel Distilling multImodal Foundation mOdel (DIFO) approach. Specifically, DIFO alternates between two steps during adaptation: (i) Customizing the ViL model by maximizing the mutual information with the target model in a prompt learning manner, (ii) Distilling the knowledge of this customized ViL model to the target model. For more fine-grained and reliable distillation, we further introduce two effective regularization terms, namely most-likely category encouragement and predictive consistency. Extensive experiments show that DIFO significantly outperforms the state-of-the-art alternatives. Code is here. Song Tang 0001, Wenxin Su, Mao Ye 0001, Xiatian Zhu |
CVPR | 1 |
| 2024 | Adversarial Experts Model for Black-box Domain AdaptationabstractBlack-box domain adaptation treats the source domain model as a black box. During the transfer process, the only available information about the target domain is the noisy labels output by the black-box model. This poses significant challenges for domain adaptation. Conventional approaches typically tackle the black-box noisy label problem from two aspects: self-knowledge distillation and pseudo-label denoising, both achieving limited performance due to limited knowledge information. To mitigate this issue, we explore the potential of off-the-shelf vision-language (ViL) multimodal models with rich semantic information for black-box domain adaptation by introducing an Adversarial Experts Model (AEM). Specifically, our target domain model is designed as one feature extractor and two classifiers, trained over two stages: In the knowledge transferring stage, with a shared feature extractor, the black-box source model and the ViL model act as two distinct experts for joint knowledge contribution, guiding the learning of one classifier each. While contributing their respective knowledge, the experts are also updated due to their own limitation and bias. In the adversarial alignment stage, to further distill expert knowledge to the target domain model, adversarial learning is conducted between the feature extractor and the two classifiers. A new consistency-max loss function is proposed to measure two classifier consistency and further improve classifier prediction certainty. Extensive experiments on multiple datasets demonstrate the effectiveness of our approach. Code is available at https://github.com/singinger/AEM. Siying Xiao, Mao Ye 0001, Qichen He, Shuaifeng Li, Song Tang 0001, Xiatian Zhu |
ACM Multimedia | 5 |
| 2024 | Cloud Object Detector Adaptation by Integrating Different Source KnowledgeabstractWe propose to explore an interesting and promising problem, Cloud Object Detector Adaptation (CODA), where the target domain leverages detections provided by a large cloud model to build a target detector. Despite with powerful generalization capability, the cloud model still cannot achieve error-free detection in a specific target domain. In this work, we present a novel Cloud Object detector adaptation method by Integrating different source kNowledge (COIN). The key idea is to incorporate a public vision-language model (CLIP) to distill positive knowledge while refining negative knowledge for adaptation by self-promotion gradient direction alignment. To that end, knowledge dissemination, separation, and distillation are carried out successively. Knowledge dissemination combines knowledge from cloud detector and CLIP model to initialize a target detector and a CLIP detector in target domain. By matching CLIP detector with the cloud detector, knowledge separation categorizes detections into three parts: consistent, inconsistent and private detections such that divide-and-conquer strategy can be used for knowledge distillation. Consistent and private detections are directly used to train target detector; while inconsistent detections are fused based on a consistent knowledge generation network, which is trained by aligning the gradient direction of inconsistent detections to that of consistent detections, because it provides a direction toward an optimal target detector. Experiment results demonstrate that the proposed COIN method achieves the state-of-the-art performance. Shuaifeng Li, Mao Ye 0001, Lihua Zhou, Nianxin Li, Siying Xiao, Song Tang 0001, Xiatian Zhu |
NeurIPS | 6 |
| 2024 | Shooting condition insensitive unmanned aerial vehicle object detection
Jinzong Cui, Mao Ye 0001, Xiatian Zhu, Song Tang 0001 |
Expert Syst. Appl. | 5 |
| 2024 | Source-Free Domain Adaptation via Target Prediction Distribution SearchingabstractAbstract Existing Source-Free Domain Adaptation (SFDA) methods typically adopt the feature distribution alignment paradigm via mining auxiliary information (eg., pseudo-labelling, source domain data generation). However, they are largely limited due to that the auxiliary information is usually error-prone whilst lacking effective error-mitigation mechanisms. To overcome this fundamental limitation, in this paper we propose a novel Target Prediction Distribution Searching (TPDS) paradigm. Theoretically, we prove that in case of sufficient small distribution shift, the domain transfer error could be well bounded. To satisfy this condition, we introduce a flow of proxy distributions that facilitates the bridging of typically large distribution shift from the source domain to the target domain. This results in a progressive searching on the geodesic path where adjacent proxy distributions are regularized to have small shift so that the overall errors can be minimized. To account for the sequential correlation between proxy distributions, we develop a new pairwise alignment with category consistency algorithm for minimizing the adaptation errors. Specifically, a manifold geometry guided cross-distribution neighbour search is designed to detect the data pairs supporting the Wasserstein distance based shift measurement. Mutual information maximization is then adopted over these pairs for shift regularization. Extensive experiments on five challenging SFDA benchmarks show that our TPDS achieves new state-of-the-art performance. The code and datasets are available at https://github.com/tntek/TPDS . Song Tang 0001, An Chang, Fabian Zhang, Xiatian Zhu, Mao Ye 0001, Changshui Zhang |
Int. J. Comput. Vis. | 1 |
| 2024 | Source-free domain adaptation with Class Prototype Discovery
Lihua Zhou, Nianxin Li, Mao Ye 0001, Xiatian Zhu, Song Tang 0001 |
Pattern Recognit. | 5 |
| 2024 | Illumination Distribution-Aware Thermal Pedestrian DetectionabstractPedestrian detection is an important task in computer vision, which is also an important part of intelligent transportation systems. For privacy protection, thermal images are widely used in pedestrian detection problems. However, thermal pedestrian detection is challenging due to the significant effect of temperature variation on the illumination of images and that fine-grained illumination annotations are difficult to be acquired. The existing methods have attempted to exploit coarse-grained day/night labels, which however even hampers the model performance. In this work, we introduce a novel idea of regressing conditional thermal-visible feature distribution, dubbed as Illumination Distribution-Aware adaptation (IDA). The key idea is to predict the conditional visible feature distribution given a thermal image, subject to their pre-computed joint distribution. Specifically, we first estimate the thermal-visible feature joint distribution by constructing feature co-occurrence matrices, offering a conditional probability distribution for any given thermal image. With this pairing information, we then form a conditional probability distribution regression task for model optimization. Critically, as a model agnostic strategy, this allows the visible feature knowledge to be transferred to the thermal counterpart implicitly for learning more discriminating feature representation. Experiment results show that our method outperforms the prior art methods, which use extra illumination annotations. Besides, as a plug-in, our method can averagely reduce about 2% MR on KAIST dataset, and improve about 1% mAP on FLIR-aligned and Autonomous Vehicles datasets without extra calculation for test. Code is available athttps://github.com/HaMeow-lst1/IDA. Mao Ye 0001, Luping Ji, Song Tang 0001, Yan Gan, Xiatian Zhu |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Progressive Source-Aware Transformer for Generalized Source-Free Domain AdaptationabstractSource-free domain adaptation (SFDA) tends to forget the source domain, suffering from limitations in real-world scenarios. Recently, generalized source-free domain adaptation (GSFDA) problem naturally emerges, aiming for good performance on both target and source domains. The existing methods attempt to retain model parameters associated with the source domain to prevent such forgetting. However, this strategy is not conducive to improving cross-domain performance on the target domain, prioritizing mitigating forgetting on the source domain. This article introduces a Progressive Source-Aware Transformer approach for GSFDA, dubbed PSAT-GDA. Our core idea is to enforce the domain adaptation process to remember the source domain by imposing source guidance, offering a target domain-centric anti-forgetting mechanism. Specifically, for each epoch, a Transformer-based deep network is adapted to do domain alignment like the traditional SFDA method, because the transformer working on the image patch sequence helps to reduce image noise caused by domain shift. Meanwhile, another Transformer is designed to generate source guidance supervising domain alignment. By augmenting target sample and mining the source information from the historical models before current epoch, source injected feature group is constructed. Based on the Transformer mechanism, the attention block can select useful source information for each target sample. From it, we devise neighbour-based and augmentation-based regularizations to shape the source guidance. Experiments on three challenging datasets show that our method can achieve evident cross-domain improvement on the target domains. Also, it can mitigate forgetting on all domains after adapting to single or multiple target domains. Song Tang 0001, Yuji Shi, Mao Ye 0001, Changshui Zhang, Jianwei Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Weakly Supervised Referring Expression Grounding via Target-Guided Knowledge DistillationabstractWeakly supervised referring expression grounding aims to train a model without the manual labels between image regions and referring expressions during the training phase. Current predominant models often adopt deep structures to reconstruct the region-expression correspondence. A crucial deficiency of the existing approaches lies in that these models neglect to exploit potential valuable information to further improve their grounding performance. To address this issue, we leverage knowledge distillation as a unique scheme to excavate and transfer helpful information for acquiring a better model. Specifically, we propose a target-guided knowledge distillation framework that accounts for region-expression pairs reconstruction and matching. We reactivate the target-related prediction information learned by a pre-trained teacher model and transfer the target-related prediction knowledge from the teacher to guide the training process and boost the performance of the student model. We conduct extensive experiments on three benchmark datasets, i.e., RefCOCO, RefCOCO+, and RefCOCOg. Without bells and whistles, our approach achieves state-of-the-art results on several splits of benchmark datasets. The implementation codes and trained models are available at: https://github.com/dami23/WREG_KD. Jinpeng Mi, Song Tang 0001, Zhiyuan Ma 0001, Qingdu Li, Jianwei Zhang 0001 |
ICRA | 2 |
| 2022 | Semantic consistency learning on manifold for source data-free unsupervised domain adaptation
Song Tang 0001, Yan Zou, Jianzhi Lyu, Mao Ye 0001, Shouming Zhong, Jianwei Zhang 0001 |
Neural Networks | 1 |
| 2021 | Model Adaptation through Hypothesis Transfer with Gradual Knowledge DistillationabstractThe ability to adapt their perception to changing environments is a core characterization of intelligent robots. At present, Unsupervised Domain Adaptation (UDA) methods are used to address this problem where the adaptation task is formulated as a transfer problem from a well-described scenario (source domain) to a new scenario (target domain). In order to implement the domain adaptation, these methods require access to the source data for achieving the distribution matching between both domains. However, in many real-world applications, the source data is inaccessible and only a source model pre-trained on the source domain is available during the transfer process. Therefore, the traditional UDA methods cannot support the challenging setting. This paper developed a new hypothesis transfer method to achieve model adaptation with gradual knowledge distillation. Specifically, we first prepare a source model through training a deep network on the labeled source domain by supervised learning. Then, we transfer the source model to the unlabeled target domain by self-training. To implement gradual knowledge distillation, we sliced the self-training into several epochs and then used the soft pseudo-labels from the latest epoch to guide the current epoch. In this process, the soft labels were generated by a semantic fusion on a proposed geometry of the neighborhood. To regulate the self-training, we developed a new objective constructed on the neighborhood. Experiments on three benchmarks have confirmed the state-of-the-art results of our method. Song Tang 0001, Yuji Shi, Zhiyuan Ma 0001, Jianzhi Lyu, Qingdu Li, Jianwei Zhang 0001 |
IROS | 1 |
| 2019 | PointNetGPD: Detecting Grasp Configurations from Point SetsabstractIn this paper, we propose an end-to-end grasp evaluation model to address the challenging problem of localizing robot grasp configurations directly from the point cloud. Compared to recent grasp evaluation metrics that are based on handcrafted depth features and a convolutional neural network (CNN), our proposed PointNetGPD is lightweight and can directly process the 3D point cloud that locates within the gripper for grasp evaluation. Taking the raw point cloud as input, our proposed grasp evaluation network can capture the complex geometric structure of the contact area between the gripper and the object even if the point cloud is very sparse. To further improve our proposed model, we generate a large-scale grasp dataset with 350k real point cloud and grasps with the YCB object set for training. The performance of the proposed model is quantitatively measured both in simulation and on robotic hardware. Experiments on object grasping and clutter removal show that our proposed model generalizes well to novel objects and outperforms state-of-the-art methods. Code and video are available at https://lianghongzhuo.github.io/PointNetGPD. Hongzhuo Liang, Xiaojian Ma 0001, Shuang Li 0014, Michael Görner, Song Tang 0001, Bin Fang 0003, Fuchun Sun 0001, Jianwei Zhang 0001 |
ICRA | 5 |
| 2019 | Visual Domain Adaptation Exploiting Confidence-SamplesabstractDomain adaptation methods are used to address a problem, in which train scenario (source domain) and test scenario (target domain) are different. The existing methods mainly perform adaptation via reducing domain discrepancy from the view of a probability distribution. However, the idea of probability distribution matching always leads to a complex optimization process. Thereby these methods are difficult to apply in some scenario like online application or fast perception in dynamic environments. In this paper, we propose a new and simple domain adaptation method that utilizes confidence- samples to facilitate the classifier training on the target domain. Here, the confidence-samples are a subset of the target samples, and they have very credibly predicted labels. In order to detect the samples, a Category Similarity Collaborative Representation (CSCR) is first developed, by which the raw labels of all target samples are predicted using the smallest projection error according to the law of category. After this, the confidence score of the raw predicted labels is evaluated by the energy context information of CSCR. Finally, the target samples with a high confidence score are selected. Because of the linearity of CSCR, our method avoids complex optimization for matching the probability distribution. Empirical studies on a standard dataset demonstrate the advantages of our method. Song Tang 0001, Yunfeng Ji, Jianzhi Lyu, Jinpeng Mi, Qingdu Li, Jianwei Zhang 0001 |
IROS | 1 |
| 2019 | Boosting VLAD with weighted fusion of local descriptors for image retrieval
Qingjie Zhao, Jimmy T. Mbelwa, Song Tang 0001, Jianwei Zhang 0001 |
Multim. Tools Appl. | 5 |
| 2019 | Adaptive pedestrian detection by predicting classifier
Song Tang 0001, Mao Ye 0001, Pei Xu 0009, Xudong Li 0001 |
Neural Comput. Appl. | 1 |
| 2019 | Weighted two-step aggregated VLAD for image retrieval
Qingjie Zhao, Jimmy T. Mbelwa, Song Tang 0001, Jianwei Zhang 0001 |
Vis. Comput. | 4 |
| 2017 | A Method of Pedestrian Re-identification Based on Multiple Saliency Features
Cailing Wang, Yechao Xu, Guangwei Gao, Song Tang 0001, Xiaoyuan Jing |
ICONIP (6) | 4 |
| 2017 | Movie Recommendation via BLSTM
Song Tang 0001, Zhiyong Wu 0001, Kang Chen 0001 |
MMM (2) | 1 |
| 2017 | Accurate object detection using memory-based models in surveillance scenes
Xudong Li 0001, Mao Ye 0001, Yiguang Liu, Feng Zhang 0052, Song Tang 0001 |
Pattern Recognit. | 6 |
| 2016 | Memory-based object detection in surveillance scenesabstractObject detection is a significant step of intelligent video surveillance. The existing methods achieve the goals by technically designing or learning special features and detection models. Conversely, we propose a method to simulate the mechanism of memory and prediction in our brain. Firstly, a fix-sized window is slid on a static image to generate sequences. Then, a convolutional neural network extracts the sequence features. Finally, a long short-term memory receives these sequence features in proper order to memorize and recognize the sequential patterns. Our contributions are 1) a memory-based classification model in which both of feature learning and sequence learning are integrated subtly, and 2) a memory-based prediction model which is specially designed to predict the potential object locations in the surveillance scene. Compared with the state-of-the-art methods, our method obtains the best performance on three surveillance datasets. Our method may give some new insights on object detection researches. Xudong Li 0001, Mao Ye 0001, Feng Zhang 0052, Song Tang 0001 |
ICME | 5 |