Yan Gan

dblp:226/5007 · DBLP profile ↗
← Back
39ranked-venue papers
13as first author
34since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 7 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 9 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-exposure high dynamic range reconstruction by incorporating imaging knowledge
Mao Ye 0001, Dengyan Luo, Yan Gan
Eng. Appl. Artif. Intell.4
2026 GS-YOLO: A lightweight and high-performance method for PCB surface defect detection
abstract
The compact layout and complex background of printed circuit boards (PCBs) pose significant challenges for surface defect detection. With limited computational resources, ensuring that PCB defect detection models are lightweight while maintaining detection performance is a persistent challenge. To address this, we propose a novel GS-YOLO network based on the YOLOv5s framework, designed to accurately identify PCB surface defects with lower computational cost. In the GS-YOLO model, we introduce the C3Ghost-S module, integrating it into the network’s backbone and neck structures. This not only reduces computational cost but also improves the model’s detection accuracy and generalization ability. Additionally, we propose the GA-SPPF module, which extracts global information and combines it with local information obtained from the SPPF module, helping the model learn both local and global features for a more comprehensive understanding of the image. Notably, the modules we designed are plug-and-play within YOLO architecture series, allowing for easy integration into different YOLO-based frameworks. Compared to YOLOv5s, GS-YOLO improves mAP by 8.1 %, reduces the parameter count by 26.49 %, and lowers computational complexity by 32.27 %. To further validate the model’s generalization capability, we conduct extensive evaluations on both the NEU-DET and GC10-DET datasets. The robustness of GS-YOLO has also been verified under various simulated degradation conditions. Through extensive experiments and comparisons with other advanced models, GS-YOLO shows significant advantages, achieving an outstanding balance between model lightweighting and superior detection performance. Our code is available at: https://github.com/niuniuhhh/GS-YOLO/tree/main .
Guoxing Li, Yan Gan, Wei Zhang 0102, Hangjun Che
Expert Syst. Appl.2
2026 Progressive category prototype optimization for black-box domain adaptation
Lihua Zhou, Song Tang 0001, Yan Gan, Mao Ye 0001
Neurocomputing4
2026 Source-Free Domain Adaptive Object Detection with semantics compensation
Song Tang 0001, Jiuzheng Yang, Mao Ye 0001, Yan Gan, Xiatian Zhu
Pattern Recognit.5
2026 Long-Short Match for Lost Control in UAV Multi-Object Tracking
abstract
Multi-Object Tracking (MOT) in Unmanned Aerial Vehicles (UAV) aims to continuously and stably detect and track objects in videos captured by UAVs. In existing MOT tracking-by-detection schemes, the tracker with a fixed step size is always employed, and a fixed length of past tracking information is input to the tracker to guide position prediction. However, the limited prediction range of a single-scale tracker leads to frequent tracking losses, and limited historical information also reduces tracking accuracy. To address these limitations, we propose a novel Long-Short Match (LSMTrack) tracking method. The key idea is to use long and short trackers and maintain a long-term motion state to improve tracking performance, thus reducing the likelihood of entering the lost status. To this end, a new Mamba-based tracker and a long-short match strategy are proposed. For long and short trackers, the same architecture is used based on Mamba. Unlike the previous Mamba-based approach, the proposed tracker maintains a long-term state while updating the state and making position predictions in each time step, so we call it a step Mamba tracker. Meanwhile, we devise a long-short match strategy at the inference stage to integrate long and short trackers, and design a lost control operation which updates the long-term states using historical state values. In this way, the matching probability and the inference efficiency are guaranteed. Experimental results on two UAV MOT datasets confirm the state-of-the-art performance. Specifically, the best results are achieved in terms of two popular MOTA and IDF1 tracking evaluation metrics.
Zi-Zhuang Zou, Mao Ye 0001, Luping Ji, Lihua Zhou, Song Tang 0001, Yan Gan, Shuai Li 0005
IEEE Trans. Multim.6
2025 Self-Prompting Analogical Reasoning for UAV Object Detection
abstract
Unmanned Aerial Vehicle Object Detection (UAVOD) presents unique challenges due to varying altitudes, dynamic backgrounds, and the small size of objects. Traditional detection methods often struggle with these challenges, as they typically rely on visual feature only and fail to extract the semantic relations between the objects. To address these limitations, we propose a novel approach named Self-Prompting Analogical Reasoning (SPAR). Our method utilizes the vision-language model (CLIP) to generate context-aware prompts based on image feature, providing rich semantic information that guides analogical reasoning. SPAR includes two main modules: self-prompting and analogical reasoning. Self-prompting module based on learnable description and CLIP-text encoder generates context-aware prompt by combining specific image feature; then an objectness prompt score map is produced by computing the similarity between pixel-level features and context-aware prompt. With this score map, multi-scale image features are enhanced and pixel-level features are chosen for graph construction. While for analogical reasoning module, graph nodes consists of category-level prompt nodes and pixel-level image feature nodes. Analogical inference is based graph convolution. Under the guidance of category-level nodes, different-scale object features have been enhanced, which helps achieve more accurate detection of challenging objects. Extensive experiments illustrate that SPAR outperforms traditional methods, offering a more robust and accurate solution for UAVOD.
Nianxin Li, Mao Ye 0001, Lihua Zhou, Song Tang 0001, Yan Gan, Zizhuo Liang, Xiatian Zhu
AAAI5
2025 Co-GNN: A Co-optimization Framework for Memory and Computation in Sampling-Based GNN Training
Yan Gan, Yujuan Tan, Yujiao Wang, Zongjie Wang, Duo Liu 0002, Ao Ren, Kan Zhong, Chaoxia Qin, Mingrui Qiang
ICIC (21)1
2025 Proxy Denoising for Source-Free Domain Adaptation
abstract
Source-Free Domain Adaptation (SFDA) aims to adapt a pre-trained source model to an unlabeled target domain with no access to the source data. Inspired by the success of large Vision-Language (ViL) models in many applications, the latest research has validated ViL's benefit for SFDA by using their predictions as pseudo supervision. However, we observe that ViL's supervision could be noisy and inaccurate at an unknown rate, potentially introducing additional negative effects during adaption. To address this thus-far ignored challenge, we introduce a novel Proxy Denoising (__ProDe__) approach. The key idea is to leverage the ViL model as a proxy to facilitate the adaptation process towards the latent domain-invariant space. Concretely, we design a proxy denoising mechanism to correct ViL's predictions. This is grounded on a proxy confidence theory that models the dynamic effect of proxy's divergence against the domain-invariant space during adaptation. To capitalize the corrected proxy, we further derive a mutual knowledge distilling regularization. Extensive experiments show that ProDe significantly outperforms the current state-of-the-art alternatives under both conventional closed-set setting and the more challenging open-set, partial-set, generalized SFDA, multi-target, multi-source, and test-time settings. Our code and data are available at https://github.com/tntek/source-free-domain-adaptation.
Song Tang 0001, Wenxin Su, Yan Gan, Mao Ye 0001, Jianwei Zhang 0001, Xiatian Zhu
ICLR3
2025 ODE-based generative modeling: Learning from a single natural image
Jian Yue, Yan Gan, Lihua Zhou, Shuaifeng Li, Mao Ye 0001
Expert Syst. Appl.2
2025 GNNBoost: Accelerating sampling-based GNN training on large scale graph by optimizing data preparation
Yujuan Tan, Yan Gan, Zhaoyang Zeng, Zhuoxin Bai, Lei Qiao 0002, Duo Liu 0002, Kan Zhong, Ao Ren
J. Syst. Archit.2
2025 LESEP: Boosting Adversarial Transferability via Latent Encoding and Semantic Embedding Perturbations
abstract
Transferability and imperceptibility of adversarial examples are pivotal for assessing the efficacy of black-box attacks. While diffusion models have been employed to generate adversarial examples, leveraging their advanced image generation capability to enhance transferability and imperceptibility, these methods typically focus only on perturbing the image or latent space. They often ignore the critical role of semantic information in the denoising process, thereby impeding the improvement of the transferability of adversarial examples. Furthermore, the modification of high-level semantics inevitably introduces image blurring. This degradation in visual quality makes the adversarial examples more susceptible to detection. To overcome the above limitations, we are the first to utilize image latent encoding and semantic embedding perturbations to enhance the performance of adversarial attacks. Then, the LESEP method is proposed. In the LESEP framework, we first apply image latent encoding attack to achieve deception of the target model. Second, the semantic embedding attack enhances the transferability of adversarial examples. Additionally, we utilize the image restoration technique to guarantee the high imperceptibility of the crafted adversarial examples. Through comprehensive experiments on diverse datasets, different network architectures and defense methods, we have demonstrated that the LESEP method achieves outstanding transferability and imperceptibility while displaying strong robustness.
Yan Gan, Chengqian Wu, Deqiang Ouyang, Song Tang 0001, Mao Ye 0001, Tao Xiang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 Quantifying Privacy Risks of Behavioral Semantics in Mobile Communication Services
abstract
Location-based mobile services, while improving user daily life, also raise significant privacy concerns in the sharing of location data. These trajectories indicate users’ traveling behavioural traces with rich semantics derived from open-source information. Behavioral-semantic analysis reveals users’ travelling motivations and underlying behavioral patterns. It contributes to attackers launching inferential attacks for behavior prediction, identity identification, or other privacy invasions, even when the location data is protected. It remains open to the issues of behavioral-semantic privacy-risk quantification and privacy-protection evaluation. This paper aims to reveal such semantic privacy risks of user behaviors arising from the publication of location trajectories in mobile scenarios. We formalize user semantic-mobility process to analyze his underlying behavior patterns. Then, we design semantic inference algorithms conditional on the released trajectory to reason about the observation-based likelihood of the user’s actual staying and transfer behaviours and behavioural-trace tracking. Extensive experiments with real-world data demonstrate their performance on inference accuracy and semantic similarity, offering a quantification criterion for deploying mobile privacy protection.
Guoying Qiu, Tiecheng Bai, Guoming Tang, Deke Guo, Chuandong Li 0001, Yan Gan, Baoping Zhou, Yulong Shen 0001
IEEE Trans. Inf. Forensics Secur.6
2025 SFCM-AEG: Source-Free Cross-Modal Adversarial Example Generation
abstract
In this paper, we present a novel task of source-free cross-modal adversarial example generation, which generates adversarial examples based on textual descriptions of attackers. This task has two challenges as follows. First, how to generate adversarial examples when the clean examples are missing or inaccessible. Second, how to achieve fine-grained custom adversarial example generation according to the semantic descriptions of the attackers. Existing adversarial example generation methods can not effectively deal with these two challenges. To address these challenges, we propose a Source-Free Cross-Modal Adversarial Example Generation framework, abbreviated as SFCM-AEG. Within the SFCM-AEG model, we firstly leverage a pre-trained GPT as a simulator to construct textual descriptions of attackers by labels. Following this, we employ a diffusion model to synthesize an image that aligns with the generated textual description. Finally, the generated images are converted into adversarial examples using an adversarial example generation method. Experimental results demonstrate that our proposed SFCM-AEG method can generate adversarial examples with customized semantic descriptions, without relying on clean examples, while achieving strong attack performance in a white-box setting.
Yan Gan, Xinyao Xiao, Tao Xiang 0001, Chengqian Wu, Deqiang Ouyang
IEEE Trans. Multim.1
2025 Generative Adversarial Networks with Learnable Auxiliary Module for Image Synthesis
abstract
Training generative adversarial networks (GANs) for noise-to-image synthesis is a challenge task, primarily due to the instability of GANs’ training process. One of the key issues is the generator’s sensitivity to input data, which can cause sudden fluctuations in the generator’s loss value with certain inputs. This sensitivity suggests an inadequate ability to resist disturbances in the generator, causing the discriminator’s loss value to oscillate and negatively impacting the discriminator. Then, the negative feedback of discriminator is also not conducive to updating generator’s parameters, leading to suboptimal image generation quality. In response to this challenge, we present an innovative GANs model equipped with a learnable auxiliary module that processes auxiliary noise. The core objective of this module is to enhance the stability of both the generator and discriminator throughout the training process. To achieve this target, we incorporate a learnable auxiliary penalty and an augmented discriminator, designed to control the generator and reinforce the discriminator’s stability, respectively. We further apply our method to the Hinge and LSGANs loss functions, illustrating its efficacy in reducing the instability of both the generator and the discriminator. The tests we conducted on LSUN, CelebA, Market-1501, and Creative Senz3D datasets serve as proof of our method’s ability to improve the training stability and overall performance of the baseline methods.
Yan Gan, Chenxue Yang, Mao Ye 0001, Renjie Huang, Deqiang Ouyang
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Temporal cues enhanced multimodal learning for action recognition in RGB-D videos
Zhiyuan Ma 0001, Jinpeng Mi, Yan Gan, Mao Ye 0001, Jianwei Zhang 0001
Neurocomputing6
2024 Behavioral-Semantic Privacy Protection for Continual Social Mobility in Mobile-Internet Services
abstract
Crowdsensing-based mobile Internet, while facilitating users’ daily life, also raises privacy concerns because of sharing user location trajectories. Combining with open-source network information, these trajectories reveal the semantics of users’ social behaviors in their travels, thus indicating their behavioral traces. Based on such social mobilities, attackers can explore users’ potential behavioral patterns and launch powerful behavioral-semantic inferential attacks for behavior prediction, identity identification, and threatening users’ location-related mobile privacy. Even through privacy protection, the released similar anonymous semantics may still bring significant privacy gains to such attacks. To the best of our knowledge, there is still no effective technique to counter such attacks and protect user behavioral semantics in mobile Internet services. To this end, this article proposes a posterior behavioral-semantic privacy-preserving solution, BSPri, by simulating the inferential attacks to eliminate the privacy risks associated with released traces. Specifically, we represent the logical association between semantic attributes and propose a similar semantic clustering and ranking method. Then, we formalize the user social-mobility stochastic process to characterize the privacy risks arising from the attacker’s observation of the released trajectory, and define a observation-based posterior privacy authentication criteria to filter anonymous semantics further. Finally, we generate synthetic trajectories with similar anonymity semantics, which bring attackers insignificant privacy gain, for users to participate in applications. Extensive experiments with the real-world data set demonstrate that our BSPri achieves an effective privacy-preserving performance, i.e., rigorous posterior-privacy constraint with limited data-availability loss, such as, distance 752 m$(47$-m closer, compared with our previous work MSP), direction deviation$39.5^{\circ }~(11.5^{\circ }$smaller), and semantic similarity$43.4\%~(8.4\%$closer).
Guoying Qiu, Guoming Tang, Chuandong Li 0001, Deke Guo, Yulong Shen 0001, Yan Gan
IEEE Internet Things J.6
2024 DSG-BTra: Differentially Semantic-Generalized Behavioral Trajectory for Privacy-Preserving Mobile Internet Services
abstract
While facilitating user daily lives, the booming development of mobile Internet services raises their privacy concerns because of the need to share travel trajectories. Due to the differences in access patterns and sensitive location attributes, behavioral semantics of user travel suffer from different degrees of leakage risks and have personalized privacy requirements. Semantic mobility-aware personalized privacy protection is still an open research issue in mobile scenarios. To this end, we propose a differentially semantic-generalized behavioral trajectory (DSG-BTra) for achieving privacy-preserving mobile Internet services. Specifically, we first explore the underlying behavioral patterns by formalizing user social mobility. Then, we evaluate the differential privacy sensitivity of user behavior to indicate the risks it faces. Finally, we generalize the behavioral semantics with a sensitivity-quantified strength and generate a DSG-BTra for the user to participate in mobile services. Extensive experiments with real-world data sets demonstrate DSG-BTra achieves flexible balance between privacy protection and application QoS, e.g., reducing the inference probability to 0.18–0.26 with a semantic similarity of 0.3–0.5.
Guoying Qiu, Guoming Tang, Chuandong Li 0001, Deke Guo, Yulong Shen 0001, Yan Gan
IEEE Internet Things J.6
2024 BGS: Accelerate GNN training on multiple GPUs
Yujuan Tan, Zhuoxin Bai, Duo Liu 0002, Zhaoyang Zeng, Yan Gan, Ao Ren, Xianzhang Chen, Kan Zhong
J. Syst. Archit.5
2024 SPGAN: Siamese projection Generative Adversarial Networks
Yan Gan, Tao Xiang 0001, Deqiang Ouyang, Mingliang Zhou 0001, Mao Ye 0001
Knowl. Based Syst.1
2024 Current status and trends of technology, methods, and applications of Human-Computer Intelligent Interaction (HCII): A bibliometric research
Zijie Ding, Yingrui Ji, Yan Gan, Yukun Xia
Multim. Tools Appl.3
2024 Illumination Distribution-Aware Thermal Pedestrian Detection
abstract
Pedestrian detection is an important task in computer vision, which is also an important part of intelligent transportation systems. For privacy protection, thermal images are widely used in pedestrian detection problems. However, thermal pedestrian detection is challenging due to the significant effect of temperature variation on the illumination of images and that fine-grained illumination annotations are difficult to be acquired. The existing methods have attempted to exploit coarse-grained day/night labels, which however even hampers the model performance. In this work, we introduce a novel idea of regressing conditional thermal-visible feature distribution, dubbed as Illumination Distribution-Aware adaptation (IDA). The key idea is to predict the conditional visible feature distribution given a thermal image, subject to their pre-computed joint distribution. Specifically, we first estimate the thermal-visible feature joint distribution by constructing feature co-occurrence matrices, offering a conditional probability distribution for any given thermal image. With this pairing information, we then form a conditional probability distribution regression task for model optimization. Critically, as a model agnostic strategy, this allows the visible feature knowledge to be transferred to the thermal counterpart implicitly for learning more discriminating feature representation. Experiment results show that our method outperforms the prior art methods, which use extra illumination annotations. Besides, as a plug-in, our method can averagely reduce about 2% MR on KAIST dataset, and improve about 1% mAP on FLIR-aligned and Autonomous Vehicles datasets without extra calculation for test. Code is available athttps://github.com/HaMeow-lst1/IDA.
Mao Ye 0001, Luping Ji, Song Tang 0001, Yan Gan, Xiatian Zhu
IEEE Trans. Intell. Transp. Syst.5
2024 Attribute-guided face adversarial example generation
Yan Gan, Xinyao Xiao, Tao Xiang 0001
Vis. Comput.1
2023 Cross-domain video action recognition via adaptive gradual learning
Zhenwei Bao, Jinpeng Mi, Yan Gan, Mao Ye 0001, Jianwei Zhang 0001
Neurocomputing4
2023 Generative adversarial networks with adaptive learning strategy for noise-to-image synthesis
Yan Gan, Tao Xiang 0001, Hangcheng Liu, Mao Ye 0001, Mingliang Zhou 0001
Neural Comput. Appl.1
2023 PFGAN: Fast transformers for image synthesis
Tianguang Zhang, Wei Zhang 0102, Yan Gan
Pattern Recognit. Lett.4
2023 Towards Query-Efficient Black-Box Attacks: A Universal Dual Transferability-Based Framework
abstract
Adversarial attacks have threatened the application of deep neural networks in security-sensitive scenarios. Most existing black-box attacks fool the target model by interacting with it many times and producing global perturbations. However, all pixels are not equally crucial to the target model; thus, indiscriminately treating all pixels will increase query overhead inevitably. In addition, existing black-box attacks take clean samples as start points, which also limits query efficiency. In this article, we propose a novel black-box attack framework, constructed on a strategy of dual transferability (DT), to perturb the discriminative areas of clean examples within limited queries. The first kind of transferability is the transferability of model interpretations. Based on this property, we identify the discriminative areas of clean samples for generating local perturbations. The second is the transferability of adversarial examples, which helps us to produce local pre-perturbations for further improving query efficiency. We achieve the two kinds of transferability through an independent auxiliary model and do not incur extra query overhead. After identifying discriminative areas and generating pre-perturbations, we use the pre-perturbed samples as better start points and further perturb them locally in a black-box manner to search the corresponding adversarial examples. The DT strategy is general; thus, the proposed framework can be applied to different types of black-box attacks. We conduct extensive experiments to show that, under various system settings, our framework can significantly improve the query efficiency of existing black-box attacks and attack success rates.
Tao Xiang 0001, Hangcheng Liu, Shangwei Guo, Yan Gan, Wenjian He, Xiaofeng Liao 0001
ACM Trans. Intell. Syst. Technol.4
2022 Source data-free domain adaptation for a faster R-CNN
Mao Ye 0001, Yan Gan, Yiguang Liu
Pattern Recognit.4
2022 HashFormer: Vision Transformer Based Deep Hashing for Image Retrieval
abstract
Deep image hashing aims to map an input image to compact binary codes by deep neural network, to enable efficient image retrieval across large-scale dataset. Due to the explosive growth of modern data, deep hashing has gained growing attention from research community. Recently, convolutional neural networks like ResNet have dominated in deep hashing. Nevertheless, motivated by the recent advancements of vision transformers, we propose a pure transformer-based framework, called as HashFormer, to tackle the deep hashing task. Specifically, we utilize vision transformer (ViT) as our backbone, and treat binary codes as the intermediate representations for our surrogate task,i.e., image classification. In addition, we observe that the binary codes suitable for classification are sub-optimal for retrieval. To mitigate this problem, we present a novel average precision loss, which enables us to directly optimize the retrieval accuracy. To the best of our knowledge, our work is one of the pioneer works to address deep hashing learning problems without convolutional neural networks (CNNs). We perform comprehensive experiments on three widely-studied datasets: CIFAR-10, NUSWIDE and ImageNet. The proposed method demonstrates promising results against existing state-of-the-art works, validating the advantages and merits of our HashFormer.
Tao Li 0063, Lishen Pei, Yan Gan
IEEE Signal Process. Lett.4
2022 EGM: An Efficient Generative Model for Unrestricted Adversarial Examples
abstract
Unrestricted adversarial examples allow the attacker to start attacks without given clean samples, which are quite aggressive and threatening. However, existing works for generating unrestricted adversary examples are quite inefficient and cannot achieve a high success rate. In this article, we explore an end-to-end and effective solution for unrestricted adversary example generation. To stabilize the training process and make our generative model converge to satisfactory results, we design a novel decoupled two-step efficient generative model (EGM), which contains a conditional reference generator and a conditional adversarial transformer. The former is responsible for generating reference samples from noises and source classes. The latter is responsible for converting the reference sample into adversarial examples corresponding to target classes. To improve the success rate, we design a new strategy, augmentation of adversarial labels to produce dynamic target labels and enhance the exploration ability of EGM. Such a strategy can be also applied to existing attacks to improve their attack success rates, which is of independent interest. We conduct extensive experiments to evaluate our proposed model and demonstrate the necessity of decoupling the generation process in EGM. Experimental results show our EGM is much faster and achieves a higher success rate than the state-of-the-art attacks.
Tao Xiang 0001, Hangcheng Liu, Shangwei Guo, Yan Gan, Xiaofeng Liao 0001
ACM Trans. Sens. Networks4
2021 Teacher-Supervised Generative Adversarial Networks
abstract
Although generative adversarial networks (GANs) show impressive effects on image generation, existing GANs suffer an unstable training process, and thus result in poor image quality sometimes. To solve this problem, we first introduce a supervision mechanism into GANs and propose a teacher-supervised GAN (GAN-T) model. Specifically, we design a teacher supervision mechanism to inspect whether the features of generated images are as real as those of real images. If not, we add action into the generator. The action takes the encoding of the real image as prior knowledge to guide the generation of samples. We then apply our proposed method to existing GANs to show its compatibility with them. Finally, we conduct extensive experiments on the tasks of noise-to-image generation and image translation, and experimental results show that our proposed method can significantly stabilize the training process of the generator and improve the quality of generated images.
Yan Gan, Tao Xiang 0001, Hangcheng Liu, Mao Ye 0001
ICME1
2021 A novel hybrid augmented loss discriminator for text-to-image synthesis
abstract
For the text-to-image synthesis task, most discriminators in existing generative adversarial networks based methods tend to fall into a local suboptimal state too early in the training process, resulting in the poor quality of generated images. To address the above problems, a hybrid augmented loss discriminator is designed. In this designed discriminator, to reduce the sensitivity of the discriminator classification recognition, make it pay attention to the semantic and structural changes, we add the loss value of the fake sample to the loss value of the real sample. Moreover, to indirectly guide the generator to generate samples, the loss value of the real sample is added to the fake sample. The loss value mixed with real and fake samples actually augments signal transmission. It perturbs parameter update of the discriminator during optimization and prevents the discriminator from falling into the local suboptimal state prematurely. Whereafter, we apply the proposed discriminator to two kinds of text-to-image synthesis tasks. Experimental results show that the proposed method can help the baseline models to improve performance.
Yan Gan, Mao Ye 0001, Shangming Yang, Tao Xiang 0001
Int. J. Intell. Syst.1
2021 Source data-free domain adaptation of object detector through domain-specific perturbation
abstract
The current unsupervised cross-domain detection methods need source domain data to retrain the detection model in target domain. However, the source domain data may be unavailable due to privacy, decentralization, or computation resource restrictions. A natural idea is to optimize the parameters of the source domain model by self-supervised learning based on pseudo labels. We propose another approach from the viewpoint of noise perturbation without pseudo-labeling. It can be assumed that the source and target domains are actually derived from a domain invariant space through domain-specific perturbations, respectively. A super target domain can be constructed by augmenting more target domain perturbations to the target domain images. The optimal direction of the target domain to the domain invariant space can be approximated as the alignment direction from the super target domain to the target domain. Based on this idea, we propose a novel method called SOAP (SOurce data-free domain Adaptation through domain Perturbation) which can remove domain perturbation from the target domain. The image-level, instance-level, and category consistency regularizations based on Mean Teacher structure are proposed to learn the correct alignment direction. Specifically, the category consistency can also further improve the classification accuracy. Extensive experiments on multiple domain adaptation scenarios demonstrate that SOAP achieves better performance surpassing the baseline (Faster R-CNN) and multiple state-of-the-art domain adaptation methods which need to access source domain data.
Mao Ye 0001, Yan Gan, Xue Li 0001, Yingying Zhu 0003
Int. J. Intell. Syst.4
2021 A new financial data forecasting model using genetic algorithm and long short-term memory network
Yelin Gao, Yan Gan, Mao Ye 0001
Neurocomputing3
2021 Domain adaptation of object detector using scissor-like networks
Mao Ye 0001, Yan Gan, Dongde Hou
Neurocomputing4
2020 Sentence guided object color change by adversarial learning
Yan Gan, Kedi Liu, Mao Ye 0001
Neurocomputing1
2020 Knowledge based domain adaptation for semantic segmentation
Mao Ye 0001, Yan Gan, Wencong Zhang
Knowl. Based Syst.3
2020 Generative adversarial networks with denoising penalty and sample augmentation
Yan Gan, Kedi Liu, Mao Ye 0001
Neural Comput. Appl.1
2019 Generative adversarial networks with augmentation and penalty
Yan Gan, Kedi Liu, Mao Ye 0001
Neurocomputing1
2018 Unpaired cross domain image translation with augmented auxiliary domain information
Yan Gan, Junxin Gong, Mao Ye 0001, Kedi Liu
Neurocomputing1