Fan Yang 0063

dblp:29/3081-63 · DBLP profile ↗
← Back
24ranked-venue papers
3as first author
21since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Redundancy Optimization via Mutual Information for Unsupervised Domain Adaptation
abstract
Unsupervised Domain Adaptation for image classification aims to adapt models trained on a labeled source domain to an unlabeled target domain, improving target domain classification performance. However, previous methods often focus solely on sample-level relationships, neglecting the problem of redundancy within samples due to the same feature information being expressed repeatedly. This oversight may reduce the effectiveness of domain adaptation. To address this problem, we propose Redundancy Optimization via Mutual Information for Unsupervised Domain Adaptation (ROMI) to enhance UDA by mitigating feature redundancy. Specifically, we first utilize mutual information to evaluate the redundancy between different feature dimensions and minimize this dimensional mutual information, thereby reducing the redundancy percentage in the features. Subsequently, we maximize the dimensional mutual information between pairs of positive samples in a contrastive learning framework to achieve a more unbiased feature representation. Our extensive experiments on three well-recognized UDA vision benchmarks provide compelling evidence of the efficacy of ROMI.
Xing Wei 0002, Dexuan Zhao, Fan Yang 0063, Taizhang Hu, Yang Lu 0015
ICME3
2025 Resource-Efficient Affordance Grounding with Complementary Depth and Semantic Prompts
abstract
Affordance refers to the functional properties that an agent perceives and utilizes from its environment, and is key perceptual information required for robots to perform actions. This information is rich and multimodal in nature. Existing multimodal affordance methods face limitations in extracting useful information, mainly due to simple structural designs, basic fusion methods, and large model parameters, making it difficult to meet the performance requirements for practical deployment. To address these issues, this paper proposes the BiT-Align image-depth-text affordance mapping framework. The framework includes a Bypass Prompt Module (BPM) and a Text Feature Guidance (TFG) attention selection mechanism. BPM integrates the auxiliary modality depth image directly as a prompt to the primary modality RGB image, embedding it into the primary modality encoder without introducing additional encoders. This reduces the model’s parameter count and effectively improves functional region localization accuracy. The TFG mechanism guides the selection and enhancement of attention heads in the image encoder using textual features, improving the understanding of affordance characteristics. Experimental results demonstrate that the proposed method achieves significant performance improvements on public AGD20K and HICO-IIF datasets. On the AGD20K dataset, compared with the current state-of-the-art method, we achieve a 6.0% improvement in the KLD metric, while reducing model parameters by 88.8%, demonstrating practical application values. The source code will be made publicly available at https://github.com/DAWDSE/BiT-Align.
Fan Yang 0063, Guoliang Zhu, Hao Shi 0004, Yukun Zuo, Wenrui Chen, Zhiyong Li 0001, Kailun Yang 0001
IROS2
2025 One-Shot Affordance Grounding of Deformable Objects in Egocentric Organizing Scenes
abstract
Deformable object manipulation in robotics presents significant challenges due to uncertainties in component properties, diverse configurations, visual interference, and ambiguous prompts. These factors complicate both perception and control tasks. To address these challenges, we propose a novel method for One-Shot Affordance Grounding of Deformable Objects (OS-AGDO) in egocentric organizing scenes, enabling robots to recognize previously unseen deformable objects with varying colors and shapes using minimal samples. Specifically, we first introduce the Deformable Object Semantic Enhancement Module (DefoSEM), which enhances hierarchical understanding of the internal structure and improves the ability to accurately identify local features, even under conditions of weak component information. Next, we propose the ORB-Enhanced Keypoint Fusion Module (OEKFM), which optimizes feature extraction of key components by leveraging geometric constraints and improves adaptability to diversity and visual interference. Additionally, we propose an instance-conditional prompt based on image data and task context, which effectively mitigates the issue of region ambiguity caused by prompt words. To validate these methods, we construct a diverse real-world dataset, AGDDO15, which includes 15 common types of deformable objects and their associated organizational actions. Experimental results demonstrate that our approach significantly outperforms state-of-the-art methods, achieving improvements of 6.2%, 3.2%, and 2.9% in KLD, SIM, and NSS metrics, respectively, while exhibiting high generalization performance. Source code and benchmark dataset are made publicly available at https://github.com/Dikay1/OS-AGDO.
Wanjun Jia, Fan Yang 0063, Mengfei Duan, Xianchi Chen, Yinxi Wang, Yiming Jiang 0001, Wenrui Chen, Kailun Yang 0001, Zhiyong Li 0001
IROS2
2025 Multi-scale pseudo-labels filtering and key pixels adversarial alignment for domain adaptive object detection
abstract
Domain Adaptive Object Detection (DAOD) aims to adapt a detector trained on the labeled source domain to perform effectively on the unlabeled target domain. Most existing DAOD methods commonly employ self-training strategy and adversarial learning. However, the methods based on the self-training strategy often generate unreliable pseudo-labels (e.g., missing detections and false positives) due to the absence of target domain semantics, resulting in suboptimal models. Meanwhile, aligning a large number of unimportant pixels hampers adversarial learning’s ability to capture domain-invariant semantic information. Therefore, we propose the Multi-scale Adversarial Teacher Framework (MATF) based on the common self-training framework (teacher–student framework). Specifically, to mitigate the impact of unreliable pseudo-labels, we propose the Multi-scale Negatives Filtering (MNF) module in the teacher model, which prevents missing detections from damaging the model by filtering out unreliable negative samples at multiple feature scales. In addition, we propose the Multi-scale Key-Pixel Discriminator (MKPD), which predicts the distributions of key pixels at multiple feature scales emphasizing the adversarial alignment of key pixels rather than unimportant pixels. Finally, to generate higher-quality pseudo-labels, we introduce the discriminator into the student model to reduce the model’s bias towards the source domain. Unlike most methods, we construct our method on a computationally efficient but less work-intensive one-stage detector. Extensive experiments conducted on DAOD benchmark datasets such as Cityscapes, FoggyCityscapes, KITTI, and Sim10k demonstrate the strong adaptability and effectiveness of MATF.
Xing Wei 0002, Jiong Xia, Fan Yang 0063, Cang Liu, Yanke Chen, Yang Lu 0015
Eng. Appl. Artif. Intell.3
2025 Consistent positive correlation sample distribution: Alleviating the negative sample noise issue in contrastive adaptation
Xing Wei 0002, Zelin Pan, Jiansheng Peng, Fan Yang 0063, Yang Lu 0015
Expert Syst. Appl.6
2025 Dual-Level Redundancy Elimination for Unsupervised Domain Adaptation
Dexuan Zhao, Fan Yang 0063, Taizhang Hu, Xing Wei 0002, Yang Lu 0015
Expert Syst. Appl.2
2025 Optimizing low-rank adaptation with decomposed matrices and adaptive rank allocation
Dacao Zhang, Fan Yang 0063, Kunwu Zhang, Xin Li 0082, Si Wei, Richang Hong, Meng Wang 0001
Frontiers Comput. Sci.2
2025 Task-Oriented Tool Manipulation With Robotic Dexterous Hands: A Knowledge Graph Approach From Fingers to Functionality
abstract
A primary challenge in robotic tool use is achieving precise manipulation with dexterous robotic hands to mimic human actions. It requires understanding human tool use and allocating specific functions to each robotic finger for fine control. Existing work has primarily focused on the overall grasping capabilities of robotic hands, often neglecting the functional allocation among individual fingers during object interaction. In response to this, we introduce a semantic knowledge-driven approach to distribute functions among fingers for tool manipulation. Central to this approach is the finger-to-function (F2F) knowledge graph, which captures human expertise in tool use and establishes relationships between tool attributes, tasks, and manipulation elements, including functional fingers, components, required force, and gestures. We also develop a manipulation element-oriented prediction algorithm using knowledge graph semantic embedding, enhancing the prediction of manipulation elements' speed and accuracy. Additionally, we propose the functionality-integrated adaptive force feedback manipulation (FAFM) module, which integrates manipulation elements with adaptive force feedback to achieve precise finger-level control. Our framework does not rely on extensive annotated data for supervision but utilizes semantic constraints from F2F to guide tool manipulation. The proposed method demonstrates superior performance and generalizability in real-world scenarios, achieving an 8% higher success rate in grasping and manipulation of representative tool instances compared to the existing state-of-the-art methods. The dataset and code are available at https://github.com/yangfan293/F2F.
Fan Yang 0063, Wenrui Chen, Sijie Wu, Xin Li 0082, Zhiyong Li 0001, Yaonan Wang 0001
IEEE Trans. Cybern.1
2025 Unsupervised domain adaptation via causal-contrastive learning
Xing Wei 0002, Fan Yang 0063, Yang Lu 0015, Benhong Zhang, Xiang Bi
J. Supercomput.3
2025 Proposal-level reliable feature-guided contrastive learning for SFOD
Xing Wei 0002, Jiong Xia, Cang Liu, Qi-wen He, Fan Yang 0063, Yang Lu 0015
J. Supercomput.8
2025 Unsupervised Domain Adaptation via Bidirectional Transmission Generator Self-Training
abstract
Unsupervised domain adaptation (UDA) aims to transfer knowledge from the labeled source domain to the fully unlabeled target domain, thus improving the classification performance of the target domain. Recently, self-training methods have shown their effectiveness on UDA. It iteratively trains target data using the generated target pseudo-labels. However, the feature space for generating pseudo-labels contains a large amount of source information, which traps the model in the source domain, making it challenging for the generator to learn discriminative features of the target domain. In this article, we propose a self-training domain adaptation (DA) model with bidirectional transmission generators (BDTGs). Specifically, we design a bidirectional transmission structure for generators, using exponential moving average (EMA) as the bridge between two generators. The structure has two advantages: 1) by transmitting weight parameters to each other during the training process, it promotes the shift of the feature space, thereby alleviating the difficulty of the model in adapting target domain features and 2) the transmission disturbs the classification boundary and is able to expose unreliable target samples near the boundary. We design a cosine similarity-based filter to identify such samples, to reduce the influence of noisy pseudo-labels with incorrect semantic information on the model. Extensive experiments conducted on five benchmark UDA datasets show that our approach has superior classification performance.
Xing Wei 0002, Zhaoxin Ji, Fan Yang 0063, Yang Lu 0015
IEEE Trans. Neural Networks Learn. Syst.3
2025 Learning Granularity-Aware Affordances From Human-Object Interaction for Tool-Based Functional Dexterous Grasping
abstract
To enable robots to use tools, the initial step is teaching robots to employ dexterous gestures for touching specific areas precisely where tasks are performed. Affordance features of objects serve as a bridge in the functional interaction between agents and objects. However, leveraging these affordance cues to help robots achieve functional tool grasping remains unresolved. To address this, we propose a granularity-aware affordance feature extraction method for locating functional affordance areas and predicting dexterous coarse gestures. We study the intrinsic mechanisms of human tool use. On the one hand, we use fine-grained affordance features of object-functional finger contact areas to locate functional affordance regions. On the other hand, we use highly activated coarse-grained affordance features in hand-object interaction regions to predict grasp gestures. Additionally, we introduce a model-based postprocessing module that transforms affordance localization and gesture prediction into executable robotic actions. This forms GAAF-Dex, a complete framework that learns granularity-aware affordances from human-object interaction to enable tool-based functional grasping with dexterous hands. Unlike fully supervised methods that require extensive data annotation, we employ a weakly supervised approach to extract relevant cues from exocentric (Exo) images of hand-object interactions to supervise feature extraction in egocentric (Ego) images. To support this approach, we have constructed a small-scale dataset, functional affordance hand (FAH)-object interaction dataset, which includes nearly 6k images of functional hand-object interaction Exo images and Ego images of 18 commonly used tools performing six tasks. Extensive experiments on the dataset demonstrate that our method outperforms state-of-the-art methods, and real-world localization and grasping experiments validate the practical applicability of our approach. The source code and the established dataset are available at https://github.com/yangfan293/GAAF-DEX.
Fan Yang 0063, Wenrui Chen, Kailun Yang 0001, Conghui Tang, Zhiyong Li 0001, Yaonan Wang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Self-Training Domain Adaptation Via Weight Transmission Between Generators
abstract
Unsupervised domain adaptation (UDA) aims to transfer knowledge from the labeled source domain to the fully-unlabeled target domain, thus improving the classification performance of the target domain. Recently, self-training has shown its effectiveness on UDA. However, the feature space for generating pseudo-labels contains a large amount of source information, making it challenging for the generator to learn discriminative features of the target domain. In this paper, we propose a self-training domain adaptation model via weight transmission between generators (WTBG). Specifically, we develop a bi-directional transmission structure for generators, using Exponential Moving Average (EMA) as the bridge between two generators. By cyclically transmitting weight parameters between them, alleviate the difficulty of generators in learning target features. And a pseudo-label filter based on cosine similarity is designed to reduce the influence of error pseudo-labels. Extensive experiments conducted on two benchmark UDA datasets show that WTBG has superior classification performance.
Xing Wei 0002, Zhaoxin Ji, Fan Yang 0063, Yang Lu 0015
ICASSP3
2024 Unsupervised Multi-Target Domain Adaptation Incremental Method Based on Contrastive Learning
abstract
Unsupervised domain adaptation (UDA) aims to alleviate the problem of distribution differences between unlabeled target domain and labeled source domain. Multi-target domain adaptation (MTDA) requires simultaneous transfer of knowledge from a single source domain to multiple target domains. Although the mixed-domain methods and some solutions that integrate single-target domain adaptation (STDA) optimization strategies have been proposed, the problem of knowledge forgetting has not been well addressed. To solve this problem, this paper proposes a multi-target domain adaptation incremental method based on contrastive learning. We construct a contrastive adaptation network and an incremental container network, which are connected by a feature discriminator. Contrastive learning ensures the adaptation effect of each source-target domain pair. Incremental module saves the weights of contrastive module, and the both networks combine feature discriminator to align the source features, alleviating the knowledge forgetting proplem. Extensive experiments on public datasets demonstrate that our method achieves excellent classification performance.
Xing Wei 0002, Zhaoxin Ji, Fan Yang 0063, Yang Lu 0015
ICME4
2024 ECCT: Efficient Contrastive Clustering via Pseudo-Siamese Vision Transformer and Multi-view Augmentation
Xing Wei 0002, Taizhang Hu, Fan Yang 0063, Yang Lu 0015
Neural Networks4
2024 DSIS-DPR:Structured Instance Segmentation and Diffusion Prior Refinement for Dental Anatomy Learning
abstract
Instance segmentation in medical imaging plays a crucial role in clinical diagnostic tasks, and have shown promising performance in practical applications. In this paper, we discuss a more fine-grained instance segmentation task: dental structured instance segmentation based on panoramic radiographs. However, direct segmentation of tooth structures encounters inherent challenges. Traditional instance segmentation networks often fall short in capturing intricate internal features, and exacerbated by the frequent blurring found in medical imaging, which can result in the deficiency of anatomical details. To deal with these problems, we propose a novel framework called DSISDPR, which combines a dental structured instance segmentation (DSIS) network with an enhanced diffusion prior refinement (DPR) method. Specifically, our innovatively designed structureaware network leverages fine-grained feature fusion, acquiring a richer representation of internal anatomical structures. With the integration of adversarial learning, the model is primed to deliver holistic and subtle predictions of tooth structures. Furthermore, taking inspiration from dentists’ inherent ability to utilize prior knowledge, such as understanding dental structures to label invisible anatomical structures, we propose a diffusion inpainting to refine the results of DSIS without additional annotations. Equipped with built-in structure learning, DPR is capable of modifying anomalies within each predicted segmentation, resulting in a more robust and complete structured segmentation result. Meanwhile, we ensure rigorous oversight over the reconstruction of areas affected by abnormalities, ensuring that any introduced adjustments minimally disrupt the wellpredicted structured segmentation results. Extensive experiments have demonstrated that our DSIS-DPR outperforms all existing classical instance segmentation networks. The collected dataset is available:https://github.com/Zzz512/TSD.
Xianyun Wang, Linhong Wang, Zhenchen Yang, Jiacong Zhou, Yuchen Zheng 0004, Richang Hong, Jun Yu 0002, Fan Yang 0063
IEEE Trans. Multim.9
2023 Task-oriented contrastive learning for unsupervised domain adaptation
Xing Wei 0002, Fan Yang 0063, Yujie Liu 0012
Expert Syst. Appl.3
2023 Multi-level uncertainty aware learning for semi-supervised dental panoramic caries segmentation
Xianyun Wang, Sizhe Gao, Kaisheng Jiang, Huicong Zhang, Linhong Wang, Jun Yu 0002, Fan Yang 0063
Neurocomputing8
2023 M3GAN: A masking strategy with a mutable filter for multidimensional anomaly detection
Yifan Li 0005, Ziyan Wu 0006, Fan Yang 0063, Zhiyong Li 0001
Knowl. Based Syst.4
2022 An Energy Harvesting Roadside Unit communication load prediction and energy scheduling based on graph convolutional neural networks for spatial-temporal vehicle data
abstract
Abstract The Energy Harvesting Roadside Unit (EH‐RSU) with self‐powered module will not only effectively reduce the communication load of regional Vehicular Ad Hoc Networks, but also enjoys a low deployment cost. Given the imbalance in communication demands invoked by transportation systems, the EH‐RSU should allocate energy appropriately in accordance with its energy harvesting rate to ensure the communication safety of vehicles within its coverage. Firstly, we propose a novel attention‐based spatial‐temporal graph convolutional network (ASTGCN) to predict the communication load around the EH‐RSU in the road network through the surrounding vehicle information. Secondly, we use the predicted communication load as part of the input parameters to neural network and leverage a double deep Q network to ameliorate the operating states switching strategy of EH‐RSUs by reinforcement learning so that they achieve a more satisfying effective time with limited resources. Finally, we built a dataset by simulation to validate the effectiveness of our model. The results show that our prediction model has a better accuracy and the improved strategy has higher efficiency compared with other methods.
Xu Ding 0001, Fan Yang 0063
IET Signal Process.5
2021 Optimal resource utilization for intra-cluster D2D retransmission and cooperative communications in VANETs
Fan Yang 0063, Jianghong Han, Xu Ding 0001
Wirel. Networks1
2020 HMM-Based Traffic State Prediction and Adaptive Routing Method in VANETs
Kaihan Gao, Xu Ding 0001, Juan Xu 0002, Fan Yang 0063
CollaborateCom (2)4
2020 A novel approach of dynamic base station switching strategy based on Markov decision process for interference alignment in VANETs
Jianghong Han, Xu Ding 0001, Fan Yang 0063
Wirel. Networks4
2018 Efficient Correlation Tracking via Center-Biased Spatial Regularization
abstract
Correlation filters (CFs) have been applied to visual tracking with success providing excellent performance in terms of accuracy and efficiency. The underlying periodic assumption of the training samples results in their great efficiency when using the fast Fourier transform (FFT), yet it also brings unwanted boundary effects. To address this issue, the recently proposed spatially-regularized discriminative CF (SRDCF) method introduces a Gaussian weight function to regularize the learning filter, yielding favorable performances in accuracy but high computational complexity because the objective of the SRDCF cannot achieve a closed solution via the FFT. Motivated by SRDCF, we present an efficient and effective CF-based tracker using center-biased constraint weights (CBCWs), which improve simultaneously speed and accuracy. Specifically, we first construct a CBCW function by exploiting the symmetry of the Fourier transform. The values of the constraint weights are real in both time and frequency domains, so that the optimization can be directly solved in the frequency domain without any data transformation, thereby greatly reducing its computational complexity. Moreover, according to the average peak-tocorrelation energy value of the CF response, we propose an efficient and effective filter update strategy to handle occlusions during tracking. Extensive experiments on the OTB-2013, OTB- 2015, and VOT2016 benchmarks demonstrate that the proposed tracker significantly outperforms the baseline SRDCF in terms of accuracy and efficiency. Moreover, the proposed method performs favorably against 16 other representative state-of-the-art methods regarding robustness and success rate.
Jianghong Han, Fan Yang 0063, Kaihua Zhang 0001, Richang Hong
IEEE Trans. Image Process.3