Yifei Gong

dblp:238/2819 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Interactive Incremental Defect Detection Framework With Macroprobability-Controlled Adaptive Learning
abstract
Visual surface defect detection plays a crucial role in industrial quality control. A series of deep learning based algorithms have been introduced into visual defect detection, achieving remarkable performance. However, these algorithms generally suffer from inadequate adaptability to the expanding of detection categories, which limits their detection capabilities when dealing with complex and dynamic real-world application scenarios. To address this issue, this paper proposes an innovative incremental defect detection model. This model is based on learnable feature fusion and dynamic category modeling, aiming to enhance the model’s online learning and adaptability to new defect categories. To effectively utilize the existing data and optimize the model training process, this study introduces a macrodata probability control strategy to guide the reasonable configuration of training data as well as the model. Furthermore, to improve the performance during the detection process, an interactive feedback mechanism is constructed. This mechanism provides effective data in real-time during the detection process, driving the model to undergo dynamic training and gradually enhancing its recognition and adaptability to newly emerging defect types. To verify the practicality and effectiveness of the proposed algorithm, comprehensive comparative experiments were conducted on multiple datasets and a real-world detection platform was established for validation. The experimental results demonstrate the superior performance and dynamic adaptability in practical applications.
Bin Gao 0003, Yifei Gong, Yukuan Kang, Wai Lok Woo
IEEE Trans. Ind. Informatics3
2024 Drug-target binding affinity prediction model based on multi-scale diffusion and interactive learning
Zhiqin Zhu, Guanqiu Qi, Yifei Gong, Neal Mazur, Baisen Cong, Xinbo Gao 0001
Expert Syst. Appl.4
2024 Brain tumour segmentation framework with deep nuanced reasoning and Swin-T
abstract
Abstract Tumour medical image segmentation plays a crucial role in clinical imaging diagnosis. Existing research has achieved good results, enabling the segmentation of three tumour regions in MRI brain tumour images. Existing models have limited focus on the brain tumour areas, and the long‐term dependency of features is weakened as the network depth increases, resulting in blurred edge segmentation of the targets. Additionally, considering the excellent segmentation performance of the Swin Transformer(Swin‐T) network, its network structure and parameters are relatively large. To address these limitations, this paper proposes a brain tumour segmentation framework with deep nuanced reasoning and Swin‐T. It is mainly composed of the backbone hybrid network (BHN) and the deep micro texture extraction module (DMTE). The BHN combines the Swin‐T stage with a new downsampling transition module called dual path feature reasoning (DPFR). The entire network framework is designed to extract global and local features from multi‐modal data, enabling it to capture and analyze deep texture features in multi‐modal images. It provides significant optimization over the Swin‐T network structure. Experimental results on the BraTS dataset demonstrate that the proposed method outperforms other state‐of‐the‐art models in terms of segmentation performance. The corresponding source codes are available at https://github.com/CurbUni/Brain‐Tumor‐Segmentation‐Framework‐with‐Deep‐Nuanced‐Reasoning‐and‐Swin‐T .
Guanqiu Qi, Yifei Gong, Xiaolong Qu, Li Yin 0011
IET Image Process.4
2023 CARESim: An integrated agent-based simulation environment for crime analysis and risk evaluation (CARE)
Yifei Gong, Mengyan Dai, Feng Gu 0001
Expert Syst. Appl.1
2022 Conditional Feature Learning Based Transformer for Text-Based Person Search
abstract
Text-based person search aims at retrieving the target person in an image gallery using a descriptive sentence of that person. The core of this task is to calculate a similarity score between the pedestrian image and description, which requires inferring the complex latent correspondence between image sub-regions and textual phrases at different scales. Transformer is an intuitive way to model the complex alignment by its self-attention mechanism. Most previous Transformer-based methods simply concatenate image region features and text features as input and learn a cross-modal representation in a brute force manner. Such weakly supervised learning approaches fail to explicitly build alignment between image region features and text features, causing an inferior feature distribution. In this paper, we present CFLT, Conditional Feature Learning based Transformer. It maps the sub-regions and phrases into a unified latent space and explicitly aligns them by constructing conditional embeddings where the feature of data from one modality is dynamically adjusted based on the data from the other modality. The output of our CFLT is a set of similarity scores for each sub-region or phrase rather than a cross-modal representation. Furthermore, we propose a simple and effective multi-modal re-ranking method named Re-ranking scheme by Visual Conditional Feature (RVCF). Benefit from the visual conditional feature and better feature distribution in our CFLT, the proposed RVCF achieves significant performance improvement. Experimental results show that our CFLT outperforms the state-of-the-art methods by 7.03% in terms of top-1 accuracy and 5.01% in terms of top-5 accuracy on the text-based person search dataset.
Chenyang Gao, Guanyu Cai, Xinyang Jiang, Feng Zheng 0001, Jun Zhang 0018, Yifei Gong, Fangzhou Lin, Xing Sun 0001, Xiang Bai
IEEE Trans. Image Process.6
2022 Conditional Feature Embedding by Visual Clue Correspondence Graph for Person Re-Identification
abstract
Although Person Re-Identification has made impressive progress, difficult cases like occlusion, change of view-point, and similar clothing still bring great challenges. In order to tackle these challenges, extracting discriminative feature representation is crucial. Most of the existing methods focus on extracting ReID features from individual images separately. However, when matching two images, we propose that the ReID features of a query image should be dynamically adjusted based on the contextual information from the gallery image it matches. We call this type of ReID features conditional feature embedding. In this paper, we propose a novel ReID framework that extracts conditional feature embedding based on the aligned visual clues between image pairs, called Clue Alignment based Conditional Embedding (CACE-Net). CACE-Net applies an attention module to build a detailed correspondence graph between crucial visual clues in image pairs and uses discrepancy-based GCN to embed the obtained complex correspondence information into the conditional features. The experiments show that CACE-Net achieves state-of-the-art performance on three public datasets.
Fufu Yu, Xinyang Jiang, Yifei Gong, Wei-Shi Zheng 0001, Feng Zheng 0001, Xing Sun 0001
IEEE Trans. Image Process.3
2021 Ask&Confirm: Active Detail Enriching for Cross-Modal Retrieval with Partial Query
abstract
Text-based image retrieval has seen considerable progress in recent years. However, the performance of existing methods suffers in real life since the user is likely to provide an incomplete description of an image, which often leads to results filled with false positives that fit the incomplete description. In this work, we introduce the partial-query problem and extensively analyze its influence on text-based image retrieval. Previous interactive methods tackle the problem by passively receiving users’ feedback to supplement the incomplete query iteratively, which is time-consuming and requires heavy user effort. Instead, we propose a novel retrieval framework that conducts the interactive process in an Ask-and-Confirm fashion, where AI actively searches for discriminative details missing in the current query, and users only need to confirm AI’s proposal. Specifically, we propose an object-based interaction to make the interactive retrieval more user-friendly and present a reinforcement-learning-based policy to search for discriminative objects. Furthermore, since fully-supervised training is often infeasible due to the difficulty of obtaining human-machine dialog data, we present a weakly-supervised training strategy that needs no human-annotated dialogs other than a text-image dataset. Experiments show that our framework significantly improves the performance of text-based image retrieval. Code is available at https://github.com/CuthbertCai/Ask-Confirm.
Guanyu Cai, Jun Zhang 0018, Xinyang Jiang, Yifei Gong, Lianghua He, Fufu Yu, Feiyue Huang, Xing Sun 0001
ICCV4
2021 Learning Canonical View Representation for 3D Shape Recognition with Arbitrary Views
abstract
In this paper, we focus on recognizing 3D shapes from arbitrary views, i.e., arbitrary numbers and positions of viewpoints. It is a challenging and realistic setting for view-based 3D shape recognition. We propose a canonical view representation to tackle this challenge. We first transform the original features of arbitrary views to a fixed number of view features, dubbed canonical view representation, by aligning the arbitrary view features to a set of learnable reference view features using optimal transport. In this way, each 3D shape with arbitrary views is represented by a fixed number of canonical view features, which are further aggregated to generate a rich and robust 3D shape representation for shape recognition. We also propose a canonical view feature separation constraint to enforce that the view features in canonical view representation can be embedded into scattered points in a Euclidean space. Experiments on the ModelNet40, ScanObjectNN, and RGBD datasets show that our method achieves competitive results under the fixed viewpoint settings, and significantly outperforms the applicable methods under the arbitrary view setting.
Yifei Gong, Fudong Wang 0001, Xing Sun 0001, Jian Sun 0009
ICCV2
2021 Learning with Instance-Dependent Label Noise: A Sample Sieve Approach
Hao Cheng 0012, Zhaowei Zhu, Yifei Gong, Xing Sun 0001, Yang Liu 0018
ICLR4
2021 Bayesian Optimization with Particle Swarm
abstract
For most machine learning models, the mapping from the hyper-parameter set to the model's generalization error can be regarded as a complex black box function. Particle swarm optimization (PSO) methods cannot be directly used in the problem of hyper-parameters estimation since the mathematical formulation of the mapping from hyper-parameters to loss function or generalization accuracy is unclear. Functions with high evaluation costs can be solved by Bayesian optimization (BO) which converting the optimization of hyper-parameters into the optimization of an acquisition function. The proposed method in this paper uses the particle swarm method to optimize the acquisition function in the BO to get better hyper-parameters. The performances of proposed method in both of the classification and regression models are evaluated and demonstrated.
Yulai Zhang, Gongxue Zhou, Yifei Gong
IJCNN4
2021 Nested Causality Extraction on Traffic Accident Texts as Question Answering
Gongxue Zhou, Weifeng Ma, Yifei Gong, Liudi Wang, Yulai Zhang
NLPCC (2)3
2020 Rethinking Temporal Fusion for Video-Based Person Re-Identification on Semantic and Time Aspect
abstract
Recently, the research interest of person re-identification (ReID) has gradually turned to video-based methods, which acquire a person representation by aggregating frame features of an entire video. However, existing video-based ReID methods do not consider the semantic difference brought by the outputs of different network stages, which potentially compromises the information richness of the person features. Furthermore, traditional methods ignore important relationship among frames, which causes information redundancy in fusion along the time axis. To address these issues, we propose a novel general temporal fusion framework to aggregate frame features on both semantic aspect and time aspect. As for the semantic aspect, a multi-stage fusion network is explored to fuse richer frame features at multiple semantic levels, which can effectively reduce the information loss caused by the traditional single-stage fusion. While, for the time axis, the existing intra-frame attention method is improved by adding a novel inter-frame attention module, which effectively reduces the information redundancy in temporal fusion by taking the relationship among frames into consideration. The experimental results show that our approach can effectively improve the video-based re-identification accuracy, achieving the state-of-the-art performance.
Xinyang Jiang, Yifei Gong, Qize Yang, Feiyue Huang, Wei-Shi Zheng 0001, Feng Zheng 0001, Xing Sun 0001
AAAI2