Ying Wei 0007

dblp:14/4899-7 · DBLP profile ↗
← Back
25ranked-venue papers
0as first author
22since 2021 · last 2026
0000-0003-0915-5378ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ESMT: Context-adaptive vision-language tracking with episodic-semantic memory
Ying Wei 0007, Gang Yang 0002
Appl. Intell.2
2026 Explicit-Implicit Prompt Injection and Semantic-Guided Latent LoRA for Vision-Language Tracking
abstract
Prompt-based learning has shown promise in visual-language tracking (VLT), yet existing methods often rely on either explicit or implicit prompting alone, limiting fine-grained cross-modal alignment. Moreover, Low-Rank Adaptation (LoRA) -based fine-tuning in prior work typically focuses on visual-only adaptation, overlooking language semantics. To address these issues, we propose a unified VLT framework that integrates Explicit-Implicit Prompt Injection (EIPI) and Semantic-Guided Latent LoRA (SGLL). EIPI introduces semantic prompts to facilitate robust and context-sensitive target modeling through two pathways. The explicit prompts are constructed by interact between multi-modal target representations with the search region, while implicit prompts are learned from linguistic features via a lightweight bottleneck network. Then, SGLL extends standard LoRA by introducing learnable queries in the latent space, allowing residual modulation based on language-visual semantics without retraining the full model. This dual design yields a parameter-efficient tracker with strong cross-modal adaptability. Extensive experiments show our method outperforms prior prompt-based approaches while maintaining high efficiency.
Ying Wei 0007, Gang Yang 0002, Qiaohong Hao
IEEE Signal Process. Lett.2
2025 Deep spiking neural networks based on model fusion technology for remote sensing image classification
Li-Ye Niu, Ying Wei 0007, Liping Zhao 0005, Keli Hu
Eng. Appl. Artif. Intell.2
2025 ESDA: Zero-shot semantic segmentation based on an embedding semantic space distribution adjustment strategy
Jiaguang Li, Ying Wei 0007, Chuyuan Wang
Image Vis. Comput.2
2025 SKDF: A Simple Knowledge Distillation Framework for Distilling Open-Vocabulary Knowledge to Open-World Object Detector
abstract
Open World Object Detection (OWOD) is a novel computer vision task with a considerable challenge, bridging the gap between classic object detection (OD) and real-world object detection. In addition to detecting and classifying seen/known objects, OWOD algorithms are expected to localize all potential unseen/unknown objects and incrementally learn them. The large pre-trained vision-language grounding models (VLM, e.g., GLIP) have rich knowledge about the open world, but are limited by text prompts and cannot localize indescribable objects. However, there are many detection scenarios in which pre-defined language descriptions are unavailable during inference. In this paper, we attempt to specialize the VLM model for OWOD tasks by distilling its open-world knowledge into a language-agnostic detector. Surprisingly, we observe that the simple knowledge distillation approach leads to unexpected performance for unknown object detection, even with a small amount of data. Unfortunately, knowledge distillation for unknown objects severely affects the learning of detectors with conventional structures, leading to catastrophic damage to the model's ability to learn about known objects. To alleviate these problems, we propose the down-weight training strategy for knowledge distillation from vision-language model to single visual modality one. Meanwhile, we propose the cascade decoupled decoders that decouple the learning of localization and recognition to reduce the impact of category interactions of known and unknown objects on the localization learning process. Ablation experiments demonstrate that both of them are effective in mitigating the impact of open-world knowledge distillation on the learning of known objects. Additionally, to alleviate the current lack of comprehensive benchmarks for evaluating the ability of the open-world detector to detect unknown objects in the open world, we refine the benchmark for evaluating the performance of unknown object detection by augmenting annotations for unknown objects which we name"IntensiveSet$\scriptstyle\spadesuit$♠". Comprehensive experiments performed on OWOD, MS-COCO, and our proposed benchmarks demonstrate the effectiveness of our methods.
Shuailei Ma, Ying Wei 0007, Enming Zhang, Peihao Chen
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 EFTNet: an efficient fine-tuning method for few-shot segmentation
Jiaguang Li, Ying Wei 0007
Appl. Intell.4
2024 FGAHOI: Fine-Grained Anchors for Human-Object Interaction Detection
abstract
Human-Object Interaction (HOI), as an important problem in computer vision, requires locating the human-object pair and identifying the interactive relationships between them. The HOI instance has a greater span in spatial, scale, and task than the individual object instance, making its detection more susceptible to noisy backgrounds. To alleviate the disturbance of noisy backgrounds on HOI detection, it is necessary to consider the input image information to generate fine-grained anchors which are then leveraged to guide the detection of HOI instances. However, it has the following challenges. i) how to extract pivotal features from the images with complex background information is still an open question. ii) how to semantically align the extracted features and query embeddings is also a difficult issue. In this paper, a novel end-to-end transformer-based framework (FGAHOI) is proposed to alleviate the above problems. FGAHOI comprises three dedicated components namely, multi-scale sampling (MSS), hierarchical spatial-aware merging (HSAM) and task-aware merging mechanism (TAM). MSS extracts features of humans, objects and interaction areas from noisy backgrounds for HOI instances of various scales. HSAM and TAM semantically align and merge the extracted features and query embeddings in the hierarchical spatial and task perspectives in turn. In the meanwhile, a novel training strategy Stage-wise Training Strategy is designed to reduce the training pressure caused by overly complex tasks done by FGAHOI. In addition, we propose two ways to measure the difficulty of HOI detection and a novel dataset, i.e., HOI-SDC for the two challenges (Uneven Distributed Area in Human-Object Pairs and Long Distance Visual Modeling of Human-Object Pairs) of HOI instances detection. Experiments are conducted on three benchmarks: HICO-DET, HOI-SDC and V-COCO. Our model outperforms the state-of-the-art HOI detection methods, and the extensive ablations reveal the merits of our proposed contribution.
Shuailei Ma, Shanze Wang, Ying Wei 0007
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 DDOWOD: DiffusionDet for open-world object detection
Enming Zhang, Ying Wei 0007, Jiakun Xia, Xinghong Liu, Shuailei Ma
Pattern Recognit. Lett.3
2023 CAT: LoCalization and IdentificAtion Cascade Detection Transformer for Open-World Object Detection
abstract
Open-world object detection (OWOD), as a more general and challenging goal, requires the model trained from data on known objects to detect both known and unknown objects and incrementally learn to identify these unknown objects. The existing works which employ standard detection framework and fixed pseudo-labelling mechanism$(PLM)$have the following problems: (i) The inclusion of detecting unknown objects substantially reduces the model's ability to detect known ones. (ii) The$PLM$does not adequately utilize the priori knowledge of inputs. (iii) The fixed selection manner of$PLM$cannot guarantee that the model is trained in the right direction. We observe that humans subconsciously prefer to focus on all foreground objects and then identify each one in detail, rather than localize and identify a single object simultaneously, for alleviating the confusion. This motivates us to propose a novel solution called CAT: LoCalization and IdentificAtion Cascade Detection Transformer which decouples the detection process via the shared decoder in the cascade decoding way. In the meanwhile, we propose the self-adaptive pseudo-labelling mechanism which combines the model-driven with input-driven$PLM$and self-adaptively generates robust pseudo-labels for unknown objects, significantly improving the ability of CAT to retrieve unknown objects. Experiments on two benchmarks, i.e., MS-COCO and PASCAL VOC, show that our model outperforms the state-of-the-art methods. The code is publicly available at https://github.com/xiaomabufei/CAT.
Shuailei Ma, Ying Wei 0007, Thomas H. Li, Fanbing Lv
CVPR3
2023 SVF-Net: spatial and visual feature enhancement network for brain structure segmentation
Ying Wei 0007, Xiang Li 0059, Chuyuan Wang, Shanze Wang
Appl. Intell.2
2023 Contextual-wise discriminative feature extraction and robust network learning for subcortical structure segmentation
Xiang Li 0059, Ying Wei 0007, Chuyuan Wang, Chengan Liu
Appl. Intell.2
2023 Discriminative-region attention and orthogonal-view generation model for vehicle re-identification
Ying Wei 0007, Ge Li 0002
Appl. Intell.3
2023 Research Progress of spiking neural network in image classification: a review
Li-Ye Niu, Ying Wei 0007, Wen-Bo Liu, Jun-Yu Long, Tian-hao Xue
Appl. Intell.2
2023 Event-driven spiking neural network based on membrane potential modulation for remote sensing image classification
Li-Ye Niu, Ying Wei 0007, Yue Liu 0003
Eng. Appl. Artif. Intell.2
2023 CIRM-SNN: Certainty Interval Reset Mechanism Spiking Neuron for Enabling High Accuracy Spiking Neural Network
Li-Ye Niu, Ying Wei 0007
Neural Process. Lett.2
2023 Discriminative Deep Non-Linear Dictionary Learning for Visual Object Tracking
Ying Wei 0007, Shengxing Shang
Neural Process. Lett.2
2022 Video-based vehicle re-identification via channel decomposition saliency region network
Benhua Gong, Ying Wei 0007, Ruipeng Ma
Appl. Intell.3
2022 Non-linear target trajectory prediction for robust visual tracking
Zhaofu Diao, Ying Wei 0007
Appl. Intell.3
2022 High-Accuracy Spiking Neural Network for Objective Recognition Based on Proportional Attenuating Neuron
Li-Ye Niu, Ying Wei 0007, Jun-Yu Long, Wen-Bo Liu
Neural Process. Lett.2
2021 MSGSE-Net: Multi-scale guided squeeze-and-excitation network for subcortical brain structure segmentation
Xiang Li 0059, Ying Wei 0007, Shidi Fu, Chuyuan Wang
Neurocomputing2
2021 Wasserstein Distance-Based Auto-Encoder Tracking
Ying Wei 0007, Chenhe Dong, Chuaqiao Xu, Zhaofu Diao
Neural Process. Lett.2
2021 Multi-Site Infant Brain Segmentation Algorithms: The iSeg-2019 Challenge
abstract
To better understand early brain development in health and disorder, it is critical to accurately segment infant brain magnetic resonance (MR) images into white matter (WM), gray matter (GM), and cerebrospinal fluid (CSF). Deep learning-based methods have achieved state-of-the-art performance; h owever, one of the major limitations is that the learning-based methods may suffer from the multi-site issue, that is, the models trained on a dataset from one site may not be applicable to the datasets acquired from other sites with different imaging protocols/scanners. To promote methodological development in the community, the iSeg-2019 challenge (http://iseg2019.web.unc.edu) provides a set of 6-month infant subjects from multiple sites with different protocols/scanners for the participating methods. T raining/validation subjects are from UNC (MAP) and testing subjects are from UNC/UMN (BCP), Stanford University, and Emory University. By the time of writing, there are 30 automatic segmentation methods participated in the iSeg-2019. In this article, 8 top-ranked methods were reviewed by detailing their pipelines/implementations, presenting experimental results, and evaluating performance across different sites in terms of whole brain, regions of interest, and gyral landmark curves. We further pointed out their limitations and possible directions for addressing the multi-site issue. We find that multi-site consistency is still an open issue. We hope that the multi-site dataset in the iSeg-2019 and this review article will attract more researchers to address the challenging and critical multi-site issue in practice.
Yue Sun 0001, Kun Gao 0002, Zhengwang Wu, Xiaopeng Zong, Zhihao Lei, Ying Wei 0007, Jun Ma 0016, Xiaoping Yang 0001, Xue Feng 0001, Li Zhao 0001, Trung Le Phan, Jitae Shin, Tao Zhong 0002, Yu Zhang 0064, Lequan Yu, Caizi Li, Ramesh Basnet, M. Omair Ahmad, M. N. S. Swamy 0001, Wenao Ma, Qi Dou 0001, Toan Duc Bui, Camilo Bermudez, Bennett A. Landman, Ian H. Gotlib, Kathryn L. Humphreys, Sarah Shultz, Longchuan Li, Sijie Niu, Weili Lin, Valerie Jewells, Dinggang Shen, Gang Li 0001, Li Wang 0026
IEEE Trans. Medical Imaging7
2020 Scale-Invariant Siamese Network For Person Re-Identification
abstract
Most existing methods for person re-identification (ReID) almost match people at a single scale and ignore that people are often distinguishable at the right spatial locations and scales. Unlike previous works designing complex convolutional neural network (CNN) architecture or concatenating multi-branch scale-specific features, we aim to employ a simple network to learn scale-invariant features. Concretely, we first propose a shared two-branch framework with two-scale images from the same identity as inputs, which is beneficial for ReID network to focus on common features in different-scale images. Furthermore, we introduce a novel attention loss to enforce discriminative regions between two branches more consistent in the visual level. Finally, we conduct extensive evaluations on three largescale datasets and report competitive performance.
Yunzhou Zhang, Shuangwei Liu, Jining Bao, Ying Wei 0007
ICIP5
2020 Vehicle re-identification based on unsupervised local area detection and view discrimination
Ying Wei 0007, Chuyuan Wang
Image Vis. Comput.3
2019 Combination of Appearance and License Plate Features for Vehicle Re-Identification
abstract
In this work, we propose a two-module framework that combines appearance and corresponding license plate features for vehicle re-identification (Re-ID). In the appearance module, we design a Two-Branch Network to extract comprehensive global features. To obtain more discriminative feature representations, we propose an enhanced triplet loss (ETL) and combine ETL with softmax loss to optimize the parameter of Two-Branch Network. In the license plate module, we present a license plate Re-ID network that incorporates the bidirectional LSTMs into CNNs, which is effective for capturing the contexts in license plate images and significantly improves the performance of license plate Re-ID. We validate our method on both VeRi-776 dataset [1] and VehicleID dataset [2]. The experimental results show that our method outperforms most state-of-the-art approaches for vehicle Re-ID, even if only the appearance module is used.
Yanguang He, Chenhe Dong, Ying Wei 0007
ICIP3