VLDB 2026 Research / reviewers in the wild / expert
Ying Wei 0007
dblp:14/4899-7
· DBLP profile ↗
25ranked-venue papers
0as first author
22since 2021 · last 2026
0000-0003-0915-5378ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ESMT: Context-adaptive vision-language tracking with episodic-semantic memory
Ying Wei 0007, Gang Yang 0002 |
Appl. Intell. | 2 |
| 2026 | Explicit-Implicit Prompt Injection and Semantic-Guided Latent LoRA for Vision-Language TrackingabstractPrompt-based learning has shown promise in visual-language tracking (VLT), yet existing methods often rely on either explicit or implicit prompting alone, limiting fine-grained cross-modal alignment. Moreover, Low-Rank Adaptation (LoRA) -based fine-tuning in prior work typically focuses on visual-only adaptation, overlooking language semantics. To address these issues, we propose a unified VLT framework that integrates Explicit-Implicit Prompt Injection (EIPI) and Semantic-Guided Latent LoRA (SGLL). EIPI introduces semantic prompts to facilitate robust and context-sensitive target modeling through two pathways. The explicit prompts are constructed by interact between multi-modal target representations with the search region, while implicit prompts are learned from linguistic features via a lightweight bottleneck network. Then, SGLL extends standard LoRA by introducing learnable queries in the latent space, allowing residual modulation based on language-visual semantics without retraining the full model. This dual design yields a parameter-efficient tracker with strong cross-modal adaptability. Extensive experiments show our method outperforms prior prompt-based approaches while maintaining high efficiency. Ying Wei 0007, Gang Yang 0002, Qiaohong Hao |
IEEE Signal Process. Lett. | 2 |
| 2025 | Deep spiking neural networks based on model fusion technology for remote sensing image classification
Li-Ye Niu, Ying Wei 0007, Liping Zhao 0005, Keli Hu |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | ESDA: Zero-shot semantic segmentation based on an embedding semantic space distribution adjustment strategy
Jiaguang Li, Ying Wei 0007, Chuyuan Wang |
Image Vis. Comput. | 2 |
| 2025 | SKDF: A Simple Knowledge Distillation Framework for Distilling Open-Vocabulary Knowledge to Open-World Object DetectorabstractOpen World Object Detection (OWOD) is a novel computer vision task with a considerable challenge, bridging the gap between classic object detection (OD) and real-world object detection. In addition to detecting and classifying seen/known objects, OWOD algorithms are expected to localize all potential unseen/unknown objects and incrementally learn them. The large pre-trained vision-language grounding models (VLM, e.g., GLIP) have rich knowledge about the open world, but are limited by text prompts and cannot localize indescribable objects. However, there are many detection scenarios in which pre-defined language descriptions are unavailable during inference. In this paper, we attempt to specialize the VLM model for OWOD tasks by distilling its open-world knowledge into a language-agnostic detector. Surprisingly, we observe that the simple knowledge distillation approach leads to unexpected performance for unknown object detection, even with a small amount of data. Unfortunately, knowledge distillation for unknown objects severely affects the learning of detectors with conventional structures, leading to catastrophic damage to the model's ability to learn about known objects. To alleviate these problems, we propose the down-weight training strategy for knowledge distillation from vision-language model to single visual modality one. Meanwhile, we propose the cascade decoupled decoders that decouple the learning of localization and recognition to reduce the impact of category interactions of known and unknown objects on the localization learning process. Ablation experiments demonstrate that both of them are effective in mitigating the impact of open-world knowledge distillation on the learning of known objects. Additionally, to alleviate the current lack of comprehensive benchmarks for evaluating the ability of the open-world detector to detect unknown objects in the open world, we refine the benchmark for evaluating the performance of unknown object detection by augmenting annotations for unknown objects which we name"IntensiveSet$\scriptstyle\spadesuit$♠". Comprehensive experiments performed on OWOD, MS-COCO, and our proposed benchmarks demonstrate the effectiveness of our methods. Shuailei Ma, Ying Wei 0007, Enming Zhang, Peihao Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | EFTNet: an efficient fine-tuning method for few-shot segmentation
Jiaguang Li, Ying Wei 0007 |
Appl. Intell. | 4 |
| 2024 | FGAHOI: Fine-Grained Anchors for Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI), as an important problem in computer vision, requires locating the human-object pair and identifying the interactive relationships between them. The HOI instance has a greater span in spatial, scale, and task than the individual object instance, making its detection more susceptible to noisy backgrounds. To alleviate the disturbance of noisy backgrounds on HOI detection, it is necessary to consider the input image information to generate fine-grained anchors which are then leveraged to guide the detection of HOI instances. However, it has the following challenges. i) how to extract pivotal features from the images with complex background information is still an open question. ii) how to semantically align the extracted features and query embeddings is also a difficult issue. In this paper, a novel end-to-end transformer-based framework (FGAHOI) is proposed to alleviate the above problems. FGAHOI comprises three dedicated components namely, multi-scale sampling (MSS), hierarchical spatial-aware merging (HSAM) and task-aware merging mechanism (TAM). MSS extracts features of humans, objects and interaction areas from noisy backgrounds for HOI instances of various scales. HSAM and TAM semantically align and merge the extracted features and query embeddings in the hierarchical spatial and task perspectives in turn. In the meanwhile, a novel training strategy Stage-wise Training Strategy is designed to reduce the training pressure caused by overly complex tasks done by FGAHOI. In addition, we propose two ways to measure the difficulty of HOI detection and a novel dataset, i.e., HOI-SDC for the two challenges (Uneven Distributed Area in Human-Object Pairs and Long Distance Visual Modeling of Human-Object Pairs) of HOI instances detection. Experiments are conducted on three benchmarks: HICO-DET, HOI-SDC and V-COCO. Our model outperforms the state-of-the-art HOI detection methods, and the extensive ablations reveal the merits of our proposed contribution. Shuailei Ma, Shanze Wang, Ying Wei 0007 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | DDOWOD: DiffusionDet for open-world object detection
Enming Zhang, Ying Wei 0007, Jiakun Xia, Xinghong Liu, Shuailei Ma |
Pattern Recognit. Lett. | 3 |
| 2023 | CAT: LoCalization and IdentificAtion Cascade Detection Transformer for Open-World Object DetectionabstractOpen-world object detection (OWOD), as a more general and challenging goal, requires the model trained from data on known objects to detect both known and unknown objects and incrementally learn to identify these unknown objects. The existing works which employ standard detection framework and fixed pseudo-labelling mechanism$(PLM)$have the following problems: (i) The inclusion of detecting unknown objects substantially reduces the model's ability to detect known ones. (ii) The$PLM$does not adequately utilize the priori knowledge of inputs. (iii) The fixed selection manner of$PLM$cannot guarantee that the model is trained in the right direction. We observe that humans subconsciously prefer to focus on all foreground objects and then identify each one in detail, rather than localize and identify a single object simultaneously, for alleviating the confusion. This motivates us to propose a novel solution called CAT: LoCalization and IdentificAtion Cascade Detection Transformer which decouples the detection process via the shared decoder in the cascade decoding way. In the meanwhile, we propose the self-adaptive pseudo-labelling mechanism which combines the model-driven with input-driven$PLM$and self-adaptively generates robust pseudo-labels for unknown objects, significantly improving the ability of CAT to retrieve unknown objects. Experiments on two benchmarks, i.e., MS-COCO and PASCAL VOC, show that our model outperforms the state-of-the-art methods. The code is publicly available at https://github.com/xiaomabufei/CAT. Shuailei Ma, Ying Wei 0007, Thomas H. Li, Fanbing Lv |
CVPR | 3 |
| 2023 | SVF-Net: spatial and visual feature enhancement network for brain structure segmentation
Ying Wei 0007, Xiang Li 0059, Chuyuan Wang, Shanze Wang |
Appl. Intell. | 2 |
| 2023 | Contextual-wise discriminative feature extraction and robust network learning for subcortical structure segmentation
Xiang Li 0059, Ying Wei 0007, Chuyuan Wang, Chengan Liu |
Appl. Intell. | 2 |
| 2023 | Discriminative-region attention and orthogonal-view generation model for vehicle re-identification
Ying Wei 0007, Ge Li 0002 |
Appl. Intell. | 3 |
| 2023 | Research Progress of spiking neural network in image classification: a review
Li-Ye Niu, Ying Wei 0007, Wen-Bo Liu, Jun-Yu Long, Tian-hao Xue |
Appl. Intell. | 2 |
| 2023 | Event-driven spiking neural network based on membrane potential modulation for remote sensing image classification
Li-Ye Niu, Ying Wei 0007, Yue Liu 0003 |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | CIRM-SNN: Certainty Interval Reset Mechanism Spiking Neuron for Enabling High Accuracy Spiking Neural Network
Li-Ye Niu, Ying Wei 0007 |
Neural Process. Lett. | 2 |
| 2023 | Discriminative Deep Non-Linear Dictionary Learning for Visual Object Tracking
Ying Wei 0007, Shengxing Shang |
Neural Process. Lett. | 2 |
| 2022 | Video-based vehicle re-identification via channel decomposition saliency region network
Benhua Gong, Ying Wei 0007, Ruipeng Ma |
Appl. Intell. | 3 |
| 2022 | Non-linear target trajectory prediction for robust visual tracking
Zhaofu Diao, Ying Wei 0007 |
Appl. Intell. | 3 |
| 2022 | High-Accuracy Spiking Neural Network for Objective Recognition Based on Proportional Attenuating Neuron
Li-Ye Niu, Ying Wei 0007, Jun-Yu Long, Wen-Bo Liu |
Neural Process. Lett. | 2 |
| 2021 | MSGSE-Net: Multi-scale guided squeeze-and-excitation network for subcortical brain structure segmentation
Xiang Li 0059, Ying Wei 0007, Shidi Fu, Chuyuan Wang |
Neurocomputing | 2 |
| 2021 | Wasserstein Distance-Based Auto-Encoder Tracking
Ying Wei 0007, Chenhe Dong, Chuaqiao Xu, Zhaofu Diao |
Neural Process. Lett. | 2 |
| 2021 | Multi-Site Infant Brain Segmentation Algorithms: The iSeg-2019 ChallengeabstractTo better understand early brain development in health and disorder, it is critical to accurately segment infant brain magnetic resonance (MR) images into white matter (WM), gray matter (GM), and cerebrospinal fluid (CSF). Deep learning-based methods have achieved state-of-the-art performance; h owever, one of the major limitations is that the learning-based methods may suffer from the multi-site issue, that is, the models trained on a dataset from one site may not be applicable to the datasets acquired from other sites with different imaging protocols/scanners. To promote methodological development in the community, the iSeg-2019 challenge (http://iseg2019.web.unc.edu) provides a set of 6-month infant subjects from multiple sites with different protocols/scanners for the participating methods. T raining/validation subjects are from UNC (MAP) and testing subjects are from UNC/UMN (BCP), Stanford University, and Emory University. By the time of writing, there are 30 automatic segmentation methods participated in the iSeg-2019. In this article, 8 top-ranked methods were reviewed by detailing their pipelines/implementations, presenting experimental results, and evaluating performance across different sites in terms of whole brain, regions of interest, and gyral landmark curves. We further pointed out their limitations and possible directions for addressing the multi-site issue. We find that multi-site consistency is still an open issue. We hope that the multi-site dataset in the iSeg-2019 and this review article will attract more researchers to address the challenging and critical multi-site issue in practice. Yue Sun 0001, Kun Gao 0002, Zhengwang Wu, Xiaopeng Zong, Zhihao Lei, Ying Wei 0007, Jun Ma 0016, Xiaoping Yang 0001, Xue Feng 0001, Li Zhao 0001, Trung Le Phan, Jitae Shin, Tao Zhong 0002, Yu Zhang 0064, Lequan Yu, Caizi Li, Ramesh Basnet, M. Omair Ahmad, M. N. S. Swamy 0001, Wenao Ma, Qi Dou 0001, Toan Duc Bui, Camilo Bermudez, Bennett A. Landman, Ian H. Gotlib, Kathryn L. Humphreys, Sarah Shultz, Longchuan Li, Sijie Niu, Weili Lin, Valerie Jewells, Dinggang Shen, Gang Li 0001, Li Wang 0026 |
IEEE Trans. Medical Imaging | 7 |
| 2020 | Scale-Invariant Siamese Network For Person Re-IdentificationabstractMost existing methods for person re-identification (ReID) almost match people at a single scale and ignore that people are often distinguishable at the right spatial locations and scales. Unlike previous works designing complex convolutional neural network (CNN) architecture or concatenating multi-branch scale-specific features, we aim to employ a simple network to learn scale-invariant features. Concretely, we first propose a shared two-branch framework with two-scale images from the same identity as inputs, which is beneficial for ReID network to focus on common features in different-scale images. Furthermore, we introduce a novel attention loss to enforce discriminative regions between two branches more consistent in the visual level. Finally, we conduct extensive evaluations on three largescale datasets and report competitive performance. Yunzhou Zhang, Shuangwei Liu, Jining Bao, Ying Wei 0007 |
ICIP | 5 |
| 2020 | Vehicle re-identification based on unsupervised local area detection and view discrimination
Ying Wei 0007, Chuyuan Wang |
Image Vis. Comput. | 3 |
| 2019 | Combination of Appearance and License Plate Features for Vehicle Re-IdentificationabstractIn this work, we propose a two-module framework that combines appearance and corresponding license plate features for vehicle re-identification (Re-ID). In the appearance module, we design a Two-Branch Network to extract comprehensive global features. To obtain more discriminative feature representations, we propose an enhanced triplet loss (ETL) and combine ETL with softmax loss to optimize the parameter of Two-Branch Network. In the license plate module, we present a license plate Re-ID network that incorporates the bidirectional LSTMs into CNNs, which is effective for capturing the contexts in license plate images and significantly improves the performance of license plate Re-ID. We validate our method on both VeRi-776 dataset [1] and VehicleID dataset [2]. The experimental results show that our method outperforms most state-of-the-art approaches for vehicle Re-ID, even if only the appearance module is used. Yanguang He, Chenhe Dong, Ying Wei 0007 |
ICIP | 3 |