VLDB 2026 Research / reviewers in the wild / expert
Youmei Zhang
dblp:131/1392
· DBLP profile ↗
15ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0003-4185-0127ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CVDII: Enhancing One-Shot Skeleton Action Recognition Through Cross-View Dynamic Information InteractionabstractOne-shot 3D skeleton action recognition task struggles with diverse intra-class action execution styles, causing excessive discriminative information to obstruct obtaining separable feature space. We innovatively propose leveraging shared information among intra-class action executions to mitigate the over-influence of discriminative information. To this end, we proposed dynamic information interaction module (DIIM) that enables shared information to effectively weaken excessive discriminative information. Specifically, DIIM facilitates effective information interaction by constructing a guided evolution pool to store execution-related shared information and ensure such information can be retrieved. We devise shared-discriminative projection strategy (SDPS) which adopts different feature extraction strategies for specific skeleton topologies to target mining discriminative and shared information from different views of skeleton data. In summary, our proposed Cross-View Dynamic Information Interaction (CVDII) framework integrates DIIM and SDPS, effectively tackles the problem of discriminative information redundancy caused by diverse intra-class action execution styles. Experiments conducted on NTU 60, NTU 120, PKU-MMD, and Kinetics datasets demonstrate that our proposed CVDII achieves remarkable performance. Youmei Zhang, Weidong Zhang 0005, Zhiheng Li 0005, Bin Li 0042, Wei Zhang 0021 |
IEEE Trans. Image Process. | 1 |
| 2025 | FTCNet: A Foreground Transformation Contrast Network for Marine Animal Segmentation
Youmei Zhang, Jingwei Guan |
PRCV (15) | 1 |
| 2025 | ViV-ReID: Bidirectional Structural-Aware Spatial-Temporal Graph Networks on Large-Scale Video-Based Vessel Re-Identification DatasetabstractVessel re-identification (ReID) serves as a foundational task for intelligent maritime transportation systems. To enhance maritime surveillance capabilities, this study investigates video-based vessel ReID, a critical yet underexplored task in intelligent transportation systems. The lack of relevant datasets has limited the progress of Video-based vessel ReID research work. We established ViV-ReID, the first publicly available large-scale video-based vessel ReID dataset, comprising 480 vessel identities captured from 20 cross-port camera views (7,165 tracklets and 1.14 million frames), establishing a benchmark for advancing vessel ReID from image to video processing. Videos offer significantly richer information than single-frame images. The dynamic nature of video often leads to fragmented spatio-temporal features causing disrupted contextual understanding, and to address this problem, we further propose a Bidirectional Structural-Aware Spatial-Temporal Graph Network (Bi-SSTN) that explicitly aligns spatio-temporal features using vessel structural priors. Extensive experiments on the ViV-ReID dataset demonstrate that image-based ReID methods often show suboptimal performance when applied to video data. Meanwhile, it is crucial to validate the effectiveness of spatio-temporal information and establish performance benchmarks for different methods. The Bidirectional Structural-Aware Spatial-Temporal Graph Network (Bi-SSTN) significantly outperforms state-of-the-art methods on ViV-ReID, confirming its efficacy in modeling vessel-specific spatio-temporal patterns. Project web page: https://vsislab.github.io/ViV_ReID/. Mingxin Zhang 0006, Fuxiang Feng, Lin Zhang 0041, Youmei Zhang, Xiaolei Li 0003, Wei Zhang 0021 |
IEEE Trans. Image Process. | 5 |
| 2025 | SLPDR: A Benchmark for Ship License Plate Detection and RecognitionabstractShip identification is a prerequisite for the intelligent management of maritime transportation, yet existing research is confined to broad ship detection and categorization, which only provides the ship’s location or type instead of its identification. Inspired by the research on the Car License Plate (CLP), we make the first attempt to propose the concept of the Ship License Plate (SLP). In addition, the limited data hinders research on ship identification. To overcome this obstacle, we construct the first large-scale Ship License Plate Detection and Recognition (SLPDR) dataset, which contains 1,472 ship identities and 88,862 images. In addition, this paper proposes an SLP detection model named YOLO-SSA and evaluates this model as well as typical detection methods on the SLPDR dataset. The experimental results demonstrate that the proposed YOLO-SSA achieves better SLP detection performance by enhancing the features where ships and SLPs are located. Furthermore, we explore the prospective applications of SLPs in intelligent maritime transportation, including ship monitoring and berth management. Project web page: https://vsislab.github.io/SLPDR/ Youmei Zhang, Ran Song 0001, Yonghuai Liu, Ardhendu Behera, Mingxin Zhang 0006, Wei Zhang 0021 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Complete contextual information extraction for self-supervised monocular depth estimation
Dazheng Zhou, Xianjie Gao, Youmei Zhang, Bin Li 0042 |
Comput. Vis. Image Underst. | 4 |
| 2024 | Dual-Path Feature Fusion Network for Semantic Segmentation of Remote Sensing ImagesabstractBoth global contextual information and local texture information are of vital importance for the semantic segmentation of remote sensing images due to the high spatial resolution of remote sensing images and large variations in intra-class object size. In this letter, we propose a novel dual-path feature fusion semantic segmentation network for remote sensing images. A pure convolutional module called dual-path feature extraction module (DPFE) is applied to model global contextual and local texture features simultaneously with low complexity. Inspired by ConvNeXt with comparable global contextual modeling capacity with Transformer, the global path of DPFE draws some successful strategies of ConvNeXt to generate powerful global feature. Meanwhile, an attention feature fusion module (AFF) is proposed, which achieves the global and local feature comprehensive fusion by exploring the correlation of channels through attention mechanism. The proposed network is evaluated on Vaihingen and Potsdam benchmarks and the quantitative results show the proposed network can achieve overall accuracy (OA) of 91.3% and 89.7%, respectively, which are better than several representative semantic segmentation approaches. Yu Zhang 0275, Youmei Zhang, Bin Li 0042 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Progressive dense feature fusion network for single image deraining
Fuxiang Feng, Youmei Zhang, Weidong Zhang 0005, Bin Li 0042 |
Pattern Recognit. Lett. | 2 |
| 2022 | 3D Layout Estimation via Weakly Supervised Learning of Plane Parameters From 2D SegmentationabstractThe task of 3D layout estimation in an indoor scene is to predict the holistic 3D structural information of the scene from an RGB image. It is costly to obtain the ground truth 3D layout, and this issue severely restricts the learning based 3D layout estimation approaches. In this paper, we present a novel weakly supervised learning framework that is able to learn the 3D layout effectively with 2D layout segmentation mask as supervision. We employ a deep neural network to predict the plane parameters and camera intrinsic parameters in the image. Based on the predicted plane instances, the 3D layout as well as the corresponding depth map and 2D segmentation can be generated. The key objectives for learning meaningful plane parameters are the label consistency of layout segmentation and depth consistency of border pixels from adjacent planes, with which the ground truth 2D layout segmentation is able to supervise the learning of the 3D layout. We further incorporate 3D geometric reasoning and prior knowledge in the learning process to ensure that the learned 3D layout is realistic and reasonable. Experimental results show that our method can produce accurate 3D layout estimates by weakly supervised learning. Weidong Zhang 0005, Youmei Zhang, Ran Song 0001, Ying Liu 0026, Wei Zhang 0021 |
IEEE Trans. Image Process. | 2 |
| 2022 | JoT-GAN: A Framework for Jointly Training GAN and Person Re-Identification ModelabstractTo cope with the problem caused by inadequate training data, many person re-identification (re-id) methods exploit generative adversarial networks (GAN) for data augmentation, where the training of GAN is typically independent of that of the re-id model. The coupling relation between them that probably brings in a performance gain of re-id is thus ignored. In this work, we propose a general framework, namely JoT-GAN, to jointly train GAN and the re-id model. It can simultaneously achieve the optima of both the generator and the re-id model, where the training is guided by each other through a discriminator. The re-id model is boosted for two reasons: (1) the adversarial training encourages it to fool the discriminator, and (2) the generated samples augment the training data. Extensive results on benchmark datasets show that for the re-id model trained with the identification loss as well as the triplet loss, the proposed joint training framework outperforms existing methods with separate training and achieves state-of-the-art re-id performance. Ran Song 0001, Qian Zhang 0076, Peng Duan 0002, Youmei Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2019 | Attention to Head Locations for Crowd Counting
Youmei Zhang, Chunluan Zhou, Faliang Chang, Alex Chichung Kot, Wei Zhang 0021 |
ICIG (2) | 1 |
| 2019 | Multi-resolution attention convolutional neural network for crowd countingabstractEstimating crowd counts remains a challenging task due to the problems of scale variations, non-uniform distribution and complex backgrounds. In this paper, we propose a multi-resolution attention convolutional neural network (MRA-CNN) to address this challenging task. Except for the counting task, we exploit an additional density-level classification task during training and combine features learned for the two tasks, thus forming multi-scale, multi-contextual features to cope with the scale variation and non-uniform distribution. Besides, we utilize a multi-resolution attention (MRA) model to generate score maps, where head locations are with higher scores to guide the network to focus on head regions and suppress non-head regions regardless of the complex backgrounds. During the generation of score maps, atrous convolution layers are used to expand the receptive field with fewer parameters, thus getting higher-level features and providing the MRA model more comprehensive information. Experiments on ShanghaiTech, WorldExpo’10 and UCF datasets demonstrate the effectiveness of our method. Youmei Zhang, Chunluan Zhou, Faliang Chang, Alex Chichung Kot |
Neurocomputing | 1 |
| 2019 | A scale adaptive network for crowd counting
Youmei Zhang, Chunluan Zhou, Faliang Chang, Alex Chichung Kot |
Neurocomputing | 1 |
| 2018 | Auxiliary learning for crowd counting via count-net
Youmei Zhang, Faliang Chang, Mengdi Wang 0007, Fulei Zhang |
Neurocomputing | 1 |
| 2016 | Deep Neural Networks for wireless localization in indoor and outdoor environments
Wei Zhang 0021, Kan Liu 0001, Weidong Zhang 0005, Youmei Zhang, Jason Gu |
Neurocomputing | 4 |
| 2015 | Multimodal learning for facial expression recognition
Wei Zhang 0021, Youmei Zhang, Lin Ma 0002, Jingwei Guan, Shijie Gong |
Pattern Recognit. | 2 |