VLDB 2026 Research / reviewers in the wild / expert
Shangzhi Teng
dblp:226/6084
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0001-7098-9932ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Small Object Detection via Frequency-Based Multi-modal Fusion
Shangzhi Teng, Yekai Li, Xi Gong, Xueqiang Lv |
MMM (1) | 1 |
| 2026 | CVAF: A CLIP-Based View-Consistent Alignment Framework for Aerial-Ground Person Re-IdentificationabstractWith the increasing adoption of UAV platforms in areas such as public safety and smart cities, Aerial-Ground Person Re-Identification (AGPReID) has emerged as a crucial yet highly challenging task, garnering growing interest from the research community. While existing approaches have leveraged identity attributes and viewpoint disentanglement strategies to improve cross-view matching, their heavy reliance on prior knowledge often compromises model generalization. Furthermore, some methods that explicitly separate viewpoints may unintentionally discard identity-related, view-invariant features, leading to incomplete identity representations. To address these limitations, we propose a CLIP-based View-Consistent Alignment Framework (CVAF) with two training stages. In the first stage, learnable text tokens are employed to represent identity-aware textual descriptions. To promote consistent alignment across varying viewpoints, we introduce a Text Consistency Loss (TCL) that regularizes the stability of text-token interactions with multi-view images. In the second stage, we present a Semantic Filtering Module (SFM) that jointly modulates image patch tokens along spatial and channel dimensions. A text-guided cross-attention mechanism generates spatial attention maps to explicitly emphasize identity-relevant regions, while semantic matching between textual features and visual tokens enables adaptive reweighting of image representations, effectively suppressing background clutter and view-specific noise. Extensive experiments on multiple AGPReID datasets demonstrate that our CVAF outperforms the state-of-the-art methods. Dongxu Mao, Shangzhi Teng, Xueqiang Lyu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Attribute correlation mask fusion network for pedestrian attribute recognition
Baoan Li, Shangzhi Teng, Xueqiang Lyu |
Vis. Comput. | 3 |
| 2024 | TL-RelD: Tight-Loose Pairwise Loss for Object Re-Identification
Changwang Mei, Xindong You, Shangzhi Teng, Xueqiang Lyu |
PRCV (12) | 3 |
| 2023 | Unsupervised Vehicle Re-Identification via Raw UAV Videos
Shangzhi Teng, Tingting Dong |
ICIG (2) | 1 |
| 2021 | Viewpoint and Scale Consistency Reinforcement for UAV Vehicle Re-Identification
Shangzhi Teng, Shiliang Zhang, Qingming Huang, Nicu Sebe |
Int. J. Comput. Vis. | 1 |
| 2021 | Multi-View Spatial Attention Embedding for Vehicle Re-IdentificationabstractVehicle Re-Identification (Re-ID) is a challenging vision task mainly because the appearance of a vehicle varies dramatically under different viewpoints. Moreover, different vehicles with the same model and color commonly show similar appearance, thus are hard to be distinguished. To alleviate negative effects of viewpoint variance, we design a multi-view branch network where each branch learns a viewpoint-specific feature without parameter sharing. Being able to focus on a limited range of viewpoints, this viewpoint-specific feature performs substantially better than the general feature learned by an uniform network. To further differentiate visually similar vehicles, we strengthen the discriminative power on their subtle local differences by introducing a spatial attention model into each feature learning branch. The multi-view feature learning and spatial attention learning compose our neural network architecture, which is trained end to end with the softmax loss and triplet loss, respectively. We evaluate our methods on two large vehicle Re-ID datasets, i.e., VehicleID and VeRi-776, respectively. Extensive experiments show that our methods achieve promising performance. For example, we achieve mAP accuracy of 76.78% and 72.53% on VehicleID and VeRi-776 dataset respectively, substantially better than current state-of-the art. Shangzhi Teng, Shiliang Zhang, Qingming Huang, Nicu Sebe |
IEEE Trans. Circuits Syst. Video Technol. | 1 |