Shangzhi Teng

dblp:226/6084 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0001-7098-9932ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Small Object Detection via Frequency-Based Multi-modal Fusion
Shangzhi Teng, Yekai Li, Xi Gong, Xueqiang Lv
MMM (1)1
2026 CVAF: A CLIP-Based View-Consistent Alignment Framework for Aerial-Ground Person Re-Identification
abstract
With the increasing adoption of UAV platforms in areas such as public safety and smart cities, Aerial-Ground Person Re-Identification (AGPReID) has emerged as a crucial yet highly challenging task, garnering growing interest from the research community. While existing approaches have leveraged identity attributes and viewpoint disentanglement strategies to improve cross-view matching, their heavy reliance on prior knowledge often compromises model generalization. Furthermore, some methods that explicitly separate viewpoints may unintentionally discard identity-related, view-invariant features, leading to incomplete identity representations. To address these limitations, we propose a CLIP-based View-Consistent Alignment Framework (CVAF) with two training stages. In the first stage, learnable text tokens are employed to represent identity-aware textual descriptions. To promote consistent alignment across varying viewpoints, we introduce a Text Consistency Loss (TCL) that regularizes the stability of text-token interactions with multi-view images. In the second stage, we present a Semantic Filtering Module (SFM) that jointly modulates image patch tokens along spatial and channel dimensions. A text-guided cross-attention mechanism generates spatial attention maps to explicitly emphasize identity-relevant regions, while semantic matching between textual features and visual tokens enables adaptive reweighting of image representations, effectively suppressing background clutter and view-specific noise. Extensive experiments on multiple AGPReID datasets demonstrate that our CVAF outperforms the state-of-the-art methods.
Dongxu Mao, Shangzhi Teng, Xueqiang Lyu
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Attribute correlation mask fusion network for pedestrian attribute recognition
Baoan Li, Shangzhi Teng, Xueqiang Lyu
Vis. Comput.3
2024 TL-RelD: Tight-Loose Pairwise Loss for Object Re-Identification
Changwang Mei, Xindong You, Shangzhi Teng, Xueqiang Lyu
PRCV (12)3
2023 Unsupervised Vehicle Re-Identification via Raw UAV Videos
Shangzhi Teng, Tingting Dong
ICIG (2)1
2021 Viewpoint and Scale Consistency Reinforcement for UAV Vehicle Re-Identification
Shangzhi Teng, Shiliang Zhang, Qingming Huang, Nicu Sebe
Int. J. Comput. Vis.1
2021 Multi-View Spatial Attention Embedding for Vehicle Re-Identification
abstract
Vehicle Re-Identification (Re-ID) is a challenging vision task mainly because the appearance of a vehicle varies dramatically under different viewpoints. Moreover, different vehicles with the same model and color commonly show similar appearance, thus are hard to be distinguished. To alleviate negative effects of viewpoint variance, we design a multi-view branch network where each branch learns a viewpoint-specific feature without parameter sharing. Being able to focus on a limited range of viewpoints, this viewpoint-specific feature performs substantially better than the general feature learned by an uniform network. To further differentiate visually similar vehicles, we strengthen the discriminative power on their subtle local differences by introducing a spatial attention model into each feature learning branch. The multi-view feature learning and spatial attention learning compose our neural network architecture, which is trained end to end with the softmax loss and triplet loss, respectively. We evaluate our methods on two large vehicle Re-ID datasets, i.e., VehicleID and VeRi-776, respectively. Extensive experiments show that our methods achieve promising performance. For example, we achieve mAP accuracy of 76.78% and 72.53% on VehicleID and VeRi-776 dataset respectively, substantially better than current state-of-the art.
Shangzhi Teng, Shiliang Zhang, Qingming Huang, Nicu Sebe
IEEE Trans. Circuits Syst. Video Technol.1