Pingyu Wang

dblp:215/7135 · DBLP profile ↗
← Back
22ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0001-9769-8035ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 7 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DSRIR: Dynamic spatial refinement learning for progressive all-in-one image restoration
Xiao Liu 0022, Yutong Yang, Zhengyong Wang, Xiaohai He, Honggang Chen, Yi Li 0069, Pingyu Wang
Inf. Process. Manag.8
2026 RelPosGAR: Hierarchical relative position-aware interaction modeling for weakly supervised skeleton-based group activity recognition
Lindong Li, Linbo Qing, Liuyi Tao, Pingyu Wang, Honggang Chen, Owen Noel Newton Fernando, Weisi Lin
Pattern Recognit.4
2026 Beyond deceptive flatness: Dual-order solution for strengthening adversarial transferability
Pingyu Wang, Xingjian Zheng, Linbo Qing, Qi Liu 0005
Pattern Recognit.2
2026 Progressive Reasoning-Based Group Activity Recognition
abstract
Group activity recognition (GAR) plays a crucial role in computer vision, enabling the exploration and comprehension of human behavior patterns. Existing methods mainly focus on dyad-level interactions within a group, but sociological studies have highlighted the importance of individual features, subgroup-level interactions, and overall group structure for understanding group activities. Therefore, we propose a new framework, the progressive group activity reasoning model (PGAR), which models these four aspects for GAR. Initially, we construct a person-person graph (PPG) using individual features to capture dyadic interactions. Subsequently, the PPG is fed into a novel ingredient graph model (Ingredient-GNN) for capturing subgroup-level interactions. Finally, we fuse the dyad-level and subgroup-level interactions with global information of group structure, obtained through an F-Formation modeling module, to form comprehensive representations for GAR. The F-Formation modeling module decouples the group structure into position, orientation, and skeleton graphs, and subsequently performs attribute recoupling at the individual level using the designed Tri-Coupling Transformer to form a global representation of the group structure. Extensive experiments on four public datasets demonstrate that our final model effectively integrates multi-level representations for group activity understanding, with our F-Formation modeling module outperforming comparable methods that rely solely on non-visual data.
Lindong Li, Linbo Qing, Wang Tang, Pingyu Wang, Haosong Gou, Ce Zhu
IEEE Trans. Circuits Syst. Video Technol.4
2025 Graph-based interactive knowledge distillation for social relation continual learning
abstract
As multimedia advances, there is a growing need for machines to adeptly understand diverse social relations . Traditional methods for recognizing these relations, which are limited to a fixed number of classes, are ill-equipped for continual learning as new social interactions emerge. To address this prob-lem, we propose a pioneering Graph-based Interactive Knowledge Distillation (GI-KD) method for social relation continual learning. GI-KD, embedded in a class incremental learning structure, creates a balanced system where previously learned social relations and new knowledge are positioned at either end of the scale. The old and new knowledge is learned dynamically by adjusting the tilt of the balance. To achieve this balance, we propose a novel Libra loss function, which evaluate the relative contribution of old and new information and thus guides the adaptive fine-tuning of the model. We evaluate the GI-KD on three public social relation recognition (SRR) datasets, under different data distribution strategies. Our method shows a remarkable average 3.6% increase in incremental accuracy over current CIL techniques, effectively reducing catastrophic forgetting. Furthermore, GI-KD improves mAP and Acc by 4.6%, 5.4%, and 4.5%, respectively, compared to current CIL techniques, highlighting its strength in both continual learning and SRR.
Wang Tang, Linbo Qing, Pingyu Wang, Lindong Li, Yonghong Peng
Neurocomputing3
2025 Spatio-temporal interactive reasoning model for multi-group activity recognition
Jianglan Huang, Lindong Li, Linbo Qing, Wang Tang, Pingyu Wang, Li Guo 0018, Yonghong Peng
Pattern Recognit.5
2025 DRFormer: A Discriminable and Reliable Feature Transformer for Person Re-Identification
abstract
As person image variations are likely to cause a part misalignment problem, most previous person Re-Identification (ReID) works may adopt local feature partition or additional landmark annotations to acquire aligned person features and boost ReID performance. However, such approaches either only achieve coarse-grained part alignments without considering detailed image variations within each part, or require extra annotated landmarks to train an available pose estimation model. In this work, we propose an effective Discriminable and Reliable Transformer (DRFormer) framework to learn part-aligned person representations with only person identity labels. Specifically, the DRFormer framework consists of Discriminable Feature Transformer (DFT) and Reliable Feature Transformer (RFT) modules, which generate discriminable and reliable high-order features, respectively. For reducing the dimension of high-order features, the DFT module utilizes a Self-Attentive Kronecker Product (SAKP) algorithm to promote the representational capabilities of compressed features via a self-attention strategy. For eliminating the background noise, the RFT module mines the foreground regions to adaptively aggregate foreground features via a Gumbel-Softmax strategy. Moreover, the proposed framework derives from an interpretable motivation and elegantly solves part misalignments without using feature partition or pose estimation. This paper theoretically and experimentally demonstrates the superiority of the proposed DRFormer framework, achieving state-of-the-art performance on various person ReID datasets.
Pingyu Wang, Xingjian Zheng, Linbo Qing, Bonan Li, Zhicheng Zhao 0001, Honggang Chen
IEEE Trans. Inf. Forensics Secur.1
2025 A Stable and Efficient Data-Free Model Attack With Label-Noise Data Generation
abstract
The objective of a data-free closed-box adversarial attack is to attack a victim model without using internal information, training datasets or semantically similar substitute datasets. Concerned about stricter attack scenarios, recent studies have tried employing generative networks to synthesize data for training substitute models. Nevertheless, these approaches concurrently encounter challenges associated with unstable training and diminished attack efficiency. In this paper, we propose a novel query-efficient data-free closed-box adversarial attack method. To mitigate unstable training, for the first time, we directly manipulate the intermediate-layer feature of a generator without relying on any substitute models. Specifically, a label noise-based generation module is created to enhance the intra-class patterns by incorporating partial historical information during the learning process. Additionally, we present a feature-disturbed diversity generation method to augment the inter-class distance. Meanwhile, we propose an adaptive intra-class attack strategy to heighten attack capability within a limited query budget. In this strategy, entropy-based distance is utilized to characterize the relative information from model outputs, while positive classes and negative samples are used to enhance low attack efficiency. The comprehensive experiments conducted on six datasets demonstrate the superior performance of our method compared to six state-of-the-art data-free closed-box competitors in both label-only and probability-only attack scenarios. Intriguingly, our method can realize the highest attack success rate on the online Microsoft Azure model under an extremely low query budget. Additionally, the proposed approach not only achieves more stable training but also significantly reduces the query count for a more balanced data generation. Furthermore, our method can maintain the best performance under the existing defense models and a limited query budget.
Xingjian Zheng, Linbo Qing, Qi Liu 0005, Pingyu Wang, Yu Liu 0123, Jiyang Liao
IEEE Trans. Inf. Forensics Secur.5
2025 Hypergraph Mamba Reasoning-Based Social Relation Recognition
abstract
Recognizing social relations from images is crucial for improving machine perception of social interactions. Current studies mainly focus on exploring single-type relation reasoning frameworks, such as the relation between father, mother and son in a family. However, real-world scenarios often involve complex hybrid relations, such as friendships and professional relations, which pose a challenge for current methods due to the difficulty of establishing robust logical connections between these relations. In fact, in this hybrid social relation recognition setting, the interactions extend beyond dyadic to multipartite structures. To effectively explore these multipartite interactions, we propose a novel Hypergraph Mamba (HGM) framework. Specifically, we construct two hypergraphs, i.e., Person-Person Hypergraphs (PPH) and Person-Object Hypergraphs (POH), to model these high-order multipartite interactions. The HGM module performs social relation reasoning within these hypergraph structures, which includes a Vertex Selection Algorithm to mitigate inference confusion by filtering out confounders, and a Vertex Interaction Operator to find optimal global vertex neighborhoods by capturing long-range vertex dependencies. In addition, a Multilevel Transformer is proposed to adaptively align the PPH and POH inferred knowledge and visual signals to facilitate information fusion. We validate the effectiveness of our proposed HGM model on several public datasets and perform extensive ablation studies to elucidate the reasons contributing to its superior performance. Experimental results indicate that our HGM model achieves superior accuracy in predicting social relations compared to the state-of-the-art methods. Codes and datasets are available at: https://github.com/tw-repository/HGM-SRR.
Wang Tang, Linbo Qing, Pingyu Wang, Lindong Li, Ce Zhu
IEEE Trans. Image Process.3
2025 GAReID: Grouped and Attentive High-Order Representation Learning for Person Re-Identification
abstract
As person parts are frequently misaligned between detected human boxes, an image representation that can handle this part misalignment is required. In this work, we propose an effective grouped attentive re-identification (GAReID) framework to learn part-aligned and background robust representations for person re-identification (ReID). Specifically, the GAReID framework consists of grouped high-order pooling (GHOP) and attentive high-order pooling (AHOP) layers, which generate high-order image and foreground features, respectively. In addition, a novel grouped Kronecker product (GKP) is proposed to use both channel group and shuffle strategies for high-order feature compression, while promoting the representational capabilities of compressed high-order features. We show that our method derives from an interpretable motivation and elegantly reduces part misalignments without using landmark detection or feature partition. This article theoretically and experimentally demonstrates the superiority of the GAReID framework, achieving state-of-the-art performance on various person ReID datasets.
Pingyu Wang, Zhicheng Zhao 0001, Yanyun Zhao, Nikolaos V. Boulgouris
IEEE Trans. Neural Networks Learn. Syst.1
2024 BlazeBVD: Make Scale-Time Equalization Great Again for Blind Video Deflickering
Xinmin Qiu, Congying Han, Bonan Li, Tiande Guo, Pingyu Wang, Xuecheng Nie
ECCV (17)6
2024 Hunting imaging biomarkers in pulmonary fibrosis: Benchmarks of the AIIB23 challenge
abstract
• This paper investigates the capacity of AI models for airway modelling on national datasets with paired clinical metadata. • We evaluated AI models against unharmonised, noisy, and out-of-distribution data, as well as the prognostication for FLD. • We found a new biomarker for mortality prediction, outperforming existing clinical measurements (FVC% and fibrosis scores). • In-depth analysis of AI models on airway modelling and prognosis, highlighting challenges and future research directions. Airway-related quantitative imaging biomarkers are crucial for examination, diagnosis, and prognosis in pulmonary diseases. However, the manual delineation of airway structures remains prohibitively time-consuming. While significant efforts have been made towards enhancing automatic airway modelling, current public-available datasets predominantly concentrate on lung diseases with moderate morphological variations. The intricate honeycombing patterns present in the lung tissues of fibrotic lung disease patients exacerbate the challenges, often leading to various prediction errors. To address this issue, the 'Airway-Informed Quantitative CT Imaging Biomarker for Fibrotic Lung Disease 2023′ (AIIB23) competition was organized in conjunction with the official 2023 International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI). The airway structures were meticulously annotated by three experienced radiologists. Competitors were encouraged to develop automatic airway segmentation models with high robustness and generalization abilities, followed by exploring the most correlated QIB of mortality prediction. A training set of 120 high-resolution computerised tomography (HRCT) scans were publicly released with expert annotations and mortality status. The online validation set incorporated 52 HRCT scans from patients with fibrotic lung disease and the offline test set included 140 cases from fibrosis and COVID-19 patients. The results have shown that the capacity of extracting airway trees from patients with fibrotic lung disease could be enhanced by introducing voxel-wise weighted general union loss and continuity loss. In addition to the competitive image biomarkers for mortality prediction, a strong airway-derived biomarker (Hazard ratio>1.5, p < 0.0001) was revealed for survival prognostication compared with existing clinical measurements, clinician assessment and AI-based biomarkers.
Yang Nan 0002, Xiaodan Xing, Zeyu Tang 0001, Federico Felder, Sheng Zhang 0024, Roberta Eufrasia Ledda, Xiaoliu Ding, Feng Shi 0001, Tianyang Sun, Zehong Cao, Yun Gu, Pingyu Wang, Wen Tang 0005, Pengxin Yu, Han Kang, Junqiang Chen, Michail Mamalakis, Francesco Prinzi, Gianluca Carlini, Lisa Cuneo, Abhirup Banerjee, Zhaohu Xing, Lei Zhu 0003, Zacharia Mesbah, Dhruv Jain, Tsiry Mayet, Hongyu Yuan, Qing Lyu 0009, Abdul Qayyum 0002, Moona Mazher, Athol Wells, Simon Walsh, Guang Yang 0006
Medical Image Anal.18
2023 LTReID: Factorizable Feature Generation With Independent Components for Long-Tailed Person Re-Identification
abstract
With the rapid increase of large-scale and real-world person datasets, it is crucial to address the problem of long-tailed data distributions,i.e., head classes have large number of images while tail classes occupy extremely few samples. We observe that the imbalanced data distribution is likely to distort the overall feature space and impair the generalization capability of trained models. Nevertheless, this long-tailed problem has been rarely investigated in previous person Re-Identification (ReID) works. In this paper, we propose a novelLong-Tailed Re-Identification(LTReID) framework to simultaneously alleviate class-imbalance and hard-imbalance problems. Specifically, each real feature is decomposed into multiple independent components with two decorrelation losses. Then these components are randomly aggregated to generate more fake features for tail classes than head ones, resulting in the class-balance between head and tail classes. For the hard-balance between easy and hard samples, we utilize adversarial learning to generate more hard features than easy ones. The proposed framework can be trained in an end-to-end manner and avoids increasing the space and time complexity of inference models. Moreover, comprehensive experiments are conducted on the four ReID datasets so as to validate the effectiveness of the overall framework and the advantage of each module. Our results show that when trained with either balanced or imbalanced datasets, the LTReID achieves superior performance over the state-of-the-art methods.
Pingyu Wang, Zhicheng Zhao 0001, Hongying Meng
IEEE Trans. Multim.1
2022 Attentive Feature Augmentation for Long-Tailed Visual Recognition
abstract
Deep neural networks have achieved great success on many visual recognition tasks. However, training data with a long-tailed distribution dramatically degenerates the performance of recognition models. In order to relieve this imbalance problem, an effective Long-Tailed Visual Recognition (LTVR) framework is proposed based on learned balance and robust features under long-tailed distribution circumstances. In this framework, a plug-and-play Attentive Feature Augmentation (AFA) module is designed to mine class-related and variation-related features of original samples via a novel hierarchical channel attention mechanism. Then, those features are aggregated to synthesize fake features to cope with the imbalance of the original dataset. Moreover, a Lay-Back Learning Schedule (LBLS) is developed to ensure a good initialization of feature embedding. Extensive experiments are conducted with a two-stage training method to verify the effectiveness of the proposed framework on both feature learning and classifier rebalancing in the long-tailed image recognition task. Experimental results show that, when trained with imbalanced datasets, the proposed framework achieves superior performance over the state-of-the-art methods.
Weiqiu Wang, Zhicheng Zhao 0001, Pingyu Wang, Hongying Meng
IEEE Trans. Circuits Syst. Video Technol.3
2021 Pruned-YOLO: Learning Efficient Object Detector Using Model Pruning
Pingyu Wang, Zhicheng Zhao 0001
ICANN (4)2
2021 Rethinking Anchor-Object Matching and Encoding in Rotating Object Detection
abstract
Rotating object detection is more challenging than horizontal object detection because of the multi-orientation of the objects involved. In the recent anchor-based rotating object detector, the IoU-based matching mechanism has some mismatching and wrong-matching problems. Moreover, the encoding mechanism does not correctly reflect the location relationships between anchors and objects. In this paper, RBox-Diff-based matching (RDM) mechanism and angle-first encoding (AE) method are proposed to solve these problems. RDM optimizes the anchor-object matching by replacing IoU (Intersection-over-Union) with a new concept called RBox-Diff, while AE optimizes the encoding mechanism to make the encoding results consistent with the relative position between objects and anchors more. The proposed methods can be easily applied to most of the anchor-based rotating object detectors without introducing extra parameters. The extensive experiments on DOTA-v1.0 dataset show the effectiveness of the proposed methods over other advanced methods.
Zhaohui Hou, Pingyu Wang, Zhicheng Zhao 0001
VCIP3
2021 HOReID: Deep High-Order Mapping Enhances Pose Alignment for Person Re-Identification
abstract
Despite the remarkable progress in recent years, person Re-Identification (ReID) approaches frequently fail in cases where the semantic body parts are misaligned between the detected human boxes. To mitigate such cases, we propose a novel High-Order ReID (HOReID) framework that enables semantic pose alignment by aggregating the fine-grained part details of multilevel feature maps. The HOReID adopts a high-order mapping of multilevel feature similarities in order to emphasize the differences of the similarities between aligned and misaligned part pairs in two person images. Since the similarities of misaligned part pairs are reduced, the HOReID enhances pose-robustness within the learned features. We show that our method derives from an intuitive and interpretable motivation and elegantly reduces the misalignment problem without using any prior knowledge from human pose annotations or pose estimation networks. This paper theoretically and experimentally demonstrates the effectiveness of the proposed HOReID, achieving superior performance over the state-of-the-art methods on the four large-scale person ReID datasets.
Pingyu Wang, Zhicheng Zhao 0001, Xingyu Zu, Nikolaos V. Boulgouris
IEEE Trans. Image Process.1
2021 Deep Multi-Patch Matching Network for Visible Thermal Person Re-Identification
abstract
Visible Thermal Person Re-Identification(VTReID) is a cross-modality retrieval problem in computer vision. Accurate VTReID is very challenging due to large modality discrepancies. In this work, we design a novelMulti-Patch Matching Network(MPMN) framework to simultaneously mitigate the heterogeneity of coarse-grained and fine-grained visual semantics. In view of cross-modality matching, we verify that aligning modality distributions of the original features is likely to suffer from the selective alignment behavior, i.e., only focuses on easiest dimensions or subspaces. Inspired by adversarial learning, we propose a newMulti-Patch Modality Alignment(MPMA) loss to jointly balance and reduce the modality discrepancies of multi-patch features by mining hard subspaces and abandoning easy subspaces. Since multi-patch features are potentially complementary to each other, the semantic correlations between different patches should be exploited during training. Motivated by knowledge distillation, we put forward a newCross-Patch Correlation Distillation(CPCD) loss to transfer the semantic knowledges across different patches. To balance multi-patch tasks, an effectivePatch-Aware Priority Attention(PAPA) method is further introduced to dynamically prioritize hard patch tasks during training. This paper experimentally demonstrates the effectiveness of the proposed methods, achieving superior performance over the state-of-the-art methods on RegDB and SYSU-MM01 datasets.
Pingyu Wang, Zhicheng Zhao 0001, Yanyun Zhao, Haiying Wang 0005, Lei Yang 0063
IEEE Trans. Multim.1
2020 Deep hard modality alignment for visible thermal person re-identification
Pingyu Wang, Zhicheng Zhao 0001, Yanyun Zhao, Lei Yang 0063
Pattern Recognit. Lett.1
2019 GRNet: An Efficient Group Ranking Network for Face Attribute Recognition
abstract
Face attribute recognition methods have been far from real-world applications despite remarkable progress in recent years. Most of these methods fail to mine attribute relationships with traditional cross entropy loss and hold considerable computational complexity. In this work, we propose a group ranking network (GRNet) to investigate substantial relationships among face attributes. First, an attribute grouping manner is designed to capture interactions among spatially related attributes and enhance computational efficiency. Second, a new supervision signal is presented to model attribute ranking relationships. Under the supervision of ranking loss, GRNet learns intra-attribute and inter-attribute ranking features to intensify the representational capability of attribute models. We evaluate our approach on aligned and unaligned CelebA datasets. Results show that the performance of the proposed approach is superior to other start-of-the-art methods on attribute recognition.
Dingchang Hu, Pingyu Wang, Zhicheng Zhao 0001
VCIP2
2019 Deep class-skewed learning for face recognition
Pingyu Wang, Zhicheng Zhao 0001, Yandong Guo, Yanyun Zhao, Bojin Zhuang
Neurocomputing1
2017 Joint multi-feature fusion and attribute relationships for facial attribute prediction
abstract
Predicting facial attributes from wild images is very challenging due to complex face variations. The key to this problem is to construct rich facial representations and take advantage of attribute relationships. In this paper, we propose a novel multi-task convolutional neural network (MTCNN) and a supervision signal called Online Batch Relation Loss (OBRL) for face attribute prediction in the wild. In particular, MTCNN builds informative facial features by embedding identity, age and race features from IdentityNet, AgeNet and RaceNet respectively. In addition, OBRL can diminish distribution shift of attribute relationships by mining attribute correlation within each minibatch, while it penalizes the probability divergence between a pair of attributes. In order to learn discriminative attribute features, we feed AttributeNet with fused facial features and partition attributes into nine groups to share intra-group features and reduce redundant computation. Finally, AttributeNet is optimized with the joint supervision of Cross Entropy Loss and OBRL. Experiments on CelebA and LFWA show that the proposed method outperforms the state-of-the-art methods with a significant margin.
Pingyu Wang, Zhicheng Zhao 0001
VCIP1