VLDB 2026 Research / reviewers in the wild / expert
Min Gao 0007
dblp:45/1016-7
· DBLP profile ↗
13ranked-venue papers
2as first author
12since 2021 · last 2025
0009-0000-0208-121XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Face, body and person analysis · 39% Efficient and distributed learning · 22% Deep learning architectures and training · 12% | |
| Computer networks
1 paper |
Internet of things and sensor networks · 100% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Face, body and person analysis
pedestrian attribute recognition |
2.3 | 3 | 2025 | Enhanced Visual-Semantic Interaction with Tailored Prompts for Pedestrian Attribute Recognition · CVPR 2025 Selective and Orthogonal Feature Activation for Pedestrian Attribute Recognition · AAAI 2024 Exponential Information Bottleneck Theory Against Intra-Attribute Variations for Pedestrian Attribute Recognition · IEEE Trans. Inf. Forensics Secur. 2023 |
Computer vision › Face, body and person analysis
person re-identification |
1.6 | 2 | 2025 | Layered Semi-Second-Order Information Bottleneck and Auxiliary Domain Classification for Person Re-Identification · Int. J. Comput. Vis. 2025 A Two-Stream Hybrid Convolution-Transformer Network Architecture for Clothing-Change Person Re-Identification · IEEE Trans. Multim. 2024 |
Machine learning › Efficient and distributed learning › federated learning
data heterogeneity |
0.9 | 1 | 2025 | Multimodal Fusion Using Multi-View Domains for Data Heterogeneity in Federated Learning · AAAI 2025 |
Machine learning › Transfer learning and domain adaptation
domain shift |
0.9 | 1 | 2025 | Multimodal Fusion Using Multi-View Domains for Data Heterogeneity in Federated Learning · AAAI 2025 |
Machine learning › Efficient and distributed learning
federated learning |
0.9 | 1 | 2025 | Multimodal Fusion Using Multi-View Domains for Data Heterogeneity in Federated Learning · AAAI 2025 |
Machine learning › Efficient and distributed learning › federated learning
multimodal federated learning |
0.9 | 1 | 2025 | Multimodal Fusion Using Multi-View Domains for Data Heterogeneity in Federated Learning · AAAI 2025 |
Computer vision › Vision and language › vision-language model
prompt learning |
0.9 | 1 | 2025 | Enhanced Visual-Semantic Interaction with Tailored Prompts for Pedestrian Attribute Recognition · CVPR 2025 |
Computer vision › Face, body and person analysis › person re-identification › long-term person re-identification
cloth-changing person re-identification |
0.8 | 1 | 2024 | A Two-Stream Hybrid Convolution-Transformer Network Architecture for Clothing-Change Person Re-Identification · IEEE Trans. Multim. 2024 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.8 | 1 | 2024 | Selective and Orthogonal Feature Activation for Pedestrian Attribute Recognition · AAAI 2024 |
Machine learning › Representation and self-supervised learning
information bottleneck |
0.7 | 1 | 2023 | Exponential Information Bottleneck Theory Against Intra-Attribute Variations for Pedestrian Attribute Recognition · IEEE Trans. Inf. Forensics Secur. 2023 |
Machine learning › Representation and self-supervised learning › representation learning › robust representation learning
robust feature learning |
0.7 | 1 | 2023 | Exponential Information Bottleneck Theory Against Intra-Attribute Variations for Pedestrian Attribute Recognition · IEEE Trans. Inf. Forensics Secur. 2023 |
Internet of things and sensor networks
multimodal sensing |
0.3 | 1 | 2025 | Multimodal Fusion Using Multi-View Domains for Data Heterogeneity in Federated Learning · AAAI 2025 |
Machine learning › Learning paradigms
class imbalance |
0.2 | 1 | 2024 | Selective and Orthogonal Feature Activation for Pedestrian Attribute Recognition · AAAI 2024 |
Methods — techniques the papers use, named apart from their topics
logit alignment · 1.7angular margin · 1.7prompt tuning · 0.9multimodal interaction · 0.9information bottleneck · 0.9auxiliary-domain classification · 0.9vision transformer · 0.8orthogonal feature activation · 0.8convolutional neural network · 0.8bilinear pooling · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multimodal Fusion Using Multi-View Domains for Data Heterogeneity in Federated LearningabstractMultimodal information plays an important role in the advanced Internet of Things (IoT) in the era of 6G, which provides reliable and comprehensive assistance for downstream tasks through further fusion and analysis via federated learning (FL). One of the primary challenges in FL is data heterogeneity, which may lead to domain shifts and sharply different local long-tailed category distribution across nodes. These issues hinder the large-scale deployment of FL in IoT applications equipped with multiple various multimodal sensors due to performance deterioration. In this paper, we propose a novel multimodal fusion framework to tackle the aforementioned coupled problems arising during the cooperative fusion of multimodal information without privacy exposure among decentralized nodes equipped with diverse sensors. Specifically, we introduce a flexible global logit alignment (GLA) method based on multi-view domains. This method enables the fusion of diverse multimodal information with the consideration of domain shifts caused by modality-based data heterogeneity. Furthermore, we propose a novel local angular margin (LAM) scheme, which dynamically adjusts decision boundaries for locally seen categories while preserving global decision boundaries for unseen categories. This effectively mitigates severe model divergence caused by significantly different category distributions. Extensive simulations demonstrate the superiority of the proposed framework, which exhibits significant merits in tackling model degeneration caused by data heterogeneity and enhancing modality-based generalization for heterogeneous scenarios. Min Gao 0007, Haifeng Zheng, Xinxin Feng |
AAAI | 1 |
| 2025 | Enhanced Visual-Semantic Interaction with Tailored Prompts for Pedestrian Attribute RecognitionabstractPedestrian attribute recognition (PAR) seeks to predict multiple semantic attributes associated with a specific pedestrian. There are two types of approaches for PAR: unimodal framework and bimodal framework. The former one is to seek a robust visual feature. However, the lack of exploiting semantic feature of linguistic modality is the main concern. The latter one utilizes prompt learning techniques to integrate linguistic data. However, static prompt templates and simple bimodal concatenation cannot to capture the extensive intra-class attribute variability and support active modalities collaboration. In this paper, we propose an Enhanced Visual-Semantic Interaction with Tailored Prompts (EVSITP) framework for PAR. We present an Image-Conditional Dual-Prompt Initialization Module (IDIM) to adaptively generate context-sensitive prompts from visual inputs. Subsequently, a Prompt Enhanced and Regularization Module (PERM) is proposed to strengthen linguistic information from IDIM. We further design a Bimodal Mutual Interaction Module (BMIM) to ensure bidirectional modalities communication. In addition, existing PAR datasets are collected over a short period in limited scenarios, which do not align with real-world scenarios. Therefore, we annotate a long-term person re-identification dataset to create a new PAR dataset, Celeb-PAR. Experiments on several challenging PAR datasets show that our method outperforms state-of-the-art approaches. Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Yuzhong Chen 0001, Qiang Wu 0001 |
CVPR | 3 |
| 2025 | Layered Semi-Second-Order Information Bottleneck and Auxiliary Domain Classification for Person Re-Identification
Anguo Zhang, Junyi Wu 0001, Yueming Gao, Min Gao 0007, Yongduan Song 0001, Sio-Hang Pun |
Int. J. Comput. Vis. | 4 |
| 2025 | Rethinking attention mechanism for enhanced pedestrian attribute recognitionabstractPedestrian Attribute Recognition (PAR) plays a crucial role in various computer vision applications, demanding precise and reliable identification of attributes from pedestrian images. Traditional PAR methods, though effective in leveraging attention mechanisms, often suffer from the lack of direct supervision on attention, leading to potential overfitting and misallocation. This paper introduces a novel and model-agnostic approach, Attention-Aware Regularization (AAR), which rethinks the attention mechanism by integrating causal reasoning to provide direct supervision of attention maps. AAR employs perturbation techniques and a unique optimization objective to assess and refine attention quality, encouraging the model to prioritize attribute-specific regions. Our method demonstrates significant improvement in PAR performance by mitigating the effects of incorrect attention and fostering a more effective attention mechanism. Experiments on standard datasets showcase the superiority of our approach over existing methods, setting a new benchmark for attention-driven PAR models. Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Yuzhong Chen 0001, Qiang Wu 0001 |
Neurocomputing | 3 |
| 2025 | High-order diversity feature learning for pedestrian attribute recognitionabstractPedestrian attribute recognition (PAR) involves accurately identifying multiple attributes present in pedestrian images. There are two main approaches for PAR: part-based method and attention-based method. The former relies on existing segmentation or region detection methods to localize body parts and learn corresponding attribute-specific feature from the corresponding regions, where the performance heavily depends on the accuracy of body region localization. The latter adopts the embedded attention modules or transformer attention to exploit detailed feature. However, it can focus on certain body regions but often provide coarse attention, failing to capture fine-grained details, the learned feature may also be interfered with by irrelevant information. Meanwhile, these methods overlook the global contextual information. This work argues for replacing coarse attention with detailed attention and integrating it with global contextual feature from ViT to jointly represent attribute-specific regions. To tackle this issue, we propose a High-order Diversity Feature Learning (HDFL) method for PAR based on ViT. We utilize a polynomial predictor to design an Attribute-specific Detailed Feature Exploration (ADFE) module, which can construct the high-order statistics and gain more fine-grained feature. Our ADFE module is a parameter-friendly method that provides flexibility in deciding its utilization during the inference phase. A Soft-redundancy Perception Loss (SPLoss) is proposed to adaptively measure the redundancy between feature of different orders, which can promote diverse characterization of features. Experiments on several PAR datasets show that our method achieves a new state-of-the-art (SOTA) performance. On the most challenging PA100K dataset, our method outperforms previous SOTA by 1.69% and achieves the highest mA of 84.92%. Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Yuzhong Chen 0001, Qiang Wu 0001 |
Neural Networks | 3 |
| 2025 | Learning Comprehensive Representation via Selective Activation and Dual-Level Orthogonality for Pedestrian Attribute RecognitionabstractMulti-label Pedestrian Attribute Recognition (PAR) involves identifying a series of semantic attributes in person images. Existing PAR solutions typically rely on CNN as the backbone network to extract pedestrian features. Unfortunately, CNNs process only one adjacent region at a time, resulting in the disappearance of long-range relations between different attribute-specific regions. To address this limitation, we adopt the Vision Transformer (ViT) instead of CNN as the backbone for PAR, aiming to build long-range relations and extract more robust features. However, PAR suffers from an inherent attribute imbalance issue, causing ViT to naturally focus more on attributes that appear frequently in the training set and ignore some pedestrian attributes that appear less. The native features extracted by ViT are not able to tolerate the imbalance attribute distribution issue. To tackle this issue, we propose a novel component and a dual-level loss: the Selective Feature Activation Method (SFAM), the Orthogonal Feature Activation Loss (OFALoss), and Orthogonal Weight Regularization Loss (OWRLoss). SFAM smartly suppresses the more informative attribute-specific features, thus compelling the PAR model to pay greater attention to attribute-specific regions that are often overlooked. The proposed OFALoss enforces an orthogonal constraint on the original feature extracted by ViT and the suppressed features from SFAM, promoting the comprehensiveness of feature representation in each attribute-specific region. Furthermore, OWRLoss is employed for decreasing correlations among entries of the last shared classification layer, which can alleviate the highly correlated of weight vectors caused by non-uniform distribution. This can prevent excessive mutual interference among different attributes during attribute recognition. Our model-agnostic approach is plug-and-play, requiring no additional training parameters in the training process. We conduct experiments on several benchmark PAR datasets, including PETA, PA100K, RAPv1, and RAPv2, demonstrating the effectiveness of our method. Specifically, our method outperforms existing state-of-the-art approaches. Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Yuzhong Chen 0001, Qiang Wu 0001, Jianqiang Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Selective and Orthogonal Feature Activation for Pedestrian Attribute RecognitionabstractPedestrian Attribute Recognition (PAR) involves identifying the attributes of individuals in person images. Existing PAR methods typically rely on CNNs as the backbone network to extract pedestrian features. However, CNNs process only one adjacent region at a time, leading to the loss of long-range inter-relations between different attribute-specific regions. To address this limitation, we leverage the Vision Transformer (ViT) instead of CNNs as the backbone for PAR, aiming to model long-range relations and extract more robust features. However, PAR suffers from an inherent attribute imbalance issue, causing ViT to naturally focus more on attributes that appear frequently in the training set and ignore some pedestrian attributes that appear less. The native features extracted by ViT are not able to tolerate the imbalance attribute distribution issue. To tackle this issue, we propose two novel components: the Selective Feature Activation Method (SFAM) and the Orthogonal Feature Activation Loss. SFAM smartly suppresses the more informative attribute-specific features, compelling the PAR model to capture discriminative features from regions that are easily overlooked. The proposed loss enforces an orthogonal constraint on the original feature extracted by ViT and the suppressed features from SFAM, promoting the complementarity of features in space. We conduct experiments on several benchmark PAR datasets, including PETA, PA100K, RAPv1, and RAPv2, demonstrating the effectiveness of our method. Specifically, our method outperforms existing state-of-the-art approaches by GRL, IAA-Caps, ALM, and SSC in terms of mA on the four datasets, respectively. Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Jianqiang Zhao |
AAAI | 3 |
| 2024 | Adaptive Decentralized Federated Learning in Resource-Constrained IoT NetworksabstractDecentralized federated learning (DFL) is a novel distributed machine-learning paradigm where participants collaborate to train machine-learning models without the assistance of the central server. The decentralized framework can effectively overcome the communication bottleneck and single-point-of-failure issues encountered in federated learning (FL). However, most existing DFL methods may ignore the communication resource constraints of the system. This may result in these methods unsuitable for many practical scenarios because the given resource constraints cannot be guaranteed. In this article, we propose a novel DFL, called DFL with adaptive compression ratio (AdapCom-DFL), that can adaptively adjust the compression ratio of transmission data to keep the communication latency within the constraint. Furthermore, we propose a communication network topology pruning approach to reduce communication overhead by pruning poor links with low data rates while ensuring the convergence. Additionally, a power allocation approach is presented to improve the performance by reallocating the power of communication links while complying with the communication energy constraint. Extensive simulation results demonstrate that the proposed AdapCom-DFL with network pruning and power allocation approach achieves better performance and requires less bandwidth under the given resource constraints compared with some existing approaches. Mengxuan Du, Haifeng Zheng, Min Gao 0007, Xinxin Feng |
IEEE Internet Things J. | 3 |
| 2024 | Multimodal Fusion With Block Term Decomposition for Asynchronous Federated LearningabstractFederated learning (FL) has been extensively studied as a means of ensuring data privacy while cooperatively training a global model across decentralized devices. Among various FL approaches, asynchronous federated learning (AFL) has distinct advantages in overcoming the straggler problem via server-side aggregation as soon as it receives a local model. However, AFL still faces several challenges in large-scale real-world applications, such as stale model problems and modality heterogeneity across geographically distributed and industrial devices with different functions. In this article, we propose a multimodal fusion framework for AFL to address the aforementioned problems. Specifically, a novel multilinear block fusion model is designed to fuse various multimodal information, which serves as an enhancement for perceiving and transmitting the important modality and block during local training. An adaptive aggregation strategy is further developed to fully utilize heterogeneous data by allowing the global model to favor the received local model based on both freshness and the importance of the local data. Extensive simulations with different data distributions demonstrate the superiority of the proposed framework in heterogeneity scenarios, which exhibits significant merits in the improvement of modality-based generalization without sacrificing convergence speed and communication consumption. Min Gao 0007, Haifeng Zheng, Mengxuan Du, Xinxin Feng |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | A Two-Stream Hybrid Convolution-Transformer Network Architecture for Clothing-Change Person Re-IdentificationabstractLong-term (also called Clothing-Change) person re-identification (CC-reID) aims at confirming the identity of pedestrians captured at diverse locations and/or times. Current CC-reID methods heavily rely on ID features learned by the CNN architecture. However, with limited receptive fields, CNN is hard to effectively explore some unique but discriminative ID features (e.g., hair style, tattoo and accessories) from small body regions. Compared with CNN, Transformer has certain merits in exploring more diverse ID-unique features1and retaining more details by the multi-head self-attention design and the removal of down-sampling operation. In this paper, a two-stream hybrid Convolution-Transformer Network (CT-Net) is proposed for CC-reID by combining both CNN and Transformer parallelly in an end-to-end learning scheme. Specifically, CT-Net contains a CNN-based stream (C-Stream) and a Transformer-based stream (T-Stream). Compared with using C-Stream only, T-Stream is used to encourage the C-Stream to explore more detailed ID-unique features when the clothing information is no reliable in CC-reID. Specifically, a Feature Supplement Module (FSM) is proposed to transfer features learned by T-Stream to C-Stream from low-level to high-level for mining more ID-unique feature. In order to further enhance the discriminability2and complementary of ID features learned by our CT-Net, we also introduce a hierarchical supervision with bilinear pooling (HSBP). Experimental results demonstrate that CT-Net performs favorably against the state-of-the-art methods over three CC-reID benchmarks. Meanwhile, CT-Net also demonstrates good generalization ability by achieving comparable performance on traditional person re-ID datasets such as Market-1501 and DukeMTMC-reID. Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Jianqiang Zhao, Huiji Zhang, Anguo Zhang |
IEEE Trans. Multim. | 3 |
| 2023 | Exponential Information Bottleneck Theory Against Intra-Attribute Variations for Pedestrian Attribute RecognitionabstractMulti-label pedestrian attribute recognition (PAR) involves assigning multiple attributes to pedestrian images captured by video surveillance cameras. Despite its importance, learning robust attribute-related features for PAR remains a challenge due to the large intra-attribute variations in the image space. These variations, which stem from changes in pedestrian poses, illumination conditions, and background noise, make extracted attribute-related features susceptible to irrelevant information or noise interference. Existing PAR methods rely on body prior extractors or attention mechanisms to locate attribute-correlation regions for extracting robust features. However, these methods may not be robust to intra-attribute variations, which limits their effectiveness. To address this challenge, we propose a novel and flexible PAR framework that leverages the exponential information bottleneck (ExpIB) approach. Our ExpIB-Net uses mutual information compression as the main penalty during the early stage of training, thereby eliminating irrelevant information. As training progresses, the mutual information penalty weakens and the Binary Cross-Entropy Loss (BCELoss) contributes to improving the PAR recognition accuracy. Our method can also be integrated into an attention module to form the AttExpIB-Net, which better handles intra-attribute variations for better performance. Additionally, our model-agnostic ExpIB approach is plug-and-play, requiring no additional computational overhead during inference. Experiments on several challenging PAR datasets show that our method outperforms state-of-the-art approaches. Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Jianqiang Zhao, Jieming Shi 0001, Anguo Zhang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | A Distributed Hierarchical Deep Computation Model for Federated Learning in Edge ComputingabstractDeep learning has recently garnered significant interest in many applications especially for big data analytics in the edge computing environment. Federated learning, as a novel machine learning technique, aims to build a shared learning model from training data on distributed edge nodes to protect data privacy. However, the model update in federated learning requires parameter exchanges among edge nodes, which is rather bandwidth-consuming. This article proposes a novel distributed hierarchical tensor deep computation model by condensing the model parameters from a high-dimensional tensor space into a set of low-dimensional subspaces to reduce the bandwidth consumption and storage requirement for federated learning. Moreover, an updating approach with a hierarchical tensor back-propagation algorithm is developed by directly computing the gradients of low-dimensional parameters to reduce the memory requirement of training for edge nodes and improve training efficiency. Finally, extensive simulations on classical datasets with different local data distributions are presented for the performance evaluation. The results demonstrate that the proposed model relieves the burden of communication bandwidth and reduces energy consumption at edge nodes for federated learning. Haifeng Zheng, Min Gao 0007, Zhizhang (David) Chen, Xinxin Feng |
IEEE Trans. Ind. Informatics | 2 |
| 2019 | An Adaptive Sampling Scheme via Approximate Volume Sampling for Fingerprint-Based Indoor LocalizationabstractIn recent years Wi-Fi fingerprinting has attracted much attention in indoor localization because of the availability of high-quality signal and pervasive deployment of wireless LANs. For fingerprint-based localization, however, offline site survey is usually time-consuming and labor-intensive. Therefore, reducing the burden of offline site survey becomes an important issue for fingerprint-based indoor localization. In this paper, using a low-tubal-rank tensor to model Wi-Fi fingerprints of all reference points (RPs), we propose an adaptive sampling scheme via approximate volume sampling to improve reconstruction accuracy of radio map with reduced expenditure. We propose a rank-increasing strategy to effectively estimate the rank of the underlying fingerprint tensor to alleviate the computation burden for tensor completion. We provide a theoretical foundation to analyze the proposed scheme and derive the performance bounds in terms of sample complexity and reconstruction error. We prove that the proposed scheme can achieve a relative error guarantee. Finally, we validate the effectiveness of the proposed scheme through extensive simulations using both synthetic and real datasets. The simulation results demonstrate that the proposed scheme is able to not only reduce reconstruction error and improve localization accuracy but also reduce running time compared to the state-of-the-art schemes. Haifeng Zheng, Min Gao 0007, Zhizhang (David) Chen, Xiao-Yang Liu, Xinxin Feng |
IEEE Internet Things J. | 2 |