VLDB 2026 Research / reviewers in the wild / expert
Yunpeng Gong
dblp:219/2474
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-6498-2555ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Face, body and person analysis · 54% Transfer learning and domain adaptation · 15% Graph learning · 15% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 61% Computer animation and physical simulation · 30% Geometric modeling and processing · 9% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Face, body and person analysis
person re-identification |
1.8 | 2 | 2026 | A Theory-Inspired Framework for Few-Shot Cross-Modal Sketch Person Re-Identification · AAAI 2026 Cross-Modality Perturbation Synergy Attack for Person Re-identification · NeurIPS 2024 |
Machine learning › Transfer learning and domain adaptation
domain generalization |
1.0 | 1 | 2026 | A Theory-Inspired Framework for Few-Shot Cross-Modal Sketch Person Re-Identification · AAAI 2026 |
Computer vision › Segmentation and scene understanding
image segmentation |
1.0 | 1 | 2026 | Physical Regularization Loss: Integrating Physical Knowledge to Image Segmentation · Int. J. Comput. Vis. 2026 |
Machine learning › Graph learning › graph neural network › graph neural network architecture
physics-informed graph neural networks |
1.0 | 1 | 2026 | PEGNet: A Physics-Embedded Graph Network for Long-Term Stable Multiphysics Simulation · AAAI 2026 |
Computer vision › Face, body and person analysis › person re-identification › multi-modal person re-identification
sketch re-identification |
1.0 | 1 | 2026 | A Theory-Inspired Framework for Few-Shot Cross-Modal Sketch Person Re-Identification · AAAI 2026 |
Computational science and engineering
multiphysics simulation |
1.0 | 1 | 2026 | PEGNet: A Physics-Embedded Graph Network for Long-Term Stable Multiphysics Simulation · AAAI 2026 |
Computational science and engineering › computational physics
physics simulation |
1.0 | 1 | 2026 | PEGNet: A Physics-Embedded Graph Network for Long-Term Stable Multiphysics Simulation · AAAI 2026 |
Visual content generation and editing
3d content generation |
0.9 | 1 | 2025 | Phys4DGen: Physics-Compliant 4D Generation with Multi-Material Composition Perception · ACM Multimedia 2025 |
Visual content generation and editing › 3d content generation
4d content generation |
0.9 | 1 | 2025 | Phys4DGen: Physics-Compliant 4D Generation with Multi-Material Composition Perception · ACM Multimedia 2025 |
Computer vision › Face, body and person analysis › person re-identification › multi-modal person re-identification
visible-infrared person re-identification |
0.8 | 1 | 2024 | Cross-Modality Perturbation Synergy Attack for Person Re-identification · NeurIPS 2024 |
Security and privacy of machine learning
adversarial attack |
0.8 | 1 | 2024 | Cross-Modality Perturbation Synergy Attack for Person Re-identification · NeurIPS 2024 |
Security and privacy of machine learning
adversarial robustness |
0.8 | 1 | 2024 | Cross-Modality Perturbation Synergy Attack for Person Re-identification · NeurIPS 2024 |
Security and privacy of machine learning › adversarial attack › universal attack
universal perturbation |
0.8 | 1 | 2024 | Cross-Modality Perturbation Synergy Attack for Person Re-identification · NeurIPS 2024 |
Computational science and engineering › scientific machine learning › physics-informed machine learning › physics-informed neural networks
partial differential equation solving |
0.3 | 1 | 2026 | PEGNet: A Physics-Embedded Graph Network for Long-Term Stable Multiphysics Simulation · AAAI 2026 |
Geometric modeling and processing
shape representation |
0.3 | 1 | 2025 | Phys4DGen: Physics-Compliant 4D Generation with Multi-Material Composition Perception · ACM Multimedia 2025 |
Methods — techniques the papers use, named apart from their topics
physics-embedded message passing · 2.0physical regularization · 2.0graph neural network · 2.0gradient-based perturbation optimization · 1.5physical regularization loss · 1.0meta-learning · 1.0contrastive alignment · 1.0adversarial perturbation · 1.0semantic segmentation · 0.9physical simulation · 0.9multimodal large language model · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Theory-Inspired Framework for Few-Shot Cross-Modal Sketch Person Re-IdentificationabstractSketch-based person re-identification aims to match hand-drawn sketches with RGB surveillance images, but remains challenging due to severe modality gaps and limited labeled data. To address this, we propose KTCAA, a theoretically inspired framework for few-shot cross-modal generalization. Drawing on generalization bounds, we identify two key factors affecting target risk: (1) domain discrepancy, reflecting the alignment difficulty between source and target distributions; and (2) perturbation invariance, measuring the model’s robustness to modality shifts. Accordingly, we design: (1) Alignment Augmentation (AA), which applies localized sketch-style transformations to simulate target distributions and guide progressive alignment; and (2) Knowledge Transfer Catalyst (KTC), which enhances perturbation invariance by introducing worst-case modality perturbations and enforcing consistency. These modules are jointly optimized within a meta-learning paradigm that transfers alignment knowledge from data-abundant RGB domains to sketch scenarios. Experiments on multiple benchmarks show that KTCAA achieves state-of-the-art performance, particularly under data-scarce conditions. Yunpeng Gong, Yongjie Hou, Jiangming Shi, Kim Long Diep, Min Jiang 0005 |
AAAI | 1 |
| 2026 | PEGNet: A Physics-Embedded Graph Network for Long-Term Stable Multiphysics SimulationabstractAccurate and efficient simulations of physical phenomena governed by partial differential equations (PDEs) are important for scientific and engineering progress. While traditional numerical solvers are powerful, they are often computationally expensive. Recently, data-driven methods have emerged as alternatives, but they frequently suffer from error accumulation and limited physical consistency, especially in multiphysics and complex geometries. To address these challenges, we propose PEGNet, a Physics-Embedded Graph Network that incorporates PDE-guided message passing to redesign the graph neural network architecture. By embedding key PDE dynamics like convection, viscosity, and diffusion into distinct message functions, the model naturally integrates physical constraints into its forward propagation, producing more stable and physically consistent solutions. Additionally, a hierarchical architecture is employed to capture multi-scale features, and physical regularization is integrated into the loss function to further enforce adherence to governing physics. We evaluated PEGNet on benchmarks, including custom datasets for respiratory airflow and drug delivery, showing significant improvements in long-term prediction accuracy and physical consistency over existing methods. Zhenzhong Wang, Junyuan Liu, Yunpeng Gong, Min Jiang 0005 |
AAAI | 4 |
| 2026 | Evaluating and predicting green economic efficiency in Chinese cities: A three-stage network SBM and machine learning approach
Yunpeng Gong, Huayong Niu |
Expert Syst. Appl. | 3 |
| 2026 | Physical Regularization Loss: Integrating Physical Knowledge to Image Segmentation
Huafeng Li 0001, Guanqiu Qi, Baisen Cong, Yunpeng Gong, Zhiqin Zhu |
Int. J. Comput. Vis. | 6 |
| 2025 | Phys4DGen: Physics-Compliant 4D Generation with Multi-Material Composition Perceptionabstract4D content generation aims to create dynamically evolving 3D content that responds to specific input objects such as images or 3D representations. Current approaches typically incorporate physical priors to animate 3D representations, but these methods suffer from significant limitations: they not only require users lacking physics expertise to manually specify material properties but also struggle to effectively handle the generation of multi-material composite objects. To address these challenges, we propose Phys4DGen, a novel 4D generation framework that integrates multi-material composition perception with physical simulation. The framework achieves automated, physically plausible 4D generation through three innovative modules: first, the 3D Material Grouping module partitions heterogeneous material regions on 3D representations' surfaces via semantic segmentation; second, the Internal Physical Structure Discovery module constructs the mechanical structure of object interiors; finally, we distill physical prior knowledge from multimodal large language models to enable rapid and automatic material properties identification for both objects' surfaces and interiors. Experiments on both synthetic and real-world datasets demonstrate that Phys4DGen can generate high-fidelity 4D content with physical realism in open-world scenarios, significantly outperforming state-of-the-art methods. Jiajing Lin, Zhenzhong Wang, Dejun Xu, Yunpeng Gong, Min Jiang 0005 |
ACM Multimedia | 5 |
| 2024 | Beyond Augmentation: Empowering Model Robustness under Extreme Capture EnvironmentsabstractPerson Re-identification (re-ID) in computer vision aims to recognize and track individuals across different cameras. While previous research has mainly focused on challenges like pose variations and lighting changes, the impact of extreme capture conditions is often not adequately addressed. These extreme conditions, including varied lighting, camera styles, angles, and image distortions, can significantly affect data distribution and re-ID accuracy.Current research typically improves model generalization under normal shooting conditions through data augmentation techniques such as adjusting brightness and contrast. However, these methods pay less attention to the robustness of models under extreme shooting conditions. To tackle this, we propose a multi-mode synchronization learning (MMSL) strategy . This approach involves dividing images into grids, randomly selecting grid blocks, and applying data augmentation methods like contrast and brightness adjustments. This process introduces diverse transformations without altering the original image structure, helping the model adapt to extreme variations. This method improves the model’s generalization under extreme conditions and enables learning diverse features, thus better addressing the challenges in re-ID. Extensive experiments on a simulated test set under extreme conditions have demonstrated the effectiveness of our method. This approach is crucial for enhancing model robustness and adaptability in real-world scenarios, supporting the future development of person re-identification technology. Yunpeng Gong, Yongjie Hou, Chuangliang Zhang, Min Jiang 0005 |
IJCNN | 1 |
| 2024 | Beyond Dropout: Robust Convolutional Neural Networks Based on Local Feature MaskingabstractIn the contemporary of deep learning, where models often grapple with the challenge of simultaneously achieving robustness against adversarial attacks and strong generalization capabilities, this study introduces an innovative Local Feature Masking (LFM) strategy aimed at fortifying the performance of Convolutional Neural Networks (CNNs) on both fronts. During the training phase, we strategically incorporate random feature masking in the shallow layers of CNNs, effectively alleviating overfitting issues, thereby enhancing the model’s generalization ability and bolstering its resilience to adversarial attacks. LFM compels the network to adapt by leveraging remaining features to compensate for the absence of certain semantic features, nurturing a more elastic feature learning mechanism. The efficacy of LFM is substantiated through a series of quantitative and qualitative assessments, collectively showcasing a consistent and significant improvement in CNN’s generalization ability and resistance against adversarial attacks—a phenomenon not observed in current and prior methodologies. The seamless integration of LFM into established CNN frameworks underscores its potential to advance both generalization and adversarial robustness within the deep learning paradigm. Through comprehensive experiments, including robust person re-identification baseline generalization experiments and adversarial attack experiments, we demonstrate the substantial enhancements offered by LFM in addressing the aforementioned challenges. This contribution represents a noteworthy stride in advancing robust neural network architectures. Yunpeng Gong, Chuangliang Zhang, Yongjie Hou, Lifei Chen, Min Jiang 0005 |
IJCNN | 1 |
| 2024 | Cross-Task Attack: A Self-Supervision Generative Framework Based on Attention ShiftabstractStudying adversarial attacks on artificial intelligence (AI) systems helps discover model shortcomings, enabling the construction of a more robust system. Most existing adversarial attack methods only concentrate on single-task single-model or single-task cross-model scenarios, overlooking the multi-task characteristic of artificial intelligence systems. As a result, most of the existing attacks do not pose a practical threat to a comprehensive and collaborative AI system. However, implementing cross-task attacks is highly demanding and challenging due to the difficulty in obtaining the real labels of different tasks for the same picture and harmonizing the loss functions across different tasks. To address this issue, we propose a self-supervised Cross-Task Attack framework (CTA), which utilizes co-attention and anti-attention maps to generate cross-task adversarial perturbation. Specifically, the co-attention map reflects the area to which different visual task models pay attention, while the anti-attention map reflects the area that different visual task models neglect. CTA generates cross-task perturbations by shifting the attention area of samples away from the co-attention map and closer to the anti-attention map. We conduct extensive experiments on multiple vision tasks and the experimental results confirm the effectiveness of the proposed design for adversarial attacks. Qingyuan Zeng, Yunpeng Gong, Min Jiang 0005 |
IJCNN | 2 |
| 2024 | Cross-Modality Perturbation Synergy Attack for Person Re-identificationabstractIn recent years, there has been significant research focusing on addressing security concerns in single-modal person re-identification (ReID) systems that are based on RGB images. However, the safety of cross-modality scenarios, which are more commonly encountered in practical applications involving images captured by infrared cameras, has not received adequate attention. The main challenge in cross-modality ReID lies in effectively dealing with visual differences between different modalities. For instance, infrared images are typically grayscale, unlike visible images that contain color information. Existing attack methods have primarily focused on the characteristics of the visible image modality, overlooking the features of other modalities and the variations in data distribution among different modalities. This oversight can potentially undermine the effectiveness of these methods in image retrieval across diverse modalities. This study represents the first exploration into the security of cross-modality ReID models and proposes a universal perturbation attack specifically designed for cross-modality ReID. This attack optimizes perturbations by leveraging gradients from diverse modality data, thereby disrupting the discriminator and reinforcing the differences between modalities. We conducted experiments on three widely used cross-modality datasets, namely RegDB, SYSU, and LLCM. The results not only demonstrate the effectiveness of our method but also provide insights for future improvements in the robustness of cross-modality ReID systems. Yunpeng Gong, Zhun Zhong, Yansong Qu, Zhiming Luo, Rongrong Ji, Min Jiang 0005 |
NeurIPS | 1 |