Yunpeng Gong

dblp:219/2474 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-6498-2555ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Face, body and person analysis · 54% Transfer learning and domain adaptation · 15% Graph learning · 15%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 61% Computer animation and physical simulation · 30% Geometric modeling and processing · 9%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis
person re-identification
1.822026
A Theory-Inspired Framework for Few-Shot Cross-Modal Sketch Person Re-Identification · AAAI 2026
Cross-Modality Perturbation Synergy Attack for Person Re-identification · NeurIPS 2024
Machine learning › Transfer learning and domain adaptation
domain generalization
1.012026
A Theory-Inspired Framework for Few-Shot Cross-Modal Sketch Person Re-Identification · AAAI 2026
Computer vision › Segmentation and scene understanding
image segmentation
1.012026
Physical Regularization Loss: Integrating Physical Knowledge to Image Segmentation · Int. J. Comput. Vis. 2026
Machine learning › Graph learning › graph neural network › graph neural network architecture
physics-informed graph neural networks
1.012026
PEGNet: A Physics-Embedded Graph Network for Long-Term Stable Multiphysics Simulation · AAAI 2026
Computer vision › Face, body and person analysis › person re-identification › multi-modal person re-identification
sketch re-identification
1.012026
A Theory-Inspired Framework for Few-Shot Cross-Modal Sketch Person Re-Identification · AAAI 2026
Computational science and engineering
multiphysics simulation
1.012026
PEGNet: A Physics-Embedded Graph Network for Long-Term Stable Multiphysics Simulation · AAAI 2026
Computational science and engineering › computational physics
physics simulation
1.012026
PEGNet: A Physics-Embedded Graph Network for Long-Term Stable Multiphysics Simulation · AAAI 2026
Visual content generation and editing
3d content generation
0.912025
Phys4DGen: Physics-Compliant 4D Generation with Multi-Material Composition Perception · ACM Multimedia 2025
Visual content generation and editing › 3d content generation
4d content generation
0.912025
Phys4DGen: Physics-Compliant 4D Generation with Multi-Material Composition Perception · ACM Multimedia 2025
Computer vision › Face, body and person analysis › person re-identification › multi-modal person re-identification
visible-infrared person re-identification
0.812024
Cross-Modality Perturbation Synergy Attack for Person Re-identification · NeurIPS 2024
Security and privacy of machine learning
adversarial attack
0.812024
Cross-Modality Perturbation Synergy Attack for Person Re-identification · NeurIPS 2024
Security and privacy of machine learning
adversarial robustness
0.812024
Cross-Modality Perturbation Synergy Attack for Person Re-identification · NeurIPS 2024
Security and privacy of machine learning › adversarial attack › universal attack
universal perturbation
0.812024
Cross-Modality Perturbation Synergy Attack for Person Re-identification · NeurIPS 2024
Computational science and engineering › scientific machine learning › physics-informed machine learning › physics-informed neural networks
partial differential equation solving
0.312026
PEGNet: A Physics-Embedded Graph Network for Long-Term Stable Multiphysics Simulation · AAAI 2026
Geometric modeling and processing
shape representation
0.312025
Phys4DGen: Physics-Compliant 4D Generation with Multi-Material Composition Perception · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

physics-embedded message passing · 2.0physical regularization · 2.0graph neural network · 2.0gradient-based perturbation optimization · 1.5physical regularization loss · 1.0meta-learning · 1.0contrastive alignment · 1.0adversarial perturbation · 1.0semantic segmentation · 0.9physical simulation · 0.9multimodal large language model · 0.9
YearPublicationVenuePosition
2026 A Theory-Inspired Framework for Few-Shot Cross-Modal Sketch Person Re-Identification
abstract
Sketch-based person re-identification aims to match hand-drawn sketches with RGB surveillance images, but remains challenging due to severe modality gaps and limited labeled data. To address this, we propose KTCAA, a theoretically inspired framework for few-shot cross-modal generalization. Drawing on generalization bounds, we identify two key factors affecting target risk: (1) domain discrepancy, reflecting the alignment difficulty between source and target distributions; and (2) perturbation invariance, measuring the model’s robustness to modality shifts. Accordingly, we design: (1) Alignment Augmentation (AA), which applies localized sketch-style transformations to simulate target distributions and guide progressive alignment; and (2) Knowledge Transfer Catalyst (KTC), which enhances perturbation invariance by introducing worst-case modality perturbations and enforcing consistency. These modules are jointly optimized within a meta-learning paradigm that transfers alignment knowledge from data-abundant RGB domains to sketch scenarios. Experiments on multiple benchmarks show that KTCAA achieves state-of-the-art performance, particularly under data-scarce conditions.
Yunpeng Gong, Yongjie Hou, Jiangming Shi, Kim Long Diep, Min Jiang 0005
AAAI1
2026 PEGNet: A Physics-Embedded Graph Network for Long-Term Stable Multiphysics Simulation
abstract
Accurate and efficient simulations of physical phenomena governed by partial differential equations (PDEs) are important for scientific and engineering progress. While traditional numerical solvers are powerful, they are often computationally expensive. Recently, data-driven methods have emerged as alternatives, but they frequently suffer from error accumulation and limited physical consistency, especially in multiphysics and complex geometries. To address these challenges, we propose PEGNet, a Physics-Embedded Graph Network that incorporates PDE-guided message passing to redesign the graph neural network architecture. By embedding key PDE dynamics like convection, viscosity, and diffusion into distinct message functions, the model naturally integrates physical constraints into its forward propagation, producing more stable and physically consistent solutions. Additionally, a hierarchical architecture is employed to capture multi-scale features, and physical regularization is integrated into the loss function to further enforce adherence to governing physics. We evaluated PEGNet on benchmarks, including custom datasets for respiratory airflow and drug delivery, showing significant improvements in long-term prediction accuracy and physical consistency over existing methods.
Zhenzhong Wang, Junyuan Liu, Yunpeng Gong, Min Jiang 0005
AAAI4
2026 Evaluating and predicting green economic efficiency in Chinese cities: A three-stage network SBM and machine learning approach
Yunpeng Gong, Huayong Niu
Expert Syst. Appl.3
2026 Physical Regularization Loss: Integrating Physical Knowledge to Image Segmentation
Huafeng Li 0001, Guanqiu Qi, Baisen Cong, Yunpeng Gong, Zhiqin Zhu
Int. J. Comput. Vis.6
2025 Phys4DGen: Physics-Compliant 4D Generation with Multi-Material Composition Perception
abstract
4D content generation aims to create dynamically evolving 3D content that responds to specific input objects such as images or 3D representations. Current approaches typically incorporate physical priors to animate 3D representations, but these methods suffer from significant limitations: they not only require users lacking physics expertise to manually specify material properties but also struggle to effectively handle the generation of multi-material composite objects. To address these challenges, we propose Phys4DGen, a novel 4D generation framework that integrates multi-material composition perception with physical simulation. The framework achieves automated, physically plausible 4D generation through three innovative modules: first, the 3D Material Grouping module partitions heterogeneous material regions on 3D representations' surfaces via semantic segmentation; second, the Internal Physical Structure Discovery module constructs the mechanical structure of object interiors; finally, we distill physical prior knowledge from multimodal large language models to enable rapid and automatic material properties identification for both objects' surfaces and interiors. Experiments on both synthetic and real-world datasets demonstrate that Phys4DGen can generate high-fidelity 4D content with physical realism in open-world scenarios, significantly outperforming state-of-the-art methods.
Jiajing Lin, Zhenzhong Wang, Dejun Xu, Yunpeng Gong, Min Jiang 0005
ACM Multimedia5
2024 Beyond Augmentation: Empowering Model Robustness under Extreme Capture Environments
abstract
Person Re-identification (re-ID) in computer vision aims to recognize and track individuals across different cameras. While previous research has mainly focused on challenges like pose variations and lighting changes, the impact of extreme capture conditions is often not adequately addressed. These extreme conditions, including varied lighting, camera styles, angles, and image distortions, can significantly affect data distribution and re-ID accuracy.Current research typically improves model generalization under normal shooting conditions through data augmentation techniques such as adjusting brightness and contrast. However, these methods pay less attention to the robustness of models under extreme shooting conditions. To tackle this, we propose a multi-mode synchronization learning (MMSL) strategy . This approach involves dividing images into grids, randomly selecting grid blocks, and applying data augmentation methods like contrast and brightness adjustments. This process introduces diverse transformations without altering the original image structure, helping the model adapt to extreme variations. This method improves the model’s generalization under extreme conditions and enables learning diverse features, thus better addressing the challenges in re-ID. Extensive experiments on a simulated test set under extreme conditions have demonstrated the effectiveness of our method. This approach is crucial for enhancing model robustness and adaptability in real-world scenarios, supporting the future development of person re-identification technology.
Yunpeng Gong, Yongjie Hou, Chuangliang Zhang, Min Jiang 0005
IJCNN1
2024 Beyond Dropout: Robust Convolutional Neural Networks Based on Local Feature Masking
abstract
In the contemporary of deep learning, where models often grapple with the challenge of simultaneously achieving robustness against adversarial attacks and strong generalization capabilities, this study introduces an innovative Local Feature Masking (LFM) strategy aimed at fortifying the performance of Convolutional Neural Networks (CNNs) on both fronts. During the training phase, we strategically incorporate random feature masking in the shallow layers of CNNs, effectively alleviating overfitting issues, thereby enhancing the model’s generalization ability and bolstering its resilience to adversarial attacks. LFM compels the network to adapt by leveraging remaining features to compensate for the absence of certain semantic features, nurturing a more elastic feature learning mechanism. The efficacy of LFM is substantiated through a series of quantitative and qualitative assessments, collectively showcasing a consistent and significant improvement in CNN’s generalization ability and resistance against adversarial attacks—a phenomenon not observed in current and prior methodologies. The seamless integration of LFM into established CNN frameworks underscores its potential to advance both generalization and adversarial robustness within the deep learning paradigm. Through comprehensive experiments, including robust person re-identification baseline generalization experiments and adversarial attack experiments, we demonstrate the substantial enhancements offered by LFM in addressing the aforementioned challenges. This contribution represents a noteworthy stride in advancing robust neural network architectures.
Yunpeng Gong, Chuangliang Zhang, Yongjie Hou, Lifei Chen, Min Jiang 0005
IJCNN1
2024 Cross-Task Attack: A Self-Supervision Generative Framework Based on Attention Shift
abstract
Studying adversarial attacks on artificial intelligence (AI) systems helps discover model shortcomings, enabling the construction of a more robust system. Most existing adversarial attack methods only concentrate on single-task single-model or single-task cross-model scenarios, overlooking the multi-task characteristic of artificial intelligence systems. As a result, most of the existing attacks do not pose a practical threat to a comprehensive and collaborative AI system. However, implementing cross-task attacks is highly demanding and challenging due to the difficulty in obtaining the real labels of different tasks for the same picture and harmonizing the loss functions across different tasks. To address this issue, we propose a self-supervised Cross-Task Attack framework (CTA), which utilizes co-attention and anti-attention maps to generate cross-task adversarial perturbation. Specifically, the co-attention map reflects the area to which different visual task models pay attention, while the anti-attention map reflects the area that different visual task models neglect. CTA generates cross-task perturbations by shifting the attention area of samples away from the co-attention map and closer to the anti-attention map. We conduct extensive experiments on multiple vision tasks and the experimental results confirm the effectiveness of the proposed design for adversarial attacks.
Qingyuan Zeng, Yunpeng Gong, Min Jiang 0005
IJCNN2
2024 Cross-Modality Perturbation Synergy Attack for Person Re-identification
abstract
In recent years, there has been significant research focusing on addressing security concerns in single-modal person re-identification (ReID) systems that are based on RGB images. However, the safety of cross-modality scenarios, which are more commonly encountered in practical applications involving images captured by infrared cameras, has not received adequate attention. The main challenge in cross-modality ReID lies in effectively dealing with visual differences between different modalities. For instance, infrared images are typically grayscale, unlike visible images that contain color information. Existing attack methods have primarily focused on the characteristics of the visible image modality, overlooking the features of other modalities and the variations in data distribution among different modalities. This oversight can potentially undermine the effectiveness of these methods in image retrieval across diverse modalities. This study represents the first exploration into the security of cross-modality ReID models and proposes a universal perturbation attack specifically designed for cross-modality ReID. This attack optimizes perturbations by leveraging gradients from diverse modality data, thereby disrupting the discriminator and reinforcing the differences between modalities. We conducted experiments on three widely used cross-modality datasets, namely RegDB, SYSU, and LLCM. The results not only demonstrate the effectiveness of our method but also provide insights for future improvements in the robustness of cross-modality ReID systems.
Yunpeng Gong, Zhun Zhong, Yansong Qu, Zhiming Luo, Rongrong Ji, Min Jiang 0005
NeurIPS1