VLDB 2026 Research / reviewers in the wild / expert
Yuanpeng Tu
dblp:301/9706
· DBLP profile ↗
11ranked-venue papers
7as first author
11since 2021 · last 2026
0009-0006-2978-666XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 7 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PlayLife: Tracking-Shot Video GenerationabstractAbstract Tracking shot videos, where a third person camera follows a moving subject, are central to games and cinematography and are increasingly useful for simulation, content creation, and embodied evaluation. However, existing controllable video generators often rely on text or coarse global cues. In addition, their motion transfer pipelines typically assume a fixed viewpoint, and game specific systems frequently discretize movement. Consequently, producing third person videos with free human level motion and consistent camera tracking from minimal inputs remains underexplored. To address these issues, we present PlayLife, a tracking shot image to video generator. Specifically, given a single exocentric image and a human motion sequence, it synthesizes third person videos in which a virtual camera consistently follows the actor while the generated motion adheres to the provided actions across AAA game and real world scenes. At the core is a geometry-aware action injection module. Concretely, articulated keypoints from the motion sequence are projected into the first frame using estimated camera parameters and are fused with video latents through cross attention. Through this design, the model achieves part level correspondence, stable trajectories, and scene aware view control. To learn from scarce and noisy supervision, we adopt hierarchical training . In the first stage, lightweight temporal adapters are pretrained on a large corpus with paired motion and video data; in the second stage, they are finetuned on a curated subset with a pose aware reconstruction loss that emphasizes articulated foreground regions, thereby improving temporal synchronization and motion fidelity. Finally, we curate tracking shot data and establish evaluation protocols in both game and real world settings. Furthermore, experiments with visualizations, automatic metrics, and user studies show strong controllability, accurate alignment between motion and camera view, high perceptual realism, and robust generalization from minimal inputs. Yuanpeng Tu |
Int. J. Comput. Vis. | 1 |
| 2026 | Memory Consistency Guided Divide-and-Conquer Learning for Generalized Category Discovery
Yuanpeng Tu, Zhun Zhong, Hengshuang Zhao |
Int. J. Comput. Vis. | 1 |
| 2025 | Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive ActivationsabstractPre-trained stable diffusion models (SD) have shown great advances in visual correspondence.
In this paper, we investigate the capabilities of Diffusion Transformers (DiTs) for accurate dense correspondence. Distinct from SD, DiTs exhibit a critical phenomenon in which very few feature activations exhibit significantly larger values than others, known as massive activations, leading to uninformative representations and significant performance degradation for DiTs.
The massive activations consistently concentrate at very few fixed dimensions across all image patch tokens, holding little local information.
We analyze these dimension-concentrated massive activations and uncover that their concentration is inherently linked to the Adaptive Layer Normalization (AdaLN) in DiTs.
Building on these findings, we propose the Diffusion Transformer Feature (DiTF), a training-free AdaLN-based framework that extracts semantically discriminative features from DiTs.
Specifically, DiTF leverages AdaLN to adaptively localize and normalize massive activations through channel-wise modulation.
Furthermore, a channel discard strategy is introduced to mitigate the adverse effects of massive activations.
Experimental results demonstrate that our DiTF outperforms both DINO and SD-based models and establishes a new state-of-the-art performance for DiTs in different visual correspondence tasks (e.g., with +9.4\% on Spair-71k and +4.4\% on AP-10K-C.S.). Chaofan Gan, Yuanpeng Tu, Tieyuan Chen, Yuxi Li 0009, Mehrtash Harandi, Weiyao Lin |
NeurIPS | 2 |
| 2025 | PlayerOne: Egocentric World SimulatorabstractWe introduce PlayerOne, the first egocentric realistic world simulator, facilitating immersive and unrestricted exploration within vividly dynamic environments. Given an egocentric scene image from the user, PlayerOne can accurately construct the corresponding world and generate egocentric videos that are strictly aligned with the real-scene human motion of the user captured by an exocentric camera. PlayerOne is trained in a coarse-to-fine pipeline that first performs pretraining on large-scale egocentric text-video pairs for coarse-level egocentric understanding, followed by finetuning on synchronous motion-video data extracted from egocentric-exocentric video datasets with our automatic construction pipeline. Besides, considering the varying importance of different components, we design a part-disentangled motion injection scheme, enabling precise control of part-level movements. In addition, we devise a joint reconstruction framework that progressively models both the 4D scene and video frames, ensuring scene consistency in the long-form video generation. Experimental results demonstrate its great generalization ability in precise control of varying human movements and world-consistent modeling of diverse scenarios. It marks the first endeavor into egocentric real-world simulation and can pave the way for the community to delve into fresh frontiers of world modeling and its diverse applications. Yuanpeng Tu, Hao Luo 0004, Xi Chen 0119, Xiang Bai, Fan Wang 0019, Hengshuang Zhao |
NeurIPS | 1 |
| 2024 | Self-Supervised Likelihood Estimation with Energy Guidance for Anomaly Segmentation in Urban ScenesabstractRobust autonomous driving requires agents to accurately identify unexpected areas (anomalies) in urban scenes. To this end, some critical issues remain open: how to design advisable metric to measure anomalies, and how to properly generate training samples of anomaly data? Classical effort in anomaly detection usually resorts to pixel-wise uncertainty or sample synthesis, which ignores the contextual information and sometimes requires auxiliary data with fine-grained annotations. On the contrary, in this paper, we exploit the strong context-dependent nature of segmentation task and design an energy-guided self-supervised frameworks for anomaly segmentation, which optimizes an anomaly head by maximizing likelihood of self-generated anomaly pixels. For this purpose, we design two estimators to model anomaly likelihood, one is a task-agnostic binary estimator and the other depicts the likelihood as residual of task-oriented joint energy. Based on proposed estimators, we devise an adaptive self-supervised training framework, which exploits the contextual reliance and estimated likelihood to refine mask annotations in anomaly areas. We conduct extensive experiments on challenging Fishyscapes and Road Anomaly benchmarks, demonstrating that without any auxiliary data or synthetic models, our method can still achieves comparable performance to supervised competitors. Code is available at https://github.com/yuanpengtu/SLEEG. Yuanpeng Tu, Yuxi Li 0009, Boshen Zhang, Liang Liu 0007, Jiangning Zhang, Yabiao Wang, Cairong Zhao |
AAAI | 1 |
| 2024 | Self-supervised Feature Adaptation for 3D Industrial Anomaly Detection
Yuanpeng Tu, Boshen Zhang, Liang Liu 0007, Yuxi Li 0009, Jiangning Zhang, Yabiao Wang, Chengjie Wang 0001, Cairong Zhao |
ECCV (2) | 1 |
| 2024 | DAC: 2D-3D Retrieval with Noisy Labels via Divide-and-Conquer Alignment and CorrectionabstractWith the recent burst of 2D and 3D data, cross-modal retrieval has attracted increasing attention recently. However, manual labeling by non-experts will inevitably introduce corrupted annotations given ambiguous 2D/3D content. Though previous works have addressed this issue by designing a naive division strategy with hand-crafted thresholds, their performance generally exhibits great sensitivity to the threshold value. Besides, they fail to fully utilize the valuable supervisory signals within each divided subset. To tackle this problem, we propose a Divide-and-conquer 2D-3D cross-modal Alignment and Correction framework (DAC), which comprises Multimodal Dynamic Division (MDD) and Adaptive Alignment and Correction (AAC). Specifically, the former performs accurate sample division by adaptive credibility modeling for each sample based on the compensation information within multimodal loss distribution. Then in AAC, samples in distinct subsets are exploited with different alignment strategies to fully enhance the semantic compactness and meanwhile alleviate over-fitting to noisy labels, where a self-correction strategy is introduced to improve the quality of representation. Moreover. To evaluate the effectiveness in real-world scenarios, we introduce a challenging noisy benchmark, namely Objaverse-N200, which comprises 200k-level samples annotated with 1156 realistic noisy labels. Extensive experiments on both traditional and the newly proposed benchmarks demonstrate the generality and superiority of our DAC, where DAC outperforms state-of-the-art models by a large margin. (i.e., with +5.9% gain on ModelNet40 and +5.8% on Objaverse-N200). Chaofan Gan, Yuanpeng Tu, Yuxi Li 0009, Weiyao Lin |
ACM Multimedia | 2 |
| 2023 | Learning from Noisy Labels with Decoupled Meta Label PurifierabstractTraining deep neural networks (DNN) with noisy labels is challenging since DNN can easily memorize inaccurate labels, leading to poor generalization ability. Recently, the meta-learning based label correction strategy is widely adopted to tackle this problem via identifying and correcting potential noisy labels with the help of a small set of clean validation data. Although training with purified labels can effectively improve performance, solving the meta-learning problem inevitably involves a nested loop of bi-level optimization between model weights and hyper-parameters (i.e., label distribution). As compromise, previous methods resort to a coupled learning process with alternating update. In this paper, we empirically find such simultaneous optimization over both model weights and label distribution can not achieve an optimal routine, consequently limiting the representation ability of backbone and accuracy of corrected labels. From this observation, a novel multi-stage label purifier named DMLP is proposed. DMLP decouples the label correction process into label-free representation learning and a simple meta label purifier, In this way, DMLP can focus on extracting discriminative feature and label correction in two distinctive stages. DMLP is a plug-and-play label purifier, the purified labels can be directly reused in naive end-to-end network retraining or other robust learning methods, where state-of-the-art results are obtained on several synthetic and real-world noisy datasets, especially under high noise levels. Code is available at https://github.com/yuanpengtu/DMLP. Yuanpeng Tu, Boshen Zhang, Yuxi Li 0009, Liang Liu 0007, Jian Li 0062, Yabiao Wang, Chengjie Wang 0001, Cairong Zhao |
CVPR | 1 |
| 2023 | Learning with Noisy labels via Self-supervised Adversarial Noisy MaskingabstractCollecting large-scale datasets is crucial for training deep models, annotating the data, however, inevitably yields noisy labels, which poses challenges to deep learning algorithms. Previous efforts tend to mitigate this problem via identifying and removing noisy samples or correcting their labels according to the statistical properties (e.g., loss values) among training samples. In this paper, we aim to tackle this problem from a new perspective, delving into the deep feature maps, we empirically find that models trained with clean and mislabeled samples manifest distinguishable activation feature distributions. From this observation, a novel robust training approach termed adversarial noisy masking is proposed. The idea is to regularize deep features with a label quality guided masking scheme, which adaptively modulates the input data and label simultaneously, preventing the model to overfit noisy samples. Further, an auxiliary task is designed to reconstruct input data, it naturally provides noise-free self-supervised signals to rein-force the generalization ability of models. The proposed method is simple yet effective, it is tested on synthetic and real-world noisy datasets, where significant improvements are obtained over previous methods. Code is available at https://github.com/yuanpengtu/SANM. Yuanpeng Tu, Boshen Zhang, Yuxi Li 0009, Liang Liu 0007, Jian Li 0062, Jiangning Zhang, Yabiao Wang, Chengjie Wang 0001, Cairong Zhao |
CVPR | 1 |
| 2023 | Content-Adaptive Auto-Occlusion Network for Occluded Person Re-IdentificationabstractThe occluded person re-identification (ReID) aims to match person images captured in severely occluded environments. Current occluded ReID works mostly rely on auxiliary models or employ a part-to-part matching strategy. However, these methods may be sub-optimal since the auxiliary models are constrained by occlusion scenes and the matching strategy will deteriorate when both query and gallery set contain occlusion. Some methods attempt to solve this problem by applying image occlusion augmentation (OA) and have shown great superiority in their effectiveness and lightness. But there are two defects that existed in the previous OA-based method: 1) The occlusion policy is fixed throughout the entire training and cannot be dynamically adjusted based on the current training status of the ReID network. 2) The position and area of the applied OA are completely random, without reference to the image content to choose the most suitable policy. To address these challenges, we propose a novel Content-Adaptive Auto-Occlusion Network (CAAO), that is able to dynamically select the proper occlusion region of an image based on its content and the current training status. Specifically, CAAO consists of two parts: the ReID network and the Auto-Occlusion Controller (AOC) module. AOC automatically generates the optimal OA policy based on the feature map extracted from the ReID network and applies occlusion on the images for ReID network training. An on-policy reinforcement learning based alternating training paradigm is proposed to iteratively update the ReID network and AOC module. Comprehensive experiments on occluded and holistic person ReID benchmarks demonstrate the superiority of CAAO. Cairong Zhao, Zefan Qu, Xinyang Jiang, Yuanpeng Tu, Xiang Bai |
IEEE Trans. Image Process. | 4 |
| 2021 | Salience-Guided Iterative Asymmetric Mutual Hashing for Fast Person Re-IdentificationabstractPerson Re-identification (ReID) aims to retrieve the pedestrian with the same identity across different views. Existing studies mainly focus on improving accuracy, while ignoring their efficiency. Recently, several hash based methods have been proposed. Despite their improvement in efficiency, there still exists an unacceptable gap in accuracy between these methods and real-valued ones. Besides, few attempts have been made to simultaneously explicitly reduce redundancy and improve discrimination of hash codes, especially for short ones. Integrating Mutual learning may be a possible solution to reach this goal. However, it fails to utilize the complementary effect of teacher and student models. Additionally, it will degrade the performance of teacher models by treating two models equally. To address these issues, we propose a salience-guided iterative asymmetric mutual hashing (SIAMH) to achieve high-quality hash code generation and fast feature extraction. Specifically, a salience-guided self-distillation branch (SSB) is proposed to enable SIAMH to generate hash codes based on salience regions, thus explicitly reducing the redundancy between codes. Moreover, a novel iterative asymmetric mutual training strategy (IAMT) is proposed to alleviate drawbacks of common mutual learning, which can continuously refine the discriminative regions for SSB and extract regularized dark knowledge for two models as well. Extensive experiment results on five widely used datasets demonstrate the superiority of the proposed method in efficiency and accuracy when compared with existing state-of-the-art hashing and real-valued approaches. The code is released at https://github.com/Vill-Lab/SIAMH. Cairong Zhao, Yuanpeng Tu, Zhihui Lai 0001, Fumin Shen, Heng Tao Shen, Duoqian Miao 0001 |
IEEE Trans. Image Process. | 2 |