Peiyuan Zhang

dblp:240/6112 · DBLP profile ↗
← Back
20ranked-venue papers
11as first author
19since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 11 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Generalization Lower Bounds for GD and SGD in Smooth Stochastic Convex Optimization
abstract
This work studies the generalization error of gradient methods. More specifically, we focus on how training steps $T$ and step-size $\eta$ might affect generalization in smooth stochastic convex optimization (SCO) problems. Recent works show that in some cases longer training can hurt generalization. Our work reexamines this for smooth SCO and find that the conclusion can be case-dependent. In particular, we first study SCO problems when the loss is \emph{realizable}, i.e. an optimal solution minimizes all the data points. Our work provides excess risk lower bounds for Gradient Descent (GD) and Stochastic Gradient Descent (SGD) and finds that longer training may not hurt generalization. In the short training scenario $\eta T = O(n)$ ($n$ is sample size), our lower bounds tightly match and certify the respective upper bounds. However, for the long training scenario where $\eta T =O(n)$, our analysis reveals a gap between the lower and upper bounds, indicating that longer training does hurt generalization for realizable objectives. A conjecture is proposed that the gap can be closed by improving upper bounds, supported by analyses in two special instances. Moreover, besides the realizable setup, we also provide first tight excess risk lower bounds for GD and SGD under the general non-realizable smooth SCO setting, suggesting that existing stability analyses are tight in step-size and iteration dependence, and that overfitting provably happens when there is no interpolating minimum.
Peiyuan Zhang, Jiaye Teng, Jingzhao Zhang
AISTATS1
2025 Point2RBox-v2: Rethinking Point-supervised Oriented Object Detection with Spatial Layout Among Instances
abstract
With the rapidly increasing demand for oriented object detection (OOD), recent research involving weakly-supervised detectors for learning OOD from point annotations has gained great attention. In this paper, we rethink this challenging task setting with the layout among instances and present Point2RBox-v2. At the core are three principles: 1) Gaussian overlap loss. It learns an upper bound for each instance by treating objects as 2D Gaussian distributions and minimizing their overlap. 2) Voronoi watershed loss. It learns a lower bound for each instance through watershed on Voronoi tessellation. 3) Consistency loss. It learns the size/rotation variation between two output sets with respect to an input image and its augmented view. Supplemented by a few devised techniques, e.g. edge loss and copy-paste, the detector is further enhanced. To our best knowledge, Point2RBox-v2 is the first approach to explore the spatial layout among instances for learning point-supervised OOD. Our solution is elegant and lightweight, yet it is expected to give a competitive performance especially in densely packed scenes: 62.61%/86.15%/34.71% on DOTA/HRSC/FAIR1M.
Yi Yu 0010, Botao Ren, Peiyuan Zhang, Shaofeng Zhang, Feipeng Da, Junchi Yan, Xue Yang 0005
CVPR3
2025 EgoLife: Towards Egocentric Life Assistant
abstract
We introduce EgoLife, a project to develop an egocentric life assistant that accompanies and enhances personal efficiency through AI-powered wearable glasses. To lay the foundation for this assistant, we conducted a comprehensive data collection study where six participants lived together for one week, continuously recording their daily activities—including discussions, shopping, cooking, social-izing, and entertainment—using AI glasses for multimodal person-view video references. This effort resulted in EgoLife Dataset, a comprehensive 300-hour egocentric, terpersonal, multiview, and multimodal daily life with intensive annotation. Leveraging this dataset, we troduce EgoLifeQA, a suite of long-context, life-oriented question-answering tasks designed to provide meaningful sistance in daily life by addressing practical questions as recalling past relevant events, monitoring health and offering personalized recommendations.To address the key technical challenges of 1) developing robust visual-audio models for egocentric data, 2) enabling identity recognition, and 3) facilitating long-context question answering over extensive temporal information, we introduce EgoBulter, an integrated system comprising EgoGPT and EgoRAG. EgoGPT is an omni-modal model trained on egocentric datasets, achieving state-of-the-art performance on egocentric video understanding. EgoRAG is a retrieval-based component that supports answering ultra-long-context questions. Our experimental studies verify their working mechanisms and reveal critical factors and bottlenecks, guiding future improvements. By releasing our datasets, models, and benchmarks, we aim to stimulate further research in egocentric AI assistants.
Shuai Liu 0002, Hongming Guo, Yuhao Dong, Xiamengwei Zhang, Pengyun Wang, Zitang Zhou, Binzhu Xie, Bei Ouyang, Zhengyu Lin, Marco Cominelli, Zhongang Cai, Bo Li 0080, Yuanhan Zhang, Peiyuan Zhang, Fangzhou Hong, Jörg Widmer, Francesco Gringoli, Lei Yang 0059, Ziwei Liu 0002
CVPR17
2025 2nd Latent in the Wild Fingerprint Recognition Competition
abstract
This paper presents a summary of the 2nd Latent in the Wild Fingerprint Recognition Competition held at the 2025 International Joint Conference on Biometrics. The competition has two tracks: latent fingerprint 1) recognition, and 2) quality assessment. It attracted a total of 12 participating teams from academia and industry for both tracks, representing 10 countries. In total, 8 valid submissions were evaluated by the organizers. The competition aimed to advance the state-of-the-art in latent fingerprint recognition and quality assessment by providing a challenging dataset of latent fingerprints collected in natural, non-ideal conditions. This paper summarizes the dataset, evaluation protocols, submitted methods, and the competition results.
Xinwei Liu 0001, Renfang Wang, Peiyuan Zhang, Tim Oblak, Lara Anzur, Peter Peer, Evaldas Borcovas, Arturas Nakvosas, Ignas Mataitis, Valdemaras Pasvenskas, Andrius Stankevicius, Marko Lange, David Stumpf, Sven Utcke, Patryk Szwargulski, Fantin Girard, Zacharie Legault, Ekansh Thakur, Jaishana Bindhu Priya, Pavan Kumar C, Ramachandra Raghavendra, Kiran B. Raja
IJCB4
2025 Temporal Reasoning Transfer from Text to Video
abstract
Video Large Language Models (Video LLMs) have shown promising capabilities in video comprehension, yet they struggle with tracking temporal changes and reasoning about temporal relationships. While previous research attributed this limitation to the ineffective temporal encoding of visual inputs, our diagnostic study reveals that video representations contain sufficient information for even small probing classifiers to achieve perfect accuracy. Surprisingly, we find that the key bottleneck in Video LLMs' temporal reasoning capability stems from the underlying LLM's inherent difficulty with temporal concepts, as evidenced by poor performance on textual temporal question-answering tasks. Building on this discovery, we introduce the Textual Temporal reasoning Transfer (T3). T3 synthesizes diverse temporal reasoning tasks in pure text format from existing image-text datasets, addressing the scarcity of video samples with complex temporal scenarios. Remarkably, without using any video data, T3 enhances LongVA-7B's temporal understanding, yielding a 5.3 absolute accuracy improvement on the challenging TempCompass benchmark, which enables our model to outperform ShareGPT4Video-8B trained on 28,000 video samples. Additionally, the enhanced LongVA-7B model achieves competitive performance on comprehensive video benchmarks. For example, it achieves a 49.7 accuracy on the Temporal Reasoning task of Video-MME, surpassing powerful large-scale models such as InternVL-Chat-V1.5-20B and VILA1.5-40B. Further analysis reveals a strong correlation between textual and video temporal task performance, validating the efficacy of transferring temporal reasoning abilities from text to video domains.
Lei Li 0039, Yuanxin Liu, Linli Yao, Peiyuan Zhang, Chenxin An, Lean Wang, Xu Sun 0001, Lingpeng Kong, Qi Liu 0049
ICLR4
2025 Fast Video Generation with Sliding Tile Attention
abstract
Diffusion Transformers (DiTs) with 3D full attention power state-of-the-art video generation, but suffer from prohibitive compute cost -- when generating just a 5-second 720P video, attention alone takes 800 out of 950 seconds of total inference time. This paper introduces sliding tile attention (STA) to address this challenge. STA leverages the observation that attention scores in pretrained video diffusion models predominantly concentrate within localized 3D windows. By sliding and attending over local spatial-temporal region, STA eliminates redundancy from full attention. Unlike traditional token-wise sliding window attention (SWA), STA operates tile-by-tile with a novel hardware-aware sliding window design, preserving expressiveness while being \emph{hardware-efficient}. With careful kernel-level optimizations, STA offers the first efficient 2D/3D sliding-window-like attention implementation, achieving 58.79\% MFU -- 7.17× faster than prior art methods. On the leading video DiT model, Hunyuan, it accelerates attention by 1.6–10x over FlashAttention-3, yielding a 1.36–3.53× end-to-end speedup with no or minimum quality loss.
Peiyuan Zhang, Runlong Su, Hangliang Ding, Ion Stoica, Zhengzhong Liu 0001, Hao Zhang 0025
ICML1
2025 An Optimized Franz-Parisi Criterion and its Equivalence with SQ Lower Bounds
abstract
Bandeira et al. (2022) introduced the Franz-Parisi (FP) criterion for characterizing the computational hard phases in statistical detection problems. The FP criterion, based on an annealed version of the celebrated Franz-Parisi potential from statistical physics, was shown to be equivalent to low-degree polynomial (LDP) lower bounds for Gaussian additive models, thereby connecting two distinct approaches to understanding the computational hardness in statistical inference. In this paper, we propose a refined FP criterion that aims to better capture the geometric ``overlap" structure of statistical models. Our main result establishes that this optimized FP criterion is equivalent to Statistical Query (SQ) lower bounds---another foundational framework in computational complexity of statistical inference. Crucially, this equivalence holds under a mild, verifiable assumption satisfied by a broad class of statistical models, including Gaussian additive models, planted sparse models, as well as non-Gaussian component analysis (NGCA), single-index (SI) models, and convex truncation detection settings. For instance, in the case of convex truncation tasks, the assumption is equivalent with the Gaussian correlation inequality (Royen, 2014) from convex geometry. In addition to the above, our equivalence not only unifies and simplifies the derivation of several known SQ lower bounds—such as for the NGCA model (Diakonikolas et al., 2017) and the SI model (Damian et al., 2024)—but also yields new SQ lower bounds of independent interest, including for the computational gaps in mixed sparse linear regression (Arpino et al., 2023) and convex truncation (De et al., 2023).
Theodor Misiakiewicz, Ilias Zadik, Peiyuan Zhang
NeurIPS4
2025 Faster Video Diffusion with Trainable Sparse Attention
abstract
Scaling video diffusion transformers (DiTs) is limited by their quadratic 3D attention, even though most of the attention mass concentrates on a small subset of positions. We turn this observation into VSA, a trainable, hardware-efficient sparse attention that replaces full attention at both training and inference. In VSA, a lightweight coarse stage pools tokens into tiles and identifies high-weight critical tokens; a fine stage computes token-level attention only inside those tiles subjecting to block computing layout to ensure hard efficiency. This leads to a single differentiable kernel that trains end-to-end, requires no post-hoc profiling, and sustains 85\% of FlashAttention3 MFU. We perform a large sweep of ablation studies and scaling-law experiments by pretraining DiTs from 60M to 1.4B parameters. VSA reaches a Pareto point that cuts training FLOPS by 2.53$\times$ with no drop in diffusion loss. Retrofitting the open-source Wan2.1-1.3B model speeds up attention time by 6$\times$ and lowers end-to-end generation time from 31s to 18s with comparable quality, while for the 14B model, end-to-end generation time is reduced from 1274s to 576s. Furthermore, we introduce a preliminary study of Sparse-Distill, the first method to enable sparse attention and distillation concurrently, achieving 50.9x speed up for Wan-1.3B while maintaining quality. These results establish trainable sparse attention as a practical alternative to full attention and a key enabler for further scaling of video diffusion models. Code is available at https://github.com/hao-ai-lab/FastVideo.
Peiyuan Zhang, Haofeng Huang, Will Lin, Zhengzhong Liu 0001, Ion Stoica, Eric P. Xing, Hao Zhang 0025
NeurIPS1
2025 Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing
abstract
Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but they still face challenges in General Visual Editing, particularly in following complex instructions, preserving appearance consistency, and supporting flexible input formats. To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE). RISEBench focuses on four key reasoning categories: Temporal, Causal, Spatial, and Logical Reasoning. We curate high-quality test cases for each category and propose an robust evaluation framework that assesses Instruction Reasoning, Appearance Consistency, and Visual Plausibility with both human judges and the LMM-as-a-judge approach. We conducted experiments evaluating nine prominent visual editing models, comprising both open-source and proprietary models. The evaluation results demonstrate that current models face significant challenges in reasoning-based editing tasks. Even the most powerful model evaluated, GPT-image-1, achieves an accuracy of merely 28.8%. RISEBench effectively highlights the limitations of contemporary editing models, provides valuable insights, and indicates potential future directions for the field of reasoning-aware visual editing. Our code and data have been released at https://github.com/PhoenixZ810/RISEBench.
Peiyuan Zhang, Kexian Tang, Xiaorong Zhu, Hao Li 0069, Wenhao Chai, Renqiu Xia, Guangtao Zhai, Junchi Yan, Hua Yang 0001, Xue Yang 0005, Haodong Duan
NeurIPS2
2025 PointOBB-v3: Expanding Performance Boundaries of Single Point-Supervised Oriented Object Detection
Peiyuan Zhang, Xue Yang 0005, Yi Yu 0010, Qingyun Li, Yue Zhou 0005, Xiaosong Jia, Jingdong Chen, Xiang Li 0041, Junchi Yan, Yansheng Li 0001
Int. J. Comput. Vis.1
2024 The Dynamic Nature of Procrastination
Peiyuan Zhang, Yijun Lin 0006, Falk Lieder, Wei Ji Ma
CogSci1
2024 Counting Stars is Constant-Degree Optimal For Detecting Any Planted Subgraph: Extended Abstract
abstract
We prove that whenever $p=\Omega(1)$ and for any graph $H$, counting $O(1)$-stars is optimal among all constant degree polynomial tests in terms of strongly separating an instance of $G(n,p),$ from the union of a random copy of $H$ with an instance of $G(n,p).$ Our work generalizes and extends multiple previous results on the inference abilities of $O(1)$-degree polynomials in the literature.
Xifan Yu, Ilias Zadik, Peiyuan Zhang
COLT3
2023 One Network, Many Masks: Towards More Parameter-Efficient Transfer Learning
abstract
Fine-tuning pre-trained language models for multiple tasks tends to be expensive in terms of storage.To mitigate this, parameter-efficient transfer learning (PETL) methods have been proposed to address this issue, but they still require a significant number of parameters and storage when being applied to broader ranges of tasks.To achieve even greater storage reduction, we propose PROPETL, a novel method that enables efficient sharing of a single PETL module which we call prototype network (e.g., adapter, LoRA, and prefix-tuning) across layers and tasks.We then learn binary masks to select different sub-networks from the shared prototype network and apply them as PETL modules into different layers.We find that the binary masks can determine crucial information from the network, which is often ignored in previous studies.Our work can also be seen as a type of pruning method, where we find that overparameterization also exists in the seemingly small PETL modules.We evaluate PROPETL on various downstream tasks and show that it can outperform other PETL methods with approximately 10% of the parameter storage required by the latter. 1
Guangtao Zeng, Peiyuan Zhang, Wei Lu 0011
ACL (1)2
2022 Procrastination and the Intention-Behavior Gap
Luísa Leonelli de Moraes, Peiyuan Zhang, Yijun Lin 0006, Wei Ji Ma
CogSci2
2022 Effects of reward schedule and pressure on procrastination
Peiyuan Zhang, Yijun Lin 0006, Wei Ji Ma
CogSci1
2022 Better Few-Shot Relation Extraction with Label Prompt Dropout
abstract
Few-shot relation extraction aims to learn to identify the relation between two entities based on very limited training examples.Recent efforts found that textual labels (i.e., relation names and relation descriptions) could be extremely useful for learning class representations, which will benefit the few-shot learning task.However, what is the best way to leverage such label information in the learning process is an important research question.Existing works largely assume such textual labels are always present during both learning and prediction.In this work, we argue that such approaches may not always lead to optimal results.Instead, we present a novel approach called label prompt dropout, which randomly removes label descriptions in the learning process.Our experiments show that our approach is able to lead to improved class representations, yielding significantly better results on the few-shot relation extraction task. 1
Peiyuan Zhang, Wei Lu 0011
EMNLP1
2021 Revisiting the Role of Euler Numerical Integration on Acceleration and Stability in Convex Optimization
abstract
Viewing optimization methods as numerical integrators for ordinary differential equations (ODEs) provides a thought-provoking modern framework for studying accelerated first-order optimizers. In this literature, acceleration is often supposed to be linked to the quality of the integrator (accuracy, energy preservation, symplecticity). In this work, we propose a novel ordinary differential equation that questions this connection: both the explicit and the semi-implicit (a.k.a symplectic) Euler discretizations on this ODE lead to an accelerated algorithm for convex programming. Although semi-implicit methods are well-known in numerical analysis to enjoy many desirable features for the integration of physical systems, our findings show that these properties do not necessarily relate to acceleration.
Peiyuan Zhang, Antonio Orvieto, Hadi Daneshmand, Thomas Hofmann 0001, Roy S. Smith
AISTATS1
2021 Designing a behavioral experiment to study the factors underlying procrastination
Peiyuan Zhang, Wei Ji Ma
CogSci1
2021 Rethinking the Variational Interpretation of Accelerated Optimization Methods
abstract
The continuous-time model of Nesterov's momentum provides a thought-provoking perspective for understanding the nature of the acceleration phenomenon in convex optimization. One of the main ideas in this line of research comes from the field of classical mechanics and proposes to link Nesterov's trajectory to the solution of a set of Euler-Lagrange equations relative to the so-called Bregman Lagrangian. In the last years, this approach led to the discovery of many new (stochastic) accelerated algorithms and provided a solid theoretical foundation for the design of structure-preserving accelerated methods. In this work, we revisit this idea and provide an in-depth analysis of the action relative to the Bregman Lagrangian from the point of view of calculus of variations. Our main finding is that, while Nesterov's method is a stationary point for the action, it is often not a minimizer but instead a saddle point for this functional in the space of differentiable curves. This finding challenges the main intuition behind the variational interpretation of Nesterov's method and provides additional insights into the intriguing geometry of accelerated paths.
Peiyuan Zhang, Antonio Orvieto, Hadi Daneshmand
NeurIPS1
2020 a process model of procrastination
Peiyuan Zhang, Wei Ji Ma
CogSci1