Peisen Zhao

dblp:119/9223 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0002-3787-193XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 GammaDiff: Deep Diffusion Models for Gamma Index Synthesis in Radiation Therapy
abstract
Radiation therapy is a cornerstone of tumor treatment, and accurate prediction of the gamma passing rate (GPR) for intensity-modulated radiation therapy (IMRT) plans is clinically critical. Existing AI-based predictors often lack locational information of dose accuracy. We propose GammaDiff, a diffusion-model framework for gamma distribution prediction, with three main contributions: (1) an advanced noise-prediction network that fuses CNNs for local features with transformers for global context, achieving multi-scale modeling with efficient computation; (2) a multi-scale fusion U-Net (MFUnet) that em-beds fluence maps structure via hierarchical feature integration into the noise-prediction process; and (3) a two-stage diffusion procedure in which the forward process progressively adds noise to form training samples, and the reverse process uses the opti-mized predictor to reconstruct high-fidelity gamma distributions. Extensive experiments show that GammaDiff outperforms prior methods on PSNR and SSIM, with notably higher sensitivity to failure cases, providing a more robust, reliable AI solution for plan-quality verification in radiation therapy.
Yuquan Wang, Peisen Zhao, Ruijie Yang, Hongxia Deng
BIBM3
2025 Prediction of Patient Plan Quality Assurance Based on a Multimodal Cascade Feature Fusion Model
Peisen Zhao, YuQuan Wang, Hongxia Deng, RuiJie Yang
ICIC (19)2
2025 Uncover the balanced geometry in long-tailed contrastive language-image pretraining
Zhihan Zhou 0002, Yushi Ye, Feng Hong 0004, Peisen Zhao, Jiangchao Yao, Ya Zhang 0002, Qi Tian 0001, Yanfeng Wang 0001
Mach. Learn.4
2025 FDBKeeper: Enabling Scalable Coordination Services for Metadata Management using Distributed Key-Value Databases
abstract
High-reliability distributed coordination services have become an indispensable part of modern large-scale distributed systems. Popular coordination services (e.g., ZooKeeper) adopt a single-writer design to provide a centralized service for managing system metadata, including various configuration information and data catalogs, and to provide distributed synchronization functions. With the continuous increase in metadata size and the scale of distributed systems, these coordination services gradually become performance bottlenecks due to their limitations in capacity, read and write performance, and scalability. To bridge the gaps, we propose FDBKeeper, a novel solution that enables scalable coordination services on distributed ACID key-value database systems. Our motivation is that transactional key-value stores (i.e., FoundationDB) meet the demands of performance and scalability required by large-scale distributed systems over coordination service. To leverage these advantages, coordination services can be implemented as an upper layer on top of distributed ACID key-value databases. Our experimental results demonstrate that FDBKeeper significantly outperforms ZooKeeper across key metrics. Additionally, FDBKeeper reduces hardware resource costs on average by 33% in the production environment, resulting in substantial monetary cost savings. We have successfully replaced ZooKeeper with FDBKeeper in the production-grade ClickHouse cluster deployment.
Jun-Peng Zhu, Peng Cai 0001, Xuan Zhou 0001, Peisen Zhao, Linpeng Tang
Proc. VLDB Endow.5
2024 UMG-CLIP: A Unified Multi-granularity Vision Generalist for Open-World Understanding
Bowen Shi 0003, Peisen Zhao, Yuhang Zhang 0012, Jin Li 0057, Wenrui Dai, Junni Zou, Hongkai Xiong, Qi Tian 0001, Xiaopeng Zhang 0008
ECCV (38)2
2023 Distilling Vision-Language Pre-Training to Collaborate with Weakly-Supervised Temporal Action Localization
abstract
Weakly-supervised temporal action localization (WTAL) learns to detect and classify action instances with only category labels. Most methods widely adopt the off-the-shelf Classification-Based Pre-training (CBP) to generate video features for action localization. However, the different optimization objectives between classification and localization, make temporally localized results suffer from the serious incomplete issue. To tackle this issue without additional annotations, this paper considers to distill free action knowledge from Vision-Language Pre-training (VLP), as we surprisingly observe that the localization results of vanilla VLP have an over-complete issue, which is just complementary to the CBP results. To fuse such complementarity, we propose a novel distillation-collaboration framework with two branches acting as CBP and VLP respectively. The framework is optimized through a dual-branch alternate training strategy. Specifically, during the B step, we distill the confident background pseudo-labels from the CBP branch; while during the F step, the confident foreground pseudo-labels are distilled from the VLP branch. As a result, the dualbranch complementarity is effectively fused to promote one strong alliance. Extensive experiments and ablation studies on THUMOS14 and ActivityNet1.2 reveal that our method significantly outperforms state-of-the-art methods.
Chen Ju, Kunhao Zheng, Jinxiang Liu, Peisen Zhao, Ya Zhang 0002, Jianlong Chang, Qi Tian 0001, Yanfeng Wang 0001
CVPR4
2023 Prune Spatio-temporal Tokens by Semantic-aware Temporal Accumulation
abstract
Transformers have become the primary backbone of the computer vision community due to their impressive performance. However, the unfriendly computation cost impedes their potential in the video recognition domain. To optimize the speed-accuracy trade-off, we propose Semanticaware Temporal Accumulation score (STA) to prune spatiotemporal tokens integrally. STA score considers two critical factors: temporal redundancy and semantic importance. The former depicts a specific region based on whether it is a new occurrence or a seen entity by aggregating token-to-token similarity in consecutive frames while the latter evaluates each token based on its contribution to the overall prediction. As a result, tokens with higher scores of STA carry more temporal redundancy as well as lower semantics thus being pruned. Based on the STA score, we are able to progressively prune the tokens without introducing any additional parameters or requiring further re-training. We directly apply the STA module to off-the-shelf ViT and VideoSwin backbones, and the empirical results on Kinetics-400 and Something-Something V2 achieve over 30% computation reduction with a negligible ~ 0.2% accuracy drop. The code is released at https://github.com/Mark12Ding/STA.
Shuangrui Ding, Peisen Zhao, Xiaopeng Zhang 0008, Rui Qian 0001, Hongkai Xiong, Qi Tian 0001
ICCV2
2023 Adaptive Mutual Supervision for Weakly-Supervised Temporal Action Localization
abstract
Weakly-supervised temporal action localization aims to localize actions from untrimmed long videos with only video-level category labels. Most previous methods ignore the incompleteness issue of Class Activation Sequences (CAS), suffering from trivial detection results. To tackle this issue, we propose a novel Adaptive Mutual Supervision (AMS) framework with two branches, where the base branch detects the most discriminative action regions, while the supplementary branch localizes the less discriminative action regions through an adaptive sampler. The sampler dynamically updates the inputs for the supplementary branch using a sampling weight sequence negatively correlated with the CAS from the base branch, thus encouraging the supplementary branch to localize the action regions underestimated by the base branch. To promote mutual enhancement between two branches, we further construct mutual location supervision. Each branch adopts the location pseudo-labels generated from the other branch as the localization supervision. By alternately optimizing two branches for multiple iterations, we progressively complete action regions. Extensive experiments on THUMOS14 and ActivityNet1.2 demonstrate that the proposed AMS method significantly outperforms state-of-the-art methods.
Chen Ju, Peisen Zhao, Siheng Chen, Ya Zhang 0002, Xiaoyun Zhang 0001, Yanfeng Wang 0001, Qi Tian 0001
IEEE Trans. Multim.2
2022 Progressive privileged knowledge distillation for online action detection
Peisen Zhao, Lingxi Xie, Ya Zhang 0002, Qi Tian 0001
Pattern Recognit.1
2022 Actionness-Guided Transformer for Anchor-Free Temporal Action Localization
abstract
Temporal action localization, detecting actions in untrimmed videos, is widely studied by anchor-based approaches that first generate excessive action proposals,i.e., temporal windows, then evaluate and classify these proposals. To reduce the number of action proposals, recent studies use an anchor-free approach that leverages each time point rather than a temporal window to represent an action instance. However, this point representation, usually modeled by temporal convolutions, may have the fixed and limited receptive field to detect an entire action. So we propose an Actionness-guided Transformer (Ag-Trans) model to learn representations for each point proposal. Ag-Trans first predicts the actionness,i.e., time sequences of the action starting, continuing, and ending phases, then the corresponding action phase can be embedded to model the point representation. Experimental results show that the Ag-Trans model outperforms the CNN-based model under the same experiment settings, especially for long-duration actions.
Peisen Zhao, Lingxi Xie, Ya Zhang 0002, Qi Tian 0001
IEEE Signal Process. Lett.1
2021 ESAD: End-to-end Semi-supervised Anomaly Detection
Chaoqin Huang, Peisen Zhao, Ya Zhang 0002, Yanfeng Wang 0001, Qi Tian 0001
BMVC3
2021 Divide and Conquer for Single-frame Temporal Action Localization
abstract
Single-frame temporal action localization (STAL) aims to localize actions in untrimmed videos with only one timestamp annotation for each action instance. Existing methods adopt the one-stage framework but couple the counting goal and the localization goal. This paper proposes a novel two-stage framework for the STAL task with the spirit of divide and conquer. The instance counting stage leverages the location supervision to determine the number of action instances and divide a whole video into multiple video clips, so that each video clip contains only one complete action instance; and the location estimation stage leverages the category supervision to localize the action instance in each video clip. To efficiently represent the action instance in each video clip, we introduce the proposal-based representation, and design a novel differentiable mask generator to enable the end-to-end training supervised by category labels. On THUMOS14, GTEA, and BEOID datasets, our method outperforms state-of-the-art methods by 3.5%, 2.7%, 4.8% mAP on average. And extensive experiments verify the effectiveness of our method.
Chen Ju, Peisen Zhao, Siheng Chen, Ya Zhang 0002, Yanfeng Wang 0001, Qi Tian 0001
ICCV2
2021 A Solution to Multi-modal Ads Video Tagging Challenge
abstract
In this paper, we present our solution to the Multi-modal Ads Video Tagging Challenge of Tencent Advertising Algorithm Competition in ACM Multimedia 2021 Grand Challenges. We extend the baseline model by redesigning the visual feature extraction procedure and we modify the loss function to cope with sparse positive targets. Moreover, we propose Semi-supervised Learning with Negative Masking to leverage both labeled data and unlabeled data from the preliminary contest which effectively enhances the training process. We further utilize Cross-Class Relevance Learning to boost the performance. We achieve 0.8237 GAP score via model ensemble and rank the second place among all submissions in the challenge.
Yuanzhe Gu, Peisen Zhao, Zhonglin Zu
ACM Multimedia4
2021 Universal-to-Specific Framework for Complex Action Recognition
abstract
Video-based action recognition has recently attracted much attention in the field of computer vision. To solve more complex recognition tasks, it has become necessary to distinguish different levels of interclass variations. Inspired by a common flowchart based on the human decision-making process that first narrows down the probable classes and then applies a "rethinking" process for finer-level recognition, we propose an effective universal-to-specific (U2S) framework for complex action recognition. The U2S framework is composed of three subnetworks: a universal network, a category-specific network, and a mask network. The universal network first learns universal feature representations. The mask network then generates attention masks for confusing classes through category regularization based on the output of the universal network. The mask is further used to guide the category-specific network for class-specific feature representations. The entire framework is optimized in an end-to-end manner. Experiments on a variety of benchmark datasets, e.g., the Something-Something, UCF101, and HMDB51 datasets, demonstrate the effectiveness of the U2S framework; i.e., U2S can focus on discriminative spatiotemporal regions for confusing categories. We further visualize the relationship between different classes, showing that U2S indeed improves the discriminability of learned features. Moreover, the proposed U2S model is a general framework and may adopt any base recognition network.
Peisen Zhao, Lingxi Xie, Ya Zhang 0002, Qi Tian 0001
IEEE Trans. Multim.1
2020 Bottom-Up Temporal Action Localization with Mutual Regularization
Peisen Zhao, Lingxi Xie, Chen Ju, Ya Zhang 0002, Yanfeng Wang 0001, Qi Tian 0001
ECCV (8)1
2012 Probabilistic Neural Network for RSS-Based Collaborative Localization
abstract
One critical challenge for accurate localization with Received Signal Strength Indicator (RSSI) is the anisotropic environment, which causes the RSS-Distance Relationship (RDR) to vary spatially. To alleviate localization error caused by RDR anisotropy, most of existing works adopt multiple RDR algorithms. However, we have found that the arbitrary RDR selection in these algorithms can lead to large localization error. Moreover, localization accuracy can be further enhanced by utilizing information provided by more Access Points (APs). To address these problems, we propose a Probabilistic Neural Network based localization algorithm in this paper. The algorithm features two steps: Global Optimization and Regional Compensation, during which all APs exchange information about the Blind Node (BN) to locate it collaboratively. Simulation result shows that the proposed algorithm can achieve a localization accuracy 35% higher than that of multiple RDR algorithms.
Peisen Zhao, Chunxiao Jiang, Hongyang Chen 0001, Yong Ren 0001
VTC Spring1