Zhenying Fang

dblp:224/9893 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
5since 2021 · last 2026
0009-0004-2002-7267ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 1Computer networks · 1
YearPublicationVenuePosition
2026 PseR: Pseudo-Label Refinement for Point-Supervised Temporal Action Detection
Zhenying Fang, Richang Hong
IEEE Trans. Multim.1
2026 PAB: Points as a Bridge for Omni-Supervised Temporal Action Detection
abstract
Omni-supervised temporal action detection aims to leverage various forms of supervised data to improve the generalization ability of action detectors. Existing methods rely on low-quality random proposals for classification and regression, resulting in suboptimal detection performance. To address this, we proposePointsAs aBridge (PAB), a unified pseudo-label generator that integrates unlabeled, fully labeled, and weakly labeled annotations (e.g., video-level and point-level labels), generating high-quality pseudo-labels and improving detection performance. PAB consists of two main modules:Multi-ModalPointMining (MMPM) andPoint-EnhancedTransformer (PET). The MMPM module uses a learnable visual prototype and a frozen textual prototype for each category. These prototypes jointly identify temporal points in the input video where action instances are likely to occur, with each point representing its temporal position and associated action category. Based on these mined points, the PET models a series of explicit action queries to predict the temporal boundaries of the action instance for each point. In this framework, points serve as a bridge, transmitting temporal positions and action categories between the two modules. Based on MMPM's predicted action categories and PET's estimated action boundaries for each point, PAB generates high-quality pseudo-labels for unlabeled and weakly labeled data, enabling a fully supervised detector to achieve state-of-the-art performance in semi- and omni-supervised settings.
Zhenying Fang, Richang Hong
IEEE Trans. Multim.1
2025 Boundary Discretization and Reliable Classification Network for Temporal Action Detection
abstract
Temporal action detection aims to recognize the action category and determine each action instance's starting and ending time in untrimmed videos. The mixed method has demonstrated notable performance by integrating both anchor-based and anchor-free approaches. However, while it leverages the strengths of each method, it also retains their respective limitations. For instance, the anchor-based approach depends on manually crafted anchors tailored to specific datasets, while the anchor-free approach predicts potential action instances at each temporal position, resulting in a significant number of false positives in category prediction. The inclusion of these limitations undermines the potential benefits of the mixed method. In this paper, we propose a novel Boundary Discretization and Reliable Classification Network (BDRC-Net) that addresses the issues above by introducing boundary discretization and reliable classification modules. Specifically, the boundary discretization module (BDM) elegantly merges anchor-based and anchor-free approaches in the form of boundary discretization, eliminating the need for the traditional handcrafted anchor design. Furthermore, the reliable classification module (RCM) predicts reliable global action categories to reduce false positives. Extensive experiments conducted on different benchmarks demonstrate that our proposed method achieves competitive detection performance.
Zhenying Fang, Jun Yu 0002, Richang Hong
IEEE Trans. Multim.1
2023 LPR: learning point-level temporal action localization through re-training
abstract
Abstract Point-level temporal action localization (PTAL) aims to locate action instances in untrimmed videos with only one timestamp annotation for each action instance. Existing methods adopt the localization-by-classification paradigm to locate action boundaries in the temporal class activation map (TCAM) by thresholding, also known as TCAM-based method. However, TCAM-based methods are limited by the gap between classification and localization tasks, since TCAM is generated by a classification network. To address this issue, we propose a re-training framework for the PTAL task, also known as LPR. This framework consists of two stages: pseudo-label generation and re-training. In the pseudo-label generation stage, we propose a feature embedding module based on a transformer encoder to capture global context features and optimize pseudo-labels’ quality by leveraging point-level annotations. In the re-training stage, LPR uses the above pseudo-labels as supervision to locate action instances with a temporal action localization network rather than generating TCAMs. Furthermore, to alleviate the effects of label noise in the pseudo-labels, we propose a joint learning classification module (JLCM) in the re-training stage. This module contains two classification sub-modules that simultaneously predict action categories and are guided by a jointly determined clean set for network training. The proposed framework achieves state-of-the-art localization performance on both the THUMOS’14 and BEOID datasets.
Zhenying Fang, Jianping Fan 0001, Jun Yu 0002
Multim. Syst.1
2021 Complementary Temporal Classification Activation Maps in Temporal Action Localization
Suguo Zhu, Zhenying Fang
PRCV (2)4
2020 Proposal Complementary Action Detection
abstract
Temporal action detection not only requires correct classification but also needs to detect the start and end times of each action accurately. However, traditional approaches always employ sliding windows or actionness to predict the actions, and it is different to train to model with sliding windows or actionness by end-to-end means. In this article, we attempt a different idea to detect the actions end-to-end, which can calculate the probabilities of actions directly through one network as one part of the results. We present PCAD, a novel proposal complementary action detector to deal with video streams under continuous, untrimmed conditions. Our approach first uses a simple fully 3D convolutional network to encode the video streams and then generates candidate temporal proposals for activities by using anchor segments. To generate more precise proposals, we also design a boundary proposal network to offer some complementary information for the candidate proposals. Finally, we learn an efficient classifier to classify the generated proposals into different activities and refine their temporal boundaries at the same time. Our model can achieve end-to-end training by jointly optimizing classification loss and regression loss. When evaluating on the THUMOS’14 detection benchmark, PCAD achieves state-of-the-art performance in high-speed models.
Suguo Zhu, Xiaoxian Yang, Jun Yu 0002, Zhenying Fang, Meng Wang 0001, Qingming Huang
ACM Trans. Multim. Comput. Commun. Appl.4
2019 PCPCAD: Proposal Complementary Action Detector
abstract
Temporal action detection is still a challenging task. This task not only requires correct classification, but also needs to accurately detect the start and end times of each action. In this paper, we present a novel proposal complementary action detector (PCAD) to deal with video streams under continuous, untrimmed conditions. Our approach first uses a simple fully 3D convolutional (Conv3D) network to encode the video streams and then generates candidate temporal proposals for activities by using anchor segments. To generate more precise proposals, we also designed a boundary proposal network (BPN) to offer some complementary information for the candidate proposals. Finally, we learn an efficient classifier to classify the generated proposals into different activities and refine their temporal boundaries at the same time. Our model can achieve end-to-end training by jointly optimizing classification loss and regression loss. When evaluating on THUMOS'14 detection benchmark, PCAD achieves the state-of-the-art performance in high-speed models.
Zhenying Fang, Suguo Zhu, Jun Yu 0002, Qi Tian 0001
ICME1
2019 Multimodal activity recognition with local block CNN and attention-based spatial weighted CNN
Suguo Zhu, Zhenying Fang, Jun Yu 0002, Junping Du 0001
J. Vis. Commun. Image Represent.2
2018 Hierarchical convolutional features for end-to-end representation-based visual tracking
Suguo Zhu, Zhenying Fang, Fei Gao 0006
Mach. Vis. Appl.2