Tianyi Zhang 0004

dblp:17/322-4 · DBLP profile ↗
← Back
13ranked-venue papers
10as first author
7since 2021 · last 2025
0000-0002-4474-914XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 9 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Edge-aware Affinity Enhancement for Image Manipulation Localization
abstract
Image manipulation localization (IML) refers to the task of identifying regions in images that have been altered by specific tampering techniques, such as copy-move, splicing, or inpainting. Transformers have been applied to IML tasks due to their ability to model long-range correlations between pixels through the self-attention mechanism. However, the inherent self-attention mechanism may not accurately model or sufficiently enhance these correlations, particularly in detailed edge traces, due to its limitations. In this paper, we propose an Edge-Aware Affinity Enhancement approach for the IML task. Specifically, we introduce an Affinity Regularization Module to establish inter-patch correlations for feature regularization via random walk propagation. Based on the extracted correlation representation, we propose an Edge-Affinity Guidance strategy to further refine the correlation accuracy, particularly in ambiguous edge regions. Extensive experimental results demonstrate that our method outperforms state-of-the-art image manipulation localization techniques in terms of localization accuracy.
Tianyi Zhang 0004, Qinglong Lin, Pengming Feng, Rubo Zhang
ACM Multimedia1
2025 PSRR-MaxpoolNMS++: Fast Non-Maximum Suppression With Discretization and Pooling
abstract
Non-maximum suppression (NMS) is an essential post-processing step for object detection. The de-facto standard for NMS, namely GreedyNMS, is not parallelizable and could thus be the performance bottleneck in object detection pipelines. MaxpoolNMS is introduced as a fast and parallelizable alternative to GreedyNMS. However, MaxpoolNMS is only capable of replacing the GreedyNMS at the first stage of two-stage detectors like Faster R-CNN. To address this issue, we observe that MaxpoolNMS employs the process of box coordinate discretization followed by local score argmax calculation, to discard the nested-loop pipeline in GreedyNMS to enable parallelizable implementations. In this paper, we introduce a simple Relationship Recovery module and a Pyramid Shifted MaxpoolNMS module to improve the above two stages, respectively. With these two modules, our PSRR-MaxpoolNMS is a generic and parallelizable approach, which can completely replace GreedyNMS at all stages in all detectors. Furthermore, we extend PSRR-MaxpoolNMS to the more powerful PSRR-MaxpoolNMS++. As for box coordinate discretization, we propose Density-based Discretization for better adherence to the target density of the suppression. As for local score argmax calculation, we propose an Adjacent Scale Pooling scheme for mining out the duplicated box pairs more accurately and efficiently. Extensive experiments demonstrate that both our PSRR-MaxpoolNMS and PSRR-MaxpoolNMS++ outperform MaxpoolNMS by a large margin. Additionally, PSRR-MaxpoolNMS++ not only surpasses PSRR-MaxpoolNMS but also attains competitive accuracy and much better efficiency when compared with GreedyNMS. Therefore, PSRR-MaxpoolNMS++ is a parallelizable NMS solution that can effectively replace GreedyNMS at all stages in all detectors.
Tianyi Zhang 0004, Chunyun Chen, Yun Liu 0011, Xue Geng, Mohamed M. Sabry, Jie Lin 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Weakly supervised temporal action localization: a survey
Ronglu Li, Tianyi Zhang 0004, Rubo Zhang
Multim. Tools Appl.2
2024 Integration of Global and Local Knowledge for Foreground Enhancing in Weakly Supervised Temporal Action Localization
abstract
Weakly Supervised Temporal Action Localization (WTAL) aims to identify the temporal duration of actions and classify the action categories with only video-level labels in the training stage. Motivated by the intuition that the attention maps generated from various views will assist in enhancing the foreground action temporal segments, in this paper we propose a WTAL pipeline based on a novel attention mechanism that effectively integrates global and local knowledge. Our attention mechanism is mainly composed of a global attention branch and a local attention branch. Specifically, the global attention branch is built on the inter-segment similarity to sparsely mine out the correlation knowledge within the entire video, while the local attention branch is built on the convolutional structure to densely aggregate the information within the fixed local respective field. Experiments on THUMOS14 and ActivityNet v1.3 datasets demonstrate the effectiveness of our proposed WTAL pipeline compared to state-of-the-art methods.
Tianyi Zhang 0004, Ronglu Li, Pengming Feng, Rubo Zhang
IEEE Trans. Multim.1
2022 Scalable Hardware Acceleration of Non-Maximum Suppression
abstract
Non-maximum Suppression (NMS) in one- and two-stage object detection deep neural networks (e.g., SSD and Faster-RCNN) is becoming the computation bottleneck. In this paper, we introduce a hardware acceleration for the scalable PSRR-MaxpoolNMS algorithm. Our architecture shows 75.0× and 305× speedups compared to the software implementation of the PSRR-MaxpoolNMS as well as the hardware implementations of GreedyNMS, respectively, while simultaneously achieving comparable Mean Average Precision (mAP) to software-based floating-point implementations. Our architecture is 13.4× faster than the state-of-the-art NMS one. Our accelerator supports both one- and two-stage detectors, while supporting very high input resolutions (i.e., FHD)—essential input size for better detection accuracy.
Chunyun Chen, Tianyi Zhang 0004, Zehui Yu, Adithi Raghuraman, Shwetalaxmi Udayan, Jie Lin 0001, Mohamed M. Sabry
DATE2
2021 PSRR-MaxpoolNMS: Pyramid Shifted MaxpoolNMS With Relationship Recovery
abstract
Non-maximum Suppression (NMS) is an essential post-processing step in modern convolutional neural networks for object detection. Unlike convolutions which are inherently parallel, the de-facto standard for NMS, namely GreedyNMS, cannot be easily parallelized and thus could be the performance bottleneck in convolutional object detection pipelines. MaxpoolNMS is introduced as a parallelizable alternative to GreedyNMS, which in turn enables faster speed than GreedyNMS at comparable accuracy. However, MaxpoolNMS is only capable of replacing the GreedyNMS at the first stage of two-stage detectors like Faster-RCNN. There is a significant drop in accuracy when applying MaxpoolNMS at the final detection stage, due to the fact that MaxpoolNMS fails to approximate GreedyNMS precisely in terms of bounding box selection. In this paper, we propose a general, parallelizable and configurable approach PSRR-MaxpoolNMS, to completely replace GreedyNMS at all stages in all detectors. By introducing a simple Relationship Recovery module and a Pyramid Shifted MaxpoolNMS module, our PSRR-MaxpoolNMS is able to approximate GreedyNMS more precisely than MaxpoolNMS. Comprehensive experiments show that our approach outperforms MaxpoolNMS by a large margin, and it is proven faster than GreedyNMS with comparable accuracy. For the first time, PSRR-MaxpoolNMS provides a fully parallelizable solution for customized hardware design, which can be reused for accelerating NMS everywhere.
Tianyi Zhang 0004, Jie Lin 0001, Peng Hu 0002, Mohamed M. Sabry
CVPR1
2021 Guided Co-Segmentation Network for Fast Video Object Segmentation
abstract
Semi-supervised video object segmentation is a task of propagating instance masks given in the first frame to the entire video. It is a challenging task since it usually suffers from heavy occlusions, large deformation, and large variations of objects. To alleviate these problems, many existing works apply time-consuming techniques such as fine-tuning, post-processing, or extracting optical flow, which makes them intractable for online segmentation. In our work, we focus on online semi-supervised video object segmentation. We propose a GCSeg (Guided Co-Segmentation) Network which is mainly composed of a Reference Module and a Co-segmentation Module, to simultaneously incorporate the short-term, middle-term, and long-term temporal inter-frame relationships. Moreover, we propose an Adaptive Search Strategy to reduce the risk of propagating inaccurate segmentation results in subsequent frames. Our GCSeg network achieves state-of-the-art performance on online semi-supervised video object segmentation on Davis 2016 and Davis 2017 datasets.
Weide Liu, Guosheng Lin, Tianyi Zhang 0004, Zichuan Liu
IEEE Trans. Circuits Syst. Video Technol.3
2020 Splitting Vs. Merging: Mining Object Regions with Discrepancy and Intersection Loss for Weakly Supervised Semantic Segmentation
Tianyi Zhang 0004, Guosheng Lin, Weide Liu, Jianfei Cai 0001, Alex Chichung Kot
ECCV (22)1
2019 Semantic Segmentation via Domain Adaptation with Global Structure Embedding
abstract
In this paper we focus on the problem of unsupervised domain adaptation for semantic segmentation. The previous works usually focus on adversarial learning either in pixel-level or feature-level. However, global structure knowledge is often neglected in the adversarial learning due to the possible reasons: First, the result of pixel-level adversarial learning does not necessarily preserve the semantic consistency of the input image. Second, global structure knowledge is not embedded to regularize the feature-level adversarial learning. In this work, we propose a framework for unsupervised domain adaptation in semantic segmentation which effectively incorporates pixel- level, feature-level adversarial learning and self-training strategy. Our framework embeds the global structure knowledge into the adversarial training step to tackle the problem of structure misalignment. Consequently, our proposed framework achieves the state-of-the-art semantic segmentation domain adaptation results on the task of transferring GTA5 to Cityscapes.
Tianyi Zhang 0004, Guosheng Lin, Jianfei Cai 0001, Alex Chichung Kot
VCIP1
2019 Task-in-all Domain Adaptation for Semantic Segmentation
abstract
In this work we tackle the problem of unsupervised domain adaptation for semantic segmentation. One pipeline is to sequentially train image-translation model and the final task segmentation model. In such pipeline, image translation is aimed to generate the translated source-domain images which are visually similar to the target-domain images and then the final task model is trained using the translated images and its corresponding groundtruth. However, the visually optimal translated-images are not necessarily optimal for the final task of segmenting the target-domain images. Thus we propose a Task-in-all pipeline for unsupervised domain adaptation on semantic segmentation, which incorporates image translation and final segmentation task into an end-to-end training pipeline. Our aim is to generate the translated images which better assists the final task, instead of just being visually similar to the target domain images. We show that in the task of adapting from GTA5 to Cityscapes dataset, the segmentation performance of our Task-in-all pipeline outperforms the sequentially training pipeline, with simpler model structure and less training complexity.
Tianyi Zhang 0004, Chuanxia Zheng, Guosheng Lin, Jianfei Cai 0001, Alex Chichung Kot
VCIP1
2019 Decoupled Spatial Neural Attention for Weakly Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation receives much research attention since it alleviates the need to obtain a large amount of dense pixel-wise ground-truth annotations for the training images. Compared with other forms of weak supervision, image labels are quite efficient to obtain. In this paper, we focus on the weakly supervised semantic segmentation with image label annotations. Recent progress for this task has been largely dependent on the quality of generated pseudo-annotations. In this paper, inspired by spatial neural-attention for image captioning, we propose a decoupled spatial neural attention network for generating pseudo-annotations. Our decoupled attention structure could simultaneously identify the object regions and localize the discriminative parts, which generates high-quality pseudo-annotations in one forward path. The generated pseudo-annotations lead to the segmentation results that achieve the state of the art in weakly supervised semantic segmentation.
Tianyi Zhang 0004, Guosheng Lin, Jianfei Cai 0001, Chunhua Shen, Alex Chichung Kot
IEEE Trans. Multim.1
2017 Action proposals using hierarchical clustering of super-trajectories
abstract
Action localization aims to determine the spatial and temporal location of certain action which appears in a video. To facilitate action localization, spatio-temporal proposals which are likely to contain the action of interest are extracted to reduce the search space of candidate locations in a video, inspired by the object proposals in images. In this paper, considering the effectiveness of spatio-temporal trajectories for video action recognition and action proposal generation, we build our unsupervised action proposal generation pipeline upon super-trajectories. Specifically, we first group trajectories into super-trajectories inspired by super-voxels, and then employ hierarchical clustering on super-trajectories by taking different aspect and temporal ratios into consideration. Comprehensive experiments on two benchmark datasets (i.e., UCF-sports and MSR-II) demonstrate that our action proposal generation pipeline not only achieves the state-of-the-art recall, but also achieves competitive results for the action localization task.
Tianyi Zhang 0004, Li Niu 0002, Jianfei Cai 0001, Alex Chichung Kot
VCIP1
2016 Efficient object feature selection for action recognition
abstract
Currently most action recognition or video classification tasks highly rely on the motion features such as state-of-the-art Improved Dense Trajectory (IDT) features. Despite the huge success, IDT features lack of rich static object-level information. In this paper, we make use of the object-level features for action recognition tasks. For efficiently and effectively processing large-scale video data, we propose a two-layer feature selection framework including local object feature selection (LS) and global feature selection (GS). Both of the selection methods can improve recognition accuracy while greatly reducing the feature dimension or feature processing complexity. Experimental results show that the selected object-level features contain complimentary information to IDT features and the combination with IDT features can further improve the recognition accuracy significantly.
Tianyi Zhang 0004, Yu Zhang 0004, Jianfei Cai 0001, Alex Chichung Kot
ICASSP1