Taiyi Su

dblp:264/6774 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0002-4357-1095ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Structure-Preserved Superpixel Perception for Referring Image Segmentation
abstract
Referring image segmentation aims to identify and segment objects in images based on linguistic expressions. Existing methods typically match individual pixels with linguistic words (pixel-word perception) and use interpolation upsampling for segmentation. However, these approaches encounter two limitations. First, as a meaningful word describes an entire entity while a pixel merely captures fragmented visual cues, the widely used pixel-word perception experiences semantic hierarchy misalignment, resulting in the misjudgment of the target entity. Second, the approximate estimation in the interpolation upsampling misleads region segmentation, struggles to preserve the target entity boundaries during feature reconstruction, and causes oversegmentation of non-target regions or undersegmentation of critical parts. To address these issues, a structure-preserved superpixel perception (SSP) framework is designed, which integrates a superpixel perception mechanism (SPM) to refine object identification and a replication upsampling strategy (RUS) to preserve appearance integrity. Specifically, SPM starts by utilizing superpixel clustering and language guidance to mine the knowledge of target object efficiently. Meanwhile, potential target regions are also extracted by SPM through traditional pixel-word associations. Then, SPM refines the recognition of target entity by establishing relationships between pixel-level representation and superpixel-level knowledge. Regarding RUS, it can directly replicate spatial and channel information for more precise feature reconstruction, serving as an alternative to interpolation upsampling. RUS consists of two complementary components: spatial replication and channel replication, which account for the assurance of target entity’s structural integrity and preservation of fine-grained details, respectively. Extensive experiments on the benchmark RefCOCO, RefCOCO+ and G-Ref datasets demonstrate that the proposed SSP framework achieves superior performance while using fewer FLOPs compared to the state-of-the-art methods. The source code of this work can be found in https://mic.tongji.edu.cn.
Taiyi Su, Shanshan Du, Hanli Wang
IEEE Trans. Circuits Syst. Video Technol.2
2026 Prompted Contrastive Learning for Skeleton-Based Action Recognition
abstract
Large pre-trained vision-language models have shown great potential across various visual understanding tasks. However, in skeleton-based action recognition, it is challenging to design suitable prompts to guide the model in understanding actions depicted by skeletal structures. Moreover, the modality gap between skeleton data and text descriptions poses obstacles to effective cross-modality contrastive learning. In this work, we propose a prompted contrastive learning framework for skeleton-based action recognition. Specifically, to automatically generate natural language descriptions of actions, we introduce prompted contrast with knowledge engine (PCKE), which utilizes a pre-trained large language model as the knowledge engine to generate text prompts for the input of text encoder. Moreover, as the prompt words generated by the language model may not always be optimal, we explore the potential of the model to autonomously learn prompts. To this end, we present prompted contrast with learning engine (PCLE), a straightforward method that replaces the prompts’ context words with learnable vectors. After separately encoding skeleton data and text prompts with the skeleton and text encoder, a skeleton-language interaction process is implemented to optimize feature alignment and reduce modality gap. The interaction module contains two components: a skeleton-language bridger to connect the representation spaces of the skeleton and text encoders, and cross-modality attention to fuse information of both modalities, facilitating cross-modal knowledge transfer and refining feature alignment. Extensive experiments on four benchmark datasets of NTU RGB+D, NTU RGB+D 120, NW-UCLA, and PKU-MMD demonstrate the effectiveness of the proposed method. The source code of this work can be found in https://mic.tongji.edu.cn.
Taiyi Su, Hanli Wang
IEEE Trans. Circuits Syst. Video Technol.1
2025 WaveCL: Wavelet Calibration Learning for Referring Video Object Segmentation
abstract
Referring video object segmentation (RVOS) focuses on segmenting target objects in a video based on natural language descriptions. However, existing methods typically rely on text cues that are unrelated to video content, and the target entity is only recognized in the pixel space. This often leads to ambiguous cross-modal understanding and fragmented perception across space and time, resulting in inaccurate or incomplete segmentation of the target objects. To address these challenges, a novel wavelet calibration learning (WaveCL) framework is proposed to unify cross-modal understanding and preserve spatial-temporal integrity of the target object. The WaveCL framework is built on two core components: semantic-calibrated entity perception (SEP) and wavelet-guided integrity perception (WIP). SEP aligns the textual semantics with video content, enabling more accurate and context-aware cross-modal understanding. WIP, on the other hand, leverages wavelet representations to capture fine-grained details of the target object from a global spatial-temporal perspective. By refining wavelet clues with the guidance of text queries, WIP enhances the integrity of segmentation. Through the collaboration of SEP and WIP, WaveCL enables precise, target-specific segmentation with detailed boundaries and consistent spatial-temporal perception. Extensive experiments on four benchmark datasets of Ref-YouTube-VOS, Ref-DAVIS17, A2D-Sentences, and JHMDB-Sentences show that WaveCL outperforms existing state-of-the-art methods. The source code of this work can be found in https://mic.tongji.edu.cn.
Taiyi Su, Hanli Wang
ACM Multimedia2
2025 Similarity Shuffled Criss-Cross Transformer With Angle Loss for Image-Text Matching
abstract
Image-text matching aims to retrieve images from the guidance of textual queries or retrieve text expressions with the help of images. Existing Transformer-based methods compute attention for all tokens and thus suffer from redundant information, resulting in inadequate focus on salient features. On the other hand, the widely adopted bidirectional ranking loss overlooks the importance of expanding the distance between positive and negative samples, leading to the misclassification of negative samples as positive ones. In this work, we propose similarity shuffled criss-cross Transformer (SSCT) with angle loss for image-text matching. Specifically, a grouping-shuffling operation is introduced to better distinguish salient features from redundant information, bypassing the need for fully connected mapping. The grouping-shuffling operation establishes channel dependencies across different groups of feature representations, enhancing salient features while suppressing unimportant ones. Then, a criss-cross attention mechanism that equips self-attention with a novel criss-cross convolution is designed to make isolated information cooperatively express integral semantics. Moreover, a novel angle loss is introduced to expand the distances between positive and negative samples. Extensive experiments on the benchmark datasets of MSCOCO and Flickr30K demonstrate that the proposed methods achieve superior performances compared to state-of-the-art methods.
Taiyi Su, Hanli Wang, Zhangkai Ni
IEEE Trans. Multim.2
2024 Memory-Based Contrastive Learning with Optimized Sampling for Incremental Few-Shot Semantic Segmentation
abstract
Incremental few-shot semantic segmentation (IFSS) aims to incrementally expand a semantic segmentation model’s ability to identify new classes based on few samples. However, it grapples with the dual challenges of catastrophic forgetting (due to feature drift in old classes) and overfitting (triggered by inadequate samples in new classes). To address these issues, a novel approach is proposed to integrate pixel-wise and region-wise contrastive learning, complemented by an optimized example and anchor sampling strategy. The proposed method incorporates a region memory and pixel memory designed to explore the high-dimensional embedding space more effectively. The memory, retaining the feature embeddings of known classes, facilitates the calibration and alignment of seen class features during the learning process of new classes. To further mitigate overfitting, the proposed approach implements an optimized example and anchor sampling strategy. Extensive experiments show the competitive performance of the proposed method. The source code of this work can be found in https://mic.tongji.edu.cn.
Miaojing Shi, Taiyi Su, Hanli Wang
ISCAS3
2024 Transductive Learning With Prior Knowledge for Generalized Zero-Shot Action Recognition
abstract
It is challenging to achieve generalized zero-shot action recognition. Different from the conventional zero-shot tasks which assume that the instances of the source classes are absent in the test set, the generalized zero-shot task studies the case that the test set contains both the source and the target classes. Due to the gap between visual feature and semantic embedding as well as the inherent bias of the learned classifier towards the source classes, the existing generalized zero-shot action recognition approaches are still far less effective than traditional zero-shot action recognition approaches. Facing these challenges, a novel transductive learning with prior knowledge (TLPK) model is proposed for generalized zero-shot action recognition. First, TLPK learns the prior knowledge which assists in bridging the gap between visual features and semantic embeddings, and preliminarily reduces the bias caused by the visual-semantic gap. Then, a transductive learning method that employs unlabeled target data is designed to overcome the bias problem in an effective manner. To achieve this, a target semantic-available approach and a target semantic-free approach are devised to utilize the target semantics in two different ways, where the target semantic-free approach exploits prior knowledge to produce well-performed semantic embeddings. By exploring the usage of the aforementioned prior-knowledge learning and transductive learning strategies, TLPK significantly bridges the visual-semantic gap and alleviates the bias between the source and the target classes. The experiments on the benchmark datasets of HMDB51 and UCF101 demonstrate the effectiveness of the proposed model compared to the state-of-the-art methods. The source code of this work can be found inhttps://mic.tongji.edu.cn
Taiyi Su, Hanli Wang, Qiuping Qi, Bin He 0003
IEEE Trans. Circuits Syst. Video Technol.1
2023 Adaptive Token Excitation with Negative Selection for Video-Text Retrieval
Juntao Yu, Zhangkai Ni, Taiyi Su, Hanli Wang
ICANN (7)3
2023 Multi-Level Content-Aware Boundary Detection for Temporal Action Proposal Generation
abstract
It is challenging to generate temporal action proposals from untrimmed videos. In general, boundary-based temporal action proposal generators are based on detecting temporal action boundaries, where a classifier is usually applied to evaluate the probability of each temporal action location. However, most existing approaches treat boundaries and contents separately, which neglect that the context of actions and the temporal locations complement each other, resulting in incomplete modeling of boundaries and contents. In addition, temporal boundaries are often located by exploiting either local clues or global information, without mining local temporal information and temporal-to-temporal relations sufficiently at different levels. Facing these challenges, a novel approach named multi-level content-aware boundary detection (MCBD) is proposed to generate temporal action proposals from videos, which jointly models the boundaries and contents of actions and captures multi-level (i.e., frame level and proposal level) temporal and context information. Specifically, the proposed MCBD preliminarily mines rich frame-level features to generate one-dimensional probability sequences, and further exploits temporal-to-temporal proposal-level relations to produce two-dimensional probability maps. The final temporal action proposals are obtained by a fusion of the multi-level boundary and content probabilities, achieving precise boundaries and reliable confidence of proposals. The extensive experiments on the three benchmark datasets of THUMOS14, ActivityNet v1.3 and HACS demonstrate the effectiveness of the proposed MCBD compared to state-of-the-art methods. The source code of this work can be found in https://mic.tongji.edu.cn.
Taiyi Su, Hanli Wang
IEEE Trans. Image Process.1
2022 Learned Image Compression with Multi-Scale Spatial and Contextual Information Fusion
abstract
Although learned image compression based on convolution neural network and hyperprior makes significant progress, the distinction between original and reconstructed images is still obvious. In order to reconstruct compressed image with higher quality, a novel model based on fusing multi-scale spatial and context information is proposed in this work. Since spatial information might be dropped during the forward propagation when neural networks go deeper, a multi-scale information fusion module is designed to help the encoder to retain the necessary spatial information while removing the redundancy in latent representation. Meanwhile, a multi-scale 3D context module with varying-sized masked 3D convolution kernels is devised to obtain multi-scale correlation in latent representation. The experiments demonstrate the superiority of the proposed approach over a number of state-of-the-art image compression methods and the versatile video coding.
Hanli Wang, Taiyi Su
ICIP3
2020 Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark Study
abstract
Existing enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions.
Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin
IEEE Trans. Image Process.24