Lingyu Si

dblp:298/0368 · DBLP profile ↗
← Back
31ranked-venue papers
5as first author
31since 2021 · last 2026
0000-0002-7735-6676ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 3 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 UVENet: A novel end-to-end model for temporal consistency in underwater video enhancement
Huijie Guo, Dazhao Du, Shouyou Huang, Changwen Zheng, Lingyu Si
Neural Networks6
2026 A dimensional structure based knowledge distillation method for cross-modal learning
Lingyu Si, Shouyou Huang, Junzhi Yu 0001, Fuchun Sun 0001
Neural Networks1
2026 On the Transferability and Discriminability of Representation Learning in Unsupervised Domain Adaptation
abstract
In this paper, we addressed the limitation of relying solely on distribution alignment and source-domain empirical risk minimization in Unsupervised Domain Adaptation (UDA). Our information-theoretic analysis showed that this standard adversarial-based framework neglects the discriminability of target-domain features, leading to suboptimal performance. To bridge this theoretical-practical gap, we defined "good representation learning" as guaranteeing both transferability and discriminability, and proved that an additional loss term targeting target-domain discriminability is necessary. Building on these insights, we proposed a novel adversarial-based UDA framework that explicitly integrates a domain alignment objective with a discriminability-enhancing constraint. Instantiated as Domain-Invariant Representation Learning with Global and Local Consistency (RLGLC), our method leverages Asymmetrically-Relaxed Wasserstein of Wasserstein Distance (AR-WWD) to address class imbalance and semantic dimension weighting, and employs a local consistency mechanism to preserve fine-grained target-domain discriminative information. Extensive experiments across multiple benchmark datasets demonstrate that RLGLC consistently surpasses state-of-the-art methods, confirming the value of our theoretical perspective and underscoring the necessity of enforcing both transferability and discriminability in adversarial-based UDA.
Wenwen Qiang, Ziyin Gu, Lingyu Si, Jiangmeng Li, Changwen Zheng, Fuchun Sun 0001, Hui Xiong 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 FreeStyle: Free lunch for text-guided style transfer using diffusion models
Feihong He, Fuhui Sun, Lingyu Si, Li Shen 0008
Pattern Recognit.5
2026 S2I-DiT: Unlocking the semantic-to-image transferability by fine-tuning large diffusion transformer models
Enze Xie, Chongjian Ge, Xiang Li 0041, Lingyu Si, Changwen Zheng, Zhenguo Li
Pattern Recognit.5
2026 A Physical Model-Guided Framework for Underwater Image Enhancement and Depth Estimation
abstract
Due to the selective absorption and scattering of light by diverse aquatic media, underwater images usually suffer from various visual degradations. Existing underwater image enhancement (UIE) approaches that combine underwater physical imaging models with neural networks often fail to accurately estimate imaging model parameters such as scene depth and veiling light, resulting in poor performance in certain scenarios. To address this issue, we propose a physical model-guided framework for jointly training a Deep Degradation Model (DDM) with any advanced UIE model. DDM includes three well-designed sub-networks to accurately estimate various imaging parameters: a veiling light estimation sub-network, a factors estimation sub-network, and a depth estimation sub-network. Based on the estimated parameters and the underwater physical imaging model, we impose physical constraints on the enhancement process by modeling the relationship between underwater images and desired clean images, i.e., outputs of the UIE model. Moreover, while our framework is compatible with any UIE model, we design a simple yet effective fully convolutional UIE model, termed UIEConv. UIEConv utilizes both global and local features through a dual-branch structure. UIEConv trained within our framework achieves remarkable enhancement results across diverse underwater scenes. Furthermore, as a byproduct of UIE, the trained depth estimation sub-network enables accurate underwater scene depth estimation. Extensive experiments conducted in various real underwater imaging scenarios, including deep-sea environments with artificial light sources, validate the effectiveness of our framework and the UIEConv model. Code is available at https://github.com/ddz16/UWEnhancer.
Dazhao Du, Lingyu Si, Fanjiang Xu, Jianwei Niu 0002, Fuchun Sun 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 SlotFusion: Object-Centric Audiovisual Feature Fusion with Slot Attention for Remote Sensing Scene Recognition
abstract
Despite significant advancements in remote sensing multimodal learning, particularly in image-image feature fusion, the exploration of audio-image feature fusion remains insufficient. Given the complexity and redundancy of ground objects in remote sensing images, accurately aligning audio features with image features during the fusion process is a critical challenge. In this paper, we introduce an object-centric feature fusion method named SlotFusion. By employing a slot attention-based feature decoupling module and a slot-based audiovisual feature fusion module, we transform modality features with complex semantic information into a set of slot features corresponding to object units and use gated activation units to adaptively implement object-centric feature fusion. Experiments on the Audio Visual Aerial Scene Recognition dataset (ADVANCE) demonstrate that the proposed SlotFusion significantly improves remote sensing scene recognition performance, with a 7.04% increase in overall accuracy compared to previous methods, achieving state-of-the-art results.
Fangzhou Han, Lamei Zhang, Lingyu Si
ICASSP4
2025 Less Yet Robust: Crucial Region Selection for Scene Recognition
abstract
Scene recognition, particularly for aerial and underwater images, often suffers from various types of degradation, such as blurring or overexposure. Previous works that focus on convolutional neural networks have been shown to be able to extract panoramic semantic features and perform well on scene recognition tasks. However, low-quality images still impede model performance due to the inappropriate use of high-level semantic features. To address these challenges, we propose an adaptive selection mechanism to identify the most important and robust regions with high-level features. Thus, the model can perform learning via these regions to avoid interference. implement a learnable mask in the neural network, which can filter high-level features by assigning weights to different regions of the feature matrix. We also introduce a regularization term to further enhance the significance of key high-level feature regions. Different from previous methods, our learnable matrix pays extra attention to regions that are important to multiple categories but may cause misclassification and sets constraints to reduce the influence of such regions. This is a plug-and-play architecture that can be easily extended to other methods. Additionally, we construct an Underwater Geological Scene Classification dataset to assess the effectiveness of our model. Extensive experimental results demonstrate the superiority and robustness of our proposed method over state-of-the-art techniques on two datasets.
Jianqi Zhang, Mengxuan Wang, Lingyu Si, Changwen Zheng, Fanjiang Xu
ICASSP4
2025 D2TR: Sea Clutter Suppression via Dynamic Dual-Tree Complex Wavelet Selection and Target-Guided Regularization
abstract
Marine radar imaging systems frequently encounter performance degradation caused by sea clutter interference in complex oceanic environments. While numerous sea clutter suppression methods have been proposed, their generalization capabilities remain substantially constrained by two critical factors: inadequate hierarchical feature extraction of both clutter patterns and target signatures, and insufficient utilization of domain-specific characteristics. Current approaches fail to comprehensively exploit the distinctive high-frequency sub-band manifestations of clutter in the frequency domain, while simultaneously neglecting the sparse distribution of targets in the spatial domain. To address these limitations, we propose a novel sea clutter suppression algorithm. First, a dynamic dual-tree complex wavelet selection mechanism is introduced to enhance frequency-domain feature extraction, effectively capturing clutter features across high-frequency subbands and achieving better separation of clutter and targets. Second, leveraging the sparse distribution of targets in PPI image, a target-guided regularization is designed to strengthen spatial-domain target feature extraction, suppress background interference and preserve target information. Finally, a comprehensive large-scale sea clutter suppression dataset, RadarP-PIV2, is developed, encompassing diverse and complex scenarios to enable systematic qualitative and quantitative evaluations. Experimental results demonstrate that the proposed algorithm significantly outperforms existing methods in sea clutter suppression tasks, showcasing superior practical value and robust generalization capabilities.
Yuezheng Lv, Lingyu Si
ICIP6
2025 Advancing Complex Wide-Area Scene Understanding with Hierarchical Coresets Selection
abstract
Scene understanding is one of the core tasks in computer vision, aiming to extract semantic information from images to identify objects, scene categories, and their interrelationships. Although advancements in Vision-Language Models (VLMs) have driven progress in this field, existing VLMs still face challenges in adaptation to unseen complex wide-area scenes. To address the challenges, this paper proposes a Hierarchical Coresets Selection (HCS) mechanism to advance the adaptation of VLMs in complex wide-area scene understanding. It progressively refines the selected regions based on the proposed theoretically guaranteed importance function, which considers utility, representativeness, robustness, and synergy. Without requiring additional fine-tuning, HCS enables VLMs to achieve rapid understandings of unseen scenes at any scale using minimal interpretable regions while mitigating insufficient feature density. HCS is a plug-and-play method that is compatible with any VLM. Experiments demonstrate that HCS achieves superior performance and universality in various tasks. The code is available at https://wangjingyao07.github.io/HCS.github.io/.
Lingyu Si, Changwen Zheng
ACM Multimedia3
2025 UIEDP: Boosting underwater image enhancement with diffusion prior
Dazhao Du, Enhan Li, Lingyu Si, Wenlong Zhai, Fanjiang Xu, Jianwei Niu 0002, Fuchun Sun 0001
Expert Syst. Appl.3
2025 Breaking the alignment barrier: A spatiotemporal alignment-free RGBT tracking approach
Meibo Lv, Daming Zhou, Lingyu Si, Ruiheng Zhang 0001
Neurocomputing4
2025 Cognition-Driven Structural Prior for Instance-Dependent Label Transition Matrix Estimation
abstract
The label transition matrix has emerged as a widely accepted method for mitigating label noise in machine learning. In recent years, numerous studies have centered on leveraging deep neural networks to estimate the label transition matrix for individual instances within the context of instance-dependent noise. However, these methods suffer from low search efficiency due to the large space of feasible solutions. Behind this drawback, we have explored that the real murderer lies in the invalid class transitions, that is, the actual transition probability between certain classes is zero but is estimated to have a certain value. To mask the invalid class transitions, we introduced a human-cognition-assisted method with structural information from human cognition. Specifically, we introduce a structured transition matrix network (STMN) designed with an adversarial learning process to balance instance features and prior information from human cognition. The proposed method offers two advantages: 1) better estimation effectiveness is obtained by sparing the transition matrix and 2) better estimation accuracy is obtained with the assistance of human cognition. By exploiting these two advantages, our method parametrically estimates a sparse label transition matrix, effectively converting noisy labels into true labels. The efficiency and superiority of our proposed method are substantiated through comprehensive comparisons with state-of-the-art methods on three synthetic datasets and a real-world dataset. Our code will be available at https://github.com/WheatCao/STMN-Pytorch.
Ruiheng Zhang 0001, Zhe Cao 0001, Shuo Yang 0006, Lingyu Si, Lixin Xu 0001, Fuchun Sun 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Rethinking Causal Relationships Learning in Graph Neural Networks
abstract
Graph Neural Networks (GNNs) demonstrate their significance by effectively modeling complex interrelationships within graph-structured data. To enhance the credibility and robustness of GNNs, it becomes exceptionally crucial to bolster their ability to capture causal relationships. However, despite recent advancements that have indeed strengthened GNNs from a causal learning perspective, conducting an in-depth analysis specifically targeting the causal modeling prowess of GNNs remains an unresolved issue. In order to comprehensively analyze various GNN models from a causal learning perspective, we constructed an artificially synthesized dataset with known and controllable causal relationships between data and labels. The rationality of the generated data is further ensured through theoretical foundations. Drawing insights from analyses conducted using our dataset, we introduce a lightweight and highly adaptable GNN module designed to strengthen GNNs' causal learning capabilities across a diverse range of tasks. Through a series of experiments conducted on both synthetic datasets and other real-world datasets, we empirically validate the effectiveness of the proposed module. The codes are available at https://github.com/yaoyao-yaoyao-cell/CRCG.
Hang Gao 0004, Chengyu Yao, Jiangmeng Li, Lingyu Si, Fengge Wu, Changwen Zheng, Huaping Liu 0001
AAAI4
2024 Self-Supervised Representation Learning with Meta Comprehensive Regularization
abstract
Self-Supervised Learning (SSL) methods harness the concept of semantic invariance by utilizing data augmentation strategies to produce similar representations for different deformations of the same input. Essentially, the model captures the shared information among multiple augmented views of samples, while disregarding the non-shared information that may be beneficial for downstream tasks. To address this issue, we introduce a module called CompMod with Meta Comprehensive Regularization (MCR), embedded into existing self-supervised frameworks, to make the learned representations more comprehensive. Specifically, we update our proposed model through a bi-level optimization mechanism, enabling it to capture comprehensive features. Additionally, guided by the constrained extraction of features using maximum entropy coding, the self-supervised learning model learns more comprehensive features on top of learning consistent features. In addition, we provide theoretical support for our proposed method from information theory and causal counterfactual perspective. Experimental results show that our method achieves significant improvement in classification, object detection and semantic segmentation tasks on multiple benchmark datasets.
Huijie Guo, Ying Ba, Jie Hu 0019, Lingyu Si, Wenwen Qiang, Lei Shi 0002
AAAI4
2024 CartoonDiff: Training-free Cartoon Image Generation with Diffusion Transformer Models
abstract
Image cartoonization has attracted significant interest in the field of image generation. However, most of the existing image cartoonization techniques require re-training models using images of cartoon style. In this paper, we present CartoonDiff, a novel training-free sampling approach which generates image cartoonization using diffusion transformer models. Specifically, we decompose the reverse process of diffusion models into the semantic generation phase and the detail generation phase. Furthermore, we implement the image cartoonization process by normalizing high-frequency signal of the noisy image in specific denoising steps. CartoonDiff doesn’t require any additional reference images, complex model designs, or the tedious adjustment of multiple parameters. Extensive experimental results show the powerful ability of our CartoonDiff. The project page is available at: https://cartoondiff.github.io/
Feihong He, Lingyu Si, Leilei Yan, Shimeng Hou, Fanzhang Li
ICASSP3
2024 Radardiff: Improving Sea Clutter Suppression Using Diffusion Models for Radar Images
abstract
Marine radar is employed across multiple fields, notably in navigation, meteorology, defense, and security. Marine radar images are highly sensitive to sea clutter, highlighting the crucial importance of sea clutter suppression in radar image processing. However, existing algorithms for sea clutter suppression often struggle to effectively generalize in complex marine environments. In this paper, we introduce RadarDiff, a novel approach that leverages diffusion models to enhance sea clutter suppression in marine radar plan-position indicator (PPI) images. We treat sea clutter suppression as an image-to-image translation task and propose a novel data augmentation method to create image pairs with and without sea clutter. Additionally, we introduce a unique loss function designed to address the challenge of small targets disappearing after suppression. To our knowledge, we are the first to utilize the diffusion-based model in sea clutter suppression for radar PPI images. Our quantitative and qualitative results demonstrate significant improvements compared to traditional denoising methods and classical GAN-based models.
Lingyu Si, Changwen Zheng, Fanjiang Xu, Fuchun Sun 0001
ICASSP1
2024 FPGNet: Single Image Deraining with High-Frequency Channel and Frequency Domain Prior Guidance
abstract
In recent years, deep learning methods have shown promising results in Single Image Deraining (SID). However, these methods still suffer from unsatisfactory residual rain streaks, primarily due to the absence of image priors embedding and limitations in modeling capacity. In this paper, we propose a joint High-Frequency Channel and Fourier Frequency Domain Guided Network (FPGNet) to remove complex rain streaks while obtaining clearer background details. Specifically, FPGNet comprises two core components: the High-frequency Guidance Module (HFGM) and the Fourier-based Multi-scale Feature Extraction Module (FMFEM). The HFGM focuses on learning the natural distribution of rain streaks from the high-frequency channel to guide the network’s modeling of rain streaks, while the FMFEM serves as the backbone of FPGNet to learn rain layers in rainy images. Additionally, we introduce an Interactive Fusion Module (IFM) to enhance the integration of the high-frequency channel into the network. Extensive experiments demonstrate that our network outperforms state-of-the-art deraining methods in effectively removing rain streaks from rainy images.
Zhaoyong Yan, Lingyu Si
ICASSP3
2024 Regularized Hypothesis-Induced Wasserstein Divergence for unsupervised domain adaptation
Lingyu Si, Wenwen Qiang, Changwen Zheng, Junzhi Yu 0001, Fuchun Sun 0001
Knowl. Based Syst.1
2024 A Novel Causal Inference-Guided Feature Enhancement Framework for PolSAR Image Classification
abstract
In recent years, there has been a prominent focus on enhancing the quality of features derived from convolutional neural networks (CNNs) within the field of polarimetric synthetic aperture radar (PolSAR) image classification. Targeting this challenge, this article first visualizes the lack of discriminability and generalizability in CNN features through several empirical observations. Subsequently, we explain why these problems arise from a causal perspective, accomplished by means of a structural causal model (SCM) constructed according to the training and testing process of CNNs. This SCM facilitates the identification of variables that affect the quality of PolSAR image feature learning, as well as an intervention on those variables using backdoor adjustment. Building upon this groundwork, a novel causal inference-guided feature enhancement framework is constructed. It can be seamlessly integrated into any CNN-based PolSAR image classifier in a plug-and-play manner, enabling the enhanced classifier to filter out interference information and prevent model overfitting. These two aspects bring better feature discriminability and generalizability, respectively, leading to improved classification performance. Experimental results on four widely-used PolSAR image datasets demonstrate the effectiveness of our proposed framework. We integrate it into several mainstream methods in the field and show that the accuracy of the enhanced classifier is improved compared to the original model.
Lingyu Si, Wenwen Qiang, Lamei Zhang, Junzhi Yu 0001, Yuquan Wu, Changwen Zheng, Fuchun Sun 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Improving SAR Automatic Target Recognition via Trusted Knowledge Distillation From Simulated Data
abstract
In recent years, significant research has been conducted on utilizing simulated data to support Synthetic Aperture Radar Automatic Target Recognition (SAR-ATR) based on deep learning techniques. By distilling the dark knowledge extracted from simulated samples, quality of the learned representations on measured samples can be effectively enhanced. However, our study highlights an important oversight in previous works: unquestioning trust on all simulated samples inevitably introduces the part of dark knowledge that is detrimental to SAR-ATR performance. To address this issue, we introduce evidential learning to estimate the confidence degree of the model after inputting simulated samples, thereby assessing the validity of the dark knowledge to be distilled. Then, the simulated-measured knowledge distillation process will be carried out in a trusted manner. Specifically, we encourage the model to prioritize distilling the dark knowledge with higher validity while avoiding the influence of inferior knowledge through a dynamic confidence weighting method. Additionally, we transform the standard logits-based knowledge distillation loss function into a feature-based one, giving the proposed method the ability to plug-and-play. The above aspects constitute the proposed trusted simulated-measured knowledge distillation method for SAR-ATR. Multiple comparative studies on the Simulated And Measured Paired Labeled Experiment (SAMPLE) dataset demonstrate the effectiveness of our proposed method, which not only achieves superior performance but also maintains the desired computational complexity in the inference phase.
Fangzhou Han, Lingyu Si, Lamei Zhang
IEEE Trans. Geosci. Remote. Sens.3
2024 A Trusted Generative-Discriminative Joint Feature Learning Framework for Remote Sensing Image Classification
abstract
Remote sensing image (RSI) classification is a popular research topic that aims to assign semantic labels to images acquired from aerial or maritime platforms. Existing deep feature learning methods for this task can be divided into two paradigms: generative and discriminative. The former methods are good at capturing every local detail of images, while the later approaches focus on the most salient area. The significant differences between the two types of methods, both in terms of their underlying mechanisms and practical implementation, motivate us to integrate information acquired by both paradigms by exploiting their complementary strengths. However, this idea faces a challenge that local information in the extracted features, especially those from generative methods, may not be reliable for RSI classification. The reason for this challenge is that, due to the characteristics of the ground observation perspective, some RSIs, while semantically different, exhibit a significant degree of similarity in local details. This phenomenon leads to insufficient discriminability of local features to separate multiple RSI categories, which implies that the classification results overly focused on local information may be unreliable. To address this issue, in this article, we propose a novel framework that integrates generative and discriminative feature learning methods with evidential learning for RSI classification. Our framework uses the Dirichlet distribution to model the predicted probabilities to be integrated, thereby collecting evidence about their reliability. This enables us to integrate multiple features at an evidence level and make reliable decisions, overcoming the unreliabilities of generative-discriminative joint feature learning induced by RSI characteristics. We evaluate the proposed framework on several satellite and shipborne RSI classification datasets. The experimental results show that our method outperforms the state-of-the-art baselines in terms of accuracy and robustness.
Lingyu Si, Wenwen Qiang, Zeen Song, Bo Du 0001, Junzhi Yu 0001, Fuchun Sun 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Part-Aware Correlation Networks for Few-Shot Learning
abstract
Few-shot learning brings the machine close to human thinking which enables fast learning with limited samples. Recent work considers local features to achieve contextual semantic complementation, while they are merely coarsened feature observations that can only extract insignificant label correlations. On the contrary, partial properties of few-shot examples significantly draw the implicit feature observations that can reveal the underlying label correlation of rare label classification. To fully explore the correlation between labels and partial features, this paper proposes a Part-Aware Correlation Network (PACNet) based on Partial Representation (PR) and Semantic Covariance Matrix (SCM). Specifically, we develop a partial representing module of an object that eliminates object-independent information and allows the model to focus on more distinctive parts. Furthermore, a semantic covariance measure function is redefined as a way to learn the semantic relationships of partial representations and to compute the partial similarity between the query sample and the support set. Experiments on three benchmark datasets consistently show that the proposed method outperforms the state-of-the-art counterparts,e.g., on the PartImageNet dataset, the performance gains of up to 12% and 5.9% are observed for the 5-way 1-shot and 5-way 5-shot settings, respectively.
Ruiheng Zhang 0001, Jinyu Tan, Zhe Cao 0001, Lixin Xu 0001, Lingyu Si, Fuchun Sun 0001
IEEE Trans. Multim.6
2023 Robust Causal Graph Representation Learning against Confounding Effects
abstract
The prevailing graph neural network models have achieved significant progress in graph representation learning. However, in this paper, we uncover an ever-overlooked phenomenon: the pre-trained graph representation learning model tested with full graphs underperforms the model tested with well-pruned graphs. This observation reveals that there exist confounders in graphs, which may interfere with the model learning semantic information, and current graph representation learning methods have not eliminated their influence. To tackle this issue, we propose Robust Causal Graph Representation Learning (RCGRL) to learn robust graph representations against confounding effects. RCGRL introduces an active approach to generate instrumental variables under unconditional moment restrictions, which empowers the graph representation learning model to eliminate confounders, thereby capturing discriminative information that is causally related to downstream predictions. We offer theorems and proofs to guarantee the theoretical effectiveness of the proposed approach. Empirically, we conduct extensive experiments on a synthetic dataset and multiple benchmark datasets. Experimental results demonstrate the effectiveness and generalization ability of RCGRL. Our codes are available at https://github.com/hang53/RCGRL.
Hang Gao 0004, Jiangmeng Li, Wenwen Qiang, Lingyu Si, Changwen Zheng, Fuchun Sun 0001
AAAI4
2023 Disentangle and Remerge: Interventional Knowledge Distillation for Few-Shot Object Detection from a Conditional Causal Perspective
abstract
Few-shot learning models learn representations with limited human annotations, and such a learning paradigm demonstrates practicability in various tasks, e.g., image classification, object detection, etc. However, few-shot object detection methods suffer from an intrinsic defect that the limited training data makes the model cannot sufficiently explore semantic information. To tackle this, we introduce knowledge distillation to the few-shot object detection learning paradigm. We further run a motivating experiment, which demonstrates that in the process of knowledge distillation, the empirical error of the teacher model degenerates the prediction performance of the few-shot object detection model as the student. To understand the reasons behind this phenomenon, we revisit the learning paradigm of knowledge distillation on the few-shot object detection task from the causal theoretic standpoint, and accordingly, develop a Structural Causal Model. Following the theoretical guidance, we propose a backdoor adjustment-based knowledge distillation method for the few-shot object detection task, namely Disentangle and Remerge (D&R), to perform conditional causal intervention toward the corresponding Structural Causal Model. Empirically, the experiments on benchmarks demonstrate that D&R can yield significant performance boosts in few-shot object detection. Code is available at https://github.com/ZYN-1101/DandR.git.
Jiangmeng Li, Wenwen Qiang, Lingyu Si, Chengbo Jiao, Changwen Zheng, Fuchun Sun 0001
AAAI4
2023 Do We Really Need Temporal Convolutions in Action Segmentation?
abstract
Recognizing and segmenting actions from long videos is a challenging problem. Most existing methods focus on designing temporal convolutional models. However, these models are limited in their flexibility and ability to model long-term dependencies. Transformers have recently been used in various tasks. But the lack of inductive bias and the inefficiency of handling long video sequences limit the application of Transformers in action segmentation. In this paper, we present a pure Transformer-based model without temporal convolutions in action segmentation, called Temporal U-Transformer. The U-Transformer architecture not only reduces complexity but also introduces an inductive bias that neighboring frames are more likely to belong to the same class. Besides, we further propose a boundary-aware loss based on the distribution of similarity scores between frames from attention modules to improve the ability to recognize boundaries. Extensive experiments show the effectiveness of our method.
Dazhao Du, Yu Li 0003, Zhongang Qi, Lingyu Si, Ying Shan
ICME5
2023 Timestamp-Supervised Action Segmentation from the Perspective of Clustering
abstract
Video action segmentation under timestamp supervision has recently received much attention due to lower annotation costs. Most existing methods generate pseudo-labels for all frames in each video to train the segmentation model. However, these methods suffer from incorrect pseudo-labels, especially for the semantically unclear frames in the transition region between two consecutive actions, which we call ambiguous intervals. To address this issue, we propose a novel framework from the perspective of clustering, which includes the following two parts. First, pseudo-label ensembling generates incomplete but high-quality pseudo-label sequences, where the frames in ambiguous intervals have no pseudo-labels. Second, iterative clustering iteratively propagates the pseudo-labels to the ambiguous intervals by clustering, and thus updates the pseudo-label sequences to train the model. We further introduce a clustering loss, which encourages the features of frames within the same action segment more compact. Extensive experiments show the effectiveness of our method.
Dazhao Du, Enhan Li, Lingyu Si, Fanjiang Xu, Fuchun Sun 0001
IJCAI3
2022 SimViT: Exploring a Simple Vision Transformer with Sliding Windows
abstract
Although vision Transformers have achieved excellent performance as backbone models in many vision tasks, most of them intend to capture global relations of all tokens in an image or a window, which disrupts the inherent spatial and local correlations between patches in 2D structure. In this paper, we introduce a simple vision Transformer named SimViT, to incorporate spatial structure and local information into the vision Transformers. Specifically, we introduce Multi-head Central Self-Attention(MCSA) instead of conventional Multi-head Self-Attention to capture highly local relations. The introduction of sliding windows facilitates the capture of spatial structure. Meanwhile, SimViT extracts multiscale hierarchical features from different layers for dense prediction tasks. Extensive experiments show the SimViT is effective and efficient as a general-purpose backbone model for various image processing tasks. Especially, our SimViT-Micro only needs 3.3M parameters to achieve 71.1% top-1 accuracy on ImageNet-1k dataset, which is the smallest size vision Transformer model by now.
Lingyu Si, Changwen Zheng
ICME4
2022 Convolutional Transformer with Similarity-based Boundary Prediction for Action Segmentation
abstract
Action classification has made great progress, but segmenting and recognizing actions from long videos remains a challenging problem. Recently, Transformer-based models with strong sequence modeling ability have succeeded in many se-quence modeling tasks. However, the lack of inductive bias and the difficulty of handling long video sequences limit the application of the Transformer in the action segmentation task. In order to explore the potential of the Transformer in this task, we replace some specific linear layers in the vanilla Transformer with dilated temporal convolution, and a sparse attention mechanism is utilized to reduce the time and space complexities to process long video sequences. Besides, directly using frame-wise classification loss to train the model will cause that frames at boundaries of actions are treated equally with those in the middle of actions, and the learned features are not sensitive to boundaries. We propose a new local log-context attention module to predict whether each frame is at the beginning, middle, or end of an action. Since boundary frames are similar to their neighboring frames of different classes, our similarity-based boundary prediction helps learn more discriminative features. Extensive experiments on three datasets show the effectiveness of our method.
Dazhao Du, Yu Li 0003, Zhongang Qi, Lingyu Si, Ying Shan
ICTAI5
2022 Bootstrapping Informative Graph Augmentation via A Meta Learning Approach
abstract
Recent works explore learning graph representations in a self-supervised manner. In graph contrastive learning, benchmark methods apply various graph augmentation approaches. However, most of the augmentation methods are non-learnable, which causes the issue of generating unbeneficial augmented graphs. Such augmentation may degenerate the representation ability of graph contrastive learning methods. Therefore, we motivate our method to generate augmented graph with a learnable graph augmenter, called MEta Graph Augmentation (MEGA). We then clarify that a "good" graph augmentation must have uniformity at the instance-level and informativeness at the feature-level. To this end, we propose a novel approach to learning a graph augmenter that can generate an augmentation with uniformity and informativeness. The objective of the graph augmenter is to promote our feature extraction network to learn a more discriminative feature representation, which motivates us to propose a meta-learning paradigm. Empirically, the experiments across multiple benchmark datasets demonstrate that MEGA outperforms the state-of-the-art methods in graph self-supervised learning tasks. Further experimental studies prove the effectiveness of different terms of MEGA. Our codes are available at https://github.com/hang53/MEGA.
Hang Gao 0004, Jiangmeng Li, Wenwen Qiang, Lingyu Si, Fuchun Sun 0001, Changwen Zheng
IJCAI4
2022 Multi-view representation learning from local consistency and global alignment
Lingyu Si, Wenwen Qiang, Jiangmeng Li, Fanjiang Xu, Funchun Sun
Neurocomputing1