Shenglan Liu 0001

dblp:34/8493-1 · also Sheng-Lan Liu 0001 · DBLP profile ↗
← Back
64ranked-venue papers
12as first author
37since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 39 · 10 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 1 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Interaction makes better segmentation: An interaction-based framework for temporal action segmentation
Minjie Xu, Jiajun Fan, Chenyu Xiao, Shenglan Liu 0001, Lin Feng 0001
Knowl. Based Syst.5
2025 DTOS: Dynamic Time Object Sensing with Large Multimodal Model
abstract
Existing multimodal large language models (MLLMs) face significant challenges in Referring Video Object Segmentation(RVOS). We identify three critical challenges: (C1) insufficient quantitative representation of textual numerical data, (C2) repetitive and degraded response templates for spatiotemporal referencing, and (C3) loss of visual information in video sampling queries lacking textual guidance. To address these, we propose a novel framework, Dynamic Time Object Sensing (DTOS), specifically designed for RVOS. To tackle (C1) and (C2), we introduce specialized tokens to construct multi-answer response templates, enabling regression of event boundaries and target localization. This approach improves the accuracy of numerical regression while mitigating the issue of repetitive degradation. To address (C3), we propose a Text-guided Clip Sampler (TCS) that selects video clips aligned with user instructions, preventing visual information loss and ensuring consistent temporal resolution. TCS is also applicable to Moment Retrieval tasks, with enhanced multimodal input sequences preserving spatial details and maximizing temporal resolution. DTOS demonstrates exceptional capability in flexibly localizing multiple spatiotemporal targets based on user-provided textual instructions. Extensive experiments validate the effectiveness of our approach, with DTOS achieving state-of-the-art performance in $\mathcal{J}{{\& }}\mathcal{F}$ scores: an improvement of +4.36 on MeViS, +4.48 on Ref-DAVIS17, and +3.02 on Ref-YT-VOS. Additionally, our TCS demonstrates exceptional performance in Moment Retrieval. The code is available at https://github.com/Maulog/OPEN-DTOS-LMM.
Jirui Tian, Jinrong Zhang 0001, Shenglan Liu 0001, Luhao Xu, Zhixiong Huang, Gao Huang 0001
CVPR3
2025 Flexible Streaming Temporal Action Segmentation with Diffusion Models
abstract
Temporal distribution shifts occur not only in low-dimensional time-series data but also in high-dimensional data like videos. This phenomenon leads to significant performance degeneration in video understanding methods such as streaming temporal action segmentation. To address this issue, we propose a flexible streaming temporal action segmentation model with diffusion models (FSTAS-DM). By utilizing streaming video clips with varying feature distributions as control conditions, our model can adapt to the shifts and inconsistency of the distribution between the training and testing domains. Additionally, we have introduced a multistage conditional control training strategy (MSCC), which enhances the temporal generalization ability of the model. Our method demonstrates commendable performance on datasets like GTEA, 50Salads, and Breakfast.
Wenjun Wen, Shenglan Liu 0001, Lin Feng 0001
ICME3
2025 Knowledge-Driven Visual Target Navigation: Dual Graph Navigation
abstract
In unknown environments, navigating a robot by a given image to a specific location or instance is critical and challenging. The existing end-to-end approaches require simultaneous implicit learning of multiple subtasks, and modular approaches depend on metric information. Both approaches face high computational demands, often leading to difficulties in real-time updates and limited generalization, making them challenging to implement on resource-constrained devices. To address these challenges, we propose Dual Graph Navigation (DGN), a knowledge-driven, lightweight image instance navigation framework. DGN builds an External Knowledge Graph (EKG) from small-scale datasets to capture prior object correlations, efficiently guiding target exploration. During exploration, DGN builds an Internal Knowledge Graph (IKG) using an instance-aware module, which records explored objects based on reachability relationships rather than precise metric information. The IKG dynamically updates the EKG, enhancing the robot's adaptability to the current environment. Together, they realize topological perception and reduce computational overhead. Furthermore, unlike approaches characterized by over-dependence between components, DGN employs a plug-and-play modular design that allows independent training and flexible replacement of functional modules, effectively enhancing generalization performance while reducing training and deployment costs. Experiments illustrate that DGN generalizes well in different simulation environments (AI2-THOR, Habitat), achieving state-of-the-art performance on the ProcTHOR-10K dataset. It is compatible with three distinct real-world robot platforms, including edge computing devices without CUDA support. It exhibits a decision-making speed of 3.8 to 5.5 times over baseline methods. Further details can be found on the project page: https://dogplanningloyo.github.io/DGN/.
Jiansong Pei, Bingcheng Dong, Guangsheng Li, Shenglan Liu 0001
ICRA7
2025 PMCFNet: Prompt-Guided Multi-scale Cross-Modal Fusion Network for Referring Remote Sensing Image Segmentation
Yuqiu Kong, Shenglan Liu 0001
PRCV (5)4
2025 Bridging the Point to Boundary Gap for Point-Supervised Temporal Action Localization with Single-Stage Inference
Junshi Yang, Shenglan Liu 0001, Xuhan Sheng, Yiheng Zhou, Lin Feng 0001, Jiajun Fan
PRCV (7)2
2025 Lightweight Edge-Guided Super-Resolution Network for Remote Sensing Images
abstract
Recently, deep learning-based remote sensing image super-resolution (RSISR) techniques have achieved significant progress, but challenges remain in preserving critical edge details essential for high-quality image reconstruction, These details are crucial for tasks like object recognition, change detection, and accurate analysis in remote sensing imagery. Furthermore, existing RSISR methods typically require substantial computational resources, making them unsuitable for resource-constrained edge devices. To address these challenges, we propose a novel Edge-Guided Super-Resolution Network (EGSRN). The network employs an Edge Extraction Module (Edge Net) to explicitly extract edge information from low-resolution images, combined with multi-layer Feature Extraction Modules (FEM) and an Edge Information Fusion (EIF) mechanism to progressively integrate edge and image features. This design enables precise recovery of edge details, significantly enhancing the overall visual quality of the reconstructed images. Edge-aware processing enhances visual fidelity while also improving the accuracy of downstream tasks, such as classification, object detection, and change analysis. Furthermore, the network incorporates lightweight designs such as depthwise separable convolutions and channel shuffling to effectively reduce computational demands. Comprehensive experiments were conducted on two remote sensing datasets, and the model’s parameter count and floating-point operations (FLOPs) were evaluated. Results demonstrate that the proposed method achieves an excellent balance between performance and model complexity, delivering superior super-resolution reconstruction quality while maintaining low computational costs, making it well-suited for resource-limited real-world applications.
Zhixiong Huang, Xinying Wang 0005, Shenglan Liu 0001, Lin Feng 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 WFA-SRNet: A Wavelet-Guided and Feature-Aware Network for Remote Sensing Image Super-Resolution
abstract
Recently, deep learning-based remote sensing image super-resolution (RSISR) methods have achieved remarkable progress. However, effectively preserving high-frequency details remains a significant challenge, as these features are critical for downstream tasks such as object detection, change analysis, and scene classification. Moreover, relying solely on the information contained in low-resolution images often results in the loss of structural details, thereby degrading reconstruction quality. To address these issues, we propose a novel Wavelet-guided and Feature-Aware Super-Resolution Network (WFA-SRNet). The proposed network adopts a dual-branch architecture, consisting of a Feature Extraction Block (FEB) and a High-Frequency Extraction Block (HFE), to collaboratively model semantic structures and fine-grained textures. Specifically, FEB integrates a Shift-Window Cross Attention (SWCA) mechanism and a dictionary-based similarity matching strategy to capture non-local self-similarities, while the HFE branch incorporates a wavelet-domain high-frequency modeling module (WD-HFE), which explicitly decomposes and reconstructs frequency components via Discrete Wavelet Transform (DWT) and Inverse DWT (IDWT) to enhance edge and texture recovery. Furthermore, a Fusion Attention (FA) module is designed to guide the integration of multi-source features from both semantic and high-frequency pathways. Extensive experiments on multiple benchmark remote sensing datasets demonstrate that WFA-SRNet achieves superior reconstruction performance, particularly in restoring structural and textural details. Additionally, the proposed method significantly improves the accuracy of downstream classification tasks, showing strong potential for practical RSISR applications.
Xinying Wang 0005, Zhixiong Huang, Shenglan Liu 0001, Lin Feng 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 DCR-SRNet: A Degradation-Contrastive and Wavelet-Guided Network for Blind Remote Sensing Image Super-Resolution
abstract
Recently, deep learning-based remote sensing image super-resolution (RSISR) has achieved remarkable progress. However, conventional super-resolution methods usually assume a fixed and known degradation process (e.g., bicubic downsampling), which often leads to significant performance degradation when applied to real-world data with diverse and unknown degradations. To overcome this limitation, we propose DCR-SRNet, a novel Degradation-Contrastive and Wavelet-Guided Network for blind RSISR. The proposed network incorporates three key innovations: First, we design a contrastive degradation representation learning strategy that disentangles degradation priors from scene semantics by pulling together representations of identical degradations across different scenes while pushing apart those of different degradations within the same scene. Second, we introduce a wavelet-guided patch-wise weighted loss module, which employs wavelet decomposition and patch-level discrimination scores to adaptively reweight the pixel-wise loss, thereby enhancing the recovery of edge and texture details. Third, we design an adaptive modulation block (AMB) that injects degradation priors into the reconstruction process through feature- and channel-wise modulation, enabling robust adaptation to diverse degradations. Extensive experiments on three benchmark remote sensing datasets demonstrate that DCR-SRNet significantly outperforms state-of-the-art methods, particularly in preserving structural and textural details.
Zhixiong Huang, Xinying Wang 0005, Shenglan Liu 0001, Lin Feng 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Multidimensional Refinement Graph Convolutional Network With Robust Decouple Loss for Fine-Grained Skeleton-Based Action Recognition
abstract
Graph convolutional networks (GCNs) have been widely used in skeleton-based action recognition. However, existing approaches are limited in fine-grained action recognition due to the similarity of interclass data. Moreover, the noisy data from pose extraction increase the challenge of fine-grained recognition. In this work, we propose a flexible attention block called channel-variable spatial-temporal attention (CVSTA) to enhance the discriminative power of spatial-temporal joints and obtain a more compact intraclass feature distribution. Based on CVSTA, we construct a multidimensional refinement GCN (MDR-GCN) that can improve the discrimination among channel-, joint-, and frame-level features for fine-grained actions. Furthermore, we propose a robust decouple loss (RDL) that significantly boosts the effect of the CVSTA and reduces the impact of noise. The proposed method combining MDR-GCN with RDL outperforms the known state-of-the-art skeleton-based approaches on fine-grained datasets, FineGym99 and FSD-10, and also on the coarse NTU-RGB + D 120 dataset and NTU-RGB + D X-view version. Our code is publicly available at https://github.com/dingyn-Reno/MDR-GCN.
Shenglan Liu 0001, Jinrong Zhang 0001, Kai-Yuan Liu, Gao Huang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Decoupling Spatio-Temporal Network for Fine-Grained Temporal Action Segmentation
abstract
Fine-grained Temporal Action Segmentation (TAS) poses greater challenges compared to general temporal action segmentation. Fine-grained TAS requires distinguishing subtle differences among similar actions and accurately modeling along with spatio-temporal attributes. However, previous methods have largely ignored the exploration of spatio-temporal properties, leading to bad performance on fine-grained datasets. In this paper, we propose a novel Decoupling Spatio-Temporal Network (DSTN) that includes the Action Segmentation Expert (ASE) and the Semantic Information Decision Map (SIDM). The DSTN aims to obtain independent action semantics by decoupling spatio-temporal features and attributes. The ASE consists of Spatial-ASE and Temporal-ASE for frame-by-frame action segmentation, while SIDM is used to align spatio-temporal attribute action labels to ensure the rationality of the results. By leveraging pre-existing spatio-temporal attribute actions, DSTN enables zero-shot TAS and the identification of new actions. Furthermore, recognizing the limitations of current datasets, we construct the FineSkating dataset specifically for Fine-grained TAS. Our model outperforms or competes with the state-of-the-art methods on three challenging datasets.
Haifei Duan, Shenglan Liu 0001, Chenwei Tan, Jirui Tian
ICME2
2024 Two-Step Temporal Divisive Clustering for Unsupervised Action Segmentation
abstract
The goal of unsupervised action segmentation (UAS) is to classify video frames into predefined action classes, which can be considered as a clustering or boundary detection problem. Previous research utilizing bottom-up agglomerative hierarchical clustering methods suffers from over-segmentation or under-segmentation. To address these problems, we propose the Two-step Temporal Divisive Clustering (TTDC) with two components. The first step of TTDC is top-down Temporal Divisive Clustering (TDC), which captures global contexts by comparing the intra-class variances of different classes, and captures local contexts through boundary detection. The second step is the Self-supervised Soft Boundary Regression Network (SS-BRN). SS-BRN is trained by soft pseudo-labels from TDC to refine the boundaries of clusters. In addition, to alleviate the issue of low confidence in pseudo-labels, we use a loss function with soft pseudo-labels. Our empirical evaluations on three benchmarks including 50Salads, Breakfast, and MPII Cooking 2 dataset demonstrate that TTDC outperforms the state-of-the-art methods.
Yule Liu, Zhuben Dong, Shenglan Liu 0001, Wujun Wen, Lin Feng 0001
ICME3
2024 2M-AF: A Strong Multi-Modality Framework For Human Action Quality Assessment with Self-supervised Representation Learning
abstract
Human Action Quality Assessment (AQA) is a prominent area of research in human action analysis. Current mainstream methods only consider the RGB modality which results in limited feature representation and insufficient performance due to the complexity of the AQA task. In this paper, we propose a simple and modular framework called the Two-Modality Assessment Framework (2M-AF), which comprises a skeleton stream, an RGB stream and a regression module. For the skeleton stream, we develop the Self-supervised Mask Encoder Graph Convolution Network (SME-GCN) to achieve representation learning, and further implement score assessment. Additionally, we propose a Preference Fusion Module (PFM) to fuse features, which can effectively avoid the disadvantages of different modalities. Our experimental results demonstrate the superiority of the proposed 2M-AF over current state-of-the-art methods on three publicly available datasets: AQA-7, UNLV-Diving, and MMFS-63.
Shenglan Liu 0001, Jinrong Zhang 0001, Wenyue Chen, Haifei Duan, Bingcheng Dong
ACM Multimedia3
2024 Identifying the potential miRNA biomarkers based on multi-view networks and reinforcement learning for diseases
abstract
MicroRNAs (miRNAs) play important roles in the occurrence and development of diseases. However, it is still challenging to identify the effective miRNA biomarkers for improving the disease diagnosis and prognosis. In this study, we proposed the miRNA data analysis method based on multi-view miRNA networks and reinforcement learning, miRMarker, to define the potential miRNA disease biomarkers. miRMarker constructs the cooperative regulation network and functional similarity network based on the expression data and known miRNA-disease relations, respectively. The cooperative regulation of miRNAs was evaluated by measuring the changes of relative expression. Natural language processing was introduced for calculating the miRNA functional similarity. Then, miRMarker integrates the multi-view miRNA networks and defines the informative miRNA modules through a reinforcement learning strategy. We compared miRMarker with eight efficient data analysis methods on nine transcriptomics datasets to show its superiority in disease sample discrimination. The comparison results suggested that miRMarker outperformed other data analysis methods in receiver operating characteristic analysis. Furthermore, the defined miRNA modules of miRMarker on colorectal cancer data not only show the excellent performance of cancer sample discrimination but also play significant roles in the cancer-related pathway disturbances. The experimental results indicate that miRMarker can build the robust miRNA interaction network by integrating the multi-view networks. Besides, exploring the miRNA interaction network using reinforcement learning favors defining the important miRNA modules. In summary, miRMarker can be a hopeful tool in biomarker identification for human diseases.
Benzhe Su, Xiaohui Lin 0002, Shenglan Liu 0001, Xin Huang 0015
Briefings Bioinform.4
2024 Involving Distinguished Temporal Graph Convolutional Networks for Skeleton-Based Temporal Action Segmentation
abstract
For RGB-based temporal action segmentation (TAS), excellent methods that capture frame-level features have achieved remarkable performance. However, for motion-centered TAS, it is still challenging for existing methods that ignore the extraction of spatial features of joints. In addition, inaccurate action boundaries caused by the frames of similar motion destroy the integrity of the action segments. To alleviate the issues, an end-to-end Involving Distinguished Temporal Graph Convolutional Networks called IDT-GCN is proposed. First, we construct an enhanced spatial graph structure that adaptively captures the similar and differential dependencies between joints in a single topology through learning two independent correlation modeling functions. Then, the proposed Involving Distinguished Graph Convolutional (ID-GC) models the spatial correlations of different actions in a video by using multiple enhanced topologies on the corresponding channels. Furthermore, we design a generic modeling temporal action regression network, termed Temporal Segment Regression (TSR), to extract segmented encoding features and action boundary representations by modeling action sequences. Combining them with label smoothing modules, we develop powerful spatial-temporal graph convolutional networks (IDT-GCN) for fine-grained TAS, which notably outperforms state-of-the-art methods on the MCFS-22 and MCFS-130 datasets. Adding TSR to TCN-based baseline methods achieves competitive performance compared with the state-of-the-art transformer-based methods on RGB-based datasets, i.e., Breakfast and 50Salads. Further experimental results on the action recognition task verify the superiority of the enhanced spatial graph structure over the previous graph convolutional networks.
Kai-Yuan Liu, Shenglan Liu 0001, Lin Feng 0001, Hong Qiao
IEEE Trans. Circuits Syst. Video Technol.3
2024 "Where Does the Devil Lie?": Multimodal Multitask Collaborative Revision Network for Trusted Road Segmentation
abstract
Road segmentation is an essential component of navigation systems. Although recent advancements in road segmentation, the occurrence of failure segmentations remains inevitable. For safety-critical tasks, e.g., navigation, knowing when and where road segmentation fails is crucial. In this paper, we propose a novel trusted road segmentation architecture, namely Multimodal Multitask Collaborative Revision Network (M2CRN), to improve the trust of road segmentation. Our approach incorporates two strategies to predict and rectify segmentation errors. Firstly, a joint learning framework is devised to generate road segmentation results while estimating failure segmentation masks. Secondly, the road segmentation branch is equipped with an Uncertainty-Aware Revision Module (UARM), which eliminates the error in road segmentation. Additionally, we suppress the response of error regions in the road segmentation branch with an innovative design, called Adaptive Soft Error Suppression (ASES). To validate our methods, extensive experiments are conducted on three benchmark road segmentation datasets. The results demonstrate significant performance improvements with a real-time inference speed of 33.3 FPS, reaffirming the soundness of our revision model.
Guoguang Hua, Dalian Zheng, Shishun Tian, Wenbin Zou, Shenglan Liu 0001, Xia Li 0006
IEEE Trans. Multim.5
2024 Hierarchical Neighbors Embedding
abstract
Manifold learning now plays an important role in machine learning and many relevant applications. In spite of the superior performance of manifold learning techniques in dealing with nonlinear data distribution, their performance would drop when facing the problem of data sparsity. It is hard to obtain satisfactory embeddings when sparsely sampled high-dimensional data are mapped into the observation space. To address this issue, in this article, we propose hierarchical neighbors embedding (HNE), which enhances the local connections through hierarchical combination of neighbors. And three different HNE-based implementations are derived by further analyzing the topological connection and reconstruction performance. The experimental results on both the synthetic and real-world datasets illustrate that our HNE-based methods could obtain more faithful embeddings with better topological and geometrical properties. From the view of embedding quality, HNE develops the outstanding advantages in dealing with data of general distributions. Furthermore, comparing with other state-of-the-art manifold learning methods, HNE shows its superiority in dealing with sparsely sampled data and weak-connected manifolds.
Shenglan Liu 0001, Wujun Wen, Hong Qiao
IEEE Trans. Neural Networks Learn. Syst.1
2023 Reducing the Label Bias for Timestamp Supervised Temporal Action Segmentation
abstract
Timestamp supervised temporal action segmentation (TSTAS) is more cost-effective than fully supervised counterparts. However, previous approaches suffer from severe label bias due to over-reliance on sparse timestamp annotations, resulting in unsatisfactory performance. In this paper, we propose the Debiasing-TSTAS (D-TSTAS) framework by exploiting unannotated frames to alleviate this bias from two phases: 1) Initialization. To reduce the dependencies on annotated frames, we propose masked timestamp predictions (MTP) to ensure that initialized model captures more contextual information. 2) Refinement. To overcome the limitation of the expressiveness from sparsely annotated timestamps, we propose a center-oriented timestamp expansion (CTE) approach to progressively expand pseudo-timestamp groups which contain semantic-rich motion representation of action segments. Then, these pseudo-timestamp groups and the model output are used to iteratively generate pseudo-labels for refining the model in a fully supervised setup. We further introduce segmental confidence loss to enable the model to have high confidence predictions within the pseudo-timestamp groups and more accurate action boundaries. Our D-TSTAS outperforms the state-of-the-art TSTAS method as well as achieves competitive results compared with fully supervised approaches on three benchmark datasets.
Shenglan Liu 0001, Chenwei Tan, Zihang Shao
CVPR3
2023 Local Neighbor Propagation Embedding
Wenduo Ma, Hengzhi Yu, Shenglan Liu 0001
PRCV (3)3
2023 Hierarchical Spatial-Temporal Network for Skeleton-Based Temporal Action Segmentation
Chenwei Tan, Talas Fu, Minjie Xu, Shenglan Liu 0001
PRCV (10)6
2023 Skeleton-based action recognition with local dynamic spatial-temporal aggregation
Lianyu Hu 0003, Shenglan Liu 0001, Wei Feng 0005
Expert Syst. Appl.2
2023 Multimodal speech emotion recognition based on multi-scale MFCCs and multi-view attention mechanism
Lin Feng 0001, Luyao Liu 0001, Shenglan Liu 0001, Han-Qing Yang
Multim. Tools Appl.3
2023 SS-INR: Spatial-Spectral Implicit Neural Representation Network for Hyperspectral and Multispectral Image Fusion
abstract
Due to the limitation of imaging equipment, it is difficult to acquire hyperspectral images with high spatial resolution directly. Existing approaches improve the resolution of HSIs by fusing multispectral image (MSI) and hyperspectral image (HSI). However, most of them are only feed-forward. They only learn low- to high-resolution feature mappings without considering the ill-posedness of super-resolution tasks, leading to a large solution space of mapping functions and making it difficult to learn a complete mapping function. Moreover, there is a large resolution difference between HSI and MSI, and some up-sampling operations are inevitably employed in the network. Nevertheless, traditional upsampling methods only represent pixel points in a discrete way, failing to adequately restore the continuous spatial and spectral information. To this end, this paper proposes a spatial-spectral implicit neural representation network for hyperspectral and multispectral image fusion (SS-INR). Inspired by the success of implicit neural representation(INR) in continuum reconstruction, we design spatial-INR and spectral-INR for spatial and spectral resolution reconstruction, respectively. SS-INR contains two processes: forward fusion (FF) and back-projection fusion(BPF). In the FF process, the input HSI is first spatially upsampled with Spatial-INR to overcome spatial resolution differences while performing initial fusion with MSI. In the BPF process, we explore the spatial and spectral degradation processes and use them as prior knowledge for error correction. Extensive experiments on five public hyperspectral datasets demonstrate the effectiveness of SS-INR, and SS-INR achieves competitive results compared with existing state-of-the-art fusion methods. The source code for SS-INR will be released at https://github.com/wxy11-27/SS-INR.
Xinying Wang 0005, Cheng Cheng 0013, Shenglan Liu 0001, Ruoxi Song, Xiang-Hai Wang 0001, Lin Feng 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Double Attention Network Based on Sparse Sampling
abstract
Locating action segments in long untrimmed videos is a sub-task of video understanding, which more and more scholars pay attention to. Boundary ambiguity and over-segmentation errors are two difficult problems. To handle them, we propose a network called Double Attention Network based on Sparse Sampling (DASS) on the basis of MS-TCN series. First, we design a Seq2Seq Convolution Sampling Network (SCSN) to reduce feature redundancy, which also works on over-fitting. Second, we devise a Global Temporal Attention Module (GTAM) to help predict action boundaries and improve the effect of post-processing from a global perspective. Third, we propose Local Temporal Attention Module (LTAM), which both casts attention to local frames and complements details lost in high dilated layers. We perform experiments on three challenging datasets: 50Salads, GTEA and Breakfast and prove our model is state-of-the-art.
Zhuben Dong, Conghui Hao, Shenglan Liu 0001
ICME7
2022 Vision Shared and Representation Isolated Network for Person Search
abstract
Person search is a widely-concerned computer vision task that aims to jointly solve the problems of pedestrian detection and person re-identification in panoramic scenes. However, the pedestrian detection focuses on the consistency of pedestrians, while the person re-identification attempts to extract the discriminative features of pedestrians. The inevitable conflict greatly restricts the researches on the one-stage person search methods. To address this issue, we propose a Vision Shared and Representation Isolated (VSRI) network to decouple the two conflicted subtasks simultaneously, through which two independent representations are constructed for the two subtasks. To enhance the discrimination of the re-ID representation, a Multi-Level Feature Fusion (MLFF) module is proposed. The MLFF adopts the Spatial Pyramid Feature Fusion (SPFF) module to obtain diverse features from the stem network. Moreover, the multi-head self-attention mechanism is employed to construct a Multi-head Attention Driven Extraction (MADE) module and the cascaded convolution unit is adopted to devise a Feature Decomposition and Cascaded Integration (FDCI) module, which facilitates the MLFF to obtain more discriminative representations of the pedestrians. The proposed method outperforms the state-of-the-art methods on the mainstream datasets.
Yang Liu 0066, Yingping Li, Chengyu Kong, Yuqiu Kong, Shenglan Liu 0001
IJCAI5
2022 Background Suppressed and Motion Enhanced Network for Weakly Supervised Video Anomaly Detection
Yang Liu 0066, Wanxiao Yang, Hangyou Yu, Lin Feng 0001, Yuqiu Kong, Shenglan Liu 0001
PRCV (3)6
2022 Multiview nonlinear discriminant structure learning for emotion recognition
Shuai Guo 0002, Li Song 0001, Rong Xie 0004, Lin Li 0062, Shenglan Liu 0001
Knowl. Based Syst.5
2022 Spatial Focus Attention for Fine-Grained Skeleton-Based Action Tasks
abstract
Dynamic skeletal data has been widely studied for human action tasks due to its high-level semantic information and less data than RGB features. However, attention-based previous methods fail to focus on the local grouped joint dependence of the human body, which is vital to distinguishing various actions in fine-grained tasks, such as skeletal action segmentation and recognition. This work proposes spatial focus attention for the fine-grained skeleton-based action tasks. Specifically, we decouple the attention map to enhance the grouped joint dependence adaptively by the decouple probability. To further focus on local grouped dependence, the tree structural attention maps can be built by hierarchical decoupling and guide the model to focus on complementary local dependence in the different leaf nodes. Our proposed approach achieves state-of-the-art performance on fine-grained skeleton-based human action segmentation tasks (MCFS-22) and recognition tasks (FSD-10). Besides, on the coarse-grained dataset (NTU-60), the proposed spatial focus attention also achieves outstanding performance.
Yuanfeng Xu, Shenglan Liu 0001
IEEE Signal Process. Lett.5
2021 Temporal Segmentation of Fine-gained Semantic Action: A Motion-Centered Figure Skating Dataset
abstract
Temporal Action Segmentation (TAS) has achieved great success in many fields such as exercise rehabilitation, movie editing, etc. Currently, task-driven TAS is a central topic in human action analysis. However, motion-centered TAS, as an important topic, is little researched due to unavailable datasets. In order to explore more models and practical applications of motion-centered TAS, we introduce a Motion-Centered Figure Skating (MCFS) dataset in this paper. Compared with existing temporal action segmentation datasets, the MCFS dataset is fine-grained semantic, specialized and motion-centered. Besides, RGB-based and Skeleton-based features are provided in the MCFS dataset. Experimental results show that existing state-of-the-art methods are difficult to achieve excellent segmentation results (including accuracy, edit and F1 score) in the MCFS dataset. This indicates that MCFS is a challenging dataset for motion-centered TAS. The latest dataset can be downloaded at https://shenglanliu.github.io/mcfs-dataset/.
Shenglan Liu 0001, Aibin Zhang, Zhuben Dong, Renhao Zhang
AAAI1
2021 Adaptive Graph Convolutional Network with Prior Knowledge for Action Recognition
Guihong Lao, Lianyu Hu 0004, Shenglan Liu 0001, Zhuben Dong, Wujun Wen
ICONIP (2)3
2021 Spatial-Temporal Attention Network with Multi-similarity Loss for Fine-Grained Skeleton-Based Action Recognition
Shenglan Liu 0001, Hao Liu 0029, Jinjing Zhao, Lin Feng 0001, Guihong Lao, Guangzhe Li
ICONIP (2)2
2021 A Lightweight Multidimensional Self-attention Network for Fine-Grained Action Recognition
Hao Liu 0029, Shenglan Liu 0001, Lin Feng 0001, Lianyu Hu 0004, Heyu Fu
ICONIP (2)2
2021 Weighted P-Rank: a Weighted Article Ranking Algorithm Based on a Heterogeneous Scholarly Network
Shenglan Liu 0001, Lin Feng 0001, Ning Cai 0002
ICONIP (1)2
2021 Efficient Two-Step Networks for Temporal Action Segmentation
Zhuben Dong, Lin Feng 0001, Lianyu Hu 0004, Shenglan Liu 0001
Neurocomputing9
2021 Three Degree Binary Graph and Shortest Edge Clustering for re-ranking in multi-feature image retrieval
Guihong Lao, Shenglan Liu 0001, Chenwei Tan, Guangzhe Li, Lin Feng 0001
J. Vis. Commun. Image Represent.2
2021 Bottom-up broadcast neural network for music genre classification
Caifeng Liu, Lin Feng 0001, Guochao Liu, Huibing Wang, Shenglan Liu 0001
Multim. Tools Appl.5
2021 Social Neighborhood Graph and Multigraph Fusion Ranking for Multifeature Image Retrieval
abstract
A single feature is hard to describe the content of images from an overall perspective, which limits the retrieval performances of single-feature-based methods in image retrieval tasks. To fully describe the properties of images and improve the retrieval performances, multifeature fusion ranking-based methods are proposed. However, the effectiveness of multifeature fusion in image retrieval has not been theoretically explained. This article gives a theoretical proof to illustrate the role of independent features in improving the retrieval results. Based on the theoretical proof, the original ranking list generated with a single feature greatly influences the performances of multifeature fusion ranking. Inspired by the principle of three degrees of influence in social networks, this article proposes a reranking method named k -nearest neighbors' neighbors' neighbors' graph (N3G) to improve the original ranking list by a single feature. Furthermore, a multigraph fusion ranking (MFR) method motivated by the group relation theory in social networks for multifeature ranking is also proposed, which considers the correlations of all images in multiple neighborhood graphs. Evaluation experiments conducted on several representative data sets (e.g., UK-bench, Holiday, Corel-10K, and Cifar-10) validate that N3G and MFR outperform the other state-of-the-art methods.
Shenglan Liu 0001, Muxin Sun, Lin Feng 0001, Hong Qiao, Shuyuan Chen, Yang Liu 0066
IEEE Trans. Neural Networks Learn. Syst.1
2020 Skeleton-Based Action Recognition with Dense Spatial Temporal Graph Network
Lin Feng 0001, Zhenning Lu, Shenglan Liu 0001, Yang Liu 0066, Lianyu Hu 0004
ICONIP (5)3
2020 A Discriminative STGCN for Skeleton Oriented Action Recognition
Lin Feng 0001, Yang Liu 0066, Qianxin Huang, Shenglan Liu 0001, Yingping Li
ICONIP (5)5
2020 Bionic Vision Descriptor for Image Retrieval
Guangzhe Li, Shenglan Liu 0001, Lin Feng 0001
ICONIP (1)2
2020 Self-adaption neighborhood density clustering method for mixed data stream with concept drift
Shuliang Xu, Lin Feng 0001, Shenglan Liu 0001, Hong Qiao
Eng. Appl. Artif. Intell.3
2020 An incrementally cascaded broad learning framework to facial landmark tracking
Caifeng Liu, Lin Feng 0001, Shuai Guo 0002, Huibing Wang, Shenglan Liu 0001, Hong Qiao
Neurocomputing5
2020 FSD-10: A fine-grained classification dataset for figure skating
Shenglan Liu 0001, Gao Huang 0001, Hong Qiao, Lianyu Hu 0004, Aibin Zhang, Yang Liu 0066
Neurocomputing1
2020 Fuzzy granularity neighborhood extreme clustering
Shuliang Xu, Shenglan Liu 0001, Lin Feng 0001
Neurocomputing2
2020 Deep attention based music genre classification
Sen Luo, Shenglan Liu 0001, Hong Qiao, Yang Liu 0066, Lin Feng 0001
Neurocomputing3
2020 Multi-view laplacian eigenmaps based on bag-of-neighbors for RGB-D human emotion recognition
Shenglan Liu 0001, Shuai Guo 0002, Wei Wang 0036, Hong Qiao, Wenbo Luo
Inf. Sci.1
2020 Manifold graph embedding with structure information propagation for community discovery
Shuliang Xu, Shenglan Liu 0001, Lin Feng 0001
Knowl. Based Syst.2
2020 Multi-feature weighting neighborhood density clustering
Shuliang Xu, Lin Feng 0001, Shenglan Liu 0001, Hong Qiao
Neural Comput. Appl.3
2019 Rough extreme learning machine: A new classification method based on uncertainty measure
Lin Feng 0001, Shuliang Xu, Shenglan Liu 0001, Hong Qiao
Neurocomputing4
2019 Multi-view laplacian least squares for human emotion recognition
Shuai Guo 0002, Lin Feng 0001, Zhanbo Feng, Yi-Hao Li, Shenglan Liu 0001, Hong Qiao
Neurocomputing6
2018 Quasi-curvature Local Linear Projection and Extreme Learning Machine for nonlinear dimensionality reduction
Shenglan Liu 0001, Jun Wu 0008, Lin Feng 0001, Sen Luo, Deqin Yan
Neurocomputing1
2018 Perceptual uniform descriptor and ranking on manifold for image retrieval
Shenglan Liu 0001, Jun Wu 0008, Lin Feng 0001, Hong Qiao, Yang Liu 0066, Wenbo Luo, Wei Wang 0036
Inf. Sci.1
2018 Global similarity preserving hashing
Yang Liu 0066, Lin Feng 0001, Shenglan Liu 0001, Muxin Sun
Soft Comput.3
2018 Manifold Warp Segmentation of Human Action
abstract
Human action segmentation is important for human action analysis, which is a highly active research area. Most segmentation methods are based on clustering or numerical descriptors, which are only related to data, and consider no relationship between the data and physical characteristics of human actions. Physical characteristics of human motions are those that can be directly perceived by human beings, such as speed, acceleration, continuity, and so on, which are quite helpful in detecting human motion segment points. We propose a new physical-based descriptor of human action by curvature sequence warp space alignment (CSWSA) approach for sequence segmentation in this paper. Furthermore, time series-warp metric curvature segmentation method is constructed by the proposed descriptor and CSWSA. In our segmentation method, descriptor can express the changes of human actions, and CSWSA is an auxiliary method to give suggestions for segmentation. The experimental results show that our segmentation method is effective in both CMU human motion and video-based data sets.
Shenglan Liu 0001, Lin Feng 0001, Yang Liu 0066, Hong Qiao, Jun Wu 0008, Wei Wang 0036
IEEE Trans. Neural Networks Learn. Syst.1
2017 A Transferable Framework: Classification and Visualization of MOOC Discussion Threads
Lin Feng 0001, Guochao Liu, Sen Luo, Shenglan Liu 0001
ICONIP (4)4
2017 Image retrieval framework based on texton uniform descriptor and modified manifold ranking
Jun Wu 0008, Lin Feng 0001, Shenglan Liu 0001, Muxin Sun
J. Vis. Commun. Image Represent.3
2017 Multi-view spectral clustering via robust local subspace learning
Lin Feng 0001, Yang Liu 0066, Shenglan Liu 0001
Soft Comput.4
2016 Extend semi-supervised ELM and a frame work
Shenglan Liu 0001, Lin Feng 0001, Huibing Wang, Xiao Yao 0001
Neural Comput. Appl.1
2015 A novel CBIR system with WLLTSA and ULRGA
Lin Feng 0001, Shenglan Liu 0001, Xiao Yao 0001, Qiao Hong
Neurocomputing2
2015 Locality Structured Sparsity Preserving Embedding
abstract
In recent years, the theory of sparse representation (SR) has been widely exploited in sparse subspace learning (SSL). Among all these methods, SR is a parameter-free global algorithm in nature which is mostly utilized to construct the correlations between samples to avoid some negative effects incurred by k-nearest neighbor (KNN) or some other methods. However, these SSL algorithms always lack obvious discrimination because of the ignorance of samples distribution. Meanwhile, some incorrect correlations are taken into consideration owing to the global feature of SR. To solve these two problems, a new SSL algorithm called locality structured sparsity preserving embedding (LSPE) is proposed in this paper. We add the local structured information to SR and construct correlations between samples. However, LSPE is an unsupervised method which wastes all label information. Therefore, LSPE is extended to semi-supervised LSPE (SLSPE) in this paper. SLSPE not only makes good use of the label information but also enhances the discriminative power of LSPE. Extensive experiments have been performed on three image datasets (CMU, COIL20, ORL) and two UCI datasets (Glass, Segment) to prove the efficiency of the LSPE and SLSPE.
Lin Feng 0001, Huibing Wang, Shenglan Liu 0001
Int. J. Pattern Recognit. Artif. Intell.3
2015 Global Correlation Descriptor: A novel image representation for image retrieval
Lin Feng 0001, Jun Wu 0008, Shenglan Liu 0001
J. Vis. Commun. Image Represent.3
2015 Scatter Balance: An Angle-Based Supervised Dimensionality Reduction
abstract
Subspace selection is widely applied in data classification, clustering, and visualization. The samples projected into subspace can be processed efficiently. In this paper, we research the linear discriminant analysis (LDA) and maximum margin criterion (MMC) algorithms intensively and analyze the effects of scatters to subspace selection. Meanwhile, we point out the boundaries of scatters in LDA and MMC algorithms to illustrate the differences and similarities of subspace selection in different circumstances. Besides, the effects of outlier classes on subspace selection are also analyzed. According to the above analysis, we propose a new subspace selection method called angle linear discriminant embedding (ALDE) on the basis of angle measurement. ALDE utilizes the cosine of the angle to get new within-class and between-class scatter matrices and avoids the small sample size problem simultaneously. To deal with high-dimensional data, we extend ALDE to a two-stage ALDE (TS-ALDE). The synthetic data experiments indicate that ALDE can balance the within-class and between-class scatters and be robust to outlier classes. The experimental results based on UCI machine-learning repository and image databases show that TS-ALDE has a lower time complexity than ALDE while processing high-dimensional data.
Shenglan Liu 0001, Lin Feng 0001, Hong Qiao
IEEE Trans. Neural Networks Learn. Syst.1
2014 Robust activation function and its application: Semi-supervised kernel extreme learning method
Shenglan Liu 0001, Lin Feng 0001, Xiao Yao 0001, Huibing Wang
Neurocomputing1
2013 Maximal Similarity Embedding
Lin Feng 0001, Shenglan Liu 0001, Zhen Yu Wu, Bo Jin 0001
Neurocomputing2