Yang Liu 0066

dblp:51/3710-66 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0001-8257-1429ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 HFP-SAM: Hierarchical Frequency Prompted SAM for Efficient Marine Animal Segmentation
abstract
Marine Animal Segmentation (MAS) aims at identifying and segmenting marine animals from complex marine environments. Most of previous deep learning-based MAS methods struggle with the long-distance modeling issue. Recently, Segment Anything Model (SAM) has gained popularity in general image segmentation. However, it lacks of perceiving fine-grained details and frequency information. To this end, we propose a novel learning framework, named Hierarchical Frequency Prompted SAM (HFP-SAM) for high-performance MAS. First, we design a Frequency Guided Adapter (FGA) to efficiently inject marine scene information into the frozen SAM backbone through frequency domain prior masks. Additionally, we introduce a Frequency-aware Point Selection (FPS) to generate highlighted regions through frequency analysis. These regions are combined with the coarse predictions of SAM to generate point prompts and integrate into SAM's decoder for fine predictions. Finally, to obtain comprehensive segmentation masks, we introduce a Full-View Mamba (FVM) to efficiently extract spatial and channel contextual information with linear computational complexity. Extensive experiments on four public datasets demonstrate the superior performance of our approach. We will make our code publicly available upon the acceptance.
Tianyu Yan, Yang Liu 0066, Tongdan Tang, Yili Ma, Long Lv, Feng Tian 0001, Weibing Sun, Huchuan Lu
IEEE Trans. Image Process.4
2025 Equilibrium Policy Generalization: A Reinforcement Learning Framework for Cross-Graph Zero-Shot Generalization in Pursuit-Evasion Games
abstract
Equilibrium learning in adversarial games is an important topic widely examined in the fields of game theory and reinforcement learning (RL). Pursuit-evasion game (PEG), as an important class of real-world games from the fields of robotics and security, requires exponential time to be accurately solved. When the underlying graph structure varies, even the state-of-the-art RL methods require recomputation or at least fine-tuning, which can be time-consuming and impair real-time applicability. This paper proposes an Equilibrium Policy Generalization (EPG) framework to effectively learn a generalized policy with robust cross-graph zero-shot performance. In the context of PEGs, our framework is generally applicable to both pursuer and evader sides in both no-exit and multi-exit scenarios. These two generalizability properties, to our knowledge, are the first to appear in this domain. The core idea of the EPG framework is to train an RL policy across different graph structures against the equilibrium policy for each single graph. To construct an equilibrium oracle for single-graph policies, we present a dynamic programming (DP) algorithm that provably generates pure-strategy Nash equilibrium with near-optimal time complexity. To guarantee scalability with respect to pursuer number, we further extend DP and RL by designing a grouping mechanism and a sequence model for joint policy decomposition, respectively. Experimental results show that, using equilibrium guidance and a distance feature proposed for cross-graph PEG training, the EPG framework guarantees desirable zero-shot performance in various unseen real-world graphs. Besides, when trained under an equilibrium heuristic proposed for the graphs with exits, our generalized pursuer policy can even match the performance of the fine-tuned policies from the state-of-the-art PEG methods.
Runyu Lu, Peng Zhang 0127, Ruochuan Shi, Yuanheng Zhu, Dongbin Zhao, Yang Liu 0066, Dong Wang 0004, Cesare Alippi
NeurIPS6
2025 MambaVT: Spatio-Temporal Contextual Modeling for Robust RGB-T Tracking
Simiao Lai, Chang Liu 0071, Jiawen Zhu 0003, Ben Kang, Yang Liu 0066, Dong Wang 0004, Huchuan Lu
IEEE Trans. Circuits Syst. Video Technol.5
2025 EMTrack: Efficient Multimodal Object Tracking
abstract
Multi-modal object tracking has received increasing attention, given the limitations the representation ability in certain challenging scenarios of single RGB modality. Recent prompt tuning techniques enable multimodal tracking to effectively inherit knowledge from foundation models trained with a large amount of RGB tracking data and achieve parameter-efficient training. However, few works focus on the efficient inference of multimodal tracking handling multiple RGB-X (RGB-Thermal, RGB-Depth, RGB-Event, etc.) tracking tasks simultaneously, especially on resource-limited devices such as CPU. In this work, we propose an efficient multimodal tracker named EMTrack. EMTrack follows a concise and unified multimodal tracking framework with simple knowledge distillation. RGB modality and auxiliary modality are added after patch-embedding layer for fusion, reducing the computational complexity of multimodal tracking compared with that of single modality. Before fusion operation, we introduce a modal-specific spatial modulation module to exploit and realize adaptive spatial adjustment of different modality features. Multiple modal-specific experts are adopted to capture specific information for different RGB-X tracking tasks, which assists in handling such tasks in a unified model with joint training. EMTrack achieves competitive performance on various RGB-X tracking benchmarks while reaching a good balance of performance and speed on different platforms. Especially on an Intel Core i9-10850K CPU device, EMTrack achieves 29.1 fps, a real-time speed, with only 2.0G MAC computation.
Chang Liu 0071, Ziqi Guan, Simiao Lai, Yang Liu 0066, Huchuan Lu, Dong Wang 0004
IEEE Trans. Circuits Syst. Video Technol.4
2025 Enhancing the Two-Stream Framework for Efficient Visual Tracking
abstract
Practical deployments, especially on resource-limited edge devices, necessitate high speed for visual object trackers. To meet this demand, we introduce a new efficient tracker with a Two-Stream architecture, named ToS. While the recent one-stream tracking framework, employing a unified backbone for simultaneous processing of both the template and search region, has demonstrated exceptional efficacy, we find the conventional two-stream tracking framework, which employs two separate backbones for the template and search region, offers inherent advantages. The two-stream tracking framework is more compatible with advanced lightweight backbones and can efficiently utilize benefits from large templates. We demonstrate that the two-stream setup can exceed the one-stream tracking model in both speed and accuracy through strategic designs. Our methodology rejuvenates the two-stream tracking paradigm with lightweight pre-trained backbones and the proposed three efficient strategies: 1) A feature-aggregation module that improves the representation capability of the backbone, 2) A channel-wise approach for feature fusion, presenting a more effective and lighter alternative to spatial concatenation techniques, and 3) An expanded template strategy to boost tracking accuracy with negligible additional computational cost. Extensive evaluations across multiple tracking benchmarks demonstrate that the proposed method sets a new state-of-the-art performance in efficient visual tracking.
ChengAo Zong, Xin Chen 0032, Jie Zhao 0014, Yang Liu 0066, Huchuan Lu, Dong Wang 0004
IEEE Trans. Image Process.4
2022 Scale Adaptive Fusion Network for RGB-D Salient Object Detection
Yuqiu Kong, Yushuo Zheng, Cuili Yao, Yang Liu 0066
ACCV (3)4
2022 Vision Shared and Representation Isolated Network for Person Search
abstract
Person search is a widely-concerned computer vision task that aims to jointly solve the problems of pedestrian detection and person re-identification in panoramic scenes. However, the pedestrian detection focuses on the consistency of pedestrians, while the person re-identification attempts to extract the discriminative features of pedestrians. The inevitable conflict greatly restricts the researches on the one-stage person search methods. To address this issue, we propose a Vision Shared and Representation Isolated (VSRI) network to decouple the two conflicted subtasks simultaneously, through which two independent representations are constructed for the two subtasks. To enhance the discrimination of the re-ID representation, a Multi-Level Feature Fusion (MLFF) module is proposed. The MLFF adopts the Spatial Pyramid Feature Fusion (SPFF) module to obtain diverse features from the stem network. Moreover, the multi-head self-attention mechanism is employed to construct a Multi-head Attention Driven Extraction (MADE) module and the cascaded convolution unit is adopted to devise a Feature Decomposition and Cascaded Integration (FDCI) module, which facilitates the MLFF to obtain more discriminative representations of the pedestrians. The proposed method outperforms the state-of-the-art methods on the mainstream datasets.
Yang Liu 0066, Yingping Li, Chengyu Kong, Yuqiu Kong, Shenglan Liu 0001
IJCAI1
2022 Background Suppressed and Motion Enhanced Network for Weakly Supervised Video Anomaly Detection
Yang Liu 0066, Wanxiao Yang, Hangyou Yu, Lin Feng 0001, Yuqiu Kong, Shenglan Liu 0001
PRCV (3)1
2021 Social Neighborhood Graph and Multigraph Fusion Ranking for Multifeature Image Retrieval
abstract
A single feature is hard to describe the content of images from an overall perspective, which limits the retrieval performances of single-feature-based methods in image retrieval tasks. To fully describe the properties of images and improve the retrieval performances, multifeature fusion ranking-based methods are proposed. However, the effectiveness of multifeature fusion in image retrieval has not been theoretically explained. This article gives a theoretical proof to illustrate the role of independent features in improving the retrieval results. Based on the theoretical proof, the original ranking list generated with a single feature greatly influences the performances of multifeature fusion ranking. Inspired by the principle of three degrees of influence in social networks, this article proposes a reranking method named k -nearest neighbors' neighbors' neighbors' graph (N3G) to improve the original ranking list by a single feature. Furthermore, a multigraph fusion ranking (MFR) method motivated by the group relation theory in social networks for multifeature ranking is also proposed, which considers the correlations of all images in multiple neighborhood graphs. Evaluation experiments conducted on several representative data sets (e.g., UK-bench, Holiday, Corel-10K, and Cifar-10) validate that N3G and MFR outperform the other state-of-the-art methods.
Shenglan Liu 0001, Muxin Sun, Lin Feng 0001, Hong Qiao, Shuyuan Chen, Yang Liu 0066
IEEE Trans. Neural Networks Learn. Syst.6
2020 Skeleton-Based Action Recognition with Dense Spatial Temporal Graph Network
Lin Feng 0001, Zhenning Lu, Shenglan Liu 0001, Yang Liu 0066, Lianyu Hu 0004
ICONIP (5)5
2020 A Discriminative STGCN for Skeleton Oriented Action Recognition
Lin Feng 0001, Yang Liu 0066, Qianxin Huang, Shenglan Liu 0001, Yingping Li
ICONIP (5)3
2020 FSD-10: A fine-grained classification dataset for figure skating
Shenglan Liu 0001, Gao Huang 0001, Hong Qiao, Lianyu Hu 0004, Aibin Zhang, Yang Liu 0066
Neurocomputing8
2020 Deep attention based music genre classification
Sen Luo, Shenglan Liu 0001, Hong Qiao, Yang Liu 0066, Lin Feng 0001
Neurocomputing5
2018 Perceptual uniform descriptor and ranking on manifold for image retrieval
Shenglan Liu 0001, Jun Wu 0008, Lin Feng 0001, Hong Qiao, Yang Liu 0066, Wenbo Luo, Wei Wang 0036
Inf. Sci.5
2018 Global similarity preserving hashing
Yang Liu 0066, Lin Feng 0001, Shenglan Liu 0001, Muxin Sun
Soft Comput.1
2018 Manifold Warp Segmentation of Human Action
abstract
Human action segmentation is important for human action analysis, which is a highly active research area. Most segmentation methods are based on clustering or numerical descriptors, which are only related to data, and consider no relationship between the data and physical characteristics of human actions. Physical characteristics of human motions are those that can be directly perceived by human beings, such as speed, acceleration, continuity, and so on, which are quite helpful in detecting human motion segment points. We propose a new physical-based descriptor of human action by curvature sequence warp space alignment (CSWSA) approach for sequence segmentation in this paper. Furthermore, time series-warp metric curvature segmentation method is constructed by the proposed descriptor and CSWSA. In our segmentation method, descriptor can express the changes of human actions, and CSWSA is an auxiliary method to give suggestions for segmentation. The experimental results show that our segmentation method is effective in both CMU human motion and video-based data sets.
Shenglan Liu 0001, Lin Feng 0001, Yang Liu 0066, Hong Qiao, Jun Wu 0008, Wei Wang 0036
IEEE Trans. Neural Networks Learn. Syst.3
2017 Multi-view spectral clustering via robust local subspace learning
Lin Feng 0001, Yang Liu 0066, Shenglan Liu 0001
Soft Comput.3
2016 Metric learning with geometric mean for similarities measurement
Huibing Wang, Lin Feng 0001, Yang Liu 0066
Soft Comput.3
2016 Semantic Discriminative Metric Learning for Image Similarity Measurement
abstract
Along with the arrival of multimedia time, multimedia data has replaced textual data to transfer information in various fields. As an important form of multimedia data, images have been widely utilized by many applications, such as face recognition and image classification. Therefore, how to accurately annotate each image from a large set of images is of vital importance but challenging. To perform these tasks well, it is crucial to extract suitable features to character the visual contents of images and learn an appropriate distance metric to measure similarities between all images. Unfortunately, existing feature operators, such as histogram of gradient, local binary pattern, and color histogram, care more about the visual character of images and lack the ability to distinguish semantic information. Similarities between those features cannot reflect the real category correlations due to the well-known semantic gap. In order to solve this problem, this paper proposes a regularized distance metric framework called semantic discriminative metric learning (SDML). SDML combines geometric mean with normalized divergences and separates images from different classes simultaneously. The learned distance metric can treat all images from different classes equally. And distinctions between similar classes with entirely different semantic contents are emphasized by SDML. This procedure ensures the consistency between dissimilarities and semantic distinctions and avoids inaccuracy similarities incurred by unbalanced locations of samples. Various experiments on benchmark image datasets show the excellent performance of the novel method.
Huibing Wang, Lin Feng 0001, Jing Zhang 0028, Yang Liu 0066
IEEE Trans. Multim.4