Lei Zhou 0008

dblp:72/5749-8 · DBLP profile ↗
← Back
23ranked-venue papers
9as first author
15since 2021 · last 2026
0000-0003-0860-0724ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SCRNet: A spatial-channel reconstruction network for multi-task pineapple detection and a novel pineapple dataset
Songwei Lian, Zimin (Max) Yang, Lei Zhou 0008, Cong Lin 0004
Expert Syst. Appl.4
2025 ROS-SAM: High-Quality Interactive Segmentation for Remote Sensing Moving Object
abstract
The availability of large-scale remote sensing video data underscores the importance of high-quality interactive segmentation. However, challenges such as small object sizes, ambiguous features, and limited generalization make it difficult for current methods to achieve this goal. In this work, we propose ROS-SAM, a method designed to achieve high-quality interactive segmentation while preserving generalization across diverse remote sensing data. The ROS-SAM is built upon three key innovations: 1) LoRA-based fine-tuning, which enables efficient domain adaptation while maintaining SAM’s generalization ability, 2) Enhancement of network deep layers to improve the discriminability of extracted features, thereby reducing misclassifications, and 3) Integration of global context with local boundary details in the mask decoder to generate high-quality segmentation masks. Additionally, we redesign the data pipeline to ensure the model learns to better handle objects at varying scales during training while focusing on high-quality predictions during inference. Experiments on remote sensing video datasets show that the data pipeline boosts the IoU by 6%, while ROS-SAM increases the IoU by 13%. Finally, when evaluated on existing remote sensing object tracking datasets, ROS-SAM demonstrates impressive zero-shot capabilities, generating masks that closely resemble manual annotations. These results confirm ROS-SAM as a powerful tool for fine-grained segmentation in remote sensing applications. Code is available at: https://github.com/ShanZard/ROS-SAM.
Yang Liu 0357, Lei Zhou 0008
CVPR3
2025 Learning Visual Proxy for Compositional Zero-Shot Learning
Yang Liu 0357, Chenchen Jing, Lei Zhou 0008, Wenjun Wang 0002
ICCV5
2025 Jpeg stereo image lossy recompression with mutual information enhancement
Junwei Zhou 0002, Benyi Zhang, Shengping Wu, Lei Zhou 0008, Yanchao Yang 0002, Jianwen Xiang
Multim. Syst.4
2024 Lightweight Autoencoder with Hierarchical Priors for Learned Image Compression
abstract
Image compression has become an important task for reducing storage and transmission costs. However, recent models for learned image compression have been developed to increase the network’s number of layers and channels to achieve better visual effects. This resulted in higher computing and memory resources, making deploying the model on compute-constrained platforms such as wireless devices impractical. In this paper, we propose a lightweight autoencoder with hierarchical priors. The lightweight autoencoder reduces the model’s parameter size and calculation amount based on ensuring high fidelity and low bit rates of the image. Simulation results indicate the proposed model yields a smaller size: the parameters are reduced by 81.66%, and the calculation amount is reduced by 94.7% over the benchmark. Besides, the proposed model results in a speed improvement of 200 times. At the same time, our model achieves nearly the same performance as the baseline on MS-SSIM and LPIPS distortion metrics.
Junwei Zhou 0002, Lei Zhou 0008, Yanchao Yang 0002, Jianwen Xiang
HPCC4
2024 SMTCNN - A global spatio-temporal texture convolutional neural network for 3D dynamic texture recognition
Liangliang Wang 0007, Lei Zhou 0008, Peidong Liang, Ke Wang 0028, Lianzheng Ge
Image Vis. Comput.2
2024 Learning adversarial semantic embeddings for zero-shot recognition in open worlds
Guansong Pang, Xiao Bai 0001, Lei Zhou 0008, Xin Ning 0001
Pattern Recognit.5
2023 Generalized Zero-Shot Learning via Implicit Attribute Composition
abstract
Zero-shot learning (ZSL) is an important but challenging task in computer vision that aims to identify unseen classes without matching training samples. Current cutting-edge ZSL methods based on locality focus on acquiring the explicit locality of distinguishing characteristics, which could face a lack of adequate supervision at the class attribute level. This paper introduces a novel approach called IAC, which aims to learn Implicit Attribute Composition for ZSL. This method is more comprehensive compared to attribute localization that solely focuses on class-level attribute supervision. IAC utilizes subspace representations that efficiently capture the inherent structure of high-dimensional image features. Then, we learn implicit attribute composition through subspace representation learning. The superiority of the proposed IAC compared to the state-of-the-art is demonstrated through sufficient experiments conducted on three commonly used ZSL datasets, CUB, SUN, and AwA2.
Lei Zhou 0008, Yang Liu 0357, Qiang Li 0060
SMC1
2023 Information bottleneck and selective noise supervision for zero-shot learning
Lei Zhou 0008, Yang Liu 0357, Pengcheng Zhang 0003, Xiao Bai 0001, Lin Gu 0003, Jun Zhou 0001, Yazhou Yao, Tatsuya Harada, Edwin R. Hancock
Mach. Learn.1
2023 Attribute subspaces for zero-shot learning
Lei Zhou 0008, Yang Liu 0357, Xiao Bai 0001, Na Li 0014, Xiaohan Yu 0001, Jun Zhou 0001, Edwin R. Hancock
Pattern Recognit.1
2023 Aerial image recognition in discriminative bi-transformer
Yichen Zhao, Yaxiong Chen, Xiongbo Lu, Lei Zhou 0008, Shengwu Xiong 0001
Signal Process.4
2022 Where to Focus: Investigating Hierarchical Attention Relationship for Fine-Grained Visual Classification
Yang Liu 0357, Lei Zhou 0008, Pengcheng Zhang 0003, Xiao Bai 0001, Lin Gu 0003, Xiaohan Yu 0001, Jun Zhou 0001, Edwin R. Hancock
ECCV (24)2
2022 Learning Prototype via Placeholder for Zero-shot Recognition
abstract
Zero-shot learning (ZSL) aims to recognize unseen classes by exploiting semantic descriptions shared between seen classes and unseen classes. Current methods show that it is effective to learn visual-semantic alignment by projecting semantic embeddings into the visual space as class prototypes. However, such a projection function is only concerned with seen classes. When applied to unseen classes, the prototypes often perform suboptimally due to domain shift. In this paper, we propose to learn prototypes via placeholders, termed LPL, to eliminate the domain shift between seen and unseen classes. Specifically, we combine seen classes to hallucinate new classes which play as placeholders of the unseen classes in the visual and semantic space. Placed between seen classes, the placeholders encourage prototypes of seen classes to be highly dispersed. And more space is spared for the insertion of well-separated unseen ones. Empirically, well-separated prototypes help counteract visual-semantic misalignment caused by domain shift. Furthermore, we exploit a novel semantic-oriented fine-tuning method to guarantee the semantic reliability of placeholders. Extensive experiments on five benchmark datasets demonstrate the significant performance gain of LPL over the state-of-the-art methods.
Zaiquan Yang, Yang Liu 0357, Wenjia Xu, Lei Zhou 0008
IJCAI5
2021 Goal-Oriented Gaze Estimation for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) aims to recognize novel classes by transferring semantic knowledge from seen classes to unseen classes. Since semantic knowledge is built on attributes shared between different classes, which are highly local, strong prior for localization of object attribute is beneficial for visual-semantic embedding. Interestingly, when recognizing unseen images, human would also automatically gaze at regions with certain semantic clue. Therefore, we introduce a novel goal-oriented gaze estimation module (GEM) to improve the discriminative attribute localization based on the class-level attributes for ZSL. We aim to predict the actual human gaze location to get the visual attention regions for recognizing a novel object guided by attribute description. Specifically, the task-dependent attention is learned with the goal-oriented GEM, and the global image features are simultaneously optimized with the regression of local attribute features. Experiments on three ZSL benchmarks, i.e., CUB, SUN and AWA2, show the superiority or competitiveness of our proposed method against the state-of-the-art ZSL methods. The ablation analysis on real gaze data CUB-VWSW also validates the benefits and accuracy of our gaze estimation module. This work implies the promising benefits of collecting human gaze dataset and automatic gaze estimation algorithms on high-level computer vision tasks. The code is available at https://github.com/osierboy/GEM-ZSL.
Yang Liu 0357, Lei Zhou 0008, Xiao Bai 0001, Yifei Huang 0002, Lin Gu 0003, Jun Zhou 0001, Tatsuya Harada
CVPR2
2021 Relation-Aware Reasoning with Graph Convolutional Network
Lei Zhou 0008, Yang Liu 0357, Xiao Bai 0001, Xiang Wang 0014, Chen Wang 0026, Liang Zhang 0044, Lin Gu 0003
ICIG (1)1
2020 Matrix Classifier On Dynamic Functional Connectivity For Mci Identification
abstract
One of the most popular method for Alzheimer's disease (AD) diagnosis is exploring the Brain functional connectivity (FC) from resting-state functional magnetic resonance imaging (RS-fMRI). To early prevent AD, it is crucial to distinguish AD and and its preclinical stage, mild cognitive impairment (MCI) and early MCI (eMCI). In many existing works, dynamic functional connectivity (dFC) which contains rich spatiotemporal information has been exploited for the MCI and eMCI identification. However, most of these dFC based methods only consider the correlation between discrete brain status while ignore the valuable spatiotemporal information contained in dFC. To overcome this limitation, we propose a matrix classifier based method on the dFC signal for MCI and eMCI identification. Specifically, we first represent the dFC correlations by matrix features which contain rich spatiotemporal information and then learn the support matrix machines (SMM) to classify AD and its preclinical stage. Experiments on 600 real people data provide by the Alzheimer's Disease Neuroimaging Initiative (ADNI) demonstrate that our proposed matrix classifier based method outperforms other FC and dFC based methods for both normal controls (NC)/MCI identification and NC/eMCI identification.
Lei Zhou 0008, Liang Zhang 0044, Xiao Bai 0001, Jun Zhou 0001
ICIP1
2020 Fast Subspace Clustering Based on the Kronecker Product
abstract
Subspace clustering is a useful technique for many computer vision applications in which the intrinsic dimension of high-dimensional data is often smaller than the ambient dimension. Spectral clustering, as one of the main approaches to subspace clustering, often takes on a sparse representation or a low-rank representation to learn a block diagonal self-representation matrix for subspace generation. However, existing methods require solving a large scale convex optimization problem with a large set of data, with computational complexity reaches O(N3) for N data points. Therefore, the efficiency and scalability of traditional spectral clustering methods can not be guaranteed for large scale datasets. In this paper, we propose a subspace clustering model based on the Kronecker product. Due to the property that the Kronecker product of a block diagonal matrix with any other matrix is still a block diagonal matrix, we can efficiently learn the representation matrix which is formed by the Kronecker product of k smaller matrices. By doing so, our model significantly reduces the computational complexity to O(kN3/k). Furthermore, our model is general in nature, and can be adapted to different regularization based subspace clustering methods. Experimental results on two public datasets show that our model significantly improves the efficiency compared with several state-of-the-art methods. Moreover, we have conducted experiments on synthetic data to verify the scalability of our model for large scale datasets.
Lei Zhou 0008, Xiao Bai 0001, Liang Zhang 0044, Jun Zhou 0001, Edwin R. Hancock
ICPR1
2020 Learning binary code for fast nearest subspace search
Lei Zhou 0008, Xiao Bai 0001, Xianglong Liu 0001, Jun Zhou 0001, Edwin R. Hancock
Pattern Recognit.1
2019 A One-step Pruning-recovery Framework for Acceleration of Convolutional Neural Networks
abstract
Acceleration of convolutional neural network has received increasing attention during the past several years. Among various acceleration techniques, filter pruning has its inherent merit by effectively reducing the number of convolution filters. However, most filter pruning methods resort to tedious and time-consuming layer-by-layer pruning-recovery strategy to avoid a significant drop of accuracy. In this paper, we present an efficient filter pruning framework to solve this problem. Our method accelerates the network in one-step pruning-recovery manner with a novel optimization objective function, which achieves higher accuracy with much less cost compared with existing pruning methods. Furthermore, our method allows network compression with global filter pruning. Given a global pruning rate, it can adaptively determine the pruning rate for each single convolutional layer, while these rates are often set as hyper-parameters in previous approaches. Evaluated on VGG- 16 and ResNet-50 using ImageNet, our approach outperforms several state-of-the-art methods with less accuracy drop under the same and even much fewer floating-point operations (FLOPs).
Xiao Bai 0001, Lei Zhou 0008, Jun Zhou 0001
ICTAI3
2019 Hyperspectral Image Classification Based on Non-Local Neural Networks
abstract
Deep convolutional neural network has been used for pixel-wise hyperspectral image classification. However, convolutional operations only extract features from local neighborhood at a time, which is inefficient to capture long-range dependencies. On the other hand, the lack of training samples often leads to over-fitting problem. In this paper, we proposed a neural network which is formed by sequential local and non-local operation blocks. The proposed network takes hyperspectral image as input and outputs the class inference of each pixel. The local operation module extracts local spatial and spectral features. The non-local operation module computes the response at a position as a weighted sum of the features at all positions. So it can capture long-range dependencies without stacking deep layers. Experiments on two public datasets show that our proposed method outperforms several state-of-the-art methods using limited number of training samples.
Chen Wang 0031, Xiao Bai 0001, Lei Zhou 0008, Jun Zhou 0001
IGARSS3
2019 Latent Distribution Preserving Deep Subspace Clustering
abstract
Subspace clustering is a useful technique for many computer vision applications in which the intrinsic dimension of high-dimensional data is smaller than the ambient dimension. Traditional subspace clustering methods often rely on the self-expressiveness property, which has proven effective for linear subspace clustering. However, they perform unsatisfactorily on real data with complex nonlinear subspaces. More recently, deep autoencoder based subspace clustering methods have achieved success owning to the more powerful representation extracted by the autoencoder network. Unfortunately, these methods only considering the reconstruction of original input data can hardly guarantee the latent representation for the data distributed in subspaces, which inevitably limits the performance in practice. In this paper, we propose a novel deep subspace clustering method based on a latent distribution-preserving autoencoder, which introduces a distribution consistency loss to guide the learning of distribution-preserving latent representation, and consequently enables strong capacity of characterizing the real-world data for subspace clustering. Experimental results on several public databases show that our method achieves significant improvement compared with the state-of-the-art subspace clustering methods.
Lei Zhou 0008, Xiao Bai 0001, Xianglong Liu 0001, Jun Zhou 0001, Edwin R. Hancock
IJCAI1
2019 Deep supervised hashing using symmetric relative entropy
Xueni Zhang, Lei Zhou 0008, Xiao Bai 0001, Xiushu Luan, Jie Luo 0004, Edwin R. Hancock
Pattern Recognit. Lett.2
2018 Binary Coding by Matrix Classifier for Efficient Subspace Retrieval
abstract
Fast retrieval in large-scale database with high-dimensional subspaces is an important task in many applications, such as image retrieval, video retrieval and visual recognition. This can be facilitated by approximate nearest subspace (ANS) retrieval which requires effective subspace representation. Most of the existing methods for this problem represent subspace by point in the Euclidean space or the Grassmannian space before applying the approximate nearest neighbor (ANN) search. However, the efficiency of these methods can not be guaranteed because the subspace representation step can be very time consuming when coping with high dimensional data. Moreover, the transforming process for subspace to point will cause subspace structural information loss which influence the retrieval accuracy. In this paper, we present a new approach for hashing-based ANS retrieval. The proposed method learns the binary codes for given subspace set following a similarity preserving criterion. It simultaneously leverages the learned binary codes to train matrix classifiers as hash functions. This method can directly binarize a subspace without transforming it into a vector. Therefore, it can efficiently solve the large-scale and high-dimensional multimedia data retrieval problem. Experiments on face recognition and video retrieval show that our method outperforms several state-of-the-art methods in both efficiency and accuracy.
Lei Zhou 0008, Xiao Bai 0001, Xianglong Liu 0001, Jun Zhou 0001
ICMR1