Linna Zhang

dblp:86/8620 · DBLP profile ↗
← Back
22ranked-venue papers
3as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Twin cross contrastive learning with multi-modality fusion for drug-target affinity prediction
Linna Zhang, Zhaowei Wang 0005, Wuhao Liu, Xiaodong Duan, Qiguo Dai
Artif. Intell. Medicine1
2026 Vision-Semantics-Label: A New Two-Step Paradigm for Action Recognition With Large Language Model
abstract
In recent years, the rapid advancement of multi-modal large language models has propelled the development of video-based conversation models. Due to their exceptional video understanding capabilities, there is often an expectation that these models can handle all video-related tasks, including action recognition. However, because action recognition datasets typically lack semantic information, limiting the performance of dialogue models. Additionally, as these dialogue models are designed for video understanding, they frequently overlook critical information required for action recognition—continuous motion—in their model architecture and training dataset configurations. To address these challenges, we first propose a novel two-step mapping framework based on large language models, termed “Vision-Semantics-Label” mapping, to better adapt video-based large language models for action recognition. In the first step, we proposed a visual-skeletal collaborative learning large language model (VS-LLM), which utilizes human keypoints to compensate for the missing motion details without increasing the input token length of the large language model. In the second step, we designed two mapping methods: verb noun match (VN-Match) and all text match (ALL-Match), which can effectively extract relevant action descriptions from the text. Finally, we construct semantic action recognition datasets to ensure that the training data inherently contains action details, enabling the model to better achieve action recognition. We evaluate our approach on five benchmark datasets, demonstrating the state-of-the-art performance of large language models in action recognition. The source code and dataset are publicly available at https://github.com/xiaoyu92568/VS-LLM.
Wanru Xu, Shichao Kan, Linna Zhang, Yi Jin 0001, Yi-Gang Cen, Yidong Li
IEEE Trans. Circuits Syst. Video Technol.4
2025 Injecting Cross-modal Fine-Grained Perception into LLMs for 3D Object-of-Interest Understanding
abstract
Recent advancements in 3D Large Language Models (LLMs) have revealed significant potential in enhancing the understanding of 3D scenes. However, previous methods have struggled with extracting and utilizing fine-grained information of 3D objects for the coarsness of point clouds, resulting in limitations in understanding object-of-interested (OoI) within the scene. To address this issue, we introduce the object-centric 2D-3D interaction module for enhancing the ability of LLMs for 3D understanding tasks, which consists of the fine-grained 2D representation perception and the object-centric 3D scene representation perception. Specifically, the 2D representation associated with 3D objects is captured based on cross-modal semantic consistency without any spatial projector. Experimental results show that our model significantly outperforms existing methods on benchmarks including ScanRefer and ScanQA.
Qianqian Sun, Lu Shi 0004, Linna Zhang, Gaoyun An, Yi Jin 0001, Yidong Li, Yi-Gang Cen
ICME3
2025 Hierarchical Meta-prototypes Network for Few-shot Action Recognition
abstract
Existing few-shot action recognition (FSAR) studies predominantly follow a metric learning framework, where prototypes are generated directly from features extracted by an encoder, and classification is performed via distance-based matching. However, due to the limited number of available samples, significant variations exist between different video features of the same class. As a result, the same query video may yield different classification results when matched against different sets of support videos. To address this issue, we propose a novel Hierarchical Meta-Prototypes Network (HMP-Net). The key innovation of our approach lies in the introduction of a category-agnostic and feature-agnostic meta-prototype module, which guides video feature mapping into a more suitable feature space. To optimize this meta-prototype, we design an alternating meta-prototype training strategy, where the model first learns to transform features under a fixed meta-prototype, and then the meta-prototype is refined to better guide feature mapping. Additionally, to adapt image-based metric learning models to video-based FSAR tasks, we introduce a series of lightweight adaptation modules. Specifically, we integrate an adapter into the encoder to improve video frame feature extraction, design a hierarchical prototype generation mechanism to enhance overall video understanding, and incorporate a task-specific perception module to extract unique features for each task. These adaptations make our model better suited for FSAR, significantly improving performance. We evaluate HMP-Net on five challenging benchmarks, and experimental results demonstrate that our model achieves new state-of-the-art performance on HMDB51, UCF101, Kinetics, and SthSthV2-Small. Extensive empirical evaluations further highlight the effectiveness and robustness of HMP-Net.
Yi-Gang Cen, Wanru Xu, Yue Zhang 0065, Yi Jin 0001, Yidong Li, Linna Zhang
ACM Multimedia7
2025 Model Evaluation-Driven Aggregation: FedMEDA for Robust Non-IID Federated Learning
abstract
Federated learning aims to jointly train a general and robust global model through users’ local training models (not local private data). Therefore, a crucial step is how to aggregate local models, which poses a challenge when the data among users is non-independent and identically distributed (non-i.i.d.). This paper proposes a new federated model-pooling algorithm, FedMEDA. From the perspective of model evaluation, this algorithm monitors the model pool, screens out high-quality local models beneficial to the construction of the global model, and combines them to achieve a more robust aggregation. We verify that personalized models can be accurately evaluated through a simple weighted F1-score. Our experiments validate the superior performance of FedMEDA in scenarios where the Dirichlet distribution is used to simulate non-i.i.d. user data. In addition, FedMEDA is compatible with recent work on standardizing user-model training. Only the aggregation strategy needs to be changed while keeping the other parts of the federated learning algorithm unchanged.
Bosong Zhang, Linna Zhang, Hai Wang 0010
TrustCom2
2025 Feature Transformation Reconstruction (FTR) Network for Unsupervised Anomaly Detection
abstract
The goal of the feature reconstruction network based on an autoencoder in the training phase is to force the network to reconstruct the input features well. The network tends to learn shortcuts of “identity mapping,” which leads to the network outputting abnormal features as they are in the inference phase. As such, the abnormal features based on reconstruction error cannot be distinguished from normal features, significantly limiting the detection performance of such methods. To address this issue, we propose a feature transformation reconstruction (FTR) network, which can avoid the identity mapping problem. Specifically, we use a normalizing flow model as a feature transformation (FT) network to transform input features into other forms. The training goal of the feature reconstruction (FR) network is no longer to reconstruct the input features but to reconstruct the transformed features, effectively avoiding the shortcut of learning the “identity map.” Furthermore, this paper proposes a masked convolutional attention (MCA) module, which randomly masks the input features in the training phase and reconstructs the input features in a self‐supervised manner. In the testing phase, the MCA can effectively suppress the excessive reconstruction of abnormal features and further improve anomaly detection performance. FTR achieves the scores of the area under the receiver operating characteristic curve (AUROC) at 99.5% and 97.8% on the MVTec AD and BTAD datasets, respectively, outperforming other state‐of‐the‐art methods. Moreover, FTR is faster than the existing methods, with a high speed of 137 frames per second (FPS) on a 3080ti GPU.
Linna Zhang, Lanyao Zhang, Qi Cao 0002, Shichao Kan, Yi-Gang Cen, Fugui Zhang, Yansen Huang
Int. J. Intell. Syst.1
2025 Multi-Modal Self-Perception Enhanced Large Language Model for 3D Region-of-Interest Captioning With Limited Data
abstract
3D Region-of-Interest (RoI) Captioning involves translating a model's understanding of specific objects within a complex 3D scene into descriptive captions. Recent advancements in Large Language Models (LLMs) have shown great potential in this area. Existing methods capture the visual information from RoIs as input tokens for LLMs. However, this approach may not provide enough detailed information for LLMs to generate accurate region-specific captions. In this paper, we introduce Self-RoI, a Large Language Model with multi-modal self-perception capabilities for 3D RoI captioning. To ensure LLMs receive more precise and sufficient information, Self-RoI incorporates Implicit Textual Info. Perception to construct a multi-modal vision-language information. This module utilizes a simple mapping network to generate textual information about basic properties of RoI from vision-following response of LLMs. This textual information is then integrated with the RoI's visual representation to form a comprehensive multi-modal instruction for LLMs. Given the limited availability of 3D RoI-captioning data, we propose a two-stage training strategy to optimize Self-RoI efficiently. In the first stage, we align 3D RoI vision and caption representations. In the second stage, we focus on 3D RoI vision-caption interaction, using a disparate contrastive embedding module to improve the reliability of the implicit textual information and employing language modeling loss to ensure accurate caption generation. Our experiments demonstrate that Self-RoI significantly outperforms previous 3D RoI captioning models. Moreover, the Implicit Textual Info. Perception can be integrated into other multi-modal LLMs for performance enhancement. We will make our code available for further research.
Lu Shi 0004, Shichao Kan, Yi Jin 0001, Linna Zhang, Yi-Gang Cen
IEEE Trans. Multim.4
2025 Low-Shot Unsupervised Visual Anomaly Detection via Sparse Feature Representation
abstract
Visual anomaly detection is an essential component in modern industrial manufacturing. Existing studies using notions of pairwise similarity distance between a test feature and nominal features have achieved great breakthroughs. However, the absolute similarity distance lacks certain generalizations, making it challenging to extend the comparison beyond the available samples. This limitation could potentially hamper anomaly detection performance in scenarios with limited samples. This article presents a novel sparse feature representation anomaly detection (SFRAD) framework, which formulates the anomaly detection as a sparse feature representation problem; and notably proposes an anomaly score by orthogonal matching pursuit (ASOMP) as a novel detection metric. Specifically, SFRAD calculates the Gaussian kernel distance between the test feature and its sparse representation in the nominal feature space for anomaly detection. Here, the orthogonal matching pursuit (OMP) algorithm is adopted to achieve the sparse feature representation. Moreover, to construct a low-redundancy memory bank storing the basis features for sparse representation, a novel basis feature sampling (BFS) algorithm is proposed by considering both the maximum coverage and the optimum feature representation simultaneously. As a result, SFRAD incorporates both the advantages of absolute similarity and linear representation; and this enhances the generalization in low-shot scenarios. Extensive experiments on the MVTec anomaly detection (MVTec AD), Kolektor surface-defect dataset (KolektorSDD), Kolektor surface-defect dataset 2 (KolektorSDD2), MVTec logical constraints anomaly detection (MVTec LOCO AD), Visual anomaly (VISA), Modified national institute of standards and technology (MNIST), and CIFAR-10 datasets demonstrate that our proposed SFRAD outperforms the previous methods and achieves state-of-the-art unsupervised anomaly detection performance. Notably, significantly improved outcomes and results have also been achieved on low-shot anomaly detection. Code is available at https://github.com/fanghuisky/SFRAD.
Fanghui Zhang, Haiyue Zhu, Yi-Gang Cen, Shichao Kan, Linna Zhang, Prahlad Vadakkepat, Tong Heng Lee
IEEE Trans. Neural Networks Learn. Syst.5
2024 Federated Learning Greedy Aggregation Optimization for Non-Independently Identically Distributed Data
abstract
In the domain of federated learning, traditional federated averaging algorithms encounter difficulties in maximizing accuracy in non-IID data circumstances due to the large data volume and the low model efficiency ratio of participants. To address the problem of non-IID data among participants in FL, this paper puts forward a novel greedy aggregation algorithm based on the Fapmodel evaluation, named FedEGA. FedEGA incorporates models from each participant into the global model successively as potential components and builds the global model by evenly distributing multiple model weights that are fine-tuned with various hyperparameter configurations. During the construction process, a dynamic weight parameter mechanism is adopted to balance the accuracy and precision evaluation of the model. By using validation set scores on the central server, FedEGA ranks models in descending order, ensuring that the global model does not perform worse than the best individual model on the retained validation set, thereby alleviating issues caused by client drift and the hindrance to maximizing accuracy due to non-IID data. Experiments on public datasets and real-world data show that our algorithm surpasses classic FL algorithms such as Federated Averaging (FedAvg), Federated Proximal (FedProx) optimization, and Federated Self-Regularization (FedSR) in terms of maximum accuracy and communication efficiency in non-IID scenarios. Additionally, experiments with imbalanced data confirm its stronger robustness and generalization capabilities.
Bosong Zhang, Hai Wang 0010, Linna Zhang
TrustCom4
2023 RA-KD: Random Attention Map Projection for Knowledge Distillation
Linna Zhang, Yuehui Chen, Yaou Zhao
ICIC (4)1
2023 End-to-end feature diversity person search with rank constraint of cross-class matrix
Yue Zhang 0065, Shuqin Wang 0001, Shichao Kan, Yi-Gang Cen, Linna Zhang
Neurocomputing5
2023 Multiscale spatial temporal attention graph convolution network for skeleton-based anomaly behavior detection
Shichao Kan, Fanghui Zhang, Yi-Gang Cen, Linna Zhang, Damin Zhang
J. Vis. Commun. Image Represent.5
2023 A graph model-based multiscale feature fitting method for unsupervised anomaly detection
Fanghui Zhang, Shichao Kan, Damin Zhang, Yi-Gang Cen, Linna Zhang, Vladimir Mladenovic
Pattern Recognit.5
2022 Nonconvex low-rank and sparse tensor representation for multi-view subspace clustering
Shuqin Wang 0001, Yongyong Chen, Yi-Gang Cen, Linna Zhang, Hengyou Wang, Viacheslav V. Voronin
Appl. Intell.4
2021 Cross-domain Person Re-identification Based on the Sample Relation Guidance
Yue Zhang 0065, Fanghui Zhang, Shichao Kan, Linna Zhang, Jiaping Zong, Yi-Gang Cen
ICIG (2)4
2021 Low-Rank And Sparse Tensor Representation For Multi-View Subspace Clustering
abstract
Learning an effective affinity matrix as the input of spectral clustering to achieve promising multi-view clustering is a key issue of subspace clustering. In this paper, we propose a low-rank and sparse tensor representation (LRSTR) method that learns the affinity matrix through a self-representation tensor and retains the similarity information of the view dimensions for multi-view subspace clustering. Specifically, the proposed LRSTR method imposes the tensor nuclear norm and tensor sparse constraints on self-representation tensor to characterize the relationship between views. The optimization model is solved under the framework of alternating direction method of multiplier. Experimental results on four datasets show that the proposed LRSTR method is better than several state-of-the-art methods.
Shuqin Wang 0001, Yongyong Chen, Yigang Ce, Linna Zhang, Viacheslav V. Voronin
ICIP4
2021 Error-robust low-rank tensor approximation for multi-view clustering
Shuqin Wang 0001, Yongyong Chen, Yi Jin 0001, Yi-Gang Cen, Yidong Li, Linna Zhang
Knowl. Based Syst.6
2020 Metric learning-based kernel transformer with triplets and label constraints for feature fusion
Shichao Kan, Linna Zhang, Zhihai He, Yi-Gang Cen, Shiming Chen 0001, Jikun Zhou
Pattern Recognit.2
2020 Abnormal event detection in surveillance videos based on low-rank and compact coefficient dictionary learning
Zhenjiang Miao, Yi-Gang Cen, Xiao-Ping Zhang 0002, Linna Zhang, Shiming Chen 0001
Pattern Recognit.5
2019 Weighted-learning-instance-based retrieval model using instance distance
Hao Wu 0022, Yueli Li, Xiaohan Bi, Linna Zhang, Rongfang Bie, Junqi Guo
Mach. Vis. Appl.5
2019 Supervised Deep Feature Embedding With Handcrafted Feature
abstract
Image representation methods based on deep convolutional neural networks (CNNs) have achieved the state-of-the-art performance in various computer vision tasks, such as image retrieval and person re-identification. We recognize that more discriminative feature embeddings can be learned with supervised deep metric learning and handcrafted features for image retrieval and similar applications. In this paper, we propose a new supervised deep feature embedding with a handcrafted feature model. To fuse handcrafted feature information into CNNs and realize feature embeddings, a general fusion unit is proposed (called Fusion-Net). We also define a network loss function with image label information to realize supervised deep metric learning. Our extensive experimental results on the Stanford online products' data set and the in-shop clothes retrieval data set demonstrate that our proposed methods outperform the existing state-of-the-art methods of image retrieval by a large margin. Moreover, we also explore the applications of the proposed methods in person re-identification and vehicle re-identification; the experimental results demonstrate both the effectiveness and efficiency of the proposed methods.
Shichao Kan, Yi-Gang Cen, Zhihai He, Zhi Zhang 0005, Linna Zhang
IEEE Trans. Image Process.5
2018 Joint entropy based learning model for image retrieval
Hao Wu 0022, Yueli Li, Xiaohan Bi, Linna Zhang, Rongfang Bie, Yingzhuo Wang
J. Vis. Commun. Image Represent.4