Shuyi Li 0003

dblp:79/9577-3 · DBLP profile ↗
← Back
34ranked-venue papers
13as first author
33since 2021 · last 2026
0000-0001-6264-9006ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 14 · 4 first-author · 14 since 2021Security and privacy · 6 · 6 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Scene adaptive dynamic multi-modal knowledge for video captioning
Haoying Sun, Shuyi Li 0003, Zeyu Xi, Yunhao Zhao, Lifang Wu
Expert Syst. Appl.2
2026 Towards generalized video captioning: An effective multi-modal knowledge graph perspective
Haoying Sun, Shuyi Li 0003, Zeyu Xi, Lifang Wu
Knowl. Based Syst.2
2026 DpFedFKP: Dynamic personalized federated learning for finger knuckle print recognition
Shuyi Li 0003, Jianian Hu, Bob Zhang 0001, Shanping Yu, Lifang Wu
Pattern Recognit.1
2026 Palmprint de-identification via diffusion model for high-quality and diverse synthesis
Licheng Yan, Bob Zhang 0001, Andrew Beng Jin Teoh, Lu Leng, Shuyi Li 0003, Ziyuan Yang 0001
Pattern Recognit.5
2025 Multi-Scale Parallel Hybrid Network for Palmprint Recognition
abstract
Due to its rich characteristic and high accuracy, palmprint patterns have played an important role in identity recognition. However, current palmprint recognition techniques often concentrate on extracting salient features, neglecting the interactions between local and global textural attributes in palmprint images. To address this issue, we propose a novel multi-scale network capable of integrating local and global features for palmprint recognition, dubbed Multi-Scale Parallel Hybrid Network (MSPHNet). Specifically, we designed a Parallel Hybrid Feature Extraction Block (PHEB), which includes a parallel Convolutional Neural Network (CNN) -based branch for local feature extraction and a Transformer-based branch for capturing global features. Furthermore, acknowledging the uneven distribution of critical texture and its details in palmprint images, we address the current research shortfall in pixel relationship analysis by introducing the Comprehensive Attention Block (CAB). This block integrates Spatial Attention (SA) and Pixel Attention (PA) to effectively leverage the imbalanced pixel distributions within the CNN branch. Extensive experimental results on numerous public palmprint datasets demonstrated that our proposed method achieves remarkable results.
Hao Yang 0056, Shuyi Li 0003, Bob Zhang 0001
ICASSP2
2025 DSSM-KG: Dual-Stream State-Space Modeling with Adaptive Knowledge Injection for Video Captioning
abstract
Video captioning aims to generate natural language descriptions of video content. Recent methods extract temporal and spatial information separately and use dataset-specific prior knowledge to enhance caption quality. However, they may be inadequate in joint spatiotemporal modeling and lack the utilization of commonsense knowledge, making it difficult to fully understand the video. To address these issues, this paper proposes a dual-stream state-space model (DSSM-KG) based on cross-modal knowledge injection. Specifically, by integrating the heterogeneous Mamba with the Transformer in both parallel and sequential manners, we construct the spatially enhanced dual-stream state-space module (S-DSSM) and the temporally enhanced dual-stream state-space module (T-DSSM) to strengthen joint spatiotemporal modeling. Additionally, a knowledge graph that integrates both commonsense and dataset-specific information is constructed and adaptively injected into the decoder to furnish the model with extensive video-related knowledge. Experimental results indicate that the structural designs of DSSM-KG, together with the knowledge injection mechanism, demonstrate significant efficacy, yielding competitive performance on mainstream video captioning datasets such as MSVD and MSR-VTT.
Haoying Sun, Shuyi Li 0003, Zeyu Xi, Lifang Wu
ICMR2
2025 Graph Contrastive Learning with Mutual Information Maximization for Incomplete Multimodal Recognition
Jianian Hu, Hao Yang 0056, Shuyi Li 0003, Lifang Wu
PRCV (15)5
2025 IdTrPalm: Identity-Traceable Stylized Palmprint Image Generation
Longfa Liu, Lunke Fei, Shuyi Li 0003, Jian Zhu 0001, Yuanrong Xu, Shaohua Teng
PRCV (15)3
2025 EIKA: Explicit & Implicit Knowledge-Augmented Network for entity-aware sports video captioning
Zeyu Xi, Ge Shi 0002, Haoying Sun, Shuyi Li 0003, Lifang Wu
Expert Syst. Appl.5
2025 Dual-Cohesion Metric Learning for Few-Shot Hand-Based Multimodal Recognition
abstract
Hand-based multimodal biometrics has garnered significant attention in information security and identity authentication. However, prevalent multimodal recognition techniques often extract the discriminant features from different modalities separately, ignoring the structural consistency between various modalities of the same class. Moreover, these methods generally focus on specific-scenarios, where recognition performance will be compromised when faced with different databases or multiple application scenarios. To solve these limitations, we present an innovative Dual-Cohesion Metric Learning (DCML) framework embedded in noise decomposition for few-shot hand multimodal biometrics. This approach comprehensively exploits multimodal features from both intra-modal and inter-modal structural consistency to improve its robustness across multiple applications. Specifically, DCML imposes a dual-cohesion mechanism to pull in the cross-modal distance of the same label and the within-class distance for each modal, while concurrently pushing away the between-class distance in the projected space. Furthermore, in the procedure of feature learning, the proposed DCML incorporates the low-rank constraint to mitigate the interference of noise in the raw data and enforces a sparsity constraint to extract more salient and compact features. Notably, our DCML can be flexibly extended to other multimodal biometrics. Extensive experimental results on six multimodal datasets demonstrated that our DCML outperforms the latest approaches in multiple multimodal recognition scenarios and has strong generalization ability even when the training samples are small.
Shuyi Li 0003, Bob Zhang 0001, Qinghua Hu
IEEE Trans. Inf. Forensics Secur.1
2025 Dynamic Personalized Federated Learning for Cross-Spectral Palmprint Recognition
abstract
Palmprint recognition has recently garnered attention due to its high accuracy, strong robustness, and high security. Existing deep learning-based palmprint recognition methods usually require large amounts of data for centralized training, facing the challenge of privacy disclosure. In addition, the non-independent and identically distributed (non-IID) issue in the multi-spectral palmprint images generally leads to the degradation of recognition performance. To tackle these problems, this paper proposes a dynamic personalized federated learning model for cross-spectral palmprint recognition, called DPFed-Palm. Specifically, for each client's local training, we present a new combination of loss functions to enforce the constraints of local models and effectively enhance the feature representation capability of models. Subsequently, DPFed-Palm aggregates the above-trained local models by using the combined aggregation strategies of the Federated Averaging (FedAvg) and Personalized Federated Learning (PFL) to obtain the best personalized global model of each client. For the selection of the best personalized global model, we develop a dynamic weight selection strategy to obtain the optimal weights of the local and global models by cross-spectral (cross-client) testing. Extensive experimental results on three public PolyU multispectral, IITD, and CASIA datasets show that the proposed method outperforms the existing techniques in privacy-preserving and recognition performance.
Shuyi Li 0003, Jianian Hu, Bob Zhang 0001, Xin Ning 0001, Lifang Wu
IEEE Trans. Image Process.1
2024 A Generative Method for Finger Knuckle Print Recognition
Bob Zhang 0001, Shuyi Li 0003, Hao Yang 0056
ICPR (28)3
2024 Dual low-rank structure embedding for robust visual information processing
Jianhang Zhou, Hengmin Zhang, Shuyi Li 0003, Bob Zhang 0001, Leyuan Fang, David Zhang 0001
Knowl. Based Syst.3
2024 Joint Discriminative Analysis With Low-Rank Projection for Finger Vein Feature Extraction
abstract
Over the last decades, finger vein biometric recognition has generated increasing attention because of its high security, accuracy, and natural anti-counterfeiting. However, most of the existing finger vein recognition approaches rely on image enhancement or require much prior knowledge, which limits their generalization ability to different databases and different scenarios. Additionally, these methods rarely take into account the interference of noise elements in feature representation, which is detrimental to the final recognition results. To tackle these problems, we propose a novel jointly embedding model, called Joint Discriminative Analysis with Low-Rank Projection (JDA-LRP), to simultaneously extract noise component and salient information from the raw image pixels. Specifically, JDA-LRP decomposes the input image into noise and clean components via low-rank representation and transforms the clean data into a subspace to adaptively learn salient features. To further extract the most representative features, the proposed JDA-LRP enforces the discriminative class-induced constraint of the training samples as well as the sparse constraint of the embedding matrix to aggregate the embedded data of each class in their respective subspace. In this way, the discriminant ability of the jointly embedding model is greatly improved, such that JDA-LRP can be adapted to multiple scenarios. Comprehensive experiments conducted on three commonly used finger vein databases and four palm-based biometric databases illustrate the superiority of our proposed model in recognition accuracy, computational efficiency, and domain adaptation.
Shuyi Li 0003, Ruijun Ma 0001, Jianhang Zhou, Bob Zhang 0001, Lifang Wu
IEEE Trans. Inf. Forensics Secur.1
2024 Robust and Sparse Least Square Regression for Finger Vein and Finger Knuckle Print Recognition
abstract
Due to their high reliability, security, and anti-counterfeiting, finger-based biometrics (such as finger vein and finger knuckle print) have recently received considerable attention. Despite recent advances in finger-based biometrics, most of these approaches leverage much prior information and are non-robust for different modalities or different scenarios. To address this problem, we propose a structured Robust and Sparse Least Square Regression (RSLSR) framework to adaptively learn discriminative features for personal identification. To achieve the powerful representation capacity of the input data, RSLSR synchronously integrates robust projection learning, noise decomposition, and discriminant sparse representation into a unified learning framework. Specifically, RSLSR jointly learns the most discriminative information from the original pixels of the finger images by introducing the$l_{2,1}$norm. A sparse transformation matrix and reconstruction error are simultaneously enforced to enhance its robustness to noise, thus making RSLSR adaptable to multi-scenarios. Extensive experiments on five contact-based and contactless-based finger databases demonstrate the clear superiority of the proposed RSLSR in terms of recognition accuracy and computational efficiency.
Shuyi Li 0003, Bob Zhang 0001, Lifang Wu, Ruijun Ma 0001, Xin Ning 0001
IEEE Trans. Inf. Forensics Secur.1
2024 Structure Suture Learning-Based Robust Multiview Palmprint Recognition
abstract
Low-quality palmprint images will degrade the recognition performance, when they are captured under the open, unconstraint, and low-illumination conditions. Moreover, the traditional single-view palmprint representation methods have been difficult to express the characteristics of each palm strongly, where the palmprint characteristics become weak. To tackle these issues, in this article, we propose a structure suture learning-based robust multiview palmprint recognition method (SSL_RMPR), which comprehensively presents the salient palmprint features from multiple views. Unlike the existing multiview palmprint representation methods, SSL_RMPR introduces a structure suture learning strategy to produce an elastic nearest neighbor graph (ENNG) on the reconstruction errors that simultaneously exploit the label information and the latent consensus structure of the multiview data, such that the discriminant palmprint representation can be adaptively enhanced. Meanwhile, a low-rank reconstruction term integrating with the projection matrix learning is proposed, in such a manner that the robustness of the projection matrix can be improved. Particularly, since no extra structure capture term is imposed into the proposed model, the complexity of the model can be greatly reduced. Experimental results have proven the superiority of the proposed SSL_RMPR by achieving the best recognition performances on a number of real-world palmprint databases.
Shuping Zhao, Lunke Fei, Jie Wen 0001, Bob Zhang 0001, Pengyang Zhao, Shuyi Li 0003
IEEE Trans. Neural Networks Learn. Syst.6
2023 Preference Contrastive Learning for Personalized Recommendation
Yulong Bai 0002, Meng Jian, Shuyi Li 0003, Lifang Wu
PRCV (9)3
2023 Linear discriminant analysis with generalized kernel constraint for robust image classification
Shuyi Li 0003, Hengmin Zhang, Ruijun Ma 0001, Jianhang Zhou, Jie Wen 0001, Bob Zhang 0001
Pattern Recognit.1
2023 Efficient and Effective Nonconvex Low-Rank Subspace Clustering via SVT-Free Operators
abstract
With the growing interest in convex and nonconvex low-rank matrix learning problems, the widely used singular value thresholding (SVT) operators associated with rank relaxation functions often face higher computational complexity, particularly for large-scale data matrices. To improve the efficacy of low-rank subspace clustering and overcome the issue of high computational complexity, this work proposes an efficient and effective method that avoids the need for singular value decomposition (SVD) computations in the iteration scheme. This can be achieved through the use of a computationally efficient and compact formulation, as well as automatic removal of the optimal mean, which reduces time consumption and enhances evaluation performance. A unified clustering framework based on Schatten-$p$norm regularized by$\ell _{2,q}$-norm can be formulated using this processing way, where inner element suppression can be achieved by choosing appropriate$p$,$q \in (0,1)$. Additionally, calculating the optimal mean enhances the robustness of the proposed method in the presence of outliers. Unlike the general iteration scheme of the alternating direction method of multiplier (ADMM) algorithms that introduce auxiliary splitting variables, the proposed alternating re-weighted least square (ARwLS) algorithm uses matrix inverse and multiplication computations to obtain analytic solutions, resulting in faster processing speeds for each sub-problem. To further investigate, we provide the computational complexity of each iteration and the theoretical analysis of the convergence property, where the derived solution is a stationary point. Experimental results on synthetic data and several benchmark datasets demonstrate the promising efficiency and efficacy of the proposed clustering method compared to classical and competing algorithms.
Hengmin Zhang, Shuyi Li 0003, Jing Qiu 0002, Yang Tang 0001, Jie Wen 0001, Zhiyuan Zha, Bihan Wen
IEEE Trans. Circuits Syst. Video Technol.2
2023 Flexible and Generalized Real Photograph Denoising Exploiting Dual Meta Attention
abstract
Supervised deep learning techniques have been widely explored in real photograph denoising and achieved noticeable performances. However, being subject to specific training data, most current image denoising algorithms can easily be restricted to certain noisy types and exhibit poor generalizability across testing sets. To address this issue, we propose a novel flexible and well-generalized approach, coined as dual meta attention network (DMANet). The DMANet is mainly composed of a cascade of the self-meta attention blocks (SMABs) and collaborative-meta attention blocks (CMABs). These two blocks have two forms of advantages. First, they simultaneously take both spatial and channel attention into account, allowing our model to better exploit more informative feature interdependencies. Second, the attention blocks are embedded with the meta-subnetwork, which is based on metalearning and supports dynamic weight generation. Such a scheme can provide a beneficial means for self and collaborative updating of the attention maps on-the-fly. Instead of directly stacking the SMABs and CMABs to form a deep network architecture, we further devise a three-stage learning framework, where different blocks are utilized for each feature extraction stage according to the individual characteristics of SMAB and CMAB. On five real datasets, we demonstrate the superiority of our approach against the state of the art. Unlike most existing image denoising algorithms, our DMANet not only possesses a good generalization capability but can also be flexibly used to cope with the unknown and complex real noises, making it highly competitive for practical applications.
Ruijun Ma 0001, Shuyi Li 0003, Bob Zhang 0001, Leyuan Fang
IEEE Trans. Cybern.2
2023 Learning Sparse and Discriminative Multimodal Feature Codes for Finger Recognition
abstract
Compared with uni-modal biometrics systems, multimodal biometrics systems using multiple sources of information for establishing an individual’s identity have received considerable attention recently. However, most traditional multimodal biometrics techniques generally extract features from each modality independently, ignoring the implicit associations between different modalities. In addition, most existing work uses hand-crafted descriptors that are difficult to capture the latent semantic structure. This paper proposes to learn the sparse and discriminative multimodal feature codes (SDMFCs) for multimodal finger recognition, which simultaneously takes into account the specific and common information among inter-modality and intra-modality. Specifically, given the multimodal finger images, we first establish the local difference matrix to capture informative texture features in local patches. Then, we aim to jointly learn discriminative and compact binary codes by constraining the observations from multiple modalities. Finally, we develop a novel SDMFC-based multimodal finger recognition framework, which integrates the local histograms of each division block in the learned binary codes together for classification. Experimental results on three commonly used finger databases demonstrate the effectiveness and robustness of the proposed framework in multimodal biometrics tasks.
Shuyi Li 0003, Bob Zhang 0001, Lunke Fei, Shuping Zhao, Yicong Zhou
IEEE Trans. Multim.1
2023 Adaptive Graph Embedded Preserving Projection Learning for Feature Extraction and Selection
abstract
Preserving projection learning has been widely used in feature extraction and selection for unsupervised image classification. Generally, some related methods constructed a graph to represent the nearest neighbor relationships of the data based on the Euclidean distances among different samples, which used 0 or 1 to predefine whether two samples are from the same class. Since a simple Euclidean distance is sensitive to noise, the predefined graph cannot produce exact correlations between the two samples. What is more, the predefined graph cannot reflect the structure of the projected data on a latent subspace when the projection matrix is learned. To solve these problems, in this article a novel adaptive graph embedded preserving projection learning (AGE_PPL) method is proposed, first combining the sparsity-based graph learning and the projection learning as an integral framework for feature extraction and feature selection. In particular, a sparse representation term with$l_{1}$-norm is exploited in AGE_PPL to achieve the adaptive graph of the data to preserve the local structures among different samples while the projection matrix is learned. Meanwhile, a global-scale constraint is imposed to preserve the global structure of the data on a latent subspace. Therefore, the transformed samples will be more discriminative, allowing margins of the same class to be reduced, and margins among different classes to be enlarged. Experimental results proved the effectiveness of the proposed algorithm by obtaining competitive performances over other baseline and state-of-the-art methods. In addition, the proposed method is very flexible for feature selection and dimensionality reduction.
Shuping Zhao, Jigang Wu, Bob Zhang 0001, Lunke Fei, Shuyi Li 0003, Pengyang Zhao
IEEE Trans. Syst. Man Cybern. Syst.5
2022 Generative Adaptive Convolutions for Real-World Noisy Image Denoising
abstract
Recently, deep learning techniques are soaring and have shown dramatic improvements in real-world noisy image denoising. However, the statistics of real noise generally vary with different camera sensors and in-camera signal processing pipelines. This will induce problems of most deep denoisers for the overfitting or degrading performance due to the noise discrepancy between the training and test sets. To remedy this issue, we propose a novel flexible and adaptive denoising network, coined as FADNet. Our FADNet is equipped with a plane dynamic filter module, which generates weight filters with flexibility that can adapt to the specific input and thereby impedes the FADNet from overfitting to the training data. Specifically, we exploit the advantage of the spatial and channel attention, and utilize this to devise a decoupling filter generation scheme. The generated filters are conditioned on the input and collaboratively applied to the decoded features for representation capability enhancement. We additionally introduce the Fourier transform and its inverse to guide the predicted weight filters to adapt to the noisy input with respect to the image contents. Experimental results demonstrate the superior denoising performances of the proposed FADNet versus the state-of-the-art. In contrast to the existing deep denoisers, our FADNet is not only flexible and efficient, but also exhibits a compelling generalization capability, enjoying tremendous potential for practical usage.
Ruijun Ma 0001, Shuyi Li 0003, Bob Zhang 0001
AAAI2
2022 Learning Unified Binary Feature Codes for Cross-Illumination Palmprint Recognition
Wei Jia 0001, Lunke Fei, Shuping Zhao, Shuyi Li 0003, Jie Wen 0001, Jinrong Cui
CGI4
2022 Row-sparsity Binary Feature Learning for Open-set Palmprint Recognition
abstract
Binary feature representation methods have received increasing attention due to their high efficiency and great robustness to illumination variation. However, most of them are hand-designed feature descriptors that generally require much prior knowledge in their design. This paper introduces a Row-sparsity Binary Feature Learning (Rs-BFL) method to adaptively learn and encode palmprint features for open-set palmprint recognition. Given the training palmprint images, RsBFL jointly learns a bank of linear projection functions that transform the informative texture features into discriminative binary codes. Afterwards, we calculate the block-wise histograms of each feature map and concatenate them as the final feature representation. Based on the pre-trained projection matrix, we mapped the palmprint texture features of the test samples into binary features for matching. For RsBFL, we enforce three criteria: 1) the quantization error between the projected real-valued features and the binary features is minimized, at the same time, the projection noise is minimized; 2) the latent label semantic information is utilized to minimize the distance of the within-class samples and simultaneously maximize the distance of the between-class samples; 3) the$l_{2,1}$norm is used to make the projection matrix to extract more discriminative features. Extensive experimental results on two publicly accessible palmprint datasets demonstrated the effectiveness and powerful learning capability of the proposed method.
Shuyi Li 0003, Ruijun Ma 0001, Jianhang Zhou, Bob Zhang 0001
IJCB1
2022 Learning Compact Multirepresentation Feature Descriptor for Finger-Vein Recognition
abstract
Due to its high anti-counterfeiting and universality, the use of finger-vein pattern for identity authentication has recently attracted extensive attention in academia and industry. Despite recent advances in finger-vein recognition, most of the hand-crafted descriptors require strong prior knowledge, which may be ineffective in expressing its distinctiveness. In this paper, we present a novel compact multi-representation feature descriptor (CMrFD) with visual and semantic consistency, for finger-vein feature representation. Given the finger-vein images, we first form two-view representations to describe the informative vein features in local patches. Then, we jointly learn a feature transformation to map the two-view representations into discriminative binary codes. For the projection function, we linearly combine multi-view information and minimize the quantization error between the projected binary features and the original real-valued features. In terms of visual consistency, we minimize the Euclidean distance of each representation from the same class, at the same time, maximize the Euclidean distance from different classes in the projected space. Semantic consistency is used to ensure that similar images have compact multi-representation combined projection features. Lastly, we calculate the block-wise histograms as the final extracted features for finger-vein recognition. Experimental results on four widely used finger-vein databases demonstrate that the proposed method outperforms the state-of-the-art finger-vein recognition methods.
Shuyi Li 0003, Ruijun Ma 0001, Lunke Fei, Bob Zhang 0001
IEEE Trans. Inf. Forensics Secur.1
2022 Meta PID Attention Network for Flexible and Efficient Real-World Noisy Image Denoising
abstract
Recent deep convolutional neural networks for real-world noisy image denoising have shown a huge boost in performance by training a well-engineered network over external image pairs. However, most of these methods are generally trained with supervision. Once the testing data is no longer compatible with the training conditions, they can exhibit poor generalization and easily result in severe overfitting or degrading performances. To tackle this barrier, we propose a novel denoising algorithm, dubbed as Meta PID Attention Network (MPA-Net). Our MPA-Net is built based upon stacking Meta PID Attention Modules (MPAMs). In each MPAM, we utilize a second-order attention module (SAM) to exploit the channel-wise feature correlations with second-order statistics, which are then adaptively updated via a proportional-integral-derivative (PID) guided meta-learning framework. This learning framework exerts the unique property of the PID controller and meta-learning scheme to dynamically generate filter weights for beneficial update of the extracted features within a feedback control system. Moreover, the dynamic nature of the framework enables the generated weights to be flexibly tweaked according to the input at test time. Thus, MPAM not only achieves discriminative feature learning, but also facilitates a robust generalization ability on distinct noises for real images. Extensive experiments on ten datasets are conducted to inspect the effectiveness of the proposed MPA-Net quantitatively and qualitatively, which demonstrates both its superior denoising performance and promising generalization ability that goes beyond those of the state-of-the-art denoising methods.
Ruijun Ma 0001, Shuyi Li 0003, Bob Zhang 0001, Haifeng Hu 0001
IEEE Trans. Image Process.2
2022 Towards Fast and Robust Real Image Denoising With Attentive Neural Network and PID Controller
abstract
With the development of deep learning technologies, recent research on real-world noisy image denoising has achieved a considerable improvement in performance. However, a common limitation for existing approaches is the imbalanced trade-off between denoising accuracy and efficiency. To address this problem, we propose a robust and efficient denoiser, called a hierarchical-based PID-attention denoising network (HPDNet), to flexibly deal with the sophisticated noise. The core of our algorithm is the PID-attentive recurrent network (PAR-Net) whose framework mainly consists of the LSTM network and PID controller. PAR-Net inherits the advantages of both the attentive recurrent network and control action, which can encourage more discriminatory feature representations. This learning procedure is implemented within a feedback control system, allowing a faster and more robust means to enhance feature discriminability. Furthermore, by decomposing the noisy image and stacking the PAR-Nets, our PAR-Net can work on a progressively hierarchical framework, and hence obtain multi-scale features and manageable successive refinements. On several widely used datasets, the proposed HPDNet demonstrates high efficiency, while delivering a better perceptually appealing image quality over state-of-the-art image denoising methods.
Ruijun Ma 0001, Shuyi Li 0003, Bob Zhang 0001
IEEE Trans. Multim.2
2021 An Adaptive Discriminant and Sparsity Feature Descriptor for Finger Vein Recognition
abstract
The use of finger-vein (FV) trait for the purpose of identity authentication has attracted much attention in recent years. However, most of the conventional FV recognition methods are hand-crafted by design and require strong prior knowledge, which are ineffective at expressing the distinctiveness of the FV images. In this paper, we propose an adaptive discriminant and sparsity feature descriptor (DSFD) for FV feature extraction and recognition. Specifically, we first form a direction difference vector (DDV) to better represent the direction feature of the FV images. Afterwards, the DSFD adaptively projects the DDVs into a feature space with discriminative binary codes in which the distance of the within-class samples is minimized and simultaneously the distance of the between-class samples is maximized. Lastly, we concatenate the block-wise histograms into a global histogram as the final feature descriptor for FV recognition. Experimental results on two publicly available FV databases demonstrate the effectiveness of the proposed method.
Shuyi Li 0003, Bob Zhang 0001
ICASSP1
2021 Local discriminant coding based convolutional feature representation for multimodal finger recognition
Shuyi Li 0003, Bob Zhang 0001, Shuping Zhao, Jinfeng Yang
Inf. Sci.1
2021 Jointly learning compact multi-view hash codes for few-shot FKP recognition
Lunke Fei, Bob Zhang 0001, Jie Wen 0001, Shaohua Teng, Shuyi Li 0003, David Zhang 0001
Pattern Recognit.5
2021 Joint discriminative feature learning for multimodal finger recognition
Shuyi Li 0003, Bob Zhang 0001, Lunke Fei, Shuping Zhao
Pattern Recognit.1
2021 Joint Discriminative Sparse Coding for Robust Hand-Based Multimodal Recognition
abstract
Multimodal biometrics recognition has recently attracted much interest for its higher security and effectiveness compared with unimodal biometrics recognition. However, most of the conventional multimodal recognition approaches generally focus on extracting semantic information from different modalities independently, while ignoring the implicit correlations among inter-modality. In this paper, we propose a simple yet effective supervised multimodal feature learning method, called joint discriminative sparse coding (JDSC), which is applied for hand-based multimodal recognition including finger-vein and finger-knuckle-print fusion, palm-vein and palmprint fusion, as well as palm-vein and dorsal-hand-vein fusion. Considering that relevant samples from different modalities have semantic correlations, JDSC projects the raw data into a shared space in which the distance of the between-class is maximized and the distance of the within-class is minimized, at the same time, the correlation among the inter-modality of the within-class is maximized. Therefore, sparse binary codes quantified by the obtained projection matrix can have more discriminative power for multimodal recognition tasks. Thorough experiments on six commonly used multimodal datasets demonstrate the superiority of our proposed method over several state-of-the-art techniques.
Shuyi Li 0003, Bob Zhang 0001
IEEE Trans. Inf. Forensics Secur.1
2020 Discriminant and Sparsity Based Least Squares Regression with l1 Regularization for Feature Representation
abstract
Least squares regression (LSR) has two main issues that greatly limits the improvement of performance: 1) The target matrix is too rigid leading to a large regression error; 2) the underlying geometric structure of the training data is often ignored to learn a more discriminative projection matrix. To solve these dilemmas, this paper presents a discriminant and sparsity based least squares regression with l1regularization (DS_LSR). In DS_LSR, the sparse coefficient matrix of the training data with l1regularization is jointly learned with the projection matrix to make the projection matrix discriminative. In addition, an orthogonal relaxed term is introduced to hold the structure of regression targets while relaxing the rigid label matrix. Extensive experimental results demonstrate the effectiveness of the proposed method in classification accuracy.
Shuping Zhao, Bob Zhang 0001, Shuyi Li 0003
ICASSP3