VLDB 2026 Research / reviewers in the wild / expert
Ruijun Ma 0001
dblp:199/8810-1
· DBLP profile ↗
18ranked-venue papers
10as first author
14since 2021 · last 2026
0000-0001-6876-8153ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 first-author · 6 since 2021Security and privacy · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploiting anchor-free and graph reasoning framework for dense tea bud detection and picking point identification
Zhiye Shen, Yinghu Cai, Kaile Yuan, Wenbin Zhen, Ruijun Ma 0001 |
Eng. Appl. Artif. Intell. | 6 |
| 2024 | Joint Discriminative Analysis With Low-Rank Projection for Finger Vein Feature ExtractionabstractOver the last decades, finger vein biometric recognition has generated increasing attention because of its high security, accuracy, and natural anti-counterfeiting. However, most of the existing finger vein recognition approaches rely on image enhancement or require much prior knowledge, which limits their generalization ability to different databases and different scenarios. Additionally, these methods rarely take into account the interference of noise elements in feature representation, which is detrimental to the final recognition results. To tackle these problems, we propose a novel jointly embedding model, called Joint Discriminative Analysis with Low-Rank Projection (JDA-LRP), to simultaneously extract noise component and salient information from the raw image pixels. Specifically, JDA-LRP decomposes the input image into noise and clean components via low-rank representation and transforms the clean data into a subspace to adaptively learn salient features. To further extract the most representative features, the proposed JDA-LRP enforces the discriminative class-induced constraint of the training samples as well as the sparse constraint of the embedding matrix to aggregate the embedded data of each class in their respective subspace. In this way, the discriminant ability of the jointly embedding model is greatly improved, such that JDA-LRP can be adapted to multiple scenarios. Comprehensive experiments conducted on three commonly used finger vein databases and four palm-based biometric databases illustrate the superiority of our proposed model in recognition accuracy, computational efficiency, and domain adaptation. Shuyi Li 0003, Ruijun Ma 0001, Jianhang Zhou, Bob Zhang 0001, Lifang Wu |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Robust and Sparse Least Square Regression for Finger Vein and Finger Knuckle Print RecognitionabstractDue to their high reliability, security, and anti-counterfeiting, finger-based biometrics (such as finger vein and finger knuckle print) have recently received considerable attention. Despite recent advances in finger-based biometrics, most of these approaches leverage much prior information and are non-robust for different modalities or different scenarios. To address this problem, we propose a structured Robust and Sparse Least Square Regression (RSLSR) framework to adaptively learn discriminative features for personal identification. To achieve the powerful representation capacity of the input data, RSLSR synchronously integrates robust projection learning, noise decomposition, and discriminant sparse representation into a unified learning framework. Specifically, RSLSR jointly learns the most discriminative information from the original pixels of the finger images by introducing the$l_{2,1}$norm. A sparse transformation matrix and reconstruction error are simultaneously enforced to enhance its robustness to noise, thus making RSLSR adaptable to multi-scenarios. Extensive experiments on five contact-based and contactless-based finger databases demonstrate the clear superiority of the proposed RSLSR in terms of recognition accuracy and computational efficiency. Shuyi Li 0003, Bob Zhang 0001, Lifang Wu, Ruijun Ma 0001, Xin Ning 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Learning Attention in the Frequency Domain for Flexible Real Photograph DenoisingabstractRecent advancements in deep learning techniques have pushed forward the frontiers of real photograph denoising. However, due to the inherent pooling operations in the spatial domain, current CNN-based denoisers are biased towards focusing on low-frequency representations, while discarding the high-frequency components. This will induce a problem for suboptimal visual quality as the image denoising tasks target completely eliminating the complex noises and recovering all fine-scale and salient information. In this work, we tackle this challenge from the frequency perspective and present a new solution pipeline, coined as frequency attention denoising network (FADNet). Our key idea is to build a learning-based frequency attention framework, where the feature correlations on a broader frequency spectrum can be fully characterized, thus enhancing the representational power of the network across multiple frequency channels. Based on this, we design a cascade of adaptive instance residual modules (AIRMs). In each AIRM, we first transform the spatial-domain features into the frequency space. Then, a learning-based frequency attention framework is devised to explore the feature inter-dependencies converted in the frequency domain. Besides this, we introduce an adaptive layer by leveraging the guidance of the estimated noise map and intermediate features to meet the challenges of model generalization in the noise discrepancy. The effectiveness of our method is demonstrated on several real camera benchmark datasets, with superior denoising performance, generalization capability, and efficiency versus the state-of-the-art. Ruijun Ma 0001, Yaoxuan Zhang, Bob Zhang 0001, Leyuan Fang, Dong Huang 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Linear discriminant analysis with generalized kernel constraint for robust image classification
Shuyi Li 0003, Hengmin Zhang, Ruijun Ma 0001, Jianhang Zhou, Jie Wen 0001, Bob Zhang 0001 |
Pattern Recognit. | 3 |
| 2023 | Flexible and Generalized Real Photograph Denoising Exploiting Dual Meta AttentionabstractSupervised deep learning techniques have been widely explored in real photograph denoising and achieved noticeable performances. However, being subject to specific training data, most current image denoising algorithms can easily be restricted to certain noisy types and exhibit poor generalizability across testing sets. To address this issue, we propose a novel flexible and well-generalized approach, coined as dual meta attention network (DMANet). The DMANet is mainly composed of a cascade of the self-meta attention blocks (SMABs) and collaborative-meta attention blocks (CMABs). These two blocks have two forms of advantages. First, they simultaneously take both spatial and channel attention into account, allowing our model to better exploit more informative feature interdependencies. Second, the attention blocks are embedded with the meta-subnetwork, which is based on metalearning and supports dynamic weight generation. Such a scheme can provide a beneficial means for self and collaborative updating of the attention maps on-the-fly. Instead of directly stacking the SMABs and CMABs to form a deep network architecture, we further devise a three-stage learning framework, where different blocks are utilized for each feature extraction stage according to the individual characteristics of SMAB and CMAB. On five real datasets, we demonstrate the superiority of our approach against the state of the art. Unlike most existing image denoising algorithms, our DMANet not only possesses a good generalization capability but can also be flexibly used to cope with the unknown and complex real noises, making it highly competitive for practical applications. Ruijun Ma 0001, Shuyi Li 0003, Bob Zhang 0001, Leyuan Fang |
IEEE Trans. Cybern. | 1 |
| 2022 | Generative Adaptive Convolutions for Real-World Noisy Image DenoisingabstractRecently, deep learning techniques are soaring and have shown dramatic improvements in real-world noisy image denoising. However, the statistics of real noise generally vary with different camera sensors and in-camera signal processing pipelines. This will induce problems of most deep denoisers for the overfitting or degrading performance due to the noise discrepancy between the training and test sets. To remedy this issue, we propose a novel flexible and adaptive denoising network, coined as FADNet. Our FADNet is equipped with a plane dynamic filter module, which generates weight filters with flexibility that can adapt to the specific input and thereby impedes the FADNet from overfitting to the training data. Specifically, we exploit the advantage of the spatial and channel attention, and utilize this to devise a decoupling filter generation scheme. The generated filters are conditioned on the input and collaboratively applied to the decoded features for representation capability enhancement. We additionally introduce the Fourier transform and its inverse to guide the predicted weight filters to adapt to the noisy input with respect to the image contents. Experimental results demonstrate the superior denoising performances of the proposed FADNet versus the state-of-the-art. In contrast to the existing deep denoisers, our FADNet is not only flexible and efficient, but also exhibits a compelling generalization capability, enjoying tremendous potential for practical usage. Ruijun Ma 0001, Shuyi Li 0003, Bob Zhang 0001 |
AAAI | 1 |
| 2022 | Row-sparsity Binary Feature Learning for Open-set Palmprint RecognitionabstractBinary feature representation methods have received increasing attention due to their high efficiency and great robustness to illumination variation. However, most of them are hand-designed feature descriptors that generally require much prior knowledge in their design. This paper introduces a Row-sparsity Binary Feature Learning (Rs-BFL) method to adaptively learn and encode palmprint features for open-set palmprint recognition. Given the training palmprint images, RsBFL jointly learns a bank of linear projection functions that transform the informative texture features into discriminative binary codes. Afterwards, we calculate the block-wise histograms of each feature map and concatenate them as the final feature representation. Based on the pre-trained projection matrix, we mapped the palmprint texture features of the test samples into binary features for matching. For RsBFL, we enforce three criteria: 1) the quantization error between the projected real-valued features and the binary features is minimized, at the same time, the projection noise is minimized; 2) the latent label semantic information is utilized to minimize the distance of the within-class samples and simultaneously maximize the distance of the between-class samples; 3) the$l_{2,1}$norm is used to make the projection matrix to extract more discriminative features. Extensive experimental results on two publicly accessible palmprint datasets demonstrated the effectiveness and powerful learning capability of the proposed method. Shuyi Li 0003, Ruijun Ma 0001, Jianhang Zhou, Bob Zhang 0001 |
IJCB | 2 |
| 2022 | Towards Broad Learning Networks on Unmanned Mobile Robot for Semantic SegmentationabstractThis article investigates the real-time semantic segmentation in robot engineering applications based on the Broad Learning System (BLS), and a novel Multi-level Enhancement Layers Network (MELNet) based on BLS framework is proposed for real-time vision tasks in a complex street scene on the unmanned mobile robot. This network mainly solves two problems: (1) mitigating the contradiction between accuracy and speed while maintaining low model complexity, and (2) accurately describing objects based on their shape despite their different sizes. Firstly, the BLS architecture is expanded to the deep network with trainable parameters. This trainable network could adjust its weights in a complex environment, and mitigate the adverse impact of the environment on the complex tasks. Secondly, enhancement layers with the extended enhancement layers could extract both detailed information and semantic information. Moreover, an Upsampling Atrous Spatial Pyramid Pooling (UPASPP) is designed to fuse detail and semantic information to describe object features properly. Finally, in the case of the MNIST dataset and Cityscapes dataset, we get high accuracy with 8.01M parameters and quicker inference speed on a single GTX 1070 Ti card. At the same time, the unmanned mobile robot (BIT-NAZA) is employed to evaluate semantic performance in real-world situations. This reveals that MELNet could be run adequately on the embedded device and effectively operate in the real-robot system. Jiehao Li, Yingpeng Dai, Xiaohang Su, Ruijun Ma 0001 |
ICRA | 5 |
| 2022 | Learning Compact Multirepresentation Feature Descriptor for Finger-Vein RecognitionabstractDue to its high anti-counterfeiting and universality, the use of finger-vein pattern for identity authentication has recently attracted extensive attention in academia and industry. Despite recent advances in finger-vein recognition, most of the hand-crafted descriptors require strong prior knowledge, which may be ineffective in expressing its distinctiveness. In this paper, we present a novel compact multi-representation feature descriptor (CMrFD) with visual and semantic consistency, for finger-vein feature representation. Given the finger-vein images, we first form two-view representations to describe the informative vein features in local patches. Then, we jointly learn a feature transformation to map the two-view representations into discriminative binary codes. For the projection function, we linearly combine multi-view information and minimize the quantization error between the projected binary features and the original real-valued features. In terms of visual consistency, we minimize the Euclidean distance of each representation from the same class, at the same time, maximize the Euclidean distance from different classes in the projected space. Semantic consistency is used to ensure that similar images have compact multi-representation combined projection features. Lastly, we calculate the block-wise histograms as the final extracted features for finger-vein recognition. Experimental results on four widely used finger-vein databases demonstrate that the proposed method outperforms the state-of-the-art finger-vein recognition methods. Shuyi Li 0003, Ruijun Ma 0001, Lunke Fei, Bob Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | Meta PID Attention Network for Flexible and Efficient Real-World Noisy Image DenoisingabstractRecent deep convolutional neural networks for real-world noisy image denoising have shown a huge boost in performance by training a well-engineered network over external image pairs. However, most of these methods are generally trained with supervision. Once the testing data is no longer compatible with the training conditions, they can exhibit poor generalization and easily result in severe overfitting or degrading performances. To tackle this barrier, we propose a novel denoising algorithm, dubbed as Meta PID Attention Network (MPA-Net). Our MPA-Net is built based upon stacking Meta PID Attention Modules (MPAMs). In each MPAM, we utilize a second-order attention module (SAM) to exploit the channel-wise feature correlations with second-order statistics, which are then adaptively updated via a proportional-integral-derivative (PID) guided meta-learning framework. This learning framework exerts the unique property of the PID controller and meta-learning scheme to dynamically generate filter weights for beneficial update of the extracted features within a feedback control system. Moreover, the dynamic nature of the framework enables the generated weights to be flexibly tweaked according to the input at test time. Thus, MPAM not only achieves discriminative feature learning, but also facilitates a robust generalization ability on distinct noises for real images. Extensive experiments on ten datasets are conducted to inspect the effectiveness of the proposed MPA-Net quantitatively and qualitatively, which demonstrates both its superior denoising performance and promising generalization ability that goes beyond those of the state-of-the-art denoising methods. Ruijun Ma 0001, Shuyi Li 0003, Bob Zhang 0001, Haifeng Hu 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | Towards Fast and Robust Real Image Denoising With Attentive Neural Network and PID ControllerabstractWith the development of deep learning technologies, recent research on real-world noisy image denoising has achieved a considerable improvement in performance. However, a common limitation for existing approaches is the imbalanced trade-off between denoising accuracy and efficiency. To address this problem, we propose a robust and efficient denoiser, called a hierarchical-based PID-attention denoising network (HPDNet), to flexibly deal with the sophisticated noise. The core of our algorithm is the PID-attentive recurrent network (PAR-Net) whose framework mainly consists of the LSTM network and PID controller. PAR-Net inherits the advantages of both the attentive recurrent network and control action, which can encourage more discriminatory feature representations. This learning procedure is implemented within a feedback control system, allowing a faster and more robust means to enhance feature discriminability. Furthermore, by decomposing the noisy image and stacking the PAR-Nets, our PAR-Net can work on a progressively hierarchical framework, and hence obtain multi-scale features and manageable successive refinements. On several widely used datasets, the proposed HPDNet demonstrates high efficiency, while delivering a better perceptually appealing image quality over state-of-the-art image denoising methods. Ruijun Ma 0001, Shuyi Li 0003, Bob Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | PID Controller-Guided Attention Neural Network Learning for Fast and Effective Real Photographs DenoisingabstractReal photograph denoising is extremely challenging in low-level computer vision since the noise is sophisticated and cannot be fully modeled by explicit distributions. Although deep-learning techniques have been actively explored for this issue and achieved convincing results, most of the networks may cause vanishing or exploding gradients, and usually entail more time and memory to obtain a remarkable performance. This article overcomes these challenges and presents a novel network, namely, PID controller guide attention neural network (PAN-Net), taking advantage of both the proportional-integral-derivative (PID) controller and attention neural network for real photograph denoising. First, a PID-attention network (PID-AN) is built to learn and exploit discriminative image features. Meanwhile, we devise a dynamic learning scheme by linking the neural network and control action, which significantly improves the robustness and adaptability of PID-AN. Second, we explore both the residual structure and share-source skip connections to stack the PID-ANs. Such a framework provides a flexible way to feature residual learning, enabling us to facilitate the network training and boost the denoising performance. Extensive experiments show that our PAN-Net achieves superior denoising results against the state-of-the-art in terms of image quality and efficiency. Ruijun Ma 0001, Bob Zhang 0001, Yicong Zhou, Fangyuan Lei |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Pseudo 3D Auto-Correlation Network for Real Image DenoisingabstractThe extraction of auto-correlation in images has shown great potential in deep learning networks, such as the self-attention mechanism in the channel domain and the self-similarity mechanism in the spatial domain. However, the realization of the above mechanisms mostly requires complicated module stacking and a large number of convolution calculations, which inevitably increases model complexity and memory cost. Therefore, we propose a pseudo 3D auto-correlation network (P3AN) to explore a more efficient way of capturing contextual information in image de-noising. On the one hand, P3AN uses fast 1D convolution instead of dense connections to realize criss-cross interaction, which requires less computational resources. On the other hand, the operation does not change the feature size and makes it easy to expand. It means that only a simple adaptive fusion is needed to obtain contextual information that includes both the channel domain and the spatial domain. Our method built a pseudo 3D auto-correlation attention block through 1D convolutions and a lightweight 2D structure for more discriminative features. Extensive experiments have been conducted on three synthetic and four real noisy datasets. According to quantitative metrics and visual quality evaluation, the P3AN shows great superiority and surpasses state-of-the-art image denoising methods. Xiaowan Hu, Ruijun Ma 0001, Yuanhao Cai, Xiaole Zhao, Yulun Zhang 0001, Haoqian Wang |
CVPR | 2 |
| 2020 | Gaussian Pyramid of Conditional Generative Adversarial Network for Real-World Noisy Image Denoising
Ruijun Ma 0001, Bob Zhang 0001, Haifeng Hu 0001 |
Neural Process. Lett. | 1 |
| 2020 | Efficient and Fast Real-World Noisy Image Denoising by Combining Pyramid Neural Network and Two-Pathway Unscented Kalman FilterabstractRecently, image prior learning has emerged as an effective tool for image denoising, which exploits prior knowledge to obtain sparse coding models and utilize them to reconstruct the clean image from the noisy one. Albeit promising, these prior-learning based methods suffer from some limitations such as lack of adaptivity and failed attempts to improve performance and efficiency simultaneously. With the purpose of addressing these problems, in this paper, we propose a Pyramid Guided Filter Network (PGF-Net) integrated with pyramid-based neural network and Two-Pathway Unscented Kalman Filter (TP-UKF). The combination of pyramid network and TP-UKF is based on the consideration that the former enables our model to better exploit hierarchical and multi-scale features, while the latter can guide the network to produce an improved (a posteriori) estimation of the denoising results with fine-scale image details. Through synthesizing the respective advantages of pyramid network and TP-UKF, our proposed architecture, in stark contrast to prior learning methods, is able to decompose the image denoising task into a series of more manageable stages and adaptively eliminate the noise on real images in an efficient manner. We conduct extensive experiments and show that our PGF-Net achieves notable improvement on visual perceptual quality and higher computational efficiency compared to state-of-the-art methods. Ruijun Ma 0001, Haifeng Hu 0001, Songlong Xing |
IEEE Trans. Image Process. | 1 |
| 2019 | Photorealistic Face Completion with Semantic Parsing and Face Identity-Preserving FeaturesabstractTremendous progress on deep learning has shown exciting potential for a variety of face completion tasks. However, most learning-based methods are limited to handle general or structure specified face images (e.g., well-aligned faces). In this article, we propose a novel face completion algorithm, called Learning and Preserving Face Completion Network (LP-FCN), which simultaneously parses face images and extracts face identity-preserving (FIP) features. By tackling these two tasks in a mutually boosting way, the LP-FCN can guide an identity preserving inference and ensure pixel faithfulness of completed faces. In addition, we adopt a global discriminator and a local discriminator to distinguish real images from synthesized ones. By training with a combined identity preserving, semantic parsing and adversarial loss, the LP-FCN encourages the completion results to be semantically valid and visually consistent for more complicated image completion tasks. Experiments show that our approach obtains similar visual quality, but achieves better performance on unaligned faces completion and fine detailed synthesis against the state-of-the-art methods. Ruijun Ma 0001, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2018 | Perceptual Face Completion using a Local-Global Generative Adversarial NetworkabstractFace completion is one of the most challenging problems, as the reconstruction algorithm should render the missing pixels with semantically plausible contents. Recent methods have achieved promising advances in photorealistic human face synthesis. However, these approaches are limited to deal with general or structure specified faces. In this paper, we propose a Two-Pathway Perceptual Generative Adversarial Network (TPP-GAN) for face completion by perceiving semantic representations from both global structures and local details of a face. We combine a reconstruction network and a perceptual network containing two pathway adversarial networks (local and global) into our framework to efficiently ensure the transfer of the prominent facial features to the occluded parts, which encourages a visually high-quality image completion results. Experimental results well demonstrate that our proposed framework not only generates locally semantic and globally consistent fragments, but also outperforms existing methods on unaligned faces and synthesis of part components. Ruijun Ma 0001, Haifeng Hu 0001 |
ICPR | 1 |