VLDB 2026 Research / reviewers in the wild / expert
Rui Zhu 0006
dblp:72/1974-6
· DBLP profile ↗
39ranked-venue papers
8as first author
27since 2021 · last 2026
0000-0002-9944-0369ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 12 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dawid-Skene-model-based label-noise mitigation for federated learningabstractFederated learning (FL) enables collaborative model training without centralising raw data, but its performance is susceptible to label noise from clients. A common mitigation strategy involves using a clean, labelled public dataset at the server to assess client reliability. However, this approach is impractical due to the unrealistic assumption of availability of a clean, labelled public dataset. To address this issue, we propose FedDS, a novel approach that brings the Dawid-Skene model from statistical analysis to FL, which enables the estimation of the reliability of each client in FL without requiring any labelled data at the server. This approach effectively mitigates the adverse impact of heterogeneous label noise under a weaker and more practical assumption, offering a robust aggregation strategy for real-world FL scenarios with label noise. The code is available at https://github.com/Gia99999/FedDS . Jia Dong, Rui Zhu 0006, Xinyi Shang, Jing-Hao Xue |
Inf. Sci. | 2 |
| 2026 | Federated learning with noisy labels: A comprehensive and concise review of current methodologies and future directionsabstractFederated learning, a vital paradigm in modern machine learning, enables private and decentralised training of models that is crucial for learning from sensitive data. Noisy label learning, another vital paradigm in modern machine learning, addresses the training of models from the data with potentially incorrect labels. Their integration, namely federated learning with noisy labels (FLNL), is an emerging but challenging topic arising from the practice of machine learning, which, however, still lacks a review of its research progress. The aim of this paper is to fill in this gap. We first summarise four core challenges to FLNL: localised label noise, across-client heterogeneity of label noise, localised overfitting to label noise, and inadequate benchmarking. We then propose a taxonomy to categorise current FLNL studies into four types that address the four challenges correspondingly: sample-wise methods, client-wise methods, model-wise methods, and benchmark-wise studies. This work offers the first comprehensive and concise review dedicated to FLNL; moreover, we also provide future research directions for this rapidly evolving and practically significant field. Jia Dong, Rui Zhu 0006, Xinyi Shang, Jing-Hao Xue |
Neural Networks | 2 |
| 2026 | UC-PUAL: A universally consistent classifier of positive-unlabelled dataabstractPositive-unlabelled (PU) learning is a challenging task in pattern recognition, as there are only labelled-positive instances and unlabelled instances available for the training of a classifier. The task becomes even harder when the PU data show an underlying trifurcate pattern that positive instances roughly distribute on both sides of ground-truth negative instances. To address this issue, we propose a universally consistent PU classifier with asymmetric loss (UC-PUAL) on positive instances. We also propose two three-block algorithms for non-convex optimisation to enable UC-PUAL to obtain linear and kernel-induced non-linear decision boundaries, respectively. Theoretical and experimental results verify the superiority of UC-PUAL. The code for UC-PUAL is available at https://github.com/tkks22123/UC-PUAL. Rui Zhu 0006, Jing-Hao Xue |
Pattern Recognit. | 2 |
| 2026 | FAFN: Feature alignment and filtering network for fine-grained few-shot image classification
Jijie Wu, Qiyu Yin, Rui Zhu 0006 |
Pattern Recognit. | 3 |
| 2026 | Fine-Tuning via Linked Domains: A Closed-Form Dual Alignment Mechanism for Transferring Vision-Language ModelsabstractAdapters and prompt learning have become two de facto strategies to fine-tune pre-trained vision-language models, mitigating the high computational cost of fine-tuning an entire model for downstream tasks. They can align the prediction from the fine-tuned model with that from the pre-trained model. However, the existing methods of these strategies primarily focus on aligning within a single modality, and the exploration of bidirectional interactions between modalities remains limited. To address this issue, we propose a closed-form dual alignment mechanism (DAM) thatnot only ensures the consistency in predictions within a single modality but also achieves the alignment of features across different modalities. In DAM, all alignments are achieved by closed-form solutions to ridge regression, without inducing a massive number of learnable parameters. Experimental results demonstrate that DAM outperforms the state-of-the-art methods on 11 benchmarks over various evaluation metrics. Our codes are available at https://github.com/Peiy-Lu/DAM. Peiyu Lu, Rui Zhu 0006, Zhanyu Ma, Jie Cao 0014, Jing-Hao Xue |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Identity-Preserving Diffusion for Face RestorationabstractFace restoration is a critical task in computer vision, aiming to restore high-quality facial images from degraded inputs. In existing diffusion models, identity information is not well preserved when confronted with severely degradation. To address this challenge, we propose a Local Patch-Based Identity-Preserving Diffusion (LPIP-Diff) framework. Our local patch-based strategy leverages the interrelationships between neighboring patches to model highly structured facial context, which facilitates the restoration of fine-grained details and the preservation of identity-related features. We also introduce a fusion degradation estimation method that makes each overlapping area restored multiple times by adjacent patches, effectively restoring local details. The experimental results of LPIP-Diff on three publicly available datasets, including one severely degraded dataset, consistently demonstrate its superiority over the state-of-the-art methods in terms of both quantitative and qualitative evaluations, strikes a good balance between realism and fidelity, and enhances robustness against degradation. Xiaying Bai, Wenming Yang, Rui Zhu 0006, Jing-Hao Xue |
ICASSP | 4 |
| 2025 | Oracle-MoE: Locality-preserving Routing in the Oracle Space for Memory-constrained Large Language Model InferenceabstractMixture-of-Experts (MoE) is widely adopted to deploy Large Language Models (LLMs) on edge devices with limited memory budgets. Although MoE is, in theory, an inborn memory-friendly architecture requiring only a few activated experts to reside in the memory for inference, current MoE architectures cannot effectively fulfill this advantage and will yield intolerable inference latencies of LLMs on memory-constrained devices. Our investigation pinpoints the essential cause as the remarkable temporal inconsistencies of inter-token expert activations, which generate overly frequent expert swapping demands dominating the latencies. To this end, we propose a novel MoE architecture, Oracle-MoE, to fulfill the real on-device potential of MoE-based LLMs. Oracle-MoE route tokens in a highly compact space suggested by attention scores, termed the oracle space, to effectively maintain the semantic locality across consecutive tokens to reduce expert activation variations, eliminating massive swapping demands. Theoretical analysis proves that Oracle-MoE is bound to provide routing decisions with better semantic locality and, therefore, better expert activation consistencies. Experiments on the pretrained GPT-2 architectures of different sizes (200M, 350M, 790M, and 2B) and downstream tasks demonstrate that without compromising task performance, our Oracle-MoE has achieved state-of-the-art inference speeds across varying memory budgets, revealing its substantial potential for LLM deployments in industry. Jixian Zhou, Ruijun Huang, Hengjie Cao, Mengyi Chen, Anrui Chen, Mingzhi Dong, Yujiang Wang 0001, Dongsheng Li 0002, David A. Clifton, Qin Lv, Rui Zhu 0006, Fan Yang 0001, Tun Lu, Ning Gu 0001, Li Shang 0002 |
ICML | 13 |
| 2025 | PUAL: A classifier on trifurcate positive-unlabelled data
Rui Zhu 0006, Jing-Hao Xue |
Neurocomputing | 3 |
| 2025 | GKF-PUAL: A group kernel-free approach to positive-unlabeled learning with variable selectionabstractVariable selection is important for classification of data with many irrelevant predicting variables, but it has not yet been well studied in positive-unlabeled (PU) learning, where classifiers have to be trained without labelled-negative instances. In this paper, we propose a group kernel-free PU classifier with asymmetric loss (GKF-PUAL) to achieve quadratic PU classification with group-lasso regularisation embedded for variable selection. We also propose a five-block algorithm to solve the optimization problem of GKF-PUAL. Our experimental results reveal the superiority of GKF-PUAL in both PU classification and variable selection, improving the baseline PUAL by more than 10% in F1-score across four benchmark datasets and removing over 70% of irrelevant variables on six benchmark datasets. The code for GKF-PUAL is at https://github.com/tkks22123/GKF-PUAL . • We propose a group kernel-free PU classifier (GKF-PUAL) with variable selection. • We propose a five-block algorithm for optimization of GKF-PUAL. • Experimental results verify the superiority of GKF-PUAL. Rui Zhu 0006, Jing-Hao Xue |
Inf. Sci. | 2 |
| 2025 | Clarity in chaos: Boosting few-shot classification through information suppression and sparsificationabstractThe advance of deep learning has invigorated the research of few-shot classification. However, the interference of non-target information in feature representations hampers classification generalization. To tackle this issue, we propose an irrelevant information suppression (IIS) module, which is focused on suppressing the weight of unimportant information and elevating the sparsity of feature representations . An IIS network with three consecutive IIS modules is developed, to illustrate the progressive suppression of unimportant information and highlighting of key discriminative features of the target. Extensive experiments showcase the superior performance of our IIS network on five widely-used benchmark datasets. Furthermore, we show that the IIS module can be readily used as a plug-in module by state-of-the-art few-shot classifiers, and can clearly further improve their performance. Our code is available on GitHub at https://github.com/LC4188/IISNet . • We propose an IIS module to progressively suppress non-target information. • The IIS module can be readily used as a plug-in module. • The IIS module can clearly improve the performance of few-shot classifiers. Luchen Ji, Rui Zhu 0006, Zhanyu Ma, Jing-Hao Xue |
Pattern Recognit. | 3 |
| 2025 | SRML: Structure-relation mutual learning network for few-shot image classificationabstractFew-shot image classification aims at tackling a challenging but practical classification setting, where only few labelled images are available for training. Metric-based methods are main-stream solutions for few-shot image classification, but many of them extract features that are either irrelevant to target objects in the query images or insufficient to describe the local shape or structural patterns within images, which can lead to mis-identification of the target objects, especially when the images are of multiple objects. To resolve this issue, we propose the structure-relation mutual learning (SRML) network, which first learns both the intra-image structural features and the inter-image relational features in a parallel fashion via two parallel branches, the structural feature extractor (SFE) and the relational feature extractor (RFE), and then harnesses mutual learning to enable knowledge exchange between them. In such a manner, the structural features learnt from the SFE branch not only contain the structural patterns within the images, but also focus more on the target objects, guided by the relational knowledge from the RFE branch. In return, the RFE branch can exploit the more-focused structural knowledge to better match the target objects in the support and query images. We conduct extensive experiments on four few-shot classification benchmark datasets to showcase the superior classification of the proposed SRML network, achieving a 3.17% improvement in classification accuracy over the leading competitor, RENet Kang et al. (2021). The code of this work can be found in https://github.com/Rilliant7/SRML . Rui Zhu 0006, Zhanyu Ma, Jie Cao 0014, Jing-Hao Xue |
Pattern Recognit. | 3 |
| 2025 | Rise by Lifting Others: Interacting Features to Uplift Few-Shot Fine-Grained ClassificationabstractFew-shot fine-grained classification entails notorious subtle inter-class variation. Recent works address this challenge by developing attention mechanisms, such as the task discrepancy maximization (TDM) that can highlight discriminative channels. This paper, however, aims to reveal that, besides designing sophisticated attention modules, a well-designed input scheme, which simply blends two types of features and their interactions capturing different properties of the target object, can also greatly promote the quality of the learnt weights. To illustrate, we design a bi-feature interactive TDM (BiFI-TDM) module to serve as a strong foundation for TDM to discover the most discriminative channels with ease. Specifically, we design a novel mixing strategy to produce four sets of channel weights with different focuses, reflecting the properties of the corresponding input features and their interactions, as well as a proper feature re-weighting scheme. Extensive experiments on four benchmark fine-grained image datasets showcase superior performance of BiFI-TDM in metric-based few-shot methods. Our codes are available athttps://github.com/Peiy-Lu/BiFI-TDM. Peiyu Lu, Rui Zhu 0006, Zhanyu Ma, Jie Cao 0014, Jing-Hao Xue |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Selectively Augmented Attention Network for Few-Shot Image ClassificationabstractFew-shot image classification is a challenging task that aims to learn from a limited number of labelled training images a classification model that can be generalised to unseen classes. Two strategies are usually taken to improve the classification performances of few-shot image classifiers: either applying data augmentation to enlarge the sample size of the training set and reduce overfitting, or involving attention mechanisms to highlight discriminative spatial regions or channels. However, naively applying them to few-shot classifiers directly and separately may lead to undesirable results; for example, some augmented images may focus majorly on the background rather than the object, which brings additional noises to the training process. In this paper, we propose a unified framework, the selectively augmented attention (SAA) network, that carefully integrates the best of the two approaches in an end-to-end fashion via a selective best match module to select the most representative images from the augmented training set. The selected images tend to concentrate on the objects with less irrelevant background, which can assist the subsequent calculation of attentions by alleviating the interference from background. Moreover, we design a joint attention module to jointly learn both the spatial and channel-wise attentions. Experimental results on four benchmark datasets showcase the superior classification performance of the proposed SAA network compared with the state-of-the-arts. Rui Zhu 0006, Zhanyu Ma, Jie Cao 0014, Jing-Hao Xue |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Denoising Reuse: Exploiting Inter-Frame Motion Consistency for Efficient Video GenerationabstractDenoising-based diffusion models have attained impressive image synthesis; however, their applications on videos can lead to unaffordable computational costs due to the per-frame denoising operations. In pursuit of efficient video generation, we present a Diffusion Reuse MOtion (Dr. Mo) network to accelerate the video-based denoising process. Our crucial observation is that the latent representations in early denoising steps between adjacent video frames exhibit high consistencies with motion clues. Inspired by the discovery, we propose to accelerate the video denoising process by incorporating lightweight, learnable motion features. Specifically, Dr. Mo will only compute all denoising steps for base frames. For a non-based frame, Dr. Mo will propagate the pre-computed based latents of a particular step with inter-frame motions to obtain a fast estimation of its coarse-grained latent representation, from which the denoising will continue to obtain more sensitive and fine-grained representations. On top of this, Dr. Mo employs a meta-network named Denoising Step Selector (DSS) to dynamically determine the step to perform motion-based propagations for each frame, ensuring the correct transformation of multi-granularity visual features. Extensive evaluations on video generation and editing tasks indicate that Dr. Mo delivers widely applicable acceleration for diffusion-based video generations while effectively retaining the visual quality and style. Video generation and visualization results can be found athttps://drmo-denoising-reuse.github.io. Yixuan Chen 0003, Yujiang Wang 0001, Mingzhi Dong, Dongsheng Li 0002, Rui Zhu 0006, David A. Clifton, Robert P. Dick, Qin Lv, Fan Yang 0001, Tun Lu, Ning Gu 0001, Li Shang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2025 | Query-Aware Cross-Mixup and Cross-Reconstruction for Few-Shot Fine-Grained Image ClassificationabstractFew-shot fine-grained image classification is prominent but challenging in computer vision, aiming to distinguish sub-classes under the same parent class but with only a few labeled support samples. Data augmentation techniques were explored to address the few-shot issue, but they often fail to mitigate the bias between support and query samples. Therefore, in this paper we propose a query-aware cross-mixup and cross-reconstruction method to address both few-shot and fine-grained issues. Specifically, in the training phase, we randomly select query samples and mix them with the support samples from the same class to augment the support set. This first strategy ensures the augmented support set query-aware within each sub-class. Then, we reconstruct both query samples and support samples from both original and cross-mixed support samples, thus leveraging both cross-reconstruction and self-reconstruction to enhance classification. This second strategy, enabling the reconstruction also query-aware, further mitigates the bias between support and query samples, leading to more reliable generalization. We evaluate our proposed method on four widely used few-shot fine-grained image classification datasets, and experimental results demonstrate its effectiveness in achieving the state-of-the-art classification performance. Dongliang Chang, Rui Zhu 0006, Zhanyu Ma, Jing-Hao Xue |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Improving Diffusion-Based Image Restoration with Error Contraction and Error CorrectionabstractGenerative diffusion prior captured from the off-the-shelf denoising diffusion generative model has recently attained significant interest. However, several attempts have been made to adopt diffusion models to noisy inverse problems either fail to achieve satisfactory results or require a few thousand iterations to achieve high-quality reconstructions. In this work, we propose a diffusion-based image restoration with error contraction and error correction (DiffECC) method. Two strategies are introduced to contract the restoration error in the posterior sampling process. First, we combine existing CNN-based approaches with diffusion models to ensure data consistency from the beginning. Second, to amplify the error contraction effects of the noise, a restart sampling algorithm is designed. In the error correction strategy, the estimation-correction idea is proposed on both the data term and the prior term. Solving them iteratively within the diffusion sampling framework leads to superior image generation results. Experimental results for image restoration tasks such as super-resolution (SR), Gaussian deblurring, and motion deblurring demonstrate that our approach can reconstruct high-quality images compared with state-of-the-art sampling-based diffusion models. Qiqi Bao 0001, Zheng Hui, Rui Zhu 0006, Peiran Ren, Xuansong Xie, Wenming Yang |
AAAI | 3 |
| 2024 | Once Read is Enough: Domain-specific Pretraining-free Language Models with Cluster-guided Sparse Experts for Long-tail Domain KnowledgeabstractLanguage models (LMs) only pretrained on a general and massive corpus usually cannot attain satisfying performance on domain-specific downstream tasks, and hence, applying domain-specific pretraining to LMs is a common and indispensable practice.
However, domain-specific pretraining can be costly and time-consuming, hindering LMs' deployment in real-world applications.
In this work, we consider the incapability to memorize domain-specific knowledge embedded in the general corpus with rare occurrences and long-tail distributions as the leading cause for pretrained LMs' inferior downstream performance.
Analysis of Neural Tangent Kernels (NTKs) reveals that those long-tail data are commonly overlooked in the model's gradient updates and, consequently, are not effectively memorized, leading to poor domain-specific downstream performance.
Based on the intuition that data with similar semantic meaning are closer in the embedding space, we devise a Cluster-guided Sparse Expert (CSE) layer to actively learn long-tail domain knowledge typically neglected in previous pretrained LMs.
During pretraining, a CSE layer efficiently clusters domain knowledge together and assigns long-tail knowledge to designate extra experts. CSE is also a lightweight structure that only needs to be incorporated in several deep layers.
With our training strategy, we found that during pretraining, data of long-tail knowledge gradually formulate isolated, outlier clusters in an LM's representation spaces, especially in deeper layers. Our experimental results show that only pretraining CSE-based LMs is enough to achieve superior performance than regularly pretrained-finetuned LMs on various downstream tasks, implying the prospects of domain-specific-pretraining-free language models. Mengyi Chen, Jixian Zhou, Yubin Shi, Yixuan Chen 0003, Mingzhi Dong, Yujiang Wang 0001, Dongsheng Li 0002, Rui Zhu 0006, Robert P. Dick, Qin Lv, Fan Yang 0001, Tun Lu, Ning Gu 0001, Li Shang 0002 |
NeurIPS | 10 |
| 2024 | A simple scheme to amplify inter-class discrepancy for improving few-shot fine-grained image classificationabstractFew-shot image classification is a challenging topic in pattern recognition and computer vision. Few-shot fine-grained image classification is even more challenging, due to not only the few shots of labelled samples but also the subtle differences to distinguish subcategories in fine-grained images. A recent method called task discrepancy maximisation (TDM) can be embedded into the feature map reconstruction network (FRN) to generate discriminative features, by preserving the appearance details through reconstructing the query image and then assigning higher weights to more discriminative channels, producing the state-of-the-art performance for few-shot fine-grained image classification. However, due to the small inter-class discrepancy in fine-grained images and the small training set in few-shot learning, the training of FRN+TDM can result in excessively flexible boundaries between subcategories and hence overfitting. To resolve this problem, we propose a simple scheme to amplify inter-class discrepancy and thus improve FRN+TDM. To achieve this aim, instead of developing new modules, our scheme only involves two simple amendments to FRN+TDM: relaxing the inter-class score in TDM, and adding a centre loss to FRN. Extensive experiments on five benchmark datasets showcase that, although embarrassingly simple, our scheme is quite effective to improve the performance of few-shot fine-grained image classification. The code is available at https://github.com/Airgods/AFRN.git. Zijie Guo, Rui Zhu 0006, Zhanyu Ma, Jun Guo 0002, Jing-Hao Xue |
Pattern Recognit. | 3 |
| 2024 | DSR-Diff: Depth map super-resolution with diffusion model
Huiyun Cao, Bin Xia 0014, Rui Zhu 0006, Qingmin Liao, Wenming Yang |
Pattern Recognit. Lett. | 4 |
| 2023 | ReNAP: Relation network with adaptiveprototypical learning for few-shot classification
Yalan Li, Yixiao Zheng, Rui Zhu 0006, Zhanyu Ma, Jing-Hao Xue, Jie Cao 0014 |
Neurocomputing | 4 |
| 2023 | Statistical hypothesis testing as a novel perspective of pooling for image quality assessmentabstractImage quality assessment is usually achieved by pooling local quality scores. However, commonly used pooling strategies, based on simple sample statistics, are not always sensitive to distortions. In this short communication, we propose a novel perspective of pooling: reliable pooling through statistical hypothesis testing, which enables effective detection of subtle changes of population parameters when the underlying distribution of local quality scores is affected by distortions. To illustrate the significance of this novel perspective, we design a new pooling strategy utilising simple one-sided one-sample t-test. The experiments on benchmark databases show the reliability of hypothesis testing-based pooling, compared with state-of-the-art pooling strategies. Rui Zhu 0006, Fei Zhou 0001, Wenming Yang, Jing-Hao Xue |
Signal Process. Image Commun. | 1 |
| 2023 | Locally-Enriched Cross-Reconstruction for Few-Shot Fine-Grained Image ClassificationabstractFew-shot fine-grained image classification has attracted considerable attention in recent years for its realistic setting to imitate how humans conduct recognition tasks. Metric-based few-shot classifiers have achieved high accuracies. However, their metric function usually requires two arguments of vectors, while transforming or reshaping three-dimensional feature maps to vectors can result in loss of spatial information. Image reconstruction is thus involved to retain more appearance details: the test images are reconstructed by different classes and then classified to the one with the smallest reconstruction error. However, discriminative local information, vital to distinguish sub-categories in fine-grained images with high similarities, is not well elaborated when only the base features from a usual embedding module are adopted for reconstruction. Hence, we propose the novel local content-enriched cross-reconstruction network (LCCRN) for few-shot fine-grained classification. In LCCRN, we design two new modules: the local content-enriched module (LCEM) to learn the discriminative local features, and the cross-reconstruction module (CRM) to fully engage the local features with the appearance details obtained from a separate embedding module. The classification score is calculated based on the weighted sum of reconstruction errors of the cross-reconstruction tasks, with weights learnt from the training process. Extensive experiments on four fine-grained datasets showcase the superior classification performance of LCCRN compared with the state-of-the-art few-shot classification methods. Codes are available at:https://github.com/lutsong/LCCRN. Jijie Wu, Rui Zhu 0006, Zhanyu Ma, Jing-Hao Xue |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Sain: Similarity-Aware Video Frame InterpolationabstractVideo frame interpolation (VFI) aims to synthesize an intermediate frame between two consecutive original frames. Most existing methods simply linearly combine the warped frames, leading to a loss of image texture. Since moving objects usually have similarities in consecutive frames, we propose a similarity-aware video frame interpolation method (SAIN) that searches patches with similar texture in the embedding space from input frames to extract features and capture image details. To gather the frame details and restore image texture, SAIN incorporates an implicit neural representation learning from similar patches to enrich image details and refine outputs in frame synthesis networks. Experiments demonstrate that SAIN preserves image texture and enhances interpolated image quality significantly. Yue Lv, Wenming Yang, Wangmeng Zuo, Qingmin Liao, Rui Zhu 0006 |
ICASSP | 5 |
| 2022 | Quality-Oriented Feature Regression for Robust Image Similarity MetricabstractFull-reference image quality assessment aims to predict the perceptual quality of a distorted image based on its similarity to the pristine reference. In this paper, we propose a robust image similarity metric by fully exploring the representation power of deep learning-based features. A convolutional neu-ral network (CNN) is adopted to extract deep features from multiple scales. We show that such CNN features that con-tain multi -scale visual information are comprehensive and ro-bust enough for quality assessment. We further propose a quality-oriented feature regression (QOFR) module based on the multi-layer perceptron architecture. The QOFR module can efficiently integrate hierarchy CNN features and generate the final quality score. Extensive experiments on the bench-mark datasets demonstrate that our method achieves state-of-the-art performance with outstanding robustness and general-ization ability. Qiqi Bao 0001, Rui Zhu 0006, Wenming Yang, Qingmin Liao |
ICME | 3 |
| 2022 | Distilling Resolution-robust Identity Knowledge for Texture-Enhanced Face HallucinationabstractThe main focus of most existing face hallucination methods is to generate visually pleasing results. However, in many applications, the final goal is to identify the person in the low-resolution (LR) image. In this paper, we propose a texture and identity integration network (TIIN) to effectively incorporate identity information into face hallucination tasks. TIIN consists of an identity-preserving denormalization module (IDM) and an equalized texture enhance module (ETEM). The IDM exploits the identity prior and the ETEM improves image quality through histogram equalization. To extract identity information effectively, we propose a resolution-robust identity knowledge distillation network (RIKDN). RIKDN is specifically designed for LR face recognition and can be of independent interest. It employs two teacher-student streams. One stream narrows the performance gap between high-resolution (HR) and LR images. The other distills correlation information from the HR-HR teacher stream to guide learning in the LR-HR student stream. We conduct extensive experiments on multiple datasets to demonstrate the effectiveness of our methods. Qiqi Bao 0001, Rui Zhu 0006, Bowen Gang, Pengyang Zhao, Wenming Yang, Qingmin Liao |
ACM Multimedia | 2 |
| 2022 | Constrained mutual convex cone method for image set based recognition
Naoya Sogi, Rui Zhu 0006, Jing-Hao Xue, Kazuhiro Fukui |
Pattern Recognit. | 2 |
| 2021 | Generalisations of stochastic supervision models
Xiaoou Lu, Yangqi Qiao, Rui Zhu 0006, Guijin Wang, Zhanyu Ma, Jing-Hao Xue |
Pattern Recognit. | 3 |
| 2020 | Generalization Bound of Gradient Descent for Non-Convex Metric LearningabstractMetric learning aims to learn a distance measure that can benefit distance-based methods such as the nearest neighbour (NN) classifier. While considerable efforts have been made to improve its empirical performance and analyze its generalization ability by focusing on the data structure and model complexity, an unresolved question is how choices of algorithmic parameters, such as the number of training iterations, affect metric learning as it is typically formulated as an optimization problem and nowadays more often as a non-convex problem. In this paper, we theoretically address this question and prove the agnostic Probably Approximately Correct (PAC) learnability for metric learning algorithms with non-convex objective functions optimized via gradient descent (GD); in particular, our theoretical guarantee takes the iteration number into account. We first show that the generalization PAC bound is a sufficient condition for agnostic PAC learnability and this bound can be obtained by ensuring the uniform convergence on a densely concentrated subset of the parameter space. We then show that, for classifiers optimized via GD, their generalizability can be guaranteed if the classifier and loss function are both Lipschitz smooth, and further improved by using fewer iterations. To illustrate and exploit the theoretical findings, we finally propose a novel metric learning method called Smooth Metric and representative Instance LEarning (SMILE), designed to satisfy the Lipschitz smoothness property and learned via GD with an early stopping mechanism for better discriminability and less computational cost of NN. Mingzhi Dong, Rui Zhu 0006, Yujiang Wang 0001, Jing-Hao Xue |
NeurIPS | 3 |
| 2020 | Deep learning for image super-resolution
Wenming Yang, Fei Zhou 0001, Rui Zhu 0006, Kazuhiro Fukui, Guijin Wang, Jing-Hao Xue |
Neurocomputing | 3 |
| 2020 | Adjusting the imbalance ratio by the dimensionality of imbalanced data
Rui Zhu 0006, Yiwen Guo, Jing-Hao Xue |
Pattern Recognit. Lett. | 1 |
| 2020 | Special Issue on Advances in Statistical Methods-based Visual Quality Assessment
Fei Zhou 0001, Wenming Yang, Xinbo Gao 0001, Hantao Liu, Rui Zhu 0006, Jing-Hao Xue |
Signal Process. Image Commun. | 5 |
| 2020 | A Novel Separating Hyperplane Classification Framework to Unify Nearest-Class-Model Methods for High-Dimensional DataabstractIn this article, we establish a novel separating hyperplane classification (SHC) framework to unify three nearest-class-model methods for high-dimensional data: the nearest subspace method (NSM), the nearest convex hull method (NCHM), and the nearest convex cone method (NCCM). Nearest-class-model methods are an important paradigm for the classification of high-dimensional data. We first introduce the three nearest-class-model methods and then conduct dual analysis for theoretically investigating them, to understand deeply their underlying classification mechanisms. A new theorem for the dual analysis of NCCM is proposed in this article by discovering the relationship between a convex cone and its polar cone. We then establish the new SHC framework to unify the nearest-class-model methods based on the theoretical results. One important application of this new SHC framework is to help explain empirical classification results: why one class model has a better performance than others on certain data sets. Finally, we propose a new nearest-class-model method, the soft NCCM, under the novel SHC framework to solve the overlapping class model problem. For illustrative purposes, we empirically demonstrate the significance of our SHC framework and the soft NCCM through two types of typical real-world high-dimensional data: the spectroscopic data and the face image data. Rui Zhu 0006, Ziyu Wang 0003, Naoya Sogi, Kazuhiro Fukui, Jing-Hao Xue |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Learning distance to subspace for the nearest subspace methods in high-dimensional data classification
Rui Zhu 0006, Mingzhi Dong, Jing-Hao Xue |
Inf. Sci. | 1 |
| 2018 | MvSSIM: A quality assessment index for hyperspectral images
Rui Zhu 0006, Fei Zhou 0001, Jing-Hao Xue |
Neurocomputing | 1 |
| 2018 | LRID: A new metric of multi-class imbalance degree based on likelihood-ratio testabstractIn this paper, we introduce a new likelihood ratio imbalance degree (LRID) to measure the class-imbalance extent of multi-class data. Imbalance ratio (IR) is usually used to measure class-imbalance extent in imbalanced learning problems. However, IR cannot capture the detailed information in the class distribution of multi-class data, because it only utilises the information of the largest majority class and the smallest minority class. Imbalance degree (ID) has been proposed to solve the problem of IR for multi-class data. However, we note that improper use of distance metric in ID can have harmful effect on the results. In addition, ID assumes that data with more minority classes are more imbalanced than data with less minority classes, which is not always true in practice. Thus ID cannot provide reliable measurement when the assumption is violated. In this paper, we propose a new metric based on the likelihood-ratio test, LRID, to provide a more reliable measurement of class-imbalance extent for multi-class data. Experiments on both simulated and real data show that LRID is competitive with IR and ID, and can reduce the negative correlation with F1 scores by up to 0.55. Rui Zhu 0006, Ziyu Wang 0003, Zhanyu Ma, Guijin Wang, Jing-Hao Xue |
Pattern Recognit. Lett. | 1 |
| 2018 | Cone-based joint sparse modelling for hyperspectral image classification
Ziyu Wang 0003, Rui Zhu 0006, Kazuhiro Fukui, Jing-Hao Xue |
Signal Process. | 2 |
| 2017 | Building a discriminatively ordered subspace on the generating matrix to classify high-dimensional spectral data
Rui Zhu 0006, Kazuhiro Fukui, Jing-Hao Xue |
Inf. Sci. | 1 |
| 2017 | On the orthogonal distance to class subspaces for high-dimensional data classification
Rui Zhu 0006, Jing-Hao Xue |
Inf. Sci. | 1 |
| 2017 | Matched Shrunken Cone Detector (MSCD): Bayesian Derivations and Case Studies for Hyperspectral Target DetectionabstractHyperspectral images (HSIs) possess non-negative properties for both hyperspectral signatures and abundance coefficients, which can be naturally modeled using cone-based representation. However, in hyperspectral target detection, cone-based methods are barely studied. In this paper, we propose a new regularized cone-based representation approach to hyperspectral target detection, as well as its two working models by incorporating into the cone representation l2-norm and l1-norm regularizations, respectively. We call the new approach the matched shrunken cone detector (MSCD). Also important, we provide principled derivations of the proposed MSCD from the Bayesian perspective: we show that MSCD can be derived by assuming a multivariate half-Gaussian distribution or a multivariate half-Laplace distribution as the prior distribution of the coefficients of the models. In the experimental studies, we compare the proposed MSCD with the subspace methods and the sparse representation-based methods for HSI target detection. Two real hyperspectral data sets are used for evaluating the detection performances on sub-pixel targets and full-pixel targets, respectively. Results show that the proposed MSCD can outperform other methods in both cases, demonstrating the competitiveness of the regularized cone-based representation. Ziyu Wang 0003, Rui Zhu 0006, Kazuhiro Fukui, Jing-Hao Xue |
IEEE Trans. Image Process. | 2 |